Dynamic selection of candidate bitrates for video encoding
The dynamic selection of candidate average bitrates based on video characteristics addresses the issue of inconsistent quality and resource inefficiency in existing methods, resulting in optimized video encoding that maintains quality and reduces resource usage.
Patent Information
- Application Number
- JP2025084633
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-06
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-11-24
AI Technical Summary
Existing video encoding methods using static lists of candidate average bitrates fail to optimize video quality and resource usage due to varying video characteristics, leading to suboptimal transcoding and inconsistent quality gaps between profiles.
A pre-analysis optimization process dynamically selects candidate average bitrates based on video characteristics, using rate-distortion curves and machine learning to generate an optimized list for each video portion, ensuring consistent quality and efficient resource use.
This approach enhances video quality by optimizing segment selection for profiles, minimizing storage and distribution footprint while maintaining consistent quality levels, thereby improving the viewing experience.
Smart Images

Figure 2025128149000001_ABST
Abstract
Description
[Background technology]
[0001] One method of delivering video to client devices uses adaptive bitrate streaming (ABR). Adaptive bitrate streaming is based on providing multiple streams (often called variants or profiles) encoded with different levels of video attributes, such as different levels of bitrate and / or quality. A profile ladder lists different profiles available for a client to use when streaming segments of video. The client can dynamically select a profile based on network conditions and other factors. The video is segmented (e.g., divided into individual segments, typically several seconds in length), and the client can switch from one profile to another at segment boundaries when network conditions change. For example, when a network condition with higher available bandwidth is experienced, a video delivery system may want to provide the client with a profile with a higher bitrate that improves the quality of the video being streamed. When a network condition with lower available bandwidth is experienced, the video delivery system may want to provide the client with a profile with a lower bitrate so that the client can play the video without any playback issues, such as rebuffering or download failures. Summary of the Invention
[0002] The included drawings are for illustrative purposes and merely serve to provide examples of possible structures and operations of the disclosed inventive systems, apparatus, methods, and computer program products. These drawings in no way limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed implementations. [Brief explanation of the drawings]
[0003] [Figure 1]
[0003] FIG. 1 illustrates a system for dynamically selecting a list of candidate average bit rates according to some embodiments. [Figure 2]
[0004] FIG. 1 illustrates an example of a portion of a video according to some embodiments. [Figure 3A]
[0005] FIG. 10 illustrates an example of generating a list of candidate average bit rates according to some embodiments. [Figure 3B]
[0006] FIG. 10 illustrates an example of generating encoded segments according to some embodiments. [Figure 4]
[0007] FIG. 10 illustrates an example of clustering encoded segments into multiple pools according to some embodiments. [Figure 5]
[0008] FIG. 10 illustrates an example of a selection process according to some embodiments. [Figure 6]
[0009] FIG. 10 illustrates an example of a graph of rate-distortion curves that may be used to select coded segments for pooling according to some embodiments. [Figure 7]
[0010] FIG. 10 illustrates an example of selected encoded segments for a per-segment profile according to some embodiments. [Figure 8]
[0011] 1A and 1B are diagrams illustrating examples of different rate-distortion curves for video content according to some embodiments. [Figure 9]
[0012] 10A-10C illustrate different characteristics using different encoding configurations according to some embodiments. [Figure 10]
[0013] FIG. 10 illustrates an example of using static candidate average bit rates for different rate-distortion curves according to some embodiments. [Figure 11]
[0014] FIG. 10 illustrates an optimized candidate average bitrate list according to some embodiments. [Figure 12]
[0015] FIG. 1 illustrates a more detailed example of a segment quality-driven adaptive (SQA) system and pre-analysis optimization process according to some embodiments. [Figure 13]
[0016] FIG. 1 illustrates a more detailed example of a rate-distortion (RD) prediction system according to some embodiments. [Figure 14]
[0017] FIG. 10 illustrates the output of a predictive network according to some embodiments. [Figure 15]
[0018] FIG. 1 illustrates a simplified flowchart of a method for performing an optimization process for selecting a list of candidate average bit rates according to some embodiments. [Figure 16]
[0019] 4A and 4B illustrate an example of determining boundaries for a list of candidate average bit rates according to some embodiments. [Figure 17]
[0020] 10 illustrates an example of eliminating candidate average bit rates based on quality according to some embodiments. [Figure 18]
[0021] FIG. 10 illustrates an example in which a minimum gap is used to eliminate candidate average bit rates according to some embodiments. [Figure 19]
[0022] 6 is a graph illustrating when adding candidate average bit rates may be advantageous according to some embodiments. [Figure 20]
[0023] 10A and 10B illustrate the decision to add candidate average bit rates according to some embodiments. [Figure 21]
[0024] FIG. 1 illustrates a video streaming system that communicates with multiple client devices over one or more communication networks according to one embodiment. [Figure 22]
[0025] FIG. 1 shows a schematic diagram of a device for viewing video content and advertisements. DETAILED DESCRIPTION OF THE INVENTION
[0004]
[0026]
[0013] Techniques for video distribution systems are described herein. In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of some embodiments. Some embodiments, as defined by the claims, may include some or all of the features in these examples, alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0005]
[0027] The system can adaptively generate a list of bitrates to be used to encode a video. The list of bitrates may be referred to as candidate average bitrates (CABs). The encoder transcodes segments of the video using each bitrate in the list of candidate average bitrates. In some embodiments, the system can dynamically select a list of candidate average bitrates for different portions of the video, such as for different chunks of the video. A chunk may be an independent coding unit that the encoder encodes using the same settings. A video may include one or more chunks, and each chunk may include multiple segments. In some embodiments, the list of candidate average bitrates may be set at the chunk level. Although the list of candidate average bitrates is described as being set at the chunk level, the list of candidate average bitrates may be set for different portions of the video.
[0006]
[0028] The encoder can encode segments of the video using bit rates in a list of candidate average bit rates to generate multiple candidate segments. A segment quality-driven adaptation (SQA) process can select a segment from the candidate segments to use for a profile in a profile ladder. The goal of the process is to optimize (e.g., minimize) the storage or distribution footprint of a portion of the video while maintaining similar quality.
[0007]
[0029] Each video may have different characteristics. Similarly, different portions of the same video may have different characteristics. Using a static list of candidate average bit rates for all portions of a video or for multiple videos may not provide optimal results. For example, a static list of candidate average bit rates may encode a video with simple video content with a bit rate higher than required. Also, a video with complex video content may be encoded at low quality due to insufficient bit rate. Furthermore, a static list of candidate average bit rates may generate segments with irregular quality gaps from an encoding perspective. For example, adjacent profiles may have similar video quality that is redundant with each other or may have an unacceptably large quality gap. Having similar video quality for adjacent profiles may be unnecessary and may not provide much benefit in viewing quality. For example, if two bit rates in the list of candidate average bit rates result in an encoded segment with similar quality, transcoding the segment using those two bit rates may be redundant and waste resources. Also, having a large quality gap may result in a poor viewing experience during playback, as the quality may change abruptly when playback switches from one profile to another.
[0008]
[0030] To overcome the above drawbacks, the pre-analysis optimization process can dynamically select bit rates within a list of candidate average bit rates for a video. To select the list of candidate average bit rates, the pre-analysis optimization process can analyze portions of the video and output an optimized list of candidate average bit rates for the portions. For example, the pre-analysis optimization process can analyze characteristics of each portion and output a list of candidate average bit rates for each portion. In some embodiments, the pre-analysis optimization process can predict characteristics of the portions, such as a rate-distortion curve that describes quality versus bit rate for the portion. The pre-analysis optimization process uses the respective rate-distortion curve to determine the optimal list of bit rates for each portion.
[0009]
[0031] The optimization process provides many advantages. For example, it provides an optimal selection of transcoded segments to select when selecting segments for a profile in the profile ladder. If the list of candidate average bit rates is set at a static value for the entire video and / or is the same for multiple different videos, suboptimal transcoding may occur. Different videos, and even different portions of the same video, may have diverse characteristics. Thus, a static list of candidate average bit rates may be suboptimal for some videos or portions of videos. The use of a dynamic list of candidate average bit rates based on the characteristics of portions of videos may result in higher quality video and viewing experiences, since the segment quality-driven adaptation process may have a better selection of encoded segments to select to form profiles for the profile ladder.
[0010]
[0032] system
[0033] FIG. 1 illustrates a system 100 for dynamically selecting a list of candidate average bit rates, according to some embodiments. The system 100 includes a content delivery network 102, a client 104, and a video delivery system 106. The source files may contain different types of content, such as video, audio, or other types of content information. Video may be used for illustrative purposes, but other types of content may be understood. In some embodiments, the source files may be received in a format that requires encoding into another format, as described below. For example, the source file may be a mezzanine file containing compressed video. The mezzanine file may be encoded to generate other files, such as different profiles of the video.
[0011]
[0034] A content provider may operate the video delivery system 106 to provide a content distribution service that allows entities to request and receive media content. The content provider may use the video delivery system 106 to coordinate the distribution of media content to the clients 104. While a single client 104 is described, multiple clients 104 may be using the service. The media content may be different types of content, such as on-demand video from a library of videos and live video. In some embodiments, live video may be where video is available based on a linear schedule. Video may also be provided on-demand. On-demand video may be content that can be requested at any time and is not limited to viewing on a linear schedule. The video may be a program, such as a movie, a show, or an advertisement.
[0012]
[0035] The clients 104 may include different computing devices such as smartphones, living room devices, televisions, set-top boxes, tablet devices, etc. The clients 104 include a media player 112 capable of playing content such as videos. In some embodiments, the media player 112 may receive segments of videos and play these segments. The clients 104 may send a request for a segment to one of the content delivery networks 102 and then receive the requested segment for playback on the media player 112. A segment may be a portion of a video, such as 6 seconds of a video.
[0013]
[0036] A video may be encoded into a profile ladder including multiple profiles. Each profile may correspond to a different configuration, which may be a different level of bitrate and / or quality, but may also include other characteristics such as codec type, computing resource type (e.g., computer processing unit), etc. Each video may have an associated profile with a different configuration. The profiles may be categorized into different levels, and each level may be associated with a different configuration. For example, a level may be a combination of bitrate, resolution, codec, etc., and each level may be associated with a different bitrate, such as 400 kilobytes per second (kbps), 650 kbps, 1000 kbps, 1500 kbps, ...12000 kbps. Each level may also be associated with another characteristic, such as a quality characteristic (e.g., resolution). Profile levels may be referred to as higher or lower, such that a profile with a higher bitrate or quality may be rated higher than a profile with a lower bitrate or quality. An encoder may use the characteristics to encode the source video. For example, an encoder may encode a source video at a target bitrate of 1500 kbps.
[0014]
[0037] The content delivery network 102 includes a server that can deliver video to the client 104. The content delivery network 102 receives requests for segments of video from the client 104 and delivers the segments of video to the client 104. The client 104 may request a segment of video from one of the profile levels based on the current playback state. The playback state may be any state experienced based on the playback of the video, such as available bandwidth, buffer length, etc. For example, the client 104 may use an adaptive bitrate algorithm to select a profile for the video based on the current available bandwidth, buffer length, or other playback state. The client 104 may continuously evaluate the current playback state and switch between profiles during playback of a segment of video. For example, during playback, the media player 112 may request a different profile for the video asset. For example, if low-bandwidth playback conditions are being experienced, the media player 112 may request a lower profile associated with a lower bitrate for the upcoming segment of the video. However, if playback conditions of higher available bandwidth are being experienced, the media player 112 may request a higher level profile associated with higher bandwidth for the upcoming segment of video.
[0015]
[0038] A segment quality-driven adaptive processing system (SQA system) 108 can encode segments using a list of candidate average bit rates. The SQA system 108 then selects segments for each profile using an optimization process. For example, the SQA system 108 can adaptively select segments with optimal bit rates for each profile in a profile ladder while maintaining similar quality levels. The SQA system 108 enables the system to maintain quality similar to or matching the target bit rate while minimizing the number of bits required to store or distribute the content.
[0016]
[0039] The pre-analysis optimization process 110 can dynamically generate a list of candidate average bit rates for the portions of the video. In some embodiments, the pre-analysis optimization process 110 can predict characteristics of each of the portions of the video, such as a rate-distortion curve. The pre-analysis optimization process 110 then selects candidate average bit rates for the portions of the video based on analyzing the characteristics of each of the portions of the video.
[0017]
[0040] The following first describes the segment quality-driven adaptive processing process, and then describes the dynamic selection of the list of candidate average bit rates in more detail.
[0018]
[0041] Segment Quality Driven Adaptation Process
[0042] As described above, the optimization process 110 can dynamically select a list of candidate average bit rates for portions of a video. The portions of a video may be different sizes. FIG. 2 shows an example of a portion of a video according to some embodiments. A video 200 may be divided into different portions at the segment level and the chunk level. In some embodiments, at 202, a video 200 may be divided into chunk-level portions. For example, multiple chunks, chunk_0, chunk_1, ..., chunk_m, may be included in the video 200. Each respective chunk may be divided into smaller portions, which may be called segments. For example, at 204, chunk_0 is divided into segments, segment_0, segment_1, ..., segment_n. Similarly, although not shown, chunk_1 may be divided into its own respective segments, segment_0, segment_1, ..., segment_n. A segment may be shorter in length than a chunk. For example, a chunk may be a two-minute video, and a segment may be a five-second video.
[0019]
[0043] In a segment quality-driven adaptation process, the SQA system 108 can process each segment of the video 200 to generate multiple encodings of each respective segment based on a list of candidate average bit rates. For illustrative purposes, the optimization process 110 selects a list of candidate average bit rates for each chunk, but lists of candidate average bit rates may be selected for different portion sizes, such as per segment, for multiple chunks, etc. The bit rates included in each respective list of candidate average bit rates can be optimized based on characteristics associated with each portion (e.g., chunk and / or segment) of the video using the list of candidate average bit rates. Given different characteristics for different chunks, the respective lists of candidate average bit rates may be different. However, it may be possible for the bit rates for multiple chunks in each respective list of candidate average bit rates to be the same.
[0020]
[0044] FIG. 3A illustrates an example of generating a list of candidate average bit rates according to some embodiments. Chunks chunk_0, chunk_1, chunk_2, ..., chunk_n are shown at 202. At 302, optimization system 110 has a list of candidate average bit rates selected for each chunk based on characteristics of each respective chunk. For example, for chunk_0, list #0 of candidate average bit rates is based on characteristics for chunk_0. Also, list #1 of candidate average bit rates is based on characteristics for chunk_1, and so on. In some examples, for chunk_0, list #0 of candidate average bit rates may include bit rates of 8500, 7750, 7000, 6250, 5500, 4750, 4000, and 3250 kilobytes per second (Kbps). For chunk_1, list #1 of candidate average bit rates may include bit rates of 7000, 6250, 4750, 4000, 3250, 2000, and 1250 Kbps.
[0021]
[0045] The list of candidate average bit rates may include bit rates to be used by the encoder to encode each segment. Traditionally, the candidate average bit rates may have statically included the same bit rates. Sometimes, two types of bit rates are used for all chunks. The first type may be a target average bit rate, and the second type may be an intermediate average bit rate. The target average bit rate may be a base bit rate associated with a profile in a profile ladder for adaptive bit rate encoding. The intermediate average bit rate may be supplemental to the target average bit rate. For example, additional bit rates between the target average bit rates may be added. The use of the intermediate average bit rate may provide additional bit rates for encoding additional encoding segments that may have different characteristics, such as quality, from the segments encoded from the target average bit rate. In some cases, the optimization process 110 may include bit rates from the target average bit rate and / or the intermediate average bit rate in the list of candidate average bit rates. For example, the optimization process 110 may include the target average bit rate in the list of candidate average bit rates, but may dynamically select other bit rates. In another example, the optimization process may dynamically select a bitrate in the list of candidate average bitrates based solely on the characteristics of the chunk.
[0022]
[0046] As described above, the encoder generates encoded segments for the chunks. Figure 3B shows an example of generating encoded segments according to some embodiments. At 204, the segments for chunks in chunk_0 are denoted segment_0, segment_1, segment_2, ..., segment_n. At 304, a list of candidate average bitrates (CAB list) for chunk_0 is used. In some embodiments, the same list of candidate average bitrates for chunk_0 is used for all segments of the chunk. However, multiple different lists of candidate average bitrates may be used for different segments of the chunk. The encoder then encodes the segments of chunk_0 using the list of candidate average bitrates.
[0023]
[0047] At 306, the encoded segments for each segment are listed. For each segment, the encoder encodes the segment using an average bit rate in the list of candidate average bit rates. The encoder can target each average bit rate when encoding the segment. This results in a set of encoded segments for each segment of the chunk, such as ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, ..., ENC_S0_CAB_n for segment_0. In the notation, ENC_S0 represents the encoded segment for segment_0, and CAB_0, CAB_1, CAB_2, etc. represent candidate average bit rates. For example, CAB_0 may be 8500 Kbps, CAB_1 may be 7750 Kbps, and CAB_2 may be 7000 Kbps. Each encoded segment may be encoded at the same quality level, such as 1080p. The process may be repeated for another quality level using the list of candidate average bit rates.
[0024]
[0048] For each segment, the optimization process 110 clusters the encoded segments into multiple pools. Each pool may correspond to one profile. FIG. 4 illustrates an example of clustering encoded segments into multiple pools according to some embodiments. At 402, multiple encoded segments are shown for candidate average bit rates. Each segment may have an associated value for a quality indicator. For example, encoded segment ENC_S0_CAB_0 may have a quality of quality_S0_c0, encoded segment ENC_S0_CAB_1 may have a quality of quality_S0_c1, etc. In the notation, quality_S0 represents the encoded segment of segment_0, and c0, c1, c2, etc. represent the quality for this encoded segment.
[0025]
[0049] Different methods may be used to include encoded segments in pools 401-1, 404-2, and 404-p. For example, each pool may have or be associated with a profile. Each profile may be associated with a target bitrate, which may be the maximum bitrate that can be used to encode segments for the associated profile. The SQA system 108 may include encoded segments starting with the highest average bitrate that can be used for the associated profile for the pool. The SQA system 108 may then add other encoded segments at other bitrates less than the maximum bitrate. This may result in different encoded segments being included in each pool. For example, pool S0_Pool_0 may include segments ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, etc. Pool S0_Pool_1 may also include encoded segments ENC_S0_CAB_2, ENC_S0_CAB_3, ENC_S0_CAB_4, etc. Thus, pool S0_Pool_1 may contain coded segments that start at a bit rate lower than the maximum bit rate in pool S0_Pool_0. If the coded segments are coded at bit rates of 8500, 7750, 7000, 6250, 5500, 4750, 4000, 3250 Kbps, pool S0_pool_0 may start with coded segments with average bit rates of 8500, 7750, 7000, etc., and pool S0_pool_1 may start with coded segments with average bit rates of 7000, 6250, 5500, etc. In some examples, exemplary bit rates for pools may be pool_0: 8500, 7700, 7000, 6250, 5500, 4750, pool_1: 7000, 6250, 5500, 4750, 4000, and pool_p: 5500, 4750, 4000, 3250.
[0026]
[0050] From each pool, the SQA system 108 may select one encoded segment based on using a selection process. FIG. 5 shows an example of a selection process according to some embodiments. The following process may be performed for each pool. At 404-1, pool S0_pool_0 from FIG. 4 is shown with its respective encoded segments. The SQA system 108 may use one or more rules to select an encoded segment for each pool. At 502, the SQA system 108 selects an encoded segment ENC_S0_CAB_1 for pool S0_pool_0. In some embodiments, the SQA system 108 may attempt to select an encoded segment with the minimum bitrate that has a quality value that meets a criterion. In some examples, the SQA system 108 may start with the first encoded segment in the pool, such as the segment with the highest bitrate. Then, the SQA system 108 selects a neighboring encoded segment in the pool, such as the encoded segment with the next highest bitrate. If the first and second encoded segments have similar quality (e.g., within a threshold), the SQA system 108 selects the encoded segment with the lowest bitrate. The SQA system 108 may continue the comparison using adjacent encoded segments in the pool, such as the second and third encoded segments. If the adjacent encoded segments do not have similar quality, the process may terminate. Other methods may be used, such as starting with the encoded segment with the lowest bitrate. The process may also select the segment with the lowest bitrate that has a quality within the threshold of another segment, such as the segment with the highest bitrate. The following describes an example of a process using a rate-distortion curve.
[0027]
[0051] 6 shows an example of a rate-distortion curve graph 600 that may be used to select coding segments for pooling according to some embodiments. In graph 600, the Y-axis is quality and the X-axis is bitrate. Curve 602 defines the relationship between quality and bitrate. For example, the curve may plot the rate and distortion of a segment or chunk, although the curve may also plot other characteristics of quality and bitrate.
[0028]
[0052] The coded segments may be listed as A, B, C, D, E, and F on the curve 602 based on their respective rates and distortions. At 604, an example of coded segments having similar quality is shown. In this case, coded segment C and coded segment D have similar bit rates and similar quality. For example, the quality difference between coded segment C and coded segment D may meet a threshold value min_gap (e.g., equal and / or smaller). In this case, the SQA system 108 may select coded segment D because the quality difference is minimal, since this coded segment has a lower bit rate compared to coded segment C, but segment D provides similar quality compared to segment C.
[0029]
[0053] The SQA system 108 can also collapse coded segments whose quality exceeds an upper boundary. For example, the upper boundary at 606 may be the boundary used to determine coded segments as candidates for collapse. In this case, the SQA system 108 may select one or more of the segments above the upper threshold, such as selecting only one segment (e.g., segment B), or selecting fewer segments found to exceed the upper threshold (e.g., selecting two of four segments). In another example, coded segments A and B may be removed. The SQA system 108 may also remove coded segments whose quality falls below a lower boundary. For example, a lower threshold is shown at 608. The SQA system 108 may select one or more of the segments below the lower threshold, such as selecting only one segment (e.g., segment F), or selecting fewer segments found to fall below the lower threshold. In another example, coded segments E and F may be removed. The upper and lower thresholds may be used to restrict segments for a profile that exceed or fall below a desired bit rate or quality. One reason for using an upper limit is to restrict the bit rate used to encode a segment, and one reason for using a lower limit is to restrict a bit rate from being too low. After processing the encoded segments to remove unencoded segments, the SQA system 108 may select a segment for the profile. For example, the SQA system 108 may select the encoded segment with the lowest bit rate that has a quality level that meets the threshold, such as within the gap with the highest-quality segment. In this case, the SQA system 108 may select encoded segment D.
[0030]
[0054] While the above rules may be used to select segments, other processes may also be used. For example, the selection of an encoded segment may be based on which encoded segments have been selected for other profiles. In some examples, the selected segment may be based on reducing the storage of encoded segments where a profile may reuse segments from other profiles. Thus, the SQA system 108 may optimize quality while minimizing the bitrate used for encoded segments that fall between the lower bound and the ceiling.
[0031]
[0055] FIG. 7 shows examples of selected encoded segments for each segment's profile, according to some embodiments. At 702, 704, 706, and 708, encoded segments are shown for Profile_0, Profile_1, Profile_2, and Profile_P, respectively. Within a profile, the SQA system 108 can select different encoded segments having different candidate average bit rates for different segments. For example, for Profile_0, Segment_0 was encoded using candidate average bit rate CAB_1, Segment_1 was encoded using candidate average bit rate CAB_0, Segment_2 was encoded using candidate average bit rate CAB_0, etc. In some examples, in Profile_0, Segment_0 was encoded using a bit rate of 7750 Kbps, Segment_1 was encoded using a bit rate of 8500 Kbps, and Segment_2 was encoded using a bit rate of 8500 Kbps. For profile_1, segment_0 was coded using CAB_4, segment_1 was coded using CAB_2, and segment_2 was coded using CAB_3. For example, for profile 1, segment_0 was coded using a bitrate of 5500 Kbps, segment_1 was coded using a bitrate of 7000 Kbps, and segment_2 was coded using a bitrate of 6250 Kbps.
[0032]
[0056] The following then describes an optimization process for dynamically generating a list of candidate average bit rates.
[0033]
[0057] Optimization Process
[0058] As mentioned above, video content may have diverse characteristics, such that content in different videos may have different characteristics, and content within the same video may also have different characteristics. For example, some content, such as cartoons or news, may be easy to encode. However, some content, such as live-action movies or sports, may be difficult to encode. The encoding characteristics may vary. The following describes different characteristics for content.
[0034]
[0059] 8 shows an example of different rate-distortion curves for video content according to some embodiments. The rate-distortion curve is used to indicate the relationship between quality and bitrate, although other metrics may be used to indicate the relationship between quality and bitrate for video content. Different rate-distortion curves may be shown for different chunks of video, although the rate-distortion curves may be different for different portions of the video, such as segments, chunks, multiple chunks, or different videos.
[0035]
[0060] Three chunks, chunk_A, chunk_B, and chunk_C, are shown along with graphs 802, 804, and 806 of rate-distortion curves for the chunks, respectively. In graph 802, quality varies steeply at lower bit rates, but does not vary much at higher bit rates. In graph 804, quality varies as the bit rate increases with a constant correlation. In graph 806, quality at lower bit rates may vary only minimally, while quality increases steeply at higher bit rates.
[0036]
[0061] In addition to different content generating different rate-distortion curves, different encoding configurations may also generate different encoding results. Different encoding configurations may include using different encoders (e.g., x264, x265, etc.) or different encoding parameters (e.g., rate-distortion optimization (RDO) level, B-frames, number of references, etc.). Figure 9 illustrates different characteristics using different encoding configurations according to some embodiments. For the same segment or chunk, a first encoding configuration in 902 results in different characteristics compared to a second encoding configuration shown in 904. Encoding configuration A results in a rate-distortion curve similar to chunk_A above, and encoding configuration B results in a rate-distortion curve similar to chunk_B above, even though these rate-distortion curves are for the same content.
[0037]
[0062] Considering that the rate-distortion curves may be different, using a static list of candidate average bit rates may not be optimal. For example, using the same list of candidate average bit rates for different rate-distortion curves may not provide optimal results. FIG. 10 shows an example of using static candidate average bit rates for different rate-distortion curves according to some embodiments. Graphs 802, 804, and 806 show different rate-distortion curves for the different chunks shown in FIG. 8. The dotted lines in each graph indicate different bit rates in the list of candidate average bit rates. Some problems may arise when using a fixed list of candidate average bit rates. For example, in graph 802, the two highest candidate average bit rates at 1008 may be redundant because they have similar quality to the third candidate average bit rate at 1010. That is, to provide an encoded segment with similar quality, only one bit rate, such as the bit rates listed at 1010, may need to be encoded.
[0038]
[0063] In graph 804, at 1012, two candidate average bit rates may be redundant because these two coded segments have similar quality compared to the coded segment with the next lowest bit rate shown at 1014. As above, only one bit rate may need to be coded, such as the lowest bit rate at 1014, to give coded segments with similar quality.
[0039]
[0064] In graph 806, the three lowest candidate average bit rates may produce encoded segments with similar quality at 1016. Also, the candidate average bit rates may be too far apart, so that the difference in quality between the encoded segments may be too large at 1018. That is, to minimize the difference in quality between the candidate average bit rates, it may be more desirable to have more candidate average bit rates with smaller quality differences.
[0040]
[0065] 11 illustrates an optimized candidate average bitrate list according to some embodiments. In graph 802, the SQA system 108 can dynamically select candidate average bitrates to optimize the quality observed in the encoded segment. For example, in 1102, the SQA system 108 can increase the number of candidate average bitrates at bitrates where the curve is steep. Also, in 1103, the SQA system 108 can decrease the number of candidate average bitrates where the curve does not change the quality much.
[0041]
[0066] In graph 804, at 1104, the SQA system 108 may remove candidate average bitrates from the lowest bitrate where quality may be redundant, and at 1106, the SQA system 108 may add additional bitrates to obtain varying quality at higher bitrates.
[0042]
[0067] In graph 806, at 1108, the SQA system 108 may eliminate bit rates at the lower end of the curve. Also, at 1110, the SQA system 108 may space the candidate average bit rates more evenly to capture different quality levels in more even increments.
[0043]
[0068] Pre-analysis optimization process design
[0069] 12 shows a more detailed example of the SQA system 108 and pre-analysis optimization process 110 according to some embodiments. A chunk to be encoded is received. An encoding configuration may also be received that defines settings for encoding the chunk. The encoding configuration may include an encoder type, a quality level, etc.
[0044]
[0070] The pre-analysis optimization process 110 can receive chunks and encoding configurations and output an optimized list of candidate average bit rates. The RD prediction system 1202 can predict rate-distortion curves for segments and / or chunks within the chunks. While predicting a rate-distortion curve for a segment or chunk may be described, rate-distortion curves may be generated for different portions of the video, such as for multiple chunks and / or multiple segments. As described in more detail below, the RD prediction system 1202 can use machine learning logic to generate a prediction of the rate-distortion curve for a segment.
[0045]
[0071] The predicted rate-distortion curve is output to the CAB list optimization system 1204. The CAB list optimization system 1204 can optimize a list of candidate average bit rates for the chunk based on, for example, the predicted rate-distortion curves for the segments within the chunk. The optimized list of candidate average bit rates may be based on the characteristics of each chunk and may be different for chunks with content having different characteristics. This process is described in more detail below.
[0046]
[0072] The CAB list optimization system 1204 outputs the optimized list of candidate average bit rates to the SQA system 108. The SQA system 108 includes an encoding system 1206 that receives the optimized list of encoding configurations, chunks, and candidate average bit rates. The encoding system 1206 then uses each candidate average bit rate in the list to encode each segment of the chunk. After encoding each segment using the list of candidate average bit rates, the selection system 1208 selects an encoded segment for each profile in the profile ladder using the selection process described above. The selection system 1208 outputs the selected encoded segments for the profiles in the profile ladder.
[0047]
[0073] The following describes the prediction of segment characteristics and then optimization to select a list of candidate average bit rates.
[0048]
[0074] RD prediction system
[0075] 13 shows a more detailed example of the RD prediction system 1202 according to some embodiments. A feature extraction system 1302 receives chunks of video. The feature extraction system 1302 may then extract values for features that may convey information related to video transcoding. Some examples of features may be about the video content, encoding settings, etc. The extracted features may provide a better prediction of the characteristics of the segments of the chunks. The values for the features are output to a prediction network 1304.
[0049]
[0076] The prediction network 1304 can use the trained model to generate characteristics for segments of chunks, such as predicted rate-distortion curves. The prediction network 1304 may use different machine learning algorithms, such as support vector machine (SVM) regression, convolutional neural networks (CNN), boosting, etc. The trained model can be trained based on a particular machine learning algorithm.
[0050]
[0077] The prediction network 1304 can receive values for the features in addition to other inputs such as a segment location, an encoding configuration, and a target bit rate. The segment location may be the segment location (e.g., that segment within the video) for which to generate a rate-distortion curve, the encoding configuration may include the configuration used to encode the segment, and the target bit rate may include an output bit rate range for the segment. The prediction network 1304 can output a rate-distortion curve for the segment between the output bit rate ranges based on the features.
[0051]
[0078] FIG. 14 illustrates the output of the prediction network 1304 according to some embodiments. At 204, the segments for the chunk include segment_0, segment_1, segment_2, ..., segment_n. A rate-distortion curve may be generated for each segment within each chunk of video. For example, at 1402, a rate-distortion curve is output for each segment. The rate-distortion curve for segment_0, the rate-distortion curve for segment_1, etc. are shown. Each rate-distortion curve is based on characteristics for the respective segment. A list of candidate average bitrates for the chunk may be generated based on the rate-distortion curves. Chunk-level rate-distortion curves may also be output.
[0052]
[0079] List of candidate average bitrate optimizations
[0080] 15 shows a simplified flowchart 1500 of a method for performing an optimization process for selecting a list of candidate average bit rates according to some embodiments. At 1502, the CAB list optimization system 1204 determines boundaries for the list of candidate average bit rates. For example, the boundaries may be maximum and minimum bit rates that may be used for the list of candidate average bit rates. Different methods may be used to determine the boundaries, and are described in more detail in FIG. 16.
[0053]
[0081] At 1504, the CAB list optimization system 1204 generates a list of potential candidate average bitrates with optimal bitrate allocations. In some embodiments, one list of potential candidate average bitrates is generated for the chunk based on the maximum and minimum bitrates determined in 502. The list of potential candidate average bitrates may be generated using different methods. One method may be to use a predetermined list that falls between the minimum and maximum bitrates. For example, the predetermined list may include bitrates from the target average bitrate and the median average bitrate. For example, a bitrate from a predetermined list within a minimum and maximum range may be used. Another method may determine the total number of potential candidate average bitrates and divide the bitrate range between the minimum and maximum bitrates into intervals. Different examples may be used, such as:
[0054]
number
[0055]
number
[0056]
number
[0057]
number
[0058] where interval_i is the interval value of i, interval_(i+1) is the interval value + 1, interval_(i+2) is the interval value + 2, and delta is a predetermined value.
[0059]
[0082] The total number of intervals may be set to a number, such as 10. The intervals for interval_i may be set based on the above method by dividing the range into total numbers. The CAB list optimization system 1204 then selects bit rates based on the interval values to divide the range of bit rates between the minimum and maximum bit rates into a list of bit rates. For example, a minimum bit rate of 2000 and a maximum bit rate of 10,000 with an interval of 1500 and a total number of bit rates of 5 may result in a list of bit rates of 10,000, 7500, 5000, 3500, and 2000 when using equal division.
[0060]
[0083] At 1506, the CAB list optimization system 1204 refines the list of potential candidate average bit rates using an optimal quality allocation to generate an optimized list of candidate average bit rates. The quality allocation may examine the quality for each segment and determine whether the quality satisfies one or more rules. For example, redundant candidate average bit rates, such as candidate average bit rates with similar quality, may be eliminated. Additionally, additional candidate average bit rates may be added as needed, such as when neighboring candidate average bit rates have a quality gap that exceeds a threshold, such as a too large difference. The process is described in more detail in Figures 17, 18, and 19.
[0061]
[0084] As illustrated in 1502 of FIG. 15, the CAB list optimization system 1204 determines boundaries for a list of candidate average bit rates. FIG. 16 shows an example of determining boundaries for a list of candidate average bit rates according to some embodiments. The following process is described, but other processes may be understood. For example, settings may be used to determine minimum and maximum bit rates. In this example, the CAB list optimization system 1204 may analyze the minimum and maximum bit rates for the rate-distortion curve for each segment in the chunk and determine what the minimum and maximum bit rates should be at the chunk level.
[0062]
[0085] At 1602, the rate-distortion curves for each segment are received and analyzed. The CAB list optimization system 1204 may then select a minimum bitrate and a maximum bitrate for each segment based on the respective rate-distortion curve for the segment. For example, for segment_0, the minimum and maximum bitrates may be selected based on characteristics of the rate-distortion curve for segment_0. For example, the CAB list optimization system 1204 may set a maximum quality threshold and a minimum quality threshold and use the rate-distortion curve to determine a minimum bitrate corresponding to the minimum quality threshold and a maximum bitrate corresponding to the maximum quality threshold. For segment_1, the CAB list optimization system 1204 selects a minimum and maximum bitrate based on characteristics of the rate-distortion curve for segment_1, and so on.
[0063]
[0086] The above analysis was performed at the segment level. The CAB list optimization system 1204 then analyzes the segment-level results to determine minimum and maximum values at the chunk level. At 1606, the CAB list optimization system 1204 determines a maximum value from among the values for the maximum bitrate for the segment, from max_bitrate_0, max_bitrate_1, max_bitrate_2, ..., max_bitrate_n, etc. The CAB list optimization system 1204 also determines a minimum value from among the values for the minimum bitrate for the segment, from among min_bitrate_0, min_bitrate_1, min_bitrate_2, ..., min_bitrate_n, etc.
[0064]
[0087] At 1608, the CAB list optimization system 1204 outputs minimum and maximum bit rates for the chunk. In this case, the lowest minimum bit rate is selected from the minimum bit rates for the segment, and the highest maximum bit rate is selected from the maximum bit rates for the segment. The selection process may take into account the individual characteristics of the rate-distortion curve for the segment and select minimum and maximum bit rates that may include all of the minimum and maximum bit rates determined at the segment level. For example, if the minimum bit rates are 2000, 3000, and 3500, the minimum bit rate selected will be 2000. Similarly, if the maximum bit rates are 10000, 9000, and 8500, the maximum bit rate selected will be 10000. While the above process may be used, other methods of selecting minimum and maximum bit rates may be understood, such as taking an average of the values.
[0065]
[0088] As illustrated in 1506 of FIG. 15, the CAB list optimization system 1204 defines a list of candidate average bit rates with optimal quality allocations. Part of the allocation includes eliminating candidate average bit rates based on similar quality. Quality similarity may be defined in different ways. For example, the CAB list optimization system 1204 determines the distance between quality values to determine whether some candidate average bit rates should be eliminated. FIG. 17 shows an example of eliminating candidate average bit rates based on quality according to some embodiments. In this example, the CAB list optimization system 1204 may determine whether the quality levels of two adjacent candidate average bit rates meet a threshold, such as being within the threshold. The candidate with the higher bit rate may then be eliminated.
[0066]
[0089] At 1702, each segment may have an associated potential removal list of encoded segments that may potentially be removed. As shown, for segment_0, the CAB list optimization system 1204 has determined that the candidate average bit rates of S0_CAB_0, S0_CAB_3, and S0_CAB_4 may be removed. These candidate average bit rates may be removed because the encoded segments may have similar quality levels that meet the threshold with respect to adjacent encoded segments. Similarly, for segment_1, the CAB list optimization system 1204 has determined that the candidate average bit rates of S1_CAB_0 and S1_CAB_2 may be removed, and for segment_n, the CAB list optimization system 1204 has determined that the candidate average bit rates for Sn_CAB_0 and Sn_CAB_3 may be removed. The segment is not removed for segment_2 because the segments are not determined to have similar quality levels within the threshold.
[0067]
[0090] The above analysis was at the segment level. Then, at 1704, the CAB list optimization system 1204 can use the segment-level candidate average bit rates to determine candidate average bit rates to be removed at the chunk level. For example, the CAB list optimization system 1204 may select a candidate average bit rate for the chunk level based on the occurrence of the candidate average bit rate in different segments of the potential removal list. In some embodiments, the CAB list optimization system 1204 can select a candidate average bit rate and calculate the total number of occurrences in the potential removed candidate pool. If the total number of this candidate average bit rate meets a threshold, such as at or above the threshold, the CAB list optimization system 1204 places this candidate average bit rate in the removal candidate list at the chunk level. For example, the candidate average bit rate CAB_0 is found in three of the above-mentioned segments (e.g., segment_0, segment_1, and segment_n), meeting the threshold of "3." The CAB list optimization system 1204 then places the candidate average bit rate of CAB_0 in the removal candidate list. Candidate average bit rates CAB_2, CAB_3, and CAB_4 may not meet the threshold because the bit rates occur in two or fewer segments in the potential removal list. Therefore, the CAB list optimization system 1204 does not include these candidate average bit rates in the removal candidate list. Other methods of selecting which candidate average bit rates to remove may be understood.
[0068]
[0091] The above analysis was performed at the segment level and merged to the chunk level. However, the process may be performed at a different level. For example, the analysis may be used to merge candidate average bitrates from multiple chunks to a portion of the video covering multiple chunk levels, or from multiple chunks to the video level.
[0069]
[0092] The following describes an example of removing candidate average bit rates. Figure 18 shows an example of using a minimum gap to remove candidate average bit rates according to some embodiments. For example, in 1802, candidate average bit rates C and D have similar quality levels that satisfy a threshold, such as threshold min_gap. In this case, this candidate average bit rate is adjacent to candidate average bit rate D, and candidate average bit rate C has a larger bit rate than candidate average bit rate D but has a minimum quality advantage, so the CAB list optimization system 1204 determines that one of the candidate average bit rates, such as candidate average bit rate C, should be removed. In this case, the CAB list optimization system 1204 may compare the difference between the quality values for the adjacent candidate average bit rates with threshold min_gap and remove one of the candidate average bit rates when the threshold is met.
[0070]
[0093] Another part of quality allocation involves adding candidate average bit rates based on gaps in quality. Figure 19 shows a graph 1900 illustrating when it may be advantageous to add candidate average bit rates according to some embodiments. The CAB list optimization system 1204 may use a threshold, such as the maximum gap max_gap in 1902, to determine when to add a candidate average bit rate. For example, if there is a gap in quality values between adjacent candidate average bit rates that is greater than the threshold max_gap, such as between candidate average bit rates C and D in graph 1900, the CAB list optimization system 1204 may add a candidate average bit rate between candidate average bit rates C and D on the rate-distortion curve.
[0071]
[0094] Different methods may be used to determine how many new candidate average bit rates should be added. The CAB list optimization system 1204 may add "i" new candidates when two candidates are separated by a ratio-based threshold. For example, different ratios may constitute the gap between the added candidates, such as 1:1, meaning each gap is equal, 1:1.5, meaning each gap is 1.5, or other ratios.
[0072]
[0095] In one possible process, the variable i is set to i=1, and the CAB list optimization system 1204 adds i new candidate average bit rates based on the ratio. For example, one candidate average bit rate, labeled "F," may be added between points C and D. Then, if all gaps between the new adjacent candidate average bit rates are smaller than the threshold max_gap, the process ends. However, if this is not the case, the value of the variable i is incremented, such as by "2," and two new candidates are added between the candidate average bit rates based on the ratio. For example, two or more candidate average bit rates may be added between points C and F and between points F and D. Then, the process continues as described above. Once the candidate average bit rates have been added such that there is no gap between the candidate average bit rates D and C that is larger than the threshold, the candidate average bit rate is output.
[0073]
[0096] The above process is determined for each segment. The CAB list optimization system 1204 may then take the potential added candidate average bit rates at the segment level and merge the candidate average bit rates at the chunk level. Figure 20 illustrates the determination of adding candidate average bit rates according to some embodiments. In 2002, each segment may have candidate average bit rates that can potentially be added. For segment_0, two candidate average bit rates may be added midway between the candidate average bit rates CAB_0 and CAB_1. Also, one candidate average bit rate may be added between CAB_4 and CAB_5. For segment_1, two candidate average bit rates may be added between the candidate average bit rates CAB_0 and CAB_1. For segment_n, two candidate average bit rates may be added between the candidate average bit rates CAB_0 and CAB_1. Thus, three segments add two candidate average bit rates between candidate average bit rates CAB_0 and CAB_1, and one segment adds one candidate average bit rate between CAB_4 and CAB_5.
[0074]
[0097] The CAB list optimization system 1204 may use the segment-level candidates to determine an additional candidate list for the chunk. For example, to be added to the additional candidate list at the chunk level, the CAB list optimization system 1204 may determine whether a potential added candidate average bit rate at the segment level is found within a threshold, such as the number of segments. If the threshold is 70% of the segments, the CAB list optimization system 1204 adds two candidates between candidate average bit rates CAB_0 and CAB_1 because these candidates are found in more than 70% of the segments (three out of four segments). The CAB list optimization system 1204 does not add a candidate between CAB_4 and CAB_5 because this addition is found only in segment_0, which is less than 70% of the segments. In this case, adding an additional candidate average bit rate may not be needed because only one segment requires addition, and adding candidate average bit rates for all other segments of the chunk when only one segment is affected may not be useful. However, adding two candidate average bit rates between candidate average bit rates CAB_0 and CAB_1 may be beneficial since over 70% of the segments had potential additions.
[0075]
[0098] The output of the CAB list optimization system 1204 is a list of candidate average bit rates per chunk. For example, a list of candidate average bit rates per chunk, such as that depicted in FIG. 3A, is output by the CAB list optimization system 1204 according to some embodiments.
[0076]
[0099] conclusion
[0100] Therefore, the list of candidate average bit rates can be optimized based on the characteristics found in each segment. This produces an improved list of candidate average bit rates for each chunk that is optimized for the characteristics of each chunk. The candidate average bit rates can improve the quality of the selection of encoded segments available for each chunk profile selection. This can improve the quality of the video in addition to improving the playback experience.
[0077]
[0101] system
[0102] Features and aspects disclosed herein may be implemented in conjunction with a video streaming system 2100 that communicates with multiple client devices over one or more communications networks, as shown in Figure 21. Aspects of the video streaming system 2100 are described merely to provide one example of an application for enabling distribution and delivery of content prepared in accordance with the present disclosure. It should be understood that the present technology is not limited to streaming video applications and may be adapted to other applications and delivery mechanisms.
[0078]
[0103] In one embodiment, a media program provider may include a library of media programs. For example, the media programs may be aggregated and provided through a site (e.g., a website), an application, or a browser. A user may access the media program provider's site or application and request a media program. The user may be limited to requesting only media programs provided by the media program provider.
[0079]
[0104] In system 2100, video data may be obtained from one or more sources, e.g., video source 2110, for use as input to video content server 2102. The input video data may comprise raw or edited frame-based video data in any suitable digital format, e.g., Moving Picture Experts Group (MPEG)-1, MPEG-2, MPEG-4, VC-1, H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or other formats. Alternatively, the video may be provided in a non-digital format and converted to a digital format using a scanner or transcoder. The input video data may include various types of video clips or programs, e.g., television episodes, movies, and other content generated as primary content of interest to consumers. The video data may also include audio, or audio alone may be used.
[0080]
[0105] The video streaming system 2100 may include one or more computer servers or modules 2102, 2104, and 2107 distributed across one or more computers. Each server 2102, 2104, 2107 may include or be operatively coupled to one or more data stores 2109, such as databases, indexes, files, or other data structures. The video content server 2102 may access a data store (not shown) of various video segments. The video content server 2102 may provide video segments as directed by a user interface controller that communicates with client devices. As used herein, a video segment refers to a distinct portion of frame-based video data, such as a television episode, a motion picture, a recorded live performance, or other video content that may be used in a streaming video session for viewing.
[0081]
[0106] In some embodiments, the video ad server 2104 may access a data store of relatively short videos (e.g., 10-second, 30-second, or 60-second video ads) configured as advertisements for particular advertisers or messages. The advertisements may be provided to advertisers in exchange for some type of payment, or may include promotional messages, public service messages, or some other information for the system 2100. The video ad server 2104 may serve video ad segments as directed by a user interface controller (not shown).
[0082]
[0107] The video streaming system 2100 may also include a pre-analysis optimization process 110 .
[0083]
[0108] The video streaming system 2100 may further include an integration and streaming component 2107 that integrates video content and video advertisements into streaming video segments. For example, the streaming component 2107 may be a content server or a streaming media server. A controller (not shown) can determine the selection or configuration of advertisements within the streaming video based on any suitable algorithm or process. The video streaming system 2100 may include other modules or units not shown in FIG. 21 , such as a management server, a commerce server, a network infrastructure, an advertisement selection engine, etc.
[0084]
[0109] The video streaming system 2100 may be connected to a data communications network 2112. The data communications network 2112 may include a local area network (LAN), a wide area network (WAN), such as the Internet, a telephone network, a wireless network 2114 (e.g., a wireless cellular telecommunications network (WCS)), or some combination of these or similar networks.
[0085]
[0110] One or more client devices 2120 can communicate with the video streaming system 2100 via the data communications network 2112, the wireless network 2114, or another network. Such client devices may include, for example, one or more laptop computers 2120-1, desktop computers 2120-2, “smart” mobile phones 2120-3, tablet devices 2120-4, network-enabled televisions 2120-5, or combinations thereof, via a router 2118 for a LAN, a base station 2117 for the wireless network 2114, or some other connection. In operation, such client devices 2120 may send and receive data or instructions to the system 2100 in response to user input or other input received from a user input device. In response, the system 2100 can provide video segments and metadata from the data store 2109 to the client device 2120 in response to a selection of a media program. The client device 2120 can use a display screen, projector, or other video output device to output video content from streaming video segments in a media player and receive user input to interact with the video content.
[0086]
[0111] Delivery of audio-video data from the streaming component 2107 to remote client devices via computer networks, telecommunications networks, and combinations of such networks can be implemented using various methods, such as streaming. In streaming, a content server continuously streams audio-video data to a media player component running at least partially on the client device, and the media player component can play the audio-video data simultaneously as it receives the streaming data from the server. Although streaming is described, other delivery methods can be used. The media player component can begin playing the video data immediately after receiving the first portion of data from the content provider. Traditional streaming technologies use a single provider that delivers a stream of data to a set of end users. Delivering a single stream to a large audience can require high bandwidth and processing power, and the provider's required bandwidth can increase as the number of end users increases.
[0087]
[0112] Streaming media can be delivered on-demand or live. Streaming allows for instant playback at any point within a file. End users can skip through a media file and start playback or change playback to any point within the media file. Thus, end users do not have to wait for a file to download incrementally. Streaming media is typically delivered from a small number of dedicated servers with high bandwidth capabilities through dedicated devices that accept requests for video files and use information about the format, bandwidth, and structure of those files to deliver only the amount of data needed to play the video, at the speed required to play the video. Streaming media servers can also take into account the transmission bandwidth and capabilities of the media player on the destination client. The streaming component 2107 can communicate with the client device 2120 using control and data messages to adapt to changing network conditions as the video is played. These control messages can include commands to enable control functions such as fast-forwarding, fast-rewinding, pausing, or seeking to a specific part of the file at the client.
[0088]
[0113] Because the streaming component 2107 transmits video data only when needed and at the required rate, precise control over the number of streams served can be maintained. Viewers cannot watch high data rate video over a lower data rate transmission medium. However, a streaming media server (1) provides users with random access to video files, (2) allows monitoring of who is watching which video programs and for how long, (3) uses transmission bandwidth more efficiently because only the amount of data needed to support the viewing experience is transmitted, and (4) video files are not stored on the viewer's computer but are discarded by the media player, thus allowing for more control over the content.
[0089]
[0114] The streaming component 2107 may use TCP-based protocols such as Hypertext Transfer Protocol (HTTP) and Real-Time Messaging Protocol (RTMP). The streaming component 2107 can also deliver live webcasts and can multicast, allowing two or more clients to tune into a single stream, thus conserving bandwidth. Streaming media players may not rely on buffering the entire video to provide random access to any point in a media program. Instead, this is achieved using control messages sent from the media player to the streaming media server. Other protocols used for streaming are HTTP Live Streaming (HLS) or Dynamic Adaptive Streaming over HTTP (DASH). The HLS and DASH protocols deliver video over HTTP via a playlist of small segments, typically made available at various bitrates from one or more content delivery networks (CDNs). This allows the media player to switch both bitrate and content source for each segment. This switching helps compensate for network bandwidth fluctuations and infrastructure failures that may occur during video playback.
[0090]
[0115] Delivery of video content via streaming can be accomplished under a variety of models. In one model, users pay to view video programs, for example, by paying a fee for access to a library of media programs or a limited portion of a media program, or by using a pay-per-view service. In another model, widely adopted by broadcast television shortly after its inception, sponsors pay for the presentation of media programs in exchange for the right to present advertisements during or adjacent to the presentation of the program. In some models, advertisements are inserted at predetermined times, sometimes called "ad slots" or "ad breaks," within the video program. In streaming video, media players can be configured to prevent client devices from playing video without playing predetermined advertisements during designated ad slots.
[0091]
[0116] Referring to Figure 22, a schematic diagram of an apparatus 2200 for viewing video content and advertisements is shown. In selected embodiments, the apparatus 2200 may include a processor (CPU) 2202 operably coupled to processor memory 2204, which holds binary-coded functional modules for execution by the processor 2202. Such functional modules may include an operating system 2206 for handling system functions such as input / output and memory access, a browser 2208 for displaying web pages, and a media player 2210 for playing videos. The memory 2204 may hold additional modules not shown in Figure 22, for example, modules for performing other operations described elsewhere herein.
[0092]
[0117] The bus 2214 or other communication components may support communication of information within the device 2200. The processor 2202 may be a dedicated or special-purpose microprocessor configured or operable to perform particular tasks in accordance with the features and aspects disclosed herein by executing machine-readable software code that defines those tasks. The processor memory 2204 (e.g., random access memory (RAM) or other dynamic storage device) may be coupled to the bus 2214 or directly to the processor 2202 and may store information and instructions executed by the processor 2202. The memory 2204 may also store temporary variables or other intermediate information during the execution of such instructions.
[0093]
[0118] A computer-readable medium in storage device 2224 is connected to bus 2214 and can store static information and instructions for processor 2202; for example, storage device (CRM) 2224 can store modules for operating system 2206, browser 2208, and media player 2210 when device 2200 is powered off, and from which the modules can be loaded into processor memory 2204 when device 2200 is powered on. Storage device 2224 can include a non-transitory computer-readable storage medium that holds information, instructions, or some combination thereof, for example, instructions that, when executed by processor 2202, configure or enable device 2200 to perform one or more operations of the methods described herein.
[0094]
[0119] A network communications (comm.) interface 2216 may also be connected to the bus 2214. The network communications interface 2216 may provide or support bidirectional data communications between the device 2200 and one or more external devices, such as the streaming system 2100, optionally via a router / modem 2226 and a wired or wireless connection 2225. Alternatively, or additionally, the device 2200 may include a transceiver 2218 connected to an antenna 2229, through which the device 2200 may communicate wirelessly with a base station for a wireless communications system or with the router / modem 2226. Alternatively, the device 2200 may communicate with the video streaming system 2100 via a local area network, a virtual private network, or other network. In another alternative, the device 2200 may be incorporated as a module or component of the system 2100 and communicate with other components via the bus 2214 or by some other modality.
[0095]
[0120] Device 2200 may be connected (e.g., via bus 2214 and graphics processing unit 2220) to display unit 2228. Display 2228 may include any suitable configuration for displaying information to an operator of device 2200. For example, display 2228 may include or utilize a liquid crystal display (LCD), a touchscreen LCD (e.g., a capacitive display), a light emitting diode (LED) display, a projector, or other display device to present information to a user of device 2200 in a visual display.
[0096]
[0121] One or more input devices 2230 (e.g., an alphanumeric keyboard, microphone, keypad, remote control, game controller, camera, or camera array) may be connected to bus 2214 via user input port 2222 for communicating information and commands to device 2200. In selected embodiments, input device 2230 may provide or support control over cursor positioning. Such cursor control devices, also referred to as pointing devices, may be configured as mice, trackballs, trackpads, touchscreens, cursor direction keys, or other devices for receiving or tracking physical movements and converting the movements into electrical signals indicative of cursor movement. A cursor control device may be integrated into display unit 2228, for example, using a touch-sensitive screen. The cursor control device may communicate directional information and command selections to processor 2202 and control cursor movement on display 2228. A cursor control device may have two or more degrees of freedom, for example, allowing the device to specify a cursor position in a plane or three-dimensional space.
[0097]
[0122] Some embodiments may be embodied in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform methods described by some embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured or operable to perform those described in some embodiments.
[0098]
[0123] As used in this description and throughout the claims that follow, "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in this description and throughout the claims that follow, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0099]
[0124] The above description shows various implementations, along with examples of how aspects of some embodiments may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of some embodiments as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents may be employed without departing from the scope of the invention as defined by the claims.
Claims
1. generating, by a computing device, a first representation of a first association between bitrate and quality based on a first characteristic of a first portion of the video; generating, by the computing device, a second representation of a second association between bitrate and quality based on a second characteristic of a second portion of the video; analyzing, by the computing device, the first representation to determine a first list of bit rates for the first portion of video, and analyzing the second representation to determine a second list of bit rates for the second portion of video, wherein the first list of bit rates is different from the second list of bit rates; outputting, by the computing device, the first list of bit rates to use in encoding the first portion of video and the second list of bit rates to use in encoding the second portion of video; A method comprising:
2. generating the first representation comprises generating a first prediction of the first relevance for bit rate and quality; generating the second representation comprises generating a second prediction of the second relevance for bit rate and quality. The method of claim 1.
3. Generating the first representation includes: inputting the first features into a prediction network; generating a first prediction of the first representation based on the first features; Generating the second representation includes: inputting the second features into the prediction network; generating a second prediction of the second representation based on the second features. The method of claim 1.
4. Analyzing the first representation or analyzing the second representation may include: generating a list of potential bit rates based on the first representation or the second representation; and refining the list of potential bit rates based on qualities associated with the potential bit rates to determine the first list of bit rates or the second list of bit rates.
5. Refining said list of potential bit rates includes: The method of claim 4 , comprising removing a first potential bit rate from the list of potential bit rates.
6. Removing the potential bitrate determining a second potential bit rate; comparing a first quality of the first potential bitrate to a second quality of the second potential bitrate; and determining whether to remove the first potential bit rate based on the comparing.
7. The method of claim 6 , wherein the first potential bitrate is removed from the list of potential bitrates when a difference between the first quality and the second quality meets a threshold.
8. Refining said list of potential bit rates includes: The method of claim 4 , comprising adding a first potential bit rate to the list of potential bit rates.
9. Adding the potential bit rate determining a second potential bit rate and a third potential bit rate within the list of potential bit rates; comparing a first quality of the second potential bitrate to a second quality of the third potential bitrate; and determining whether to add the first potential bit rate based on the comparing.
10. The method of claim 9 , wherein the first potential bit rate is added when a difference between the first quality and the second quality meets a threshold.
11. Analyzing the first representation or analyzing the second representation may include: analyzing a plurality of first representations for a plurality of segments in the first portion of video or the second portion of video; determining a first minimum bit rate and a first maximum bit rate for each of the plurality of first representations; and determining a second minimum bit rate and a second maximum bit rate for the first portion of video or the second portion of video based on the first minimum bit rate and the first maximum bit rate for each of the plurality of first representations.
12. Analyzing the first representation or analyzing the second representation may include:
12. The method of claim 11, comprising generating a list of potential bit rates based on the second minimum bit rate and the second maximum bit rate, wherein potential bit rates in the list of potential bit rates are analyzed to determine the first list of bit rates or the second list of bit rates.
13. Analyzing the first representation or analyzing the second representation may include: determining potential candidate bit rates for removal for each of the plurality of first representations, wherein the potential candidate bit rates for removal are potential removals from the list of potential bit rates; and determining whether to remove potential candidate for removal bit rates based on the potential candidate for removal bit rates included as potential candidates for removal in one or more of each of the plurality of first representations.
14. Analyzing the first representation or analyzing the second representation may include: determining potential additional candidate bit rates for each of the plurality of first representations, wherein the potential additional candidates are potential additions to the list of potential bit rates; and determining whether to add a potential bitrate based on the potential additional candidate bitrates included as potential additional candidates in one or more of each of the plurality of first representations.
15. the first list of bit rates includes bit rates that are not included in the second list of bit rates. The method of claim 1.
16. encoding the first portion of video using the first list of bit rates to generate a plurality of first encoded portions for the first portion of video; 10. The method of claim 1, further comprising: encoding the second portion of video using the second list of bit rates to generate a plurality of second encoded portions for the second portion of video.
17. selecting an encoded segment from the plurality of first encoded portions for a first profile in a profile ladder; selecting an encoded segment from the plurality of second encoded portions for the first profile in the profile ladder; 17. The method of claim 16, further comprising:
18. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a computing device, generating a first representation of a first association between bitrate and quality based on a first characteristic of a first portion of the video; generating a second representation of a second association between bitrate and quality based on a second characteristic of a second portion of the video; analyzing the first representation to determine a first list of bit rates for the first portion of video, and analyzing the second representation to determine a second list of bit rates for the second portion of video, wherein the first list of bit rates is different from the second list of bit rates; and outputting the first list of bit rates to use for encoding the first portion of video and the second list of bit rates to use for encoding the second portion of video.
19. generating the first representation comprises generating a first prediction of the first relevance for bit rate and quality; generating the second representation comprises generating a second prediction of the second relevance for bit rate and quality.
20. The non-transitory computer-readable storage medium of claim 18.
20. one or more computer processors; generating a first representation of a first association between bitrate and quality based on a first characteristic of a first portion of the video; generating a second representation of a second association between bitrate and quality based on a second characteristic of a second portion of the video; analyzing the first representation to determine a first list of bit rates for the first portion of video, and analyzing the second representation to determine a second list of bit rates for the second portion of video, wherein the first list of bit rates is different from the second list of bit rates; and a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable to output the first list of bit rates to use for encoding the first portion of video and the second list of bit rates to use for encoding the second portion of video; and An apparatus comprising:
Citation Information
Patent Citations
Techniques to optimize bitrate and resolution during encoding
JP2018513604A
Segment quality-guided adaptive stream creation
JP2022008071A