Dynamic Selection of Candidate Bitrates for Video Encoding
Patent Information
- Application Number
- JP2023199246
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-06
- Filing Date
- 2023-11-24
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2043-11-24
AI Technical Summary
Existing adaptive bitrate streaming methods use static lists of candidate average bitrates, which fail to account for the varying characteristics of different video segments, leading to suboptimal encoding and playback quality issues such as redundant encoding, unnecessary quality gaps, and poor viewing experiences.
A pre-analysis optimization process dynamically selects candidate average bitrates based on the characteristics of each video segment or chunk, using rate-distortion curves and machine learning to generate an optimized list of bitrates for encoding, ensuring consistent quality and minimizing resource waste.
This approach improves video quality by optimizing the selection of transcoded segments for adaptive bitrate streaming, reducing redundant encoding and quality gaps, resulting in a better viewing experience.
Smart Images

Figure 00000029_0000 
Figure 00000030_0000 
Figure 00000031_0000
Abstract
Description
[Background technology]
[0001] One method of delivering video to client devices uses adaptive bitrate streaming (ABR). Adaptive bitrate streaming is based on providing multiple streams (often called variants or profiles) encoded with different levels of video attributes, such as different levels of bitrate and / or quality. A profile ladder lists different profiles available for a client to use when streaming segments of a video. The client can dynamically select a profile based on network conditions and other factors. The video is segmented (e.g., divided into separate segments, typically several seconds long), and the client can switch from one profile to another at segment boundaries when network conditions change. For example, a video delivery system may wish to give a client a profile with a higher bitrate that improves the quality of the video being streamed when network conditions with higher available bandwidth are experienced. When network conditions with lower available bandwidth are experienced, the video delivery system may wish to give a client a profile with a lower bitrate so that the client can play the video without having any playback issues, such as rebuffering or download failures. Summary of the Invention
[0002]
[0002] The included drawings are for illustrative purposes and merely serve to provide examples of possible structures and operations of the disclosed inventive systems, apparatus, methods, and computer program products. These drawings in no way limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed implementations. [Brief description of the drawings]
[0003] [Figure 1]
[0003] FIG. 1 illustrates a system for dynamically selecting a list of candidate average bit rates according to some embodiments. [Diagram 2]
[0004] 1 illustrates an example of a portion of a video according to some embodiments. [Figure 3A]
[0005] 4 illustrates an example of generating a list of candidate average bitrates according to some embodiments. [Figure 3B]
[0006] FIG. 2 illustrates an example of generating encoded segments according to some embodiments. [Figure 4]
[0007] FIG. 1 illustrates an example of clustering encoded segments into multiple pools according to some embodiments. [Diagram 5]
[0008] FIG. 1 illustrates an example of a selection process according to some embodiments. [Figure 6]
[0009] 4A-4C are diagrams illustrating example rate-distortion curve graphs that may be used to select encoded segments for pooling according to some embodiments. [Figure 7]
[0010] FIG. 13 illustrates an example of selected encoded segments for a per-segment profile according to some embodiments. [Figure 8]
[0011] 4A-4C are diagrams illustrating examples of different rate-distortion curves for video content in accordance with some embodiments. [Figure 9]
[0012] 4A-4C illustrate different characteristics using different encoding configurations according to some embodiments. [Figure 10]
[0013] 4 illustrates an example of using static candidate average bit rates for different rate-distortion curves according to some embodiments. [Figure 11]
[0014] 4 illustrates an optimized candidate average bitrate list according to some embodiments. [Figure 12]
[0015] FIG. 2 illustrates a more detailed example of a segment quality-driven adaptive (SQA) system and pre-analysis optimization process according to some embodiments. [Figure 13]
[0016] FIG. 1 illustrates a more detailed example of a rate-distortion (RD) prediction system in accordance with some embodiments. [Figure 14]
[0017] FIG. 13 illustrates the output of a predictive network according to some embodiments. [Figure 15]
[0018] FIG. 2 illustrates a simplified flowchart of a method for performing an optimization process for selecting a list of candidate average bit rates according to some embodiments. [Figure 16]
[0019] 4 illustrates an example of determining bounds for a list of candidate average bitrates according to some embodiments. [Figure 17]
[0020] 4 illustrates an example of eliminating candidate average bitrates based on quality according to some embodiments. [Figure 18]
[0021] FIG. 13 illustrates an example in which a minimum gap is used to eliminate candidate average bit rates according to some embodiments. [Figure 19]
[0022] 6 is a graph illustrating when adding candidate average bit rates may be advantageous in accordance with some embodiments. [Figure 20]
[0023] 4 illustrates a decision to add candidate average bit rates according to some embodiments. [Figure 21]
[0024] FIG. 1 illustrates a video streaming system in communication with multiple client devices over one or more communication networks according to one embodiment. [Figure 22]
[0025] FIG. 1 shows a schematic diagram of a device for viewing video content and advertisements. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0004]
[0026] Techniques for a video distribution system are described herein. In the following description, for purposes of explanation, numerous examples and specific details are set forth to provide a thorough understanding of some embodiments. Some embodiments, as defined by the claims, may include some or all of the features in these examples, either alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0005]
[0027] The system can adaptively generate a list of bitrates used to encode the video. The list of bitrates may be referred to as candidate average bitrates (CAB). The encoder transcodes a segment of the video using each bitrate in the list of candidate average bitrates. In some embodiments, the system can dynamically select the list of candidate average bitrates for different portions of the video, such as for different chunks of the video. A chunk may be an independent coding unit that the encoder encodes using the same settings. A video may include one or more chunks, and each chunk may include multiple segments. In some embodiments, the list of candidate average bitrates may be set at the chunk level. Although the list of candidate average bitrates is described as being set at the chunk level, the list of candidate average bitrates may be set for different portions of the video.
[0006]
[0028] The encoder can encode a segment of the video using a bitrate in a list of candidate average bitrates to generate a number of candidate segments. A segment quality-driven adaptation (SQA) process can select a segment from the candidate segments to use for a profile in the profile ladder. The target of the process is to optimize (e.g., minimize) the storage or delivery footprint of a portion of the video while maintaining similar quality.
[0007]
[0029] Each video may have different characteristics. Similarly, different parts in the same video may have different characteristics. Using a static list of candidate average bit rates for all parts of a video or for multiple videos may not give optimal results. For example, a static list of candidate average bit rates may encode a video with simple video content with more bit rates than required. Also, a video with complex video content may be encoded with low quality due to insufficient bit rates. Furthermore, a static list of candidate average bit rates may generate segments with irregular quality gaps from an encoding point of view. For example, adjacent profiles may have similar video quality that is redundant with each other or may have unacceptably large quality gaps. Having similar video quality for adjacent profiles may be unnecessary and may not provide much advantage in viewing quality. For example, if two bit rates in the list of candidate average bit rates result in an encoded segment with similar quality, transcoding the segment with those two bit rates may be redundant and may waste resources. Also, having a large quality gap may result in a poor viewing experience during playback, as the quality may change abruptly when playback switches from one profile to another.
[0008]
[0030] To overcome the above drawbacks, the pre-analysis optimization process can dynamically select bit rates in a list of candidate average bit rates for a video. To select the list of candidate average bit rates, the pre-analysis optimization process can analyze a portion of the video and output an optimized list of candidate average bit rates for the portion. For example, the pre-analysis optimization process can analyze characteristics of each portion and output a list of candidate average bit rates for each portion. In some embodiments, the pre-analysis optimization process can predict characteristics of the portion, such as a rate-distortion curve that describes quality versus bit rate for the portion. The pre-analysis optimization process uses the respective rate-distortion curves to determine an optimal list of bit rates for each portion.
[0009]
[0031] The optimization process provides many advantages. For example, the process provides an optimal selection of transcoded segments to select when selecting segments for a profile in the profile ladder. If the list of candidate average bit rates is set at a static value for the entire video and / or is the same for several different videos, suboptimal transcoding may result. Different videos, and even different parts of the same video, may have diverse characteristics. Thus, a static list of candidate average bit rates may be suboptimal for some videos or parts of videos. The use of a dynamic list of candidate average bit rates based on the characteristics of parts of videos may result in higher quality videos and viewing experiences, since the segment quality-driven adaptation process may have a better selection of encoded segments to select to form a profile for the profile ladder.
[0010]
[0032] system
[0033] FIG. 1 illustrates a system 100 for dynamically selecting a list of candidate average bit rates, according to some embodiments. The system 100 includes a content delivery network 102, a client 104, and a video delivery system 106. The source files may include different types of content, such as video, audio, or other types of content information. Video may be used for illustration purposes, but other types of content may be understood. In some embodiments, the source files may be received in a format that requires encoding into another format, as described below. For example, the source file may be a mezzanine file that includes compressed video. The mezzanine file may be encoded to generate other files, such as different profiles of the video.
[0011]
[0034] A content provider may operate the video delivery system 106 to provide a content delivery service that allows entities to request and receive media content. The content provider may use the video delivery system 106 to coordinate the distribution of media content to the clients 104. Although a single client 104 is described, multiple clients 104 may be using the service. The media content may be different types of content, such as on-demand video from a library of videos, and live video. In some embodiments, live video may be where video is available based on a linear schedule. Video may also be provided on-demand. On-demand video may be content that can be requested at any time and is not limited to viewing on a linear schedule. Video may be a program, such as a movie, a show, an advertisement, etc.
[0012]
[0035] The clients 104 may include different computing devices such as smartphones, living room devices, televisions, set-top boxes, tablet devices, etc. The clients 104 include a media player 112 capable of playing content such as videos. In some embodiments, the media player 112 may receive segments of videos and play these segments. The client 104 may send a request for a segment to one of the content delivery networks 102 and then receive the requested segment for playback on the media player 112. A segment may be a portion of a video, such as 6 seconds of a video.
[0013]
[0036] A video may be encoded into a profile ladder that includes multiple profiles. Each profile may correspond to a different configuration, which may be a different level of bitrate and / or quality, but may also include other characteristics, such as codec type, computing resource type (e.g., computer processing units), etc. Each video may have an associated profile with a different configuration. The profiles may be classified at different levels, and each level may be associated with a different configuration. For example, a level may be a combination of bitrate, resolution, codec, etc., and each level may be associated with a different bitrate, such as 400 kilobytes per second (kbps), 650 kbps, 1000 kbps, 1500 kbps, ... 12000 kbps, etc. Also, each level may be associated with another characteristic, such as a quality characteristic (e.g., resolution). The profile levels may be referred to as higher or lower, such that a profile with a higher bitrate or quality may be rated higher than a profile with a lower bitrate or quality. An encoder may use the characteristics to encode the source video. For example, an encoder may encode a source video at a target bitrate of 1500 kbps.
[0014]
[0037] The content delivery network 102 includes a server that can deliver videos to the clients 104. The content delivery network 102 receives requests for segments of videos from the clients 104 and delivers the segments of videos to the clients 104. The clients 104 may request segments of videos from one of the profile levels based on current playback conditions. The playback conditions may be any conditions experienced based on the playback of the video, such as available bandwidth, buffer length, etc. For example, the clients 104 may use an adaptive bitrate algorithm to select a profile for the video based on the current available bandwidth, buffer length, or other playback conditions. The clients 104 may continually evaluate the current playback conditions and switch between profiles during playback of a segment of the video. For example, during playback, the media player 112 may request a different profile for the video asset. For example, if low bandwidth playback conditions are being experienced, the media player 112 may request a lower profile associated with a lower bitrate for the upcoming segment of the video. However, if playback conditions of higher available bandwidth are being experienced, the media player 112 may request a higher level profile associated with a higher bandwidth for the upcoming segment of the video.
[0015]
[0038] A segment quality-driven adaptive processing system (SQA system) 108 can encode segments using a list of candidate average bit rates. The SQA system 108 then selects a segment for each profile using an optimization process. For example, the SQA system 108 can adaptively select a segment with an optimal bit rate for each profile in a profile ladder while maintaining a similar quality level. The SQA system 108 enables the system to maintain a quality similar to or matching the target bit rate while minimizing the number of bits required to store or distribute the content.
[0016]
[0039] The pre-analysis optimization process 110 can dynamically generate a list of candidate average bit rates for the portions of the video. In some embodiments, the pre-analysis optimization process 110 can predict characteristics of each of the portions of the video, such as a rate-distortion curve. The pre-analysis optimization process 110 then selects candidate average bit rates for the portions of the video based on analyzing the characteristics of each of the portions of the video.
[0017]
[0040] The following first describes the segment quality-driven adaptive processing process, and then describes the dynamic selection of the list of candidate average bit rates in more detail.
[0018]
[0041] Segment quality driven adaptation process
[0042] As discussed above, the optimization process 110 can dynamically select a list of candidate average bit rates for portions of a video. The portions of a video may be of different sizes. FIG. 2 illustrates an example of a portion of a video according to some embodiments. A video 200 may be divided into different portions at a segment level and a chunk level. In some embodiments, at 202, a video 200 may be divided into chunk level portions. For example, a number of chunks may be included in the video 200, chunk_0, chunk_1, ... chunk_m. Each respective chunk may be divided into smaller portions, which may be referred to as segments. For example, at 204, chunk_0 is divided into segments, segment_0, segment_1, ... segment_n. Similarly, although not shown, chunk_1 may be divided into its own respective segments, segment_0, segment_1, ... segment_n. A segment may be shorter in length than a chunk. For example, a chunk may be a two minute video and a segment may be a five second video.
[0019]
[0043] In a segment quality-driven adaptation process, the SQA system 108 may process each segment of the video 200 to generate multiple encodings of each respective segment based on a list of candidate average bit rates. For illustrative purposes, the optimization process 110 selects a list of candidate average bit rates for each chunk, but the list of candidate average bit rates may be selected for different portion sizes, such as per segment, for multiple chunks, etc. The bit rates included in each respective list of candidate average bit rates may be optimized based on characteristics associated with the respective portion (e.g., chunk and / or segment) of the video using the list of candidate average bit rates. Given different characteristics for different chunks, the respective lists of candidate average bit rates may be different. However, it may be possible for the bit rates for multiple chunks in the respective lists of candidate average bit rates to be the same.
[0020]
[0044] 3A illustrates an example of generating a list of candidate average bitrates according to some embodiments. At 202, chunks chunk_0, chunk_1, chunk_2, ... chunk_n are shown. At 302, the optimization system 110 has a list of candidate average bitrates selected for each chunk based on characteristics of each respective chunk. For example, for chunk_0, list #0 of candidate average bitrates is based on characteristics for chunk_0, and list #1 of candidate average bitrates is based on characteristics for chunk_1, etc. In some examples, for chunk_0, list #0 of candidate average bitrates may include bitrates of 8500, 7750, 7000, 6250, 5500, 4750, 4000, 3250 kilobytes per second (Kbps). For chunk_1, list #1 of candidate average bitrates may include bitrates of 7000, 6250, 4750, 4000, 3250, 2000, and 1250 Kbps.
[0021]
[0045] The list of candidate average bit rates may include bit rates to be used by the encoder to encode the respective segments. Traditionally, the candidate average bit rates may have statically included the same bit rates. Sometimes, two types of bit rates were used for all chunks. The first type may be a target average bit rate and the second type may be an intermediate average bit rate. The target average bit rate may be a base bit rate associated with a profile in a profile ladder for adaptive bit rate encoding. The intermediate average bit rate may be a supplement to the target average bit rate. For example, additional bit rates between the target average bit rates may be added. The use of the intermediate average bit rate may provide additional bit rates for encoding additional encoding segments that may have different characteristics, such as quality, than the encoded segments from the target average bit rate. In some cases, the optimization process 110 may include bit rates from the target average bit rate and / or the intermediate average bit rate in the list of candidate average bit rates. For example, the optimization process 110 may include the target average bit rate in the list of candidate average bit rates, but may dynamically select other bit rates. In another example, the optimization process may dynamically select a bitrate in the list of candidate average bitrates based solely on the characteristics of the chunk.
[0022]
[0046] As described above, the encoder generates encoded segments for the chunks. Figure 3B shows an example of generating encoded segments according to some embodiments. At 204, the segments for chunks of chunk_0 are shown as segment_0, segment_1, segment_2, ..., segment_n. At 304, a list of candidate average bitrates (list of CABs) for chunk_0 is used. In some embodiments, the same list of candidate average bitrates for chunk_0 is used for all segments of the chunk. However, multiple different lists of candidate average bitrates may be used for different segments of the chunk. The encoder then encodes the segments of chunk_0 using the list of candidate average bitrates.
[0023]
[0047] At 306, the encoded segments for each segment are listed. For each segment, the encoder uses an average bit rate in the list of candidate average bit rates to encode the segment. The encoder can target each average bit rate when encoding the segment. This results in a set of encoded segments for each segment of the chunk, such as ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, ..., ENC_S0_CAB_n encoded segments for segment_0. In the notation, ENC_S0 represents the encoded segment for segment_0, and CAB_0, CAB_1, CAB_2, etc. represent candidate average bit rates. For example, CAB_0 may be 8500 Kbps, CAB_1 may be 7750 Kbps, and CAB_2 may be 7000 Kbps. Each encoded segment may be encoded at the same quality level, such as 1080p. The process may be repeated for another quality level using the list of candidate average bit rates.
[0024]
[0048] For each segment, the optimization process 110 clusters the encoded segments into multiple pools. Each pool may correspond to one profile. FIG. 4 illustrates an example of clustering encoded segments into multiple pools according to some embodiments. At 402, multiple encoded segments are shown for candidate average bit rates. Each segment may have an associated value for a quality indicator. For example, encoded segment ENC_S0_CAB_0 may have a quality of quality_S0_c0, encoded segment ENC_S0_CAB_1 may have a quality of quality_S0_c1, etc. In the notation, quality_S0 represents the encoded segment of segment_0, and c0, c1, c2, etc. represent the quality for this encoded segment.
[0025]
[0049] Different methods may be used to include the encoded segments in the pools 401-1, 404-2, 404-p. For example, each pool may have or be associated with a profile. Each profile may be associated with a target bitrate, which may be the maximum bitrate that may be used to encode the segments in the associated profile. The SQA system 108 may include the encoded segments starting with the highest average bitrate that may be used in the associated profile for the pool. The SQA system 108 may then add other encoded segments at other bitrates that are smaller than the maximum bitrate. This may result in different encoded segments being included in each pool. For example, pool S0_Pool_0 may include segments ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, etc. And pool S0_Pool_1 may include encoded segments ENC_S0_CAB_2, ENC_S0_CAB_3, ENC_S0_CAB_4, etc. Thus, pool S0_Pool_1 may contain encoded segments starting at a bitrate less than the maximum bitrate in pool S0_Pool_0. If the encoded segments are encoded at bitrates of 8500, 7750, 7000, 6250, 5500, 4750, 4000, 3250 Kbps, pool S0_pool_0 may start with encoded segments at average bitrates of 8500, 7750, 7000, etc., and pool S0_pool_1 may start with segments encoded at average bitrates of 7000, 6250, 5500, etc. In some examples, example bitrates for pools may be pool_0: 8500, 7700, 7000, 6250, 5500, 4750, pool_1: 7000, 6250, 5500, 4750, 4000, and pool_p: 5500, 4750, 4000, 3250.
[0026]
[0050] From each pool, the SQA system 108 may select one encoded segment based on using a selection process. FIG. 5 shows an example of a selection process according to some embodiments. The following process may be performed for each pool. At 404-1, pool S0_pool_0 from FIG. 4 is shown with its respective encoded segments. The SQA system 108 may use one or more rules to select an encoded segment for each pool. At 502, the SQA system 108 selects an encoded segment ENC_S0_CAB_1 for pool S0_pool_0. In some embodiments, the SQA system 108 may attempt to select an encoded segment with a minimum bitrate that has a quality value that meets a criterion. In some examples, the SQA system 108 may start with the first encoded segment in the pool, such as the segment with the highest bitrate. The SQA system 108 then selects a neighboring encoded segment in the pool, such as the encoded segment with the next highest bitrate. If the first and second encoded segments have similar quality (e.g., within a threshold), the SQA system 108 selects the encoded segment with the lowest bitrate. The SQA system 108 may continue the comparison using adjacent encoded segments in the pool, such as the second and third encoded segments. When the adjacent encoded segments do not have similar quality, the process may end. Other methods may also be used, such as starting with the encoded segment with the lowest bitrate. The process may also select the segment with the lowest bitrate that has a quality within the threshold of another segment, such as the segment with the highest bitrate. The following describes an example of a process using a rate-distortion curve.
[0027]
[0051] 6 shows an example of a rate-distortion curve graph 600 that may be used to select encoding segments for a pool according to some embodiments. In the graph 600, the Y-axis is quality and the X-axis is bitrate. The curve 602 defines the relationship between quality and bitrate. For example, the curve may plot the rate and distortion of a segment or chunk, although the curve may plot other characteristics of quality and bitrate.
[0028]
[0052] The coding segments may be listed as A, B, C, D, E, F on the curve 602 based on their respective rates and distortions. At 604, an example of coded segments having similar quality is shown. In this case, coding segment C and coding segment D have similar bitrates and similar quality. For example, the quality difference between coded segment C and coded segment D may meet a threshold value min_gap (e.g., equal and / or smaller). In this case, the SQA system 108 may select coding segment D since the quality difference is minimal since this coding segment has a lower bitrate compared to coding segment C, but segment D provides similar quality compared to segment C.
[0029]
[0053] The SQA system 108 can also collapse coded segments whose quality exceeds an upper boundary. For example, the upper boundary at 606 can be the boundary used to determine coded segments as candidates to collapse. In this case, the SQA system 108 can select one or more of the segments that are above the upper threshold, such as to select only one segment (e.g., segment B) or to select fewer segments found to be above the upper threshold (e.g., to select two of the four segments). In other examples, coded segments A and B can be removed. The SQA system 108 may also remove coded segments whose quality falls below a lower boundary. For example, a lower threshold is shown at 608. The SQA system 108 can select one or more of the segments that are below the lower threshold, such as to select only one segment (e.g., segment F) or to select fewer segments found to be below the lower threshold. In other examples, coded segments E and F can be removed. The upper and lower thresholds may be used to restrict segments for a profile that exceed or are lower than a desired bit rate or quality. One reason an upper limit is used is to restrict the bit rate used to encode a segment, and one reason a lower limit is used is to restrict a bit rate used from being too low. After processing the encoded segments to remove encoded segments, the SQA system 108 may select a segment for the profile. For example, the SQA system 108 may select the encoded segment with the lowest bit rate that has a quality level that meets the thresholds, such as within a gap with the highest quality segment. In this case, the SQA system 108 may select encoded segment D.
[0030]
[0054] While the above rules may be used to select the segments, other processes may be used. For example, the selection of the encoded segments may be based on which encoded segments have been selected for other profiles. In some examples, the segments selected may be based on reducing the storage of encoded segments where the profile may reuse segments from other profiles. Thus, the SQA system 108 may optimize the quality while minimizing the bitrate used for encoded segments that are found between the lower bound and the ceiling.
[0031]
[0055] FIG. 7 illustrates examples of selected encoded segments for each segment profile, according to some embodiments. At 702, 704, 706, and 708, encoded segments are shown for profile_0, profile_1, profile_2, and profile_P, respectively. Within a profile, the SQA system 108 may select different encoded segments with different candidate average bit rates for different segments. For example, for profile_0, segment_0 was encoded using candidate average bit rate CAB_1, segment_1 was encoded using candidate average bit rate CAB_0, segment_2 was encoded using candidate average bit rate CAB_0, etc. In some examples, in profile_0, segment_0 was encoded using a bit rate of 7750 Kbps, segment_1 was encoded using a bit rate of 8500, and segment_2 was encoded using a bit rate of 8500 Kbps. For profile_1, segment_0 was coded using CAB_4, segment_1 was coded using CAB_2, and segment_2 was coded using CAB_3. For example, in profile 1, segment_0 was coded using a bitrate of 5500 Kbps, segment_1 was coded using a bitrate of 7000 Kbps, and segment_2 was coded using a bitrate of 6250 Kbps.
[0032]
[0056] The following then describes an optimization process for dynamically generating a list of candidate average bit rates.
[0033]
[0057] Optimization Process
[0058] As mentioned above, video content may have diverse characteristics, such that content in different videos may have different characteristics, and content in the same video may also have different characteristics. For example, some content, such as cartoons and news, may be easy to encode. However, some content, such as live action movies or sports, may be difficult to encode. The encoding characteristics may be different. The following describes different characteristics for content.
[0034]
[0059] 8 illustrates an example of different rate-distortion curves for video content according to some embodiments. The rate-distortion curves are used to indicate the relationship between quality and bitrate, although other metrics may be used to indicate the relationship between quality and bitrate for video content. Different rate-distortion curves may be shown for different chunks of a video, although the rate-distortion curves may be different for different portions of a video, such as segments, chunks, multiple chunks, or different videos.
[0035]
[0060] Three chunks, chunk_A, chunk_B, and chunk_C, are shown with graphs 802, 804, and 806 of rate-distortion curves for the chunks, respectively. In graph 802, the quality changes steeply at lower bit rates, but at higher bit rates, the quality does not change much. In graph 804, the quality changes as the bit rate increases with a constant correlation. In graph 806, the quality at lower bit rates may change only minimally, while the quality increases steeply at higher bit rates.
[0036]
[0061] In addition to different content producing different rate-distortion curves, different encoding configurations may also produce different encoding results. Different encoding configurations may include using different encoders (e.g., x264, x265, etc.) or different encoding parameters (RDO level, B-frames, number of references, etc.). Figure 9 illustrates different characteristics using different encoding configurations according to some embodiments. For the same segment or chunk, a first encoding configuration in 902 results in different characteristics compared to a second encoding configuration shown in 904. Encoding configuration A results in a rate-distortion curve similar to chunk_A above, and encoding configuration B results in a rate-distortion curve similar to chunk_B above, even though these rate-distortion curves are for the same content.
[0037]
[0062] Considering that the rate-distortion curves above may be different, using a static list of candidate average bitrates may not be optimal. For example, using the same list of candidate average bitrates for different rate-distortion curves may not give optimal results. Figure 10 shows an example of using static candidate average bitrates for different rate-distortion curves according to some embodiments. Graphs 802, 804, and 806 show different rate-distortion curves for different chunks shown in Figure 8. The dotted lines in each graph indicate different bitrates in the list of candidate average bitrates. Some problems may arise when using a fixed list of candidate average bitrates. For example, in graph 802, at 1008, the two highest candidate average bitrates may be redundant since they have similar quality as the third candidate average bitrate at 1010. That is, to give an encoded segment with similar quality, only one bitrate may need to be encoded, such as the bitrates listed in 1010.
[0038]
[0063] In graph 804, at 1012, two candidate average bit rates may be redundant because these two encoded segments have similar quality compared to the encoded segment having the next lower bit rate shown at 1014. As above, only one bit rate may need to be encoded, such as the lowest bit rate at 1014, to give encoded segments with similar quality.
[0039]
[0064] In graph 806, the three lowest candidate average bit rates may produce encoded segments with similar quality, at 1016. Also, the candidate average bit rates may be too far apart, since the quality difference between the encoded segments may be too large, at 1018. That is, to minimize the quality difference between the candidate average bit rates, it may be more desirable to have more candidate average bit rates with less quality difference.
[0040]
[0065] 11 shows an optimized candidate average bitrate list according to some embodiments. In graph 802, the SQA system 108 can dynamically select candidate average bitrates to optimize the quality found in the encoded segment. For example, in 1102, the SQA system 108 can increase the number of candidate average bitrates at bitrates where the curve is steep. Also, in 1103, the SQA system 108 can decrease the number of candidate average bitrates where the curve does not change the quality much.
[0041]
[0066] In the graph 804, at 1104, the SQA system 108 may remove candidate average bitrates from the lowest bitrates where quality may be redundant, and at 1106, the SQA system 108 may add additional bitrates to capture the varying quality at higher bitrates.
[0042]
[0067] In graph 806, at 1108, the SQA system 108 may eliminate bitrates at the low end of the curve. Also, at 1110, the SQA system 108 can space the candidate average bitrates more evenly to capture different quality levels in more even increments.
[0043]
[0068] Pre-analysis optimization process design
[0069] 12 shows a more detailed example of the SQA system 108 and pre-analysis optimization process 110 according to some embodiments. A chunk to be encoded is received. Also, an encoding configuration may be received that defines settings for encoding the chunk. The encoding configuration may include an encoder type, a quality level, etc.
[0044]
[0070] The pre-analysis optimization process 110 can receive the chunks and the encoding configurations and output an optimized list of candidate average bit rates. The RD prediction system 1202 can predict rate-distortion curves for segments and / or chunks within the chunks. Although predicting a rate-distortion curve for a segment or chunk may be described, rate-distortion curves may be generated for different portions of the video, such as for multiple chunks and / or multiple segments. As described in more detail below, the RD prediction system 1202 can use machine learning logic to generate a prediction of the rate-distortion curve of a segment.
[0045]
[0071] The predicted rate-distortion curve is output to the CAB list optimization system 1204. The CAB list optimization system 1204 can optimize a list of candidate average bitrates for the chunk based on the predicted rate-distortion curves for segments within the chunk, etc. The optimized list of candidate average bitrates may be based on characteristics of the respective chunk and may be different for chunks having content with different characteristics. This process is described in more detail below.
[0046]
[0072] The CAB list optimization system 1204 outputs the optimized list of candidate average bit rates to the SQA system 108. The SQA system 108 includes an encoding system 1206 that receives the encoding configuration, chunk, and optimized list of candidate average bit rates. The encoding system 1206 then uses each candidate average bit rate in the list to encode each segment of the chunk. After encoding each segment using the list of candidate average bit rates, the selection system 1208 selects an encoded segment for each profile in the profile ladder using a selection process as described above. The selection system 1208 outputs the selected encoded segments for the profiles in the profile ladder.
[0047]
[0073] The following describes the prediction of the characteristics of a segment and then an optimization for selecting a list of candidate average bit rates.
[0048]
[0074] RD prediction system
[0075] 13 shows a more detailed example of the RD prediction system 1202 according to some embodiments. The feature extraction system 1302 receives a chunk of video. The feature extraction system 1302 may then extract values for features that may convey information related to video transcoding. Some examples of features may be about the video content, encoding settings, etc. The extracted features may give a better prediction of the characteristics of the segments of the chunk. The values for the features are output to a prediction network 1304.
[0049]
[0076] The prediction network 1304 can use the trained model to generate characteristics for segments of chunks, such as predicted rate-distortion curves. The prediction network 1304 may use different machine learning algorithms, such as support vector machine (SVM) regression, convolutional neural network (CNN), boosting, etc. The trained model can be trained based on a particular machine learning algorithm.
[0050]
[0077] The prediction network 1304 may receive values for the features in addition to other inputs such as a segment location, an encoding configuration, and a target bit rate. The segment location may be the segment location (e.g., that segment in the video) for which to generate a rate-distortion curve, the encoding configuration may include the configuration used to encode the segment, and the target bit rate may include an output bit rate range for the segment. The prediction network 1304 may output a rate-distortion curve for the segment between the output bit rate ranges based on the features.
[0051]
[0078] FIG. 14 illustrates an output of the prediction network 1304 according to some embodiments. At 204, the segments for the chunk include segment_0, segment_1, segment_2, ..., segment_n. A rate-distortion curve may be generated for each segment in each chunk of the video. For example, at 1402, a rate-distortion curve is output for each segment. A rate-distortion curve for segment_0, a rate-distortion curve for segment_1, etc. are shown. Each rate-distortion curve is based on characteristics for the respective segment. A list of candidate average bit rates for the chunk may be generated based on the rate-distortion curves. A chunk-level rate-distortion curve may also be output.
[0052]
[0079] List of candidate average bitrate optimizations
[0080] 15 shows a simplified flowchart 1500 of a method for performing an optimization process for selecting a list of candidate average bit rates according to some embodiments. At 1502, the CAB list optimization system 1204 determines boundaries for the list of candidate average bit rates. For example, the boundaries may be maximum and minimum bit rates that may be used for the list of candidate average bit rates. Different methods may be used to determine the boundaries and are described in more detail in FIG. 16.
[0053]
[0081] At 1504, the CAB list optimization system 1204 generates a list of potential candidate average bitrates with optimal bitrate allocation. In some embodiments, one list of potential candidate average bitrates is generated for the chunk based on the maximum and minimum bitrates determined at 502. The list of potential candidate average bitrates may be generated using different methods. One method may be to use a predefined list that falls between the minimum and maximum bitrates. For example, the predefined list may include bitrates from the target average bitrate and the median average bitrate. For example, a bitrate from a predefined list within a minimum and maximum range may be used. Another method may determine a total number of potential candidate average bitrates and divide the bitrate range between the minimum and maximum bitrates into intervals. Different examples may be used, such as:
[0054]
number
[0055]
number
[0056]
number
[0057]
number
[0058] where interval_i is the interval value of i, interval_(i+1) is the interval value +1, interval_(i+2) is the interval value +2, and delta is a predetermined value.
[0059]
[0082] The total number of intervals may be set to a number, such as 10. The intervals for interval_i may be set based on the above method by dividing the range into the total number. The CAB list optimization system 1204 then selects bitrates based on the interval value to divide the range of bitrates between the minimum and maximum bitrates into a list of bitrates. For example, a minimum bitrate of 2000 and a maximum bitrate of 10,000 with an interval of 1500 and a total number of bitrates of 5 may result in a list of bitrates of 10,000, 7500, 5000, 3500, and 2000 when using equal division.
[0060]
[0083] At 1506, the CAB list optimization system 1204 refines the list of potential candidate average bit rates with an optimal quality allocation to generate an optimized list of candidate average bit rates. The quality allocation may inspect the quality for each segment and determine whether the quality satisfies one or more rules. For example, redundant candidate average bit rates, such as candidate average bit rates with similar quality, may be removed. Also, additional candidate average bit rates may be added as needed, such as when adjacent candidate average bit rates have a quality gap that exceeds a threshold, such as a too large difference. The process is described in more detail in Figures 17, 18, and 19.
[0061]
[0084] As illustrated in 1502 of FIG. 15, the CAB list optimization system 1204 determines the boundaries for the list of candidate average bit rates. FIG. 16 illustrates an example of determining the boundaries for the list of candidate average bit rates according to some embodiments. The following process is described, but other processes may be understood. For example, settings may be used to determine the minimum and maximum bit rates. In this example, the CAB list optimization system 1204 may analyze the minimum and maximum bit rates for the rate distortion curve for each segment in the chunk and determine what the minimum and maximum bit rates should be at the chunk level.
[0062]
[0085] At 1602, the rate-distortion curves for each segment are received and analyzed. The CAB list optimization system 1204 may then select a minimum bitrate and a maximum bitrate for each segment based on the respective rate-distortion curve for the segment. For example, for segment_0, a minimum bitrate and a maximum bitrate are selected based on characteristics of the rate-distortion curve for segment_0. For example, the CAB list optimization system 1204 may set a maximum quality threshold and a minimum quality threshold and use the rate-distortion curve to determine a minimum bitrate corresponding to the minimum quality threshold and a maximum bitrate corresponding to the maximum quality threshold. For segment_1, the CAB list optimization system 1204 selects a minimum bitrate and a maximum bitrate based on characteristics of the rate-distortion curve for segment_1, and so on.
[0063]
[0086] The above analysis was performed at the segment level. The CAB list optimization system 1204 then analyzes the segment level results to determine minimum and maximum values at the chunk level. At 1606, the CAB list optimization system 1204 determines a maximum value from the values for the maximum bitrate for the segment from max_bitrate_0, max_bitrate_1, max_bitrate_2, ..., max_bitrate_n, etc. The CAB list optimization system 1204 also determines a minimum value from the values for the minimum bitrate for the segment from min_bitrate_0, min_bitrate_1, min_bitrate_2, ..., min_bitrate_n, etc.
[0064]
[0087] At 1608, the CAB list optimization system 1204 outputs minimum and maximum bit rates for the chunk. In this case, the lowest minimum bit rate is selected from the minimum bit rates for the segment, and the highest maximum bit rate is selected from the maximum bit rates for the segment. The selection process may take into account the individual characteristics of the rate distortion curve for the segment and select minimum and maximum bit rates that may include all of the minimum and maximum bit rates determined at the segment level. For example, if the minimum bit rates are 2000, 3000, and 3500, the minimum bit rate selected will be 2000. Similarly, if the maximum bit rates are 10000, 9000, and 8500, the maximum bit rate selected will be 10000. Although the above process may be used, other methods of selecting minimum and maximum bit rates may be understood, such as taking an average of the values.
[0065]
[0088] As illustrated in 1506 of FIG. 15, the CAB list optimization system 1204 defines a list of candidate average bitrates with optimal quality allocation. Part of the allocation includes removing candidate average bitrates based on similar quality. The similarity of quality may be defined in different ways. For example, the CAB list optimization system 1204 determines the distance between quality values to determine whether some candidate average bitrates should be removed. FIG. 17 shows an example of removing candidate average bitrates based on quality according to some embodiments. In this example, the CAB list optimization system 1204 may determine whether the quality levels of two adjacent candidate average bitrates meet a threshold, such as being within the threshold. Then, the candidate with the higher bitrate may be removed.
[0066]
[0089] At 1702, each segment may have an associated potential removal list of encoded segments that may potentially be removed. As shown, for segment_0, the CAB list optimization system 1204 has determined that the candidate average bit rates of S0_CAB_0, S0_CAB_3, and S0_CAB_4 may be removed. These candidate average bit rates may be removed because the encoded segments may have similar quality levels that meet the threshold with respect to adjacent encoded segments. Similarly, for segment_1, the CAB list optimization system 1204 has determined that the candidate average bit rates of S1_CAB_0 and S1_CAB_2 may be removed, and for segment_n, the CAB list optimization system 1204 has determined that the candidate average bit rates for Sn_CAB_0 and Sn_CAB_3 may be removed. The segment is not removed for segment_2 because the segments are not determined to have similar quality within the threshold.
[0067]
[0090] The above analysis was at the segment level. Then, at 1704, the CAB list optimization system 1204 can use the segment level candidate average bitrates to determine candidate average bitrates to be removed at the chunk level. For example, based on the occurrence of the candidate average bitrate in different segments of the potential removal list, the CAB list optimization system 1204 can select a candidate average bitrate for the chunk level. In some embodiments, the CAB list optimization system 1204 can select a candidate average bitrate and calculate the total number of occurrences in the potential removed candidate pool. If the total number of this candidate average bitrate meets a threshold, such as at or above the threshold, the CAB list optimization system 1204 places this candidate average bitrate in the removal candidate list at the chunk level. For example, the candidate average bitrate CAB_0 is found in three of the above-mentioned segments (e.g., segment_0, segment_1, and segment_n), meeting the threshold of "3". Then, the CAB list optimization system 1204 places the candidate average bitrate of CAB_0 in the removal candidate list. Candidate average bit rates CAB_2, CAB_3, and CAB_4 may not meet the threshold because the bit rates occur in two or fewer segments in the potential removal list. Thus, the CAB list optimization system 1204 does not place these candidate average bit rates in the removal candidate list. Other methods of selecting which candidate average bit rates to remove may be understood.
[0068]
[0091] The above analysis was performed at the segment level and merged to the chunk level. However, the process may be performed at a different level. For example, the analysis may be used to merge candidate average bitrates from multiple chunks to a portion of the video covering multiple chunk levels, or from multiple chunks to the video level.
[0069]
[0092] The following describes an example of removing candidate average bitrates. Figure 18 shows an example of using a minimum gap to remove candidate average bitrates according to some embodiments. For example, at 1802, candidate average bitrates C and D have similar quality levels that meet a threshold, such as the threshold min_gap. In this case, the CAB list optimization system 1204 determines that one of the candidate average bitrates, such as candidate average bitrate C, should be removed, since this candidate average bitrate is adjacent to candidate average bitrate D, and candidate average bitrate C has a larger bitrate than candidate average bitrate D, but has a minimum quality advantage. In this case, the CAB list optimization system 1204 may compare the difference between the quality values for the adjacent candidate average bitrates to the threshold min_gap, and remove one of the candidate average bitrates when the threshold is met.
[0070]
[0093] Another part of quality allocation involves adding candidate average bit rates based on quality gaps. Figure 19 shows a graph 1900 illustrating when adding candidate average bit rates may be advantageous according to some embodiments. The CAB list optimization system 1204 may use a threshold, such as the maximum gap max_gap in 1902, to determine when to add a candidate average bit rate. For example, if there is a gap in quality values between adjacent candidate average bit rates that is greater than the threshold max_gap, such as between candidate average bit rates C and D in graph 1900, the CAB list optimization system 1204 may add a candidate average bit rate between candidate average bit rates C and D on the rate-distortion curve.
[0071]
[0094] Different methods may be used to determine how many new candidate average bitrates should be added. The CAB list optimization system 1204 may add "i" new candidates when two candidates are separated by a ratio-based threshold. Different ratios may configure the gap between the added candidates, such as 1:1, meaning each gap is equal, 1:1.5, meaning each gap is 1.5, or other ratios.
[0072]
[0095] In one possible process, the variable i is set to i=1, and the CAB list optimization system 1204 adds i new candidate average bitrates based on the ratio. For example, one candidate average bitrate named "F" may be added between points C and D. Then, if all gaps between the new adjacent candidate average bitrates are smaller than the threshold max_gap, the process ends. However, if not, the value of the variable i is incremented, such as "2", and two new candidates are added between the candidate average bitrates based on the ratio. For example, two or more candidate average bitrates may be added between points C and F and between points F and D. Then, the process continues as described above. Once the candidate average bitrates are added such that there is no gap between the candidate average bitrates D and C that is larger than the threshold, the candidate average bitrate is output.
[0073]
[0096] The above process is determined for each segment. Then, the CAB list optimization system 1204 may take the potential added candidate average bitrates at the segment level and merge the candidate average bitrates at the chunk level. Figure 20 shows the determination of adding a candidate average bitrate according to some embodiments. In 2002, each segment may have a candidate average bitrate that may potentially be added. For segment_0, two candidate average bitrates may be added in the middle between the candidate average bitrates CAB_0 and CAB_1. Also, one candidate average bitrate may be added between CAB_4 and CAB_5. For segment_1, two candidate average bitrates may be added between the candidate average bitrates CAB_0 and CAB_1. For segment_n, two candidate average bitrates may be added between the candidate average bitrates CAB_0 and CAB_1. Thus, three segments add two candidate average bitrates between candidate average bitrates CAB_0 and CAB_1, and one segment adds one candidate average bitrate between CAB_4 and CAB_5.
[0074]
[0097] The CAB list optimization system 1204 may use the segment level candidates to determine the additional candidate list for the chunk. For example, to be added to the additional candidate list at the chunk level, the CAB list optimization system 1204 may determine whether the potential added candidate average bitrate at the segment level is found within a threshold, such as the number of segments. If the threshold is 70% of the segments, the CAB list optimization system 1204 adds two candidates between the candidate average bitrates CAB_0 and CAB_1, because these candidates are found in more than 70% of the segments (three out of four segments). The CAB list optimization system 1204 does not add a candidate between CAB_4 and CAB_5, because this addition is found only in segment_0, which is less than 70% of the segments. In this case, adding additional candidate average bitrates may not be required since only one segment requires addition, and adding candidate average bitrates for all other segments of the chunk may not be useful if only one segment is affected. However, adding two candidate average bitrates between candidate average bitrates CAB_0 and CAB_1 may be beneficial since over 70% of the segments had potential additions.
[0075]
[0098] The output of the CAB list optimization system 1204 is a list of candidate average bit rates per chunk. For example, a list of candidate average bit rates per chunk, such as that described in FIG. 3A, is output by the CAB list optimization system 1204 according to some embodiments.
[0076]
[0099] conclusion
[0100] Thus, the list of candidate average bit rates can be optimized based on the characteristics found in each segment. This produces an improved list of candidate average bit rates per chunk that is optimized for the characteristics of each chunk. The candidate average bit rates can improve the quality of the selection of encoded segments available for the selection of each chunk's profile. This can improve the quality of the video in addition to improving the playback experience.
[0077]
[0101] system
[0102] Features and aspects disclosed herein may be implemented in conjunction with a video streaming system 2100 communicating with multiple client devices over one or more communication networks, as shown in Figure 21. Aspects of the video streaming system 2100 are described merely to provide one example of an application for enabling distribution and delivery of content prepared in accordance with the present disclosure. It should be understood that the technology is not limited to streaming video applications and may be adapted to other applications and delivery mechanisms.
[0078]
[0103] In one embodiment, a media program provider may include a library of media programs. For example, the media programs may be aggregated and provided through a site (e.g., a website), an application, or a browser. A user may access a site or application of the media program provider and request a media program. The user may be limited to requesting only media programs provided by the media program provider.
[0079]
[0104] In the system 2100, video data may be obtained from one or more sources, e.g., video source 2110, for use as input to the video content server 2102. The input video data may comprise raw or edited frame-based video data in any suitable digital format, e.g., Moving Picture Experts Group (MPEG)-1, MPEG-2, MPEG-4, VC-1, H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or other formats. Alternatively, the video may be provided in a non-digital format and converted to a digital format using a scanner or transcoder. The input video data may include various types of video clips or programs, e.g., television episodes, movies, and other content generated as the primary content of interest to consumers. The video data may also include audio, or only audio may be used.
[0080]
[0105] The video streaming system 2100 may include one or more computer servers or modules 2102, 2104, and 2107 distributed across one or more computers. Each server 2102, 2104, 2107 may include or be operatively coupled to one or more data stores 2109, such as databases, indexes, files, or other data structures. The video content server 2102 may access a data store (not shown) of various video segments. The video content server 2102 may serve video segments as directed by a user interface controller that communicates with a client device. As used herein, a video segment refers to a distinct portion of frame-based video data, such as a television episode, a motion picture, a recorded live performance, or other video content that may be used in a streaming video session to view the video content.
[0081]
[0106] In some embodiments, the video ad server 2104 may access a data store of relatively short videos (e.g., 10-second, 30-second, or 60-second video ads) configured as advertisements for particular advertisers or messages. The advertisements may be provided to advertisers in exchange for some type of payment, or may include promotional messages, public service messages, or some other information for the system 2100. The video ad server 2104 may serve video ad segments as directed by a user interface controller (not shown).
[0082]
[0107] The video streaming system 2100 may also include a pre-analysis optimization process 110 .
[0083]
[0108] The video streaming system 2100 may further include an integration and streaming component 2107 that integrates video content and video advertisements into streaming video segments. For example, the streaming component 2107 may be a content server or a streaming media server. A controller (not shown) may determine the selection or configuration of advertisements in the streaming video based on any suitable algorithm or process. The video streaming system 2100 may include other modules or units not shown in FIG. 21, such as a management server, a commerce server, a network infrastructure, an advertisement selection engine, etc.
[0084]
[0109] The video streaming system 2100 may be connected to a data communications network 2112. The data communications network 2112 may include a local area network (LAN), a wide area network (WAN), such as the Internet, a telephone network, a wireless network 2114 (e.g., a wireless cellular telecommunications network (WCS)), or some combination of these or similar networks.
[0085]
[0110] One or more client devices 2120 can communicate with the video streaming system 2100 via the data communications network 2112, the wireless network 2114, or another network. Such client devices may include, for example, one or more laptop computers 2120-1, desktop computers 2120-2, “smart” mobile phones 2120-3, tablet devices 2120-4, network-enabled televisions 2120-5, or combinations thereof, via a router 2118 for a LAN, via a base station 2117 for the wireless network 2114, or via some other connection. In operation, such client devices 2120 may send and receive data or instructions to the system 2100 in response to user input or other input received from a user input device. In response, the system 2100 can provide video segments and metadata from the data store 2109 to the client device 2120 in response to a selection of a media program. The client device 2120 can use a display screen, projector, or other video output device to output video content from streaming video segments in a media player and receive user input for interacting with the video content.
[0086]
[0111] The delivery of the audio-video data may be implemented using various methods, e.g., streaming, from the streaming component 2107 to the remote client device via computer networks, telecommunications networks, and combinations of such networks. In streaming, a content server continuously streams the audio-video data to a media player component operating at least partially on the client device, and the media player component can play the audio-video data simultaneously as it receives the streaming data from the server. Although streaming is described, other delivery methods may be used. The media player component may begin playing the video data immediately after receiving the first portion of the data from the content provider. Traditional streaming techniques use a single provider that delivers a stream of data to a set of end users. Delivering a single stream to a large audience may require high bandwidth and processing power, and the provider's required bandwidth may increase as the number of end users increases.
[0087]
[0112] Streaming media may be delivered on demand or live. Streaming allows for instant playback at any point in the file. An end user can skip through a media file and start playback or change playback to any point in the media file. Thus, the end user does not have to wait for the file to download incrementally. Typically, streaming media is delivered from a small number of dedicated servers with high bandwidth capabilities through dedicated devices that accept requests for video files and, using information about the format, bandwidth, and structure of those files, deliver only the amount of data needed to play the video at the speed needed to play the video. The streaming media server may also take into account the transmission bandwidth and capabilities of the media player on the destination client. The streaming component 2107 may communicate with the client device 2120 using control and data messages to adapt to changing network conditions as the video is played. These control messages may include commands to enable control functions such as fast forward, fast rewind, pause, or seek to a particular part of the file at the client.
[0088]
[0113] Because the streaming component 2107 transmits video data only when needed and at the rate required, precise control over the number of streams served can be maintained. Viewers cannot watch high data rate video over a lower data rate transmission medium. However, a streaming media server (1) provides users with random access to video files, (2) allows monitoring of who is watching which video programs and for how long, (3) uses transmission bandwidth more efficiently because only the amount of data required to support the viewing experience is transmitted, and (4) video files are not stored on the viewer's computer but discarded by the media player, thus allowing more control over the content.
[0089]
[0114] The streaming component 2107 may use TCP-based protocols such as Hypertext Transfer Protocol (HTTP) and Real-Time Messaging Protocol (RTMP). The streaming component 2107 may also deliver live webcasts and may multicast, which allows more than one client to tune into a single stream, thus conserving bandwidth. A streaming media player may not rely on buffering the entire video to provide random access to any point in a media program. Instead, this is accomplished using control messages sent from the media player to the streaming media server. Other protocols used for streaming are HTTP Live Streaming (HLS) or Dynamic Adaptive Streaming over HTTP (DASH). The HLS and DASH protocols deliver video over HTTP via a playlist of small segments that are typically made available at various bitrates from one or more content delivery networks (CDNs). This allows the media player to switch both bitrate and content source for each segment. This switching helps to compensate for network bandwidth fluctuations and infrastructure failures that may occur during the playback of the video.
[0090]
[0115] Delivery of video content through streaming can be accomplished under a variety of models. In one model, a user pays to view a video program, e.g., pays a fee for access to a library of media programs or a restricted portion of a media program, or uses a pay-per-view service. In another model, widely adopted by broadcast television shortly after its inception, sponsors pay for the presentation of a media program in exchange for the right to present advertisements during or adjacent to the presentation of the program. In some models, advertisements are inserted at predetermined times, sometimes called "ad slots" or "ad breaks," within the video program. In streaming video, a media player can be configured to prevent a client device from playing a video without playing a predetermined advertisement during a designated ad slot.
[0091]
[0116] Referring to Figure 22, a schematic diagram of an apparatus 2200 for viewing video content and advertisements is shown. In selected embodiments, the apparatus 2200 may include a processor (CPU) 2202 operatively coupled to a processor memory 2204, which holds binary coded functional modules for execution by the processor 2202. Such functional modules may include an operating system 2206 for handling system functions such as input / output and memory access, a browser 2208 for displaying web pages, and a media player 2210 for playing videos. The memory 2204 may hold additional modules not shown in Figure 22, for example, modules for performing other operations described elsewhere herein.
[0092]
[0117] The bus 2214 or other communication components may support communication of information within the device 2200. The processor 2202 may be a dedicated or special purpose microprocessor configured or operable to perform particular tasks in accordance with the features and aspects disclosed herein by executing machine-readable software code that defines the particular tasks. The processor memory 2204 (e.g., a random access memory (RAM) or other dynamic storage device) may be coupled to the bus 2214 or directly to the processor 2202 and may store information and instructions executed by the processor 2202. The memory 2204 may also store temporary variables or other intermediate information during the execution of such instructions.
[0093]
[0118] A computer readable medium in storage 2224 is connected to bus 2214 and can store static information and instructions for processor 2202, for example, storage (CRM) 2224 can store modules for operating system 2206, browser 2208, and media player 2210 when device 2200 is powered off, and from which the modules can be loaded into processor memory 2204 when device 2200 is powered on. Storage device 2224 can include a non-transitory computer readable storage medium that holds information, instructions, or some combination thereof, for example, instructions that, when executed by processor 2202, configure or enable device 2200 to perform one or more operations of the methods described herein.
[0094]
[0119] A network communication (comm.) interface 2216 may also be connected to the bus 2214. The network communication interface 2216 may provide or support bidirectional data communication between the device 2200 and one or more external devices, such as the streaming system 2100, optionally via a router / modem 2226 and a wired or wireless connection 2225. Alternatively or additionally, the device 2200 may include a transceiver 2218 connected to an antenna 2229, through which the device 2200 may communicate wirelessly with a base station for a wireless communication system or with the router / modem 2226. Alternatively, the device 2200 may communicate with the video streaming system 2100 via a local area network, a virtual private network, or other network. In another alternative, the device 2200 may be incorporated as a module or component of the system 2100 and communicate with other components via the bus 2214 or by some other modality.
[0095]
[0120] The device 2200 may be connected (e.g., via the bus 2214 and the graphics processing unit 2220) to a display unit 2228. The display 2228 may include any suitable configuration for displaying information to an operator of the device 2200. For example, the display 2228 may include or utilize a liquid crystal display (LCD), a touchscreen LCD (e.g., a capacitive display), a light emitting diode (LED) display, a projector, or other display device to present information to a user of the device 2200 in a visual display.
[0096]
[0121] One or more input devices 2230 (e.g., an alphanumeric keyboard, a microphone, a keypad, a remote controller, a game controller, a camera, or a camera array) may be connected to the bus 2214 via a user input port 2222 for communicating information and commands to the device 2200. In selected embodiments, the input device 2230 may provide or support control over cursor positioning. Such cursor control devices, also referred to as pointing devices, may be configured as mice, trackballs, trackpads, touch screens, cursor direction keys, or other devices for receiving or tracking physical movements and converting the movements into electrical signals indicative of cursor movement. The cursor control device may be incorporated into the display unit 2228, for example with a touch-sensitive screen. The cursor control device may communicate directional information and command selections to the processor 2202 and control cursor movement on the display 2228. The cursor control device may have two or more degrees of freedom, for example, allowing the device to specify a cursor position in a plane or three-dimensional space.
[0097]
[0122] Some embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform methods described by some embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured or operable to perform those described in some embodiments.
[0098]
[0123] As used in this description and throughout the claims which follow, "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in this description and throughout the claims which follow, the meaning of "in" includes "in" and "on," unless the context clearly dictates otherwise.
[0099]
[0124] The above description shows various implementations, along with examples of how aspects of some embodiments can be implemented. The above examples and embodiments should not be considered as the only embodiments, but are presented to illustrate the flexibility and advantages of some embodiments as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents may be employed without departing from the scope of the invention as defined by the claims.
Claims
1. generating, by the computing device, a first representation of a first association between bitrate and quality based on a first characteristic of a first portion of the video; generating, by the computing device, a second representation of a second association between bitrate and quality based on a second characteristic of a second portion of the video; analyzing, by the computing device, the first representation to determine a first list of bitrates for the first portion of a video, and analyzing the second representation to determine a second list of bitrates for the second portion of a video, wherein the first list of bitrates is different from the second list of bitrates; outputting, by the computing device, the first list of bit rates to use in encoding the first portion of a video and the second list of bit rates to use in encoding the second portion of a video; A method comprising:
2. generating the first representation comprises generating a first prediction of the first relevance for bit rate and quality; generating the second representation comprises generating a second prediction of the second relevance for bit rate and quality. The method of claim 1.
3. Generating the first representation includes: inputting the first features into a predictive network; generating a first prediction of the first representation based on the first features; Generating the second representation includes: inputting the second features into the prediction network; and generating a second prediction of the second representation based on the second features. The method of claim 1.
4. Analysing the first representation or the second representation may include: generating a list of potential bit rates based on the first representation or the second representation; refining the list of potential bitrates based on qualities associated with the potential bitrates to determine the first list of bitrates or the second list of bitrates; The method of claim 1 , comprising:
5. Refining the list of potential bit rates includes: removing a first potential bitrate from said list of potential bitrates. The method of claim 4 comprising:
6. removing the potential bit rate determining a second potential bit rate; comparing a first quality of the first potential bitrate to a second quality of the second potential bitrate; determining whether to remove the first potential bit rate based on the comparing; and The method of claim 5 comprising:
7. The method of claim 6 , wherein the first potential bitrate is removed from the list of potential bitrates when a difference between the first quality and the second quality meets a threshold.
8. Refining the list of potential bit rates includes: adding a first potential bitrate to said list of potential bitrates; The method of claim 4 comprising:
9. Adding the potential bit rate determining a second potential bitrate and a third potential bitrate in the list of potential bitrates; comparing a first quality of the second potential bitrate to a second quality of the third potential bitrate; determining whether to add the first potential bit rate based on said comparing; and The method of claim 8 comprising:
10. The method of claim 9 , wherein the first potential bitrate is added when a difference between the first quality and the second quality meets a threshold.
11. Analysing the first representation or the second representation may include: analyzing a plurality of first representations for a plurality of segments in the first portion of a video or the second portion of a video; determining a first minimum bitrate and a first maximum bitrate for each of the plurality of first representations; determining a second minimum bitrate and a second maximum bitrate for the first portion of a video or the second portion of a video based on the first minimum bitrate and the first maximum bitrate for each of the plurality of first representations; The method of claim 1 , comprising:
12. Analysing the first representation or the second representation may include:
12. The method of claim 11, comprising generating a list of potential bitrates based on the second minimum bitrate and the second maximum bitrate, wherein potential bitrates in the list of potential bitrates are analyzed to determine the first list of bitrates or the second list of bitrates.
13. Analysing the first representation or the second representation may include: determining potential removal candidate bitrates for each of the plurality of first representations, where the potential removal candidate bitrates are potential removals from the list of potential bitrates; determining whether to remove potential candidate for removal bit rates based on the potential candidate for removal bit rates included as potential candidates for removal in one or more of each of the plurality of first representations; The method of claim 12 comprising:
14. Analysing the first representation or the second representation may include: determining potential addition candidate bitrates for each of the plurality of first representations, where the potential addition candidates are potential additions to the list of potential bitrates; determining whether to add a potential bitrate based on the potential additional candidate bitrates included as potential additional candidates in one or more of each of the plurality of first representations; The method of claim 12 comprising:
15. the first list of bit rates includes bit rates that are not included in the second list of bit rates. The method of claim 1.
16. encoding the first portion of a video using the first list of bit rates to generate a plurality of first encoded portions for the first portion of a video; encoding the second portion of video using the second list of bit rates to generate a plurality of second encoded portions for the second portion of video; The method of claim 1 further comprising:
17. selecting an encoded segment from the plurality of first encoded portions for a first profile in a profile ladder; selecting an encoded segment from the plurality of second encoded portions for the first profile in the profile ladder; 20. The method of claim 16, further comprising:
18. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by a computing device, generating a first representation of a first association between bitrate and quality based on a first characteristic of a first portion of the video; generating a second representation of a second association between bitrate and quality based on a second characteristic of a second portion of the video; analyzing the first representation to determine a first list of bit rates for the first portion of a video, and analyzing the second representation to determine a second list of bit rates for the second portion of a video, wherein the first list of bit rates is different from the second list of bit rates; outputting the first list of bit rates to use for encoding the first portion of video and the second list of bit rates to use for encoding the second portion of video; A non-transitory computer-readable storage medium enabling the computing device to:
19. generating the first representation comprises generating a first prediction of the first relevance for bit rate and quality; generating the second representation comprises generating a second prediction of the second relevance for bit rate and quality.
20. The non-transitory computer-readable storage medium of claim 18.
20. one or more computer processors; generating a first representation of a first association between bitrate and quality based on a first characteristic of a first portion of the video; generating a second representation of a second association between bitrate and quality based on a second characteristic of a second portion of the video; analyzing the first representation to determine a first list of bit rates for the first portion of a video and analyzing the second representation to determine a second list of bit rates for the second portion of a video, where the first list of bit rates is different from the second list of bit rates; and outputting the first list of bit rates to use for encoding the first portion of video and the second list of bit rates to use for encoding the second portion of video. a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for: An apparatus comprising: