Coding operation point selection for video coding
By dynamically selecting the encoding operation point, utilization rate distortion curve and optimization process, the encoding suboptimal problem caused by the diversity of video content is solved, and more efficient resource use and stable video quality experience are achieved.
Patent Information
- Application Number
- CN202411729163.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-07
- Filing Date
- 2024-11-28
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, the static list of video encoding operation points cannot adapt to the diversity of video content, resulting in suboptimal encoding results, which may lead to redundant resource usage and unstable video quality experience.
Through the pre-analysis optimization process, the encoding operation points are dynamically selected to generate a list of encoding operation points based on the characteristics of the video, including local and global optimization processes, the video clip characteristics are analyzed using the rate distortion curve, the clustered segments are clustered, and the best encoding operation points are selected.
Improves video encoding quality, reduces resource usage, provides a more stable video quality experience and more efficient configuration file generation.
Smart Images

Figure CN120455685A_ABST
Abstract
Description
[0001] [Cross-reference to related applications]
[0002] This application is related to U.S. application No. 18 / 179,281, filed on March 6, 2023, entitled “DYNAMIC SELECTION OF CANDIDATEBITRATES FOR VIDEO ENCODING,” and U.S. application No. 18 / 295,184, filed on April 3, 2023, entitled “PREDICTION OF RATE DISTORTION CURVES FOR VIDEO ENCODING,” the entire contents of which are incorporated herein by reference for all purposes.
Technical field
[0003] Embodiments of the present application relate to encoding systems. [Background Technology]
[0004] A method for delivering video to a client device uses Adaptive Bitrate Streaming (ABR). Adaptive bitrate streaming is based on providing multiple streams (often referred to as variants or profiles) that are encoded with different video attribute levels (e.g., different bitrates and / or quality levels). A profile ladder lists the different profiles available to the client when streaming segments of a video. The client can dynamically select a profile based on network conditions and other factors. The video is segmented (e.g., divided into discrete segments, typically each a few seconds long), and the client can switch from one profile to another at the segment boundaries as network conditions change. For example, a video delivery system wants to provide a profile with a higher bitrate to the client when experiencing network conditions with higher available bandwidth, which improves the quality of the streamed video. When experiencing network conditions with lower available bandwidth, the video delivery system wishes to provide a profile with a lower bitrate to the client so that the client can play the video without any playback issues such as rebuffering or downloading failures.
[0005] Encoding operational points can be a list of significant encoding bitrates that can be used to encode segments of a video during the encoding process. Encoding operational points are set as defaults for multiple videos. However, videos may have different content with different characteristics. Furthermore, a single video may include segments or shots with different characteristics. Thus, applying a default or universal encoding operational point to content with different characteristics across videos or even within the same video may be suboptimal. [Summary of the invention]
[0006] In some embodiments, a method includes receiving a plurality of representations of a relationship between bitrate and quality for a first portion of content. A representation in the plurality of representations is based on respective second portions of content included in a first portion of a video. Clusters of the plurality of representations are generated, and the clusters are analyzed to determine a first list of encoding operation points for each cluster. The method analyzes the first list of encoding operation points for each cluster to determine a second list of encoding operation points. The second list of encoding operation points is output for encoding the first portion of content.
Brief Description of the Drawings
[0007] The included drawings are for illustrative purposes and are only used to provide examples of possible structures and operations of the disclosed inventive systems, devices, methods, and computer program products. These drawings do not limit any changes in form and details that may be made by those skilled in the art without departing from the spirit and scope of the disclosed embodiments.
[0008] Figure 1 A system for dynamically selecting a list of encoding operation points is described in accordance with some embodiments.
[0009] Figure 2 An example of a portion of a video is shown in accordance with some embodiments.
[0010] Figure 3A Depicted is an example of generation of a list of encoding operation points according to some embodiments.
[0011] Figure 3B Depicted are examples of generation of encoded segments according to some embodiments.
[0012] Figure 4 Depicted is an example of clustering encoded segments into multiple pools according to some embodiments.
[0013] Figure 5 An example of a selection process according to some implementations is depicted.
[0014] Figure 6Depicted is an example of a graph of a rate-distortion curve that may be used to select encoded segments for a pool, according to some embodiments.
[0015] Figure 7 Depicted are examples of selected encoded segments with a profile for each segment in accordance with some embodiments.
[0016] Figure 8 Depicted are examples of different rate-distortion curves for video content according to some embodiments.
[0017] Figure 9 Different characteristics using different encoding configurations according to some embodiments are shown.
[0018] Figure 10 Depicted is an example of using static encoding operating points for different rate-distortion curves in accordance with some embodiments.
[0019] Figure 11 Depicted is an optimized list of encoding operation points according to some embodiments.
[0020] Figure 12 Depicted are more detailed examples of encoding systems and pre-analysis optimization processes according to some embodiments.
[0021] Figure 13 A more detailed example of an RD prediction system according to some embodiments is depicted.
[0022] Figure 14 Depicted is the output of a prediction network according to some embodiments.
[0023] Figure 15 Depicted is a simplified flow chart of a method for selecting an encoding operation point according to some embodiments.
[0024] Figure 16 Depicted is a simplified flow chart for selecting a representative curve according to some embodiments.
[0025] Figure 17A Depicted are examples of RD curves for tiled segments according to some embodiments.
[0026] Figure 17B Depicted is a graph for selecting representative RD curves for various clusters, according to some embodiments.
[0027] Figure 18 Depicted is a simplified flow chart of a local optimization process for selecting an optimal range according to some embodiments.
[0028] Figure 19A Depicted is a graph showing an inflection point according to some embodiments.
[0029] Figure 19B Examples of lower limit points, inflection points, and upper limit points are shown according to some embodiments.
[0030] Figure 20 Depicted is a graph illustrating an example of selecting a locally optimal encoding operating point, according to some embodiments.
[0031] Figure 21 Depicted is a graph depicting selection of a locally optimal encoding operation point based on constraints, according to some embodiments.
[0032] Figure 22 Depicted are different examples of locally optimal encoding operating points for representative RD curves according to some embodiments.
[0033] Figure 23 Depicted is a graph illustrating a globally optimal encoding operating point according to some embodiments.
[0034] Figure 24 A video streaming system that communicates with multiple client devices via one or more communication networks is described according to one embodiment.
[0035] Figure 25 A schematic diagram depicting a device for viewing video content and advertisements. [Specific implementation method]
[0036] This document describes technologies for a video delivery system. In the following description, for purposes of explanation, numerous examples and specific details are set forth to provide a thorough understanding of some embodiments. Some embodiments, as defined by the claims, may include some or all of the features described in these examples, alone or in combination with other features described below, and may also include modifications and equivalents of the features and concepts described herein.
[0037] System Overview
[0038] A system can adaptively generate a list of encoding operation points for use when encoding a video. An encoding operation point can be information used by the encoding system to encode the video, such as bitrate, resolution, or other information. The encoding operation points can be used in various processes, such as when generating profiles for a profile ladder or as candidate average bitrates for generating multiple candidate segments during the Segment Quality Driven Adaptation (SQA) process. In the SQA process, the list of encoding operation points can be referred to as a Candidate Average Bitrate (CAB). The encoder uses the information in the encoding operation point list (e.g., the respective bitrates) to transcode the video segment. In some embodiments, the system can dynamically select a list of encoding operation points for different portions of the video (e.g., different chunks of the video). A chunk can be an independent coding unit that the encoder encodes using the same settings. A video can include one or more chunks, and each chunk can include multiple portions, such as segments. A segment can be an independent unit of analysis. For example, a segment can be independently decodable. Some examples of a segment can include a group of pictures that can be decodable together. In some embodiments, the list of encoding operation points can be set at the chunk level. Although the list of encoding operation points is discussed as being set at the tile level, the list of encoding operation points can be set for different portions of the video. Furthermore, for other types of content, segments and tiles can be used as units. However, there may be situations where these units cannot be used, but other portion sizes can be defined. For example, a first portion corresponding to a tile can be defined to include multiple second portions corresponding to segments.
[0039] The encoder can use a list of encoding operation points to encode segments of a video to generate multiple candidate segments. The SQA process selects segments from the candidate segments for use in a profile in a profile ladder. The goal of this process is to optimize (e.g., minimize) the storage or delivery footprint of a portion of the video while maintaining similar quality. The profile ladder can be a list of profiles associated with different bit rates from the encoding operation points. When generating the profile ladder, the encoder can generate segments using the bit rates of the encoding operation points to generate segments for multiple profiles. The profile ladder can be used for adaptive bitrate streaming playback.
[0040] Each video can have different characteristics. Furthermore, different parts within the same video can also have different characteristics. Using a static list of encoding operation points for all parts of a video, or for multiple videos, may not provide optimal results. For example, a static list of encoding operation points may encode a video with simple video content at a higher bitrate than required. Furthermore, due to insufficient bitrate, a video with complex video content may be encoded with poor quality. Furthermore, a static list of encoding operation points may generate segments with irregular quality gaps from an encoding perspective. For example, adjacent profiles may have similar video quality that is redundant with each other, or they may have unacceptably large quality gaps. Having similar video quality for adjacent profiles may be unnecessary and may not provide much benefit in terms of viewing quality. For example, if two bitrates in the encoding operation point list produce encoded segments with similar quality, encoding the segments at both bitrates may be redundant and may waste resources. Furthermore, having large quality gaps may result in a poor viewing experience during playback, as the quality may change significantly when switching from one profile to another.
[0041] To overcome the above shortcomings, the pre-analysis optimization process can dynamically select encoding operation points from a list of encoding operation points for a video. To select the list of encoding operation points, the pre-analysis optimization process can analyze a portion of the video and output an optimized list of encoding operation points for that portion. For example, the pre-analysis optimization process can analyze the characteristics of each portion and output a list of encoding operation points for each portion. In some embodiments, the pre-analysis optimization process can predict the characteristics of the portion, such as a rate-distortion curve that describes the quality and bitrate of the portion. The pre-analysis optimization process uses each rate-distortion curve to determine the optimal list of encoding operation points for each portion.
[0042] A partition can have a large number of segments. Moreover, the segments can have content with different characteristics. For example, some content of a segment in a partition may be easy to encode, such as stationary or slow motion. However, some content of another segment in the partition may be difficult to encode, such as quick motion or a detailed scene. If the same settings or the same RD curve are used for the partition to determine the encoding operation point, it is impossible to achieve optimal encoding for all of the segments in the partition. Therefore, the system can perform an optimization process to select the encoding operation point based on the RD curves of the segments from the partition.
[0043] The system can generate characteristics such as RD curves for video segments. The system can then perform a clustering process to cluster the segments into clusters. In some embodiments, a representative curve for each cluster can be determined. Thereafter, the system can perform a local optimization process and a global optimization process. The local optimization process can determine the optimal range for each representative RD curve. The system can then determine a local optimal encoding operation point for each representative RD curve within the optimal range. In the global optimization process, the system can calculate a global optimal encoding operation point from the local optimal encoding operation points of each representative RD curve. For example, the system can use the representative RD curves of the segments in the block to determine the global optimal encoding operation point for the block.
[0044] Once the global optimal encoding operation point is determined, the system can use the global encoding operation point. For example, the system can use the global encoding operation point to generate profiles in a profile ladder or to generate individual candidate segments for the SQA process.
[0045] The optimization process provides many advantages. If the list of encoding operation points is set with static values for the entire video and / or is the same for multiple different videos, suboptimal transcoding may result. Different videos, as well as different parts of the same video, can have different characteristics. Therefore, a static list of encoding operation points may be suboptimal for some videos or parts of videos. Using a dynamic list of encoding operation points based on the characteristics of the video portion can produce a higher quality video and viewing experience. For example, a segment quality driven adaptation process can have a better selection of encoded segments to choose from to form the profiles of the profile ladder. Moreover, the encoding operation points can generate better profiles for the profile ladder.
[0046] Using clusters and representative RD curves can also result in reduced resource usage. For example, fewer RD curves can be analyzed, and results can be determined more quickly. Furthermore, local and global optimization processes are used significantly more than analyzing RD curves. This process generates an optimal range of encoding operation points for a block, which represents the characteristics of each segment of the block.
[0047] system
[0048] Figure 1A system 100 for dynamically selecting a list of encoding operation points according to some embodiments is depicted. System 100 includes a content delivery network 102, a client 104, and a video delivery system 106. Source files can include different types of content, such as video, audio, or other types of content information. The term video may be used for discussion purposes, but other types of content, such as audio, images, etc., are also understood. In some embodiments, a source file may be received in a format that requires encoding into another format, as discussed below. For example, a source file may be a mezzanine file that includes compressed video. The mezzanine file can be encoded to create other files, such as different profiles of the video.
[0049] A content provider may operate a video delivery system 106 to provide a content delivery service that allows entities to request and receive media content. A content provider may use the video delivery system 106 to coordinate the distribution of media content to clients 104 via the content delivery network 102. Although a single client 104 has been discussed, multiple clients 104 may use this service. The media content may be of different types, such as video on demand and live video from a video library. In some embodiments, live video may be where video is available based on a linear schedule. Video may also be provided on demand. Video on demand may be content that can be requested at any time and is not limited to viewing on a linear schedule. Video may be a program, such as a movie, a show, an advertisement, or the like.
[0050] Clients 104 can include various computing devices, such as smartphones, living room devices, televisions, set-top boxes, tablet devices, and the like. Clients 104 include a media player 112 that can play content, such as videos. In some implementations, media player 112 receives video clips and can play these clips. Clients 104 can send requests for the clips to a content delivery network 102 and then receive the requested clips for playback on media player 112. These clips can be portions of a video, such as six seconds of video.
[0051] Videos can be encoded in a profile ladder that includes multiple profiles. Each profile can correspond to a different configuration, which can be different levels of bit rate or quality, but can also include other characteristics, such as codec type, computing resource type (e.g., computer processing unit), etc. Each video can have an associated profile with different configurations. Profiles can be categorized at different levels, and each level can be associated with a different configuration. For example, a level can be a combination of bit rate, resolution, codec, etc. In some embodiments, each level can be associated with a different bit rate, such as 400 kilobytes per second (kbps), 650kbps, 1000kbps, 1500kbps...12000kbps. Moreover, each level can be associated with another characteristic, such as a quality characteristic (e.g., resolution). Profile levels can be referred to as higher or lower, for example, a profile with a higher bit rate or quality can be rated higher than a profile with a lower bit rate or quality. The encoder can use this characteristic to encode the source video. For example, the encoder can encode the source video at a target bit rate of 1500kbps.
[0052] Content delivery network 102 includes a server that can deliver videos received from video delivery system 106 to client 104. Content delivery network 102 receives a request for a video clip from client 104 and delivers the video clip to client 104. Client 104 can request a video clip from one of the profile levels based on current playback conditions. Playback conditions can be any conditions experienced during video playback, such as available bandwidth, buffer length, etc. For example, client 104 can use an adaptive bitrate algorithm to select a profile for the video based on the current available bandwidth, buffer length, or other playback conditions. Client 104 can continuously evaluate the current playback conditions during playback of a video clip and switch between profiles. For example, during playback, media player 112 can request different profiles for a video asset. For example, if low-bandwidth playback conditions are experienced, media player 112 can request a lower profile associated with a lower bitrate for the upcoming video clip. However, if higher-available bandwidth playback conditions are experienced, media player 112 can request a higher-level profile associated with a higher bandwidth for the upcoming video clip.
[0053] The encoding system 108 can encode the clips using a list of encoding operation points. When using the SQA process, the encoding system 108 uses an optimization process to select clips for each profile. For example, the encoding system 108 can adaptively select clips with the best bitrate for each profile in the profile ladder while maintaining a similar quality level. The SQA process allows the system to maintain a quality similar to or matching the target bitrate while minimizing the number of bits required to store or deliver the content. Alternatively, when generating the profile ladder, the encoding system 108 encodes the clips using the encoding operation points for each partition with which the clip is associated. The encoded clips then form the profiles in the profile ladder for the video. The profiles formed for the clips in different partitions can be optimized based on the characteristics of the clips in each partition.
[0054] The pre-analysis optimization process 110 can dynamically generate a list of encoding operation points for a portion of a video (e.g., a partition). In some embodiments, the pre-analysis optimization process 110 can predict various characteristics of a portion of the video, such as a rate-distortion curve, based on characteristics of various segments of the partition. The pre-analysis optimization process 110 then selects an encoding operation point for the partition based on analyzing the various characteristics of the segments of the partition.
[0055] As described above, the optimization process 110 may dynamically select a list of encoding operation points for a portion of the video. The portion of the video may be of different sizes. Figure 2 An example of portions of a video according to some embodiments is shown. A video 200 can be divided into different portions at the segment level and the chunk level. In some embodiments, at 202, the video 200 can be divided into chunk-level portions. For example, the video 200 can include multiple chunks chunk_0, chunk_1...chunk_m. Each of the chunks can be divided into smaller portions, which can be referred to as segments. For example, at 204, chunk_0 is divided into segments segment_0, segment_1...segment_n. Similarly, although not shown, chunk_1 can be divided into its own segments segment_0, segment_1...segment_n. The length of a segment can be shorter than a chunk.
[0056] In the segment quality driven adaptation process, the encoding system 108 may process each segment of the video 200 to generate multiple encodings for each of the segments based on a list of encoding operation points. For the purposes of this discussion, the optimization process 110 selects a list of encoding operation points per partition; however, lists of encoding operation points may be selected for different portion sizes (e.g., per segment, multiple partitions, etc.). The bitrates included in each of the respective lists of encoding operation points may be optimized based on characteristics associated with the respective portions (e.g., partitions and / or segments) of the video for which the list of encoding operation points will be used. Given the different characteristics of different partitions, the respective lists of encoding operation points may be different. However, it is possible that the bitrates for multiple partitions in the respective lists of encoding operation points are the same.
[0057] The fragment quality driven adaptive processing will now be described in more detail below, followed by the dynamic selection of the encoding operation point list.
[0058] Fragment quality-driven adaptive process
[0059] Figure 3A An example of generating a list of encoding operation points according to some embodiments is depicted. At 202, chunks chunk_0, chunk_1, chunk_2, ..., chunk_n are shown. At 302, the optimization system 110 selects a list of encoding operation points for each chunk based on characteristics of each of the respective chunks. For example, for chunk_0, list #0 of encoding operation points is based on the characteristics of chunk_0. Furthermore, list #1 of encoding operation points is based on the characteristics of chunk_1, and so on. In some examples, for chunk_0, list #0 of encoding operation points may include bit rates of 8500, 7750, 7000, 6250, 5500, 4750, 4000, and 3250 kilobytes per second (Kbps). For chunk_1, list #1 of encoding operation points may include bit rates of 7000, 6250, 4750, 4000, 3250, 2000, and 1250 Kbps.
[0060] The list of encoding operation points may include the bitrates used by the encoder to encode each segment. Conventionally, encoding operation points may statically include the same bitrate. Sometimes, two types of bitrates are used for all segments. The first type may be a target average bitrate, while the second type may be an intermediate average bitrate. The target average bitrate may be the base bitrate associated with a profile in a profile ladder used for adaptive bitrate encoding. The intermediate average bitrate may be in addition to the target average bitrate. For example, additional bitrates between the target average bitrates may be added. The use of intermediate average bitrates may provide additional bitrates for encoding additional encoded segments, which may have different characteristics, such as quality, than the encoded segments at the target average bitrate. In some cases, the optimization process 110 may include bitrates from the target average bitrate and / or the intermediate average bitrates in the list of encoding operation points. For example, the optimization process 110 may include the target average bitrate in the list of encoding operation points but dynamically select a different bitrate. In other examples, the optimization process may dynamically select a bit rate from a list of encoding operation points based solely on characteristics of the partitions.
[0061] As described above, the encoder generates encoded segments in blocks. Figure 3B An example of generating encoded segments according to some embodiments is depicted. At 204, the segments segment_0, segment_1, segment_2, ..., segment_n of a partition chunk_0 are shown. At 304, a list of encoding operation points (a list of CABs) for chunk_0 is used. In some embodiments, the same list of encoding operation points used for chunk_0 is used for all segments of the partition. However, multiple different lists of encoding operation points can be used for different segments of the partition. The encoder then encodes the segments of chunk_0 using the list of encoding operation points.
[0062] At 306, the encoded segments for each of the segments are listed. For each segment, the encoder encodes the segment using the average bitrate from the list of encoding operation points. The encoder can target each average bitrate when encoding the segment. This produces a set of encoded segments for each segment of the partition, such as the encoded segments ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, ... ENC_S0_CAB_n for segment_0. In this notation, ENC_S0 represents the encoded segment of segment_0, and CAB_0, CAB_1, CAB_2, etc. represent the encoding operation points. For example, CAB_0 can be 8500 Kbps, CAB_1 can be 7750 Kbps, and CAB_2 can be 7000 Kbps. Each encoded segment can be encoded at the same quality level (e.g., 1080p). This process can be repeated for another quality level using the list of encoding operation points.
[0063] For each segment, the optimization process 110 clusters the encoded segments into a plurality of pools. Each pool may correspond to a profile. Figure 4 An example of clustering coded segments into multiple pools according to some embodiments is depicted. At 402, multiple coded segments for an encoding operation point are shown. Each segment can have an associated value for a quality metric. For example, coded segment ENC_S0_CAB_0 can have a quality of quality_S0_c0, coded segment ENC_S0_CAB_1 can have a quality of quality_S0_c1, and so on. In this notation, quality_S0 represents the coded segment of segment_0, while c0, c1, c2, etc. represent the quality of the coded segment.
[0064] Different methods can be used to include encoded segments in pools 401-1, 404-2, ... 404-p. For example, each pool can have a profile or be associated with a profile. Each profile can be associated with a target bitrate, which can be the maximum bitrate that can be used to encode the segments in the associated profile. The encoding system 108 can include encoded segments starting with the highest average bitrate available for the associated profile of the pool. The encoding system 108 can then add other encoded segments with other bitrates less than the maximum bitrate. This can result in different encoded segments included in each pool. For example, pool S0_Pool_0 can include segments ENC_S0_CAB_0, ENC_S0_CAB_1, ENC_S0_CAB_2, etc. Furthermore, pool S0_Pool_1 can include encoded segments ENC_S0_CAB_2, ENC_S0_CAB_3, ENC_S0_CAB_4, etc. Therefore, pool S0_Pool_1 may include encoded segments starting at a bitrate less than the maximum bitrate in pool S0_pool_0. If the encoded segments are encoded at a bitrate of 8500, 7750, 7000, 6250, 5500, 4750, 4000, 3250 Kbps, pool S0_pool_0 may start with encoded segments having an average bitrate of 8500, 7750, 7000, etc., and pool S0_pool_1 may start with encoded segments having an average bitrate of 7000, 6250, 5500, etc. In some examples, example bit rates for the pools may be pool_0: 8500, 7700, 7000, 6250, 5500, 4750, pool_1: 7000, 6250, 5500, 4750, 4000, and pool_p: 5500, 4750, 4000, 3250.
[0065] From each pool, the encoding system 108 may select an encoded segment based on a usage selection process. Figure 5 An example of a selection process according to some embodiments is depicted. The following process may be performed for each pool. At 404-1, a pool from Figure 4The encoding system 108 may use one or more rules to select encoded segments for each of the pools. At 502, the encoding system 108 selects the encoded segment ENC_S0_CAB_1 from pool S0_pool_0. In some embodiments, the encoding system 108 may attempt to select the encoded segment with the lowest bitrate that has a quality value that meets a criterion. In some examples, the encoding system 108 may start with the first encoded segment in the pool (e.g., the segment with the highest bitrate). The encoding system 108 then selects adjacent encoded segments in the pool, e.g., the segment with the next highest bitrate. If the first and second encoded segments have similar quality (e.g., within a threshold), the encoding system 108 selects the encoded segment with the lowest bitrate. The encoding system 108 may continue comparing the second encoded segment with adjacent encoded segments in the pool (e.g., the third encoded segment). The process may end if the adjacent encoded segments do not have similar quality. Other approaches may also be used, such as starting with the encoded segment with the lowest bitrate. Furthermore, the process may select the segment with the lowest bitrate that has a quality within a threshold of another segment (eg, the segment with the highest bitrate).An example of a process using rate-distortion curves will be described below.
[0066] Figure 6 An example of a graph 600 of a rate-distortion curve that can be used to select encoded segments for a pool according to some embodiments is depicted. In graph 600, the Y-axis is quality and the X-axis is bitrate. Curve 602 defines the relationship between quality and bitrate. For example, the curve can plot rate and distortion for a segment or partition, but the curve can also plot other characteristics of quality and bitrate.
[0067] Based on the respective rates and distortions of the encoded segments, the encoded segments may be ranked as A, B, C, D, E, and F on curve 602. At 604, an example of encoded segments with similar quality is shown. In this case, encoded segment C and encoded segment D have similar bit rates and similar quality. For example, the quality difference between encoded segment C and encoded segment D may satisfy a threshold value, min_gap (e.g., be equal to and / or less than). In this case, since the quality difference is minimal, encoding system 108 may select encoded segment D because it has a lower bit rate than encoded segment C, but provides similar quality as segment C.
[0068] The encoding system 108 may also collapse encoded segments whose quality exceeds an upper bound. For example, the upper bound at 606 may be the bound used to determine encoded segments as candidates for collapse. In this case, the encoding system 108 may select one or more segments above the upper threshold, such as selecting only one segment (e.g., segment B) or selecting fewer of the segments found to be above the upper threshold (e.g., selecting two of four segments). In other examples, encoded segments A and B may be removed. Furthermore, the encoding system 108 may remove encoded segments whose quality falls below a lower bound. For example, a lower threshold is shown at 608. The encoding system 108 may select one or more segments below the lower threshold, such as selecting only one segment (e.g., segment F) or selecting fewer of the segments found to be below the lower threshold. In other examples, encoded segments E and F may be removed. The upper and lower thresholds may be used to limit segments of a profile that exceed or fall below a desired bitrate or quality. One reason for using an upper bound is to limit the bitrate used to encode a segment, and one reason for using a lower bound is to limit the bitrate used to be too low. After processing the encoded segments to remove the encoded segments, the encoding system 108 can select segments for the profile. For example, the encoding system 108 can select the encoded segments with the lowest bitrate that have a quality level that meets a threshold (e.g., within a certain range of the highest quality segment). In this case, the encoding system 108 can select encoded segment D.
[0069] While the above-described rules can be used to select segments, other processes can be used. For example, the selection of encoded segments can be based on which encoded segments have already been selected for other profiles. In some examples, the selected segments can be based on reducing the storage of encoded segments, where a profile can reuse segments from other profiles. Thus, the encoding system 108 can optimize the quality and minimize the bitrate for encoded segments found to be between the lower and upper bounds.
[0070] Figure 7An example of selected encoded segments for a profile for each segment according to some embodiments is depicted. Encoded segments for Profile_0, Profile_1, Profile_2, and Profile_p are shown at 702, 704, 706, and 708, respectively. Within a profile, the encoding system 108 can select different encoded segments with different encoding operation points for different segments. For example, for profile_0, segment_0 is encoded using encoding operation point CAB_1, segment_1 is encoded using encoding operation point CAB_0, segment_2 is encoded using encoding operation point CAB_0, and so on. In some examples, in profile_0, segment_0 is encoded using a bitrate of 7750 Kbps, segment_1 is encoded using a bitrate of 8500, and segment_2 is encoded using a bitrate of 8500 Kbps. For profile_1, segment_0 is encoded using CAB_4, segment_1 is encoded using CAB_2, and segment_2 is encoded using CAB_3. For example, in profile_1, segment_0 is encoded using a bitrate of 5500 Kbps, segment_1 is encoded using a bitrate of 7000 Kbps, and segment_2 is encoded using a bitrate of 6250 Kbps.
[0071] Now, the optimization process of dynamically generating the encoding operation point list will be described below.
[0072] Optimization process
[0073] As mentioned above, video content can have different characteristics, for example, the content in different videos can have different characteristics, and the content in the same video can also have different characteristics. For example, some content may be easy to encode, such as cartoons or news. Cartoons may have the characteristics of large color blocks and simple / obvious textures / edges. News may have the characteristics of still scenes or moderately moving shots. However, some content may be difficult to encode, such as in live-action movies or sports. Sports are opposite to news because sports may have the characteristics of fast movement throughout the video, not only for the camera but also for the athletes in the game. Live action may have the characteristics of many rich details, colors, and other characteristics such as film grain. Therefore, the characteristics of encoding may be different. The different characteristics of content will be described below.
[0074] Figure 8Depicted are examples of different rate-distortion curves for video content according to some embodiments. Rate-distortion curves are used to illustrate the relationship between quality and bitrate, but other metrics can be used to illustrate the relationship between the quality and bitrate of video content. Different rate-distortion curves can be illustrated for different partitions of a video; however, the rate-distortion curves can be different for different portions (e.g., segments, partitions, multiple partitions) of a video or for different videos.
[0075] The three chunks, chunk_A, chunk_B, and chunk_C, are shown in graphs 802, 804, and 806 of the rate-distortion curves for the chunks, respectively. In graph 802, the quality changes at a steeper slope at lower bit rates, but at higher bit rates, the quality does not change much. In graph 804, the quality changes with increasing bit rate in a stable relationship. In graph 806, the quality at lower bit rates may change only minimally, while the quality increases at higher bit rates with a steeper slope.
[0076] In addition to producing different content with different rate-distortion curves, different encoding configurations can also produce different encoding results. Different encoding configurations can include using different encoders (e.g., x264, x265, etc.) or different encoding parameters (rate-distortion optimization (RDO) level, B frames, reference numbering, etc.). Figure 9 Different characteristics using different encoding configurations according to some embodiments are shown. For the same segment or chunk, the first encoding configuration 902 produces different characteristics compared to the second encoding configuration shown in 904. Encoding configuration A produces a rate-distortion curve similar to chunk_A above, while encoding configuration B produces a rate-distortion curve similar to chunk_B above, even though these rate-distortion curves are for the same content.
[0077] Considering that the above rate-distortion curves may be different, using a static list of coding operation points may not be optimal. For example, using the same list of coding operation points for different rate-distortion curves may not produce the best results. Figure 10 Depicts an example of using static coding operating points for different rate-distortion curves according to some embodiments. Graphs 802, 804, and 806 depict Figure 8 802 , the two highest encoding operation points at 1008 may be redundant because they have similar quality to the third encoding operation point at 1010. That is, only one encoding bitrate, such as the bitrate listed at 1010, may be required to provide encoded segments with similar quality.
[0078] In graph 804, at 1012, two encoding operation points may be redundant because the two encoded segments have similar quality compared to the encoded segment having the next lowest bit rate shown at 1014. Similar to the above, encoding at only one bit rate, such as the lowest bit rate of 1014, may be necessary to provide encoded segments with similar quality.
[0079] In graph 806, at 1016, the lowest three encoding operation points may produce encoded segments with similar quality. Furthermore, at 1018, the encoding operation points may be too far apart because the quality differences between the encoded segments may be too large. That is, more encoding operation points with smaller quality differences may be more desirable to minimize the quality differences between the encoding operation points.
[0080] Figure 11 An optimized list of encoding operation points according to some embodiments is depicted. In graph 802, the encoding system 108 can dynamically select encoding operation points to optimize the quality found in the encoded segment. For example, at 1102, the encoding system 108 can increase the number of encoding operation points at bitrates where the curve is steep. Furthermore, at 1103, the encoding system 108 can decrease the number of encoding operation points where the curve does not change the quality much.
[0081] In graph 804, encoding system 108 may remove encoding operation points from the lowest bitrate where the quality may be redundant at 1104. Also, encoding system 108 may add additional bitrates at 1106 to capture the changed quality at the higher bitrates.
[0082] In graph 806, encoding system 108 may remove the bitrate at the lower end of the curve at 1108. Also, encoding system 108 may more evenly space encoding operation points at 1110 to capture different quality levels in more even increments.
[0083] Pre-analysis optimization process design
[0084] Figure 12 A more detailed example of an encoding system 108 and a pre-analysis optimization process 110 according to some embodiments is depicted. A partition to be encoded is received. Furthermore, an encoding configuration defining settings for encoding the partition can be received. The encoding configuration can include an encoder type, a quality level, and the like.
[0085] The pre-analysis optimization process 110 may receive a partition and encoding configuration and output an optimized list of encoding operation points. The RD prediction system 1202 may predict rate-distortion curves for segments within a partition and / or for the segments. Although rate-distortion curves may be described as being predicted for segments or segments, rate-distortion curves may be generated for different portions of a video (e.g., for multiple partitions and / or multiple segments). In some embodiments, the RD prediction system 1202 may use machine learning logic to generate predictions of rate-distortion curves for segments.
[0086] The predicted rate-distortion curve is output to the coding point optimization system 1204. The coding point optimization system 1204 can, for example, generate a list of coding operation points for the partition based on the predicted rate-distortion curves of the segments in the partition. The optimized list of coding operation points can be based on the characteristics of each partition and can be different for partitions with content having different characteristics. A partition can have a large number of segments, and these segments can have content having different characteristics. For example, some content of a segment in a partition may be easy to encode, such as stills or slow motion. However, some content of another segment may be difficult to encode, such as fast motion or detailed scenes. If the same RD curve is used for the partition to determine the coding operation point, optimal encoding of the partition cannot be achieved. Therefore, the coding point optimization system 1204 can perform an optimization process to select the coding operation point based on the RD curves of the segments from the partition. This process will be described in more detail below.
[0087] The encoding point optimization system 1204 outputs an optimized list of encoding operation points to the encoding system 108. The encoding system 108 receives the encoding configuration, the partition, and the optimized list of encoding operation points. The encoder 1206 then encodes each segment of the partition using each encoding operation point in the list. That is, the encoder 1206 encodes the segment for each bitrate of the encoding operation point. After encoding each segment using the list of encoding operation points, the selection system 1208 uses the selection process described above to select encoded segments for each profile in the profile ladder for the SQA process. The selection system 1208 outputs the encoded segments selected for the profiles in the profile ladder. When the process generates the profile ladder, the encoded segments of the encoding operation points output by the encoder 1206 can be used in the profile ladder.
[0088] The prediction of characteristics of a segment will be described below, followed by the optimization of the list for selecting encoding operation points.
[0089] RD prediction system
[0090] Figure 13A more detailed example of RD prediction system 1202 according to some embodiments is depicted. Feature extraction system 1302 receives video segments. Feature extraction system 1302 can then extract values for features that can convey information related to video transcoding. Some examples of features can relate to video content, encoding settings, etc. The extracted features can provide better predictions of the characteristics of the segments. The feature values are output to prediction network 1304.
[0091] The prediction network 1304 can use the trained model to generate characteristics of the segmented segments, such as predicted rate-distortion curves. The prediction network 1304 can use different machine learning algorithms, such as support vector machine (SVM) regression, convolutional neural network (CNN), boosting, etc. The trained model can be trained based on a specific machine learning algorithm.
[0092] In addition to other inputs (e.g., segment position, encoding configuration, and target bitrate), the prediction network 1304 may also receive the value of a feature. The segment position may be the segment position where the rate-distortion curve is generated (e.g., which segment in the video), the encoding configuration may include the configuration to be used to encode the segment, and the target bitrate may include an output bitrate range for the segment. The prediction network 1304 may output a rate-distortion curve for the segment within the output bitrate range based on the feature.
[0093] Figure 14 The output of prediction network 1304 according to some embodiments is depicted. At 204, the segments of the partition include segment_0, segment_1, segment_2, ... segment_n. A rate-distortion curve can be generated for each segment in each partition of the video. For example, at 1402, a rate-distortion curve is output for each of the segments. A rate-distortion curve for segment_0, a rate-distortion curve for segment_1, and so on are shown. Each rate-distortion curve is based on characteristics of the respective segment. A list of encoding operation points for the partition can be generated based on the rate-distortion curves. Furthermore, a partition-level rate-distortion curve can be output.
[0094] Encoding operating point selection
[0095] Different segments may have different optimal encoding operation points. However, for encoder 1206, one set of encoding operation points may be used for a portion, such as each block. The following process may optimize encoding operation points based on a first portion of the video (e.g., segment-level optimization) and a second portion of the video (e.g., segment-level optimization). Segment-level optimization may be referred to as local optimization, while segment-level optimization may be referred to as global optimization. Coding point optimization system 1204 may determine the best result based on balancing local optimization and global optimization.
[0096] Figure 15 A simplified flowchart 1500 of a method for selecting an encoding operation point according to some embodiments is depicted. The following process may be performed for each partition in a video. The encoding point optimization system 1204 receives RD curves for segments of the partition. Then, at 1502, the encoding point optimization system 1204 selects a representative RD curve. The representative RD curve may be based on a cluster of RD curves determined from the RD curves of the segments in the partition. Figure 16 A selection of representative RD curves is described in more detail.
[0097] At 1504, the coding point optimization system 1204 performs a local segment-level coding operation point selection process at the segment level. The local coding operation point selection process may determine the local optimal coding operation point for each of the representative RD curves of each segment of the partition. Figure 18 Let's start by describing the local optimization process in more detail.
[0098] At 1506, the encoding point optimization system 1204 may perform a global block-level optimization process to select an optimized encoding operation point for each block. For example, the encoding point optimization system 1204 may combine the locally optimal encoding operation points selected at the segment level to form an optimized encoding operation point for the block. In some embodiments, the combination may weight the locally optimal encoding operation points differently, for example, based on the cluster to which the locally optimal encoding operation point is associated. Clusters that are considered more important may be weighted higher, and their locally optimal encoding operation points may have a greater impact on the optimized encoding operation point for the block. Figure 22 The global optimization process is described in more detail below.Once the encoding operation points for the partitions are determined, the encoder 1208 can use the encoding operation points to encode segments of the partitions.
[0099] Now, the representative curve selection process will be described below.
[0100] Representative curve selection process
[0101] The following process can be performed for each segment in the video. There may be a large number of segment-level RD curves. The representative curve selection process can select a representative curve that can best represent the RD curve of most segment-level content characteristics in the segment. Using representative curves can improve the process because fewer RD curves can be analyzed while still representing most of the content characteristics of the segment. Moreover, since fewer RD curves can be analyzed, the process can be performed faster. Although representative curves are described, they may not be used. For example, every RD curve can be analyzed.
[0102] Figure 16A simplified flowchart 1600 for selecting a representative curve according to some embodiments is depicted. At 1602, the code point optimization system 1204 receives graph information for an RD curve. The graph information can be based on characteristics of the RD curves of the segments from the block. For example, the graph information can be characteristics of each RD curve, such as the coordinates and slope of the sampling points of the RD curve.
[0103] At 1604, the code point optimization system 1204 uses the graph information to cluster the RD curves into clusters. The clusters may include RD curves with similar characteristics. The clustering process may use different clustering algorithms, such as K-means clustering, BIRCH clustering, or other clustering methods. The clustering process will be in Figure 17A Described in more detail in .
[0104] At 1606, the code point optimization system 1204 selects a representative RD curve for each cluster. Various methods can be used to select a representative RD curve for each cluster, such as selecting an RD curve based on the average of the cluster's RD curves or as the average, or creating a new RD curve representing the cluster. The representative RD curve can be analyzed to replace the RD curve in the cluster. Figure 17B The selection process is described in more detail.
[0105] Figure 17A An example of an RD curve for a segment of a block according to some embodiments is depicted. The Y-axis shows quality and the X-axis shows bit rate. Graph 1700 depicts RD curves having different characteristics. For example, the characteristics of the RD curve may include different quality values at different bit rate values. The coding point optimization system 1204 may determine a cluster of RD curves having similar characteristics within a range. For example, clusters may be formed at 1702-1, 1702-2, and 1702-3. Each RD curve within each cluster is shown with a different dotted line (cluster 1702-1), dashed line (cluster 1702-2), or dashed-dot line (cluster 1702-3).
[0106] Each cluster 1702 may include a different number of RD curves. For example, cluster 1702-1 includes four RD curves, cluster 1702-2 includes five RD curves, and cluster 1702-3 includes three RD curves. Note that this example may be simplified, and the number of RD curves in each cluster may be different. In another example, cluster 1702-3 may be the smallest cluster with five RD curves, cluster 1702-1 may be a medium-sized cluster with 20 RD curves, and cluster 1702-2 may be the largest cluster with 150 RD curves.
[0107] The number of clusters determined can be influenced based on configuration settings, such as the maximum number of clusters that can be formed and a distance threshold between all RD curves in a cluster. For example, a setting for the maximum number of clusters and a setting for a loss threshold can be used. The maximum number of clusters can be the maximum number of clusters that can be formed, where the number of clusters can be equal to or less than the maximum number of clusters. The loss threshold can be a distance threshold between all RD curves in each cluster. That is, an RD curve cannot have a value greater than the distance threshold. The code point optimization system 1204 can analyze the characteristics of the RD curves to determine the clusters based on the configuration settings.
[0108] Figure 17B A graph 1704 is depicted for selecting representative RD curves for each cluster according to some embodiments. The graph 1704 includes Figure 17A The same RD curve and cluster shown. The code point optimization system 1204 can use different processes to select a representative curve. For example, the code point optimization system 1204 can select a representative RD curve as the curve closest to the center of the cluster, which curve can have the smallest total distance to each curve in the cluster. Moreover, the code point optimization system 1204 can create an RD curve as a representative RD curve based on the RD curves of the cluster, for example, an RD curve that is the average value of all RD curves in the cluster. Other methods can also be used. In this example, a representative RD curve is shown for each cluster with a solid line, and other RD curves in the cluster are shown with dotted lines, short dashes, or dot-dash lines. For clusters 1702-1, 1702-2, and 1702-3, representative RD curves are shown at 1706-1, 1706-2, and 1706-3, respectively. Although one representative RD curve is selected in this example, the code point optimization system 1204 can select multiple RD curves for a cluster.
[0109] Table 1 describes pseudocode that can be used to determine representative RD curves for a cluster. The pseudocode can determine representative RD curves using the settings for the maximum number of clusters, max_clustering_num, and the loss threshold, loss_threshold. The function clustering_func(rd_curves) can receive RD curves and generate clusters as clustering_results. Representative RD curves, representative_rd_curves, are generated by finding RD curves with a distance less than the loss threshold, where the number of clusters is less than the maximum number of clusters.
[0110]
[0111] Table I
[0112] After determining the representative RD curve, a local optimization process is performed.
[0113] Local optimization process
[0114] Figure 18 A simplified flowchart 1800 depicts a local optimization process for selecting an optimal range according to some embodiments. The local optimization process can generate a locally optimal encoding operation point for each of the representative RD curves. The process can determine an optimal range within which the encoding operation point can be included, and then determine the best locally optimal encoding operation point within the optimal range. The following process can be performed for each of the representative RD curves.
[0115] At 1802, the code point optimization system 1204 selects a point of the RD curve, such as a knee point. The knee point may be a point to be included in an optimal range, which may be based on changes in a characteristic of the RD curve (e.g., quality or bit rate or both). In some embodiments, when both quality and bit rate change at a faster rate, the RD curve may provide more information than when one or both of the quality and bit rate do not change much. For example, as described above in Figure 10 and Figure 11 As described in , different characteristics of the RD curve can provide more information. If the bitrate changes rapidly, but the quality is stable, having multiple operating code points in the range will not provide much information because the quality has not changed much. On the other hand, if the quality changes rapidly, but the bitrate is similar, having operating code points in the range will not provide much information because the bitrate is similar. At the inflection point, the range around the inflection point can provide more information because both the quality and the bitrate are changing at a faster rate, and having more operating code points around the inflection point is better. In some embodiments, the inflection point can be the point where the maximum curvature of the RD curve is calculated. The inflection point can be calculated using different methods. In some embodiments, the code point optimization system 1204 can calculate the slope of multiple sampling points of the RD curve and select the point that includes the fastest slope change as the inflection point. Other algorithms can also be used to select the inflection point, such as calculating the rate of change of a characteristic and selecting a point based on the rate of change. Examples of inflection points will be described in . Figure 19A described in .
[0116] At 1804, the coding point optimization system 1204 selects a ceiling point of the optimal range. The ceiling point may be the highest vertex where operational coding points may be inserted. Conceptually, after the ceiling point, one of the bit rate or the quality may be similar. For example, as the bit rate increases after the ceiling point, the quality may remain similar. In some embodiments, the coding point optimization system 1204 may use the slope of the sampling points to calculate the starting point of an area where the bit rate changes but the quality remains similar within a threshold. The coding point optimization system 1204 may determine the starting point of a flat area as the ceiling point, meaning that the curve will be flat after this point. In other words, after this point, as the bit rate increases, the quality values will be similar, for example, after Figure 10 However, if the curve shape is different, the upper limit point can be determined differently, for example, at Figure 11 At 1110 in the code point, the upper limit point can be near the highest bit rate because the RD curve changes at the fastest rate as the bit rate increases. The code point optimization system 1204 can use the slope of each point to calculate the flat area starting point bitrate flat_area_starting_point And make sure that this point is greater than the inflection point bitrate knee_point The following formula can be used to calculate the upper limit bitrate ceiling_point :
[0117] bitrate ceiling_point =max(bitrate flat_area_starting_point ,bitrate knee_point )
[0118] The upper bitrate limit can be the bitrate at the inflection point or the maximum bitrate at the start of the flat region. Other constraints can be used. For example, a bitrate threshold can be used, where the bitrate cannot exceed this threshold for different requirements, such as content delivery network bandwidth costs, device playback considerations, etc. Also, a quality threshold can be used, where the quality cannot exceed this threshold because after this point, the quality will be similar due to limitations (such as human eye perception of quality).
[0119] At 1806, the code point optimization system 1204 selects a floor point of the optimal range. The floor point of the optimal range can be the lowest point that can include the optimal code point. In some embodiments, the floor point can be selected where the bit rate does not change much compared to the quality. However, other curve shapes may indicate that other floor points may be needed. Different methods can be used to determine the floor point. In some embodiments, the following formula can calculate the floor point, but other methods can be used:
[0120] bitrate floor_point
[0121] =quality2bitrate(quality knee_point -k·(quality ceiling_point
[0122] -quality knee_point ))(k>0)
[0123] Lower limit bitrate floor_point , can be based on the quality of the inflection point knee_point , minus the difference between the mass of the upper limit point and the mass of the inflection point Multiply by a constant k. Other methods can also be used. Further, a constraint can be used to select the lower limit point, such as a bit rate threshold, which indicates that the bit rate cannot fall below the threshold. Also, a quality threshold can be used, where the quality cannot fall below the threshold because lower quality may be unacceptable. A quality gap threshold can be used, where the quality gap between the upper and lower limit points should not be greater than the threshold. Figure 19B Examples of the lower limit point, the inflection point, and the upper limit point of the RD curve are shown, which will be described below.
[0124] At 1808, the encoding point optimization system 1204 selects a locally optimal encoding operation point within the optimal range. In some embodiments, this process can determine a solution under constraints that consider quality, bit rate, or both quality and bit rate. The optimization goal can use the constraints to distribute the quality of the locally optimal encoding operation point as close to the rule as possible. A score can be used to indicate how close the locally optimal encoding operation point is to the rule. The higher the score, the better the optimization result. The following may be examples of rules, but other rules may be used:
[0125] ·Equal division: interval_(i+1) / interval_i=1.
[0126] Progressive partitioning: interval_(i+1)=interval_i+delta.
[0127] Geometric partitioning: interval_(i+1) / interval_i=1.5,1.618.
[0128] Here, interval_i is the interval value of i, interval_(i+1) is the interval value+1, and delta is a predetermined value.
[0129] The total number of intervals can be set to a number, such as 10. The intervals of interval_i can be set based on the above method by dividing the range into a total number. The code point optimization system 1204 then selects bit rates based on the interval values to divide the bit rate range between the minimum bit rate and the maximum bit rate into a bit rate list. For example, when using equal division, a minimum bit rate of 2000 kbps and a maximum bit rate of 10000 kbps with an interval of 1500 kbps and a total number of bit rates of five can produce a bit rate list of 10000, 7500, 5000, 3500, and 2000 kbps.
[0130] The above rules can generate local optimal encoding operation points via equal intervals, gradually, via geometric intervals, or using other rules. The following may be examples of constraints, but other constraints may be used:
[0131] Keep the bit rate interval between two adjacent points larger than the minimum bit rate gap min_bitrate_gap.
[0132] Keep the bitrate gap between two adjacent points smaller than the maximum bitrate gap max_bitrate_gap.
[0133] During the local optimization process for selecting the locally optimal coding operation point, the following process may be used in some embodiments. First, candidate internal points are generated within the optimal range. This process may utilize a threshold, such as a minimum quality difference threshold, which specifies the difference between two adjacent candidate points and should be equal to the minimum quality difference. The code point optimization system 1204 then generates a list of candidate internal points on the RD curve based on this difference. In some embodiments, the code point optimization system 1204 generates different combinations of candidate internal points in the candidate internal point list. This process defines multiple candidate internal points in a list and enumerates all possible combinations of candidate internal points in the list. For example, if the code point optimization system 1204 generates five candidate internal points (A, B, C, D, E), and the number of points in a candidate internal point list is three, the candidate internal point list may be: {(A, B, C), (A, B, D), (A, B, E), (A, C, D), (A, C, E), (A, D, E), (B, C, D), (B, C, E), (B, D, E), (C, D, E)}. In some implementations, the upper and lower limit points may always be in all candidate lists, but are not in all candidate lists in the above example.
[0134] The code point optimization system 1204 then selects the optimal local coding operation point from all possible combinations of the candidate internal point list. The selection of the locally optimal coding operation point can be an optimal solution problem under constraints. In this example, the code point optimization system 1204 uses the equalization rule as the optimization target, and the constraint is that the bit rate gap between two adjacent points must be greater than the minimum bit rate gap min_bitrate_gap and less than the maximum bit rate gap max_bitrate_gap. The following is a method for evaluating the equalization score of a list of numbers:
[0135] Step 1: Sort all values in the list from smallest to largest, i=0, 1, 2, 3...n.
[0136] Step 2: Calculate anchor_step.
[0137] Step 3: Get each anchor position. i =value0+i*anchor_step
[0138] Step 4: Calculate the average offset of all points.
[0139] Step 5: Output the score.
[0140] Table II contains the pseudocode for the above generation of the balance score. All candidate point results are stored in the array bitrate[N][M]. The variable N is the number of candidate point lists and the combination ID combination_ID used for indexing, while the variable M is the number of points in a candidate point list. The variable variable_ID is used for indexing. The pseudocode finds the optimal combination ID for the candidate interior point list.
[0141]
[0142] Table II
[0143] Figure 20 and Figure 21 Locally optimal encoding operating point selection is described in more detail.
[0144] As discussed above in 1802, an inflection point is determined. Figure 19A A graph 1900 is depicted showing an inflection point 1902 on an RD curve 1901 according to some embodiments. The Y-axis of the graph 1900 shows quality, and the X-axis shows bit rate. The inflection point 1902 can be selected as described above. As can be seen, both the quality and the bit rate can change more quickly than at other points on the RD curve 1901.
[0145] Figure 19B An example 1904 of a lower limit point 1906, an inflection point 1902, and an upper limit point 1908 according to some embodiments is shown. As shown, the upper limit point 1908 is at a higher bit rate than the inflection point 1902. Points on the RD curve 1901 after the upper limit point 1908 may include higher bit rates, but the quality remains similar. For the lower limit point 1906, points before the lower and upper limit points 1906 of the RD curve 1901 may include changed quality values, but the bit rate remains similar. However, the bit rate and quality of points around the inflection point 1902 may change more rapidly than the lower limit point 1906 and the upper limit point 1908, and the inflection point 1902 is between the lower limit point 1906 and the upper limit point 1908.
[0146] As discussed above in 1808, a locally optimal encoding operation point is determined. Figure 20 Graph 2000 illustrates an example of selecting a locally optimal encoding operation point, according to some embodiments. The selected point can be based on an equal division process. Here, the quality within the optimal range is divided equally, and the six points corresponding to the dotted line at 2002 of RD curve 1901 can be locally optimal encoding operation points. The bitrate values of the locally optimal encoding operation points can be used to encode the segmented video. Other methods can select other locally optimal encoding operation points.
[0147] Figure 21 A graph 2100 depicting selection of a local optimal encoding operation point based on constraints according to some embodiments is shown. Constraints for a minimum bitrate gap, min_bitrate_gap, at 2102 and a maximum bitrate gap, max_bitrate_gap, at 2104 are shown on the RD curve 1901. The minimum bitrate gap forces the local optimal encoding operation point to be at a distance greater than the minimum bitrate gap. Furthermore, the maximum bitrate gap forces the local optimal encoding operation point to be no further away than the maximum bitrate gap.
[0148] Figure 22 Different examples of locally optimal encoding operation points for representative RD curves according to some embodiments are depicted. Graphs 2200, 2204, and 2208 illustrate different examples of locally optimal encoding operation points 2202-1, 2202-2, and 2202-3 from three representative RD curves 1901-1, 1901-2, and 1901-3, respectively. It can be seen that locally optimal encoding operation points can be found within different optimal ranges for different quality and bitrate values. These values can be located at points on each RD curve 1901 that provide optimal information for the RD curve as described above.
[0149] Once the locally optimal encoding operation point is determined, the encoding point optimization system 1204 determines the globally optimal encoding operation point. Figure 23 A graph 2300 illustrating globally optimal coding operation points according to some embodiments is depicted. RD curves for all segments of a partition are shown. The combination of locally optimal coding operation points from representative RD curves yields four globally optimal coding operation points at 2302. For example, the dotted lines at 2302 illustrate the four globally optimal coding operation points for the partition at corresponding bit rates. These globally optimal coding operation points may adhere to set constraints and also optimally represent the segment. For example, the coding point optimization system 1204 may optimally combine the locally optimal coding operation points to form the globally optimal coding operation point. In some embodiments, the coding point optimization system 1204 may use a weighted combination based on the characteristics of the clusters 1700. The weighting values may be based on the number of RD curves in the cluster. For example, cluster 1702-2 may include the largest number of RD curves, and the locally optimal coding operation point from this cluster may be weighted the highest. Cluster 1702-3 may include the smallest number of RD curves, and the locally optimal coding operation point from this cluster may be weighted the lowest. Furthermore, the weighted combination may also take into account other information, such as the importance of the segments (e.g., segments containing important content). Furthermore, user preferences or preferred settings may be used, such as using local optimal encoding operation points for segments at the beginning and end of the video. An example of using a weighted average of cluster sizes is shown below:
[0150]
[0151] The cluster number cluster_number may be a cluster identifier. n The value of can be based on the number of RD curves in the cluster. Other methods, such as a winner-takes-all method, can also be used. In this example, one cluster wins, and the local optimal encoding operation point from that cluster is used as the global optimal encoding operation point. In this example, the winning cluster may have a large number of RD curves compared to other clusters, and the local optimal encoding operation point of that cluster may best represent the segment. A combination of these two options can also be used. If the weighted value of a cluster is large enough, the code point optimization system 1204 can use the winner-takes-all method; otherwise, the code point optimization system 1204 can use a weighted average method.
[0152] Once the global optimal encoding operation point is determined, the encoding system 108 can use the global optimal encoding operation point. For example, the encoder 1206 can use the global optimal encoding operation point to encode the segmented segments. The segmented encoded segments can be used to generate profiles in the profile ladder or to select segments for profiles in the SQA process, such as Figure 12 discussed in .
[0153] in conclusion
[0154] Therefore, the above process uses local optimization processes and global optimization processes to generate optimized encoding operation points. The optimized encoding operation points can take into account the segment-level RD curves within the block. This improves the selection process of the encoding operation points. Furthermore, the encoding process can be improved, wherein the encoded segments can provide segments that better represent the segments of the block. Since higher-quality segments using less bitrate can be provided to the client for playback, and a more diverse set of segments for the profile is used, the playback process can be improved. The process can also use clusters and representative RD curves to use fewer computing resources.
[0155] system
[0156] The features and aspects disclosed herein may be implemented in conjunction with a video streaming system 2400 that communicates with multiple client devices via one or more communication networks, such as Figure 24 The aspects of the video streaming system 2400 are described merely to provide an example of an application for implementing the distribution and delivery of content prepared according to the present disclosure. It should be understood that the present technology is not limited to streaming video applications and may be applicable to other applications and delivery mechanisms.
[0157] In one embodiment, a media program provider may include a media program library. For example, media programs may be aggregated and provided through a site (e.g., a website), an application, or a browser. A user may access the site or application of a media program provider and request media programs. A user may be restricted to requesting only media programs provided by the media program provider.
[0158] In system 2400, video data can be obtained from one or more sources, such as video source 2410, to be used as input to video content server 2402. The input video data can include original or edited frame-based video data in any suitable digital format, such as Moving Picture Experts Group (MPEG)-1, MPEG-2, MPEG-4, VC-1, H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or other formats. In an alternative embodiment, the video can be provided in a non-digital format and converted to a digital format using a scanner or transcoder. The input video data can include various types of video clips or programs, such as television series, movies, and other content produced as primary content of interest to consumers. The video data can also include audio, or only audio can be used.
[0159] The video streaming system 2400 may include one or more computer servers or modules 2402, 2404, and 2407 distributed across one or more computers. Each server 2402, 2404, 2407 may include, or be operatively coupled to, one or more data stores 2409, such as databases, indexes, files, or other data structures. The video content server 2402 may access a data store (not shown) for various video clips. The video content server 2402 may provide video clips as directed by a user interface controller in communication with a client device. As used herein, a video clip refers to a defined portion of frame-based video data, such as may be used in a streaming video session to view a television series, movie, recorded live performance, or other video content.
[0160] In some embodiments, the video ad server 2404 can access a data store of relatively short videos (e.g., 10-second, 30-second, or 60-second video ads) configured for advertisements or messages from specific advertisers. The ads can be provided to the advertiser in exchange for some payment or can include promotional messages, public service messages, or some other information from the system 2400. The video ad server 2404 can provide the video ad clips as directed by a user interface controller (not shown).
[0161] The video streaming system 2400 may also include a pre-analysis optimization process 110 .
[0162] The video streaming system 2400 may also include an integration and streaming component 2407 that integrates video content and video advertisements into the streaming video clips. For example, the streaming component 2407 may be a content server or a streaming media server. A controller (not shown) may determine the selection or placement of advertisements in the streaming video based on any suitable algorithm or process. The video streaming system 2400 may include Figure 24 Other modules or units not depicted in the , such as management servers, business servers, network infrastructure, advertisement selection engine, etc.
[0163] The video streaming system 2400 can be connected to a data communication network 2412. The data communication network 2412 can include a local area network (LAN), a wide area network (WAN) (e.g., the Internet), a telephone network, a wireless network 2414 (e.g., a wireless cellular telecommunications network (WCS)), or some combination of these or similar networks.
[0164] One or more client devices 2420 can communicate with the video streaming system 2400 via a data communications network 2412, a wireless network 2414, or another network. Such client devices may include, for example, one or more laptop computers 2420-1, desktop computers 2420-2, "smart" phones 2420-3, tablet devices 2420-4, network-enabled televisions 2420-5, or a combination thereof, via a router 2418 for a LAN, via a base station 2417 for the wireless network 2414, or via some other connection. In operation, such client devices 2420 can send and receive data or instructions to the system 2400 in response to user input received from a user input device or other input. In response, the system 2400 can provide video clips and metadata from the data storage 2409 to the client devices 2420 in response to selections of media programs. The client devices 2420 can output video content from the streaming video clips in a media player using a display screen, projector, or other video output device, and receive user input for interacting with the video content.
[0165] The distribution of audio-video data can be implemented from streaming components 2407 to remote client devices using various methods (e.g., streaming) through computer networks, telecommunication networks, and combinations of these networks. In streaming, a content server continuously streams audio-video data to a media player component that operates at least partially on a client device, and the client device can play the audio-video data simultaneously with receiving the streaming data from the server. Although streaming has been discussed, other delivery methods can be used. The media player component can initiate playback of the data immediately after receiving the initial portion of the video data from the content provider. Traditional streaming technology uses a single provider to deliver a data stream to a group of end users. High bandwidth and processing power may be required to deliver a single stream to a large number of listeners, and the required bandwidth of the provider may increase as the number of end users increases.
[0166] Streaming media can be delivered on demand or live. Streaming enables immediate playback at any point within a file. End users can skip media files to start playback or change playback to any point in the media file. Therefore, end users do not need to wait for files to download progressively. Typically, streaming media is delivered from several dedicated servers with high bandwidth capabilities via dedicated devices. This dedicated device accepts requests for video files and uses information about the format, bandwidth, and structure of those files to deliver only the amount of data necessary to play the video at the rate required to play the video. The streaming media server can also take into account the transmission bandwidth and the capabilities of the media player on the destination client. Streaming component 2407 can communicate with client device 2420 using control messages and data messages to adapt to changing network conditions when playing the video. These control messages can include commands for enabling control functions such as fast forward, fast rewind, pause, or seeking a specific part of a file at the client.
[0167] Because the streaming component 2407 sends video data only when needed and at the required rate, precise control over the number of streams served can be maintained. Viewers would not be able to watch high data rate video over a lower data rate transmission medium. However, the streaming media server (1) provides users with random access to video files, (2) allows monitoring of who is watching what video programs and for how long, (3) uses transmission bandwidth more efficiently because only the amount of data required to support the viewing experience is transmitted, and (4) the video files are not stored on the viewer's computer but are discarded by the media player, thereby allowing more control over the content.
[0168] Streaming component 2407 can use TCP-based protocols, such as Hypertext Transfer Protocol (HTTP) and Real-time Messaging Protocol (RTMP). Streaming component 2407 can also deliver live broadcasts on the Internet and can perform multicasting, which allows more than one client to tune into a single stream, thereby saving bandwidth. Streaming media players can provide random access to any point in a media program without relying on buffering the entire video. On the contrary, this is achieved using a control message sent from a media player to a streaming media server. Other protocols for streaming are HTTP Live (HLS) or Dynamic Adaptive Streaming (DASH) over HTTP. HLS and DASH protocols deliver video over HTTP via a playlist of small segments, which are typically available from one or more content delivery networks (CDNs) with various bit rates. This allows media players to switch both bit rates and content sources on a segment-by-segment basis. Switching helps compensate for network bandwidth changes and infrastructure failures that may occur during video playback.
[0169] The delivery of video content via streaming can be accomplished under a variety of models. In one model, a user pays for viewing a video program, for example, by paying for access to a library of media programs or a limited portion of media programs, or by using a pay-per-view service. In another model, which became widely adopted by broadcast television shortly after its inception, sponsors pay for the presentation of a media program in exchange for the right to present advertisements during or adjacent to the presentation of the program. In some models, advertisements are inserted into a video program at predetermined times, which may be referred to as "ad spots" or "ad breaks." For streaming video, a media player may be configured so that a client device cannot play the video without playing a predetermined advertisement during a designated ad spot.
[0170] refer to Figure 25 , illustrates a diagrammatic view of a device 2500 for viewing video content and advertisements. In selected embodiments, the device 2500 may include a processor (CPU) 2502 operatively coupled to a processor memory 2504 that holds binary-coded functional modules for execution by the processor 2502. Such functional modules may include an operating system 2506 for handling system functions such as input / output and memory access, a browser 2508 for displaying web pages, and a media player 2510 for playing videos. The memory 2504 may hold Figure 25 Additional modules not shown in the figure, such as modules for performing other operations described elsewhere in this document.
[0171] The bus 2514 or other communication components may support the communication of information within the device 2500. The processor 2502 may be a specialized or dedicated microprocessor that is configured or operable to perform specific tasks in accordance with the features and aspects disclosed herein by executing machine-readable software code that defines the specific tasks. A processor memory 2504 (e.g., a random access memory (RAM) or other dynamic storage device) may be connected to the bus 2514 or directly to the processor 2502 and store information and instructions to be executed by the processor 2502. The memory 2504 may also store temporary variables or other intermediate information during the execution of such instructions.
[0172] The computer-readable medium in the storage device 2524 can be connected to the bus 2514 and store static information and instructions for the processor 2502; for example, the storage device (CRM) 2524 can store modules of the operating system 2506, the browser 2508, and the media player 2510 when the device 2500 is powered off, and when the device 2500 is powered on, the modules can be loaded from the storage device into the processor memory 2504. The storage device 2524 may include a non-transitory computer-readable storage medium that retains information, instructions, or some combination thereof, such as instructions that, when executed by the processor 2502, cause the device 2500 to be configured or operable to perform one or more operations of the methods described herein.
[0173] A network communication (comm.) interface 2516 may also be connected to the bus 2514. The network communication interface 2516 may optionally provide or support two-way data communication between the device 2500 and one or more external devices (e.g., the streaming system 2400) via a router / modem 2526 and a wired or wireless connection 2525. In an alternative embodiment, or in addition, the device 2500 may include a transceiver 2518 connected to an antenna 2529, through which the device 2500 may communicate wirelessly with a base station of a wireless communication system or with the router / modem 2526. In an alternative embodiment, the device 2500 may communicate with the video streaming system 2400 via a local area network, a virtual private network, or other network. In another alternative embodiment, the device 2500 may be incorporated as a module or component of the system 2400 and communicate with other components via the bus 2514 or through some other modality.
[0174] The device 2500 may be connected (e.g., via the bus 2514 and the graphics processing unit 2520) to a display unit 2528. The display 2528 may include any suitable configuration for displaying information to an operator of the device 2500. For example, the display 2528 may include or utilize a liquid crystal display (LCD), a touch screen LCD (e.g., a capacitive display), a light emitting diode (LED) display, a projector, or other display device to present information to a user of the device 2500 in a visual display.
[0175] One or more input devices 2530 (e.g., an alphanumeric keyboard, microphone, keypad, remote control, game controller, camera, or camera array) can be connected to bus 2514 via user input port 2522 to transmit information and commands to device 2500. In selected embodiments, input device 2530 can provide or support control of cursor positioning. Such a cursor control device, also known as a pointing device, can be configured as a mouse, trackball, trackpad, touch screen, cursor direction keys, or other device for receiving or tracking physical movement and converting the movement into electrical signals indicating cursor movement. The cursor control device can be incorporated into display unit 2528, for example, using a touch-sensitive screen. The cursor control device can transmit direction information and command selections to processor 2502 and control cursor movement on display 2528. The cursor control device can have two or more degrees of freedom, for example, allowing the device to specify a cursor position in a plane or three-dimensional space.
[0176] Some embodiments may be implemented in a non-transitory computer-readable storage medium for use with or in conjunction with an instruction execution system, device, system, or machine. The computer-readable storage medium contains instructions for controlling a computer system to perform the methods described in some embodiments. The computer system may include one or more computing devices. When executed by one or more computer processors, the instructions may be configured or operable to perform the operations described in some embodiments.
[0177] As used in the description herein and throughout the appended claims, "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in the description herein and throughout the appended claims, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0178] The above description illustrates various embodiments and examples of how aspects of some embodiments may be implemented. The above examples and embodiments should not be considered the only embodiments and are presented to illustrate the flexibility and advantages of some embodiments as defined by the appended claims. Based on the above disclosure and the appended claims, other arrangements, embodiments, implementations, and equivalents may be employed without departing from the scope of the present invention as defined by the claims.
Claims
1. A method comprising: receiving a plurality of representations of a relationship between bitrate and quality of a first content portion, wherein representations of the plurality of representations are based on respective second portions of the content included in the first content portion; generating clusters of the plurality of representations; analyzing the clusters to determine a first list of encoding operation points for each cluster; analyzing the first list of encoding operation points for each cluster to determine a second list of encoding operation points; and A second list of the encoding operation points is output for use in encoding the first content portion.
2. The method according to claim 1, wherein The first content portion includes chunks, and The second portion of the content includes segments, wherein the chunk includes a plurality of segments.
3. The method according to claim 2, wherein: The segments include independent analysis units in which the segments can be decoded independently, and The partition includes independent coding units in which the second list of encoding operation points of the encoder is not changed to encode the partition.
4. The method according to claim 1, wherein Generating the plurality of clusters of representations includes: The representations are analyzed based on a clustering process to form the clusters, wherein each cluster includes one or more of the representations.
5. The method according to claim 1, further comprising: Determine the representative relationships of each cluster; as well as The representative relationships of the clusters are analyzed to determine a first list of encoding operation points.
6. The method according to claim 5, wherein: Each representative relation represents the representation in each cluster, and The individual representative relations, rather than the representations in the cluster, are analyzed to determine the first list of encoding operation points.
7. The method according to claim 5, wherein: The respective representative relations are representations in the cluster or new representations generated from the representations in the cluster.
8. The method according to claim 1, wherein Analyzing the clusters to determine the first list of encoding operation points for each cluster includes: A range of the encoding operation points in the first list that should include the encoding operation points is determined.
9. The method according to claim 8, wherein Determining the scope includes: Determine the first point; Determine the second point; and An encoding operation point in the first list of encoding operation points is determined based on the first point and the second point.
10. The method according to claim 9, wherein: Determining the scope includes: A third point is determined, wherein the third point is based on a rate of change of the representation, and wherein the first point and the second point include the third point.
11. The method according to claim 9, wherein: Conditions related to bit rate or quality are applied to determine the first point or the second point.
12. The method according to claim 9, wherein Determining the encoding operation point includes: The encoding operation point is selected based on a spacing between encoding operation points within the range.
13. The method according to claim 12, wherein: Select encoding operation points include: Identify constraints; and An encoding operation point that satisfies the constraint condition is selected.
14. The method according to claim 13, wherein The constraints are based on quality or bit rate.
15. The method according to claim 1, wherein Analyzing the first list of encoding operation points for respective representative relations to determine the second list of encoding operation points includes: Determine the weight of each cluster; applying the weights for the respective clusters to the first list of encoding operation points for the respective clusters to generate a first list of weighted encoding operation points; and The first lists of weighted encoding operation points for the clusters are combined to determine the second list of encoding operation points.
16. The method according to claim 15, wherein The weight of each cluster is based on the number of representations found in each cluster.
17. The method according to claim 1, wherein Analyzing the first list of encoding operation points for respective representative relations to determine the second list of encoding operation points includes: Determining the weights of each cluster; and One of the encoding operation points in the first list of encoding operation points is selected based on the weight of the cluster.
18. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions, which, when executed by a computing device, cause the computing device to: A plurality of representations of a relationship between bit rate and quality of a first content portion is received, wherein: representations of the plurality of representations are based on respective second portions of the content included in the first content portion; generating clusters of the plurality of representations; analyzing the clusters to determine a first list of encoding operation points for each cluster; analyzing the first list of encoding operation points for each cluster to determine a second list of encoding operation points; as well as A second list of the encoding operation points is output for use in encoding the first content portion.
19. The non-transitory computer-readable storage medium of claim 18, wherein: The first content portion includes chunks, and The second portion of the content includes segments, wherein the chunk includes a plurality of segments.
20. A device comprising: one or more computer processors; and A computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable to: receiving a plurality of representations of a relationship between bitrate and quality of a first content portion, wherein representations of the plurality of representations are based on respective second portions of the content included in the first content portion; generating clusters of the plurality of representations; analyzing the clusters to determine a first list of encoding operation points for each cluster; analyzing the first list of encoding operation points for each cluster to determine a second list of encoding operation points; and A second list of the encoding operation points is output for use in encoding the first content portion.
Citation Information
Patent Citations
Dynamic selection of candidate bitrates for video encoding
US12225252B2
Prediction of rate distortion curves for video encoding
US20240305788A1