Resource-conscious video processor
The resource allocator dynamically allocates CPU cycles to digital video processing modules, addressing inefficiencies in existing systems by maintaining high video quality and preventing frame loss through real-time adjustments.
Patent Information
- Application Number
- DE112016004532
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-12-08
- Filing Date
- 2016-12-08
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2036-12-08
AI Technical Summary
Existing digital video processing systems inefficiently allocate CPU cycles, leading to potential frame loss and suboptimal video quality due to temporary peaks in CPU usage without dynamic adjustment.
A resource allocator dynamically allocates CPU cycles to digital video processing modules based on real-time availability and content complexity, adjusting operation modes to maintain high video quality while optimizing density and minimizing frame loss.
Ensures consistent high video quality by adaptively reallocating CPU resources, preventing frame loss and maintaining optimal video processing even during fluctuations in available cycles.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD OF THE INVENTION
[0001] Embodiments of the invention relate to optimizing the allocation of hardware resources to software responsible for processing digital video. BACKGROUND
[0002] With statistical multiplexing, the number of bits allocated to each of a plurality of digital video channels is dynamically adjusted several times per second based on the complexity of the digital video being carried by that particular channel. The complexity of digital video is a measure of how many data (or "bits") are required to describe how the digital video is to be displayed. If a particular channel requires an increase in the number of bits to adequately describe the complexity of the digital video being carried by it, additional bits can be allocated to that channel from another channel that is not currently using all of its allocated bits.
[0003] EP 1 503 595 A2 discloses a system for encoding / decoding videos with real-time complexity adaptation. The system comprises an encoder / decoder (codec) configured to cause the encoding / decoding algorithms used by the codec to dynamically adapt to the available computing resources in response to complexity measurements performed at runtime.
[0004] US 2005 / 0 024 487 A1 discloses video encoding / decoding with real-time complexity adjustment and region-of-interest encoding. In a videoconferencing system where multiple video encoders / decoders (video codecs) operate simultaneously to transmit video, audio, and other data between participants in real time, sharing the system's available resources, each codec is designed for complexity and distortion control and is capable of intelligently finding trade-offs between complexity, data rate, and distortion.
[0005] WO 2008 / 079330 A1 discloses a system for video compression with complexity throttling. Feedback control for a video compression unit is enabled by reading a current compression time used in compression of a frame of a video signal by the video compression unit at a compression complexity value, adjusting the compression complexity value in response to the current compression time, and providing the adjusted compression complexity value to the video compression unit for use in compression of at least one other frame of the video signal.
[0006] US 2015 / 0 046 927 A1 discloses a method for allocating resources of a processor executing a first real-time code component for processing a first sequence of data sections and a second code component for processing a second sequence of data sections, wherein at least the second code component has a configurable complexity.
[0007] US 2015 / 0 326 888 A1 discloses a system for controlling coding time in parallel real-time video coding. A resource controller component dynamically allocates computing resources to an estimation component and an encoding component. The estimation component generates an initial motion estimate of a raw video frame of a sequence of raw video frames based on a previous raw video frame. The encoding component encodes the previous raw video frame to generate a reconstructed video frame in parallel with generating the initial motion estimate.
[0008] US 2010 / 0 104 017 A1 discloses an apparatus for encoding multiple information signals using shared computing power. Multiple encoders are configured to each encode one of the information signals using the shared computing power, wherein each encoder is controllable with respect to its encoding complexity / encoding distortion using at least one corresponding encoding parameter.
[0009] LEE, Younghoon; KIM, Jungsoo; KYUNG, Chong-Min: “Energy-aware video encoding for image quality improvement in battery-operated surveillance camera”, in: IEEE Transactions on Very Large Scale Integration (VLSI) Systems, Volume 20, 2012, No. 2, pages 310-318, ISSN 1063-8210, describes an algorithm for planning video coding configurations in a battery-operated surveillance system to reduce image distortion while ensuring uninterrupted operation.
[0010] KAMINSKY, E.: “Dynamic computational complexity and bit allocation for optimizing H.264 / AVC video compression”, in: Journal of Visual Communication & Image Representation, Volume 19, 2008, pages 56-74, ISSN 1047-3203, discloses an approach to optimizing H.264 / AVC video compression by dynamically allocating computational complexity (e.g., a number of CPU clock cycles) and bits for encoding each coding element (basic unit) in a video sequence according to its predicted mean absolute difference.
[0011] SEMSARZADEH, Mehdi [et al.]: “A fine-grain distortion and complexity aware parameter tuning model for the H.264 / AVC encoder”, in: Signal Processing: Image Communication, Volume 28, 2013, pages 441-457, ISSN 0923-5965, discloses an approach to control the distortion and complexity of an H.264 / AVC encoder by considering a subset of higher-level coding parameters consisting of search range, number of reference frames and motion vector resolution. OVERVIEW
[0012] The present invention is defined by the independent claims. The dependent claims relate to optional features of some embodiments of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Embodiments of the invention are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals designate similar elements; in the drawings: Fig. 1 is a block diagram of a system including a resource allocator according to an embodiment of the invention; Fig. 2 shows a delivery buffer in an encoding / transcoding system according to an embodiment of the invention; Fig. 3 is a representation of the values of ActualVbvDelay, Delta, ActualVbvDelay, and TSMOB, according to an embodiment of the invention; Fig. 4 is a diagram showing an arrangement of cycle profiles 150 according to an embodiment of the invention; and Fig. 5 is a block diagram illustrating a computer system on which an embodiment of the invention may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0014] Approaches for dynamically allocating CPU cycles for use in processing digital video are set forth herein. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a basic understanding of embodiments of the invention described herein. It should be understood, however, that the embodiments of the invention described herein may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form or are discussed at a higher level in order to avoid unnecessarily obscuring the teachings of embodiments of the invention. FUNCTIONAL OVERVIEW
[0015] A digital video encoder is a software-configured hardware component that converts digital video from one format to another. A digital video encoder can support encoding more than one digital video stream simultaneously. An example of a digital video encoder is the Electra 8100, a single-rack unit (1-RU) encoder with multi-standard, multi-service, and multi-channel capabilities available from Harmonic Inc., San Jose, California. The Electra 8100 supports encoding four channels simultaneously per 1-RU.
[0016] A digital video encoder may have multiple central processing units (CPUs or cores). For example, the Electra 8100 encoder includes four CPUs. Software responsible for encoding functionality is typically written to execute on a single CPU. Therefore, four different instances of the encoding software (each individually referred to as an "encoding module") may be executed on the Electra 8100, each configured to execute on a separate CPU. Accordingly, in the prior art, each encoding module is configured to execute instructions using a single CPU.
[0017] Embodiments of the invention enable the cycles of a respective CPU of a digital video encoder to be used more efficiently by software modules executing thereon. Embodiments of the invention optimize the use of computer resources, without significant user intervention and configuration, to adaptively achieve the best video quality for a channel for a given configuration, thereby striving to minimize or eliminate any loss of frames when CPU cycle usage temporarily peaks. If available CPU cycles are abundant, embodiments may allocate additional CPU cycles to an encoding module in addition to those generally allocated by default to support a high-video-quality real-time mode to further increase and enhance video quality.
[0018] Note that while embodiments are described primarily with reference to encoding modules, the techniques discussed herein may also be used in conjunction with other types of video processing software components, such as, but not limited to, digital video transcoder modules responsible for performing transcoding functionality and digital video decoder modules responsible for performing decoding functionality. Indeed, embodiments of the invention may be used with any type of digital video processing software. SYSTEM OVERVIEW
[0019] Fig.1 is a block diagram of a system 100 including a resource allocator 110 according to an embodiment of the invention. Resource allocator 110, as used generally herein, refers to a single software component of two or more cascaded software components responsible for allocating cycle usage from one or more CPUs 130 of a hardware unit 120 to a respective one of a plurality of video processing software modules 140 that perform processing of a digital video. Non-limiting illustrative examples of video processing software modules 140 (or simply video modules 140) include digital video encoder modules, digital video decoder modules, and digital video transcoder modules.
[0020] The hardware unit 120 may correspond to any physical device on which video processing software modules 130 are executed. For example, the hardware unit 120 may correspond to the Electra 8100, which is available from Harmonic Inc., San Jose, California. The hardware unit 120 receives a plurality of incoming channels 160 that separately transport digital video. The video modules 140 executing on the hardware unit perform processing on the incoming channels, such as encoding, decoding, or transcoding the digital video transported by the incoming channels 160. The video modules 140 may support a variety of different video protocols, including, but not limited to, H.264 / AVC, H.265 / HEVC, H.262 / MPEG2, and VP9. The digital video content processed by the video modules 140 is transported by the outgoing data stream 170.
[0021] The resource allocator 110 maintains a set of cycle profiles 150. A respective one of the set of cycle profiles 150 can be assigned by the resource allocator 110 to a respective video module 130 executing on the hardware unit 120. A cycle profile 150 specifies how many cycles of the CPUs 130 are required for a particular video module 140 to perform processing to produce various video quality levels. A cycle profile 150 can also designate or implicitly request a set of hardware and software resources associated with processing a digital video to achieve a specified level of video quality and / or density.The specific cycle profile 150 assigned by the resource allocator 110 to a particular video module 140 may vary over time depending on the content type in a particular service and changes in the amount of available CPU cycle resources. The resource allocator 110 also monitors the amount of available CPU cycle resources of the hardware device on which it is executing. RESOURCE ALLOCATION DEVICE AND CONTROL LOOP
[0022] The resource allocator 110 regulates how many CPU cycle resources are allocated to a particular video module 140 responsible for processing one of the incoming channels 160. In one embodiment, the resource allocator 110 improves the quality of digital video transported by the incoming channels 160 when cycles of the CPUs 130 (i.e., "cycle resources") are made available to a given video module 140. The resource allocator 110 and a respective video module 140 remain in constant communication with each other in a closed-loop control system because both cycle resources and video complexity fluctuate rapidly over time. Accordingly, the resource allocator 110 can make adjustments to how cycle resources are allocated to a respective video module 140 based on information received from the video modules 140.Also, based on information received from resource allocator 110, video modules 140 may make adjustments to how a digital video is processed (e.g., if additional cycle resources are available to a particular video module 140, a digital video may be processed differently by that video module 140 to utilize the additional cycle resources compared to a case where a smaller number of cycle resources are available).
[0023] Each video module 140 can be configured to operate in one of a plurality of different modes. Examples of two such modes are a high-density (HD) mode and a high-quality (VQ) mode. The HD mode may correspond to a configuration that favors using a smaller amount of data (or bits) to represent a frame (full screen) of digital video in order to minimize the bandwidth required to transport the resulting video data stream, and the HQ mode may correspond to a configuration that favors using more data (or bits) to represent a frame of digital video in order to maximize video quality. Other modes may correspond to different preferences for how digital video should be encoded or processed.The mode in which a particular video module 140 operates may be configured by a user using a user interface (UI) presented by the video module 140. In embodiments of the invention, such modes are typically static in that once a user configures a particular video module 140 to operate in a particular mode, that video module 140 will continue to operate in that mode unless reconfigured by a user.
[0024] According to one embodiment, the goal of a resource allocator 110 is to ensure that the video quality of a digital video transported by an outgoing data stream 170 is comparable to, or better than, the video quality conventionally associated with an HD mode. Previously, in the prior art, when a video encoder was instructed to encode video according to a particular mode, there was no deviation or change from the selected mode in the particular manner in which the video encoder operated. In contrast, the resource allocator 110 of one embodiment responds to the availability of cycle resources and accordingly adjusts the operation of video modules 140 based on the available cycle resources.To this end, resource allocator 110 dynamically adjusts the configuration settings of video modules 140 in real time in response to changes in cycle resources. This allows for improved single bit rate (SBR) density without unduly impacting video quality, as the video quality may be comparable to a high video quality (VQ) mode while operating in high density (HD) mode. In embodiments, dynamic adjustment of the operation of a respective video module 140 may occur so that the video module 140 dynamically changes its operating behavior in real time. For example, a specific video module 140 may change its behavior to fluctuate, for example, between operating in an HD mode and a VQ mode.
[0025] In embodiments of the invention, dynamic adjustment of the operation of a video module 140 in response to changes in available cycle resources may be accomplished in a variety of ways. The operation of a video module 140 may be adjusted or controlled by embodiments in its entirety, or solely in how the video module processes a particular digital video frame or macroblock (MB). Changing the behavior on a per-video module 140 basis or on a per-MB basis represents the extreme endpoints on either side of a control, whereas changing the behavior of how the video module 140 processes a single frame of digital video represents a middle range in terms of control and stability.Also, customizing the behavior of how the video module 140 processes a single frame of digital video allows easy control of most settings, which can be adjusted to achieve a trade-off between VQ and cycles that can be done easily and with a granularity that is neither too coarse nor too fine-grained. MEASUREMENTS
[0026] The resource allocator 110 may use one or more metrics to determine how to allocate cycle resources to a respective one of the video modules 140. The resource allocator 110 may use such metrics to adjust the density with which a video module 140 processes a video or to adjust the quality of a digital video processed by a video module 140. The density or quality of a video processed by a video module 140 may be enhanced or increased when additional cycle resources are available to the video module 140, or reduced when additional cycle resources are not available.Note that an improvement in the speed of a video module 140 that would result in a reduction in the density with which a digital video is processed need only be maintained as long as cycle resources are considered to be a problem due to a risk of quality loss in the output data stream of that video module 140.
[0027] Metrics are useful because they enable the earliest possible identification of potential problems. The more balanced a metric is, the more reliable the metric is in terms of control and stability. Transient problems can be identified and addressed by filtering out an underlying metric, although this comes at the expense of reliability. At the same time, a metric used by an embodiment should not be too transient or unreliable. Some metrics that may be used by embodiments to determine how to allocate cycle resources to each of the video modules 140 are discussed below. DELTA MEASUREMENT OF THE TRANSMITTER
[0028] The hardware unit 120 must perform a timely transfer of a processed digital video to an outgoing data stream 170. If a particular video module 140 is overloaded (i.e., the particular video module 140 cannot operate in real time or within the timeframe required by the current workload), then errors occur in that video module 140, such as an underflow in the outgoing data stream 170.
[0029] The sender's delta metric is based on the VBV (Video Buffering Verifier) model. The encoder can use this model to ensure that no underflows or overflows occur at the encoder. The VBV model typically defines the decoder's behavior in terms of three parameters: (1) DTS - the decoding timestamp, (2) PTS - the presentation timestamp, and (3) MaxVbvDelay - the maximum Vbv delay. The multiplexer output buffer fullness (TSMOB) is equal to MaxVbvDelay - ActualVbvDelay + Delta. Ideally, TSMOB should be 0 if there is no processing delay and Delta is 0.
[0030] When an encoder of a particular video module 140 falls behind, an underflow occurs in a buffer (referred to as a "transmitter buffer") physically located on a hardware device 120 that stores frames of digital video to be transported by the outgoing data stream 170. Fig. 2 shows a transmitter buffer 210 in an encoding / transcoding system according to an embodiment of the invention. As in Fig. 2, the transmitter buffer 210 stores content that was output from a multiplexer but not transmitted via the outgoing data stream 170. Thus, if an encoder of a particular video module 140 falls behind, the transmitter buffer 210 will underflow the content generated by that video module 140.
[0031] At the time the hardware unit 120 begins transmitting image (N) to the outgoing data stream 170, the amount of content stored in the sender buffer 210 should be MaxVbvDelay - ActualVbvDelay + Delta, where MaxVbvDelay is the maximum VBV (Video Buffering Verifier) delay, ActualVbvDelay is the actual VBV delay, and Delta is the amount of time we add to the end-to-end delay to neutralize processing time variations in our software system. In a practical implementation, Delta can be exactly or approximately 0.6 seconds. Typically, ActualVbvDelay is less than or equal to 1 second. Fig. 3 is a representation of the values of ActualVbvDelay, Delta, and ActualVbvDelay according to an embodiment of the invention.
[0032] Note that the desired amount of content to be stored in the transmitter buffer 210 is measured in seconds, not in bits. To calculate how much content should be stored in the transmitter buffer 210 in bits for use in a constant bit rate (CBR) mode of operation, the content to be stored in the transmitter buffer 210, as measured over time, is multiplied by the bit rate. For a variable bit rate (VBR) mode of operation, the desired amount of content stored in the transmitter buffer 210 can be calculated as follows: ActualVbvDelay(Image(N)) = DTS(N) - PCR(Start of transmission of Image(N)), where DTS is the decoding timestamp and PCR is the real-time clock. Delta(Start of transmission of image(N)) = TSMOBFilling level(Start image(N)) + DTS(N) - PCR(N) - MaxVbvDelay, where Delta(N) is the actual amount of "extra time" added to the transmission of Image(N) to neutralize processing time variations. The value of Delta(N) is a metric that can be used for the control loop of one embodiment. Typically, the value of Delta is greater than 0.6 seconds in cases where cycles are not an issue. The value of Delta begins to fall below 0.6 seconds in cases where cycle resources are a potential issue for a video module 140. When there is no more content stored in the delivery buffer (i.e., TSMOB = 0), the particular video module 140 stops sending packets. However, this condition is not tied to a specific negative value of Delta, as it is dependent on VBV. LEAKY BUCKET TIME BACK ACCUMULATION
[0033] The leaky bucket time lag metric accumulates the differences between a frame's encoding time relative to the expected frame's encoding time. The expected frame's encoding time is defined as the difference in the decoding time stamp (DTS) between two frames, in ticker units of the 27 MHz clock. The actual encoding time is measured as the encoding time using a precision clock by taking a snapshot at the location (a single point on a circle to measure the round-trip time) to measure the difference in cycles consumed from the time encoding of the last frame begins until the time encoding of the next frame begins or the time encoding of the previous frame ends. CTS(N)=Coding time stamp of image N as snapshot at clock_gettime() DTS(N)=Decoding timestamp of image N DTS(N)=DTS(N−1)+27×106 / picture_rate with DTS(0)=0
[0034] In one embodiment, the variable 'frame rate' (picture_rate) may be defined based on different formats, as shown in Table 1 below.
[0035] The changes of certain attributes can be calculated as follows: Δ DTS(N)=DTS(N)−DTS(N−1) Δ CTS(N)=CTS(N)−CTS(N−1)=encode_time[n] Time backlog ticker units=fb_ticks(N)=∑(ΔCTS(k)−ΔDTS(k)),fu¨rk=0 to N
[0036] To avoid anchoring to previous time lag values, fb-ticks[N] can be calculated by: fb−ticks[N]=(ΔCTS(N)−ΔDTS(N))+(32767*fb_ticks(N−1)) / 32768
[0037] "Time lag" indicates whether a specific video module 140 (e.g., an encoder) operates faster or slower than real time up to the current frame. For a large number of frames, the following applies: Σ fb_ticks[k]→0, fu¨rk=0 up to a large number of images.
[0038] Calculating the value of ΔCTS for a given frame represents the time spent encoding that frame. Note that fb_ticks continuously increase monotonically when a video module 140 is operating slower than real-time (i.e., a lag is accumulating), and catches up when there is no lag. Typically, fb_ticks is less than 0.1 seconds for cases with no lag. For cases with lag, fb_ticks begins to increase to over 0.1 seconds.
[0039] Table 2 elaborates how the leaky bucket lag metric can be used to measure how a particular video module 140 may lag behind its assigned workload. Table 2 All timestamps refer to, or are normalized to, a 27 MHz clock signal. CTS Encoding timestamp - A timestamp for the start of image encoding. This must be recorded at the same location, either before the start of image encoding or before the image bitstream is output. Facebook Time lag in 27 MHz clock ticker units ... ... ... ... ... ... Decoder order N-10 N-9 N-8 N-1 N N+1 DTS(N) DTS((N-10) DTS((N-9) ... ... DTS((N) ... CT(N) CT(N-10) CT(N-9) ... ... ... ... ... ... ... ... CT(N) ... Delta(CT(N)) CT(N-10)-CT(N-11) CT(N-9)-CT(N-10) ... ... CT(N)-CT(N-1) ... Delta(DTS(N)) DTS(N-10) - DTS(N-11) DTS(N-9) - DTS(N-10) ... ... DTS(N) - DTS(N-1) ... FB(N) = FB(N-1) + Delta(CT(N)) - Delta(DTS(N)) FB(N-11) + Delta(CT(N-10)) - Delta (DTS(N-10)) FB(N-10) + Delta(CT(N-9)) - Delta (DTS(N-9)) ... ... FB(N-1) + Delta(CT(N)) - Delta (DTS(N)) ... TIME DELAY COMPENSATED DELTA
[0040] Embodiments of the invention may also use another metric to measure how much a particular video module 140 (which may or may not be an encoder) is falling behind. This metric, referred to as "lag-compensated delta," can be calculated as: Delta_minus_fb_ticks[n] = Delta[n] - fb_ticks[n]. Typically, when the encoder begins to fall behind, fb_ticks and Delta move in opposite directions; however, fb_ticks and Delta do not move in exactly proportional opposite directions. Fb_ticks moves a little earlier than Delta, providing a little more of a lead. Fb_ticks is not tied to recovery in case something goes wrong with the delta recovery mechanism. “Lag-Compensated Delta” provides a compensated measure when Delta begins to decrease due to the encoder’s lag. MAKING ADJUSTMENTS BASED ON AVAILABLE CPU RESOURCES
[0041] In embodiments of the invention, video modules 140 are enabled to make adjustments in how a digital video is processed when the amount of available CPU cycle resources changes. For example, embodiments may support a plurality of different cycle profiles 150. A respective cycle profile may express a different level of quality and density at which a digital video should be processed.
[0042] Cycle profiles 150 may be arranged in a logical sequence based on the consequences of their application to processing of a digital video. Fig. 4 is a diagram illustrating an arrangement of cycle profiles 150 according to an embodiment of the invention. In Fig.4 shows a sequence of numbers ranging from -9 to 9. Each number belongs to a different cycle profile (the exact description of which is not given in Fig. 4). As shown in Fig.As shown in Figure 4, the smallest number in the array (i.e., -9) corresponds to the cycle profile that produces the highest quality and lowest density digital video, while the highest number in the array (i.e., 9) corresponds to the cycle profile that produces the lowest quality and highest density digital video. Each cycle profile, which corresponds to an integer between the two outermost values (-9 and 9), is arranged based on its relative effectiveness. For example, the settings of the cycle profile corresponding to 2, when applied to a particular video module 140, result in a digital video that is of slightly lower quality but higher density than the settings of the cycle profile corresponding to 1 when applied to a particular video module 140.As another example, the settings of the cycle profile associated with -5, when applied to a particular video module 140, result in digital video that is slightly lower quality but higher density than the settings of the cycle profile associated with -6 when applied to a particular video module 140.
[0043] Embodiments of the invention attempt to select a particular cycle profile 150 for use by a particular video module 140 such that it delivers as many CPU cycles as quickly as possible while minimizing the impact on video quality as much as possible. As a result, the best compromise is made for all video modules 140. If the amount of available cycle resources is sufficient, then the quality of a digital video is maximized for all video modules 140 within the cycle budget. However, if the amount of available cycle resources is such that two or more video modules 140 request the same cycle resources to enable these video modules 140 to produce the highest video quality, then the resource allocator 110 selects and allocates the cycle profiles representing the best compromise to maximize the video quality as much as possible for these video modules 140.
[0044] Embodiments also seek to minimize the rate at which video quality degradation occurs, such that the degradation of video quality is not as severe or perceptible to the viewer when tuning or adjusting by one or more incremental steps of the cycle profile for a particular video module 140, e.g., as in Fig. 4, a switch is made from the cycle profile 150 belonging to a 2 to the cycle profile 150 belonging to a 3, and then a switch is made to the cycle profile belonging to a 4.
[0045] Note that although Fig. 4 19 cycle profiles are shown, arranged between -9 and 9, however, the specific number of individual cycle profiles 150 and the identifiers used to designate specific cycle profiles 150 (for example, in Fig.4 numbers in the range -9 to 9 used to denote individual profiles) vary from implementation to implementation.
[0046] To provide a concrete example, certain settings and characteristics of cycle profiles of one embodiment are shown in Table 3 below. The specific cycle profiles shown in Table 3 are arranged from the highest video quality / lowest density (Cycle Profile 1) to the lowest video quality / highest density (Cycle Profile 7).
[0047] As another concrete example, certain cycle profile settings and properties of one embodiment are shown in Table 4 below. The specific cycle profiles shown in Table 4 are arranged from the lowest video quality / highest density (Cycle Profile -1) to the highest video quality / lowest density (Cycle Profile -7). Table 4 Cycle profile Settings Comments -1 Enable pre-processing filters / de-interlacing filters high quality Enables pre-processing filters. Uses higher-quality de-interlacing filters for cases where the source is interlaced and the output stream is progressive. -2 rdo_params.fast_inter = AVC_FastInter_EnhancedHighDensity rdo_params.fast_sb_me= AVC_FastSubBlock_Lev1 Use of the least aggressive EE thresholds (EE = Early Exit) in an inter-prediction, resulting in a smaller number of early exits per image. ME related: Use of more extended MV candidates, e.g., global / regional MV from PA. Use of a fast mid-level subblock ME. -3 rdo_params.fast_inter = AVC_FastInter_Plus MD related: with skip mode, and more exhaustive inter-mode decision. -4 rdo_params.fast_inter = AVC_FastInter Always perform a sub-MB ME and MD for each block. MD related: even more exhaustive inter-mode decision, e.g., ME predictor, no intra_likely flag -5 rdo_prams.rdo_mode = AVC_RDO_Accurate RDO related: with more precise RDO, e.g. "Real Rate" -6 rdo_prams.use_satd_in_intra_md = AVC_SatdIntra_on Enables use of SATD for Intra-MD -7 rdo_prams .use_satd_in_intra_md = AVC_Satd_SubPelME_Ref_on Enables use of SATD for Inter-MD for Reference MB -8 rdo_prams .use_satd_in_intra_md = AVC_Satd_SubPeIME_Ref_on rdo_prams.fast_intra = 0 Enables use of SATD for Inter-MD for all MB, and more exhaustive Intra-MD
[0048] In one embodiment, two or more configuration changes that may be made for a particular video module 140 may be dynamically sequenced based on factors such as image type, resolution, frame rate, temporal hierarchy levels, and prior knowledge of the system, e.g., filter status. To clarify, the current image type may determine the order in which changes are made to a particular video module 140. A particular type of configuration change may be associated with a particular image type for which that configuration change is most effective. Indeed, certain configuration changes may only be effective for certain image types; e.g., changing how motion estimation is performed may not be beneficial for intra-image processing.If the current picture type is P or B, then most configuration changes are beneficial with respect to how encoding is performed by a particular video module 140. However, if the current picture is an I-picture, one embodiment of the invention may change the order in which configuration changes are applied with respect to a P- or B-picture type, such that configuration changes most effective for an I-type picture are applied first to the video module 140 processing the I-type picture.
[0049] In one embodiment, prior knowledge of the current configuration state of a particular video module 140 and the image type that particular video module 140 is currently processing is used to determine how to adjust the configuration of that video module 140 given the available CPU resources. To clarify, if the baseband filters for a particular video module 140 have been disabled, then in one embodiment, no baseband filter is included in any sequence of configuration changes to be communicated to that video module 140 by the resource allocator 110.
[0050] Some types of configuration changes include one or more additional subconfiguration changes. For example, a baseband filter adjustment may be performed by adjusting various baseband filter control values, such as noise reduction and picture enhancement. If any of the configuration settings are not enabled / used by the user, then any of the subconfiguration changes will not be included in any instructions communicated by the resource allocator 110 to the corresponding video module 140. CONTROL LOOP
[0051] When system 100 is not overloaded, embodiments should be able to operate such that the cycle profiles 150 used to configure the operation of each of the video modules 140 enable each of the video modules 140 to provide an optimal video quality mode or enhance the quality of the underlying channels. Furthermore, when system 100 is overloaded, any degradation in video quality should occur as slowly and smoothly as possible.
[0052] The control loop of one embodiment may be based on a metric such as the transmitter delta metric, the leaky bucket backlog accumulation metric, or the compensated backlog delta metric. The fullness of the transmitter buffer 210 is measured by TSMOB. When the transmitter buffer 210 runs dry (i.e., TSMOB = 0), there are no packets stored in the transmit buffer 210 to transmit, causing an underflow in the time domain. One of the reasons why a run-dry of the transmitter buffer 210 might occur is due to an encoder becoming behind schedule and no longer being able to send packets to the transmitter buffer 210.
[0053] The response time to an underflow in the transmitter buffer 210 is controlled by a threshold. Typically, delta (the amount of time to be added to the end-to-end delay to neutralize processing time variations in software systems) is initialized to a value of approximately 0.6 seconds. For example, if the threshold is set to 0.5 seconds, the minimum response time to control the underflow in the transmitter buffer 210 is 0.5 seconds. The transmitter buffer 210 determines the availability of bits that the multiplexer can pack into the transport data stream. The fill level of the transmitter buffer 210 determines the amount of time required to fill the transmitter buffer 210 if it begins to approach an empty state, i.e., 0 bits.Accordingly, in this example, a maximum of 0.5 seconds is available from the time the transmitter buffer 210 was full to the time the transmitter buffer 210 becomes empty to ensure that the transmitter buffer 210 remains at an optimal level without becoming empty.
[0054] The fb_ticks are a measure of how far behind a particular video module 140 is in processing. Similar to Delta, for example, if the fb_ticks threshold is set to 0.1 seconds, an attempt is made to keep the amount of video data for which a particular video module 140 is behind in processing below 0.1 seconds. Both of these threshold specification options are limited by the degree of control provided by the cycle profiles 150 for a given situation.
[0055] In one embodiment of the invention, a control loop may be used that uses fb_ticks or a time lag as process variables, while driving the time lag setpoint toward 0. A similar approach may be used by an embodiment in which the Delta or Delta_minus_fbticks signals are used as process variables, and the setpoints for Delta and Delta_minus_fbticks are set to 0.6 seconds. CENTRAL CONTROL DEVICE
[0056] When a particular video module 140 falls behind in processing a channel, one of the potential problems is that some channels occasionally fall behind more than others. The "lead channels" typically do this because the content they carry is suited to making a larger number of computationally scalable decisions in the pipeline.
[0057] On the other hand, the channels that fall behind typically carry relatively difficult content that requires a larger number of serial coding decisions (e.g., a large number of smaller intra-modes and / or a larger number of CABAC binary decisions for renormalization) to be made within a cycle budget that cannot be extended beyond the value allocated for that particular channel, and a larger number of cycle resources cannot be acquired without obtaining them from the other channels sharing the same set of CPU cores.
[0058] One way to allocate more cycle resources to any channel that is falling behind more than the others can be done by using a central entity that can normalize the time lag of all channels and bring them to the same level of time lag, thereby freeing cycles from the faster channels in favor of the slower channels. The slower channels use the normalized value sent by the central controller to free cycles. The resource allocator 110 assigns a particular video module 140 processing a channel that is falling behind a different cycle profile 150 that is one step or increment toward higher density and lower video quality. For example, if a particular video module 140 previously used a cycle profile associated with 1, as in Fig.4, then the resource allocation device 110, upon receiving the information from the central control device, can instruct this particular video module 140 to use the cycle profile belonging to a 2, as in Fig. 4, since this cycle profile is a single step or increment in the direction of higher density and lower video quality. However, it should be noted that channels that run faster when processing less computationally intensive content are disproportionately affected by a reduction in video quality caused by adjusting the cycle profiles used by the video modules 140 processing these channels to have very low video quality and higher density to ensure that the content is processed in a timely manner.
[0059] The following set of equations describes the central controller. For each channel involved in sharing a set of cores, the average time lag value can be calculated as: Fbavg=Σ Fb[i] / N for ri=0 to N channels
[0060] The normalized feedback value is applied to each channel by the central controller. The lag of other channels can then be adjusted to allow the channels to lag further, favoring the channel(s) lagging the most. This can be expressed as: If (Fb < Fbavg) Set Fb = Fbavg or Fb = (1-bias) * Fb + bias * Fbavg For use by the control loop, where bias is: 0 < bias < 1
[0061] It should be noted that adjusting the permissible time lag values in this way will impair the video quality of channels that carry content that is less computationally expensive to process and that run faster than real time or close to real time. STATISTICAL MULTIPLEXING OF CYCLES
[0062] In U.S. patent application Ser. No. 14 / 961,239, filed December 7, 2015, granted as US 10104405 B1, entitled "DYNAMIC ALLOCATION OF CPU CYCLES IN VIDEO STREAM PROCESSING," a CPU stream load balancing module is responsible for statistically multiplexing cycles across encoders using their complexity and CPU utilization. Embodiments of the invention recognize that as the complexity of the content increases, the CPU utilization required to process that content at that time increases. For statistically multiplexed channels, typically the higher-complexity channels require a larger bit rate as well as greater CPU utilization.
[0063] In the following expressions, the complexity of a channel at a time t is denoted by X i,t where i is the channel number and t is the time. The CPU utilization at a time t is denoted as C i,twhere i is the channel number and t is the time. The CPU utilization of a given channel is proportional to its complexity at that time, as shown by: Ci,t α Xi,t
[0064] Typically, we modulate the instantaneous ratio of total complexity to CPU usage to track the proportional variable βt at time t. The total complexity at time t is defined as Xtotal,t=Σ Xi,t
[0065] The total CPU usage at time t is defined as: Ctotal,t=Σ Ci,t
[0066] The proportional CPU complexity coefficient variable for modulating the CPU usage prediction βt is calculated as: βt=Ctotal,t / Xtotal,t
[0067] The proportional CPU complexity coefficient variable is filtered with respect to previous values using IIR filtering, as follows: βfilt,t=(βt+7*βt−1) / 8
[0068] Idle CPU usage is defined as Idletotal,t=100−Ctotal,t
[0069] At time t, each encoder (or video module 140) sends the look-ahead complexity value LAX i,t for the first frame in its look-ahead pipeline. Typically, the complexity of the reported frame corresponds to several frames in the future, which are compared to the one currently being encoded. Each encoder (or video module 140) also sends the current instantaneous CPU usage C i,t and the complexity of the currently encoded image X i,t to the CPU stream load balancing module. Based on this information, the CPU stream load balancing module makes a prediction of CPU usage according to the look-ahead complexity value LAX i,t for a respective coding channel. The total look-ahead complexity value at time t is defined as LAXtotal,t=∑Xi,t
[0070] The total CPU usage for the look-ahead images is calculated as CPREDtotal,t=βfilt,t*LAXtotal,t
[0071] With this information a determination can be made for IdlePred total,t = 100 - CPREDtotal,t. Note that this value could be either positive or negative, since a negative value indicates future complexity values that are likely to cause an overload of the system CPU utilization by all encoders (or video modules 14) as a whole. This value can be used by the control loop of one embodiment, as follows:
[0072] This idleCpuPerc is used by the control loop to make adjustments to the cycle profile 150 assigned to an encoder (or video module 150). This provides the CPU utilization-based feedforward prediction mechanism to the control loop.
[0073] Also, in another embodiment, the CPU stream load balancing module performs a prediction of the time lag of individual encoder channels based on the look-ahead complexity value LAX i,t for each encoder channel. A measure of the extent of a time lag of each channel at a time t is denoted by FB i,t where i is the channel number and t is the time point. A measure of the time lag of a given channel is proportional to its complexity value at that time point, as expressed by: FBi,t α Xi,t
[0074] The total time lag at time t is defined as FBtotal,t=∑ FBi,t
[0075] A proportional time lag complexity coefficient variable to adjust the time lag prediction value Δt can be calculated as: Δt=FBtotal,t / Xtotal,t
[0076] The proportional time lag complexity coefficient variable is filtered with respect to previous values using IIR filtering, as follows: Δfilt,t=(Δt+7*Δt−1) / 8
[0077] At time t, each encoder (or video module 140) sends the look-ahead complexity value LAX i,t for the first frame in its look-ahead pipeline to the CPU stream load balancing module. Typically, the complexity of the reported frame corresponds to several frames in the future, which are compared to the one currently being encoded. Each encoder (or video module 140) also sends the current instantaneous lag FB i,t and the complexity of the currently encoded image X i,t to the CPU stream load balancing module. Based on this information, the CPU stream load balancing module makes a prediction of the backlog according to the look-ahead complexity value LAX i,tfor each coding channel as follows: FBPREDtotal,t=Δfilt,t*LAXtotal,t
[0078] The time lag for a particular coding channel can be predicted as: FBPREDi,t=FBPREDtotal,t*LAXi,t / LAXtotal,t
[0079] The time lag for a particular coding channel can be used by the control loop as a feedforward prediction. This allows the expression Δ filt,t = (Δt + 7 * Δ t-1 ) / 8 reformulate to:
[0080] The above procedure improves the stability of the control loop, and it does not wait for a deviation from the setpoint before taking corrective action. This is because the disturbance is predicted before it enters the system, and this information is used to take corrective action before the disturbance has affected the system. The effect of the disturbance is thus reduced by predicting it and generating a control signal to counteract it before its impact on the system becomes noticeable. LOAD DISTRIBUTION OF SOCKET AND THREAD
[0081] Compute units are typically grouped into groups of CPUs (sockets) with access to the same memory resources. A transcoder instance may be constrained to assign its processing threads to cores of a specific group of CPUs (sockets) to reduce potential memory copy operations across CPU groups. While such a constraint is more efficient, distributing a number of transcoder instances (of different resolution, configuration, and content) among a number of CPU groups requires load balancing during runtime to avoid endemic cases where a single CPU group (socket) is overloaded all the time while the rest are underutilized. This allows, in one embodiment, adaptation of CPU groups so that, to the best of its ability, each CPU group uses the same or a similar cycle profile 150.
[0082] After an initial allocation of transcoder instances is performed during system commissioning, the following procedure may be performed by embodiments to periodically rebalance the channels during runtime.
[0083] First, a measurement of the average identifier value of the cycle profile 150 (ie one of the Fig. 4) for a respective channel of a respective socket over a 5-minute window. LEVEL [ii] [jj] expresses the average cycle profile characteristic for channel jj, restricted to socket ii.
[0084] Next, CPU groups are arranged in descending order of their avgLEVEL[ii] (average cycle profile identifier for channels, restricted to socket ii), where avgLEVEL[0] > avgLEVEL[1] > ... > avgLEVEL[n-1], where n is the number of distinct CPU groups.
[0085] Then, the channel in each CPU group [ii] that has the lowest average tuning level during the previous 5-minute window is identified as Lowest [ii]. The channel in each CPU group [ii] that has the highest average tuning level during the previous 5-minute window is identified as Highest [ii]. The number of encoding / transcoding services running on a respective CPU group is identified as numServices [ii].
[0086] Finally, a new load balancing is described in the form of pseudo-code according to one embodiment: for (int ii=0; i<(n / 2); i++) { if (((avgLEVEL[ii] - avgLEVEL[n-1-ii]) > 2) && (numServices[ii] <= numServices[n-1-ii])) { Switch Lowest[n-1-ii] to CPU group (ii); Switch Highest[ii] to CPU group (n-1-ii);} if (((avgLEVEL[ii] - avgLEVEL[n-1-ii]) > 2) && (numServices[ii] > numServices[n-1-ii]) ) { Switch Lowest[ii] to CPU group (n-1-ii);}}
[0087] As shown above, video modules 140 can be limited by grouping or physical resource allocation (i.e., how CPUs are arranged in a socket), which limits their ability to freely access resources across sockets. Accordingly, until more resources can be allocated to video module 140 outside of that socket, video module 140 requires its assigned cycling profile 150, which governs its operation, to be adjusted to accommodate the resources available to the socket until the resources become available.
[0088] Threads are another example of resources where load balancing must be performed, namely between overallocation and associated overhead. Each cycle profile 150 requires a different allocation of threads to achieve the processing configuration defined by that cycle profile. Each cycle profile 150 specifies configuration changes for a video module 150 that can be used to adjust a variety of factors, such as bit rate, frame rate, resolution, and codec. Accordingly, embodiments of the invention also ensure that the resource allocator 110 takes thread availability, overallocation, and associated overhead into account when adjusting which cycle profile 150 to assign to a particular video module 140. HARDWARE MECHANISMS
[0089] In one embodiment, the hardware unit 120 may be Fig. 1 be implemented on a computer system or correspond to it. Fig.5 is a block diagram illustrating a computer system 500 on which an embodiment of the invention may be implemented. In one embodiment, the computer system 500 includes a processor 504, a main memory 506, a read-only memory 508, a storage device 510, and a communications interface 518. The computer system 500 includes at least one processor 504 for processing information. The computer system 500 also includes a main memory 506, such as a RAM (random access memory) or other dynamic storage device, for storing information and instructions to be executed by the processor 504. The main memory 506 may also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by the processor 504.The computer system 500 includes a ROM (Read Only Memory) 508 or other static storage device for storing static information and instructions for the processor 504. A storage device 510, such as a magnetic disk or an optical disk, is provided for storing information and instructions.
[0090] Embodiments of the invention relate to the use of a computer system 500 for implementing the methods described herein. According to one embodiment of the invention, these methods are performed by the computer system 500 in response to the processor 504 executing one or more sequences of one or more instructions located in main memory 506. Such instructions may be read into the main memory 506 from another machine-readable medium, such as the storage device 510. Executing the sequences of instructions located in the main memory 506 causes the processor 504 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used instead of, or in combination with, software instructions to implement embodiments of the invention.Thus, embodiments of the invention are not limited to any specific combination of hardware, circuitry, and software.
[0091] The term "non-transitory machine-readable storage medium," as used herein, refers to any tangible medium that participates in the persistent storage of instructions that can be provided to processor 504 for execution. Non-limiting, illustrative examples of non-transitory computer-readable media include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape or any other magnetic medium, a CD-ROM, any other optical medium, a RAM, a PROM, an EPROM, a FLASH EPROM, any other memory chip or cartridge, or any other medium from which a computer can read.
[0092] Various forms of non-transitory computer-readable media may be involved in transporting one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially reside on a magnetic disk of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions to computer system 500 over a network connection 520.
[0093] Communications interface 518 provides a two-way data communications connection to a network interface 520 connected to a local area network. For example, communications interface 518 may be an Integrated Services Digital Network (ISDN) card or a modem to provide a data communications connection to a corresponding type of telephone line. As another example, communications interface 518 may be a Local Area Network (LAN) card to provide a data communications connection to a compatible LAN. Radio links may also be implemented. In any such implementation, communications interface 518 performs transmission and reception of electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0094] Network connection 520 typically provides data communication over one or more networks to other data devices. For example, network connection 520 may provide a connection over a local area network to a host computer or data facilities operated by an Internet service provider (ISP).
[0095] Computer system 500 may send messages and receive data, including program code, over the network(s), network connection 520, and communications interface 518. For example, a server could transmit requested code for an application program over the Internet, a local ISP, a local network, and then to communications interface 518. The received code may be executed by processor 504 upon receipt and / or stored in a memory device 510 or other non-volatile memory for later execution.
[0096] In the foregoing description, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation.
Claims
[1] A non-transitory computer-readable storage medium storing one or more sequences of instructions for dynamically allocating, to a video encoder, CPU cycle resources that, when executed by one or more processors (504), cause: Estimating, by a resource allocation device (110) executing on a hardware unit (120), an amount of available CPU cycle resources of two or more groups of central processing units, CPUs (130), in the hardware unit (120); and after determining, by the resource allocation device (110), a change in the amount of available CPU cycle resources, adjusting, in real time, which particular cycle profile (150) of a plurality of cycle profiles (150) is assigned to the at least one of a plurality of video modules (140), wherein the plurality of cycle profiles (150) each allocate to elements of the plurality of video modules (140) a specific amount of CPU cycle resources for processing a digital video, wherein group members of the same group of the two or more groups of CPUs (130) have access to the same set of memory resources, and wherein the plurality of cycle profiles (150) each specify a set of configuration settings and are arranged in a sequence based on a video quality and density that can be achieved by elements of the plurality of video modules (140) using configuration settings associated with a respective one of the plurality of cycle profiles (150) when processing a digital video; and in response to assigning a different cycle profile (150) to a particular video module (140), rebalancing, during runtime, a set of channels assigned to each group of the two or more groups of CPUs (130) such that, to the extent possible, each member of each group of the two or more groups of CPUs (130) serves video modules (140) assigned to the same or a similar cycle profile (150) of the plurality of cycle profiles (150). [2] The non-transitory computer-readable storage medium of claim 1, wherein executing the one or more sequences of instructions further causes: responsive to adjusting, in real time, the particular cycle profile (150) assigned to the at least one video module (140), processing, by the at least one video module (140), a particular digital video frame in a video data stream using a different set of configuration settings than for adjacent frames. [3] The non-transitory computer-readable storage medium of claim 1, wherein executing the one or more sequences of instructions further causes: responsive to adjusting, in real time, the particular cycle profile (150) assigned to the at least one video module (140), processing, by the at least one video module (140), a particular macroblock using a different set of configuration settings than for at least one other macroblock in the same frame of digital video. [4] The non-transitory computer-readable storage medium of claim 1, wherein the special cycle profile (150) assigned to the at least one video module (140) allocates further CPU cycle resources to the at least one video module (140) in addition to those assigned by default to a mode in which the at least one video module (140) is operating, for the purpose of improving video quality beyond the quality associated with that mode. [5] The non-transitory computer-readable storage medium of claim 1, wherein executing the one or more sequences of instructions further causes: Instructing, by the resource allocator (110), a particular video module (140) of the plurality of video modules (140) to change its behavior to fluctuate between operation according to a high quality, HQ mode, and a high density, HD mode, in response to changes in available CPU cycle resources. [6] The non-transitory computer-readable storage medium of claim 1, wherein executing the one or more sequences of instructions further causes: Instructing, by the resource allocator (110), a dedicated video module (140) of the plurality of video modules (140) to change a video quality level in the processed digital video generated by the dedicated video module (140) in response to changes in available CPU cycle resources. [7] The non-transitory computer-readable storage medium of claim 1, wherein adjusting, in real time, which particular cycle profile (150) is assigned to the at least one of a plurality of video modules (140) includes: Identifying, by the resource allocator (110), that a new cycle profile (150) should be assigned to the at least one of the plurality of video modules (140) in response to detecting an underflow condition in a transmitter buffer (210) of the at least one of the plurality of video modules (140). [8] The non-transitory computer-readable storage medium of claim 1, wherein adjusting, in real time, which particular cycle profile (150) is assigned to the at least one of a plurality of video modules (140) includes: Identifying, by the resource allocator (110), that the at least one of the plurality of video modules (140) should be assigned a new cycle profile (150) in response to measuring differences between a coding time per picture relative to an expected coding time per picture for a respective one of the at least one of the plurality of video modules (140). [9] Apparatus for dynamically allocating CPU cycle resources to a video encoder, comprising: one or more processors (504); and one or more non-transitory computer-readable storage media (506, 508, 510) storing one or more sequences of instructions which, when executed, cause: Estimating, by a resource allocation device (110) executing on a hardware unit (120), an amount of available CPU cycle resources of two or more groups of central processing units, CPUs (130), in the hardware unit (120); and after determining, by the resource allocation device (110), a change in the amount of available CPU cycle resources, adjusting, in real time, which particular cycle profile (150) of a plurality of cycle profiles (150) is assigned to the at least one of a plurality of video modules (140), wherein the plurality of cycle profiles (150) each allocate to elements of the plurality of video modules (140) a specific amount of CPU cycle resources for processing a digital video, wherein group members of the same group of the two or more groups of CPUs (130) have access to the same set of memory resources, and wherein the plurality of cycle profiles (150) each specify a set of configuration settings and are arranged in a sequence based on a video quality and density that can be achieved by elements of the plurality of video modules (140) using configuration settings associated with a respective one of the plurality of cycle profiles (150) when processing a digital video; and in response to assigning a different cycle profile (150) to a particular video module (140), rebalancing, during runtime, a set of channels assigned to each group of the two or more groups of CPUs (130) such that, to the extent possible, each member of each group of the two or more groups of CPUs (130) serves video modules (140) assigned to the same or a similar cycle profile (150) of the plurality of cycle profiles (150). [10] The apparatus of claim 9, wherein executing the one or more sequences of instructions further causes: responsive to adjusting, in real time, the particular cycle profile (150) assigned to the at least one video module (140), processing, by the at least one video module (140), a particular digital video frame in a video data stream using a different set of configuration settings than for adjacent frames. [11] The apparatus of claim 9, wherein executing the one or more sequences of instructions further causes: responsive to adjusting, in real time, the particular cycle profile (150) assigned to the at least one video module (140), processing, by the at least one video module (140), a particular macroblock using a different set of configuration settings than for at least one other macroblock in the same frame of digital video. [12] The apparatus of claim 9, wherein the special cycle profile (150) assigned to the at least one video module (140) allocates further CPU cycle resources to the at least one video module (140) in addition to those assigned by default to a mode in which the at least one video module (140) is operating, for the purpose of improving video quality beyond the quality associated with that mode. [13] The apparatus of claim 9, wherein executing the one or more sequences of instructions further causes: Instructing, by the resource allocator (110), a particular video module (140) of the plurality of video modules (140) to change its behavior to fluctuate between operation according to a high quality, HQ mode, and a high density, HD mode, in response to changes in available CPU cycle resources. [14] The apparatus of claim 9, wherein executing the one or more sequences of instructions further causes: Instructing, by the resource allocator (110), a dedicated video module (140) of the plurality of video modules (140) to change a video quality level in the processed digital video generated by the dedicated video module (140) in response to changes in available CPU cycle resources. [15] The apparatus of claim 9, wherein adjusting, in real time, which particular cycle profile (150) is assigned to the at least one of a plurality of video modules (140) includes: Identifying, by the resource allocator (110), that a new cycle profile (150) should be assigned to the at least one of the plurality of video modules (140) in response to detecting an underflow condition in a transmitter buffer (210) of the at least one of the plurality of video modules (140). [16] The apparatus of claim 9, wherein adjusting, in real time, which particular cycle profile (150) is assigned to the at least one of a plurality of video modules (140) includes: Identifying, by the resource allocator (110), that the at least one of the plurality of video modules (140) should be assigned a new cycle profile (150) in response to measuring differences between a coding time per picture relative to an expected coding time per picture for a respective one of the at least one of the plurality of video modules (140). [17] A method for dynamically allocating CPU cycle resources to a video encoder, comprising: Estimating, by a resource allocation device (110) executing on a hardware unit (120), an amount of available CPU cycle resources of two or more groups of central processing units, CPUs (130), in the hardware unit (120); and after determining, by the resource allocation device (110), a change in the amount of available CPU cycle resources, adjusting, in real time, which particular cycle profile (150) of a plurality of cycle profiles (150) is assigned to the at least one of a plurality of video modules (140), wherein the plurality of cycle profiles (150) each allocate to elements of the plurality of video modules (140) a specific amount of CPU cycle resources for processing a digital video, wherein group members of the same group of the two or more groups of CPUs (130) have access to the same set of memory resources, and wherein the plurality of cycle profiles (150) each specify a set of configuration settings and are arranged in a sequence based on a video quality and density that can be achieved by elements of the plurality of video modules (140) using configuration settings associated with a respective one of the plurality of cycle profiles (150) when processing a digital video; and in response to assigning a different cycle profile (150) to a particular video module (140), rebalancing, during runtime, a set of channels assigned to each group of the two or more groups of CPUs (130) such that, to the extent possible, each member of each group of the two or more groups of CPUs (130) serves video modules (140) assigned to the same or a similar cycle profile (150) of the plurality of cycle profiles (150). [18] The method of claim 17, wherein the special cycle profile (150) assigned to the at least one video module (140) allocates further CPU cycle resources to the at least one video module (140) in addition to those assigned by default to a mode in which the at least one video module (140) is operating, for the purpose of improving video quality beyond the quality associated with that mode. [19] The method of claim 17, further comprising: Instructing, by the resource allocator (110), a particular video module (140) of the plurality of video modules (140) to change its behavior to fluctuate between operation according to a high quality, HQ mode, and a high density, HD mode, in response to changes in available CPU cycle resources. [20] The method of claim 17, further comprising: Instructing, by the resource allocator (110), a dedicated video module (140) of the plurality of video modules (140) to change a video quality level in the processed digital video generated by the dedicated video module (140) in response to changes in available CPU cycle resources.
Citation Information
Patent Citations
Video codec system with real-time complexity adaptation
EP1503595A2
Video codec system with real-time complexity adaptation and region-of-interest coding
US20050024487A1
Encoding of a Plurality of Information Signals Using a Joint Computing Power
US20100104017A1
Allocating Processor Resources
US20150046927A1
Encoding time management in parallel real-time video encoding
US20150326888A1