Method and system of optimized video encoding
The method and system for optimized video encoding stabilize encoding configurations to balance quality and bitrate, addressing inconsistent frame quality and efficiency issues, achieving consistent frame quality and improved coding efficiency using advanced codecs and hardware-accelerated encoders.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEAMR IMAGING LTD
- Filing Date
- 2025-09-29
- Publication Date
- 2026-04-23
AI Technical Summary
Existing video encoding technologies struggle to efficiently balance quality and bitrate, leading to inconsistent frame quality and reduced coding efficiency due to inconsistent encoding configurations across frames.
A method and system for optimized video encoding that determines a set of candidate encoding configurations based on frame information, stabilizes encoding quality and efficiency by applying stability rules, and iteratively evaluates and selects the best configuration to meet target criteria, using hardware-accelerated encoders like NVENC and advanced codecs like AVI.
Achieves consistent frame quality and improved coding efficiency by stabilizing encoding configurations, reducing temporal flickering and bitrate fluctuations, while supporting modern codecs and enabling live encoding with low latency.
Smart Images

Figure IL2025050863_23042026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND SYSTEM OF OPTIMIZED VIDEO ENCODING
[0002] TECHNICAL FIELD
[0003] The presently disclosed subject matter relates, in general, to the field of video encoding systems, and, more particularly, to video encoding configuration and optimization driven by quality and bitrate considerations.
[0004] BACKGROUND
[0005] Past years have seen a vast increase in the amount of video content being created, to be stored, processed, or sent over the network. The creators of such video content range from premium production studios to User Generated Content (UGC), with many new sources, such as videos created with Generative Al, cameras of autonomous cars, and Internet of Things (loT), to mention but a few. This comes with simultaneous increases in video resolutions, frame rates, and bit depths, meaning new approaches to optimizing the video compression or encoding processes are more relevant than ever.
[0006] In order to efficiently tackle video compression workflows in the face of this video growth, there are several pillars to build on. The first is using more advanced encoders, such as the AVI encoder developed by the Alliance for Open Media (AOM). Another pillar is to use faster / lower cost encoding platforms, for instance using hardware-accelerated encoding such as that offered by NVENC (i.e., NVIDIA Encoder, a dedicated hardware video encoder built into NVIDIA graphics cards that offloads the resource-intensive video encoding process from the central processing unit (CPU) to the graphics processing units (GPU)), which encodes on NVIDIA GPUs. Yet another pillar is configuring the encoder to make optimal decisions to create encoders that pack the highest quality into the lowest possible bitrate.
[0007] GENERAL DESCRIPTION
[0008] In accordance with certain aspects of the presently disclosed subject matter, there is provided a computerized method of optimized video encoding of an input video sequence comprising a plurality of input frames, the method comprising: obtaining an input frame of the plurality of input frames and frame information thereof; determining a set of candidate encoding configurations usable for encoding the input frame based on the frame information, the set of candidate encoding configurations selected for enabling stabilization of encoding quality and efficiency with respect to the input frame; encoding the input frame using a first candidate encoding configuration selected from the set of candidate encoding configurations, to obtain a candidate encoded frame; evaluating the candidate encoded frame with respect to a target criterion, and verifying if a termination criterion is met; if not, repeating the encoding, evaluating and verifying using a next candidate encoding configuration selected from the set, until the termination criterion is met; and selecting, as an output frame, a candidate encoded frame best meeting the target criterion.
[0009] In addition to the above features, the method according to this aspect of the presently disclosed subject matter can comprise one or more of features (i) to (xii) listed below, in any desired combination or permutation which is technically possible:
[0010] (i). The method further comprises adding a selected encoding configuration corresponding to the output frame to a history of selected encoding configurations, giving rise to an updated history of selected encoding configurations.
[0011] (ii). The determination is in accordance with a set of stability rules and using the history of selected encoding configurations.
[0012] (iii). The set of stability rules comprises a beyond-the-frame stability rule for selecting the set of candidate encoding configurations to be consistent with one or more selected encoding configurations for one or more previously encoded frames.
[0013] (iv). The set of candidate encoding configurations includes a set of candidate
[0014] Quantization Parameter (QP) values, and the history of selected encoding configurations includes one or more previous QP values selected for one or more previously encoded input frames. (v). The beyond-the-frame stability rule stipulates that the set of candidate QP values does not deviate from the previous QP values by more than a predetermined interval.
[0015] (vi). The target criterion is based on a bit consumption metric and / or a quality measure.
[0016] (vii). The quality measure is a per frame Video Multimethod Assessment Fusion
[0017] (VMAF) value.
[0018] (viii). The set of stability rules overrides the target criterion, and an encoded frame substantially meeting the target criterion is provided as the output frame.
[0019] (ix). The frame information comprises one or more of frame type, bit consumption, metadata associated with the input frame, indication of scene change, average level of brightness, frame dimensions, and detected objects.
[0020] (x). The input video sequence corresponds to a first codec, and the encoding of the input frame is performed for transcoding to a second codec. The set of candidate encoding configurations is determined so as to not reduce quality of the input video sequence, while benefiting from properties offered by the second codec.
[0021] (xi). The first codec is a legacy codec, the second codec is a modern codec, and the properties are related to compression efficiency.
[0022] (xii). The termination criterion includes at least one of: the target criterion is met, a predetermined number of iterations is reached, and candidate encoding configuration options are exhausted.
[0023] In accordance with other aspects of the presently disclosed subject matter, there is provided a computerized system of optimized video encoding of an input video sequence comprising a plurality of input frames, the system comprising a processing circuitry configured to perform method steps of any of the above methods.
[0024] In accordance with other aspects of the presently disclosed subject matter, there is provided a non-transitory computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method steps of any of the above methods.
[0025] BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to understand the presently disclosed subject matter and to see how it may be carried out in practice, the subject matter will now be described, by way of nonlimiting example only, with reference to the accompanying drawings, in which:
[0027] FIG. 1A and IB are functional block diagrams schematically illustrating a system of optimized video encoding in accordance with certain embodiments of the presently disclosed subject matter;
[0028] FIG. 2 is a generalized flowchart showing an example of optimized video encoding in accordance with certain embodiments of the presently disclosed subject matter;
[0029] FIG. 3 is a functional block diagram schematically illustrating a stabilizer used for optimized video encoding in accordance with certain embodiments of the presently disclosed subject matter; and
[0030] FIG. 4 is a generalized flowchart showing an example of a stabilizer used for optimized video encoding in accordance with certain embodiments of the presently disclosed subject matter.
[0031] DETAILED DESCRIPTION
[0032] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed subject matter. However, it will be understood by those skilled in the art that the present disclosed subject matter can be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the present disclosed subject matter.
[0033] In the drawings and descriptions set forth, identical reference numerals indicate those components that are common to different embodiments or configurations.
[0034] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the specification discussions utilizing terms such as "obtaining", "determining", "encoding", "evaluating", "verifying", "repeating", "selecting", "using", "adding", "transcoding", or the like, include action and / or processes of a computer that manipulate and / or transform data into other data, said data represented as physical quantities, such as electronic quantities, and / or said data representing the physical objects.
[0035] The terms "computer", "computer-based system", or "computerized system" should be expansively construed to cover any kind of hardware-based electronic device with a data processing circuitry (e.g., digital signal processor (DSP), a graphics processing unit (GPU), a field programmable gate array (FPGA), including, by way of non-limiting example, the video encoding system, and respective parts thereof disclosed in the present application. The data processing circuitry (designated also as processing circuitry) can comprise, for example, one or more processors operatively connected to computer memory, loaded with executable instructions for executing operations, as further described below. The data processing circuitry encompasses a single processor or multiple processors, which may be located in the same geographical zone, or may, at least partially, be located in different zones, and may be able to communicate with each other.
[0036] The one or more processors referred to herein can represent one or more general-purpose processing devices such as a microprocessor, a central processing unit, or the like. More particularly, a given processor may be one of a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. The one or more processors may also be one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The one or more processors are configured to execute instructions for performing the operations and steps discussed herein.
[0037] The memories referred to herein can comprise one or more of the following: internal memory, such as, e.g., processor registers and cache, etc., main memory such as, e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc. The operations in accordance with the teachings herein can be performed by a computer specially constructed for the desired purposes or by a general-purpose computer specially configured for the desired purpose by a computer program stored in a non-transitory computer-readable storage medium.
[0038] The terms "non-transitory memory", "non-transitory storage medium", and "non-transitory computer-readable storage medium" used herein should be expansively construed to cover any volatile or non-volatile computer memory suitable to the presently disclosed subject matter.
[0039] Embodiments of the presently disclosed subject matter are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the presently disclosed subject matter as described herein.
[0040] As used herein, the phrase "for example," "such as", "for instance", and variants thereof describe non-limiting embodiments of the presently disclosed subject matter. Reference in the specification to "one case", "some cases", "other cases", or variants thereof means that a particular feature, structure, or characteristic described in connection with the embodiment(s) is included in at least one embodiment of the presently disclosed subject matter. Thus, the appearance of the phrase "one case", "some cases", "other cases", or variants thereof does not necessarily refer to the same embodiment(s).
[0041] It is appreciated that, unless specifically stated otherwise, certain features of the presently disclosed subject matter, which are described in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features of the presently disclosed subject matter, which are described in the context of a single embodiment, can also be provided separately or in any suitable subcombination. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the methods and apparatus.
[0042] In embodiments of the presently disclosed subject matter, one or more stages illustrated in the figures may be executed in a different order and / or one or more groups of stages may be executed simultaneously, and vice versa. Bearing this in mind, attention is now drawn to FIG. 1, schematically illustrating a system of optimized video encoding in accordance with certain embodiments of the presently disclosed subject matter.
[0043] According to certain embodiments, there is provided a computer-based system 100 for optimized video encoding of an input video sequence, the input video sequence corresponding to a plurality of input video frames (also termed as input frames or video frames). The term "video encoding" used in this patent specification should be expansively construed to cover any kind of video compression that converts raw (i.e., uncompressed) digital video to a compressed format, as well as video recompression that converts decoded or decompressed video to a re-encoded or recompressed format.
[0044] In certain embodiments, the input video sequence can refer to an original video sequence that is not encoded or compressed, such as, e.g., an original raw video clip or part thereof. Such video clips can comprise a plurality of original video frames, and can be obtained from, e.g., a digital camera or recorder, or any other suitable devices that are capable of capturing or recording individual still images or sequences of images constituting videos or movies.
[0045] In some other embodiments, the input video sequence can refer to an input video bit-stream which has been previously encoded using a video encoder (thus it can also be referred to as "encoded video bit-stream" or "compressed video bit-stream") and includes encoded data corresponding to one or more encoded video frames. In such cases, the input video bit-stream can be first decoded or reconstructed to a decoded video sequence prior to being further encoded using the optimized video encoding scheme described in the present disclosure. In further embodiments, the input video sequence can comprise the decoded or reconstructed video sequence which was decoded from the encoded video bit-stream. Without limiting the scope of the disclosure in any way, it should be noted that the term "frame" used in the specification should be expansively construed to include a single video picture, frame, image, field, or slice of the input video sequence.
[0046] According to certain embodiments, the output of system 100 for optimized video encoding may output video frames which include a bit (or byte) sequence containing a portion or chunk of an encoded video bitstream, where the encoded bitstream chunk, upon decoding, can yield a reconstructed frame that corresponds to the input video frame.
[0047] There is now described the components or blocks in System 100 according to certain embodiments and in relation to both FIGs 1A and IB.
[0048] According to certain embodiments, the system 100 for optimized video encoding of an input video sequence may include an I / O interface 140 and External Memory 150. The system may obtain the input frame and provide the resulting output frame via this I / O interface 140. Alternatively, one or both of the input and output frames may be stored in the Memory 150.
[0049] System 100 includes a processing circuitry 102 operatively connected to a hardware-based I / O interface 140 and configured to provide processing necessary for operating the system, as further detailed with reference to FIGs. 2-4. The processing circuitry 102 can comprise one or more processors (not shown separately) and one or more memories (such as the Internal Memory 156 as illustrated in Fig. IB). The one or more processors of the processing circuitry 102 can be configured to, either separately or in any appropriate combination, execute several functional modules in accordance with computer-readable instructions implemented on a non-transitory computer- readable memory comprised in the processing circuitry. Such functional modules are referred to hereinafter as comprised in the processing circuitry.
[0050] According to certain embodiments, functional modules comprised in processing circuitry 102 can include a Video Frame Encoder 130 and a Video Controller 120, which further may comprise an Encoding Configurator 122, a Stabilizer 124 and an Evaluation Module 126 which are operatively connected with each other. The Video Controller 120 determines the encoding instructions to provide to the Video Frame Encoder 130, which can be configured to perform an encode of a single input video frame, according to the configuration supplied. Further details on these modules, the interactions between them, and the system operation as a whole, will be described below in reference to FIG. 2.
[0051] Note that the Video Frame Encoder 130 is different from the common modus operandi of a video encoder. Video encoding is by nature a sequential process that runs on a sequence of input frames, creating a corresponding sequence of encoded output frames. When a regular video encoder receives an input frame N, it encodes it to obtain an encoded frame N, which is added to the output bitstream. The next video frame the encoder receives will be encoded as output frame N+l, even if the same frame N was once again provided as input, and the encoder may even use the previously encoded frame N as a reference frame in the encoding process. The Video Frame Encoder 130 offers greater flexibility, and can receive a frame N, and encode it to obtain an encoded frame N. Then, if requested, the Video Frame Encoder 130 can receive as the next input frame the same frame N again, and encode it anew as a different option or candidate for an output frame N. This mode of frame encoding operation is only possible in cases where the encoder implementation supports this functionality and offers appropriate APIs to enable operating it in this manner, such as featured in the Video Frame Encoder 130.
[0052] Having reviewed the system components that are common to FIGs. 1A and IB, attention is directed to additional possible components introduced in FIG. IB.
[0053] FIG. IB introduces additional blocks that may be optionally used in the Optimized Video Encoding System 100. According to certain embodiments, functional modules comprised in the system can further include a Video Decoder 170. This may be required if the Input Frame received by the I / O Interface 140 is a compressed video frame. In such a case, there is a need to apply a decoding process using the Video Decoder 170 which gives rise to a decoded frame in the pixel domain. This is the frame that needs to be input into the Video Frame Encoder 130.
[0054] While in the more general case of FIG. 1A a single system Memory block 150 is referred to, in some additional embodiments brought forth in FIG. IB it can be beneficial to have two separate memory units, the system Memory 150, as well as an Internal Memory 156. To clarify, in some embodiments there may not necessarily be a physical separation or distinction between the system Memory 150 and the Internal Memory 156, and they may be integrated as a single memory unit. However, in some embodiments of the disclosed subject matter, such as (but not limited to) the case where the operations applied to video frame pixels by, e.g., the modules of Video Frame Encoder 130 and Video Frame Additional Processing 162, occur within a single processing chip such as a GPU, it can be beneficial to use the GPU internal memory to hold the video frame and share it between these modules, rather than accessing the system Memory 150 which may be a much slower operation. Further, in cases where the Input Frame received by the I / O Interface 140 is a compressed video frame, it can be beneficial to load the compressed frame into the GPU and perform the decoding using a hardware-accelerated decoder present on the GPU to accelerate the loading of the frame.
[0055] In some embodiments of the disclosed subject matter, the Video Frame Additional Processing 162 may perform an operation on the video frame that is completely unrelated to the optimized video compression, such as, by way of nonlimiting examples, object detection, face detection, scene labeling, scene description, video or image to text, and subtitle extraction, to name but a few. In yet other embodiments, the processing may be linked to the encoding process, for example performing processing to identify Region Of Interest (ROI) in the frame and use these for encoder configuration, performing global motion estimation to speed up the encoder processing, and other image processing functionalities that provide information that can be useful for the video encoder. In cases where there is an Internal Memory component 156, the Video Frame Additional Processing 162 may directly access the internal memory, and in the case of GPU encoding mentioned above, this additional processing may also be performed on the GPU for increased computational efficiency, e.g., for offering both faster or shorter processing times and lower processing costs.
[0056] Further, in some embodiments of the disclosed subject matter, it may be desired to apply some preprocessing to the video prior to encoding, using the Video Frame Pre- Processing module 160. This preprocessing may take the form of various filters to improve the perceptual quality of the encoded video, for example by filtering out noise, filtering out film grain, increasing the contrast of the image, applying histogram equalization for better spread of brightness or colors in the video, sharpening, blurring, scaling, cropping, or other resizing filters. This preprocessing may also target other goals, such as removing data from the frame, for example for censoring out problematic content, adding data to the video for example to create a personalized viewing experience, or any other frame pre-process that can be beneficial for some goal in the context of the presently disclosed subject matter. In cases where there is an Internal Memory component 156, the pre-processing operation may directly access this internal memory, and, in the case of GPU encoding, the pre-processing may also be performed on the GPU for increased computational efficiency.
[0057] Further, in some embodiments of the disclosed subject matter, it may be desired to perform resizing of the input image for various reasons. One such use case may be encoding to reduced resolution, to create a ladder of encodes, suitable ,for example, for Adaptive Bit-Rate streaming (ABR). In this case it may be desired to downsize an input using the module of Scale Down 132 to a target size or by a target resizing factor. In yet further embodiments, while the output frame will correspond to the lower resolution, downscaled input video frame, it may be desired for the Evaluation Module 126 to perform the evaluation using the original resolution, in which case the module of Scale Up 134 will be employed on the reconstructed frame, prior to performing the evaluation. In yet another embodiment, it may be desired to perform Scale Up of the input, using various approaches, including but not limited to Al-based Super Resolution approaches, in order to create higher quality video in the output, and thus increasing the quality of the output stream. In cases where there is an Internal Memory component 156, the resizing operations may directly access this internal memory, and in the case of GPU encoding mentioned above, resizing may also be performed on the GPU for increased computational efficiency.
[0058] As aforementioned, the input video sequence can refer to an original video sequence that is not encoded or compressed. Alternatively, it can comprise a previously encoded or compressed video bit-stream which needs to be decoded, for instance prior to the further encoding performed by the Video Frame Encoder 130. It can also refer to the decoded video sequence which has been decoded from the encoded video bitstream. Accordingly, the input frame can be an original video frame, or a decoded or reconstructed video frame.
[0059] The term "encoding configuration" used herein should be construed to cover one or more of the following encoder settings, parameters, or compression parameters: the quantizer or quantization parameter (Q.P), a bit-rate parameter, a compression level indicating a preconfigured set of parameters in a given encoder, as well as various parameters which control encoding decisions, such as, e.g., allowed or preferred prediction modes, Lagrange multiplier lambda value used in rate-distortion and partitioning decisions, deblocking filter configuration, delta QP values between areas in the frame or different image components, etc. For exemplary purposes, certain embodiments of the presently disclosed subject matter are described with reference to the QP. However, it should be noted that the examples of processes and operations provided herein with reference to the QP can also be applied to other types of encoding configurations or parameters.
[0060] Turning now to FIG. 2, there is shown a generalized flowchart of optimized video encoding of an input video sequence in accordance with certain embodiments of the presently disclosed subject matter.
[0061] The first step of the Optimized Video Encoding System 100 applied to an input frame is to obtain said input frame (also referred to as a current frame) from the plurality of input frames 210 (e.g., by the I / O interface 140 illustrated in FIG. 1). In some embodiments, the input video sequence or the input frame thereof can be received from a user, a third party provider, or any other system that is communicatively connected with system 100. Alternatively, or additionally, the input video sequence or the input frame thereof can be pre-stored in the Memory 150 and can be retrieved therefrom.
[0062] As mentioned above, in some embodiments this input frame can be a pixel domain frame ready for encoding, possibly after undergoing decoding external to the Optimized Video Encoding System 100. In other embodiments it may comprise a bitstream chunk corresponding to a compressed video frame, for instance compressed using H.264 / AVC video compression, in which case decoding of the frame within the Optimized Video Encoding System 100 is required, e.g., by the video decoder 170 therein.
[0063] As well as obtaining the actual frame pixels for encoding, it is also beneficial to obtain frame information corresponding to the input frame (214). This frame information may be obtained from the decoder, and may include, for example, the frame type, such as whether in the input it was encoded as an l-frame, P-frame, or B- frame, how many bits (or bytes) it consumed, any metadata associated with the frame, etc. In other embodiments of the disclosed subject matter, this frame information may be obtained by running some processing on the frame. By way of non-limiting example, this information may include determining whether there is a scene change at this frame, what the average level of brightness is, and how that compares to the brightness of preceding frames, etc. The frame information may be associated with the video encoding process, and, as such, may comprise what frame type the encoder is planning to assign to this input frame (l / P / B). In yet further embodiments the frame information may be related to frame dimensions, levels of brightness or color, edge information, extent of motion compared to previous frame, and objects that are detected within the frame, such as but not limited to, face detection, etc.
[0064] While known to those skilled in the art of video compression, for the sake of description completeness, the concept of I, P, and B frames is now outlined. In video encoding it is common practice to use motion-estimation for temporal predictions in order to predict the current frame from previously encoded frames. Generally speaking, the first frame of the video, and of each Group Of Pictures is an I frame, which does not use prediction from other frames. P frames are predicted from previously encoded frames that precede the current frame in display order. B frames are bi-directionally predicted, using frames that were previously encoded, but may appear before or after the current frame in display order. Note that this description is for exemplary purposes, and is by no means intended to limit the scope of the systems and methods described herein.
[0065] In the next step (224), the obtained frame information is provided to, e.g., the stabilizer, followed by querying the stabilizer to determine a set of allowed encoding configurations (also referred to as a set of candidate encoding configurations) for the current frame, wherein the allowed configurations are possible configurations that are in some manner consistent with the configurations used to encode one or more previously encoded frames. The consistency may be in relation to the per frame configurations, or in relation to some moving average reflecting the average of the configurations selected in the previous frames within some window, for instance for the last X frames, frames corresponding to the last Y seconds of video, frames from the beginning of this Group Of Pictures (GOP), and frames from the beginning of the clip, etc. Then, the encoding configurator can be set according to the allowed encoding configurations for the current frame (226). A first candidate encoding configuration can be obtained from the encoding configurator, and the Video Frame Encoder can be configured accordingly (228). These steps are explained in more detail below in reference to FIG. 3 and FIG. 4.
[0066] Once the Video Frame Encoder is configured, the process proceeds to encode the input frame accordingly, giving rise to a candidate encoded frame (230). An encoded frame may refer only to the bits corresponding to the compressed frame chunk, or may additionally refer to the reconstructed frame, which corresponds to the frame that will be displayed upon decoding of the video bitstream. The process then proceeds to evaluate the candidate encoded frame to determine whether it meets a target criterion (240). The target criterion may take different forms. In some embodiments of the disclosed subject matter, the criterion may relate to a specific bitrate or bit consumption used by the encoded video frame. In other embodiments, it may be related to a quality metric. Various quality metrics or quality measures can be used to provide subjective quality evaluation of the encoded frame. Examples of quality measures that can be utilized herein include any of the following, or combinations thereof: Video Multimethod Assessment Fusion (VMAF), Mean Square Error, Peak Signal to Noise Ratio (PSNR), Structural SIMilarity index (SSIM), Multi-Scale Structural SIMilarity index (MS- SSIM), Video Quality Metric (VQM), Visual information Fidelity (VIF), Motion-based Video Integrity Evaluation (MOVIE), Perceptual Video Quality Measure (PVQM), quality measure using one or more of Added Artifactual Edges, texture distortion measures, and a combined quality measure combining inter-frame and intra-frame quality measures, such as described in US patent No. 9,491,464 entitled "Controlling a video content system" issued on November 8, 2016, which is incorporated herein in its entirety by reference.
[0067] Further examples of target criteria may relate to metrics corresponding to underlying ML or Al models. For example, the target criterion may be to be within an allowable range of a Learned Perceptual Image Patch Similarity (LPIPS) range, or a Frechet Inception Distance (FID), used as a metric for evaluating image quality generated by generative models, calculated by computing the Frechet distance between two Gaussians fitted to feature representations of the Inception network, such as DINOv2 (a self-supervised vision transformer trained using self-distillation with no labels).
[0068] In yet other embodiments, the target criterion may be a combined Rate Distortion (RD) criterion, possibly taking the form of Bitrate + Lambda *Distortion, where distortion may be measured with a simple MSE (Mean Squared Error) or any other quality metric.
[0069] In some embodiments, the process can proceed to the next step (250), where it can be determined whether a termination criterion is being met. The termination criterion may take one or more of the following forms: having met the target criterion in the evaluation module, having performed the maximum number of allowed encoding attempts, candidate encoding configuration options are exhausted, and / or any other termination criterion that the system is set to observe. If the termination criterion is not met, the process reverts to step (228) to obtain the next frame configuration among the set of candidate encoding configurations, and repeats steps (230), (240), and (250). Upon the termination criterion being met, the process proceeds to provide the candidate encoded frame which best meets the target criterion as the output frame corresponding to the input frame. The encoding configuration resulting in the output frame is denoted as the selected encoding configuration for the current frame, and the history of selected encoding configuration in the stabilizer can be updated with the selected configuration (260).
[0070] Turning now to FIG. 3, there is shown a functional block diagram schematically illustrating a stabilizer used for optimized video encoding in accordance with certain embodiments of the presently disclosed subject matter.
[0071] According to certain embodiments, functional modules in the stabilizer 124 may include a Frame Information Receiver 310 configured to receive frame information on the current frame, which may be obtained from the Processing Circuitry 102, and / or may be obtained by the Video Decoder 170, and / or may be obtained from the I / O interface 140 alongside the input frame. Alternatively, or in addition, this information may pertain to the frame in relation to the Video Controller 120, such as the type of frame that will be encoded - l-frame, P-frame, or B-frame, or a bit allocation associated with the frame, or any other information pertaining to the encoding of this frame that is decided outside the scope of the stabilizer. In further embodiments, the functional modules in the stabilizer may include an Allowed Configurations Decider 320, which determines, according to all available information, which candidate configuration(s) should be allowed for encoding of the current frame for enabling stabilization of encoding quality and efficiency with respect to the input frame.
[0072] The importance of having a stable encoder is two-fold. First, when the quality of each video frame is not consistent, a temporal flickering effect is caused, which causes the overall impression of quality to be lower than the actual per frame quality. This becomes particularly important in systems where the decision of the configuration of each frame is made separately, to reach a per-frame target criterion. Without stabilizing these configurations in a 'beyond-the-frame' manner, the temporal glitches or inconsistent quality of frames will have a negative impact on the obtained quality. By way of a non-limiting example, if the configuration includes a filter strength to be applied as part of the encoding process, even if applying filters of different strengths to each of a number of adjacent frames makes sense on the per-frame level, the viewing experience in motion will suffer from fluctuation and flickering.
[0073] The second aspect is related to the nature of the video encoding process, comprising usage of inter-frame predictions and context adaptations for efficient encoding. This means that using configurations for adjacent frames that are not consistent, e.g., that exhibit large variability of configurations between adjacent frames, can be detrimental to overall coding efficiency (also referred to as encoding or compression efficiency). Coding efficiency refers to the efficiency offered by a specific video compression workflow, indicating the bitrate quality tradeoff, where higher compression efficiency corresponds to being able to lower the bitstream bitrate while maintaining the same quality, or, vice versa, being able to increase the video quality of the reconstructed video without increasing the bitstream bitrate.
[0074] By way of a non-limiting example, assume the configuration includes a Quantization Parameter (QP) to be used for encoding the frame, and a sequence of 3 P- frames P0, Pl, P2 are to be encoded. If the extent of change between the video frame corresponding to P0 and the video frame corresponding to Pl is similar to the extent of change between the video frame corresponding to Pl and the video frame corresponding to P2, the target criterion can be met using a more aggressive (higher) QP value to encode frame Pl than was used for P0, thus saving bits in the encoded size of the frame corresponding to Pl. In such cases, it is common that the penalty on the required bits to invest in encoding P2 to reach the target criterion will be much higher, making it overall more efficient to use a QP value that is more conservative for Pl, which makes it more stable across the 3 encoded frames. The role of the Allowed Configurations Decider 320 is to assist in avoiding such fluctuations that may harm the quality or compression efficiency of the encoded video created by the proposed system. This is performed by using the Stabilizer Data 330, comprising, in some embodiments, a Stability Rules 340 module, and in yet further embodiments, a history of selected configurations 350.
[0075] Turning now to FIG. 4, there is provided a generalized flowchart showing a process of using a stabilizer for optimized video encoding in accordance with certain embodiments of the presently disclosed subject matter.
[0076] Upon starting the processing of an input video frame, frame information is received (410) from the Video Controller 120. The next step is collecting relevant data from the history of selected configurations (420). In a non-limiting example, in the case the encoding configuration includes a Quantization Parameter (QP), the relevant history may include one or more previously selected QP values, which may be used as-is or after some form of averaging, such as a moving average of the previous QP values, where the moving average may be calculated, for example, over the QP values selected for the frames from the beginning of the GOP, or over the QP values selected for the frames from the beginning of the video sequence, or over the QP values selected for a certain number of previously encoded frames or the over the QP values selected for the frames corresponding to a certain time window. The next step is then applying stability rules in accordance with the received frame information and in accordance with the relevant history (430), followed by creating the set of allowed encoding configurations and providing these to the encoding configurator (440). The encoding configurator uses the set of allowed encoding configurations and selects a candidate configuration for the encoding process, following through the steps 228, 230, 240, and 250 depicted in FIG. 2. After the termination criterion is met and the frame encoding is concluded, the final step of the stabilizer is to update the history of selected configurations in the stabilizer with the selected configuration (460). The scope of FIG. 4 is dedicated to the operations applied for each input frame of the input video sequence. In some embodiments of the disclosed subject matter, additional operations may occur upon initialization of the system, such as defining and / or loading the Stability Rules into the Stabilizer Data, and / or performing a reset of the history of selected configurations.
[0077] The functionality of the stabilizer may be best explained by way of some further non-limiting examples. In some embodiments of the disclosed subject matter, it may be determined that for B-frames, it may be proven beneficial to have an encoding configuration which is confined by the P-frames present in the encoded stream directly before and after the frame. In this case, the frame information would indicate that the video frame encoder is planning to encode this input frame as a B frame, and the stabilizer will allow only configurations that are in some way similar to the configurations which were selected for those preceding and following encoded P-frames. For example, if the encoding configuration comprises a Quantization Parameter (QP) for the frame, the preceding P-frame selected a QP of 30, and the following P-frame selected a QP of 32, the stabilizer may limit the allowed QP values for the B-frame to the range of 30-34. In another example, where the encoding configuration still comprises a QP value, a moving average value of the selected QP value may be observed, and it can be determined that the allowed QP value for the current frame should not deviate from the preceding moving average value more than a pre-determined delta / interval.
[0078] In other embodiments of the disclosed subject matter, the stabilizer may be required to stabilize the behavior of the encoder configurations in order to be ready for future frames. Since the optimization approach operates in a single pass mode over the input content, even when it seems that using aggressive encoding configurations, for instance a very high QP value, is possible for a specific frame without compromising the target criterion, this may cause the next frame(s), which use this frame as reference for instance for motion or inter prediction, to not be able to meet the target criterion, even when using the most lenient encoding configuration possible, for instance a minimum QP value allowed. In yet other embodiments of the disclosed subject matter, the stabilizer may be used to overcome cases that may occur where the metric used to evaluate whether the target criterion is met suffers from some 'blind spots', causing the result for a specific frame to be unreliable. By way of non-limiting example, the VMAF result for the frame may not align with actual subjective quality, for example due to visible degradation occurring only in a very small area of the encoded (or reconstructed) frame, which may not be properly reflected in the frame VMAF score. Without a stabilizer as part of the solution, such frames could cause the resulting output video to have compromised quality, as the configuration for such a frame will continue to get more and more aggressive while the target criterion is met. Alternatively, in some cases the VMAF score may be lower than the actual perceptual frame quality of the frame, for example due to poor contrast in the frame in fade-in or fade-out scenes. Without a stabilizer as part of the solution, such frames could cause the resulting output video to have an inflated bitrate, as the configuration for such a frame will continue to adapt to try and improve the quality, causing it, for instance, to reach very low QP values and incurring significant bitrate penalties. In both these cases, by keeping the extent of compression for the frame within a stable range, as indicated by information gathered thus farfrom encoding the video sequence up to this point, the loss of quality and / or unnecessary bitrate inflation can be avoided.
[0079] Reference is now made, by way of non-limiting example, to the stability rules 340. In some embodiments of the disclosed subject matter, where the encoding configuration sets a QP value, the stability rules may comprise an allowed range for the QP values to be used in encoding the current frame. Note that these rules may override the objective of meeting the target criterion. For example, a stability rule may indicate that the QP for a current frame may not deviate from the QP selected for the previously encoded frame, or the previously encoded frame of the same type (l / P / B) by more than [+X] or [-Y] . For example, if the previously encoded frame selected a QP of 27, and the values of X and Y are set to 3 and 5 respectively, the allowed QP for the current frame is in the range of 22 - 30. In this example, if the target criterion requires that a candidate frame reaches a VMAF value above 90, and the QP value of 22 (lowest allowed) corresponds to a lower VMAF, of say 85, the frame may still be outputted, which can be remedied in the next frame(s). Additionally or alternatively, a stability rule may indicate that the QP for a current frame may not deviate from the moving average of the QP values selected for the previously encoded frames, or the previously encoded frames of the same type (l / P / B) by more than [+X2] or [-Y2] . For example, if the moving average of the previously encoded frames' selected QPs is 25.3, and the values of X and Y are set to 2 and 6 respectively, the allowed QP for the current frame is in the range of 19.3 - 27.3, or in cases of integer QP values being used, 20-27.
[0080] In yet other embodiments of the disclosed subject matter, the stability rules might take a more complex form, where they adapt according to the frame information, such as frame type and additional information related to the video content. By way of non-limiting example, it may be determined that upon a scene change being identified at the current frame, the rules are either modified or removed altogether. In another non-limiting example, in case a fade-in or fade-out is identified in the current scene, or in case the scene is recognized to be very unstable in nature, the rules may be modified to allow for better adaptation to the content variability.
[0081] There is now described a possible embodiment in relation to the disclosed subject matter. The Optimized Video Encoding System 100 can be used to perform content modernization or codec modernization. This refers to using the system to perform transcoding of content stored using a first codec, which, in some embodiments, may be a legacy codec, such as MPEG-2 or even AVC, to a second, and in some embodiments more efficient, modern codec such as HEVC, AVI, or VVC, thus enjoying lower bitrates made possible by the increased compression efficiency of modern codecs. One of the obstacles in performing this modernization is determining the encoding configuration which will not compromise / reduce the quality of the legacy videos for any of the varied content, but also benefit from properties, such as the compression efficiency offered by the modern codec. A cautious approach could be to use a very high- quality setting, such as a high target bitrate or low QP value setting for the modern encoder, resulting in larger than needed files. Alternatively, a more conservative configuration or target bitrate could be selected, which, for some particularly challenging files, may cause a drop in quality, while for other, easy to encode content, there may again be inflated files that are not fully utilizing the possible savings in file size. In order to avoid any excessive bitrates across diverse content, it would be necessary to use quite aggressive encoding configurations, which, for the more challenging content, would cause a degradation in the encoded video quality. Therefore, in some embodiments of the subject matter disclosed herein of the Optimized Video Encoding System 100, the target criterion is set to create an encode that is perceptually identical to the input content. By way of non-limiting example, the perceptual identity of the encoded frame to the input frame may be measured using the quality measures set forth above. Furthermore, it can be desirable to use a hardware-accelerated encoder as the Video Frame Encoder 130 in order to achieve a very efficient, low-cost, fully automated solution for video codec modernization.
[0082] In yet further embodiments of the subject matter disclosed herein, the Optimized Video Encoding System 100 may be used to perform encoding, to obtain a target VMAF value for the encode. In this case the target criterion may be a particular VMAF value to reach per frame. This frame value may be constant across frames, or may be modulated according to frame type and / or the VMAF values of the previously encoded frames and / or bitrate values of the previously encoded frames, or may be determined using a Rate-Distortion approach using a criterion that combines both the VMAF score of the frame and the bits or bytes needed to encode the frame in a Rate Distortion approach of R + lambda * D. Note that VMAF is brought here as a non-limiting example of a metric to try and meet, but the same system may be used mutatis mutandis using other quality metrics, or other target criteria altogether.
[0083] In yet further embodiments of the subject matter disclosed herein, the Optimized Video Encoding System 100 may be used to perform live optimized video encoding. Many approaches for optimized or content adaptive encoding require offline processing with multiple passes on the content, or at least multiple passes on each Group Of Pictures (GOP) in order to determine the best encoding configuration which meets a target criterion. Other approaches may avoid the need for this full GOP processing by applying models that are learned or developed offline, and applied in a 'best-effort' approach, with no validation that a target criterion, for instance of a certain quality level, is indeed met. Further, some solutions use software encoding running on a GPU (Graphics Processing Unit) which may struggle to support live encoding for the more complex encoders such as AVI, for high resolution and high frame rate content. The proposed system enables live optimized content adaptive encoding which ensures an average target quality, while maintaining live performance and low latency, particularly when using a hardware-accelerated encoder such as the Video Frame Encoder 130, and applying a modified controller which allows only a small number of candidate encodings per frame, such as, e.g., two, to avoid increasing stream latency. As aforementioned, in some cases, in order to use a video encoder in this manner, it must support appropriate APIs (Application Programming Interfaces) that enable encoding of the same input frame more than once, while not advancing the encoder state, such as, e.g., the Nvidia NVENC codec SDKs, which include support of this mode of operation.
[0084] The Memory 150 comprises a non-transitory computer-readable storage medium. For instance, the storage module can include a buffer that holds an input video sequence as well as an output video sequence. In another example, the buffer may also hold one or more of the intermediate results, including input video frame(s), candidate encoded frame(s), valid candidate encoded frame(s), encoding instruction(s), and parameter(s), etc.
[0085] Those versed in the art will readily appreciate that the teachings of the presently disclosed subject matter are not bound by the system illustrated in FIG. 1 and the above exemplified implementations. Equivalent and / or modified functionality can be consolidated or divided in another manner, and can be implemented in any appropriate combination of software, firmware, and hardware. By way of example, the functionalities of the video encoder 108 as described herein can be divided and implemented as separate modules operatively connected thereto. For instance, such a division can be in accordance with the Video Controller 120 and the Video Frame Encoder 130, and / or in accordance with different stages within each of these modules. By way of another example, the Evaluation Module 126, the Stabilizer 124, and the Encoding Configurator 122 can be either implemented individually and in connection with the Video Frame Encoder 130, or, alternatively, one or more of these modules can be integrated within the video encoder 130.
[0086] The system in FIG. 1 can be a standalone network entity, or integrated, fully or partly, with other network entities. Those skilled in the art will also readily appreciate that the data repositories or storage module therein can be shared with other systems or be provided by other systems, including third party equipment.
[0087] It is also noted that the system illustrated in FIG. 1 can be implemented in a distributed computing environment, in which the aforementioned functional modules shown in FIG. 1 can be distributed over several local and / or remote devices, and can be linked through a communication network.
[0088] While not necessarily so, the process of operation of system 100 can correspond to some or all of the stages of the methods described with respect to FIGs. 2 and 4. Likewise, the methods described with respect to FIGs. 2 and 4 and their possible implementations can be implemented by system 100. It is therefore noted that embodiments discussed in relation to the methods described with respect to FIGs. 2 and 4 can also be implemented, mutatis mutandis, as various embodiments of the system 100, and vice versa.
[0089] Those versed in the art will readily appreciate that the examples illustrated with reference to FIGs. 1-4 are by no means inclusive of all possible alternatives, but are intended to illustrate non-limiting examples, and, accordingly, other ways of implementation can be used in addition to or in lieu of the above.
[0090] It is to be understood that the presently disclosed subject matter is not limited in its application to the details set forth in the description contained herein or illustrated in the drawings. The presently disclosed subject matter is capable of other embodiments and of being practiced and carried out in various ways. Hence, it is to be understood that the phraseology and terminology employed herein are for the purpose of description, and should not be regarded as limiting. As such, those skilled in the art will appreciate that the conception upon which this disclosure is based can readily be utilized as a basis for designing other structures, methods, and systems for carrying out the several purposes of the presently disclosed subject matter.
[0091] It will also be understood that the system according to the presently disclosed subject matter can be implemented, at least partly, as a suitably programmed computer. Likewise, the presently disclosed subject matter contemplates a computer program being readable by a computer for executing the disclosed method. The presently disclosed subject matter further contemplates a non-transitory computer-readable memory or storage medium tangibly embodying a program of instructions executable by the computer for executing the disclosed method. Specifically, a Graphic Processing Unit, or GPU, may be used as the infrastructure for execution of at least some of the functionalities of the disclosed method. The computer-readable storage medium causing a computerto carry out aspects of the present invention can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
Claims
CLAIMS1. A computerized method of optimized video encoding of an input video sequence comprising a plurality of input frames, the method comprising: obtaining an input frame of the plurality of input frames and frame information thereof; determining a set of candidate encoding configurations usable for encoding the input frame based on the frame information, the set of candidate encoding configurations selected for enabling stabilization of encoding quality and efficiency with respect to the input frame; encoding the input frame using a first candidate encoding configuration selected from the set of candidate encoding configurations, to obtain a candidate encoded frame; evaluating the candidate encoded frame with respect to a target criterion, and verifying if a termination criterion is met; if not, repeating the encoding, evaluating, and verifying using a next candidate encoding configuration selected from the set, until the termination criterion is met; and selecting, as an output frame, a candidate encoded frame best meeting the target criterion.
2. The computerized method according to claim 1, further comprising adding a selected encoding configuration corresponding to the output frame to a history of selected encoding configurations, giving rise to an updated history of selected encoding configurations.
3. The computerized method according to claim 2, wherein the determination is in accordance with a set of stability rules, and using the history of selected encoding configurations.
4. The computerized method according to claim 3, wherein the set of stability rules comprises a beyond-the-frame stability rule for selecting the set of candidate encoding configurations to be consistent with one or more selected encoding configurations for one or more previously encoded frames.
5. The computerized method according to claim 4, wherein the set of candidate encoding configurations includes a set of candidate Quantization Parameter (QP) values, and the history of selected encoding configurations includes one or more previous QP values selected for one or more previously encoded input frames.
6. The computerized method according to claim 5, wherein the beyond-the-frame stability rule stipulates that the set of candidate QP values does not deviate from the previous QP values by more than a predetermined interval.
7. The computerized method according to claim 1, wherein the target criterion is based on a bit consumption metric and / or a quality measure.
8. The computerized method according to claim 7, wherein the quality measure is a per-frame Video Multimethod Assessment Fusion (VMAF) value.
9. The computerized method according to claim 3, wherein the set of stability rules overrides the target criterion, and an encoded frame substantially meeting the target criterion is provided as the output frame.
10. The computerized method according to claim 1, wherein the frame information comprises one or more of frame type, bit consumption, metadata associated with the input frame, indication of scene change, average level of brightness, frame dimensions, and detected objects.
11. The computerized method according to claim 1, wherein the input video sequence corresponds to a first codec, and the encoding of the input frame is performed for transcoding to a second codec, wherein the set of candidate encoding configurations is determined so as to not reduce quality of the input video sequence while benefiting from properties offered by the second codec.
12. The computerized method according to claim 11, wherein the first codec is a legacy codec, the second codec is a modern codec, and the properties are related to compression efficiency.
13. The computerized method according to claim 1, wherein the termination criterion includes at least one of: the target criterion is met, a predetermined number of iterations is reached, and candidate encoding configuration options are exhausted.
14. A computerized system of optimized video encoding of an input video sequence comprising a plurality of input frames, the system comprising a processing circuitry configured to perform method steps of any of claims 1-13.
15. A non-transitory computer-readable storage medium tangibly embodying a program of instructions that, when executed by a computer, cause the computer to perform method steps of any of claims 1-13.
Citation Information
Patent Citations
Method and device for adjusting encoder parameters, computer equipment and storage medium
CN112584147A
System and method for determining encoding parameters
US20110268187A1
Iterative techniques for encoding video content
US20210160510A1