Machine learning (ML)-based rate control algorithm for finding quantization parameters (QP)
By generating dynamic QP for the video encoder through a machine learning model, the high processing power requirements and buffer problems in the existing technology are solved, and efficient video encoding optimization and quality improvement are achieved.
Patent Information
- Application Number
- CN202480013996.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-06
- Filing Date
- 2024-01-12
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies require too much processing power when optimizing quantization parameters (QP) for video encoders and are prone to video rate buffer overflow or underflow, making them unable to effectively adapt to variable transmission channel characteristics.
A machine learning model is used to generate dynamic QP strings, and a random forest model is used to generate quantization parameters for different types of video encoders, which are optimized by combining the bit representation of video frames, DCT energy, and transmission channel rate information.
Improves video encoding efficiency, reduces processing requirements, avoids buffer overflow or underflow problems, and improves video quality.
Smart Images

Figure CN120677494A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to machine learning (ML) for finding optimal quantization parameters (QPs) for video encoders. Background Art
[0002] Video encoders use quantization parameters (QPs) (also referred to herein as quantization factors or quantization values) to compress Moving Picture Experts Group (MPEG) video frames, which include intra frames (I frames), predicted frames (P frames), and bidirectional frames (B frames). Both P frames and B frames are based on I frames and are smaller than I frames to achieve a smaller data payload than would be possible if only full-scale I frames were used. The encoder uses QPs when generating these frames to optimize video quality.
[0003] As understood herein, current techniques for generating optimal QP for video encoders fall short in optimizing the QP and / or require excessive processing power, such as is the case with so-called two-pass encoding. Further, current techniques for generating optimal QP for video encoders may result in video rate buffer overflow or underflow conditions depending on variable transmission channel characteristics. Summary of the Invention
[0004] Therefore, machine learning is used to generate dynamic QP strings for the video encoder.
[0005] In a detailed example, QPs for frames other than I-frames (also referred to herein as reference frames) are generated by an ML model. In some examples, one ML model per encoder type is provided, i.e., the ML models described herein can be customized for encoder type such that a first ML model is used to generate QPs for non-reference frames of a first encoder type, and a second ML model is used to generate QPs for non-reference frames of a second encoder type.
[0006] With this introduction in mind, a device includes at least one processor configured to input at least one non-intra frame (I-frame), at least one reference video frame, and at least a bit representation of a video frame into at least one machine learning (ML) model. The processor is configured to provide output from the ML model to at least one video encoder. The output includes at least a quantization parameter (QP). The processor is configured to execute the video encoder to output at least one video frame based at least in part on the QP.
[0007] In an example embodiment, a bit representation of a video frame is received from a video encoder.
[0008] In an example embodiment, the processor is configured to input the encoded bits of at least one block of the encoded video frame to the ML model.
[0009] In some implementations, the processor is configured to input an energy representation of a discrete cosine transform (DCT) of the block of the encoded video frame to the ML model.
[0010] The processor may be configured to input at least one non-I frame and at least one reference video frame to a video encoder.
[0011] The ML model may include at least one random forest model.
[0012] In some embodiments, the first module of the ML model outputs an estimated bit representation of at least one video frame to at least a second module of the ML model. In these embodiments, the first module is configured to receive a non-I frame, at least one reference video frame, and a bit representation of a video frame, wherein the first module is configured to output the estimated bit representation to a video encoder and a second module. The second module can be configured to receive coded bits of at least one block of the coded video frame from the encoder and output the QP to the video encoder. Furthermore, the second module can be configured to receive an energy representation of a DCT of at least one block of the coded video frame from the video encoder.
[0013] In another aspect, a method includes providing at least one indication of a video transmission channel rate to at least a first machine learning (ML) module. The method further includes providing an output of the first ML module to at least a second ML module. The output of the first ML module includes at least an estimated number of bits for at least a portion of at least one video frame. Further, the method includes providing the output of the second ML module to at least one video encoder. The output of the second ML module includes at least one quantization parameter (QP). According to the method, the video encoder uses the QP to generate at least one video frame for transmission over the video transmission channel.
[0014] In another aspect, an apparatus includes at least one computer storage medium that is not a transient signal and further includes instructions executable by at least one processor to input video encoding information to at least one machine learning (ML) model. The instructions are executable to provide output of the ML model to at least one video encoder. The output includes at least a quantization parameter (QP). The instructions are executable to execute the video encoder to generate video information using the QP.
[0015] The details of the present application, both as to its structure and operation, may be best understood with reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a block diagram of an example system according to the principles of the present invention;
[0017] Figure 2 is a block diagram of an example encoder apparatus consistent with the principles of the present invention;
[0018] Figure 3 is a block diagram of an example decoder apparatus consistent with the principles of the present invention;
[0019] Figure 4 is a schematic diagram of an overall example software architecture consistent with the principles of the present invention;
[0020] Figure 5 An example ML model architecture is shown;
[0021] Figure 6 Shown Figure 5 Details of the frame bit module shown in;
[0022] Figure 7 and Figure 8 The equations related to training are shown;
[0023] Figure 9 An example ML model architecture during the training phase is shown;
[0024] Figure 10 Example logic for encoding a video using an ML model during operation is shown in an example flowchart format;
[0025] Figure 11 Example logic for a training frame bits module is shown in example flow chart format; and
[0026] Figure 12 Example logic for training a QP module is shown in an example flow chart format. DETAILED DESCRIPTION
[0027] The present disclosure generally relates to computer ecosystems that include aspects of a network of consumer electronics (CE) devices, such as, but not limited to, a computer gaming network. The systems herein may include a server component and a client component that may be connected via a network so that data may be exchanged between the client component and the server component. The client component may include one or more computing devices, including a gaming console (such as a Sony PlayStation 4). ®or game consoles made by Microsoft or Nintendo or other manufacturers), extended reality (XR) headsets (such as virtual reality (VR) headsets, augmented reality (AR) headsets), portable televisions (e.g., smart TVs, internet-enabled TVs), portable computers (such as laptops and tablets), and other mobile devices (including smartphones and additional examples discussed below). These client devices can operate in a variety of operating environments. For example, some client computers may employ, for example, a Linux operating system, an operating system from Microsoft, or a Unix operating system, or an operating system produced by Apple, Inc. or Google, or the Berkeley Software Distribution or Berkeley Standard Distribution (BSD) OS (including descendants of BSD). These operating environments can be used to execute one or more browsing programs, such as browsers made by Microsoft or Google or Mozilla or other browser programs that can access websites hosted by internet servers discussed below. In addition, an operating environment according to the principles of the present invention can be used to execute one or more computer game programs.
[0028] A server and / or gateway may be used, which may include one or more processors that execute instructions that configure the server to receive and transmit data over a network such as the Internet. Alternatively, the client and server may be connected via a local intranet or virtual private network. The server or controller may be controlled by a game console such as a Sony PlayStation 4 or a similar game console. ® , personal computers, etc.) instantiation.
[0029] Information can be exchanged between the client and the server over the network. To this end, and for security reasons, the server and / or the client may include firewalls, load balancers, temporary storage and proxies, and other network infrastructure for reliability and security. One or more servers may form a device that implements a method for providing a secure community (such as an online social networking site or a gamer network) to network members.
[0030] The processor may be a single chip or multi-chip processor that may execute logic with the aid of various lines such as address lines, data lines, and control lines, as well as registers and shift registers. A processor comprising a digital signal processor (DSP) may be one embodiment of the circuit system.
[0031] The components included in one embodiment may be used in other embodiments in any appropriate combination. For example, any of the various components described herein and / or depicted in the accompanying drawings may be combined, interchanged, or excluded from other embodiments.
[0032] “A system having at least one of A, B, and C” (similarly, “a system having at least one of A, B, or C” and “a system having at least one of A, B, and C”) includes the following systems: having only A; having only B; having only C; having both A and B; having both A and C; having both B and C; and / or having both A, B, and C.
[0033] Now refer to Figure 1 , shows an example system 10 that may include one or more of the example devices mentioned above and described further below in accordance with the principles of the present invention. The first of the example devices included in system 10 is a consumer electronics (CE) device, such as an audio-video device (AVD) 12, such as, but not limited to, a cinema display system (which may be projector-based) or an internet-enabled TV with a TV tuner (equivalently, a set-top box that controls the TV). Alternatively, AVD 12 may also be a computerized internet-enabled ("smart") phone, a tablet computer, a laptop computer, a head-mounted device (HMD) and / or headset (such as smart glasses or a VR headset), another wearable computerized device, a computerized internet-enabled music player, a computerized internet-enabled headset, a computerized internet-enabled implantable device (such as an implantable skin device), etc. In any case, it should be understood that AVD 12 is configured to implement the principles of the present invention (e.g., communicate with other CE devices to implement the principles of the present invention, execute the logic described herein, and perform any other functions and / or operations described herein).
[0034] Thus, to implement such principles, an AVD 12 can be built with some or all of the components shown. For example, the AVD 12 can include one or more touch-enabled displays 14, which can be implemented as high-definition or ultra-high-definition "4K" or higher-definition flat-panel screens. The touch-enabled display 14 can include, for example, a capacitive or resistive touch sensing layer with an electrode grid for touch sensing consistent with the principles of the present invention.
[0035] The AVD 12 may also include one or more speakers 16 for outputting audio in accordance with the principles of the present invention, and at least one additional input device 18 (such as an audio receiver / microphone) for entering audible commands into the AVD 12 to control the AVD 12. The example AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22 (such as the Internet, a WAN, a LAN, etc.) under the control of one or more processors 24. Thus, the interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that the processor 24 controls the AVD 12 to implement the principles of the present invention, including controlling other elements of the AVD 12 described herein, such as controlling the display 14 to present images on the display and receiving input from the display. Furthermore, it should be noted that the network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless phone transceiver or the Wi-Fi transceiver mentioned above.
[0036] In addition to the foregoing, the AVD 12 may also include one or more input and / or output ports 26, such as a High Definition Multimedia Interface (HDMI) port or a Universal Serial Bus (USB) port for physically connecting to another CE device and / or a headphone port for connecting headphones to the AVD 12 to present audio from the AVD 12 to the user via the headphones. For example, the input port 26 may be connected wired or wirelessly to a cable or satellite source 26a of audio and video content. Thus, the source 26a may be a separate or integrated set-top box, or a satellite receiver. Alternatively, the source 26a may be a game console or disk player containing content. When implemented as a game console, the source 26a may include some or all of the components described below with respect to the CE device 48.
[0037] The AVD 12 may also include one or more computer memories / computer-readable storage media 28 that are not transient signals, such as disk-based storage or solid-state storage, which in some cases are embodied as a separate device in the AVD's housing, or as a personal video recorder (PVR) or video disk player for playback of AV programs inside or outside the AVD's housing, or as a removable storage medium or server as described below. Additionally, in some embodiments, the AVD 12 may include a location or positioning receiver, such as, but not limited to, a cellular telephone receiver, a GPS receiver, and / or an altimeter 30, which is configured to receive geographic location information from a satellite or cellular telephone base station and provide this information to the processor 24 and / or in conjunction with the processor 24 to determine the altitude at which the AVD 12 is located.
[0038] Continuing with the description of the AVD 12, in some embodiments, the AVD 12 may include one or more cameras 32, which may be thermal imaging cameras, digital cameras (such as webcams, IR sensors, event-based sensors), and / or cameras integrated into the AVD 12 and controllable by the processor 24 to capture pictures / images and / or videos in accordance with the principles of the present invention. The AVD 12 may also include a Bluetooth ® A transceiver 34 and other near field communication (NFC) components 36 may be provided to communicate with other devices using Bluetooth and / or NFC technology, respectively. An example NFC component may be a radio frequency identification (RFID) component.
[0039] Still further, the AVD 12 may include one or more auxiliary sensors 38 that provide input to the processor 24. For example, one or more of the auxiliary sensors 38 may include one or more pressure sensors that form a layer of the touch-enabled display 14 itself, and may be, but are not limited to, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, and the like. Other sensor examples include pressure sensors, motion sensors (e.g., accelerometers, gyroscopes, odometers, or magnetic sensors), infrared (IR) sensors, optical sensors, speed and / or cadence sensors, event-based sensors, and gesture sensors (e.g., for sensing gesture commands). Thus, the sensors 38 may be implemented by one or more motion sensors, such as individual accelerometers, gyroscopes, and magnetometers, and / or an inertial measurement unit (IMU), which typically includes a combination of accelerometers, gyroscopes, and magnetometers to determine the position and orientation of the AVD 12 in three dimensions, or by event-based sensors, such as an event detection sensor (EDS). An EDS consistent with the present disclosure provides an output indicating a change in light intensity sensed by at least one pixel of a light sensing array. For example, if the light sensed by the pixel is decreasing, the output of the EDS may be -1; if the light sensed by the pixel is increasing, the output of the EDS may be + 1. Outputting a binary signal of 0 may indicate no change in light intensity below a certain threshold.
[0040] The AVD 12 may also include an over-the-air TV broadcast port 40 that provides an input to the processor 24 for receiving OTA TV broadcasts. In addition to the foregoing, it should be noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an IR data association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, such as a kinetic energy harvester that converts kinetic energy into electricity to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field programmable gate array 46 may also be included. One or more tactile / vibration generators 47 may be provided for generating tactile signals that can be sensed by a person holding or touching the device. Thus, the tactile generator 47 may vibrate all or part of the AVD 12 using an electric motor connected to an eccentric and / or unbalanced counterweight via a rotatable shaft of the motor so that the shaft can rotate under the control of the motor (which in turn can be controlled by a processor such as processor 24) to produce vibrations of various frequencies and / or amplitudes and force simulations in various directions.
[0041] A light source may also be included, such as a projector, such as an infrared (IR) projector.
[0042] In addition to the AVD 12, the system 10 may also include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to send audio and video of a computer game to the AVD 12 via commands sent directly to the AVD 12 and / or through a server described below, while the second CE device 50 may include components similar to the first CE device 48. In the example shown, the second CE device 50 may be configured as a computer game controller manipulated by the player or a head-mounted display (HMD) worn by the player. The HMD may include a heads-up transparent or opaque display for presenting AR / MR content or VR content (more generally, extended reality (XR) content), respectively. The HMD may be configured as a glasses-type display, or as a larger VR-type display sold by computer game device manufacturers.
[0043] In the example shown, only two CE devices are shown, and it should be understood that fewer or more devices may be used. The devices herein may implement some or all of the components shown for AVD 12. Any of the components shown in the following figures may be combined with some or all of the components shown in the context of AVD 12.
[0044] Referring now to the aforementioned at least one server 52, the at least one server includes at least one server processor 54, at least one tangible computer-readable storage medium 56 (such as disk-based storage or solid-state storage), and at least one network interface 58, which, under the control of the server processor 54, allows communication with the other illustrated devices over the network 22 and, in practice, facilitates communication between the server and client devices in accordance with the principles of the present invention. Note that the network interface 58 can be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface (such as, for example, a wireless telephone transceiver).
[0045] Thus, in some embodiments, server 52 may be an internet server or an entire server "farm," and in example embodiments for, for example, an online gaming application, the server may include and perform "cloud" functionality such that the devices of system 10 may access the "cloud" environment via server 52. Alternatively, server 52 may be implemented by one or more gaming consoles or other computers in the same room or nearby as the other devices shown.
[0046] The components shown in the following figures may include some or all of the components shown herein.Any user interface (UI) described herein may be combined and / or extended, and UI elements may be mixed and matched between UIs.
[0047] The present invention principle can adopt various machine learning models, including deep learning models. The machine learning model that conforms to the present invention principle can use various algorithms trained in the following manner: supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning and other forms of learning. Examples of such algorithms that can be implemented by computer circuit systems include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and a type of RNN called long short-term memory (LSTM) networks. Support vector machines (SVMs) and Bayesian networks can also be considered as examples of machine learning models. In addition to the above-mentioned network types, the models herein can also be implemented by classifiers. A specific type of ML model that can be used is a random forest (RF) model, which is a type of non-convolutional neural network.
[0048] As understood herein, therefore, performing machine learning may involve accessing training data and then training a model on the training data so that the model can process additional data to make inferences. Thus, an artificial neural network / artificial intelligence model trained via machine learning may include an input layer, an output layer, and a plurality of hidden layers therebetween that are configured and weighted to make inferences about an appropriate output.
[0049] Now refer to Figure 2 , shows at least one example transmitter / encoder device 200. The device 200 may be provided by, for example Figure 1 The video may be created by a device (such as one or more servers that maintain and manage the VCN to provide the GOP of the video). The video may be created by computer game video or other interactive video, or even video of other audio / video content (such as a movie), Internet-based YouTube video, video embedded in a website, video streamed to a set-top box via cable and satellite connections, video provided by a dedicated software application (app), etc.
[0050] therefore, Figure 2 A video source 202 of video is shown, which can be established by persistent storage (such as a hard drive and / or solid-state drive) on the device 200 or a cloud server, or other sources. Thus, a processor (not shown) in the device 200 can access the video from the source 202 and pass the video through an encoder 204 to encode the GOPs of the video, and provide the encoded GOPs to the network interface 206 of the device 200 for transmission to one or more receiver / decoder devices. The network interface 206 can be a computer network interface, a wireless telephone interface, a television broadcast or satellite or terrestrial broadcast interface, etc.
[0051] Without limitation, the encoders and decoders of this document may operate under a variety of standards including ITU MPEG2, H.264 / AVC, H.265 / HEVC, H.266 / VVC, AVI, VP8, VP9, MPEG5-Part1 / EVC / Essential Video Coding, and MPEG5-Part2 / LCEVC / Low Complexity Enhancement Video Coding.
[0052] Now refer to Figure 3 , shows at least one example receiver / decoder device 300. The device 300 may receive the GOP from the transmission device 200 via the GOP. Figure 2 The apparatus 300 may be composed of, for example, Figure 1 The device is established, the Figure 1 The apparatus 300 may be a device such as a personal computer (e.g., a desktop computer, a laptop computer), a smart phone, a head-mounted device (e.g., smart glasses or an augmented reality or virtual reality head-mounted device), etc. Additionally or alternatively, the apparatus 300 may be established by a computer game console, a television, a set-top box, etc.
[0053] like Figure 3 As shown in , the device 300 may include a network interface 302 that may receive the video / GOP from the device 200 via a wired LAN or Wi-Fi internet connection, a cable or satellite connection, a 5G or other wireless cellular internet connection, etc. Thus, in various examples, the network interface 302 may be a computer network interface, a wireless telephone interface, a television broadcast or satellite or terrestrial broadcast interface, etc.
[0054] When interface 302 receives a GOP from device 200, interface 302 may provide the GOP to decoder 304 for decoding. Decoder 304 may then provide the decoded GOP to a central processing unit (CPU) or other processor of device 300 and / or control a display driver connected to device 300 to place the GOP into a display buffer maintained at device 300.
[0055] For I-frames, a predetermined / predefined QP can be used for the initial / first frame of a GOP (a standalone I-frame) and can be set by a designer, programmer, etc. Thus, the QP for that frame can be fixed (e.g., the frame does not undergo any processing to obtain a QP value). The remaining I-frames and non-I-frames of the GOP can go through the loop process described herein and obtain QP values. Thus, for periodic I-frames in the remaining frames of the same GOP and for the same scene in the GOP, the QP can be calculated as the average of the QP of the last frame and the QP of the last I-frame. For non-periodic I-frames in the same GOP but establishing a scene change / different scene, the QP can be calculated as the average QP of past I-frames from the same GOP.
[0056] For non-I frames, the QP can be derived from the ML model described below.
[0057] Therefore, and in order to achieve the latter purpose, Figure 4 In the example shown, video encoding information 300, such as actual or target channel rate, is input to a machine learning (ML) model, which is generally designated as 402. In the example shown, the video encoding information is input to a dynamic frame bit allocator module 404, the output of which is input to a QP module 406, which outputs a respective quantization parameter (QP) for each of a plurality of regions or slices or macroblocks or code processing tree units (CTUs) or other portions of a video frame for use by an encoder, such as Figure 2 The encoder shown in ) is used when encoding the corresponding part of the corresponding video frame to store or transmit the video frame to a player. Figure 4 One or both of the modules shown in can be implemented by a non-convolutional neural network such as a random forest network.
[0058] Now go to Figure 5 , which shows a video encoder 500, such as, but not limited to, any of the video encoders described herein. To encode a video, non-I frames 502 (such as B frames or P frames) and reference frames 504 (such as I frames) are input to the encoder 500. Furthermore, the frames 502, 504 are input to the ML model 402, which in this example is input to Figure 4 The frame bit module 404 shown in .
[0059] Additional video encoding data is input to the ML model. For example, as indicated at 505, the coded bits of the previous (nth) frame and / or a numerical value representing the number of such bits are input from the encoder 500 to the frame bits module 404. In addition, the encoder 500 may input a representation 506 of the coded bits of at least one block (such as a CTU) of the coded video frame to the QP module 406. In addition, as indicated at 508, the encoder 500 may provide an energy representation of the discrete cosine transform (DCT) of at least one block of the coded video frame.
[0060] like Figure 5 As shown at 510 in FIG, the frame bits module 404, based on its inputs, outputs a representation of the estimated (or target) bits for at least one video frame to both the encoder 500 and the QP module 406. In turn, the QP module 406, based on its inputs (506, 508, 510), outputs a QP for the current portion of the frame being encoded (such as a CTU) to the encoder 500. Figure 5 The feedback loop shown in continues for each successive frame (or portion of a frame, such as a CTU) encoded by the encoder 500.
[0061] Now it can be understood that based on training, Figure 4 and Figure 5 The ML model 402 shown in and described above outputs a QP for all or a portion of a frame (such as a CTU) used to encode the frame to the encoder, and provides a target or estimated number of bits for the frame (or portion of the frame) to the encoder.
[0062] Figure 6 Details of the frame bits module 404 during training are shown. The non-I frame 600 and the I frame 602 are summed at 604 to produce a residual frame 604 and are also directly input to the correlation coefficient block 606 which provides the correlation coefficient "r" to the correlation strength block 608. The correlation coefficient "r" is Figure 6 The quotient of the terms shown in block 606 of , and is sometimes called the Pearson correlation coefficient, which has a value between -1 and 1. In the equation for block 606, X i and Y i is the value of the i-th variable in the sample, and the x-values and y-values with lines above them represent the means of the corresponding values.
[0063] Figure 6 The correlation coefficient ranges in the left-hand column of the Correlation Strength block 608 are mapped to the subjective correlation strengths in the middle column and the correlation types (zero, positive, or negative) in the right column.
[0064] like Figure 6As shown in FIG, the correlation strengths from the strength block 608 are input to the ML model under test 610, which will establish Figure 4 and Figure 5 6. The ML model under test 610 also receives the residual frame 604, an indication 612 of the size of the frame being encoded, and a ground truth estimated (or target) number of bits from a rate buffer 614. The ML model 610 is tested on this input data to minimize the difference between the output 616 of the ML model 610 and the ground truth target from the rate buffer 614.
[0065] Figure 7 and Figure 8 Equations 700 and 800 are shown to further discuss the training of the ML frame bit module 404. As mentioned above, the equations include the correlation coefficient between video frames, frame size, and frame bits. The autocorrelation coefficient / RDO / DCT strength can also be used and Figure 8 Shown in.
[0066] Regarding the correlation coefficient, the Pearson correlation coefficient (r) is the most common way to measure linear correlation. It is a value between -1 and 1 that measures the strength and direction of the relationship between two variables. This will give the relationship between frames, which will help predict frame bits. Then, the correlation coefficient strength is converted from the coefficient, such as Figure 6 As shown in block 608 of .
[0067] Regarding frame size, a frame can be divided into multiple parts, such as slices / tiles of CTU (128×128). The video quality is determined by allocating the correct number of bits by setting the correct QP.
[0068] Regarding feedback loops (such as Figure 5 The feedback loop periodically updates the available bits after each frame is coded and guides the ML module to choose the correct number of bits required for the frame.
[0069] Regarding autocorrelation coefficient / Rate-distortion optimization (RDO) / DCT, any of these three factors contributes to Figure 6 The QP is determined based on the strength of the autocorrelation coefficient, which is a robust method for finding the spectrum of the signal. Figure 7 ) can also help to use Figure 7The Lagrangian multipliers shown in [ 404 ] are used to find the frame bits, which can be advantageous because it consumes less computation time than using the autocorrelation technique. The DCT intensity (energy) of the frame is one technique for finding the number of bits required for the frame, and it is done on a block basis (64×64). Based on computational availability / processing time, at least one of the above may be necessary. The ML module receives four inputs, as described above, and predicts a target bit allocation (TBA(n)) to be sent to the QP module 406.
[0070] Now go to Figure 9 To understand the training of QP module 406. It should be understood that Figure 9 and Figure 5 The operation of is substantially the same as that of with the following exceptions: the QP module 406 is trained on the estimated (or target) frame bits from the frame bits module 404, the portion (such as a CTU or macroblock or other portion) DCT energy from the encoder 500 (before quantization), the portion (such as a CTU) bits 506, and the RDO cost 900.
[0071] Regarding the estimated frame bits (based on ML), the frame bits module 404 provides the best available bits to the QP module. Regarding the DCT energy, the DCT intensity (energy) of the block will help find the quantization parameter and it is on a block basis. The DCT energy is sent after the DCT calculation, so no additional calculation is required. The partial bits are in a feedback loop that is periodically updated after each partial and guides the QP module to choose the correct quantization parameter. RDO also helps the QP module find the quantization parameter using one or both of the expected quality metrics, namely the structural similarity index metric (SSIM) and the peak signal-to-noise ratio (PSNR).
[0072] It can now be understood that the overall process described herein includes generating data, organizing the data, evaluating correlations between variables, and selecting an appropriate machine learning module (in one implementation, a random forest). Random forest is a supervised ML algorithm for classification and regression problems. It builds a decision tree on different samples and takes a majority vote for classification and an average in the case of regression. The present technique can employ supervised regression machine learning techniques. The technique can be supervised because both features (the data of the video) and targets (QP) are available. During training, the features and targets are input to the ML model and it learns how to map the data to predictions. Furthermore, this is a regression task because the target values are continuous (in contrast to discrete classes in classification).
[0073] Another way to verify the quality of your data is to create a basic graph. Often, it is easier to spot anomalies in a graph than in numbers.
[0074] The exact steps for preparing the data will depend on the model used and the data collected, but any machine learning application will require some amount of data manipulation.
[0075] During training and testing, the data is split into training and test sets. During training, the model "sees" the answer—in this case, the actual quantized value—so it can learn how to predict the quantized value from the features. A relationship between all features and the target value is estimated, and the model learns this relationship during training. Then, when evaluating the model, it is tasked with making predictions on the test set, where the model only has access to the features, not the final answer. The predictions are compared to the ground truth answer for the QP to determine the model's accuracy.
[0076] An example ML algorithm can be based on decision tree technology. Decision trees are a classification and regression technique. They perform well at making decisions based on data by creating branches from a root (which is essentially the conditions present in the data) and providing outputs called "leaves."
[0077] An example rate control algorithm may use channel bit rate (bps) = C R Frame rate (fps) = F R , frame bit allocation average = FBA ave = (C R / F R ) (Equation 1);
[0078] Accumulated Bit Buffer(n) = ABBuf(n) = Frame coded bits (n) (i=0, n-1) (Equation 2);
[0079] Estimated bit allocation (n) = EBA (n) = [C R - ABBuf(n)] / [F R - (n-1)] (Equation 3).
[0080] As can now be appreciated, the present principles provide a technique for finding the correct quantization parameter value using an ML module. The present principles find the optimal QP value to enhance the video quality of each frame. This eliminates the need for a two-pass computational intensive approach, which is beneficial because such a two-pass technique is computationally expensive. The present principles also control the video rate buffer and maintain it at a constant level, which eliminates buffer overflow / underflow issues.
[0081] Now go to Figures 10 to 12 The flowchart of this paper is used to further illustrate the technology of this paper. Figure 10Beginning at block 1000 in
[0065] , video encoding information is input to a trained ML model. At block 1002, a QP for a frame (or portion thereof, such as a CTU) is received, which is input to an encoder at block 1004 to encode the video using the QP at block 1006. The video can be stored and / or transmitted to a player for presentation. At block 1008, the next frame (or portion of a frame) is retrieved, and the process loops back to block 1000.
[0082] Figure 11 The training of the frame bits module 404 is shown. Beginning at block 1100, the correlation coefficient, frame size, frame bits, and correlation strength are input to the module. Moving to block 1102, a reference true target bit rate is input, and at block 1104, the model is trained based on the training data to output an optimal target bit rate.
[0083] Figure 12 The training of the QP module is shown. Beginning at block 1200, the bit rate, CTU DCT energy (or other frame subset energy), CTU bits (or other frame subset bits), and RDO (or other similarity metric) are input to the module. At block 1202, the reference true QP is input, and at block 1204, the module is trained based on the training data.
[0084] While particular embodiments have been herein shown and described in detail, it is to be understood that the subject matter which is encompassed by the present invention is limited only by the claims.
Claims
1. A device comprising: at least one processor configured to: inputting at least one non-intra frame (I-frame), at least one reference video frame, and at least a bit representation of the video frame to at least one machine learning (ML) model; providing output from the ML model to at least one video encoder, the output comprising at least a quantization parameter (QP); as well as The video encoder is executed to output at least one video frame based at least in part on the QP.
2. The apparatus of claim 1, wherein the bit representation of a video frame is received from the video encoder.
3. The apparatus of claim 1 , wherein the at least one processor is configured to input the encoded bits of at least one block of an encoded video frame to the ML model.
4. The apparatus of claim 1 , wherein the at least one processor is configured to input an energy representation of a discrete cosine transform (DCT) of the at least one block of the encoded video frame to the ML model.
5. The apparatus of claim 1, wherein the at least one processor is configured to input the at least one non-I frame and the at least one reference video frame to the video encoder.
6. The apparatus of claim 1, wherein the at least one ML model comprises at least one random forest model.
7. The apparatus of claim 1 , wherein a first module of the ML model outputs an estimated bit representation of at least one video frame to at least a second module of the ML model.
8. The apparatus of claim 7 , wherein the first module is configured to receive the bit representations of the non-I frame, the at least one reference video frame, and the video frame, and the first module is configured to output the estimated bit representations to both the video encoder and the second module.
9. The apparatus of claim 8, wherein the second module is configured to receive coded bits of at least one block of an encoded video frame from the encoder and output the QP to the video encoder.
10. The apparatus of claim 9, wherein the second module is configured to receive an energy representation of a DCT of the at least one block of the encoded video frame from the video encoder.
11. A method comprising: providing at least one indication of a video transmission channel rate to at least a first machine learning (ML) module; providing an output of the first ML module to at least a second ML module, the output of the first ML module comprising at least an estimated number of bits for at least a portion of at least one video frame; providing an output of the second ML module to at least one video encoder, the output of the second ML module comprising at least one quantization parameter (QP); as well as The QP is used by the video encoder to generate at least one video frame for transmission over a video transmission channel.
12. A device comprising: At least one computer storage medium that is not a transient signal and that includes instructions executable by at least one processor to: inputting the video encoding information into at least one machine learning (ML) model; providing outputs of the ML model to at least one video encoder, the outputs comprising at least a quantization parameter (QP); as well as The video encoder is executed to generate video information using the QP.
13. The apparatus of claim 12, wherein the output comprises an estimated number of bits for at least one video frame.
14. The apparatus of claim 12, wherein the video encoding information comprises at least a representation of encoded bits of at least one video frame.
15. The apparatus of claim 12, wherein the video encoding information includes at least one non-intra frame (I frame).
16. The apparatus of claim 12, wherein the video encoding information comprises at least one reference video frame.
17. The apparatus of claim 12, wherein the video encoding information comprises coded bits of at least one block of a coded video frame.
18. The apparatus of claim 12, wherein the video encoding information comprises an energy representation of a discrete cosine transform (DCT) of the at least one block of the encoded video frame.
19. The apparatus of claim 13, wherein the ML model generates the QP based at least in part on the estimated number of bits of at least one video frame.
20. The apparatus of claim 12, wherein the video encoding information comprises a channel rate.