A video processing method, apparatus, and computer-readable storage medium

By introducing artificial intelligence-assisted programmable hardware into video codecs, and using AI algorithms to adjust encoding or decoding processing, the problems of slow video processing speed and high power consumption in the prior art are solved, and more efficient video processing is achieved.

CN113874916BActive Publication Date: 2025-06-10ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080039129.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-26
Filing Date
2020-05-06
Publication Date
2025-06-10
Estimated Expiration
2040-05-06

AI Technical Summary

Technical Problem

Existing CPU-based video codecs are slow and have high power consumption when processing videos, making it difficult to meet the needs of high-quality video transmission.

Method used

Using artificial intelligence-assisted programmable hardware video codec, through the coupling of the controller and the programmable hardware codec, AI algorithms are used to determine the non-pixel information of the video frame and adjust the encoding or decoding process.

Benefits of technology

It improves the speed and efficiency of video processing, reduces power consumption, and can process high-quality video data more efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113874916B_ABST
    Figure CN113874916B_ABST
Patent Text Reader

Abstract

The present application discloses a video processing method, apparatus, and computer-readable storage medium. According to some embodiments, the video processing apparatus includes a programmable hardware encoder configured to perform encoding processing on a plurality of input video frames. The video processing apparatus further includes a controller coupled to the programmable hardware encoder. The controller is configured to execute an instruction set to cause the video processing device to: determine first information of the plurality of input video frames and adjust the encoding processing based on the first information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 853,049, filed on May 26, 2019; the entire content of the above - mentioned document is incorporated herein by reference in its entirety. Technical field

[0003] The present invention mainly relates to video processing, particularly an artificial intelligence (AI) - assisted programmable hardware video codec. Background art

[0004] Modern video transmission systems constantly encode (e.g., compress) and decode (de - compress) video data. For example, to provide fast and high - quality cloud video streaming services, video encoders are used to compress digital video signals to reduce the transmission bandwidth consumption associated with such signals while maintaining image quality as much as possible. At the same time, user terminals receiving video streams can use video decoders to decompress the encoded video signals and then display the decompressed video images.

[0005] Video encoders and decoders (collectively referred to as "codecs") can be implemented in software or hardware. For example, a codec can be implemented as software running on one or more central processing units (CPUs). Many commercial applications (apps) use CPU - based codecs because the CPU - based codecs do not require a specific hardware environment and can be easily designed to play high - quality videos. However, CPU - based codecs are usually slow in operation and consume high power due to frequent memory access. Summary of the invention

[0006] Embodiments disclosed in this application relate to an artificial intelligence - assisted programmable hardware video codec. In some embodiments, a video processing device is provided. The video processing device includes a programmable hardware encoder configured to perform encoding processing on a plurality of input video frames. The video processing device further includes a controller coupled to the programmable hardware encoder. The controller is configured to execute a set of instruction sets to cause the video processing device to: determine first information of the plurality of input video frames and adjust the encoding processing according to the first information.

[0007] In some embodiments, a video processing device is provided. The video processing device includes a programmable hardware decoder configured to perform decoding processing on encoded video data to generate decoded data. The video processing device further includes a controller coupled to the programmable hardware decoder. The controller is configured to execute a set of instruction sets to cause the video processing device to: determine first information of the encoded video data; and adjust the decoding process based on the first information.

[0008] Each aspect of the disclosed embodiments may include a non - transitory, tangible computer - readable medium storing software instructions and executed by one or more processors. The computer - readable medium is configured and adapted to perform and execute one or more methods, operations, etc. consistent with the foregoing disclosure. In addition, aspects of the disclosed embodiments may be executed by one or more processors that are configured as dedicated processors based on software instructions written in logic and instructions, which, when processed, execute instructions for one or more operations consistent with the disclosed embodiments.

[0009] Additional objects and advantages of the disclosed embodiments will be set forth in part in the following description, in part will be obvious from the description, or may be learned by practice of the embodiments. The objects and advantages of the disclosed embodiments may be realized and obtained by the elements and combinations set forth in the claims.

[0010] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and do not limit the disclosed embodiments, as described. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 A schematic diagram showing an AI - assisted and hardware - based video processing system, consistent with an embodiment of the present invention.

[0012] Figure 2 A schematic diagram showing an exemplary hardware video encoder, consistent with an embodiment of the present application.

[0013] Figure 3 A schematic diagram showing an exemplary hardware video decoder, consistent with an embodiment of the present application.

[0014] Figure 4A A schematic diagram showing the interaction between a controller and a programmable hardware encoder, consistent with an embodiment of the present application.

[0015] Figure 4B A schematic diagram showing the interaction between a controller and a programmable hardware encoder, consistent with an embodiment of the present application.

[0016] Figure 5 A process flow diagram showing the process of using ROI region information to guide the encoding process, consistent with an embodiment of the present application.

[0017] Figure 6 A flow chart showing the process of using semantic segmentation to guide the encoding process, consistent with the disclosed embodiments of the present application.

[0018] Figure 7 A schematic diagram showing the process of estimating computational complexity in transcoding processing, consistent with the disclosed embodiments of the present application.

[0019] Figure 8 A schematic diagram showing the process of mapping coding parameters in transcoding processing, which is consistent with the embodiments disclosed in the present application.

[0020] Figure 9 A schematic diagram showing the process of enhancing pixels in transcoding processing, which is consistent with the embodiments disclosed in the present application.

[0021] Figure 10 A schematic diagram showing the process of mapping coding parameters in parallel coding processing, which is consistent with the embodiments disclosed in the present application.

[0022] Figure 11 A schematic diagram showing the process of performing video stabilization on decoded video data, which is consistent with the embodiments disclosed in the present application.

[0023] Figure 12 A schematic diagram showing an example AI controller architecture applicable to Figure 1 an artificial intelligence-assisted and hardware-based video processing system in, which is consistent with the embodiments disclosed in the present application.

[0024] Figure 13 A schematic diagram showing an exemplary hardware accelerator core architecture consistent with the embodiments of the present disclosure.

[0025] Figure 14 A schematic diagram showing an exemplary cloud system including a neural network processing architecture, which is consistent with the embodiments disclosed in the present application. Detailed Description of the Embodiments

[0026] Reference will now be made in detail to the exemplary embodiments, which are illustrated in the accompanying drawings. The following description refers to the drawings, unless otherwise specified, where the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with aspects related to the present invention described in the appended claims. Unless otherwise specifically stated, the word "or" includes all possible combinations, except in cases where it is not feasible. For example, if it is stated that a component may include A or B, then unless otherwise specifically stated or not feasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then unless otherwise specifically stated or not feasible, the component may include, A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0027] Video is a set of static pictures (or "frames") arranged in chronological order for storing visual information. Video capture devices (such as cameras) can be used to capture and store these pictures in a time sequence, while video playback devices (such as televisions, computers, smartphones, tablets, video players, or any user terminal with a display function) can be used to display these pictures in the said time sequence. In addition, in some applications, the video capture device can transmit the captured video in real time to a video playback device (such as a computer with a monitor) for monitoring, conferencing, or live streaming.

[0028] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is usually referred to as an "encoder", and the module for decompression is usually referred to as a "decoder". Encoders and decoders can be collectively referred to as "codecs". Encoders and decoders can be implemented by any suitable hardware, software, or a combination thereof. For example, the hardware implementation of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete component logic circuits, or any combination thereof. The software implementation of encoders and decoders can include program code, computer-executable instructions, firmware, or any algorithm suitable for computer implementation or a process fixed in a computer-readable medium. Video compression and decompression can be achieved through various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be referred to as a "transcoder".

[0029] The useful information of the image being encoded (referred to as the "current image") includes the changes relative to a reference image (e.g., a previously encoded and reconstructed image). These changes include changes in the position, brightness, or color of pixels, with the position change being the most concerned. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference image and the current image.

[0030] An image that does not reference another image (i.e., it is its own reference picture) is called an "I-picture" or "I-frame" (I-frame). A picture encoded using a previous picture as a reference picture is called a "P-picture" or "P-frame" (P-frame). A picture that uses both a previous picture and a subsequent picture as reference pictures (i.e., the reference is "bidirectional") is called a "B-picture" or "B-frame" (B-frame).

[0031] As described above, video processing is limited by the capacity of CPU-based codecs. To alleviate these problems, software-based codecs can be implemented on dedicated Graphics Processing Units (GPUs). GPUs can perform parallel computations and thus can render images faster than CPUs. However, the bitrate efficiency of GPUs is physically limited and may limit the video resolution that can be displayed.

[0032] Hardware-based codecs are dedicated hardware blocks designed specifically to perform specific video encoding and / or decoding processes. Hardware-based codecs generally consume less energy and have a higher processing speed than software-based codecs. However, the functionality of traditional hardware-based codecs cannot be reprogrammed to provide new features or adapt to new requirements.

[0033] Embodiments of the present disclosure provide improvements to traditional codec designs.

[0034] According to certain disclosed embodiments, Figure 1 A schematic diagram showing an AI-assisted and hardware-based video processing system 100 is shown. As Figure 1 shown, the system 100 includes a programmable hardware codec 110, which can be a hardware encoder, a hardware decoder, or a combination of both. The programmable hardware codec 110 can include any reconfigurable hardware components, such as Field Programmable Gate Arrays (FPGAs), Programmable Logic Devices (PLDs), etc. The circuits (such as logic gates) in the programmable hardware codec 110 can be programmed with software. Consistent with the disclosed embodiments, the programmable hardware codec 110 can perform video encoding or decoding processing on the input data 130 and generate output data 140. For example, the programmable hardware codec 110 can be implemented as an encoder to encode (e.g., compress) source video data and output the encoded (e.g., compressed) data. Additionally, the programmable hardware codec 110 can be implemented as a decoder to decode (e.g., decompress) the encoded video data and output the decoded (e.g., decompressed) video data for playback.

[0035] Still referring to Figure 1, the controller 120 is coupled to and communicates with the programmable hardware codec 110. The controller 120 can be a processor configured to execute program code (e.g., the AI algorithm 122) to analyze the input data 130 or output data 140 of the programmable hardware codec 110. The AI algorithm 122 may include, but is not limited to, machine learning algorithms, artificial neural networks, convolutional neural networks, deep neural networks, etc.

[0036] In some disclosed embodiments, the controller 120 can execute the AI algorithm 122 to determine encoding or decoding decisions 150 and program the programmable hardware codec 110 based on the encoding or decoding decisions 150. For example, the controller 120 can determine an encoding mode (such as inter-frame prediction or intra-frame prediction) or a motion vector according to the video data input to the hardware encoder, and use the encoding mode to guide the encoding process of the hardware encoder.

[0037] In some disclosed embodiments, the controller 120 can also execute the AI algorithm 122 to generate video analysis information 160 from the input data 130 or output data 140. For example, the controller 120 can determine the content of the decoded video data or track objects in the decoded video data. Another example is that the controller 120 can identify the ROI region in the input video frame of the hardware encoder, and plan the hardware encoder to use higher image quality for the ROI region and lower image quality for the non-ROI region.

[0038] Figure 2 A schematic diagram showing an exemplary hardware video encoder 200 consistent with the disclosed embodiments. For example, the video encoder 200 can be part of the programmable hardware codec 110 in the system 100 ( Figure 1 ). The video encoder 200 can perform intra-frame or inter-frame encoding on blocks in a video frame, including video blocks, or partitions or sub-partitions of video blocks. In a given video frame, intra-frame encoding can rely on spatial prediction to reduce or remove spatial redundancy in the video. Inter-frame encoding can rely on temporal prediction to reduce or eliminate temporal redundancy between adjacent frames in a video sequence. Intra-frame modes may be some spatial-based compression modes, while inter-frame modes (such as single prediction or dual prediction) may be some time-based compression modes.

[0039] As Figure 2As shown, the input video signal 202 can be block - processed. For example, a video block unit may be a 16×16 pixel block (e.g., a macroblock (MB)). In HEVC, an extended block size (e.g., a coding unit (CU)) can be used to compress video signals with higher resolutions, such as 108Gp or higher. In HEVC, a CU may include up to 64x64 luminance samples and corresponding chrominance samples. In VVC, the size of the CU can be further increased to include 128x128 luminance samples and corresponding chrominance samples. A CU can be divided into multiple prediction units (PUs), and different prediction methods can be applied to each unit. Each input video block (such as MB, CU, PU, etc.) can be processed using a spatial prediction unit 260 or a temporal prediction unit 262.

[0040] The spatial prediction unit 260 performs spatial prediction (e.g., intra - prediction) on the current CU using information on the same picture / slice that contains the current CU. Spatial prediction can use the pixels of adjacent blocks that have already been encoded in the same video image / slice to predict the current video block. Spatial prediction can reduce the spatial redundancy inherent in the video signal. Temporal prediction (e.g., intra - prediction or motion - compensated prediction) can use samples from the encoded video images to predict the current video block. Temporal prediction can reduce the temporal redundancy inherent in the video signal.

[0041] The temporal prediction unit 262 performs temporal prediction (e.g., inter-frame prediction) on the current CU using picture / segment information different from the picture / segment containing the current CU. The temporal prediction of a video block may be signaled by one or more motion vectors. The motion vectors may indicate the amount of motion and the direction of motion between the current block and one or more predicted blocks in a reference frame. If multiple reference pictures are supported, one or more reference picture indices may be sent for a video block. The one or more reference indices may be used to identify which reference image from the reference image buffer (RIB) or decoded picture buffer (DPB) 264 the reference image comes from and where the temporal prediction signal may come from. After spatial or temporal prediction, the mode decision and encoder control unit 280 in the encoder may select the prediction mode, e.g., based on a rate-distortion optimization method. The predicted block may be subtracted from the current video block at the adder 216. The prediction residual may be transformed by the transform unit 204 and quantized by the quantization unit 206. The quantized residual coefficients may be dequantized at the dequantization unit 210 and inverse-transformed at the inverse-transform unit 212 to form a reconstructed residual. The reconstructed block may be added to the predicted block at the adder 226 to form a reconstructed video block. The reconstructed video block may be loop-filtered, such as by a deblocking filter and an adaptive loop filter 266, before being placed in the reference image buffer 264 and used for encoding future video blocks. To form the output video bitstream 220, the coding mode (e.g., inter-frame or intra-frame), prediction mode information, motion information, quantized residual coefficients, etc. may be sent to the entropy coding unit 208 for compression and packetization to form the bitstream 220.

[0042] Consistent with the disclosed embodiments, the units of the video encoder 200 described above are implemented as hardware components, e.g., different circuit modules for performing their respective functions.

[0043] Figure 3 FIG. shows a schematic diagram of an exemplary hardware video decoder 300 consistent with the disclosed embodiments. For example, the video decoder 300 may be used as a programmable hardware codec 110 in the system 100 ( Figure 1 ). Refer to Figure 3 , the video bitstream 302 may be opened or entropy decoded in the entropy decoding unit 308, and the coding mode or prediction information may be sent to the spatial prediction unit 360 (e.g., if it is intra-frame coding) or the temporal prediction unit 362 (e.g., if it is inter-frame coding) to form a predicted block. If it is inter-frame coding, the prediction information may include the predicted block specification, one or more motion vectors (e.g., which may indicate the direction of motion and the amount of motion) or one or more reference indices (e.g., which may indicate from which reference image the prediction signal is obtained).

[0044] Motion compensation prediction may be performed by the temporal prediction unit 362 to form a temporal prediction block. The residual transform coefficients are sent to the inverse quantization unit 310 and the inverse transform unit 312 to reconstruct the residual block. The prediction block and the residual block may be added at 326. The reconstructed block may be loop-filtered (by the loop filter 366) before being stored in the reference picture buffer 364. The reconstructed video in the reference picture buffer 364 may be used to drive a display device or for predicting subsequent video blocks. The decoded video 320 may be displayed on a display.

[0045] In some disclosed embodiments, the units of the video decoder 300 described above are implemented as hardware components, e.g., different circuit blocks for performing their respective functions.

[0046] Consistent with the disclosed embodiments, the controller 120 ( Figure 1 ) may frequently access the encoding process performed by the programmable hardware encoder 200 ( Figure 2 ) and access the decoding process performed by the programmable hardware decoder 300 ( Figure 3 ). The controller 120 may determine non-pixel information from the encoding / decoding process and program the programmable hardware encoder 200 and the programmable hardware decoder 300 based on the non-pixel information.

[0047] Figure 4A FIG. shows a schematic diagram of the interaction between the controller 120 and the programmable hardware encoder 200 consistent with the disclosed embodiments. As Figure 4A shown, the controller 120 executes an artificial intelligence algorithm to analyze the source video data 132 input to the programmable hardware encoder 200 and determine non-pixel information (such as ROI regions, segmentation information, prediction modes, motion vectors, etc.) based on the source video data 132. The controller 120 may provide the non-pixel information as the encoder input 202 to the programmable hardware encoder 200 so that the programmable hardware encoder 200 can perform an encoding process according to the encoder input 202. In addition, the controller 120 may execute an artificial intelligence algorithm to extract non-pixel information (such as prediction modes and motion vectors used in the encoding process) as the encoder output 204 for guiding subsequent encoding or decoding processes.

[0048] Figure 4B FIG. is a schematic diagram showing the interaction between the controller 120 and the programmable hardware decoder 300 consistent with the embodiments disclosed in the present application. As Figure 4BAs shown, the controller 120 can provide the non-pixel information as a decoder input 302 (e.g., information about errors occurring in an encoded video frame) to the programmable hardware encoder 200, enabling the programmable hardware decoder 300 to perform a decoding process based on the decoder input 302. Additionally, the controller 120 can extract non-pixel information (e.g., prediction modes and motion vectors determined during the encoding process) as a decoder output 304 for guiding subsequent encoding or decoding processes.

[0049] Various embodiments for guiding encoding and decoding processes using artificial intelligence are described in detail below.

[0050] In some embodiments, the controller 120 can determine parameters for a video encoding process and then send the parameters to the programmable hardware encoder 200, which uses these parameters to perform the encoding process. Consistent with these embodiments, the programmable hardware encoder 200 is only allowed to make limited encoding decisions, while most encoding decisions are made by the controller 120 with the assistance of, for example, an artificial intelligence algorithm. The controller 120 provides encoding decisions as inputs to the encoder to guide the encoding process of the programmable hardware encoder 200.

[0051] For example, the controller 120 can determine an initial motion vector for an encoding block and determine a search range for the motion vector. Then, the controller 120 can input the determined parameters into the programmable hardware encoder 200, and the encoder uses the initial motion vector and the search range to perform the encoding process.

[0052] As another example, the controller 120 can receive an estimated motion vector or an estimated rate distortion from the programmable hardware encoder 200. The controller 120 can determine an optimized encoding mode or an optimized motion vector based on the received estimated motion vector or estimated rate distortion. The controller 120 sends the optimized encoding mode or the optimized motion vector to the programmable hardware encoder 200 to guide the encoding process.

[0053] As another example, the controller 120 can provide a coding tree unit (CTU) suggestion to the programmable hardware encoder 200. In particular, the controller 120 can determine a CTU structure and send information about the CTU structure to the programmable hardware encoder 200.

[0054] In some embodiments, the controller 120 may identify the ROI regions in the input video frames for the programmable hardware encoder 200. The ROI region information is then provided as an encoder input to the programmable hardware encoder 200, which is programmed to use a higher image quality (i.e., higher computational complexity) for the ROI regions and a lower image quality (i.e., lower computational complexity) for the non-ROI regions. The computational complexity determines the amount of CPU allocated for searching for motion patterns (e.g., inter-frame prediction or intra-frame prediction), motion vectors, code unit (CU) partitioning, transform unit partitioning, etc.

[0055] In accordance with the disclosed embodiments, Figure 5 is a process flow diagram of using the ROI region information to guide the encoding process. Referring to Figure 5 , the controller 120 receives a certain input video frame (step 502) and identifies whether there are any ROI regions in the input video frame (step 504). For the ROI regions, the controller 120 may program the programmable hardware encoder 200 to allocate more computational time for searching for encoding parameters (step 506). In contrast, for the non-ROI regions, the controller 120 may program the programmable hardware encoder 200 to allocate less computational time for searching for encoding parameters (step 508).

[0056] In some embodiments, the controller 120 may determine the semantic segmentation map of the input video frame. Then, the information about the segmentation is provided as an encoder input to the programmable hardware encoder 200, which may program to select the partition size or encoding mode according to the segmentation map. For example, the controller 120 may program the programmable hardware encoder 200 to use smaller CUs / TUs around the segmentation boundaries and larger CUs / TUs inside the segmentation (unless the segmentation is a non-rigid object). As another example, the controller 120 may program the programmable hardware encoder 200 to make the encoding modes or motion vectors within the same segmentation have higher correlation.

[0057] In accordance with the disclosed embodiments, Figure 6 is a process flow diagram of using semantic segmentation to guide the encoding process. As Figure 6 shown, the controller 120 receives the input video frame (step 602) and determines the semantic segmentation map of the input video frame. The controller 120 may send the information of the segmentation map to the programmable hardware encoder 200 and program the programmable hardware encoder 200 to identify the boundaries of the segmentation (step 604). The programmable hardware encoder 200 may use a smaller CU / TU partition size around the segmentation boundaries (step 606). In addition, the programmable hardware encoder 200 may predict the encoding mode of the encoding blocks within the segmentation based on the encoding modes used by other encoding blocks within the same segmentation (step 608).

[0058] In some embodiments, during transcoding processing, the controller 120 may extract non-pixel information from the decoder and use this non-pixel information to guide the encoding process in the programmable hardware encoder 200. For example, as Figure 7 shown, the controller 120 may estimate the computational complexity suitable for the transcoded image based on the previously encoded bitstream. Specifically, when the decoder is used to decode the previously encoded bitstream, the controller 120 receives from the decoder information indicating the number of bits used for encoding a video frame and the average quantization parameter. The controller 120 may estimate the image complexity (e.g., computational complexity) based on the number of bits or the average quantization parameter. Finally, the controller 120 programs the programmable hardware encoder 200 to perform rate control based on the image complexity.

[0059] In some embodiments, during transcoding processing, the controller 120 may map the encoding parameters for the previously encoded bitstream to the encoding parameters for the newly encoded bitstream, thereby reducing the time required to search for inter-frame or intra-frame prediction during transcoding processing. For example, as Figure 8 shown, when the decoder is used to decode the previously encoded bitstream, the controller 120 receives from the decoder information indicating at least one initial motion vector or initial encoding mode for the video frame. The controller 120 then adjusts the initial motion vector or initial encoding mode to a target motion vector or target encoding mode, respectively, according to the target format or definition of the transcoded video data. Finally, the controller 120 programs the programmable hardware encoder 200 to perform motion estimation or mode determination using the target motion vector or the target encoding mode.

[0060] In some embodiments, during transcoding processing, the controller 120 may use the partially decoded information to guide subsequent encoding. For example, as Figure 9 shown, when the decoder is used to decode the previously encoded bitstream, the controller 120 receives decoded information from the decoder, such as motion vectors or quantization parameters (step 902). The controller 120 may identify the ROI regions in the decoded video frame based on the decoded information (step 904). Then, during subsequent encoding for a new video format or video clarity, the controller 120 programs the programmable hardware encoder 200 to adaptively enhance the image based on the ROI regions. Specifically, the controller 120 may program the hardware encoder 200 to allocate more computational time to enhance the pixels in the ROI regions (step 906), while using less computational time to enhance the pixels in the non-ROI regions (step 908).

[0061] In some embodiments, when multiple encoders are used to perform parallel video encoding, the controller 120 can extract the common information of the parallel encoding and share it among the multiple encoders, thereby improving the encoding efficiency. For example, as Figure 10 shown, the controller 120 can receive the encoding parameters (e.g., motion vectors or encoding modes) used by the first programmable hardware encoder involved in the parallel video encoding process. Since the encoding parameters for encoding different bitstreams may be similar, the controller 120 can map the encoding parameters received from the first programmable hardware encoder to the applicable encoding parameters in the other programmable hardware encoders involved in the parallel video encoding process. Then, the controller 120 can send the mapped encoding parameters to the other programmable hardware encoders to guide their respective encoding processes.

[0062] In some embodiments, the controller 120 can use audio cues to guide the encoding process of the programmable hardware encoder 200. For example, the controller 120 can assign weights to the input video frames based on the audio information associated with the video frames. Then, the controller 120 determines an image quality proportional to the weights and programs the programmable hardware encoder 200 according to the determined video quality to encode the input video frames. For example, when the relevant audio cue is exciting, the controller 120 may choose to perform higher-quality compression.

[0063] In some embodiments, the controller 120 can determine a group of pictures (GOP) to guide the encoding process in the programmable hardware encoder 200. Specifically, the controller 120 can determine the first non-encoded I-frame among the multiple input video frames. For example, this initial I-frame can be identified by an artificial intelligence algorithm. Or, if transcoding processing is involved, the controller 120 can determine the internal CUs of the multiple input video frames according to the decoded information and determine the first non-encoded II-frame based on the internal CUs. After determining the first non-encoded I-frame, the controller 120 determines one or more additional non-encoded I-frames among the multiple input video frames. The controller 120 further indexes the first and additional non-encoded I-frames and determines the group of pictures based on the indexes of the non-encoded I-frames. Finally, the controller 120 programs the programmable hardware encoder 200 to encode the multiple input video frames using the determined group of pictures.

[0064] In some embodiments, the controller 120 may use artificial intelligence algorithms (such as reinforcement learning algorithms) to analyze the similarity between different input video frames of the programmable hardware encoder 200 and evaluate the bit budget based on the similarity. Specifically, the controller 120 may determine the similarity of a plurality of input video frames of the encoder. Then, the controller 120 may allocate a bit budget for each input video frame according to the similarity of the plurality of input video frames. The controller 120 further sends the bit budget to the programmable hardware encoder 200, which encodes the plurality of input video frames according to the corresponding bit budget.

[0065] In some embodiments, the controller 120 may use artificial intelligence algorithms (e.g., reinforcement learning algorithms) to analyze the similarity between different input video frames of the programmable hardware encoder 200 and determine reference frames based on the similarity. Specifically, the controller 120 may determine the similarity of a plurality of input video frames of the encoder. Then, based on the similarity of the plurality of input video frames, the controller 120 may determine one or more reference frames. The controller 120 may further send the information of the one or more reference frames to the programmable hardware encoder 200, which uses the one or more reference frames to encode the plurality of input video frames.

[0066] In some embodiments, the controller 120 may use ROI regions or segmentation information to define coding units or prediction units. Specifically, the controller 120 may generate a segmentation of the input video frame of the encoder. Then, the controller 120 may set at least one of the coding units or prediction units according to the segmentation of the input video frame. The controller 120 may further send the information of at least one of the coding units or prediction units to the programmable hardware encoder 120, which performs an encoding process using at least one of the coding units or prediction units.

[0067] In some embodiments, the controller 120 may use decoded information (i.e., decoder output) to achieve video stabilization. Since camera jitter can cause global motion of the image, video stabilization can be used to correct the effects brought by camera jitter. For example, as Figure 11As shown, the controller 120 may receive a plurality of motion vectors associated with a plurality of coded blocks in a coded frame from the programmable hardware decoder 300 (step 1102). The controller 120 may determine a global motion parameter for the coded frame based on the plurality of motion vectors (step 1104). If the global motion parameter indicates no global motion, the controller 120 infers that the image corresponding to the coded frame can be normally displayed (step 1106). If the global motion parameter indicates the presence of global motion, the controller 120 may further determine whether there is camera jitter in the decoded data (step 1108). If it is determined that there is camera jitter in the coded frame, the controller 120 may perform image stabilization on the decoded data based on the global motion parameter (step 1110). If it is determined that there is no camera jitter in the coded frame, the controller 120 infers that the image corresponding to the coded frame can be normally displayed (step 1106).

[0068] In some embodiments, the controller 120 may use the decoded information (i.e., decoder output) to track objects in the decoded video data or understand the content of the decoded video data. Specifically, the controller 120 may receive decoded information from the programmable hardware decoder 300, such as motion vectors, residuals, etc. Then, based on the decoded information, the controller may use artificial intelligence algorithms to identify and track the targets represented by the decoded video data. For example, the controller 120 may perform scene extraction, face filtering, attribute extraction, etc., to monitor the image content and create annotations for the image.

[0069] In some embodiments, the controller 120 may use the decoded information to guide error concealment in the decoding process. For example, the controller 120 may receive a plurality of motion vectors and coding modes associated with a plurality of coded blocks in a coded frame from the programmable hardware decoder 300. Then, the controller 120 may determine a method for concealing errors based on the plurality of motion vectors and coding modes. For the errors determined to exist in the coded frame, the controller 120 may further program the programmable hardware decoder 300 to perform error concealment according to the determined method for concealing errors in the decoding process.

[0070] Figure 12 A typical AI controller 1200 suitable for executing AI algorithms is shown, consistent with the embodiments of the present disclosure. For example, the AI controller 1200 may be configured as the controller 120( Figure 1 ) for performing the disclosed methods. In the context of the present disclosure, the AI controller 1200 may be implemented as a dedicated hardware accelerator for executing complex AI algorithms, such as machine learning algorithms, artificial neural networks (such as convolutional neural networks), or deep learning algorithms. In some embodiments, the AI controller 1200 may be referred to as a neural network processing unit (NPU). As Figure 12As shown, the AI controller 1200 may include multiple cores 1202, a command processor 1204, a direct memory access (DMA) unit 1208, a JTAG (Joint Test Action Group) / TAP (Test Access Port) controller 1210, a peripheral interface 1212, a bus 1214, etc.

[0071] It is worth noting that the core 1202 may perform algorithmic operations based on communication data. The core 1202 may include one or more processing elements, which may include a single instruction, multiple data (SIMD) architecture, the latter including one or more processing units configured to perform one or more operations (e.g., multiplication, complex multiplication, addition, multiply-accumulate, etc.) based on commands received from the command processor 1204. To perform operations on communication data packets, the core 1202 may include one or more processing elements for processing information in the data packets. Each processing element may contain any number of processing units. According to certain embodiments of the present invention, the AI controller 1200 may include multiple cores 1202, such as four cores. In certain embodiments, the multiple cores 1202 may be communicatively coupled to each other. For example, the multiple cores 1202 may be connected to a unidirectional ring bus, which supports efficient pipelining of large neural network models. The architecture of the core 1202 will be described in Figure 13 detail.

[0072] The command processor 1204 may interact with the host 1220 and transfer commands and data to the corresponding cores 1202. In certain embodiments, the command processor 1204 may interact with the host under the supervision of a kernel mode driver (KMD). In some embodiments, the command processor 1204 may modify the commands for each core 1202 so that the cores 1202 can work in parallel as much as possible. The modified commands may be stored in the instruction buffer. In certain embodiments, the command processor 1204 may be configured to coordinate one or more cores 1202 to achieve parallel execution.

[0073] The DMA unit 1208 can assist in transferring data between the host memory 1221 and the AI controller 1200. For example, the DMA unit 1208 can assist in loading data or instructions from the host memory 1221 into the local memory of the kernel 1202. The DMA unit 1208 can also assist in transferring data between multiple AI controllers. The DMA unit 1208 can allow off-chip devices to access on-chip and off-chip memories without causing the host CPU to interrupt. In addition, the DMA unit 1208 can assist in data transfer between components of the AI controller 1200. For example, the DMA unit 1208 can help transfer data between multiple kernels 1202 or within each kernel. Therefore, the DMA unit 1208 can also generate memory addresses and initiate memory read or write cycles. The DMA unit 1208 can also include several hardware registers that can be read and written by one or more processors, including a memory address register, a byte count register, one or more control registers, and other types of registers. These registers can specify the source, destination, transfer direction (reading from or writing to an input / output (I / O) device), the size of the transfer unit, or some combination of the number of bytes to be transferred in a string. It is worth noting that the AI controller 1200 can include a second DMA unit that can be used for data transfer between other AI controllers to allow multiple AI controllers to communicate directly without involving the host CPU.

[0074] The JTAG / TAP controller 1210 can specify a dedicated debug port to implement a serial communication interface (e.g., JTAG interface) for low-overhead access to the AI controller 1200 without the need for direct external access to the system address and data buses. The JTAG / TAP controller 1210 can also have an on-chip test access port interface (e.g., TAP interface), which implements a protocol to access a set of test registers that provide the chip logic level and device performance of different components.

[0075] The peripheral interface 1212 (such as a PCIe interface), if present, serves as (usually) an inter-chip bus to provide communication between the AI controller 1200 and other devices (such as the host system).

[0076] Bus 1214 (such as I 2The C bus (i.e., including the on-chip bus and the inter-chip bus). The on-chip bus connects all internal components to each other according to the requirements of the system backbone architecture. Although not all components are connected to each other, all components do have some connection to other components with which they need to communicate. The inter-chip bus connects the AI controller 1200 to other devices, such as off-chip memory or peripherals. For example, the bus 1214 can provide high-speed communication between cores and can also connect the core 1202 to other units, such as off-chip memory or peripherals. Typically, if there is a peripheral interface 1212 (e.g., an inter-chip bus), the bus 1214 is only related to the on-chip bus, although in some embodiments it may still be related to dedicated inter-bus communication.

[0077] The AI controller 1200 can also communicate with the host unit 1220. The host unit 1220 can be one or more processing units (e.g., an X86 central processing unit). As Figure 12 shown, the host unit 1220 may be associated with the host memory 1221. In some embodiments, the host memory 1221 can be an internal memory or an external memory associated with the host unit 1220. In some embodiments, the host memory 1221 can include a host disk, which is an external memory configured to provide additional storage for the host unit 1220. The host memory 1221 can be of the type of double data rate synchronous dynamic random access memory (e.g., DDR SDRAM). Compared with the on-chip storage built into the AI controller 1200 as a higher-level cache, the host memory 1221 can be configured to store a large amount of data at a slower access speed. The data stored in the host memory 1221 can be transferred to the AI controller 1200 for executing the neural network model.

[0078] In some embodiments, the host system having the host unit 1220 and the host memory 1221 can include a compiler (not shown in the figure). A compiler is a program or computer software that converts computer code written in a programming language into instructions for the AI controller 1200 to create an executable program. In machine learning applications, the compiler can perform various operations, such as preprocessing, lexical analysis, parsing, semantic analysis, converting the input program into an intermediate representation, initializing the neural network, code optimization, and code generation, or a combination of the above. For example, the compiler can compile a neural network to generate static parameters, such as the connections between neurons and the weights of neurons.

[0079] In some embodiments, a host system including a compiler may push one or more commands to the AI controller 1200. As described above, these commands may be further processed by the command processor 1204 of the AI controller 1200, temporarily stored in the instruction buffer of the AI controller 1200, and allocated to one or more corresponding cores (such as the core 1202 in Figure 12 ) or processing units. Some commands may instruct the DMA unit (e.g., the DMA unit 1208 in Figure 12 ) to load instructions and data from the host memory (e.g., the host memory 1221 in Figure 12 ) into the AI controller 1200. Then, the loaded instructions may be allocated to each core (such as the core 1202 in Figure 12 ) assigned with the corresponding tasks, and one or more cores may process these instructions.

[0080] It is worth noting that the first few instructions received by the core 1202 may instruct 1202 to load / store data from the host memory 1221 to one or more local memories of the core (e.g., the local memory 1332 in Figure 13 ). Then, each core 1202 may start an instruction pipeline, including fetching instructions from the instruction buffer (e.g., through a sequence generator), decoding the instructions (e.g., through the DMA unit 1208 in Figure 12 ), generating local memory addresses (e.g., corresponding to an operand), reading source data, performing or loading / storing operations, and then writing back the results.

[0081] According to certain embodiments, the AI controller 1200 may further include a global memory (not shown in the figure), which has memory blocks (such as 4 blocks of 8GB second-generation high-bandwidth memory (HBM2)) as the main memory (not shown in the figure). In some embodiments, the global memory may store instructions and data from the host memory 1221 through the DMA unit 1208. Then, the instructions may be allocated to the instruction buffers of each core assigned with the corresponding tasks, and the cores may process these instructions accordingly.

[0082] In some embodiments, the AI controller 1200 may further include a memory controller (not shown in the figure), which is configured to manage data reading and writing to specific memory blocks (such as HBM2) in the global memory. For example, the memory controller may manage read / write data from the cores of another AI controller (e.g., from the DMA unit corresponding to another AI controller) or from the core 1202 (e.g., from the local memory of the core 1202). It is worth noting that the AI controller 1200 may provide multiple memory controllers. For example, each memory block (such as HBM2) in the global memory may have a memory controller.

[0083] The memory controller can generate memory addresses and initiate memory read or write cycles. The memory controller can include several hardware registers that can be read and written by one or more processors. These registers can include a memory address register, a byte count register, one or more control registers, and other types of registers. These registers can specify a combination of source, destination, transfer direction (read from or written to an input / output (I / O) device), size of the transfer unit, number of bytes in a string, or other typical characteristics of the memory controller.

[0084] In the disclosed embodiments, Figure 12 the AI controller 1200 in can be used for various neural networks, such as convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), etc. Additionally, some embodiments can be configured for various processing architectures, such as neural network processing units (NPUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), tensor processing units (TPUs), application-specific integrated circuits (ASICs), any other type of heterogeneous acceleration processing units (HAPUs), or the like.

[0085] Figure 13 illustrates a typical core architecture consistent with embodiments of the present disclosure. As Figure 13 shown, the core 1202 can include one or more arithmetic units, such as first and second arithmetic units 1320 and 1322, a memory engine 1324, a sequence generator 1326, an instruction buffer 1328, a constant buffer 1330, a local memory 1332, and the like.

[0086] The first operation unit 1320 can be configured to perform operations on the received data (e.g., feature map). In some embodiments, the first arithmetic unit 1320 can include one or more arithmetic units configured to perform one or more operations (e.g., multiplication, complex multiplication, addition, multiply-accumulate, element-wise operations, etc.). In certain embodiments, the first arithmetic unit 1320 can be configured to accelerate the execution of convolutional operations or matrix multiplication operations.

[0087] The second operation unit 1322 can be configured to perform operations such as adjusting the scaling operation as described herein, ROI region operation, etc. In some embodiments, the second operation unit 1322 can include a scaling unit, a pooling data path, etc. In some embodiments, the second operation unit 1322 can be configured to cooperate with the first operation unit 1320 to scale the feature map as described herein. The disclosed embodiments are not limited to the embodiments in which the second operation unit 1322 performs scaling: in some embodiments, such scaling can be performed by the first operation unit 1320.

[0088] The memory engine 1324 can be configured to perform data copying within the corresponding kernel 1202 or between two kernels. The DMA unit 1208 can assist in copying data within the corresponding kernel or between two kernels. For example, the DMA unit 1208 can support the memory engine 1324 in copying data from local memory (such as Figure 13 the local memory 1332 therein) to the corresponding arithmetic unit. The memory engine 1324 can also be configured to perform matrix transposition to make the matrix suitable for use in the arithmetic unit.

[0089] The sequence generator 1326 can be coupled to the instruction buffer 1328 and configured to retrieve commands and distribute the commands to components of the kernel 1202. For example, the sequence generator 1326 can distribute convolution commands or multiplication commands to the first arithmetic unit 1320, pooling commands to the second arithmetic unit 1322, or data copy commands to the memory engine 1324. The sequence generator 1326 can also be configured to monitor the execution of neural network tasks and parallelize subtasks of the neural network tasks to improve execution efficiency. In some embodiments, the first arithmetic unit 1320, the second arithmetic unit 1322, and the storage engine 1324 can run in parallel under the control of the sequence generator 1326 according to the instructions stored in the instruction buffer 1328.

[0090] The instruction buffer 1328 can be configured to store instructions belonging to the corresponding kernel 1202. In some embodiments, the instruction buffer 1328 is coupled to the sequence generator 1326 and provides instructions to the sequence generator 1326. In some embodiments, the instructions stored in the instruction buffer 1328 can be transmitted or modified by the command processor 1204.

[0091] The constant buffer 1330 can be configured to store constant values. In some embodiments, the constant values stored in the constant buffer 1330 can be used by operation units - such as the first operation unit 1320 or the second operation unit 1322 - for operations such as batch normalization, quantization, dequantization, or similar operations.

[0092] Local memory 1332 can provide fast read and write storage space. To reduce possible interactions with global memory, the storage space of local memory 1332 can be implemented with a large capacity. With such a capacity, most data accesses can be performed within the kernel 1202, reducing the latency caused by data access. In some embodiments, to minimize data loading latency and energy consumption, SRAM (Static Random Access Memory) integrated on the chip can be used as local memory 1332. In some embodiments, the capacity of local memory 1332 can reach 192MB or higher. According to some embodiments of the present disclosure, local memory 1332 is evenly distributed on the chip to alleviate dense wiring and heat generation problems

[0093] The AI computing architecture of the video processing system disclosed in the present invention is not limited to the AI controller 1200 architecture described above. Consistent with the disclosed embodiments, the artificial intelligence algorithm (e.g., artificial neural network) can reside on various electronic systems. For example, the artificial intelligence algorithm can be hosted on a server, one or more nodes in a certain data center, a desktop computer, a laptop computer, a tablet computer, a smart phone, a wearable device such as a smart watch, an embedded device, an Internet of Things device, a smart device, a sensor, an orbiting satellite, or any other electronic device with computing capabilities. In addition, the way the AI algorithm is hosted in a specific device may also vary. For example, in some embodiments, the AI algorithm may be hosted and run on the general-purpose processing unit of the device, such as a central processing unit (CPU), a graphics processing unit (GPU), or a general-purpose graphics processing unit (GPGPU). In other embodiments, the artificial neural network can be hosted and run on the hardware accelerator of the device. Such as a neural processing unit (NPU), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC).

[0094] In addition, the artificial intelligence computing of the video processing system disclosed in the present invention can also be implemented in the form of cloud computing. Figure 14 A schematic diagram showing a typical cloud system including the AI controller 1200 ( Figure 12 ), which is consistent with the embodiments of the present invention. As Figure 14 shown, the cloud system 1430 can provide cloud services with artificial intelligence (AI) capabilities and can include multiple computing servers (such as 1432, 1434). For example, in some embodiments, the computing server 1432 can include Figure 12 the AI controller 1200. As Figure 14 shown, the AI controller 1200 is shown in a simple and clear manner.

[0095] With the help of the AI controller 1200, the cloud system 1430 can provide extended AI capabilities such as image recognition, face recognition, translation, 3D modeling, etc. It should be noted that the AI controller 1200 can be deployed on computing devices in other forms. For example, the AI controller 1200 can also be integrated into computing devices such as smartphones, tablets, wearable devices, etc.

[0096] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) for performing the above method. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any punched physical medium, RAM, PROM, EPROM, FLASH-EPROM or other flash memories, NVRAM, caches, registers, any other storage chip or cartridge, and the same networked versions. The device may include one or more processors (CPUs), an input / output interface, a network interface, and / or a memory.

[0097] It should be noted that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When the software is executed by a processor, it can execute the disclosed method. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the multiple modules / units described above can be combined into one module / unit, and each module / unit can be further divided into multiple sub-modules / sub-units.

[0098] The embodiments can be further described using the following terms:

[0099] 1. A video processing device, comprising:

[0100] A programmable hardware encoder configured to perform encoding processing on a plurality of input video frames; and

[0101] A controller coupled to the programmable hardware encoder, the controller being configured to execute an instruction set to cause the video processing device to:

[0102] Determine first information of the plurality of input video frames; and

[0103] Adjust the encoding processing according to the first information.

[0104] 2. The video processing device according to item 1, wherein the first information includes non-pixel information of the plurality of input video frames.

[0105] 3. A video processing device according to any one of Article 1 and Article 2, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0106] Determine an initial motion vector of an encoding block;

[0107] Determine a search range of the motion vector; and

[0108] Send the initial motion vector and the search range to the programmable hardware encoder, wherein the programmable hardware encoder is configured to perform the encoding process using the initial motion vector and the search range.

[0109] 4. The video processing device according to any one of Articles 1-3, wherein the instruction set includes a machine learning algorithm for determining the non-pixel information.

[0110] 5. The video processing device according to any one of Articles 1-4, wherein the first information includes

[0111] ROI region information in a certain input video frame among the multiple input video frames, and the controller is configured to execute the instruction set to cause the video processing device to:

[0112] Identify one or more ROI regions in the input video frame, and

[0113] Configure the programmable hardware encoder to encode the ROI region with a first computational complexity, and encode the non-ROI region with a second computational complexity different from the first computational complexity.

[0114] 6. The video processing device according to any one of Articles 1-5, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0115] Generate a segmentation of the input video frame among the multiple input video frames, and

[0116] Send the information of the segmentation to the programmable hardware encoder, and the programmable hardware encoder selects at least one of a partition size or an encoding mode for the encoding block based on the segmentation.

[0117] 7. The video processing device according to Article 6, wherein the programmable hardware encoder is configured to use a first partition size at the boundary of the segmentation and a second partition size within the segmentation, and the first partition size is smaller than the second partition size.

[0118] 8. A video processing system according to any one of Articles 6 and 7, wherein the programmable hardware encoder is configured to determine the coding mode of the coding blocks within a segment based on one or more coding modes used by other coding blocks within the same segment.

[0119] 9. The video processing apparatus according to any one of Articles 1 - 8, wherein the plurality of input video frames are generated by a decoder in a video transcoding process, and the controller is configured to execute the instruction set to cause the video processing apparatus to:

[0120] Receive information from the decoder, the information showing at least one of the number of bits or the average quantization parameter of a certain input video frame among the plurality of input video frames;

[0121] Determine the computational complexity of the input video frame based on at least one of the number of bits or the average quantization parameter; and

[0122] Adjust the bit rate of the encoding process according to the computational complexity.

[0123] 10. The video processing apparatus according to any one of Articles 1 - 9, wherein the plurality of input video frames are generated by a decoder in a video transcoding process, and the controller is configured to execute the instruction set to cause the video processing apparatus to generate:

[0124] Receive information from the decoder, the information showing at least one of the original motion vector or the original coding mode of a certain input video frame among the plurality of input video frames;

[0125] Determine at least one of the target motion vector or the target coding mode respectively based on at least one of the initial original motion vector or the initial coding mode; and

[0126] Perform the motion estimation and mode decision of the encoding process based on at least one of the target motion vector or the target coding mode.

[0127] 11. The video processing apparatus according to any one of Articles 1 - 10, wherein the programmable hardware encoder is a first encoder, and the controller is configured to execute the instruction set to cause the video processing device to:

[0128] Receive the coding parameters used by a second encoder for coding one or more of the plurality of input video frames; and

[0129] Adjust the coding parameters and send the adjusted coding parameters to the programmable hardware encoder, and the programmable hardware encoder is configured to perform the encoding process based on the adjusted coding parameters.

[0130] 12. The video processing apparatus according to claim 11, wherein the encoding parameter includes at least one of a motion vector or an encoding mode.

[0131] 13. The video processing apparatus according to any one of claims 1-12, wherein the plurality of input video frames are generated by a decoder in a video transcoding process, and the controller is configured to execute the instruction set to cause the video processing apparatus to:

[0132] Receive information from the decoder, the information showing an encoding parameter for a certain input video frame among the plurality of input video frames;

[0133] Based on the encoding parameter, identify an ROI region in the input video frame; and

[0134] Configure a programmable hardware encoder to improve the image quality within the ROI region.

[0135] 14. The video processing apparatus according to claim 13, wherein the encoding parameter includes at least one of a motion vector or a quantization parameter.

[0136] 15. The video processing apparatus according to any one of claims 1-14, wherein the controller is configured to execute the instruction set to cause the video processing apparatus to:

[0137] Assign a weight to a certain input video frame among the plurality of input video frames based on audio information associated with the input video frame;

[0138] Determine the image quality in a manner proportional to the weight; and

[0139] Configure the programmable hardware encoder to encode the input video frame according to the determined image quality.

[0140] 16. The video processing device according to any one of claims 1-15, wherein the controller is configured to execute a set of instructions to cause the video processing apparatus to:

[0141] Receive at least one of a motion vector or estimated rate distortion information from the programmable hardware encoder;

[0142] Based on the received motion vector or the estimated rate distortion information, determine at least one of an encoding mode or a motion vector; and

[0143] Send at least one of the encoding mode or the motion vector to the programmable hardware encoder, the programmable hardware encoder being configured to perform the encoding process according to at least one of the encoding mode or the motion vector.

[0144] 17. A video processing device according to any one of clauses 1 to 16, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0145] Determine the coding tree unit structure of a certain input video frame among a plurality of input video frames; and

[0146] Send information showing the coding tree unit structure to a programmable hardware encoder, wherein the programmable hardware encoder is configured to partition the input video frame according to the coding tree unit structure.

[0147] 18. A video processing device according to any one of Articles 1 to 17, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0148] Determine the first uncoded I-frame among the plurality of input video frames;

[0149] Determine one or more additional uncoded I-frames among the plurality of input video frames;

[0150] Index the first and additional uncoded I-frames;

[0151] Determine the picture group according to the indexes of the first and additional uncoded I-frames; and

[0152] Configure the programmable hardware encoder to encode the plurality of input video frames according to the determined picture group.

[0153] 19. The video processing device according to Article 18, wherein the plurality of input video frames are generated by a decoder in a video transcoding process, and the controller is configured to execute the instruction set to cause the video processing device to:

[0154] Receive information from the decoder. This information shows the internal coding units for encoding the plurality of input video frames; and

[0155] Determine the first uncoded I-frame based on the internal coding units.

[0156] 20. The video processing device according to any one of Articles 18 and 19, wherein the instruction set includes a machine learning algorithm for determining the first uncoded I-frame.

[0157] 21. A video processing device according to any one of Articles 1 to 20, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0158] Determine the similarity of the plurality of input video frames;

[0159] Allocating a bit budget for each of the multiple input video frames based on the similarity of the multiple input video frames; and

[0160] Sending the bit budget to the programmable hardware encoder, which is configured to encode the multiple input video frames according to the corresponding bit budget.

[0161] 22. The video processing device according to Article 21, wherein the instruction set includes a reinforcement learning algorithm for determining the similarity of the multiple input video frames.

[0162] 23. The video processing device according to any one of Articles 1-22, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0163] Determine the similarity of the multiple input video frames;

[0164] Determine one or more reference frames based on the similarity of the multiple input video frames; and

[0165] Send information of the one or more reference frames to the programmable hardware encoder, which is configured to encode the multiple input video frames using the one or more reference frames.

[0166] 24. The video processing device according to Article 23, wherein the instruction set includes a machine learning algorithm for determining the similarity of the multiple input video frames.

[0167] 25. The video processing device according to any one of Articles 1-24, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0168] Identify one or more ROI regions in a certain input video frame of the multiple input video frames;

[0169] Set at least one of an encoding unit or a prediction unit based on the one or more ROI regions; and

[0170] Send information of at least one of the encoding unit or the prediction unit to the programmable hardware encoder, which is configured to perform the encoding process using at least one of the encoding unit or the prediction unit.

[0171] 26. The video processing device according to any one of Articles 1-25, wherein the controller is configured to execute the instruction set to cause the video processing device to:

[0172] Generate a segmentation of a certain input video frame of the multiple input video frames;

[0173] Based on the segmentation of the input video frame, set at least one of the coding unit or the prediction unit; and

[0174] Send information of at least one of the coding unit or the prediction unit to a programmable hardware encoder, where the programmable hardware encoder is configured to perform the encoding process using at least one of the coding unit or the prediction unit.

[0175] 27. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors coupled to a programmable hardware encoder, wherein execution of the set of instructions causes the programmable hardware encoder to:[[]]

[0176] Determine first information of a plurality of input video frames; and

[0177] Based on the first information, adjust the encoding process performed by the programmable hardware encoder on the plurality of input video frames.

[0178] 28. A computer-implemented method, comprising:[[]]

[0179] Perform an encoding process on a plurality of input video frames by a programmable hardware encoder;

[0180] Determine first information of the plurality of input video frames by a controller coupled to the programmable hardware encoder; and

[0181] The controller adjusts the encoding process according to the first information.

[0182] 29. A video processing device, comprising:

[0183] A programmable hardware decoder configured to perform a decoding process on encoded video data to generate decoded data; and

[0184] A controller coupled to the programmable hardware decoder, the controller being configured to execute a set of instructions to cause the video processing device to:[[]]

[0185] Determine first information of the encoded video data; and

[0186] Adjust the decoding process according to the first information

[0187] 30. The video processing device according to Article 29, wherein the first information includes non-pixel information of the encoded video data.

[0188] 31. The video processing device according to any one of Articles 29 and 30, wherein the set of instructions includes a machine learning algorithm for determining non-pixel information.

[0189] 32. A video processing device according to any one of Articles 29 - 31, wherein the controller is configured to execute an instruction set to cause the video processing device to:

[0190] Receive a plurality of motion vectors associated with a plurality of encoded blocks in an encoded frame from a programmable hardware decoder;

[0191] Based on the plurality of motion vectors, determine global motion parameters of the encoded frame;

[0192] Determine whether there is camera jitter in the decoded data;

[0193] When it is determined that there is camera jitter in the encoded frame, perform image stabilization on the decoded data according to the global motion parameters.

[0194] 33. A video processing device according to any one of Articles 29 - 32, wherein the controller is configured to execute a set of instructions to cause the video processing device to:

[0195] Receive a plurality of encoding parameters for encoding video data from a programmable hardware decoder; and

[0196] Execute a machine learning algorithm based on the plurality of encoding parameters to identify and track an object represented by the decoded data.

[0197] 34. The video processing device according to Article 33, wherein the controller is configured to execute an instruction set to cause the video processing device to:

[0198] Detect a plurality of attributes of an object through machine learning; and

[0199] Identify an event of the object through machine learning based on the plurality of attributes of the object.

[0200] 35. A video processing device according to any one of Articles 29 - 34, wherein the controller is configured to execute an instruction set to cause the video processing device to:

[0201] Receive a plurality of motion vectors and encoding modes associated with a plurality of encoded blocks in a certain encoded frame from a programmable hardware decoder; determine an error concealment method based on the plurality of motion vectors and encoding modes; and, for an error determined to exist in the encoded frame, perform error concealment during decoding according to the determined method.

[0202] In addition to implementing the above - mentioned methods using computer - readable program code, the above - mentioned methods can also be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be regarded as a hardware component, and the devices contained and configured in the controller to implement various functions can also be regarded as the structure inside the hardware component. Alternatively, the devices configured to implement various functions can even be regarded as software modules configured to implement the methods and the structure inside the hardware component.

[0203] This disclosure can be described in the general context of computer - executable instructions executed by a computer (such as program modules). Generally, program modules include routines, programs, objects, assemblies, data structures, classes, or the like for performing specific tasks or implementing specific abstract data types. Embodiments of the present disclosure can also be implemented in a distributed computing environment. In a distributed computing environment, tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0204] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, words such as "including", "having", "containing", and "comprising" and other similar forms are intended to be equivalent in meaning and open - ended, and any one of these words does not mean that it is an exhaustive list of items or plural items, or that it is limited to the above - mentioned items or plural items.

[0205] In the foregoing specification, embodiments have been described with reference to numerous specific details, which may vary according to implementation. Certain adaptations and modifications can be made to the described embodiments. From the considerations of the specification and practice disclosed herein, other embodiments will be apparent to those skilled in the art. The present specification and examples are only regarded as examples, and the following claims indicate the true scope and spirit of the disclosure. The order of steps shown in the figures is also for illustrative purposes only and is not intended to be limited to any specific order of steps. Similarly, those skilled in the art can understand that these steps can be executed in a different order when performing the same method.

Claims

1. A video processing device, comprising: a programmable hardware encoder configured to perform encoding processing on a plurality of input video frames; and a controller coupled to the programmable hardware encoder, the controller being configured to execute an instruction set to cause the video processing device to: determine first information of the plurality of input video frames, the first information including non-pixel information of the plurality of input video frames, the non-pixel information including at least one of segmentation information, prediction mode, or motion vector; and adjust the encoding processing according to the first information; wherein the programmable hardware encoder is a first encoder, and the controller is configured to execute the instruction set to cause the video processing device to: receive encoding parameters used by a second encoder for encoding one or more of the plurality of input video frames; and adjust the encoding parameters and send the adjusted encoding parameters to the programmable hardware encoder, the programmable hardware encoder being configured to perform the encoding processing based on the adjusted encoding parameters.

2. The video processing device according to claim 1, wherein the controller is configured to execute the instruction set to cause the video processing device to: determine an initial motion vector of an encoding block; determine a search range of the motion vector; and send the initial motion vector and the search range to the programmable hardware encoder, wherein the programmable hardware encoder is configured to perform the encoding processing using the initial motion vector and the search range.

3. The video processing device according to claim 1, wherein the instruction set includes a machine learning algorithm for determining the non-pixel information.

4. The video processing device according to claim 1, wherein the first information includes ROI region information in a certain input video frame among the plurality of input video frames, and the controller is configured to execute the instruction set to cause the video processing device to: identify one or more ROI regions in the input video frame, and configure the programmable hardware encoder to encode the ROI region with a first computational complexity and encode the non-ROI region with a second computational complexity different from the first computational complexity.

5. The video processing device according to claim 1, wherein the controller is configured to execute the instruction set to cause the video processing device to: generate a segmentation of a certain input video frame among the plurality of input video frames, and send information of the segmentation to the programmable hardware encoder, and the programmable hardware encoder selects at least one of a partition size or an encoding mode for an encoding block based on the segmentation.

6. The video processing device according to claim 5, wherein the programmable hardware encoder is configured to use a first partition size at the boundary of the segmentation and use a second partition size within the segmentation, the first partition size being smaller than the second partition size.

7. The video processing device according to claim 5, wherein The programmable hardware encoder is configured to determine the coding mode of a coding block within a partition based on one or more coding modes used by other coding blocks within the same partition.

8. The video processing apparatus according to claim 1, wherein, the plurality of input video frames are generated by a decoder during video transcoding processing, and the controller is configured to execute the instruction set to cause the video processing apparatus to: receive information from the decoder that indicates at least one of the number of bits or the average quantization parameter of a certain input video frame among the plurality of input video frames; determine the computational complexity of the input video frame based on at least one of the number of bits or the average quantization parameter; and adjust the bit rate of the coding processing according to the computational complexity.

9. The video processing apparatus according to claim 1, wherein, the plurality of input video frames are generated by a decoder during video transcoding processing, and the controller is configured to execute the instruction set to cause the video processing apparatus to: receive information from the decoder that indicates at least one of the initial motion vector or the initial coding mode of a certain input video frame among the plurality of input video frames; determine at least one of a target motion vector or a target coding mode respectively based on at least one of the initial motion vector or the initial coding mode; and perform the motion estimation and mode decision of the coding processing based on at least one of the target motion vector or the target coding mode.

10. The video processing apparatus according to claim 1, wherein, the coding parameter includes at least one of a motion vector or a coding mode.

11. The video processing apparatus according to claim 1, wherein, the plurality of input video frames are generated by a decoder during video transcoding, and the controller is configured to execute the instruction set to cause the video processing apparatus to: receive information from the decoder that indicates the coding parameter for a certain input video frame among the plurality of input video frames; identify the ROI region in the input video frame based on the coding parameter; and configure the programmable hardware encoder to improve the image quality within the ROI region.

12. The video processing apparatus according to claim 11, wherein, the coding parameter includes at least one of a motion vector or a quantization parameter.

13. The video processing apparatus according to claim 1, wherein, the controller is configured to execute the instruction set to cause the video processing apparatus to: assign a weight to a certain input video frame among the plurality of input video frames based on the audio information associated with the input video frame; determine the image quality in a manner proportional to the weight; and configure the programmable hardware encoder to encode the input video frame according to the determined image quality.

14. The video processing apparatus according to claim 1, wherein, the controller is configured to execute the instruction set to cause the video processing apparatus to: receive at least one of a motion vector or estimated rate distortion information from the programmable hardware encoder; determine at least one of a coding mode or a motion vector based on the received motion vector or the estimated rate distortion information; and Send at least one of the coding mode or the motion vector to the programmable hardware encoder, which is configured to perform the encoding process according to at least one of the coding mode or the motion vector.

15. The video processing device according to claim 1, wherein, the controller is configured to execute the instruction set to cause the video processing device to: determine the coding tree unit structure of a certain input video frame among a plurality of input video frames; and send information indicating the coding tree unit structure to the programmable hardware encoder, wherein the programmable hardware encoder is configured to partition the input video frame according to the coding tree unit structure.

16. The video processing device according to claim 1, wherein the controller is configured to execute the instruction set to cause the video processing device to: determine a first uncoded I-frame among the plurality of input video frames; determine one or more additional uncoded I-frames among the plurality of input video frames; index the first and additional uncoded I-frames; determine a picture group according to the indexes of the first and additional uncoded I-frames; and configure the programmable hardware encoder to encode the plurality of input video frames according to the determined picture group.

17. The video processing device according to claim 16, wherein, the plurality of input video frames are generated by a decoder in a video transcoding process, and the controller is configured to execute the instruction set to cause the video processing device to: receive information from the decoder, the information indicating internal coding units for encoding the plurality of input video frames; and determine the first uncoded I-frame based on the internal coding units.

18. The video processing device according to claim 16, wherein, the instruction set includes a machine learning algorithm for determining the first uncoded I-frame.

19. The video processing device according to claim 1, wherein, the controller is configured to execute the instruction set to cause the video processing device to: determine the similarity of the plurality of input video frames; allocate a bit budget for each of the plurality of input video frames based on the similarity of the plurality of input video frames; and send the bit budget to the programmable hardware encoder, which is configured to encode the plurality of input video frames according to the corresponding bit budget.

20. The video processing device according to claim 19, wherein, the instruction set includes a reinforcement learning algorithm for determining the similarity of the plurality of input video frames.

21. The video processing device according to claim 1, wherein, the controller is configured to execute the instruction set to cause the video processing device to: determine the similarity of the plurality of input video frames; determine one or more reference frames based on the similarity of the plurality of input video frames; and send information of the one or more reference frames to the programmable hardware encoder, which is configured to encode the plurality of input video frames using the one or more reference frames.

22. The video processing device according to claim 21, wherein, The instruction set described above includes a machine learning algorithm for determining the similarity of the multiple input video frames.

23. The video processing apparatus according to claim 1, wherein, the controller is configured to execute the instruction set to cause the video processing apparatus to: identify one or more ROI regions in a certain input video frame of the multiple input video frames; set at least one of an encoding unit or a prediction unit based on the one or more ROI regions; and send information of at least one of the encoding unit or the prediction unit to the programmable hardware encoder, and the programmable hardware encoder is configured to perform the encoding process using at least one of the encoding unit or the prediction unit.

24. The video processing apparatus according to claim 1, wherein the controller is configured to execute the instruction set to cause the video processing apparatus to: generate a segmentation of a certain input video frame of the multiple input video frames; set at least one of an encoding unit or a prediction unit based on the segmentation of the input video frame; and send information of at least one of the encoding unit or the prediction unit to a programmable hardware encoder, where the programmable hardware encoder is configured to perform the encoding process using at least one of the encoding unit or the prediction unit.

25. A non-transitory computer-readable storage medium storing a set of instruction sets executable by one or more processors coupled to a programmable hardware encoder, wherein, execution of the instruction set causes the programmable hardware encoder to: determine first information of multiple input video frames, the first information including non-pixel information of the multiple input video frames, and the non-pixel information including at least one of segmentation information, a prediction mode, or a motion vector; and adjust the encoding process performed by the programmable hardware encoder on the multiple input video frames based on the first information; wherein the programmable hardware encoder is a first encoder, and the controller is configured to execute the instruction set to cause the video processing device to: receive encoding parameters used by a second encoder for encoding one or more of the multiple input video frames; and adjust the encoding parameters and send the adjusted encoding parameters to the programmable hardware encoder, and the programmable hardware encoder is configured to perform the encoding process based on the adjusted encoding parameters.

26. A computer-implemented method, comprising: performing an encoding process on multiple input video frames by a programmable hardware encoder; determining, by a controller coupled to the programmable hardware encoder, first information of the multiple input video frames, the first information including non-pixel information of the multiple input video frames, and the non-pixel information including at least one of segmentation information, a prediction mode, or a motion vector; and the controller adjusting the encoding process according to the first information; wherein the programmable hardware encoder is a first encoder, and the controller is configured to execute an instruction set to cause the video processing device to: receive encoding parameters used by a second encoder for encoding one or more of the multiple input video frames; and Adjust the encoding parameters and send the adjusted encoding parameters to the programmable hardware encoder, which is configured to perform the encoding process based on the adjusted encoding parameters.

Citation Information

Patent Citations

  • Analytics-modulated coding of surveillance video

    US20180139456A1

  • Machine learning video processing systems and methods

    US20190075301A1