Video coding method and corresponding device
Patent Information
- Application Number
- CN202311849660.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
Smart Images

Figure CN120238659A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of coding technologies, and particularly relates to a method and corresponding apparatus for video coding. Background Art
[0002] Video coding is a compression technology. Through video coding, a video file in one format can be converted into a video file in another format. For example: through video coding, a video file in YUV format can be converted into a video file in H.264 format or other formats, thereby achieving video compression.
[0003] During video coding, the encoder encodes each video frame in the video sequence frame by frame according to the global quantization parameter. Since video frames are continuous, if the correlation between video frames is not considered during video coding, it will cause improper use of coding resources. Summary of the Invention
[0004] This application provides a method for video coding, which is used to improve the effectiveness of coding resource allocation during the video coding process. This application also provides a corresponding apparatus, a computer-readable storage medium, and a computer program product, etc.
[0005] In a first aspect of this application, a method for video coding is provided. When encoding a video sequence, the method includes: predicting at least one target region in a pre-encoding frame according to the motion information of a first moving object in an encoded frame, where the at least one target region is related to the first moving object; encoding the video data corresponding to the at least one target region using a target quantization parameter, and the target quantization parameter is different from the global quantization parameter of the video sequence.
[0006] In this application, the video sequence may be a YUV sequence. The encoded frame and the pre-encoding frame may be two consecutive video frames, or may be two video frames separated by one or more video frames.
[0007] In this application, the pre-encoding frame refers to a frame to be encoded or a frame ready to be encoded, usually referring to the next frame that is about to be encoded immediately after the previous frame is encoded.
[0008] In this application, the first moving object may be one or more moving objects in the encoded frame. At least one target region in the pre-encoding frame is part or all of the region corresponding to the pre-encoding frame.
[0009] In this application, the global quantization parameter is the quantization parameter (QP) configured for encoding a video sequence. In the prior art, each video frame in the video sequence needs to be encoded according to this global quantization parameter. However, in this application, for a target region related to a first moving object, a target quantization parameter different from the global quantization parameter is used for encoding. This target quantization parameter can be greater than the global quantization parameter or less than the global quantization parameter. Moreover, when there are multiple target regions, the target quantization parameter corresponding to some target regions can be greater than the global quantization parameter, and the target quantization parameter corresponding to some target regions can be less than the global quantization parameter. During encoding, more encoding resources can be allocated to the target region with a smaller quantization parameter, and fewer encoding resources can be allocated to the target region with a larger quantization parameter. In this way, the effectiveness of encoding resource allocation during video encoding can be improved, and the efficiency and quality of video encoding can be enhanced.
[0010] In a possible implementation, when there is one target region, the target region is the first region or the second region; when there are multiple target regions, the multiple target regions include the first region and the second region; wherein, the first region is the motion region of the first moving object in the pre-coded frame, and the target quantization parameter corresponding to the first region is less than the global quantization parameter; the second region is a region in the pre-coded frame, and the second region is occluded by the first moving object in the next frame of the pre-coded frame, and the target quantization parameter corresponding to the second region is greater than the global quantization parameter.
[0011] In this possible implementation, the first region is the region where the first moving object is located in the predicted pre-coded frame. Since high-quality display of the moving object is usually required during playback, when encoding the video data corresponding to the first region, a smaller global quantization parameter is used. Allocating more encoding resources to this first region can improve the encoding quality of the first moving object in the pre-coded frame. The second region is the region in the predicted pre-coded frame that will be occluded by the first moving object in the next frame. Since it will be occluded by the first moving object in the next frame of this pre-coded frame, then encoding the second region with high quality in this pre-coded frame doesn't have much value. Therefore, when encoding the video data corresponding to the second region, a larger global quantization parameter is used, and fewer encoding resources are allocated to this second region, which can reduce the waste of encoding resources and improve the utilization rate of encoding resources.
[0012] In a possible implementation, the encoded frame and the pre-coded frame are consecutive frames or non-consecutive frames in the video sequence.
[0013] In this possible implementation, since the movement of an object is continuous, if the encoded frame and the pre-encoded frame are two consecutive video frames, the accuracy of predicting the target region in the pre-encoded frame can be improved. Of course, if the encoded frame and the pre-encoded frame are two video frames with a certain interval, it is also possible to implement the prediction of the target region in the pre-encoded frame, which improves the diversity of the prediction of the target region in the pre-encoded frame.
[0014] In one possible implementation, each video frame in the video sequence includes multiple blocks. The above step: predicting at least one target region in the pre-encoded frame according to the motion information of the first moving object in the encoded frame includes: predicting the position information of the motion block in the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-encoded frame, and the position information of the motion block in the encoded frame; wherein, the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the motion block in the pre-encoded frame is the block corresponding to the predicted first moving object in the pre-encoded frame; determining the first region according to the position information of the motion block in the pre-encoded frame.
[0015] In this application, the block can be a macro block (MB) in H.264, or the block in each video frame is a prediction unit (PU) in the H.265 scenario. Of course, the compression format of this application is not limited to H.264 and H.265, and can also be applied to other compression formats, such as: H.266, etc.
[0016] In this application, the interval between the encoded frame and the pre-encoded frame can indicate the relationship between the motion vector of the j-th block in the pre-encoded frame and the first-order motion vector of the j-th block in the encoded frame.
[0017] In this possible implementation, between the encoded frame and the pre-encoded frame, predicting the motion region of the first moving object in the pre-encoded frame, that is, the first region, according to the motion relationship of the corresponding blocks between the two video frames can improve the accuracy of predicting the first region.
[0018] In one possible implementation, each video frame in the video sequence includes multiple blocks. The above step: predicting at least one target region in the pre-encoded frame according to the motion information of the first moving object in the encoded frame includes: predicting the position information of the occluded block in the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-encoded frame, and the position information of the motion block in the encoded frame; wherein, the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the occluded block in the pre-encoded frame is the block predicted to be occluded by the first moving object in the next frame of the pre-encoded frame; determining the second region according to the position information of the occluded block in the pre-encoded frame.
[0019] In this possible implementation manner, between the encoded frame and the pre-encoded frame, the motion area of the first moving object in the next frame of the pre-encoded frame is predicted according to the motion relationship of the corresponding blocks between two video frames, so as to determine the occluded area in the pre-encoded frame, that is, the second area, which can improve the accuracy of predicting the second area.
[0020] In a possible implementation manner, the above steps: predicting the position information of the motion block in the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-encoded frame, and the position information of the motion block in the encoded frame, include: determining the motion vector of the motion block in the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame and the interval between the encoded frame and the pre-encoded frame; wherein, the interval between the encoded frame and the pre-encoded frame is used to indicate the multiple relationship between the motion vector of the motion block in the pre-encoded frame and the first-order motion vector; determining the position information of the motion block in the pre-encoded frame according to the position information of the motion block in the encoded frame and the motion vector of the motion block in the pre-encoded frame.
[0021] In this possible implementation manner, in the present application, if the interval between the encoded frame and the pre-encoded frame is 0, the motion vector of the motion block in the pre-encoded frame is twice the first-order motion vector. If the interval between the encoded frame and the pre-encoded frame is x, the motion vector of the motion block in the pre-encoded frame is (2 + x) times the first-order motion vector. In the present application, the position of the motion block in the pre-encoded frame is determined through the position information of the encoded frame and the motion vector of the motion block in the pre-encoded frame, which can improve the speed and accuracy of determining the position of the motion block in the pre-encoded frame.
[0022] In a possible implementation manner, predicting the position information of the occluded block in the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-encoded frame, and the position information of the motion block in the encoded frame, includes: determining the motion vector of the motion block in the next frame of the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame and the interval between the encoded frame and the pre-encoded frame; wherein, the interval between the encoded frame and the pre-encoded frame is used to indicate the multiple relationship between the motion vector of the motion block in the next frame of the pre-encoded frame and the first-order motion vector; determining the position information of the occluded block in the pre-encoded frame according to the position information of the motion block in the encoded frame and the motion vector of the motion block in the next frame of the pre-encoded frame.
[0023] In this possible implementation manner, in the present application, if the interval between the encoded frame and the pre-encoded frame is 0, the motion vector of the motion block in the next frame of the pre-encoded frame is three times that of the first-order motion vector. If the interval between the encoded frame and the pre-encoded frame is x, the motion vector of the motion block in the next frame of the pre-encoded frame is (3 + x) times that of the first-order motion vector. In the present application, by using the position information of the encoded frame and the motion vector of the motion block in the next frame of the pre-encoded frame, the position of the occluded block in the pre-encoded frame can be determined, which can improve the speed and accuracy of determining the position of the occluded block in the pre-encoded frame.
[0024] In a possible implementation manner, when there are multiple target regions, the method further includes: determining a first quantity of the coordinate vectors of the motion region pointing to the first block according to the position information of the motion blocks in the pre-encoded frame, where the first block is any block in the encoded frame; determining a second quantity of the coordinate vectors of the occluded region pointing to the first block according to the position information of the occluded blocks in the pre-encoded frame; and determining the quantization parameter of the second block corresponding to the first block in the pre-encoded frame according to the first quantity and the second quantity, where the quantization parameter of the second block is the target quantization parameter.
[0025] In this possible implementation manner, there may be multiple motion blocks around the first block. According to the position information of the motion blocks in the pre-encoded frame that has been determined previously, the quantity of the coordinate vectors of the motion region pointing to the first block in the motion blocks corresponding to the first moving object can be determined, that is, the first quantity. Similarly, according to the position information of the occluded blocks in the pre-encoded frame, the quantity of the coordinate vectors of the occluded region pointing to the first block can also be determined, that is, the second quantity. By determining the quantization parameter of the second block in the pre-encoded frame corresponding to the first block according to the first quantity and the second quantity, the accuracy of configuring the quantization parameter of the second block can be improved, and further the effectiveness of allocating the encoding resources of the block can be improved.
[0026] In a possible implementation manner, the above step: determining the quantization parameter of the second block corresponding to the first block in the pre-encoded frame according to the first quantity and the second quantity includes: determining the change amount of the quantization parameter of the second block according to the first quantity and the second quantity; and determining the quantization parameter of the second block according to the change amount of the quantization parameter of the second block and the global quantization parameter.
[0027] In this possible implementation manner, by first determining the change amount of the quantization parameter through the first quantity and the second quantity, and then combining the global quantization parameter to determine the quantization parameter of the second block, the accuracy of configuring the quantization parameter of the second block can be improved, and further the effectiveness of allocating the encoding resources of the block can be improved.
[0028] In a possible implementation manner, determining the change amount of the quantization parameter of the second block according to the first quantity and the second quantity includes: determining the change amount of the quantization parameter of the second block according to the following relational expression;
[0029]
[0030] wherein, deltaQP j represents the change amount of the quantization parameter of the second block, A is a constant, and ε is a constant. is the first quantity, is the second quantity.
[0031] In this possible implementation manner, the speed and accuracy of determining the change amount of the quantization parameter of the second block can be improved through the above relational expression.
[0032] In a possible implementation manner, the method further includes: embedding a descriptor in the second block, where the descriptor is used to indicate that the second block is a motion block or an occluded block. When the second block is a motion block, the target quantization parameter used for encoding the second block is less than the global quantization parameter. When the second block is an occluded block, the target quantization parameter used for encoding the second block is greater than the global quantization parameter.
[0033] In this possible implementation manner, the descriptor can be 0 or 1, or other forms of indication identifiers. For example: using 0 to indicate that the second block is a motion block, and using 1 to indicate that the second block is an occluded block; or, using 1 to indicate that the second block is a motion block, and using 0 to indicate that the second block is an occluded block. The present application does not make any limitation in this regard. By using the descriptor to indicate the target quantization parameter to be used when encoding the second block, the rationality of the allocation of encoding resources for the second block can be improved.
[0034] In a possible implementation manner, in a hard encoder scenario, the method further includes:
[0035] embedding deltaQP j in the second block, and deltaQP j is used to determine the corresponding target quantization parameter when encoding the second block.
[0036] In this possible implementation manner, embedding deltaQP j in the second block can improve the rationality of the allocation of encoding resources for the second block.
[0037] In a possible implementation manner, the video sequence is included in the decoded video stream.
[0038] In this possible implementation manner, the decoded video stream can be encoded according to the above video encoding process, and secondary compression of the video data can be achieved.
[0039] In a second aspect of the present application, there is provided an apparatus for video encoding. When encoding a video sequence, the apparatus includes:
[0040] A prediction unit, configured to predict at least one target region in a pre-coded frame according to the motion information of a first moving object in an encoded frame, where the at least one target region is related to the first moving object;
[0041] An encoding unit, configured to encode the video data corresponding to the at least one target region by using a target quantization parameter, where the target quantization parameter is different from the global quantization parameter of the video sequence.
[0042] In a possible implementation, when there is one target region, the target region is the first region or the second region; when there are multiple target regions, the multiple target regions include the first region and the second region; where the first region is the motion region of the first moving object in the pre-coded frame, and the target quantization parameter corresponding to the first region is less than the global quantization parameter; the second region is a region in the pre-coded frame, and the second region is occluded by the first moving object in the next frame of the pre-coded frame, and the target quantization parameter corresponding to the second region is greater than the global quantization parameter.
[0043] In a possible implementation, the encoded frame and the pre-coded frame are consecutive frames or non-consecutive frames in the video sequence.
[0044] In a possible implementation, the prediction unit is specifically configured to:
[0045] When each video frame in the video sequence includes multiple blocks, according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-coded frame, and the position information of the motion block in the encoded frame, predict the position information of the motion block in the pre-coded frame; where the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the motion block in the pre-coded frame is the predicted block corresponding to the first moving object in the pre-coded frame;
[0046] Determine the first region according to the position information of the motion block in the pre-coded frame.
[0047] In a possible implementation, the prediction unit is specifically configured to:
[0048] When each video frame in the video sequence includes multiple blocks, according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-coded frame, and the position information of the motion block in the encoded frame, predict the position information of the occluded block in the pre-coded frame; where the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the occluded block in the pre-coded frame is the predicted block that is occluded by the first moving object in the next frame of the pre-coded frame;
[0049] Determine the second region according to the position information of the occluded block in the pre-coded frame.
[0050] In a possible implementation, the prediction unit is specifically configured to:
[0051] Determine the motion vector of the motion block in the precoded frame according to the first-order motion vector of the motion block in the coded frame and the interval between the coded frame and the precoded frame; wherein, the interval between the coded frame and the precoded frame is used to indicate the multiple relationship between the motion vector of the motion block in the precoded frame and the first-order motion vector;
[0052] Determine the position information of the motion block in the precoded frame according to the position information of the motion block in the coded frame and the motion vector of the motion block in the precoded frame.
[0053] In a possible implementation manner, the prediction unit is specifically configured to:
[0054] Determine the motion vector of the motion block in the next frame of the precoded frame according to the first-order motion vector of the motion block in the coded frame and the interval between the coded frame and the precoded frame; wherein, the interval between the coded frame and the precoded frame is used to indicate the multiple relationship between the motion vector of the motion block in the next frame of the precoded frame and the first-order motion vector;
[0055] Determine the position information of the occluded block in the precoded frame according to the position information of the motion block in the coded frame and the motion vector of the motion block in the next frame of the precoded frame.
[0056] In a possible implementation manner, the prediction unit is further configured to: when there are multiple target regions;
[0057] Determine the first quantity of the coordinate vectors of the motion regions pointing to the first block according to the position information of the motion blocks in the precoded frame, where the first block is any block in the coded frame;
[0058] Determine the second quantity of the coordinate vectors of the occluded regions pointing to the first block according to the position information of the occluded blocks in the precoded frame;
[0059] Determine the quantization parameter of the second block corresponding to the first block in the precoded frame according to the first quantity and the second quantity, and the quantization parameter of the second block is the target quantization parameter.
[0060] In a possible implementation manner, the prediction unit is specifically configured to:
[0061] Determine the change amount of the quantization parameter of the second block according to the first quantity and the second quantity;
[0062] Determine the quantization parameter of the second block according to the change amount of the quantization parameter of the second block and the global quantization parameter.
[0063] In a possible implementation manner, the prediction unit is specifically configured to:
[0064] Determine the change amount of the quantization parameter of the second block according to the following relational expression;
[0065]
[0066] Among them, deltaQP j represents the change amount of the quantization parameter of the second block, A is a constant, and ε is a constant. is the first quantity, is the second quantity.
[0067] In a possible implementation, the prediction unit is further configured to:
[0068] Embed a descriptor in the second block, where the descriptor is used to indicate that the second block is a motion block or an occluded block. When the second block is a motion block, the target quantization parameter used for encoding the second block is less than the global quantization parameter. When the second block is an occluded block, the target quantization parameter used for encoding the second block is greater than the global quantization parameter.
[0069] In a possible implementation, in a hard encoder scenario, the prediction unit is further configured to: Embed deltaQP in the second block j , where deltaQP j is used to determine the corresponding target quantization parameter when encoding the second block.
[0070] In a possible implementation, the video sequence is included in the decoded video stream.
[0071] In a possible implementation, the blocks in each video frame are macroblocks MB in the H.264 scenario, or the blocks in each video frame are prediction units PU in the H.265 scenario.
[0072] The third aspect of the present application provides a video encoding device, and the video encoding device has the function of implementing the video encoding method in the above first aspect or any possible implementation of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions, such as: a prediction unit and an encoding unit.
[0073] The fourth aspect of the present application provides an electronic device, including a communication interface, an encoder, a processor, and a memory. The communication interface is coupled to the encoder, the processor, and the memory. The memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes the method in the above first aspect or any possible implementation of the first aspect.
[0074] In this application, the processor may include at least one of a central processing unit (CPU) and a graphics processing unit (GPU); wherein, both the CPU and the GPU may execute the video encoding process described in the first aspect or any possible implementation manner of the first aspect, or the CPU and the GPU cooperate to execute the video encoding process described in the first aspect or any possible implementation manner of the first aspect.
[0075] A fifth aspect of this application provides a chip system, which includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from the memory of the electronic device and send signals to the processors, and the signals include computer instructions stored in the memory; when the processors execute the computer instructions, the electronic device executes the method according to the first aspect or any possible implementation manner of the first aspect, and the processors are at least one of a CPU and a GPU.
[0076] A sixth aspect of this application provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on an electronic device, the electronic device is caused to execute the method according to the first aspect or any possible implementation manner of the first aspect.
[0077] A seventh aspect of this application provides a computer program product, which includes computer program code, and when the computer program code runs on a computer, the computer is caused to execute the method according to the first aspect or any possible implementation manner of the first aspect.
[0078] An eighth aspect of this application provides a video processing system, including: an electronic device, and the electronic device is used to execute the method according to the first aspect or any possible implementation manner of the first aspect.
[0079] A ninth aspect of this application provides a device for storing a bitstream, which includes at least one storage medium and a communication interface; the communication interface is used to receive or send the bitstream; the at least one storage medium is used to store the bitstream; the bitstream is encoded by an encoder executing the method according to the first aspect or any possible implementation manner of the first aspect.
[0080] A tenth aspect of this application provides a method for storing a bitstream, including: receiving the bitstream through the communication interface; storing the bitstream in one or more storage mediums, and the bitstream is encoded by an encoder executing the method according to the first aspect or any possible implementation manner of the first aspect.
[0081] The eleventh aspect of the present application provides a system for distributing a bitstream, including at least one storage medium and a video stream device; the at least one storage medium is used for storing the bitstream, and the bitstream is encoded by an encoder implementing the method according to the first aspect or any possible implementation manner of the first aspect;
[0082] The video stream device is configured to, in response to a request from a decoder, cause the target bitstream in the at least one storage medium to be sent to the decoder.
[0083] The twelfth aspect of the present application provides a method for distributing a bitstream, including: receiving a first request; in response to the first request, selecting a target bitstream from at least one storage medium; sending the target bitstream to a destination device; the at least one storage medium is used for storing the bitstream, and the target bitstream is encoded by an encoder implementing the method according to the first aspect or any possible implementation manner of the first aspect.
[0084] The thirteenth aspect of the present application provides a system for processing a bitstream, including an image source device, an encoder device, one or more storage media, and a destination device;
[0085] The image source device is configured to provide image data;
[0086] The encoder device is configured to obtain the image data of the image source device through an interface and encode the image data to obtain one or more bitstreams, and the bitstreams are encoded by the encoder device implementing the method according to the first aspect or any possible implementation manner of the first aspect;
[0087] The encoder device is configured to store one or more bitstreams into one or more storage media; or the encoder device is configured to encapsulate one or more bitstreams to obtain a transport bitstream;
[0088] The encoder device is configured to transmit the transport bitstream to the destination device through a communication link or a communication network; the destination device is configured to de-encapsulate the transport bitstream to obtain one or more bitstreams;
[0089] The destination device is configured to decode one or more bitstreams to obtain decoded data.
[0090] For the relevant features and effects of the second aspect of the present application and any possible implementation manners from the second aspect to the thirteenth aspect, reference may be made to the corresponding descriptions in the first aspect or any possible implementation manner of the first aspect for understanding. Description of the Drawings
[0091] Figure 1A is a schematic diagram of a scenario architecture provided by an embodiment of the present application;
[0092] Figure 1B is another schematic diagram of a scenario architecture provided by an embodiment of the present application;
[0093] Figure 2 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0094] Figure 3 It is a schematic diagram of an embodiment of a video encoding method provided by an embodiment of the present application;
[0095] Figure 4 It is a schematic diagram of a scenario example provided by an embodiment of the present application;
[0096] Figure 5A It is a schematic diagram of another embodiment of a video encoding method provided by an embodiment of the present application;
[0097] Figure 5B It is a schematic example diagram of region prediction provided by an embodiment of the present application;
[0098] Figure 6 It is a schematic diagram of another embodiment of a video encoding method provided by an embodiment of the present application;
[0099] Figure 7 It is a schematic structural diagram of a video encoding device provided by an embodiment of the present application. Detailed implementation manners
[0100] Next, with reference to the accompanying drawings, the embodiments of the present application will be described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Those of ordinary skill in the art can know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0101] Terms such as "first" and "second" in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0102] An embodiment of the present application provides a video encoding method for improving the effectiveness of encoding resource allocation during video encoding. The present application also provides corresponding devices, computer-readable storage media, computer program products, etc. The following will be described in detail respectively.
[0103] For ease of understanding, the following briefly introduces the technical terms involved in the embodiments of the present application:
[0104] YUV format: The YUV format is a picture format composed of three parts: Y, U, and V. Y represents luminance, and U and V represent the chrominance of color.
[0105] H.264: H.264 is a new generation of digital video compression format jointly proposed by the International Organization for Standardization and the International Telecommunication Union, and is one of the video codec technology standards named after the H.26x series.
[0106] H.265: H.265 is also one of the video codec technology standards named after the H.26x series.
[0107] Video stream or video sequence: Composed of multiple video frames, and the video frames can include I-frames, P-frames, and B-frames.
[0108] I-frame: Also known as a complete frame or key frame, and the content of the I-frame needs to be saved during encoding.
[0109] P-frame: Also called a forward reference frame, and uses inter-frame compression technology. The P-frame only needs to save the data that is different from the previous frame, and refers to the previous frame during compression.
[0110] B-frame: Also called a bidirectional reference frame, and uses inter-frame compression technology. During compression, it refers to both the previous frame and the next frame.
[0111] Each video frame can be further divided into slices, and the slices can be further divided into blocks.
[0112] Slice: It is the carrier of macroblocks. The purpose of setting slices is to limit the spread and transmission of errors. The encoded slices are independent of the project. The prediction of one slice cannot use the macroblocks in other slices as the reference image. This ensures that the prediction error of a certain slice will not spread to other slices. We can understand that a picture can contain one or more slices, and each slice contains an integer number of macroblocks, that is, each slice has at least one macroblock, and at most each slice contains all the macroblocks of the image. The slice can be subdivided into "slice header + slice data" because the frame data may not be transmitted completely at one time, so the header information needs to be recorded.
[0113] Macro block (MB): The main carrier of video information, which contains the luminance and chrominance information of each pixel. A macro block usually consists of a 16×16 luminance pixel and an additional 8×8 Cb and an 8×8 Cr color pixel block. In each image, several macro blocks are arranged in the form of slices. The macro block contains information such as macro block type, prediction type, coded block pattern, quantization parameter (QP), luminance and chrominance data sets of pixels, etc. MB is usually the term in H.264.
[0114] In H.265, the specific division method and naming of video frames are slightly different from those in H.264. In H.265, the corresponding unit to MB is the prediction unit (PU).
[0115] The video encoding method provided by the embodiments of the present application can be applied to various scenarios that require video encoding, especially suitable for scenarios with moving objects. Taking the intelligent transportation scenario as an example, through the solution provided by the embodiments of the present application, the video captured by the camera (also called the camera head) can be encoded and then sent to the device for storing the bitstream. The device for storing the bitstream can be a local storage device or a cloud device. The encoder for encoding the video captured by the camera can be configured on the camera or on a proximal device separated from the camera, and this proximal device is usually installed at a position relatively close to the camera.
[0116] The scenario where the encoder is configured on the camera can be referred to Figure 1A for understanding. As Figure 1A shown, taking an intersection as an example, there are four cameras at this intersection. Each camera can capture the traffic conditions within a certain range of this intersection. The video data captured by each camera will enter the encoder of this camera in the form of a video sequence. The encoder will encode each video frame in the video sequence frame by frame and then output the encoded video stream. As Figure 1A shown, the encoded video streams output by the four cameras are respectively the encoded video stream 1, the encoded video stream 2, the encoded video stream 3, and the encoded video stream 4. Each camera can send the encoded video stream to the cloud storage through the network or send the encoded video stream to the local storage device for storage. The present application does not make any limitations on this.
[0117] The scenario where the encoder is configured on the proximal device of the camera can be referred to Figure 1B for understanding. As Figure 1BAs shown, taking an intersection as an example, there are four cameras at this intersection. Each camera can capture the traffic conditions within a certain range of the intersection. The video data captured by each camera will first be transmitted to a proximal device, and the encoders configured on this proximal device will encode the video data transmitted by each camera respectively. It should be noted that one or more encoders can be configured on the proximal device. If one encoder is configured, this encoder needs to encode the video data transmitted by multiple cameras respectively. If multiple encoders are configured, different encoders can encode the video data transmitted by different cameras. After the proximal device encodes the video data of the four cameras, four encoded video streams can be obtained, namely: encoded video stream 1, encoded video stream 2, encoded video stream 3, and encoded video stream 4. The proximal device can send the encoded video streams to the cloud for storage through the network, or send the encoded video streams to a local storage device for storage. This application does not make any limitations on this.
[0118] It should be noted that the encoding scheme provided in the embodiments of this application can also be encoding in the cloud. In this case, the camera transmits the captured video data to the cloud through the network, and the cloud can encode the video data of each camera to obtain the corresponding encoded video stream.
[0119] The above Figure 1A or Figure 1B In the video encoding scenarios introduced above, whether it is the camera for encoding, the proximal device for encoding, or the device in the cloud for encoding, the structures of these electronic devices that may complete the video encoding provided in the embodiments of this application can be referred to Figure 2 for understanding.
[0120] Figure 2 FIG. is a schematic diagram of a possible logical structure of the cloud device / roadside device provided in the embodiments of this application. As Figure 2 shown, the electronic device 20 provided in the embodiments of this application includes: an encoder 200, a processor 201, a communication interface 202, a memory 203, and a bus 204. The encoder 200, the processor 201, the communication interface 202, and the memory 203 are interconnected through the bus 204. In the embodiments of this application, the processor 201 is used to control and manage the actions of the electronic device 20, and the encoder 200 is used to encode the video data obtained by the electronic device 20. The communication interface 202 is used to support the communication of the electronic device 20. For example: the communication interface 202 can obtain video data and send the encoded video data. The memory 203 is used to store the program code and data of the electronic device 20.
[0121] Among them, the processor 201 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 204 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 2 only a thick line is used to represent it in Figure 2 , but it does not mean that there is only one bus or one type of bus.
[0122] The method for video coding provided by the embodiments of this application will be described below. This method can be executed by an electronic device or by a component of an electronic device (for example: an encoder, a processor, a chip, or a chip system, etc.).
[0123] As Figure 3 shown, an embodiment of the method for video coding provided by the embodiments of this application includes:
[0124] 301. Predict at least one target region in a pre-coded frame according to the motion information of a first moving object in an encoded frame, where the at least one target region is related to the first moving object.
[0125] In this application, the video sequence can be a YUV sequence. The encoded frame and the pre-coded frame can be two consecutive video frames, or can be two video frames separated by one or more video frames.
[0126] In this application, the pre-coded frame refers to a frame to be encoded or a frame ready to be encoded, usually referring to the next frame that is about to be encoded immediately after the previous frame is encoded.
[0127] In this application, the first moving object can be one or more moving objects in the encoded frame. At least one target region in the pre-coded frame is part or all of the region corresponding to the pre-coded frame.
[0128] 302. Encode the video data corresponding to the at least one target region using a target quantization parameter, where the target quantization parameter is different from the global quantization parameter of the video sequence.
[0129] In this application, the global quantization parameter is a quantization parameter configured for encoding the video sequence.
[0130] In the solution provided by the embodiment of the present application, for a target region related to a first moving object, a target quantization parameter different from the global quantization parameter is used for encoding. The target quantization parameter can be greater than the global quantization parameter or less than the global quantization parameter. Moreover, when there are multiple target regions, the target quantization parameter corresponding to some target regions can be greater than the global quantization parameter, and the target quantization parameter corresponding to some target regions can be less than the global quantization parameter. During encoding, more encoding resources can be allocated to the target region with a smaller quantization parameter, and fewer encoding resources can be allocated to the target region with a larger quantization parameter. In this way, compared with the prior art in which each video frame in a video sequence needs to be encoded according to the global quantization parameter, the effectiveness of encoding resource allocation during video encoding can be improved, and the efficiency and quality of video encoding can be enhanced.
[0131] Optionally, in the embodiment of the present application, when there is one target region, the target region can be the first region or the second region; when there are multiple target regions, the multiple target regions include the first region and the second region. Among them, the first region is the moving region of the first moving object in the pre-encoded frame, and the target quantization parameter corresponding to the first region is less than the global quantization parameter; the second region is the region in the pre-encoded frame, and the second region is occluded by the first moving object in the next frame of the pre-encoded frame, and the target quantization parameter corresponding to the second region is greater than the global quantization parameter.
[0132] In the embodiment of the present application, the encoded frame and the pre-encoded frame can be consecutive frames or non-consecutive frames in a video sequence. When the encoded frame and the pre-encoded frame are consecutive frames, the interval between the encoded frame and the pre-encoded frame is 0. When the encoded frame and the pre-encoded frame are non-consecutive frames, the interval between the encoded frame and the pre-encoded frame is the number of video frames in between.
[0133] Regarding the moving region and the occluded region, Figure 4 Three consecutive video frames in Figure 4 are taken as an example for illustration. As Figure 4 shown, in the (n - 1)-th frame, the region where the first moving object is located is the moving region, and the occluded region in the (n - 1)-th frame is the region predicted that the first moving object will move to in the n-th frame; similarly, in the n-th frame, the region where the first moving object is located is the moving region, and the occluded region in the n-th frame is the region predicted that the first moving object will move to in the (n + 1)-th frame. Inside each frame, the motion vectors of the moving region and the occluded region can both be In the (n - 1)-th frame, the motion vectors of the moving region and the occluded region are represented as In the n-th frame, the motion vectors of the moving region and the occluded region are represented as In the (n + 1)-th frame, the motion vectors of the moving region and the occluded region are represented as Between frames, because the three video frames are consecutive, so That is to say, if the interval between the encoded frame and the pre-encoded frame is 0, the motion vector of the motion block in the pre-encoded frame is twice the first-order motion vector of the corresponding motion block in the encoded frame, and the motion vector of the motion block in the next frame of the pre-encoded frame is three times the first-order motion vector of the corresponding motion block in the encoded frame.
[0134] The Figure 4 As shown, for three consecutive frames, if the interval between the encoded frame and the pre-encoded frame is x (x is a positive integer), the motion vector of the motion block in the pre-encoded frame is (2 + x) times the first-order motion vector of the corresponding motion block in the encoded frame, and the motion vector of the motion block in the next frame of the pre-encoded frame is (3 + x) times the first-order motion vector of the corresponding motion block in the encoded frame.
[0135] In the embodiment of the present application, the process of predicting the first region (i.e., the motion region) may be: predicting the position information of the motion block in the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-encoded frame, and the position information of the motion block in the encoded frame; wherein, the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the motion block in the pre-encoded frame is the block corresponding to the predicted first moving object in the pre-encoded frame; determining the first region according to the position information of the motion block in the pre-encoded frame.
[0136] In the embodiment of the present application, the process of predicting the second region (i.e., the occluded region) may be: predicting the position information of the occluded block in the pre-encoded frame according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-encoded frame, and the position information of the motion block in the encoded frame; wherein, the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the occluded block in the pre-encoded frame is the block predicted to be occluded by the first moving object in the next frame of the pre-encoded frame; determining the second region according to the position information of the occluded block in the pre-encoded frame.
[0137] In the embodiment of the present application, for the process of determining the first region and the second region, reference may also be made to Figure 5A or Figure 6 for understanding.
[0138] As Figure 5A shown, the encoded frame and the pre-encoded frame are consecutive frames. The encoded frame may be represented by the (n - 1)-th frame, the pre-encoded frame may be represented by the n-th frame, and the next frame of the pre-encoded frame may be represented by the (n + 1)-th frame, where n is an integer greater than 1.
[0139] 501. Obtain the first-order motion vector and the position information of the i-th block in the (n - 1)-th frame.
[0140] The first-order motion vector of the i-th block in the (n - 1)-th frame may be represented by and the position information of this block may be represented by representation
[0141] 502. Determine the running vector and position information of the i-th block in the n-th frame, and the running vector and position information of the i-th block in the (n + 1)-th frame according to the first-order motion vector of the i-th block in the (n - 1)-th frame.
[0142] The running vector of the i-th block in the n-th frame can be represented as The position information can be represented by representation. Since in the embodiments of the present application, the encoded frame and the pre-encoded frame are consecutive frames, the interval between the encoded frame and the pre-encoded frame is 0, so Furthermore, the position of the corresponding block in the n-th frame is obtained
[0143] The running vector of the i-th block in the (n + 1)-th frame can be represented as The position information can be represented by representation. Furthermore, the position of the corresponding block in the (n + 1)-th frame is obtained That is, the information of the blocks in the occluded area in the n-th frame.
[0144] If there are M blocks in each frame, repeating the above steps 401 and 402, the position information set of the M blocks in the n-th frame can be obtained, which can be represented as and where represents all the motion block positions predicted in the n-th frame, represents all the occluded block positions predicted in the n-th frame.
[0145] Through and the first area and the second area in the n-th frame can be determined.
[0146] For an example of predicting the motion area and the occluded area in the n-th frame, reference can be made to Figure 5B for understanding. As Figure 5B shown, motion area prediction: The motion vector field of the current unencoded n-th frame, that is, the second-order motion vector field, can be predicted through the motion vector field of the encoded (n - 1)-th frame, that is, the first-order motion vector field. The second-order motion vector field describes the motion characteristics of the current unencoded frame, that is, the n-th frame. Using the number of second-order motion vectors pointing to the current block, the motion intensity of the current block is predicted, and the motion area is discriminated. Taking the b1 block in the (n - 1)-th frame as an example, the number M1 of second-order motion vectors pointing to the b1 block is counted block by block. The larger M1 is, the higher the possibility of the motion area in the next frame. For the b1 block in the n-th frame, a smaller QP coding can be used.
[0147] Occluded Region Prediction: Predict the motion vector field of the next uncoded (n + 1)-th frame, i.e., the third-order motion vector field, through the motion vector field of the already-coded (n - 1)-th frame, i.e., the first-order motion vector field. The third-order motion vector field describes the motion characteristics and motion regions of the next uncoded frame, i.e., the (n + 1)-th frame, and corresponds to the occluded region of the current uncoded frame, i.e., the n-th frame, using the reference relationship. Using the number of third-order motion vectors pointing to the current block, predict whether the current block will be occluded in the next frame. Taking blocks b2 - b4 as an example, count the number M2 of third-order motion vectors pointing to blocks b2 - b4 one by one. The larger M2 is, the higher the probability that the block will be occluded in the next frame. For blocks b2 - b4 in the n-th frame, a larger QP coding needs to be used.
[0148] 503. Determine a first quantity according to the position information of the motion blocks in the pre-coded frame, and determine a second quantity according to the position information of the occluded blocks in the pre-coded frame.
[0149] Wherein, the first quantity is the number of coordinate vectors of the motion region pointing to the first block, the second quantity is the number of coordinate vectors of the occluded region pointing to the first block, and the first block is any block in the already-coded frame.
[0150] Taking the j-th block in the (n - 1)-th frame as an example, it can be based on Count the number of coordinate vectors of the moving object pointing to this macro block It can be based on Count the number of coordinate vectors of the occluded region pointing to this macro block
[0151] After determining according to the (n - 1)-th frame It can be passed j = 1, 2, 3…, K, to the n-th frame.
[0152] 504. Determine the quantization parameter of the second block corresponding to the first block in the pre-coded frame according to the first quantity and the second quantity, and the quantization parameter of the second block is the target quantization parameter.
[0153] This step 504 can be: Determine the change amount of the quantization parameter of the second block according to the first quantity and the second quantity; Determine the quantization parameter of the second block according to the change amount of the quantization parameter of the second block and the global quantization parameter.
[0154] Wherein, the second block is the j-th block in the n-th frame, and the change amount of the quantization parameter of the second block can be determined by the following relational expression:
[0155]
[0156] Wherein, deltaQP j represents the change amount of the quantization parameter of the second block, A is a constant, ε is a constant, is the first quantity, is the second quantity.
[0157] If the global quantization parameter is 23, then the QP of the j-th block in the n-th frame j = 23 + deltaQP j , and this QP j is the target quantization parameter of the j-th block in the n-th frame.
[0158] When is greater than , deltaQP j is negative, and QP j will be less than the global quantization parameter 23, that is, a smaller QP j is used for encoding the high-motion area; when is less than , deltaQP j is positive, and QP j will be greater than the global quantization parameter 23, that is, a larger QP j is used for encoding the occluded area. When is equal to , deltaQP j is 0, that is, QP j is equal to the global quantization parameter 23.
[0159] 505. Use QP j to encode the j-th block in the n-th frame, and embed the information related to QP j into the bitstream of the j-th block.
[0160] The information related to QP j can be a descriptor, which is used to indicate that the j-th block is a motion block or an occluded block. When the j-th block is a motion block, the target quantization parameter used for encoding the j-th block is less than the global quantization parameter. When the j-th block is an occluded block, the target quantization parameter used for encoding the second block is greater than the global quantization parameter.
[0161] The descriptor can be 0 or 1, or other forms of indication identifiers. For example: 0 is used to indicate that the second block is a motion block, and 1 is used to indicate that the second block is an occluded block; or, 1 is used to indicate that the second block is a motion block, and 0 is used to indicate that the second block is an occluded block. This application does not make any limitations in this regard. By using the descriptor to indicate the target quantization parameter that should be used when encoding the second block, the rationality of the allocation of encoding resources for the second block can be improved.
[0162] Specifically, it can be: If then the descriptor representing motion can be embedded into the bitstream of the j-th block for subsequent decoding. If It indicates that the j-th block is an occluded descriptor to indicate that the j-th block is an occluded block, and the descriptor of the occluded block is embedded into the bitstream.
[0163] Repeat the above operations for each block in the n-th frame, and the encoding of the n-th frame can be completed.
[0164] Repeat the above process for each video frame in the video sequence, and the encoding of all frames in the video sequence can be completed.
[0165] It should be noted that the process described in the above embodiments can be applied to an H.264 encoder. Input a YUV sequence into the H.264 encoder. The resolution of this sequence can be 1920x1080, the frame rate is 30fps, the IPPP prediction structure, and the global QP can be set to 23. Through the video encoding method provided above, each frame in the YUV sequence can be encoded.
[0166] In the case of a hard encoder scenario, the information related to QP j can be deltaQP j , and deltaQP j is embedded into the bitstream of the j-th block for subsequent decoding use.
[0167] The above video encoding scheme provided by the embodiments of the present application can also be applied to the scenario of secondary compression. It is possible to first decode the already compressed video data, and then input the decoded YUV sequence into the encoder to execute the video encoding process provided by the embodiments of the present application. For example, the input is an H.264 standard bitstream v, which is decoded into a YUV sequence of 300 frames. According to the GOP length of 60 frames, it is split into 5 sub-video sequences, v1 - v5. Perform secondary encoding on v1 - v5, and the global QP can be set to 23. During the secondary encoding process, it is possible to first encode the v1 sub-video sequence to obtain the bitstream v'1 after secondary encoding, and then encode each video sequence of v2 - v5 one by one to obtain the bitstreams v'2 to v'5 after secondary encoding. After merging with v'1, the final bitstream v' after secondary compression is obtained.
[0168] The above encoding scheme provided by the embodiments of the present application can also be applied to an H.265 encoder. The process of applying it to an H.265 encoder is the same as the above process, except that the blocks in the previous frame can be macro blocks (MB) in H.264, and in the H.265 encoder, the blocks in the previous frame can be prediction units (PU). Of course, the compression format of the present application is not limited to H.264 and H.265, and can also be applied to other compression formats, such as: H.266, etc.
[0169] In the embodiments of the present application, since the movement of an object is continuous, if the encoded frame and the pre-encoded frame are two consecutive video frames, the accuracy of predicting the target area in the pre-encoded frame can be improved.
[0170] Figure 5A What is introduced is the scenario where the encoded frame and the pre-encoded frame are consecutive frames. Next, refer to Figure 6 Introduce the scenario where the encoded frame and the pre-encoded frame are non-consecutive frames. Taking the example that the encoded frame and the pre-encoded frame are separated by x video frames, where x is a positive integer, the encoded frame can be represented by the (m - 1)-th frame, the pre-encoded frame can be represented by the (m + x)-th frame, and the next frame of the pre-encoded frame can be represented by the (m + x + 1)-th frame.
[0171] 601. Obtain the first-order motion vector and position information of the i-th block in the (m - 1)-th frame.
[0172] The first-order motion vector of the i-th block in the (m - 1)-th frame can be represented by and the position information of this block can be represented by represented.
[0173] 602. Determine the motion vector and position information of the i-th block in the (m + x)-th frame and the motion vector and position information of the i-th block in the (m + x + 1)-th frame according to the first-order motion vector of the i-th block in the (m - 1)-th frame.
[0174] The motion vector of the i-th block in the (m + x)-th frame can be represented as The position information can be represented by represented. Since in the embodiments of the present application, the interval between the encoded frame and the pre-encoded frame is x, so Furthermore, the position of the corresponding block in the (m + x)-th frame is obtained
[0175] The motion vector of the i-th block in the (m + x + 1)-th frame can be represented as The position information can be represented by represented. For Furthermore, the position of the corresponding block in the (m + x + 1)-th frame is obtained That is, the information of the blocks in the occluded area in the (m + x)-th frame.
[0176] If there are M blocks in each frame, repeating the above steps 601 and 602 can obtain the set of position information of M blocks in the (m + x)-th frame, which can be represented as and wherein, represents all the motion block positions predicted in the (m + x)-th frame, represents all the occluded block positions predicted in the (m + x)-th frame.
[0177] By and the first region and the second region in the (m + x)-th frame can be determined.
[0178] 603. Determine a first quantity according to the position information of the motion blocks in the precoded frame, and determine a second quantity according to the position information of the occluded blocks in the precoded frame.
[0179] Wherein, the first quantity is the number of coordinate vectors of the motion region pointing to the first block, the second quantity is the number of coordinate vectors of the occluded region pointing to the first block, and the first block is any block in the coded frame.
[0180] Taking the j-th block in the (m - 1)-th frame as an example, it can be based on count the number of coordinate vectors of the moving object pointing to this macro block It can be based on count the number of coordinate vectors of the occluded region pointing to this macro block
[0181] After determining according to the (m - 1)-th frame then be passed to the (m + x)-th frame.
[0182] 604. Determine the quantization parameter of the second block corresponding to the first block in the precoded frame according to the first quantity and the second quantity, and the quantization parameter of the second block is the target quantization parameter.
[0183] This step 404 can be: determine the change amount of the quantization parameter of the second block according to the first quantity and the second quantity; determine the quantization parameter of the second block according to the change amount of the quantization parameter of the second block and the global quantization parameter.
[0184] Wherein, the second block is the j-th block in the (m + x)-th frame, and the change amount of the quantization parameter of the second block can be determined by the following relational expression:
[0185]
[0186] Wherein, deltaQP j represents the change amount of the quantization parameter of the second block, A is a constant, ε is a constant, is the first quantity, is the second quantity.
[0187] If the global quantization parameter is 23, then the QP of the j-th block in the (m + x)-th frame j = 23 + deltaQP j and this QP j is the target quantization parameter of the j-th block in the (m + x)-th frame.
[0188] When Greater than When deltaQP j is negative, QP j will be less than the global quantization parameter 23, that is, a smaller QP is used for the high-motion region for encoding; when j When less than When deltaQP j is positive, QP j will be greater than the global quantization parameter 23, that is, a larger QP is used for the occluded region for encoding. When j When equal to When deltaQP j is 0, that is, QP j is equal to the global quantization parameter 23.
[0189] 605. Use QP j to encode the j-th block in the (m + x)-th frame, and embed the information related to QP j into the bitstream of the j-th block.
[0190] The information related to QP j can be a descriptor, which is used to indicate that the j-th block is a motion block or an occluded block. When the j-th block is a motion block, the target quantization parameter used for encoding the j-th block is less than the global quantization parameter. When the j-th block is an occluded block, the target quantization parameter used for encoding the second block is greater than the global quantization parameter.
[0191] The descriptor can be 0 or 1, or other forms of indication identifiers. For example, the second block is indicated as a motion block by 0, and the second block is indicated as an occluded block by 1; or, the second block is indicated as a motion block by 1, and the second block is indicated as an occluded block by 0. This application does not make any limitations in this regard. By using the descriptor to indicate the target quantization parameter to be used when encoding the second block, the rationality of the allocation of encoding resources for the second block can be improved.
[0192] Specifically, it can be: if then the descriptor representing motion can be embedded into the bitstream of the j-th block for subsequent decoding. If then it represents the occluded descriptor of the j-th block to indicate that the j-th block is an occluded block, and the descriptor of the occluded block is embedded into the bitstream.
[0193] Repeating the above operations for each block in the (m + x)-th frame can complete the encoding of the (m + x)-th frame.
[0194] Repeating the above process for each video frame in the video sequence can complete the encoding of all frames in the video sequence.
[0195] It should be noted that the process described in the above embodiments can be applied to an H.264 encoder. A YUV sequence is input to the H.264 encoder. The resolution of this sequence can be 1920x1080, the frame rate is 30fps, the IPPP prediction structure, and the global QP can be set to 23. By using the video coding method provided above, each frame in the YUV sequence can be encoded.
[0196] In the case of a hardware encoder scenario, the information related to QP j can be deltaQP j , and deltaQP j is embedded in the bitstream of the j-th block for subsequent decoding use.
[0197] The above video coding solution provided by the embodiments of the present application can also be applied to the scenario of secondary compression. The already compressed video data can be decoded first, and the decoded YUV sequence is then input into the encoder to execute the video coding process provided by the embodiments of the present application. For example, the input is an H.264 standard bitstream v, which is decoded into a YUV sequence of 300 frames. According to the GOP length of 60 frames, it is split into 5 sub-video sequences, v1-v5. The v1-v5 are secondarily encoded, and the global QP can be set to 23. During the secondary encoding process, the v1 sub-video sequence can be encoded first to obtain the secondarily encoded bitstream v'1, and then the v2-v5 video sequences are encoded one by one to obtain the secondarily encoded bitstreams v'2~v'5. After merging with v'1, the final secondarily compressed bitstream v' is obtained.
[0198] The above encoding solution provided by the embodiments of the present application can also be applied to an H.265 encoder. The process of applying it to an H.265 encoder is the same as the above process, except that the blocks in the previous frame can be macro blocks (MB) in H.264, and in the H.265 encoder, the blocks in the previous frame can be prediction units (PU). Of course, the compression format of the present application is not limited to H.264 and H.265, and can also be applied to other compression formats, such as: H.266, etc.
[0199] For the encoding solution provided by the embodiments of the present application, the encoded frame and the pre-encoded frame are two video frames with a certain interval, and it is also possible to achieve the prediction of the target area in the pre-encoded frame, which improves the diversity of the prediction of the target area in the pre-encoded frame.
[0200] The video encoding method provided in the embodiments of the present application relies on frame - to - frame motion vector tracing to adaptively adjust quantization parameters, and then uses the adaptively adjusted quantization parameters for encoding. The encoding efficiency of video sequences containing moving objects is significantly improved. To better illustrate the significant improvement in the encoding efficiency of the present application, the applicant used four video sequences and conducted experiments for each video sequence using the first scheme and the second scheme respectively. The first scheme is to encode using the adaptive quantization parameters of the present application, and the second scheme is to encode using global quantization parameters; when the global quantization parameters take values of 22, 27, 32, and 37 respectively, when the video sequence is a high - motion video sequence, after encoding the four video sequences using the first scheme and the second scheme respectively, the average speed increase reached 23.67%. It can be seen that the encoding efficiency is greatly improved when using the scheme of the present application for encoding.
[0201] The video encoding scheme provided by the present application significantly improves the encoding efficiency when encoding high - motion video sequences, and there is also a slight improvement in the encoding efficiency when encoding stationary or low - motion video sequences. It can be seen that whether it is a high - motion video sequence or a stationary or low - motion video sequence, the encoding efficiency will be improved when using the scheme provided by the present application for video encoding.
[0202] For the encoding scheme provided in the embodiments of the present application above, whether the encoded frame and the pre - encoded frame are consecutive video frames or non - consecutive video frames, through the above - mentioned scheme, the adaptive adjustment of the quantization parameters for each frame or each block can be achieved, thereby improving the effectiveness of encoding resource allocation. Moreover, the scheme algorithm provided in the embodiments of the present application has a low complexity, does not require a large amount of frame buffering and pre - encoding, not only improves the encoding efficiency, but also reduces the encoding delay and improves the real - time performance of video processing.
[0203] The video encoding method is introduced above. Next, in combination with Figure 7 The video encoding device 70 provided in the embodiments of the present application is introduced. When encoding a video sequence, the device 70 includes:
[0204] A prediction unit 701, configured to predict at least one target region in the pre - encoded frame according to the motion information of the first moving object in the encoded frame, and the at least one target region is related to the first moving object.
[0205] An encoding unit 702, configured to encode the video data corresponding to the at least one target region using target quantization parameters, and the target quantization parameters are different from the global quantization parameters of the video sequence.
[0206] Optionally, when there is one target region, the target region is the first region or the second region; when there are multiple target regions, the multiple target regions include the first region and the second region; wherein, the first region is the motion region of the first moving body in the precoded frame, and the target quantization parameter corresponding to the first region is smaller than the global quantization parameter; the second region is the region in the precoded frame, and the second region is occluded by the first moving body in the next frame of the precoded frame, and the target quantization parameter corresponding to the second region is greater than the global quantization parameter.
[0207] Optionally, the encoded frame and the precoded frame are consecutive frames or non-consecutive frames in a video sequence.
[0208] Optionally, the prediction unit 701 is specifically configured to:
[0209] When each video frame in the video sequence includes multiple blocks, according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the precoded frame, and the position information of the motion block in the encoded frame, predict the position information of the motion block in the precoded frame; wherein, the motion block in the encoded frame is the block corresponding to the first moving body in the encoded frame, and the motion block in the precoded frame is the predicted block corresponding to the first moving body in the precoded frame;
[0210] Determine the first region according to the position information of the motion block in the precoded frame.
[0211] Optionally, the prediction unit 701 is specifically configured to:
[0212] When each video frame in the video sequence includes multiple blocks, according to the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the precoded frame, and the position information of the motion block in the encoded frame, predict the position information of the occluded block in the precoded frame; wherein, the motion block in the encoded frame is the block corresponding to the first moving body in the encoded frame, and the occluded block in the precoded frame is the predicted block that is occluded by the first moving body in the next frame of the precoded frame;
[0213] Determine the second region according to the position information of the occluded block in the precoded frame.
[0214] Optionally, the prediction unit 701 is specifically configured to:
[0215] According to the first-order motion vector of the motion block in the encoded frame and the interval between the encoded frame and the precoded frame, determine the motion vector of the motion block in the precoded frame; wherein, the interval between the encoded frame and the precoded frame is used to indicate the multiple relationship between the motion vector of the motion block in the precoded frame and the first-order motion vector;
[0216] According to the position information of the motion block in the encoded frame and the motion vector of the motion block in the precoded frame, determine the position information of the motion block in the precoded frame.
[0217] Optionally, the prediction unit 701 is specifically configured to:
[0218] Determine the motion vector of the motion block in the next frame of the precoded frame according to the first-order motion vector of the motion block in the coded frame and the interval between the coded frame and the precoded frame; wherein, the interval between the coded frame and the precoded frame is used to indicate the multiple relationship between the motion vector of the motion block in the next frame of the precoded frame and the first-order motion vector;
[0219] Determine the position information of the occluded block in the precoded frame according to the position information of the motion block in the coded frame and the motion vector of the motion block in the next frame of the precoded frame.
[0220] Optionally, the prediction unit 701 is further configured to: when there are multiple target regions;
[0221] Determine the first quantity of the coordinate vectors of the motion regions pointing to the first block according to the position information of the motion blocks in the precoded frame, where the first block is any block in the coded frame;
[0222] Determine the second quantity of the coordinate vectors of the occluded regions pointing to the first block according to the position information of the occluded blocks in the precoded frame;
[0223] Determine the quantization parameter of the second block corresponding to the first block in the precoded frame according to the first quantity and the second quantity, and the quantization parameter of the second block is the target quantization parameter.
[0224] Optionally, the prediction unit 701 is specifically configured to:
[0225] Determine the change amount of the quantization parameter of the second block according to the first quantity and the second quantity;
[0226] Determine the quantization parameter of the second block according to the change amount of the quantization parameter of the second block and the global quantization parameter.
[0227] Optionally, the prediction unit 701 is specifically configured to:
[0228] Determine the change amount of the quantization parameter of the second block according to the following relational expression;
[0229]
[0230] where deltaQP j represents the change amount of the quantization parameter of the second block, A is a constant, ε is a constant, is the first quantity, is the second quantity.
[0231] Optionally, the prediction unit 701 is further configured to:
[0232] Embed a descriptor in the second block, where the descriptor is used to indicate that the second block is a motion block or an occluded block. When the second block is a motion block, the target quantization parameter used for encoding the second block is less than the global quantization parameter. When the second block is an occluded block, the target quantization parameter used for encoding the second block is greater than the global quantization parameter.
[0233] Optionally, in a hard encoder scenario, the prediction unit 701 is further configured to: embed deltaQP in the second block j , delaQP j which is used to determine the corresponding target quantization parameter when encoding the second block.
[0234] Optionally, the video sequence is included in the decoded video stream.
[0235] Optionally, the blocks in each video frame are macroblocks MB in the H.264 scenario, or the blocks in each video frame are prediction units PU in the H.265 scenario.
[0236] The functions of the units in the video encoding apparatus 70 described in this application can be understood by referring to the corresponding content in the foregoing method embodiments, and will not be repeated here.
[0237] In an embodiment of this application, a computer-readable storage medium is further provided. A program is stored in the computer-readable storage medium. When the program runs on a computer, the computer is caused to execute as described in the foregoing Figures 3 - 6 illustrated embodiments.
[0238] In an embodiment of this application, a video encoding apparatus is further provided. The video encoding apparatus may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is configured to execute the foregoing Figures 3 - 6 method shown in any one of the embodiments.
[0239] In an embodiment of this application, a digital processing chip is further provided. Circuits for implementing the functions of the foregoing encoder 200, processor 201, or encoder 200, processor 201 and one or more interfaces are integrated in the digital processing chip. When a memory is integrated in the digital processing chip, the digital processing chip can complete the method steps in any one or more of the foregoing embodiments. When a memory is not integrated in the digital processing chip, it can be connected to an external memory through the communication interface. The digital processing chip implements the method in the foregoing embodiments according to the program code stored in the external memory.
[0240] In an embodiment of this application, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium. When the instructions run on an electronic device, the electronic device is caused to execute as described aboveFigures 3 - 6 The steps in the method described in any of the foregoing embodiments.
[0241] An embodiment of the present application further provides a computer program product, which, when running on a computer, causes the computer to execute the steps in the method described in any of the foregoing Figures 3 - 6 embodiments.
[0242] The video encoding device provided in the embodiment of the present application may be a chip, which includes a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the computer device executes the video encoding method described in the foregoing Figures 3 - 6 embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0243] Specifically, the foregoing processing unit or processor may include a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0244] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0245] In addition, an embodiment of the present application further provides a device for storing a bitstream, including at least one storage medium and a communication interface; the communication interface is used to receive or send a bitstream; the at least one storage medium is used to store the bitstream; the bitstream is obtained by encoding using the video encoding method described in the above Figures 3 - 6 illustrated embodiment.
[0246] A tenth aspect of the present application provides a method for storing a bitstream, including: receiving a bitstream through a communication interface; storing the bitstream in one or more storage mediums, where the bitstream is obtained by encoding using the video encoding method described in the above Figures 3 - 6 illustrated embodiment.
[0247] An eleventh aspect of the present application provides a system for distributing a bitstream, including at least one storage medium and a video stream device; the at least one storage medium is used to store the bitstream, and the bitstream is obtained by encoding using the video encoding method described in the above Figures 3 - 6 illustrated embodiment;
[0248] The video stream device is configured to, in response to a request from a decoder, send the target bitstream in at least one storage medium to the decoder.
[0249] A twelfth aspect of the present application provides a method for distributing a bitstream, including: receiving a first request; in response to the first request, selecting a target bitstream from at least one storage medium; sending the target bitstream to a destination device; the at least one storage medium is used to store the bitstream, and the target bitstream is obtained by encoding using the video encoding method described in the above Figures 3 - 6 illustrated embodiment.
[0250] A thirteenth aspect of the present application provides a system for processing a bitstream, including an image source device, an encoder device, one or more storage mediums, and a destination device;
[0251] The image source device is used to provide image data;
[0252] The encoder device is used to obtain the image data of the image source device through an interface, and encode the image data to obtain one or more bitstreams, where the bitstreams are obtained by encoding using the video encoding method described in the embodiments shown in Figures 3 - 6 ;
[0253] The encoder device is used to store one or more bitstreams into one or more storage media; or the encoder device is used to encapsulate one or more bitstreams to obtain a transport bitstream;
[0254] The encoder device is used to transmit the transport bitstream to the destination device through a communication link or a communication network; the destination device is used to de-encapsulate the transport bitstream to obtain one or more bitstreams;
[0255] The destination device is used to decode one or more bitstreams to obtain decoded data.
[0256] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, read only memory (ROM), random access memory (RAM), magnetic disk or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.
[0257] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0258] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired means (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless means (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
Claims
1. A method for video encoding, characterized in that, When encoding a video sequence, the method includes: Predicting at least one target region in a pre-coded frame based on the motion information of a first moving object in an encoded frame, where the at least one target region is related to the first moving object; Encoding the video data corresponding to the at least one target region using a target quantization parameter that is different from the global quantization parameter of the video sequence.
2. The method according to claim 1, wherein When there is one target region, the target region is the first region or the second region; when there are multiple target regions, the multiple target regions include the first region and the second region; where The first region is the motion region of the first moving object in the pre-coded frame, and the target quantization parameter corresponding to the first region is less than the global quantization parameter; The second region is a region in the pre-coded frame and is occluded by the first moving object in the next frame of the pre-coded frame, and the target quantization parameter corresponding to the second region is greater than the global quantization parameter.
3. The method according to claim 2, wherein The encoded frame and the pre-coded frame are consecutive frames or non-consecutive frames in the video sequence.
4. The method according to claim 2 or 3, characterized in that, Each video frame of the video sequence includes multiple blocks. Predicting at least one target region in a pre-coded frame based on the motion information of a first moving object in an encoded frame includes: Predicting the position information of a motion block in the pre-coded frame based on the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-coded frame, and the position information of the motion block in the encoded frame; where the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the motion block in the pre-coded frame is the predicted block corresponding to the first moving object in the pre-coded frame; Determining the first region based on the position information of the motion block in the pre-coded frame.
5. The method according to claim 2 or 3, characterized in that, Each video frame of the video sequence includes multiple blocks. Predicting at least one target region in a pre-coded frame based on the motion information of a first moving object in an encoded frame includes: Predicting the position information of an occluded block in the pre-coded frame based on the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-coded frame, and the position information of the motion block in the encoded frame; where the motion block in the encoded frame is the block corresponding to the first moving object in the encoded frame, and the occluded block in the pre-coded frame is the predicted block that is occluded by the first moving object in the next frame of the pre-coded frame; Determining the second region based on the position information of the occluded block in the pre-coded frame.
6. The method according to claim 4, wherein Predicting the position information of a motion block in the pre-coded frame based on the first-order motion vector of the motion block in the encoded frame, the interval between the encoded frame and the pre-coded frame, and the position information of the motion block in the encoded frame includes: Determine the motion vector of the motion block in the pre-coded frame according to the first-order motion vector of the motion block in the coded frame and the interval between the coded frame and the pre-coded frame; wherein, the interval between the coded frame and the pre-coded frame is used to indicate the multiple relationship between the motion vector of the motion block in the pre-coded frame and the first-order motion vector; Determine the position information of the motion block in the pre-coded frame according to the position information of the motion block in the coded frame and the motion vector of the motion block in the pre-coded frame.
7. The method according to claim 5, characterized in that The predicting the position information of the occluded block in the pre-coded frame according to the first-order motion vector of the motion block in the coded frame, the interval between the coded frame and the pre-coded frame, and the position information of the motion block in the coded frame includes: Determine the motion vector of the motion block in the next frame of the pre-coded frame according to the first-order motion vector of the motion block in the coded frame and the interval between the coded frame and the pre-coded frame; wherein, the interval between the coded frame and the pre-coded frame is used to indicate the multiple relationship between the motion vector of the motion block in the next frame of the pre-coded frame and the first-order motion vector; Determine the position information of the occluded block in the pre-coded frame according to the position information of the motion block in the coded frame and the motion vector of the motion block in the next frame of the pre-coded frame.
8. The method according to any one of claims 4 to 7, characterized in that When there are multiple target regions, the method further includes: Determine the first quantity of the coordinate vectors of the motion regions pointing to the first block according to the position information of the motion blocks in the pre-coded frame, where the first block is any block in the coded frame; Determine the second quantity of the coordinate vectors of the occluded regions pointing to the first block according to the position information of the occluded blocks in the pre-coded frame; Determine the quantization parameter of the second block corresponding to the first block in the pre-coded frame according to the first quantity and the second quantity, and the quantization parameter of the second block is the target quantization parameter.
9. The method according to claim 8, wherein The determining the quantization parameter of the second block corresponding to the first block in the pre-coded frame according to the first quantity and the second quantity includes: Determine the change amount of the quantization parameter of the second block according to the first quantity and the second quantity; Determine the quantization parameter of the second block according to the change amount of the quantization parameter of the second block and the global quantization parameter.
10. The method according to claim 9, wherein The determining the change amount of the quantization parameter of the second block according to the first quantity and the second quantity includes: Determine the change amount of the quantization parameter of the second block according to the following relational expression; where deltaQP j represents the variation of the quantization parameter of the second block, A is a constant, ε is a constant, is the first quantity, N j cculusion is the second quantity.
11. The method according to any one of claims 8 to 10, characterized in that, The method further includes: Embed a descriptor in the second block, where the descriptor is used to indicate that the second block is a motion block or an occluded block, and when the second block is a motion block, the target quantization parameter used for encoding the second block is less than the global quantization parameter, and when the second block is an occluded block, the target quantization parameter used for encoding the second block is greater than the global quantization parameter.
12. The method according to claim 10, wherein In the scenario of a hard encoder, the method further includes: Embed deltaQP in the second block j , where the deltaQP j is used to determine the corresponding target quantization parameter when encoding the second block.
13. The method according to any one of claims 1-12, characterized in that, The video sequence is included in the decoded video stream.
14. The method according to any one of claims 4-12, characterized in that, Each block in each video frame is a macroblock MB in the H.264 scenario, or each block in each video frame is a prediction unit PU in the H.265 scenario.
15. A video encoding device, characterized in that, When encoding a video sequence, the apparatus includes: A prediction unit, configured to predict at least one target region in a pre-encoded frame according to motion information of a first moving object in an encoded frame, where the at least one target region is related to the first moving object; An encoding unit, configured to encode video data corresponding to the at least one target region using a target quantization parameter, where the target quantization parameter is different from a global quantization parameter of the video sequence.
16. An electronic device, characterized in that, Including: A communication interface, an encoder, a processor, and a memory. The communication interface is coupled to the encoder, the processor, and the memory. The memory is configured to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium. When the instructions run on an electronic device, the electronic device executes the method according to any one of claims 1 to 14.
18. A computer program product, characterized in that, The computer program product includes computer program code. When the computer program code runs on a computer, the computer executes the method according to any one of claims 1 to 14.
19. A chip system, characterized in that, The chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are configured to receive signals from a memory of an electronic device and send signals to the processors, where the signals include computer instructions stored in the memory; when the processors execute the computer instructions, the electronic device executes the method according to any one of claims 1 to 14.
20. A video processing system, characterized in that, Including: An electronic device, where the electronic device is configured to execute the method according to any one of claims 1 to 14.
21. A device for storing a bitstream, characterized in that, Including at least one storage medium and a communication interface; the communication interface is configured to receive or send a bitstream; the at least one storage medium is configured to store the bitstream; the bitstream is Encoded by an encoder according to the encoding method according to any one of claims 1 to 14.
22. A method for storing a bitstream, characterized in that, Including: Receiving a bitstream through the communication interface; Storing the bitstream in one or more storage mediums, where the bitstream is Encoded by an encoder according to the encoding method according to any one of claims 1 to 14.
23. A system for distributing a bitstream, characterized in that, Including at least one storage medium and a video stream device; the at least one storage medium is configured to store a bitstream, where the bitstream is Encoded by an encoder according to the encoding method according to any one of claims 1 to 14; The video stream device is configured to, in response to a request from a decoder, enable the target bitstream in the at least one storage medium to be sent to the decoder.
24. A method for distributing a bitstream, characterized in that, Including: Receiving a first request; In response to the first request, selecting a target bitstream from at least one storage medium; sending the target bitstream to a destination device; The at least one storage medium is configured to store a bitstream, where the target bitstream is encoded by an encoder according to the encoding method according to any one of claims 1 to 14.
25. A system for processing a bitstream, characterized in that, Including an image source device, an encoder device, one or more storage mediums, and a destination device; The image source device is used to provide image data; The encoder device is used to obtain the image data of the image source device through an interface, and encode the image data to obtain one or more bitstreams, where the bitstream is encoded by the encoder device according to any one of the encoding methods in claims 1 to 14; The encoder device is used to store the one or more bitstreams into one or more storage media; Alternatively, the encoder device is used to encapsulate the one or more bitstreams to obtain a transport stream; The encoder device is used to transmit the transport stream to the destination device through a communication link or a communication network; The destination device is used to de-encapsulate the transport stream to obtain the one or more bitstreams; The destination device is used to decode the one or more bitstreams to obtain decoded data.