Image prediction method and apparatus
By employing an affine motion model for inter prediction based on adjacent image unit information, the method addresses inefficiencies in conventional video compression, enhancing coding efficiency and reducing bitrate.
Patent Information
- Application Number
- JP2023176443
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-09-29
- Filing Date
- 2023-10-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2036-09-08
AI Technical Summary
Existing video compression technologies face inefficiencies due to the use of conventional translational motion models that fail to accurately capture complex motions in video sequences, leading to incomplete removal of frame correlation and increased bitrate for affine motion models.
The implementation of an affine motion model for inter prediction, where the prediction mode of an image unit is determined based on information from adjacent image units, reducing the need for explicit encoding of motion information and optimizing the set of candidate prediction modes to improve coding efficiency.
This approach enhances coding efficiency by reducing the bitrate required for encoding prediction modes, thereby improving the accuracy of motion compensation and overall video compression performance.
Smart Images

Figure 0007703829000001 
Figure 0007703829000002 
Figure 0007703829000003
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video encoding and compression, and more particularly, to an image prediction method and apparatus.
Background Art
[0002] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, video conferencing devices, video streaming devices, and the like. Digital video devices implement video compression techniques such as those described by the MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10: the Advanced Video Coding (AVC) standard, and the ITU-T H.265: the High Efficiency Video Coding (HEVC) standard, and extensions to such standards, to more efficiently transmit and receive digital video information. By implementing such video encoding techniques, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.
[0003] Video compression technology includes spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, which reduces or removes the inherent redundancy in a video sequence. For block-based video coding, a video slice (i.e., a video frame or a part of a video frame) may be partitioned into several video blocks. A video block may also be referred to as a tree block, a coding unit (CU), and / or a coding node. Video blocks in an intra-coded (I) slice of a picture are coded by spatial prediction against reference samples in adjacent blocks of the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction against reference samples in adjacent blocks of the same picture or temporal prediction against reference samples in another reference picture. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame.
[0004] Spatial or temporal prediction results in a predicted block of the block to be coded. Residual data indicates the pixel difference between the original block to be coded and the predicted block. An inter-coded block is coded according to a motion vector indicating a block of reference samples forming the predicted block and residual data indicating the difference between the coded block and the predicted block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from a pixel domain to a transform domain, thereby generating residual transform coefficients. The residual transform coefficients may then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, may be sequentially scanned to generate a one-dimensional vector of transform coefficients, and entropy coding may be applied to achieve further compression. SUMMARY OF THE INVENTION
[0005] The present invention describes an image prediction method for improving coding efficiency. The prediction mode of the image unit to be processed is derived according to the prediction information or unit size of the adjacent image unit of the image unit to be processed, or a set of candidates for the prediction mode for marking the region level. Since the previous information is provided for the coding of the prediction mode, the bit rate for coding the prediction mode is reduced, thereby improving the coding efficiency.
[0006] According to the technology of the present invention, a method for decoding a predicted image includes determining whether a set of candidates for the prediction mode of the image unit to be processed includes an affine merge mode according to information about adjacent image units adjacent to the image unit to be processed, where the affine merge mode indicates that the predicted images of the image unit to be processed and each of the adjacent image units of the image unit to be processed are obtained by using the same affine mode; analyzing a bitstream to obtain first indication information; determining the prediction mode of the image unit to be processed according to the first indication information in the set of candidates for the prediction mode; and determining the predicted image of the image unit to be processed according to the prediction mode.
[0007] The adjacent image units of the image unit to be processed include at least the adjacent image units above, to the left, upper right, lower left, and upper left of the image unit to be processed.
[0008] According to the technology of the present invention, the step of determining whether a set of candidates for the prediction mode of the image unit to be processed includes an affine merge mode according to information about adjacent image units adjacent to the image unit to be processed includes the following implementation methods.
[0009] The first implementation method includes a step of analyzing a bitstream to obtain second instruction information when at least one prediction mode of adjacent image units obtains a predicted image by using an affine model. When the second instruction information is 1, the set of candidate prediction modes includes an affine merge mode, or when the second instruction information is 0, the set of candidate prediction modes does not include an affine merge mode, or otherwise, the set of candidate prediction modes does not include an affine merge mode.
[0010] The second implementation method includes that when at least one prediction mode of adjacent image units obtains a predicted image by using an affine model, the set of candidate prediction modes includes an affine merge mode, or otherwise, the set of candidate prediction modes does not include an affine merge mode.
[0011] The third implementation method includes the following. The prediction mode includes at least a first affine mode for obtaining a predicted image by using at least a first affine model, or a second affine mode for obtaining a predicted image by using a second affine model. Correspondingly, the affine merge mode includes at least a first affine merge mode for merging the first affine mode, or a second affine merge mode for merging the second affine mode. Correspondingly, the step of determining whether the set of candidate prediction modes of the image unit to be processed includes the affine merge mode according to the information about the adjacent image unit adjacent to the image unit to be processed is as follows: When the first affine mode ranks first in terms of quantity among the prediction modes of the adjacent prediction units, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode; when the second affine mode ranks first in terms of quantity among the prediction modes of the adjacent prediction units, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode; or when the non-affine mode ranks first in terms of quantity among the prediction modes of the adjacent prediction units, the set of candidate prediction modes does not include the affine merge mode.
[0012] The third implementation method further includes the following. Among the prediction modes of the adjacent prediction unit, when the first affine mode ranks first in terms of quantity, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode, or when the second affine mode ranks first in terms of quantity among the prediction modes of the adjacent prediction unit, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode, or when the non-affine mode ranks first in terms of quantity among the prediction modes of the adjacent prediction unit and the first affine mode ranks second in terms of quantity, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode, or when the non-affine mode ranks first in terms of quantity among the prediction modes of the adjacent prediction unit and the second affine mode ranks second in terms of quantity, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode.
[0013] The fourth implementation method includes a step of analyzing a bitstream to obtain third instruction information when at least one prediction mode of at least one of the adjacent image units obtains a predicted image by using an affine model and at least one of the widths and heights of at least one of the adjacent image units is smaller than the width and height of the image unit to be processed, and when the third instruction information is 1, the set of candidate prediction modes includes the affine merge mode, or when the third instruction information is 0, the set of candidate prediction modes does not include the affine merge mode, or otherwise, the set of candidate prediction modes does not include the affine merge mode.
[0014] The fifth implementation method includes the following. At least one prediction mode of the adjacent image units obtains a predicted image by using an affine model, and when at least one of the widths and heights of the adjacent image units is smaller than the width and height of the image unit to be processed respectively, the set of candidate prediction modes includes the affine merge mode, or otherwise, the set of candidate prediction modes does not include the affine merge mode.
[0015] According to the technology of the present invention, a method for decoding a predicted image includes: analyzing a bitstream to obtain first indication information; determining a set of candidate modes of the first image region to be processed according to the first indication information, wherein when the first indication information is 0, a set of candidate translational modes indicating a prediction mode of obtaining a predicted image by using a translational model is used as the set of candidate modes of the first image region to be processed, or when the first indication information is 1, a set of candidate translational modes and a set of candidate affine modes indicating a prediction mode of obtaining a predicted image by using an affine model are used as the set of candidate modes of the first image region to be processed; analyzing the bitstream to obtain second indication information; determining a prediction mode of the image unit to be processed according to the second indication information in the set of candidate prediction modes of the first image region to be processed, wherein the image unit to be processed belongs to the first image region to be processed; and determining a predicted image of the image unit to be processed according to the prediction mode.
[0016] The first image region to be processed includes one of an image frame group, an image frame, an image tile set, an image slice set, an image tile, an image slice, an image coding unit set, or an image coding unit.
[0017] In one example, a method for decoding a predicted image includes determining, according to information about an adjacent image unit adjacent to an image unit to be processed, whether a set of candidate prediction modes of the image unit to be processed includes an affine merge mode, where the affine merge mode indicates that the predicted image of each of the image unit to be processed and the adjacent image unit adjacent to the image unit to be processed is obtained by using the same affine mode; parsing a bitstream to obtain first indication information; determining, in the set of candidate prediction modes, the prediction mode of the image unit to be processed according to the first indication information; and determining the predicted image of the image unit to be processed according to the prediction mode.
[0018] In another example, a method for encoding a predicted image includes determining, according to information about an adjacent image unit adjacent to an image unit to be processed, whether a set of candidate prediction modes of the image unit to be processed includes an affine merge mode, where the affine merge mode indicates that the predicted image of each of the image unit to be processed and the adjacent image unit adjacent to the image unit to be processed is obtained by using the same affine mode; determining, in the set of candidate prediction modes, the prediction mode of the image unit to be processed; determining the predicted image of the image unit to be processed according to the prediction mode; and encoding first indication information into a bitstream, where the first indication information indicates the prediction mode.
[0019] In another example, a method for decoding a predicted image includes: analyzing a bitstream to obtain first instruction information; determining, according to the first instruction information, a set of candidate modes of the first image region to be processed, wherein when the first instruction information is 0, a set of candidate translational modes indicating a prediction mode for obtaining a predicted image by using a translational model is used as the set of candidate modes of the first image region to be processed, or when the first instruction information is 1, a set of candidate translational modes and a set of candidate affine modes indicating a prediction mode for obtaining a predicted image by using an affine model are used as the set of candidate modes of the first image region to be processed; analyzing the bitstream to obtain second instruction information; determining, according to the second instruction information, a prediction mode of an image unit to be processed in the set of candidate prediction modes of the first image region to be processed, wherein the image unit to be processed belongs to the first image region to be processed; and determining a predicted image of the image unit to be processed according to the prediction mode.
[0020] In another example, a method for encoding a predicted image includes: when a set of translational mode candidates indicating a prediction mode for obtaining a predicted image by using a translational model is used as a set of mode candidates for a first image region to be processed, setting first indication information to 0 and encoding the first indication information into a bitstream; or when the set of translational mode candidates and a set of affine mode candidates indicating a prediction mode for obtaining a predicted image by using an affine model are used as a set of mode candidates for a first image region to be processed, setting the first indication information to 1 and encoding the first indication information into a bitstream; determining a prediction mode of an image unit to be processed in a set of mode candidates for a prediction mode of the first image region to be processed, where the image unit to be processed belongs to the first image region to be processed; determining a predicted image of the image unit to be processed according to the prediction mode; and encoding second indication information into the bitstream, where the second indication information indicates the prediction mode.
[0021] In another example, an apparatus for decoding a predicted image includes: a first determination module configured to determine whether a set of mode candidates for a prediction mode of an image unit to be processed includes an affine merge mode according to information about adjacent image units adjacent to the image unit to be processed, where the affine merge mode indicates that predicted images of the image unit to be processed and each of the adjacent image units adjacent to the image unit to be processed are obtained by using the same affine mode; an analysis module configured to analyze a bitstream to obtain first indication information; a second determination module configured to determine a prediction mode of the image unit to be processed according to the first indication information in the set of mode candidates for the prediction mode; and a third determination module configured to determine a predicted image of the image unit to be processed according to the prediction mode.
[0022] In another example, an apparatus for encoding a predicted image includes a first determination module configured to determine whether a set of candidate prediction modes of an image unit to be processed includes an affine merge mode according to information about an adjacent image unit adjacent to the image unit to be processed, where the affine merge mode indicates that the predicted images of the image unit to be processed and the adjacent image unit of the image unit to be processed are obtained by using the same affine mode; a second determination module configured to determine a prediction mode of the image unit to be processed in the set of candidate prediction modes; a third determination module configured to determine a predicted image of the image unit to be processed according to the prediction mode; and an encoding module configured to encode first indication information into a bitstream, where the first indication information indicates the prediction mode.
[0023] In another example, an apparatus for decoding a predicted image includes a first analysis module configured to analyze a bitstream to obtain first instruction information, and a first determination module configured to determine a set of candidate modes of a first image region to be processed according to the first instruction information. When the first instruction information is 0, a set of candidate translational modes indicating a translational mode for obtaining a predicted image by using a translational model is used as the set of candidate modes of the first image region to be processed, or when the first instruction information is 1, a set of candidate translational modes and a set of candidate affine modes indicating an affine mode for obtaining a predicted image by using an affine model are used as the set of candidate modes of the first image region to be processed. A second analysis module configured to analyze the bitstream to obtain second instruction information, and a second determination module configured to determine a predicted mode of an image unit to be processed according to the second instruction information in the set of candidate predicted modes of the first image region to be processed, where the image unit to be processed belongs to the first image region to be processed. And a third determination module configured to determine a predicted image of the image unit to be processed according to the predicted mode.
[0024] In another example, when a set of candidates for a translational mode indicating a prediction mode for obtaining a predicted image by using a translational model is used as a set of candidates for the mode of a first image region to be processed, a first encoding module is configured to set first instruction information to 0 and encode the first instruction information into a bitstream, or when a set of candidates for the translational mode and a set of candidates for an affine mode indicating a prediction mode for obtaining a predicted image by using an affine model are used as a set of candidates for the mode of the first image region to be processed, the first encoding module is configured to set the first instruction information to 1 and encode the first instruction information into the bitstream; a first determination module configured to determine a prediction mode of an image unit to be processed in a set of candidates for the prediction mode of the first image region to be processed, where the image unit to be processed belongs to the first image region to be processed; a second determination module configured to determine a predicted image of the image unit to be processed according to the prediction mode; and a second encoding module configured to encode second instruction information into the bitstream, where the second instruction information indicates the prediction mode.
[0025] In another example, a device for decoding video data is provided. The device is configured to perform an operation of determining whether a set of candidates for the prediction mode of an image unit to be processed includes an affine merge mode according to information about adjacent image units adjacent to the image unit to be processed, where the affine merge mode indicates that the prediction images of the image unit to be processed and each of the adjacent image units adjacent to the image unit to be processed are obtained by using the same affine mode; an operation of analyzing a bitstream to obtain first instruction information; an operation of determining the prediction mode of the image unit to be processed according to the first instruction information in the set of candidates for the prediction mode; and an operation of determining the predicted image of the image unit to be processed according to the prediction mode, and includes a video decoder.
[0026] In another example, a device for encoding video data is provided. The device includes an operation of determining whether a set of candidate prediction modes of an image unit to be processed includes an affine merge mode according to information about an adjacent image unit adjacent to the image unit to be processed, where the affine merge mode indicates that respective prediction images of the image unit to be processed and the adjacent image unit to be processed are obtained by using the same affine mode; an operation of determining a prediction mode of the image unit to be processed in the set of candidate prediction modes; an operation of determining a prediction image of the image unit to be processed according to the prediction mode; and an operation of encoding first indication information into a bitstream, where the first indication information indicates the prediction mode, and is configured to include a video encoder that executes the operations.
[0027] In another example, a device for decoding video data is provided. The device includes an operation of analyzing a bitstream to obtain first indication information; an operation of determining a set of candidate modes of a first image area to be processed according to the first indication information, where when the first indication information is 0, a set of candidate translational modes indicating a prediction mode of obtaining a prediction image by using a translational model is used as the set of candidate modes of the first image area to be processed, or when the first indication information is 1, the set of candidate translational modes and a set of candidate affine modes indicating a prediction mode of obtaining a prediction image by using an affine model are used as the set of candidate modes of the first image area to be processed; an operation of analyzing the bitstream to obtain second indication information; an operation of determining a prediction mode of an image unit to be processed according to the second indication information in the set of candidate prediction modes of the first image area to be processed, where the image unit to be processed belongs to the first image area to be processed; and an operation of determining a prediction image of the image unit to be processed according to the prediction mode, and is configured to include a video decoder that executes the operations.
[0028] In another example, a device for encoding video data is provided. When the set of candidates for the translational mode indicating a prediction mode for obtaining a predicted image by using a translational model is used as the set of candidates for the mode of the first image region to be processed, the device sets the first instruction information to 0 and encodes the first instruction information into a bitstream, or when the set of candidates for the translational mode and the set of candidates for the affine mode indicating a prediction mode for obtaining a predicted image by using an affine model are used as the set of candidates for the mode of the first image region to be processed, the device sets the first instruction information to 1 and encodes the first instruction information into a bitstream, and in the set of candidates for the prediction mode of the first image region to be processed, an operation of determining the prediction mode of the image unit to be processed, wherein the image unit to be processed belongs to the first image region to be processed, an operation of determining the predicted image of the image unit to be processed according to the prediction mode, and an operation of encoding second instruction information into a bitstream, wherein the second instruction information indicates the prediction mode. The device includes a video encoder configured to execute the operations.
[0029] In another example, a computer-readable storage medium storing instructions is provided. When executed, the instructions cause one or more processors of a device for decoding video data to perform an operation of determining whether the set of candidates for the prediction mode of the image unit to be processed includes an affine merge mode according to information about adjacent image units adjacent to the image unit to be processed, where the affine merge mode indicates that the predicted images of the image unit to be processed and each of the adjacent image units to be processed are obtained by using the same affine mode, an operation of parsing a bitstream to obtain first instruction information, an operation of determining the prediction mode of the image unit to be processed according to the first instruction information in the set of candidates for the prediction mode, and an operation of determining the predicted image of the image unit to be processed according to the prediction mode.
[0030] In another example, a computer-readable storage medium storing instructions is provided. When executed, the instructions perform an operation of determining whether a set of candidate prediction modes of an image unit to be processed includes an affine merge mode according to information about an adjacent image unit adjacent to the image unit to be processed, where the affine merge mode indicates that respective predicted images of the image unit to be processed and the adjacent image unit adjacent to the image unit to be processed are obtained by using the same affine mode; an operation of determining a prediction mode of the image unit to be processed in the set of candidate prediction modes; an operation of determining a predicted image of the image unit to be processed according to the prediction mode; and an operation of encoding first instruction information indicating the prediction mode into a bitstream, and cause one or more processors of a device for encoding video data to execute the operations.
[0031] In another example, a computer-readable storage medium storing instructions is provided. When executed, the instructions perform operations of analyzing a bitstream to obtain first instruction information, and determining a set of candidate modes of a first image region to be processed according to the first instruction information, wherein when the first instruction information is 0, a set of candidate translational modes indicating a prediction mode of obtaining a predicted image by using a translational model is used as the set of candidate modes of the first image region to be processed, or when the first instruction information is 1, a set of candidate translational modes and a set of candidate affine modes indicating a prediction mode of obtaining a predicted image by using an affine model are used as the set of candidate modes of the first image region to be processed; operations of analyzing the bitstream to obtain second instruction information; an operation of determining a prediction mode of an image unit to be processed according to the second instruction information in the set of candidate prediction modes of the first image region to be processed, wherein the image unit to be processed belongs to the first image region to be processed; and an operation of determining a predicted image of the image unit to be processed according to the prediction mode, and causing one or more processors of a device for decoding video data to execute the operations.
[0032] In another example, a computer-readable storage medium for storing instructions is provided. When executed, the instructions set the first instruction information to 0 and encode the first instruction information into a bitstream if a set of candidates for a translational mode indicating a prediction mode for obtaining a predicted image by using a translational model is used as a set of candidates for the mode of a first image region to be processed, or set the first instruction information to 1 and encode the first instruction information into a bitstream if a set of candidates for a translational mode and a set of candidates for an affine mode indicating a prediction mode for obtaining a predicted image by using an affine model are used as a set of candidates for the mode of a first image region to be processed, and an operation of determining a prediction mode of an image unit to be processed in a set of candidates for the prediction mode of a first image region to be processed, where the image unit to be processed belongs to the first image region to be processed, an operation of determining a predicted image of the image unit to be processed according to the prediction mode, and an operation of encoding second instruction information into a bitstream, where the second instruction information indicates the prediction mode, are executed by one or more processors of a device for encoding video data.
Brief Description of the Drawings
[0033] To more clearly explain the technical solution means in the embodiments of the present invention, the following briefly explains the attached drawings required for explaining the embodiments or the prior art. Obviously, the attached drawings in the following description only show some embodiments of the present invention, and those skilled in the art can derive other drawings from these attached drawings without creative efforts.
[0034]
Figure 1
[0035]
Figure 2
[0036]
Figure 3
[0037]
Figure 4
[0038]
Figure 5
[0039]
Figure 6
[0040]
Figure 7
[0041]
Figure 8
[0042]
Figure 9
[0043]
Figure 10
[0044]
Figure 11
[0045]
Figure 12
Embodiments for Carrying Out the Invention
[0046] Hereinafter, with reference to the accompanying drawings of the embodiments of the present invention, the technical solution in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative efforts based on the embodiments of the present invention shall be within the protection scope of the present invention.
[0047] Motion compensation is one of the important techniques for improving compression efficiency in video coding. Conventional motion compensation based on block matching is a method widely applied to mainstream video encoders, especially in video coding standards. In the motion compensation method based on block matching, an inter-prediction block uses a translational motion model, and the translational motion model assumes that the motion vectors at all pixel positions of a block are equal. However, this assumption is invalid in many cases. In fact, the motion of an object in a video is usually a complex combination of motions such as translational motion, rotational motion, and zoom. When a pixel block includes these complex motions, the prediction signal obtained using the conventional motion compensation method based on block matching is inaccurate. As a result, the correlation between frames cannot be completely removed. To solve this problem, a higher-order motion model is introduced into the motion compensation of video coding. The higher-order motion model has a higher degree of freedom than the translational motion model and enables the pixels of the inter-prediction block to have different motion vectors. That is, the motion vector field generated by the higher-order motion model is more accurate.
[0048] The affine motion model described based on control points is a typical type of the upper motion model. Different from the conventional translational motion model, the value of the motion vector of each pixel point in the block is related to the position of the pixel point and is a first-order linear equation of the coordinate position. The affine motion model enables distortion transformations such as rotation or zoom of the reference block, and through motion compensation, a more accurate predicted block can be obtained.
[0049] Through motion compensation, the above-mentioned type of inter prediction that obtains a predicted block by using the affine motion model is generally referred to as the affine mode. In the current mainstream video compression encoding standard, the types of inter prediction include two modes such as the Advanced motion vector prediction (AMVP) mode and the Merge mode. In AMVP, for each encoded block, the prediction direction, the reference frame index, and the difference between the actual motion vector and the predicted motion vector need to be explicitly transferred. However, in the Merge mode, the motion information of the current encoded block is directly derived from the motion vectors of adjacent blocks. The affine mode and the inter prediction methods such as AMVP or Merge based on the translational motion model can be combined well to form a new inter prediction mode such as AMVP or Merge based on the affine motion model. For example, the Merge mode based on the affine motion model may be referred to as the Affine Merge mode. In the process of selecting the prediction mode, the new prediction mode and the prediction modes in the current standard both participate in the comparison process of "performance / cost ratio", select the optimal mode as the prediction mode, and generate the predicted image of the block to be processed. Generally, the prediction mode selection result is encoded, and the encoded prediction mode selection result is transmitted to the decoding side.
[0050] The affine mode can further improve the accuracy of the prediction block and can improve the coding efficiency. However, on the other hand, for the affine mode, in order to encode the motion information of the control points, it is necessary to consume more bitrate than the bitrate required for the uniform motion information based on the translational motion model. In addition, since the number of candidate prediction modes increases, the bitrate used to encode the prediction mode selection result also increases. All such additional bitrate consumption affects the improvement of the coding efficiency.
[0051] According to the technical solution of the present invention, on the one hand, whether the set of candidates for the prediction mode of the image unit to be processed includes the affine merge mode is determined according to the prediction mode information or size information of the adjacent image unit of the image unit to be processed. To obtain the instruction information, the bitstream is analyzed. In the set of candidates for the prediction mode, the prediction mode of the image unit to be processed is determined according to the instruction information, and the predicted image of the image unit to be processed is determined according to the prediction mode. On the other hand, the bitstream is analyzed, and whether to use a set of candidates for the prediction mode including the affine mode in a specific region is determined by using the instruction information. According to the set of candidates for the prediction mode and other received instruction information, the prediction mode is determined, and a predicted image is generated.
[0052] Therefore, the prediction mode information or size information of the adjacent image unit of the image unit to be processed may be used as preliminary knowledge for encoding the prediction information of the block to be processed. The instruction information including the set of candidates for the prediction mode in the region may also be used as preliminary knowledge for encoding the prediction information of the block to be processed. The preliminary knowledge instructs the encoding of the prediction mode and reduces the bitrate of the information for encoding mode selection, thereby improving the coding efficiency.
[0053] In addition, for example, there are a plurality of solutions for improving the efficiency in encoding the motion information of the affine model in Patent Application Nos. CN201010247275.7, CN201410584175.1, CN201410526608.8, CN201510085362.X, PCT / CN2015 / 073969, CN201510249484.8, CN201510391765.7, and CN201510543542.8, etc. The entire contents of these applications are incorporated herein by reference. Since the specific technical problems to be solved are different, the technical solution of the present invention may be applied to the above-mentioned solutions, and further, it should be understood that the encoding efficiency is improved.
[0054] It should be further understood that the affine model is a general term for a non-translational motion model. All actual motions including rotation, zoom, deformation, perspective, and the like may be used for motion estimation and motion compensation in inter-prediction by establishing different motion models, and are separately referred to as, for example, the first affine model and the second affine model.
[0055] FIG. 1 is a schematic block diagram of a video encoding system 10 according to an embodiment of the present invention. As described herein, the term "video coder" generally refers to both a video encoder and a video decoder. In the present invention, the term "video encoding" or "encoding" may generally refer to video encoding or video decoding.
[0056] As shown in FIG. 1, video encoding system 10 includes a source device 12 and a destination device 14. The source device 12 generates encoded video data. Thus, the source device 12 may be referred to as a video encoding device or a video encoding apparatus. The destination device 14 can decode the encoded video data generated by the source device 12. Thus, the destination device 14 may be referred to as a video decoding device or a video decoding apparatus. The source device 12 and the destination device 14 may be examples of a video encoding device or a video encoding apparatus. The source device 12 and the destination device 14 may include a wide range of devices including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or the like.
[0057] The destination device 14 may receive the encoded video data from the source device 12 by using a channel 16. The channel 16 may include one or more media and / or devices capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the channel 16 may include one or more communication media that enable the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. In this example, the source device 12 may modulate the encoded video data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated video data to the destination device 14. The one or more communication media may include wireless and / or wired communication media such as, for example, the radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from the source device 12 to the destination device 14.
[0058] In another example, channel 16 may include a storage medium that stores the encoded video data generated by source device 12. In this example, destination device 14 may access the storage medium by disk access or card access. The storage medium may include various locally accessible data recording media such as Blu-ray Disc (registered trademark), DVD, CD-ROM, flash memory, or other suitable digital storage media for storing the encoded video data.
[0059] In another example, channel 16 may include a file server or another intermediate storage device that stores the encoded video data generated by source device 12. In this example, destination device 14 may access the encoded video data stored in the file server or another intermediate storage device by streaming or downloading. The file server may be a type of server that stores the encoded video data and can transmit the encoded video data to destination device 14. Examples of file servers include web servers (e.g., of a website), File Transfer Protocol (FTP) servers, Network Attached Storage (NAS) devices, and local disk drives.
[0060] Destination device 14 may access the encoded video data via a standard data connection (such as an Internet connection). Exemplary types of data connections may include wireless channels (e.g., Wi-Fi (registered trademark) connection), wired connections (e.g., DSL or cable modem), or a combination of both suitable for accessing the encoded video data stored in the file server. The transmission of the encoded video data from the file server may be streaming, downloading, or a combination of both.
[0061] The technology of the present invention is not limited to wireless applications or settings. The technology may be applied to video encoding, for example, in support of various multimedia applications such as terrestrial television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., over the Internet), encoding of video data stored on a data recording medium, decoding of video data stored on a data recording medium, or other applications. In some examples, the video encoding system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0062] In the example of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. In some examples, the output interface 22 may include a modulator / demodulator (modem) and / or a transmitter. The video source 18 may include a video capture device (such as a video camera), a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources of video data.
[0063] The video encoder 20 may encode the video data from the video source 18. In some examples, the source device 12 directly transmits the encoded video data to the destination device 14 by using the output interface 22. Alternatively, the encoded video data may be stored on a storage medium or a file server for later access by the destination device 14 for decoding and / or playback.
[0064] In the example of FIG. 1, the destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. In some examples, the input interface 28 includes a receiver and / or a modem. The input interface 28 may receive encoded video data by using channel 16. The display device 32 may be integrated with the destination device 14 or may be external to the destination device 14. Generally, the display device 32 displays the decoded video data. The display device 32 may include various display devices such as, for example, a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0065] The video encoder 20 and the video decoder 30 may operate according to a video compression standard (such as the High Efficiency Video Coding (H.265) standard) and comply with the HEVC Test Model (HM). The text description of the H.265 standard, ITU-T H.265 (V3) (04 / 2015), was published on April 29, 2015 and can be downloaded from http: / / handle.itu.int / 11.1002 / 1000 / 12455. The entire content of the file is incorporated herein by reference.
[0066] Alternatively, the video encoder 20 and the video decoder 30 may operate according to other proprietary or industry standards. The standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262, or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, and ITU-T H.264 (also referred to as ISO / IEC MPEG-4 AVC), and their scalable video coding (SVC) and multi-view video coding (MVC) extensions. However, the technology of the present invention is not limited to any specific coding standard or technology.
[0067] In addition, FIG. 1 is only an example of the technology of the present invention and may be applied to video encoding settings (for example, video encoding or video decoding) that do not necessarily include any data communication between the encoding device and the decoding device. In other examples, the data may be read from local memory, streamed across a network, or operated in a similar manner. The encoding device may encode the data and store the encoded data in memory, and / or the decoding device may read the data from memory and decode the data. In many examples, encoding and decoding are performed by multiple devices that do not communicate with each other but simply encode data into memory and / or read data from memory and decode the data.
[0068] The video encoder 20 and the video decoder 30 may each be implemented as any of a variety of suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology is implemented partially in software, the device may store software instructions in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware by using one or more processors to perform the technology of the present invention. Any of the above (including hardware, software, combinations of hardware and software, or the like) may be considered as one or more processors. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and either the video encoder 20 or the video decoder 30 may be integrated as part of an encoder / decoder (CODEC) combined in each device.
[0069] The present invention generally may refer to a video encoder 20 "signaling" certain information to another device (such as a video decoder 30). The term "signaling" generally may refer to the communication of syntax elements and / or other data representing encoded video data. Such communication may occur in real time or near real time. Alternatively, such communication may occur over a period of time. For example, communication may occur when syntax elements are stored as an encoded bitstream on a computer-readable storage medium while being encoded. The syntax elements may then be read by a decoding device at any time after being stored on this medium.
[0070] As briefly mentioned above, the video encoder 20 encodes video data. The video data may include one or more pictures. Each of the pictures may be a still image. In some examples, the pictures may be referred to as video "frames". The video encoder 20 may generate a bitstream, which includes a sequence of bits forming an encoded representation of the video data. The encoded representation of the video data may include encoded pictures and associated data. The encoded pictures are the encoded representations of the pictures. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and another syntax structure. The SPS may include parameters applicable to zero or more sequences of pictures. The PPS may include parameters applicable to zero or more pictures. The syntax structure may be a set of zero or more syntax elements that are displayed together in a specified order in the bitstream.
[0071] To generate an encoded representation of a picture, video encoder 20 may partition the picture into a grid of coding tree blocks (CTBs). In some examples, a CTB may be referred to as a "tree block", "largest coding unit" (LCU), or "coding tree unit". The CTBs of HEVC may be roughly similar to the macroblocks of prior standards (such as H.264 / AVC). However, a CTB is not necessarily limited to a specific size and may include one or more coding units (CUs).
[0072] Each of the CTBs may be associated with different pixels of blocks of the same size within the picture. Each pixel may include a luminance sample and two chrominance samples. Thus, each CTB may be associated with one block of luminance samples and two blocks of chrominance samples. For ease of explanation, in the present invention, a two-dimensional pixel array may be referred to as a pixel block, and a two-dimensional sample array may be referred to as a sample block. Video encoder 20 may partition the pixel blocks associated with the CTBs into pixel blocks associated with CUs by a quadtree partition, and thus they are named "coding tree blocks".
[0073] The CTBs of a picture may be grouped into one or more slices. In some examples, each of the slices includes an integer number of CTBs. As part of picture encoding, video encoder 20 may generate an encoded representation (i.e., the encoded slice of each slice of the picture). To generate the encoded slice, video encoder 20 may encode each CTB of the slice to generate an encoded representation (i.e., the encoded CTB of each CTB of the slice).
[0074] To generate the symbolized CTB, the video encoder 20 may recursively perform a quadtree partitioning on the pixel block associated with the CTB to partition it into pixel blocks that become progressively smaller. Each of the smaller pixel blocks may be associated with a CU. The partitioned CU may be a CU that is partitioned into pixel blocks where the pixel blocks are associated with another CU. The non-partitioned CU may be a CU that is not partitioned into pixel blocks where the pixel blocks are associated with another CU.
[0075] The video encoder 20 may generate one or more prediction units (PUs) for each non-partitioned CU. Each of the PUs of the CU may be associated with different pixel blocks in the pixel block of the CU. The video encoder 20 may generate a predicted pixel block for each PU of the CU. The predicted pixel block of the PU may be a pixel block.
[0076] The video encoder 20 may generate the predicted pixel block of the PU by intra prediction or inter prediction. When the video encoder 20 generates the predicted pixel block of the PU by intra prediction, the video encoder 20 may generate the predicted pixel block of the PU based on the decoded pixels of the picture associated with the PU. When the video encoder 20 generates the predicted pixel block of the PU by inter prediction, the video encoder 20 may generate the predicted pixel block of the PU based on the decoded pixels of one or more pictures other than the picture associated with the PU.
[0077] The video encoder 20 may generate a residual pixel block of the CU based on the predicted pixel block of the PU of the CU. The residual pixel block of the CU may indicate the difference between the samples in the predicted pixel block of the PU of the CU and the corresponding samples in the original pixel block of the CU.
[0078] In addition, as part of the encoding of a non-partitioned CU, the video encoder 20 may perform a recursive quadtree partitioning on the residual pixel block of the CU to partition the residual pixel block of the CU into one or more smaller residual pixel blocks associated with the transform unit (TU) of the CU. Since each pixel in the pixel block associated with the TU includes a luminance sample and two chrominance samples, each of the TUs may be associated with a residual sample block of luminance samples and two residual sample blocks of chrominance samples.
[0079] The video coder 20 may apply one or more transforms to the residual sample block associated with the TU to generate a coefficient block (i.e., a block of coefficients). The video encoder 20 may perform quantization processing on each of the coefficient blocks. Quantization generally refers to the process of quantizing the coefficients to reduce, as much as possible, the amount of data used to represent the coefficients, thereby performing further compression.
[0080] The video encoder 20 may generate a set of syntax elements representing the coefficients in the quantized coefficient block. The video encoder 20 may apply an entropy encoding operation (such as a context adaptive binary arithmetic coding (CABAC) operation) to at least some of these syntax elements.
[0081] To apply CABAC encoding to the syntax elements, the video encoder 20 may binarize the syntax elements to form a binary digit sequence including one or more bits (referred to as "bins"). The video encoder 20 may encode some of the bins by regular CABAC encoding and other of the bins by bypass encoding.
[0082] When video encoder 20 encodes a sequence of bins by regular CABAC coding, video encoder 20 may first identify an encoding context. The encoding context may identify the probability of an encoded bin having a particular value. For example, the encoding context may indicate that the probability of encoding a bin with a value of 0 is 0.7 and the probability of encoding a bin with a value of 1 is 0.3. After identifying the encoding context, video encoder 20 may divide the interval into a lower sub-interval and an upper sub-interval. One sub-interval may be associated with the value 0 and the other sub-interval may be associated with the value 1. The width of the sub-intervals may be proportional to the probabilities indicated for the associated values by the identified encoding context.
[0083] If the bin of the syntax element has a value associated with the lower sub-interval, the encoded value may be equal to the lower boundary of the lower sub-interval. If the same bin of the syntax element has a value associated with the upper sub-interval, the encoded value may be equal to the lower boundary of the upper sub-interval. To encode the next bin of the syntax element, video encoder 20 may repeat these steps for the interval that contains the sub-interval associated with the value of the encoded bit. When video encoder 20 repeats these steps for the next bin, video encoder 20 may use the probabilities modified based on the probabilities indicated by the identified encoding context and the actual value of the encoded bin.
[0084] When the video encoder 20 encodes a bin sequence by bypass encoding, the video encoder 20 may be able to encode several bins in a single cycle. However, when the video encoder 20 encodes a bin sequence by regular CABAC encoding, the video encoder 20 may be able to encode only a single bin in one cycle. Bypass encoding may be relatively simple because in bypass encoding, the video encoder 20 does not need to select a context, and the video encoder 20 can assume that the probabilities for both symbols (0 and 1) are 1 / 2 (50%). Therefore, in bypass encoding, the interval is directly divided in half. In fact, bypass encoding bypasses the context adaptation part of the arithmetic coding engine.
[0085] The execution of bypass encoding for a bin requires less computation compared to the execution of regular CABAC encoding for the bin. In addition, the execution of bypass encoding can enable higher parallelism and higher throughput. The bins encoded by bypass encoding may be referred to as "bypass encoded bins".
[0086] In addition to performing entropy encoding on the syntax elements in the coefficient block, the video encoder 20 may apply inverse quantization and inverse transformation to the transform block so as to reconstruct the residual sample block from the transform block. The video encoder 20 may add the reconstructed residual sample block to the corresponding samples from one or more prediction sample blocks to generate a reconstructed sample block. By reconstructing the sample blocks of each color component, the video encoder 20 may reconstruct the pixel block associated with the TU. By reconstructing the pixel block of each TU of the CU in this way, the video encoder 20 may reconstruct the pixel block of the CU.
[0087] After the video encoder 20 reconstructs the pixel block of the CU, the video encoder 20 may perform a deblocking operation to reduce the blocking artifacts associated with the CU. After the video encoder 20 performs the deblocking operation, the video encoder 20 may modify the reconstructed pixel block of the CTB of the picture by using sample adaptive offset (SAO). Generally, adding an offset value to the pixels of a picture can improve the coding efficiency. After performing these operations, the video encoder 20 may store the reconstructed pixel block of the CU in the decoded picture buffer for use in generating the predicted pixel block of another CU.
[0088] The video decoder 30 may receive a bitstream. The bitstream may include an encoded representation of video data encoded by the video encoder 20. The video decoder 30 may parse the bitstream to extract syntax elements from the bitstream. As part of extracting at least some of the syntax elements from the bitstream, the video decoder 30 may entropy decode the data in the bitstream.
[0089] When video decoder 30 performs CABAC decoding, video decoder 30 may perform regular CABAC decoding for some bins and may perform bypass decoding for other bins. When video decoder 30 performs regular CABAC decoding for a syntax element, video decoder 30 may identify an encoding context. Video decoder 30 may then divide the interval into a lower sub-interval and an upper sub-interval. One sub-interval may be associated with the value 0 and the other sub-interval may be associated with the value 1. The width of the sub-interval may be proportional to the probability for the associated value indicated by the identified encoding context. If the encoded value is within the lower sub-interval, video decoder 30 may decode the bin having the value associated with the lower sub-interval. If the encoded value is within the upper sub-interval, video decoder 30 may decode the bin having the value associated with the upper sub-interval. To decode the next bin of the syntax element, video decoder 30 may repeat these steps for the interval containing the sub-interval that contains the encoded value. When video decoder 30 repeats these steps for the next bin, video decoder 30 may use the probabilities changed based on the identified encoding context and the probabilities indicated by the decoded bins. Video decoder 30 may then de-binarize the bins to recover the syntax element. De-binarize may mean that the syntax element value can be selected according to the mapping between the binary digit sequence and the syntax element value.
[0090] When the video decoder 30 performs bypass decoding, the video decoder 30 may be able to decode several bins in a single cycle. However, when the video decoder 30 performs regular CABAC decoding, the video decoder 30 may generally only be able to decode a single bin per cycle, or may require more than one cycle for a single bin. Since the video decoder 30 does not need to select a context and it can be assumed that the probabilities for both symbols (0 and 1) are 1 / 2, bypass decoding can be simpler than regular CABAC decoding. Thus, performing bypass encoding and / or decoding for a bin may require less computation compared to performing regular encoding for the bin and can enable higher parallelism and higher throughput.
[0091] The video decoder 30 may reconstruct a picture of video data based on syntax elements extracted from the bitstream. The process of reconstructing the video data based on the syntax elements may generally be contrary to the process performed by the video encoder 20 that generates the syntax elements. For example, the video decoder 30 may generate a predicted pixel block of a PU of a CU based on the syntax elements associated with the CU. In addition, the video decoder 30 may inverse quantize a coefficient block associated with a TU of the CU. The video decoder 30 may perform an inverse transform on the coefficient block to reconstruct a residual pixel block associated with the TU of the CU. The video decoder 30 may reconstruct a pixel block of the CU based on the predicted pixel block and the residual pixel block.
[0092] After the video decoder 30 reconstructs the pixel block of the CU, the video decoder 30 may perform a deblocking operation to reduce the blocking artifacts associated with the CU. Additionally, based on one or more SAO syntax elements, the video decoder 30 may apply SAO applied by the video encoder 20. After the video decoder 30 performs these operations, the video decoder 30 may store the pixel block of the CU in the decoded picture buffer. The decoded picture buffer may provide a reference picture for subsequent motion compensation, intra prediction, and presentation on a display device.
[0093] FIG. 2 is a block diagram showing an example of a video encoder 20 configured to implement the technology of the present invention. FIG. 2 is provided for illustrative purposes and should not be construed as limiting the technology widely exemplified and described in the present invention. For illustrative purposes, in the present invention, the video encoder 20 is described in the context of HEVC encoded image prediction. However, the technology of the present invention is applicable to other encoding standards or methods.
[0094] In the example of FIG. 2, the video encoder 20 includes a prediction processing unit 100, a residual generation unit 102, a transform processing unit 104, a quantization unit 106, an inverse quantization unit 108, an inverse transform processing unit 110, a reconstruction unit 112, a filter unit 113, a decoded picture buffer 114, and an entropy encoding unit 116. The entropy encoding unit 116 includes a regular CABAC encoding engine 118 and a bypass encoding engine 120. The prediction processing unit 100 includes an inter prediction processing unit 121 and an intra prediction processing unit 126. The inter prediction processing unit 121 includes a motion estimation unit 122 and a motion compensation unit 124. In another example, the video encoder 20 may include more, fewer, or different functional components.
[0095] Video encoder 20 receives video data. To encode the video data, video encoder 20 may encode each slice of each picture of the video data. As part of the encoding of the slice, video encoder 20 may encode each CTB of the slice. As part of the encoding of the CTB, prediction processing unit 100 may perform a quadtree partitioning on the pixel block associated with the CTB to partition the pixel block into progressively smaller pixel blocks. The smaller pixel blocks may be associated with CUs. For example, prediction processing unit 100 may partition the pixel block of the CTB into four sub-blocks of the same size, and one or more of the sub-blocks may be partitioned into four sub-sub-blocks of the same size, and so on.
[0096] Video encoder 20 may encode the CUs of the CTBs of the picture to generate an encoded representation of the CUs (i.e., the encoded CUs). Video encoder 20 may encode the CUs of the CTB in a zigzag scan order. In other words, video encoder 20 may sequentially encode the top-left CU, the top-right CU, the bottom-left CU, and the bottom-right CU. When video encoder 20 encodes a partitioned CU, video encoder 20 may encode the CUs associated with the sub-blocks of the pixel block of the partitioned CU in a zigzag scan order.
[0097] In addition, as part of the encoding of the CU, prediction processing unit 100 may partition the pixel block of the CU among one or more PUs of the CU. Video encoder 20 and video decoder 30 may support various PU sizes. Assuming that the size of a particular CU is 2N×2N, video encoder 20 and video decoder 30 may support a PU size of 2N×2N or N×N for intra prediction, and may support a symmetric PU size of 2N×2N, 2N×N, N×2N, N×N, or the like for inter prediction. Video encoder 20 and video decoder 30 may also support asymmetric partitioning for PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.
[0098] The inter prediction processing unit 121 may generate prediction data for each PU of the CU by performing inter prediction on each PU of the CU. The prediction data of the PU may include a predicted pixel block corresponding to the PU and motion information of the PU. The slice may be an I slice, a P slice, or a B slice. The inter prediction unit 121 may perform different operations on the PUs of the CU depending on whether the PU is in an I slice, a P slice, or a B slice. In an I slice, all PUs are intra predicted. Therefore, when the PU is in an I slice, the inter prediction unit 121 does not perform inter prediction on the PU.
[0099] When the PU is in a P slice, the motion estimation unit 122 may search for a reference picture in a list of reference pictures (such as "list 0") for the reference block of the PU. The reference block of the PU may be the pixel block that most closely corresponds to the pixel block of the PU. The motion estimation unit 122 may generate a reference picture index indicating the reference picture in list 0 that includes the reference block of the PU, and a motion vector indicating the spatial displacement between the pixel block of the PU and the reference block. The motion estimation unit 122 may output the reference picture index and the motion vector as the motion information of the PU. The motion compensation unit 124 may generate a predicted pixel block of the PU based on the reference block indicated by the motion information of the PU.
[0100] When the PU is in a B slice, the motion estimation unit 122 may perform unidirectional inter prediction or bidirectional inter prediction on the PU. To perform unidirectional inter prediction on the PU, for the reference block of the PU, the motion estimation unit 122 may search for a reference picture in the first reference picture list ("list 0") or the second reference picture list ("list 1"). The motion estimation unit 122 may output, as the motion information of the PU, a reference picture index indicating the position of the reference picture including the reference block in list 0 or list 1, a motion vector indicating the spatial displacement between the pixel block of the PU and the reference block, and a prediction direction indicator indicating whether the reference picture is in list 0 or list 1.
[0101] To perform bidirectional inter prediction on the PU, the motion estimation unit 122 may search for a reference picture in list 0 for the reference block of the PU, and may also search for a reference picture in list 1 for another reference block of the PU. The motion estimation unit 122 may generate a reference picture index indicating the positions of the reference picture including the reference block in lists 0 and 1. In addition, the motion estimation unit 122 may generate a motion vector indicating the spatial displacement between the reference block and the pixel block of the PU. The motion information of the PU may include the reference picture index and the motion vector of the PU. The motion compensation unit 124 may generate a predicted pixel block of the PU based on the reference block indicated by the motion information of the PU.
[0102] The intra prediction processing unit 126 may generate prediction data of the PU by performing intra prediction on the PU. The prediction data of the PU may include the predicted pixel block of the PU and various syntax elements. The intra prediction processing unit 126 may perform intra prediction on PUs in I slices, P slices, and B slices.
[0103] To perform intra prediction on a PU, the intra prediction processing unit 126 may generate a plurality of sets of prediction data for the PU by using a plurality of intra prediction modes. To generate a set of prediction data for the PU by using an intra prediction mode, the intra prediction processing unit 126 may extend samples from the sample blocks of adjacent PUs over the sample block of the PU in a direction associated with the intra prediction mode. Assuming that the encoding order from left to right and from top to bottom is used for PUs, CUs, and CTBs, the adjacent PU may be above, top - right, top - left, or to the left of the PU. The intra prediction processing unit 126 may use various quantities of intra prediction modes, for example, 33 directional intra prediction modes. In some examples, the quantity of intra prediction modes may depend on the size of the pixel block of the PU.
[0104] The prediction processing unit 100 may select the prediction data of the PU of the CU from the prediction data of the PU generated by the inter prediction processing unit 121 or from the prediction data of the PU generated by the intra prediction processing unit 126. In some examples, the prediction processing unit 100 selects the prediction data of the PU of the CU based on the rate / distortion metrics of the set of prediction data. The predicted pixel block of the selected prediction data may be referred to herein as the selected predicted image block.
[0105] The residual generation unit 102 may generate a residual pixel block for the CU based on the pixel block of the CU and the selected predicted image block of the PU of the CU. For example, the residual generation unit 102 may generate a residual pixel block for the CU, whereby each sample in the residual pixel block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the selected predicted image block of the PU of the CU.
[0106] The prediction processing unit 100 may partition the residual pixel block of the CU into sub-blocks by performing a quadtree partition. Each unpartitioned residual pixel block may be associated with different TUs of the CU. The size and position of the residual pixel block associated with the TU of the CU may or may not be based on the size and position of the pixel block of the PU of the CU.
[0107] Since each pixel of the residual pixel block of the TU may include one luminance sample and two chrominance samples, each TU may be associated with one block of luminance samples and two blocks of chrominance samples. The transformation processing unit 104 may generate a coefficient block for each TU of the CU by applying one or more transformations to the residual sample block associated with the TU. The transformation processing unit 104 may apply various transformations to the residual sample block associated with the TU. For example, the transformation processing unit 104 may apply a discrete cosine transform (DCT), a directional transform, or a conceptually similar transform to the residual sample block.
[0108] The quantization unit 106 may quantize the coefficients in the coefficient block. The quantization process may reduce the bit depth associated with some or all of the coefficients. For example, an n-bit coefficient may be truncated to an m-bit coefficient during quantization, where n is greater than m. The quantization unit 106 may quantize the coefficient block associated with the TU of the CU based on the quantization parameter (QP) value associated with the CU. The video encoder 20 may adjust the degree of quantization applied to the coefficient block associated with the CU by adjusting the QP value associated with the CU.
[0109] To reconstruct the residual sample block from the coefficient block, the inverse quantization unit 108 and the inverse transform processing unit 110 may separately apply inverse quantization and inverse transform to the coefficient block. The reconstruction unit 112 may add samples of the reconstructed residual sample block to corresponding samples from one or more prediction sample blocks generated by the prediction processing unit 100 to generate a reconstruction sample block associated with the TU. By reconstructing the sample block for each TU of the CU in this way, the video encoder 20 may reconstruct the pixel block of the CU.
[0110] The filter unit 113 may perform a deblocking operation to reduce blocking artifacts in the pixel block associated with the CU. In addition, the filter unit 113 may apply the SAO offset determined by the prediction processing unit 100 to the reconstructed sample block to restore the pixel block. The filter unit 113 may generate a sequence of SAO syntax elements of the CTB. The SAO syntax elements may include regular CABAC coded bins and bypass coded bins. According to the technology of the present invention, within the sequence, the bypass coded bins of a color component do not exist between two of the regular CABAC coded bins of the same color component.
[0111] The decoded picture buffer 114 may store the reconstructed pixel block. The inter prediction unit 121 may perform inter prediction on a PU of another picture by using a reference picture including the reconstructed pixel block. In addition, the intra prediction processing unit 126 may perform intra prediction on another PU in the same picture as the CU by using the reconstructed pixel block in the decoded picture buffer 114.
[0112] Entropy encoding unit 116 may receive data from another functional component of video encoder 20. For example, entropy encoding unit 116 may receive coefficient blocks from quantization unit 106 and may receive syntax elements from prediction processing unit 100. Entropy encoding unit 116 may perform one or more entropy encoding operations on the data to generate entropy encoded data. For example, entropy encoding unit 116 may perform context adaptive variable length coding (CAVLC) operations, CABAC operations, variable length to variable length (V2V) coding operations, syntax based context adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, or another type of entropy encoding operation on the data. In a particular example, entropy encoding unit 116 may encode the SAO syntax elements generated by filter unit 113. As part of the encoding of the SAO syntax elements, entropy encoding unit 116 may encode the regular CABAC coded bins of the SAO syntax elements by using regular CABAC coding engine 118 and may encode the bypass coded bins by using bypass coding engine 120.
[0113] According to the technology of the present invention, inter prediction unit 121 determines a set of candidates for inter-frame prediction modes. Thus, video encoder 20 is an example of a video encoder. According to the technology of the present invention, the video encoder determines whether the set of candidates for the prediction mode of the image unit to be processed includes an affine merge mode indicating that the prediction images of the image unit to be processed and the adjacent image units adjacent to the image unit to be processed use the same affine mode according to the information about the adjacent image units adjacent to the image unit to be processed, determines the prediction mode of the image unit to be processed in the set of candidates for the prediction mode, determines the prediction image of the image unit to be processed according to the prediction mode, and is configured to encode first indication information indicating the prediction mode into the bitstream.
[0114] Figure 3 is a flowchart showing an exemplary operation 200 of a video encoder for encoding video data according to one or more techniques of the present invention. Figure 3 is provided by way of example. In another example, the techniques of the present invention may be implemented by using more, fewer, or different steps than those shown in the example of Figure 3. According to the exemplary method of Figure 3, video encoder 20 performs the following steps.
[0115] S210. Determine whether a set of candidate prediction modes of the image unit to be processed includes the affine merge mode according to information about adjacent image units adjacent to the image unit to be processed.
[0116] Specifically, as shown in Figure 4, blocks A, B, C, D, and E are adjacent to the reconstructed block of the block to be currently encoded and are located above, to the left, top-right, bottom-left, and top-left of the block to be encoded, respectively. According to the encoding information of the adjacent reconstructed blocks, it may be determined whether the set of candidate prediction modes of the currently encoded block includes the affine merge mode.
[0117] In this embodiment of the present invention, it should be understood that Figure 4 shows the quantity and positions of adjacent reconstructed blocks of the block to be encoded for illustrative purposes. The quantity of adjacent reconstructed blocks may be more or less than 5 and is not limited in this regard.
[0118] In the first possible implementation method, it is determined whether there is a block whose prediction type is affine prediction among adjacent reconstruction blocks. If there is no block whose prediction type is affine prediction among adjacent reconstruction blocks, the set of candidate prediction modes of the block to be encoded does not include the affine merge mode, or if there is a block whose prediction type is affine prediction among adjacent reconstruction blocks, the encoding process shown in FIG. 2 is executed separately according to two cases: when the set of candidate prediction modes of the block to be encoded includes the affine merge mode, and when the set of candidate prediction modes of the block to be encoded does not include the affine merge mode. When the encoding performance in the first case is better, the set of candidate prediction modes of the block to be encoded includes the affine merge mode, and the indication information that can be assumed as the second indication information is set to 1 and encoded into the bitstream. Otherwise, the set of candidate prediction modes of the block to be encoded does not include the affine merge mode, and the second indication information is set to 0 and encoded into the bitstream.
[0119] In the second possible implementation method, it is determined whether there is a block whose prediction type is affine prediction among adjacent reconstruction blocks. If there is no block whose prediction type is affine prediction among adjacent reconstruction blocks, the set of candidate prediction modes of the block to be encoded does not include the affine merge mode, or if there is a block whose prediction type is affine prediction among adjacent reconstruction blocks, the set of candidate prediction modes of the block to be encoded includes the affine merge mode.
[0120] In a third possible implementation manner, the adjacent reconstruction block includes a plurality of affine modes, and the affine modes include, for example, a first affine mode or a second affine mode. Correspondingly, the affine merge mode includes a first affine merge mode for merging the first affine mode or a second affine merge mode for merging the second affine mode. Among the adjacent reconstruction blocks, the quantities of the first affine mode, the second affine mode, and the non-affine mode are separately obtained through statistical collection. When the first affine mode ranks first in quantity among the adjacent reconstruction blocks, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode. When the second affine mode ranks first in quantity among the adjacent reconstruction blocks, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode. When the non-affine mode ranks first in quantity among the adjacent reconstruction blocks, the set of candidate prediction modes does not include the affine merge mode.
[0121] Alternatively, in a third possible implementation manner, among adjacent reconstruction blocks, when the first affine mode ranks first in terms of quantity, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode. Among adjacent reconstruction blocks, when the second affine mode ranks first in terms of quantity, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode. Among adjacent reconstruction blocks, when the non-affine mode ranks first in terms of quantity, whether the first affine mode or the second affine mode ranks second in terms of quantity among adjacent reconstruction blocks is obtained through statistical collection. Among adjacent reconstruction blocks, when the first affine mode ranks second in terms of quantity, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode. Among adjacent reconstruction blocks, when the second affine mode ranks second in terms of quantity, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode.
[0122] In a fourth possible implementation method, it is determined whether two conditions are satisfied: (1) whether there is a block with a prediction type of affine mode among adjacent reconstruction blocks, and (2) whether the width and height of adjacent blocks in the affine mode are smaller than the width and height of the block to be coded. If neither condition is satisfied, the set of candidate prediction modes for the block to be coded does not include the affine merge mode. If both conditions are satisfied, the coding process shown in FIG. 2 is executed separately according to two cases: when the set of candidate prediction modes for the block to be coded includes the affine merge mode, and when the set of candidate prediction modes for the block to be coded does not include the affine merge mode. When the coding performance in the first case is better, the set of candidate prediction modes for the block to be coded includes the affine merge mode, and the indication information that can be assumed as the third indication information is set to 1 and coded into the bitstream. Otherwise, the set of candidate prediction modes for the block to be coded does not include the affine merge mode, and the third indication information is coded into the bitstream with 0 set.
[0123] It should be understood that the determination condition (2) in this embodiment of the present invention means that the width of the adjacent block in the affine mode is smaller than the width of the block to be coded, and the height of the adjacent block in the affine mode is smaller than the height of the block to be coded. In another embodiment, alternatively, the determination condition may be that the width of the adjacent block in the affine mode is smaller than the width of the block to be coded, or the height of the adjacent block in the affine mode is smaller than the height of the block to be coded, and it is not limited thereto.
[0124] In a fifth possible implementation manner, it is determined whether two conditions are satisfied, such as (1) whether there is a block with a prediction type of affine mode among adjacent reconstruction blocks, and (2) whether the width and height of adjacent blocks in the affine mode are smaller than the width and height of the block to be coded. If neither condition is satisfied, the set of candidate prediction modes of the block to be coded does not include the affine merge mode. If both conditions are satisfied, the set of candidate prediction modes of the block to be coded includes the affine merge mode.
[0125] In this embodiment of the present invention, it should be understood that the prediction type and size of adjacent reconstruction blocks are used as a basis for determining the set of candidate prediction modes of the currently to-be-coded block, and the attribute information of adjacent reconstruction blocks obtained by analysis is further used for determination. This is not limited in this specification.
[0126] In various possible implementation manners such as the second possible implementation manner in this embodiment of the present invention, for illustrative purposes, it should be further understood that the following decision criteria may be used to determine whether there is a block with a prediction type of affine prediction among adjacent reconstruction blocks. For illustrative purposes, if the prediction types of at least two adjacent blocks are in the affine mode, the set of candidate prediction modes of the block to be coded includes the affine merge mode; otherwise, the set of candidate prediction modes of the block to be coded does not include the affine merge mode. Alternatively, the number of adjacent blocks with a prediction type of affine mode may be at least three, or at least four, and this is not limited.
[0127] In various possible implementation manners in this embodiment of the present invention, for illustrative purposes, for example, in the fifth possible implementation manner, it should be further understood that it is determined whether two conditions are satisfied, such as (1) whether there is a block with a prediction type of affine mode among adjacent reconstruction blocks, and (2) whether the width and height of adjacent blocks in the affine mode are smaller than the width and height of the block to be coded. For illustrative purposes, the second determination condition may be whether the width and height of adjacent blocks in the affine mode are less than 1 / 2, 1 / 3, or 1 / 4 of the width and height of the block to be coded, and this is not limited thereto.
[0128] In this embodiment of the present invention, it should be further understood that the indication information is set to 0 or 1 for illustrative purposes. Alternatively, the reverse setting may be executed. For illustrative purposes, for example, in the first possible implementation manner, it may be determined whether there is a block with a prediction type of affine prediction among adjacent reconstruction blocks. If there is no block with a prediction type of affine prediction among adjacent reconstruction blocks, the set of candidate prediction modes of the block to be coded does not include the affine merge mode, or if there is a block with a prediction type of affine prediction among adjacent reconstruction blocks, the coding process shown in FIG. 2 is executed separately according to two cases: when the set of candidate prediction modes of the block to be coded includes the affine merge mode, and when the set of candidate prediction modes of the block to be coded does not include the affine merge mode. When the coding performance in the first case is better, the set of candidate prediction modes of the block to be coded includes the affine merge mode, and the indication information assumed as the second indication information is set to 0 and coded into the bitstream. Otherwise, the set of candidate prediction modes of the block to be coded does not include the affine merge mode, and the second indication information is set to 1 and coded into the bitstream.
[0129] S220. This is the stage of determining the prediction mode of the image unit to be processed in the set of candidate prediction modes.
[0130] The set of candidate prediction modes is the set of candidate prediction modes determined in S210. In order to select the mode with the optimal coding performance as the prediction mode of the block to be coded, each prediction mode in the set of candidate prediction modes is sequentially used to execute the coding process shown in FIG. 2.
[0131] In this embodiment of the present invention, it should be understood that the purpose of executing the coding process shown in FIG. 2 is to select the prediction mode with the optimal coding performance. In the selection process, the performance / cost ratio of the prediction mode may be compared. The performance is indicated by the quality of the image restoration, and the cost is indicated by the coding bit rate. Alternatively, only the performance or cost of the prediction mode may be compared. Correspondingly, all the coding stages shown in FIG. 2 may be completed, or the coding process may be stopped after the indicator that needs to be compared is obtained. For example, when the prediction mode is compared only with respect to performance, the coding process can be stopped after the prediction unit executes that stage, and this is not limited thereto.
[0132] S230. This is the stage of determining the predicted image of the image unit to be processed according to the prediction mode.
[0133] The above-cited H.265 standard and application documents such as No. CN201010247275.7 have detailed the process in which the predicted image of the block to be coded is generated according to the prediction mode including the translational model prediction mode, the affine prediction mode, the affine merge mode, or the like, and the details will not be described again here.
[0134] S240. This is the stage of coding the first instruction information into the bit stream.
[0135] The prediction mode determined in S220 is encoded into the bitstream. This step may be executed at any point in time after S220, and it should be understood that there are no specific limitations on the order of the steps as long as the step corresponds to the step where the decoder decodes the first instruction information.
[0136] FIG. 5 is a block diagram showing an example of another video encoder 40 for encoding video data according to one or more techniques of the present invention.
[0137] The video encoder 40 includes a first determination module 41, a second determination module 42, a third determination module 43, and an encoding module 44.
[0138] The first determination module 41 is configured to execute S210 that determines whether a set of candidates for the prediction mode of the image unit to be processed includes the affine merge mode according to information about adjacent image units adjacent to the image unit to be processed.
[0139] The second determination module 42 is configured to execute S220 that determines the prediction mode of the image unit to be processed in the set of candidates for the prediction mode.
[0140] The third determination module 43 is configured to execute S230 that determines the predicted image of the image unit to be processed according to the prediction mode.
[0141] The encoding module 44 is configured to execute S240 that encodes the first instruction information into the bitstream.
[0142] Since the motion information of adjacent blocks is correlated, the current block and the adjacent blocks are very likely to have the same or similar prediction modes. In this embodiment of the present invention, the prediction mode information of the current block is derived by determining information about the adjacent blocks, reducing the bit rate for encoding the prediction mode, thereby improving the encoding efficiency.
[0143] FIG. 6 is a flowchart showing an exemplary operation 300 of a video encoder for encoding video data in accordance with one or more techniques of the present invention. FIG. 5 is provided as an example. In another example, the techniques of the present invention may be implemented by using more, fewer, or different steps than those shown in the example of FIG. 5. According to the exemplary method of FIG. 5, video encoder 20 performs the following steps.
[0144] S310. Encoding indication information of a set of candidates for the prediction mode of the first image region to be processed.
[0145] When the set of translational mode candidates is used as the set of mode candidates for the first image region to be processed, the first indication information is set to 0, and the first indication information is encoded in the bitstream, where the translational mode indicates a prediction mode for obtaining a predicted image by using a translational model. When the set of translational mode candidates and the set of affine mode candidates are used as the set of mode candidates for the first image region to be processed, the first indication information is set to 1, and the first indication information is encoded in the bitstream, where the affine mode indicates a prediction mode for obtaining a predicted image by using an affine model. The first image region to be processed may be any one of an image frame group, an image frame, a set of image tiles, a set of image slices, an image tile, an image slice, a set of image coding units, or an image coding unit. Correspondingly, the first indication information is encoded in the header of an image frame group, such as, for example, a video parameter set (VPS), a sequence parameter set (SPS), supplementary enhancement information (SEI), or an image frame header, or, for example, an image parameter set (PPS), a header of a set of image tiles, a header of a set of image slices, or a header of an image tile, or, for example, an image tile header (tile header), an image slice header (slice header), a header of a set of image coding units, or a header of an image coding unit.
[0146] It should be understood that the first image region to be processed at this stage may be pre-configured or may be adaptively determined during the encoding process. The range indication of the first image region to be processed can be obtained from the protocol on the encoding / decoding side. Alternatively, the range of the first image region to be processed may be encoded in the bitstream for transmission, but this is not limited thereto.
[0147] It should be further understood that the set of prediction mode candidates may be pre-configured or may be determined after comparing encoding performance, but this is not limited thereto.
[0148] It should be further understood that in this embodiment of the present invention, the indication information being set to 0 or 1 is for illustrative purposes. Alternatively, the reverse setting may be performed.
[0149] S320. A step of determining a prediction mode of an image unit to be processed in a first image area to be processed from a set of candidate prediction modes of the first image area to be processed for the unit to be processed.
[0150] The specific method is the same as that of S220, and details will not be described again here.
[0151] S330. A step of determining a predicted image of an image unit to be processed according to the prediction mode.
[0152] The specific method is the same as that of S230, and details will not be described again here.
[0153] S340. A step of encoding the prediction mode selected for the unit to be processed into a bitstream.
[0154] The specific method is the same as that of S240, and details will not be described again here.
[0155] FIG. 7 is a block diagram showing an example of another video encoder 50 for encoding video data according to one or more techniques of the present invention.
[0156] The video encoder 50 includes a first encoding module 51, a first determination module 52, a second determination module 53, and a second encoding module 54.
[0157] The first encoding module 51 is configured to execute S310 for encoding indication information of a set of candidate prediction modes of a first image area to be processed.
[0158] The first decision module 52 is configured to execute S320 that determines the prediction mode of the image unit to be processed in the first image area to be processed from the set of candidate prediction modes of the first image area to be processed for the unit to be processed.
[0159] The second decision module 53 is configured to execute S330 that determines the predicted image of the image unit to be processed according to the prediction mode.
[0160] The second encoding module 54 is configured to execute S340 that encodes the selected prediction mode for the unit to be processed into a bitstream.
[0161] Since the motion information of adjacent blocks is correlated, there is only translational motion and there is a very high possibility that there is no affine motion in the same area. In this embodiment of the present invention, a set of candidate prediction modes for marking the region level is set to avoid the bit rate for encoding redundant modes, thereby improving the encoding efficiency.
[0162] FIG. 8 is a block diagram showing an example of a video decoder 30 configured to implement the technology of the present invention. FIG. 8 is provided for illustrative purposes and should not be construed as limiting the technology widely exemplified and described in the present invention. For illustrative purposes, the video decoder 30 is described in the present invention in the image prediction of HEVC encoding. However, the technology of the present invention is applicable to other encoding standards or methods.
[0163] In the example of FIG. 8, video decoder 30 includes an entropy decoding unit 150, a prediction processing unit 152, an inverse quantization unit 154, an inverse transform processing unit 156, a reconstruction unit 158, a filter unit 159, and a decoded picture buffer 160. Prediction processing unit 152 includes a motion compensation unit 162 and an intra prediction processing unit 164. Entropy decoding unit 150 includes a regular CABAC encoding engine 166 and a bypass encoding engine 168. In another example, video decoder 30 may include more, fewer, or different functional components.
[0164] Video decoder 30 may receive a bitstream. Entropy decoding unit 150 may analyze the bitstream to extract syntax elements from the bitstream. As part of analyzing the bitstream, entropy decoding unit 150 may entropy decode the entropy-encoded syntax elements in the bitstream. Prediction processing unit 152, inverse quantization unit 154, inverse transform processing unit 156, reconstruction unit 158, and filter unit 159 may generate decoded video data based on the syntax elements extracted from the bitstream.
[0165] The bitstream may include a sequence of encoded SAO syntax elements for CTBs. The SAO syntax elements may include regular CABAC encoded bins and bypass encoded bins. According to the technique of the present invention, in the sequence of encoded SAO syntax elements, the bypass encoded bins do not exist between two of the regular CABAC encoded bins. Entropy decoding unit 150 may decode the SAO syntax elements. As part of encoding the SAO syntax elements, entropy decoding unit 150 may decode the regular CABAC encoded bins by using regular CABAC encoding engine 166, and may decode the bypass encoded bins by using bypass encoding engine 168.
[0166] In addition, the video decoder 30 may perform a reconstruction operation on a non-partitioned CU. To perform a reconstruction operation on a non-partitioned CU, the video decoder 30 may perform a reconstruction operation on each TU of the CU. By performing a reconstruction operation on each TU of the CU, the video decoder 30 may reconstruct the residual pixel block associated with the CU.
[0167] As part of the execution of the reconstruction operation on the TU of the CU, the inverse quantization unit 154 may inverse-quantize (i.e., dequantize) the coefficient block associated with the TU. The inverse quantization unit 154 may determine the degree of quantization by using the QP value associated with the CU of the TU, and may determine the degree of inverse quantization to be applied by the inverse quantization unit 154.
[0168] After the inverse quantization unit 154 inverse-quantizes the coefficient block, the inverse transform processing unit 156 may apply one or more inverse transforms to the coefficient block to generate the residual sample block associated with the TU. For example, the inverse transform processing unit 156 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse directional transform, or another inverse transform to the coefficient block.
[0169] When the PU is encoded by intra prediction, the intra prediction processing unit 164 may perform an intra prediction to generate a prediction sample block of the PU. The intra prediction processing unit 164 may use an intra prediction mode to generate a prediction pixel block of the PU based on the pixel blocks of spatially adjacent PUs. The intra prediction processing unit 164 may determine the intra prediction mode of the PU based on one or more syntax elements obtained from the bitstream by parsing.
[0170] The motion compensation unit 162 may construct a first reference picture list (list 0) and a second reference picture list (list 1) based on syntax elements extracted from the bitstream. Additionally, when a PU is encoded by inter prediction, the entropy decoding unit 150 may extract the motion information of the PU. The motion compensation unit 162 may determine one or more reference blocks of the PU based on the motion information of the PU. The motion compensation unit 162 may generate a predicted pixel block of the PU based on one or more reference blocks of the PU.
[0171] The reconstruction unit 158 may use, when applicable, the residual pixel block associated with the TU of the CU and the predicted pixel block of the PU of the CU (i.e., intra prediction data or inter prediction data) to reconstruct the pixel block of the CU. In particular, the reconstruction unit 158 may add samples of the residual pixel block to the corresponding samples of the predicted pixel block to reconstruct the pixel block of the CU.
[0172] The filter unit 159 may perform a deblocking operation to reduce blocking artifacts associated with the pixel block of the CU of the CTB. Additionally, the filter unit 159 may modify the pixel block of the CTB based on the SAO syntax elements parsed from the bitstream. For example, the filter unit 159 may determine a value based on the SAO syntax elements of the CTB and add the determined value to the samples in the reconstructed pixel block of the CTB. By modifying at least a part of the pixel blocks of the CTB of the picture, the filter unit 159 may modify the reconstructed picture of the video data based on the SAO syntax elements.
[0173] The video decoder 30 may store the pixel block of the CU in the decoded picture buffer 160. The decoded picture buffer 160 may provide a reference picture for subsequent motion compensation, intra prediction, and presentation on a display device (such as the display device 32 in FIG. 1). For example, the video decoder 30 may perform an intra prediction or an inter prediction operation on the PU of another CU based on the pixel block in the decoded picture buffer 160.
[0174] According to the technology of the present invention, the prediction processing unit 152 determines a set of candidates for the inter prediction mode. In this way, the video decoder 30 is an example of a video decoder. According to the technology of the present invention, the video decoder determines whether the set of candidates for the prediction mode of the image unit to be processed includes an affine merge mode indicating that the prediction images of the image unit to be processed and the adjacent image units adjacent to the image unit to be processed use the same affine mode according to the information about the adjacent image units adjacent to the image unit to be processed, analyzes the bitstream to obtain the first instruction information, determines the prediction mode of the image unit to be processed according to the first instruction information in the set of candidates for the prediction mode, and is configured to determine the prediction image of the image unit to be processed according to the prediction mode.
[0175] FIG. 9 is a flowchart showing an exemplary operation 400 of a video decoder for decoding video data according to one or more technologies of the present invention. FIG. 9 is provided as an example. In another example, the technology of the present invention may be implemented by using more, fewer, or different steps than those shown in the example of FIG. 9. According to the exemplary method of FIG. 9, the video decoder 30 performs the following steps.
[0176] S410. A step of determining whether the set of candidates for the prediction mode of the image unit to be processed includes an affine merge mode according to the information about the adjacent image units adjacent to the image unit to be processed.
[0177] Specifically, as shown in FIG. 4, blocks A, B, C, D, and E are adjacent reconstruction blocks of the currently block to be encoded, and are located above, to the left, upper right, lower left, and upper left of the block to be encoded, respectively. According to the encoding information of the adjacent reconstruction blocks, it may be determined whether the set of candidate prediction modes of the currently block to be encoded includes the affine merge mode.
[0178] In this embodiment of the present invention, it should be understood that FIG. 4 shows the amount and position of the adjacent reconstruction blocks of the block to be encoded for illustrative purposes. The amount of adjacent reconstruction blocks may be more or less than 5, and is not limited thereto.
[0179] In the first possible implementation manner, it is determined whether there is a block among the adjacent reconstruction blocks whose prediction type is affine prediction. When at least one prediction mode of the adjacent image units obtains a predicted image by using an affine model, the bitstream is analyzed to obtain second indication information. When the second indication information is 1, the set of candidate prediction modes includes the affine merge mode, or when the second indication information is 0, the set of candidate prediction modes does not include the affine merge mode, or otherwise, the set of candidate prediction modes does not include the affine merge mode.
[0180] In the second possible implementation manner, it is determined whether there is a block among the adjacent reconstruction blocks whose prediction type is affine prediction. If there is no block among the adjacent reconstruction blocks whose prediction type is affine prediction, the set of candidate prediction modes of the block to be encoded does not include the affine merge mode, or if there is a block among the adjacent reconstruction blocks whose prediction type is affine prediction, the set of candidate prediction modes of the block to be encoded includes the affine merge mode.
[0181] In a third possible implementation manner, the adjacent reconstruction block includes a plurality of affine modes, the affine modes include, for example, a first affine mode or a second affine mode, and correspondingly, the affine merge mode includes a first affine merge mode for merging the first affine mode or a second affine merge mode for merging the second affine mode. Among the adjacent reconstruction blocks, the quantities of the first affine mode, the second affine mode, and the non-affine mode are separately obtained through statistical collection. When the first affine mode ranks first in terms of quantity among the adjacent reconstruction blocks, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode. When the second affine mode ranks first in terms of quantity among the adjacent reconstruction blocks, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode. When the non-affine mode ranks first in terms of quantity among the adjacent reconstruction blocks, the set of candidate prediction modes does not include the affine merge mode.
[0182] Alternatively, in a third possible implementation manner, among adjacent reconstruction blocks, when the first affine mode ranks first in terms of quantity, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode. Among adjacent reconstruction blocks, when the second affine mode ranks first in terms of quantity, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode. Among adjacent reconstruction blocks, when the non-affine mode ranks first in terms of quantity, whether the first affine mode or the second affine mode ranks second in terms of quantity among adjacent reconstruction blocks is obtained through statistical collection. Among adjacent reconstruction blocks, when the first affine mode ranks second in terms of quantity, the set of candidate prediction modes includes the first affine merge mode and does not include the second affine merge mode. Among adjacent reconstruction blocks, when the second affine mode ranks second in terms of quantity, the set of candidate prediction modes includes the second affine merge mode and does not include the first affine merge mode.
[0183] In a fourth possible implementation manner, it is determined whether two conditions are satisfied, namely, (1) whether there is a block whose prediction type is the affine mode among adjacent reconstruction blocks, and (2) whether the width and height of adjacent blocks in the affine mode are smaller than the width and height of the block to be coded. If neither condition is satisfied, the set of candidate prediction modes for the block to be coded does not include the affine merge mode. If both conditions are satisfied, the bitstream is analyzed to obtain the third indication information. When the third indication information is 1, the set of candidate prediction modes includes the affine merge mode, or when the third indication information is 0, the set of candidate prediction modes does not include the affine merge mode, or otherwise, the set of candidate prediction modes does not include the affine merge mode.
[0184] In this embodiment of the present invention, it should be understood that the determination condition (2) means that the width of the adjacent block in the affine mode is smaller than the width of the block to be encoded, and the height of the adjacent block in the affine mode is smaller than the height of the block to be encoded. In another embodiment, alternatively, the determination condition may be that the width of the adjacent block in the affine mode is smaller than the width of the block to be encoded, or the height of the adjacent block in the affine mode is smaller than the height of the block to be encoded, and it is not limited thereto.
[0185] In the fifth possible implementation manner, it is determined whether two conditions are satisfied, namely, (1) whether there is a block whose prediction type is the affine mode among the adjacent reconstruction blocks, and (2) whether the width and height of the adjacent blocks in the affine mode are smaller than the width and height of the block to be encoded. If neither condition is satisfied, the set of candidate prediction modes for the block to be encoded does not include the affine merge mode. If both conditions are satisfied, the set of candidate prediction modes for the block to be encoded includes the affine merge mode.
[0186] In this embodiment of the present invention, the prediction type and size of the adjacent reconstruction blocks are used as the basis for determining the set of candidate prediction modes of the currently encoded block. It should be understood that as long as the method corresponds to the encoding side, the attribute information of the adjacent reconstruction blocks obtained by analysis may be further used by determination. It is not limited in this specification.
[0187] In various possible implementation manners in this embodiment of the present invention, for illustrative purposes, for example, in the second possible implementation manner, the following determination criteria may be further understood to be used to determine whether there is a block whose prediction type is affine prediction among adjacent reconstruction blocks. For illustrative purposes, when the prediction types of at least two adjacent blocks are in the affine mode, the set of candidate prediction modes of the block to be coded includes the affine merge mode. Otherwise, the set of candidate prediction modes of the block to be coded does not include the affine merge mode. Alternatively, the number of adjacent blocks whose prediction type is in the affine mode may be at least three, or at least four as long as this corresponds to the coding side, and this is not limited in this regard.
[0188] In various possible implementation manners in this embodiment of the present invention, for illustrative purposes, for example, in the fifth possible implementation manner, it should be further understood that it is determined whether two conditions are satisfied, namely, (1) whether there is a block whose prediction type is in the affine mode among adjacent reconstruction blocks, and (2) whether the width and height of the adjacent blocks in the affine mode are smaller than the width and height of the block to be coded. For illustrative purposes, the second determination condition may be whether the width and height of the adjacent blocks in the affine mode are less than 1 / 2, 1 / 3, or 1 / 4 of the width and height of the block to be coded as long as this corresponds to the coding side, and this is not limited in this regard.
[0189] In this embodiment of the present invention, it should be further understood that the indication information set to 0 or 1 corresponds to the coding side.
[0190] S420. This is a step of analyzing the bitstream to obtain the first indication information.
[0191] The first instruction information indicates the index information of the prediction mode of the block to be decoded. This stage corresponds to stage S240 on the encoding side.
[0192] S430. In the set of candidate prediction modes, according to the first instruction information, it is a stage of determining the prediction mode of the image unit to be processed.
[0193] Different sets of candidate prediction modes correspond to different lists of prediction modes. The list of prediction modes corresponding to the set of candidate prediction modes determined in S410 is searched according to the index information obtained in S420, whereby the prediction mode of the block to be decoded can be found.
[0194] S440. According to the prediction mode, it is a stage of determining the predicted image of the image unit to be processed.
[0195] The specific method is the same as S230, and the details will not be described again here.
[0196] FIG. 10 is a block diagram showing an example of another video decoder 60 for decoding video data according to one or more techniques of the present invention.
[0197] The video decoder 60 includes a first determination module 61, an analysis module 62, a second determination module 63, and a third determination module 64.
[0198] The first determination module 61 is configured to execute S410 that determines whether the set of candidate prediction modes of the image unit to be processed includes the affine merge mode according to the information about the adjacent image units adjacent to the image unit to be processed.
[0199] The analysis module 62 is configured to execute S420 that analyzes the bitstream to obtain the first instruction information.
[0200] The second determination module 63 is configured to execute S430 that determines a prediction mode for an image unit to be processed according to the first instruction information in the set of candidate prediction modes.
[0201] The third determination module 64 is configured to execute S440 that determines a predicted image of an image unit to be processed according to the prediction mode.
[0202] Since the motion information of adjacent blocks is correlated, there is a very high possibility that the current block and the adjacent blocks have the same or similar prediction modes. In this embodiment of the present invention, the prediction mode information of the current block is derived by determining information about adjacent blocks, reducing the bit rate for encoding the prediction mode, and thereby improving the encoding efficiency.
[0203] FIG. 11 is a flowchart showing an exemplary operation 500 of a video decoder for decoding video data according to one or more techniques of the present invention. FIG. 11 is provided by way of example. In another example, the techniques of the present invention may be implemented by using more, fewer, or different steps than those shown in the example of FIG. 11. According to the exemplary method of FIG. 11, the video decoder 20 performs the following steps.
[0204] S510. It is a step of analyzing a bitstream to obtain first instruction information.
[0205] The first instruction information indicates whether the set of candidate modes of the first image area to be processed includes an affine motion model. This step corresponds to step S310 on the encoding side.
[0206] S520. It is a step of determining a set of candidate modes of the first image area to be processed according to the first instruction information.
[0207] When the first instruction information is 0, the set of candidates for the translation mode is used as the set of candidates for the mode of the first image area to be processed, where the translation mode indicates a prediction mode for obtaining a predicted image by using a translation model. When the first instruction information is 1, the set of candidates for the translation mode and the set of candidates for the affine mode are used as the set of candidates for the mode of the first image area to be processed, where the affine mode indicates a prediction mode for obtaining a predicted image by using an affine model. The first image area to be processed may be any one of an image frame group, an image frame, a set of image tiles, a set of image slices, an image tile, an image slice, a set of image coding units, or an image coding unit. Correspondingly, the first instruction information is encoded in the header of an image frame group, such as, for example, a video parameter set (VPS), a sequence parameter set (SPS), supplementary enhancement information (SEI), or an image frame header, or, for example, an image parameter set (PPS), a header of a set of image tiles, a header of a set of image slices, or a header of an image tile, or, for example, an image tile header (tile header), an image slice header (slice header), a header of a set of image coding units, or a header of an image coding unit.
[0208] It should be understood that the first image area to be processed at this stage may be configured in advance or may be determined adaptively in the encoding process. The display of the range of the first image area to be processed can be obtained from the protocol on the encoding / decoding side. Alternatively, as long as this corresponds to the encoding side, the range of the first image area to be processed can be received in the bitstream from the encoding side, and this is not limited in this regard.
[0209] In this embodiment of the present invention, it should be further understood that setting the instruction information to 0 or 1 is for illustrative purposes as long as this corresponds to the encoding side.
[0210] S530. This is the stage of analyzing the bitstream to obtain the second instruction information.
[0211] The second instruction information indicates the prediction mode of the block to be processed in the first image area to be processed. This stage corresponds to the stage S340 on the encoding side.
[0212] S540. This is the stage of determining the prediction mode of the image unit to be processed according to the second instruction information in the set of candidates for the prediction mode of the first image area to be processed.
[0213] The specific method is the same as that of S320, and the details will not be described again here.
[0214] S550. This is the stage of determining the predicted image of the image unit to be processed according to the prediction mode.
[0215] The specific method is the same as that of S330, and the details will not be described again here.
[0216] FIG. 12 is a block diagram showing an example of another video decoder 70 for decoding video data according to one or more techniques of the present invention.
[0217] The video decoder 70 includes a first analysis module 71, a first determination module 72, a second analysis module 73, a second determination module 74, and a third determination module 75.
[0218] The first analysis module 71 is configured to execute S510 that analyzes the bitstream to obtain the first instruction information.
[0219] The first determination module 72 is configured to execute S520 that determines a set of candidates for the mode of the first image area to be processed according to the first instruction information.
[0220] The second analysis module 73 is configured to execute S530 that analyzes the bit stream to obtain the second instruction information.
[0221] The second determination module 74 is configured to execute S540 that determines the prediction mode of the image unit to be processed according to the second instruction information in the set of candidates for the prediction mode of the first image area to be processed.
[0222] The third determination module 75 is configured to execute S550 that determines the predicted image of the image unit to be processed according to the prediction mode.
[0223] Since the motion information of adjacent blocks is correlated, there is only translational motion and there is a very high possibility that there is no affine motion in the same area. In this embodiment of the present invention, a set of candidates for the prediction mode that marks the region level is set to avoid the bit rate for encoding the redundant mode, thereby improving the encoding efficiency.
[0224] In one or more embodiments, the described functionality may be implemented by hardware, software, firmware, or any combination thereof. If the functionality is implemented by software, the functionality may be stored as one or more instructions on a computer-readable medium, or encoded or transmitted by a computer-readable medium and executed by a processing unit based on hardware. The computer-readable medium may include a computer-readable storage medium (corresponding to a tangible medium such as a data recording medium) or a communication medium. The communication medium may include, for example, any medium that facilitates the transmission of data from one location to another according to a communication protocol by using a computer program. In this manner, the computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or a carrier. The data recording medium may be any available medium that can be accessed by one or more computers or one or more processors to read instructions, code, and / or data structures for implementing the techniques described in the present invention. A computer program product may include a computer-readable medium.
[0225] By way of example and not limitation, some computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, other optical disk storage, or magnetic disk storage, other magnetic storage devices, flash memory, or any other medium that can store program code in the form of instructions or data structures and can be accessed by a computer. Additionally, any connection may be properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or another remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, or microwave), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, or microwave) is included in the definition of the medium. However, computer-readable storage media and data recording media may not include connections, carriers, signals, or other transient media and should be understood to be non-transitory tangible storage media. As used herein, disk and optical disk include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy (registered trademark) disk, and Blu-ray disk (registered trademark), where disk generally copies data magnetically and optical disk copies data optically by using a laser. Combinations of the above-described items are also included within the scope of computer-readable media.
[0226] The commands may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein may refer to the structures described above, or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined coder-decoder. Additionally, the techniques may be implemented entirely in one or more circuits or logic elements.
[0227] The technology in the present invention may be widely implemented in multiple apparatuses or devices. The apparatuses or devices include wireless handsets, integrated circuits (ICs), or IC sets (e.g., chip sets). In the present invention, various components, modules, and units are described to emphasize the functions of apparatuses configured to implement the disclosed technology, but the functions do not necessarily have to be implemented by different hardware units. Specifically, as described above, the various units may be combined into the hardware units of a coder-decoder, or provided by a set of mutually usable hardware units (including one or more of the processors described above) in combination with appropriate software and / or firmware.
[0228] It should be understood that the "one embodiment" or "embodiment" referred to throughout the specification means that the specific functions, structures, or features related to the embodiment are included in at least one embodiment of the present invention. Therefore, the "one embodiment" or "embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific functions, structures, or features may be combined in one or more embodiments by any suitable method.
[0229] In various embodiments of the present invention, it should be understood that the sequence numbers of the above processes do not mean the execution sequence and should not be construed as any limitation to the process of the implementation manner of the embodiments of the present invention. The execution sequence of the process should be determined according to the functions and internal logics of the process.
[0230] In addition, the terms "system" and "network" may be used interchangeably in this specification. The term "and / or" in this specification only describes the relevant relationship for explaining the relevant objects and represents that three relationships may exist. For example, A and / or B may represent three cases: when only A exists, when both A and B exist, and when only B exists. In addition, the symbol " / " in this specification generally indicates the "or" relationship between the relevant objects.
[0231] In the embodiments of this application, it should be understood that "B corresponding to A" indicates that B is associated with A and B can be determined according to A. However, it should be further understood that determining A according to B does not mean that B is determined only according to A, and B can also be determined according to A and / or other information.
[0232] Those skilled in the art may recognize that, in combination with the examples described in the embodiments disclosed herein, the units and algorithm steps may be implemented by electronic hardware, computer software, or a combination thereof. To clearly illustrate the compatibility between hardware and software, the above generally describes the composition and steps of each example according to the functions. Whether the function is executed by hardware or software depends on the specific application of the technical solution and the conditions of design constraints. Those skilled in the art may use different methods to implement the functions described for each specific application, but the implementation method should not be regarded as exceeding the scope of the present invention.
[0233] As can be clearly understood by those skilled in the art, for the purpose of simplicity and conciseness, for the detailed operation processes of the above-described systems, apparatuses, and units, the corresponding processes in the embodiments of the above-described methods may be referred to, and the details will not be described again here.
[0234] It should be understood that in some embodiments provided in this application, the disclosed systems, apparatuses, and methods may be implemented in other ways. For example, the described embodiments of the apparatus are merely examples. For example, the division of units is only a division of logical functions, and in the actual implementation method, it may be other divisions. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the shown or described interconnection, direct connection, or communication connection may be implemented by using some interfaces. The indirect connection or communication connection between apparatuses or units may be implemented in electronic, mechanical, or other forms.
[0235] The units described as separate parts may or may not be physically separate. The parts shown as units may or may not be physical units, may be located in one place, or may be distributed across multiple network units. Some or all of the units may be selected according to actual requirements to achieve the objectives of the solution means of the embodiment.
[0236] In addition, the functional units in the embodiments of the present invention may be integrated into one processing unit, or each of the units may exist physically alone, or two or more units may be integrated into one unit.
[0237] The above description is only a specific implementation manner of the present invention and is not intended to limit the protection scope of the present invention. Any deformation or substitution that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention shall be included within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
Claim 1 A method for transmitting a bitstream, comprising: generating a bitstream including first indication information and second indication information, wherein the first indication information is included in a sequence parameter set (SPS) of the bitstream, the first indication information indicates whether there is an affine mode in a set of candidate prediction modes of a first image region to be processed, the second indication information indicates a prediction mode of an image unit to be processed among the set of candidate prediction modes, and the image unit to be processed belongs to the first image region to be processed; transmitting the bitstream; and comprising: wherein the first indication information being 0 indicates that a set of candidate translational modes is used as the set of candidate prediction modes of the first image region to be processed, the translational mode indicates a prediction mode in which a predicted image is obtained using a translational model, or the first indication information being 1 indicates that a set of candidate translational modes and a set of candidate affine modes are used as the set of candidate prediction modes of the first image region to be processed, and the affine mode indicates a prediction mode in which a predicted image is obtained using an affine model. Claim 2 When the first indication information is 0, the first indication information indicates that there is no affine mode in the set of candidate prediction modes, or when the first indication information is 1, the first indication information indicates that there is an affine mode in the set of candidate prediction modes, the method according to claim 1. Claim 3 The method according to claim 1 or 2, wherein the first image region to be processed has one of a group of picture frames, a picture frame, a set of picture tiles, a set of picture slices, a picture tile, a picture slice, a set of picture coding units, or a picture coding unit. Claim 4 An apparatus for transmitting a bitstream, comprising: one or more processors, wherein the one or more processors are Generating a bitstream including first instruction information and second instruction information, wherein the first instruction information is included in a sequence parameter set (SPS) of the bitstream, the first instruction information indicates whether there is an affine mode in a set of candidates for a prediction mode of a first image region to be processed, the second instruction information indicates a prediction mode of an image unit to be processed among the set of candidates for the prediction mode, and the image unit to be processed belongs to the first image region to be processed, Transmitting the bitstream, configured to perform, The first instruction information being 0 indicates that a set of candidates for a translational mode is used as the set of candidates for the prediction mode of the first image region to be processed, the translational mode indicating a prediction mode in which a predicted image is obtained using a translational model, or, The first instruction information being 1 indicates that a set of candidates for a translational mode and a set of candidates for an affine mode are used as the set of candidates for the prediction mode of the first image region to be processed, the affine mode indicating a prediction mode in which a predicted image is obtained using an affine model. **Claim 5** When the first instruction information is 0, the first instruction information indicates that there is no affine mode in the set of candidates for the prediction mode, or, When the first instruction information is 1, the first instruction information indicates that there is an affine mode in the set of candidates for the prediction mode, the apparatus according to claim 4. **Claim 6** The first image region to be processed has one of an image frame group, an image frame, an image tile set, an image slice set, an image tile, an image slice, an image coding unit set, or an image coding unit, the apparatus according to claim 4 or 5.
Citation Information
Patent Citations
JPP7368414B
JPP6882560B