Hash-based Motion Search
By adopting a hash-based motion search method in video encoding, the problem of low motion estimation efficiency of non-square and non-rectangular video blocks in the prior art is solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202080006849.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-22
- Filing Date
- 2020-01-02
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-01-02
AI Technical Summary
When existing video encoding technology deals with non-square and non-rectangular video blocks, the motion estimation efficiency is low, resulting in low encoding efficiency.
The hash-based motion search method is used to determine the motion information by using hash values on non-square and non-rectangular areas of the video block, thereby predicting and converting.
Improves the efficiency of video encoding, especially when dealing with non-square and non-rectangular video blocks, reduces computational complexity and improves encoding performance.
Smart Images

Figure CN113170177B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] In accordance with applicable patent laws and / or the rules applicable to the Paris Convention, this application timely claims the priority and benefits of International Patent Application No. PCT / CN2019 / 070049 filed on January 2, 2019 and International Patent Application No. PCT / CN2019 / 087969 filed on May 22, 2019. For all purposes of U.S. law, the entire disclosure of the above applications is incorporated by reference as part of the disclosure of this application. Technical field
[0003] This document relates to video and image encoding and decoding technologies. Background art
[0004] Despite the progress in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth required for digital video usage is expected to continue to grow. Summary of the invention
[0005] Devices, systems, and methods related to digital video encoding are described, which include hash - based motion estimation. The described methods can be applied to existing video coding standards (e.g., High Efficiency Video Coding (HEVC) and / or Versatile Video Coding (VVC)) and future video coding standards or video codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: using hash - based motion search to determine motion information associated with a current block for the conversion between the current block of a video and the bit - stream representation of the video, where the size of the current block is M×N, where M and N are positive integers and M is not equal to N; applying prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: using hash - based motion search over the region of a non - rectangular and non - square current block to determine motion information associated with the current block for the conversion between the current block of a video and the bit - stream representation of the video; applying prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0008] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for the conversion between a current block of a video and a bitstream representation of the video, determining motion information associated with the current block by using hash-based motion search on a fixed subset of the samples of the current block; applying prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0009] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: as part of the conversion between a current block of a video and a bitstream representation of the video, performing hash-based motion search; for the conversion, determining the rate-distortion cost of the skip mode for each of one or more coding modes of the current block according to the hash-based motion search finding a hash match and the quantization parameter (QP) of a reference frame being no greater than the QP of the current block; applying prediction to the current block based on the rate-distortion cost and a video picture including the current block; and performing the conversion based on the prediction.
[0010] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for the conversion between a current block of a video and a bitstream representation of the video, using hash-based motion search to determine motion information associated with the current block, the hash-based motion search being based on the hash values of square sub-blocks of the current block, where the size of the current block is M×N, and where M and N are positive integers; applying prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0011] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: for the conversion between a current block of a video and a bitstream representation of the video, using hash-based motion search to determine motion information associated with the current block, the hash-based motion search including performing a K-pixel integer motion vector (MV) accuracy check for the hash-based motion search, where K is a positive integer; applying prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0012] In yet another representative aspect, the above method is embodied in processor-executable code and stored in a computer-readable program medium.
[0013] In yet another representative aspect, a device configured to or operable to perform the above method is disclosed. The device may include a processor programmed to implement the method.
[0014] In yet another representative aspect, a video decoder device may implement the methods as described herein.
[0015] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the specification, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example of bottom-up hash calculation is shown.
[0017] Figure 2A With Figure 2B An example of checking for row / column identical information of a 2N×2N block based on the results of six N×N blocks is shown.
[0018] Figure 3 An example of current picture reference is shown.
[0019] Figures 4-9 Is a flowchart of an example for a video processing method.
[0020] Figure 10 Is a block diagram of an example of a hardware platform for implementing the visual media decoding or visual media encoding techniques described herein.
[0021] Figure 11 Is a block diagram of an example video processing system in which the disclosed technology may be implemented. DETAILED DESCRIPTION
[0022] This document provides various techniques that can be used by a decoder of an image or video bitstream to improve the quality of the decompressed or decoded digital video or image. For the sake of brevity, the term "video" is used herein to include picture sequences (traditionally called video) and individual images. Additionally, a video encoder may also implement these techniques during the encoding process in order to reconstruct decoded frames for further encoding.
[0023] The use of section headings in this document is for ease of understanding and does not limit the embodiments and techniques to the corresponding sections. Thus, the embodiments of one section may be combined with the embodiments of other sections.
[0024] 1. Overview
[0025] The invention relates to video coding techniques. Specifically, it relates to motion estimation in video coding. It can be applied to existing video coding standards (such as HEVC), or to the finalized standard (Universal Video Coding). It may also be applicable to future video coding standards or video codecs.
[0026] 2. Background
[0027] Video coding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, and ISO / IEC produced MPEG-1 and MPEG-4 Visual. The two organizations jointly produced the H.262 / MPEG-2 video, H.264 / PEG-4 Advanced Video Coding (AVC), and H.265 / HEVC [1] standards. Starting from H.262, video coding standards are based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the "Joint Exploration Model" (JEM) [2][3]. In April 2018, the Joint Video Exploration Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of reducing the bitrate by 50% compared to HEVC.
[0028] The latest version of the VVC draft, namely the Versatile Video Coding (Draft 2), can be found at:
[0029] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 11_ Ljubljana / wg11 / JVET-K1001-v7.zip
[0030] The latest reference software for VVC called VTM can be found at:
[0031] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-2.1
[0032] Figure 5 is a block diagram of an example implementation of a video encoder. Figure 5 It shows that the encoder implementation has a built-in feedback path within which the video encoder also performs a video decoding function (reconstructing the compressed representation of video data for the encoding of the next video data).
[0033] 2.1 Example of Hash-Based Motion Search
[0034] Hash-based search is applied to 2Nx2N blocks. An 18-bit hash based on the original pixels is used. The first 2 bits are determined by the block size, e.g., 00 represents 8x8, 01 represents 16x16, 10 represents 32x32, and 11 represents 64x64. The next 16 bits are determined by the original pixels.
[0035] For a block, we use a similar approach but different CRC truncation polynomials to calculate two hash values. The first hash value is used for retrieval, and the second hash value is used to exclude certain hash collisions.
[0036] The hash value is calculated as follows:
[0037] For each line, calculate the 16-bit CRC value for all pixels Hash[i].
[0038] Group the line hash values together (Hash[0]Hash[1]…), and then calculate the 24-bit CRC value H.
[0039] The lower 16 bits of H will be used as the lower 16 bits of the hash value of the current block.
[0040] To avoid having one hash value corresponding to too many entries, we avoid adding blocks that meet the following conditions to the hash table
[0041] There is only one pixel value in each line; or
[0042] There is only one pixel value in each column.
[0043] Because for such blocks, blocks with a one-pixel shift can have the same content. Therefore, we exclude these blocks from the hash table. And such blocks can be perfectly predicted by horizontal prediction or vertical prediction, so excluding them from the hash table may not cause a significant performance degradation.
[0044] After building the hash table, motion search is performed as follows.
[0045] For 2Nx2N blocks,
[0046] First, perform hash-based search.
[0047] If a hash match is found, skip the regular integer pixel motion search and fractional pixel motion search. Otherwise, perform the regular motion search.
[0048] Save the integer MV found for the 2Nx2N block.
[0049] For non-2Nx2N blocks, compare the cost of using the MVP as the ME starting point and the cost of using the stored MV of the 2Nx2N block with the same reference picture as the ME starting point. Then select the one with the lower cost as the motion search starting point.
[0050] We also developed an early termination algorithm based on hash search. If all of the following conditions are met, the RDO process will be terminated without checking other modes and CU partitions.
[0051] A hash match is found.
[0052] The quality of the reference block is not worse than the expected quality of the current block (the QP of the reference block is not greater than the QP of the current block).
[0053] The current CU depth is 0.
[0054] 2.2 Additional Examples of Hash-Based Motion Search
[0055] 1. Bottom-Up Hash Value Calculation
[0056] To speed up the calculation of hash values, we propose to calculate the block hash values in a bottom-up manner. The same method as in [1] is used, but different CRC truncation polynomials are used to calculate these two hash values. And the CRC truncation polynomial used is the same as that in the current SCM.
[0057] For a 2Nx2N (N = 2, 4, 8, 16, or 32) block, the hash value is calculated based on the hash results of its four sub NxN blocks rather than pixels. Therefore, intermediate data (hash value results of smaller blocks) can be utilized as much as possible to reduce complexity. For N = 1, i.e., a 2×2 block, the block hash value is directly calculated from pixels. The whole process is as Figure 1 shown.
[0058] 2. Bottom-Up Validity Check of the Process of Adding Blocks to the Hash Table
[0059] In the proposed scheme, the condition for determining whether a block is valid to be added to the hash table remains unchanged. The only change is the step of checking whether a block is valid. Similar to the block hash value calculation, the validity check can also be performed in a bottom-up manner to reuse intermediate results and reduce complexity.
[0060] For all 2x2 blocks in the picture, row / column equality checks are performed based on pixels. The decision results of row / column equality are stored for the validity check of 4x4 blocks. For a 2Nx2N (N = 2, 4, 8, 16, or 32) block, the validity can be simply determined by the row / column equality information of 6 sub NxN blocks in each direction.
[0061] Taking the row equality check as an example. In Fig. 2(a), the four sub NxN blocks are marked as In addition, to check whether the boundary pixels between block 0 and 1 and between block 2 and 3 are the same, for the upper half block and the lower half block, the row equality information should be used for the intermediate NxN blocks 4 and 5 respectively. In other words, the 2N×2N block row equality information can be determined by the row equality information of the 6 sub N×N blocks that have been recorded in the previous step. In this way, pixel checks for different block sizes can be avoided.
[0062] 2.3 Current Picture Reference
[0063] In the High Efficiency Video Coding (HEVC) Screen Content Coding extension (HEVC-SCC) [1] and the current Versatile Video Coding (VVC) Test Model (VTM-3.0) [2], the concept of Current Picture Reference (CPR) (or once called Intra Block Copy (IBC)) has been adopted. IBC extends the concept of motion compensation from inter-frame coding to intra-frame coding. As Figure 3 shown, when applying CPR, the current block is predicted by a reference block in the same picture. Before encoding or decoding the current block, the samples in the reference block must have been reconstructed. Although CPR is not efficient for most sequences captured by cameras, for screen content, it shows significant coding gain. The reason is that there are many repetitive patterns in screen content pictures, such as icons and text characters. CPR can effectively eliminate the redundancy between these repetitive patterns. In HEVC-SCC, if an inter-coded Coding Unit (CU) selects the current picture as its reference picture, the inter-coded CU can apply CPR. In this case, the Motion Vector (MV) is renamed as the Block Vector (BV), and the BV always has integer pixel precision. To be compatible with the main profile of HEVC, the current picture is marked as a "long-term" reference picture in the Decoded Picture Buffer (DPB). It should be noted that similarly, in the multi-view / 3D video coding standard, the inter-view reference pictures are also marked as "long-term" reference pictures.
[0064] After the BV finds its reference block, the prediction can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixels from the original signal. Then, transformation and quantization can be applied as in other coding modes.
[0065] However, when the reference block is outside the picture, or overlaps with the current block, or is outside the reconstructed area, or is outside the valid area restricted by certain constraints, some or all of the pixel values are undefined. Basically, there are two ways to solve this problem. One is to prohibit such situations, for example, in terms of bitstream consistency. The other is to apply padding to those undefined pixel values. The following subsections describe the solutions in detail.
[0066] 2.4 Hash-based Search for Intra Block Copy / Current Picture Reference
[0067] Hash-based search is applied to 4x4, 8x8, and 16x16 blocks. For a block, we use a similar method but different CRC truncation polynomials to calculate two hash values. The first hash value is used for retrieval, and the second hash value is used to exclude some hash collisions. The calculation of the hash values is as follows:
[0068] For each row, calculate the 16-bit CRC value for all pixels Hash[i].
[0069] Group the row hash values together (Hash[0] Hash[1]...), and then calculate the 24-bit CRC value H.
[0070] The lower 16 bits of H will be used as the lower 16 bits of the hash value of the current block.
[0071] To avoid having one hash value corresponding to too many entries, we avoid adding blocks that meet the following conditions to the hash table.
[0072] There is only one pixel value in each row, or
[0073] There is only one pixel value in each column.
[0074] Because for such blocks, the blocks with a one-pixel shift can have the same content. Therefore, we exclude these blocks from the hash table. And such blocks can be perfectly predicted by horizontal prediction or vertical prediction.
[0075] 3. Disadvantages of existing implementations
[0076] Current hash-based search is limited to NxN (square) blocks, which limits its performance.
[0077] 4. Examples of embodiments
[0078] The following detailed techniques should be regarded as examples to explain the general concept. These techniques should not be interpreted narrowly. In addition, these techniques can be combined in any way.
[0079] 1. It is proposed that hash-based motion search can be based on blocks of size MxN, where M is not equal to N.
[0080] a. In one example, M = 2N, 4N, 8N, 16N or 32N.
[0081] b. In one example, N = 2M, 4M, 8M, 16M or 32M.
[0082] c. The hash value of a non-square block can be derived from the hash value of a square block.
[0083] d. In one example, if M is greater than N, the hash value of the MxN block can be derived from the hash values of M / N (N, N) blocks.
[0084] e. Alternatively, if M is less than N, the hash value of the MxN block can be derived from the hash values of N / M (M, M) blocks.
[0085] 2. It is proposed that hash-based motion search can be based on a triangular region.
[0086] a. In one example, the region is defined as x < ky; x, y = 0, 1, 2, … M-1, where x and y are coordinates relative to the block origin, and k is a predefined value.
[0087] b. Alternatively, the region is defined as x > ky, x, y = 0, 1, 2, … M-1
[0088] c. Alternatively, the region is defined as x < M-ky, x, y = 0, 1, 2, … M-1
[0089] d. Alternatively, the region is defined as x > M-ky, x, y = 0, 1, 2, … M-1
[0090] 3. It is proposed that hash-based motion search can be based on a fixed subset of blocks.
[0091] a. In one example, the subset is defined as a set of predefined coordinates relative to the origin of the block.
[0092] 4. Hash-based motion vectors can be added as additional starting points in the MV estimation process.
[0093] a. In integer motion estimation (such as xTZSearch or xTZSearchSelective), the hash-based motion vectors will be checked for best starting point initialization.
[0094] b. If hash-based motion vectors exist, fractional motion estimation can be skipped.
[0095] 5. Early termination
[0096] a. When a hash match is found and the quantization parameter of the reference frame is not greater than that of the current block, only the rate-distortion cost of the skip mode is checked for the ETM_MERGE_SKIP, ETM_AFFINE, and ETM_MERGE_TRIANGLE modes.
[0097] i. Additionally, alternatively, the rate-distortion cost of other modes (such as AMVP (advanced motion vector prediction)) can be skipped.
[0098] ii. Additionally, alternatively, the check for finer-grained block partitioning can be terminated.
[0099] iii. Additionally, alternatively, if the best mode is ETM_HASH_INTER after checking those skip modes, the check for finer-grained block partitioning can be terminated. Otherwise, the check for finer-grained block partitioning cannot be terminated.
[0100] b. In one example, when a hash match is found and the quantization parameter of the reference frame is not greater than that of the current block, the rate-distortion cost of the skip mode is checked only for the ETM_MERGE_SKIP, ETM_AFFINE, ETM_MERGE_TRIANGLE modes, and the ETM_INTER_ME (e.g., AMVR (adaptive motion vector resolution)) mode.
[0101] i. Additionally, alternatively, the rate-distortion cost of other modes (e.g., ETM_INTRA) is skipped from being checked.
[0102] ii. Additionally, alternatively, the finer-grained block partition check is terminated.
[0103] c. Whether to adopt the above method may depend on the size of the block.
[0104] i. In one example, it can be invoked only when the block size is greater than or equal to M x N. ii. In one example, it can be invoked only when the block size is equal to 64x64.
[0105] 6. A method of checking the hash mode of a rectangular block by the sub-square hash value of the rectangular block is proposed.
[0106] a. In one example, this method can also be applied to square blocks.
[0107] 7. A K-pixel integer MV accuracy check is proposed for the hash MV
[0108] a. In one example, K is set to 4.
[0109] 8. The above method can be applied to a part of the allowed block sizes.
[0110] a. In one example, for square blocks, the hash-based method can be applied to N×N blocks where N is less than or equal to a threshold (e.g., 64).
[0111] b. In one example, for non-square (rectangular) blocks, the hash-based method can be applied to M×N blocks where M*N is less than or equal to a threshold (e.g., 64).
[0112] c. In one example, for non-square (rectangular) blocks, the hash-based method can be applied only to 4x8 and 8x4.
[0113] d. The above method can be applied only to certain color components, such as the luminance color component / base color component (e.g., G).
[0114] Figure 4It is a flowchart of method 400 for video processing. Method 400 includes, at operation 410, for the conversion between the current block of a video and the bitstream representation of the video, using hash-based motion search to determine motion information associated with the current block, where the size of the current block is M×N, where M and N are positive integers and M is not equal to N.
[0115] Method 400 includes, at operation 420, applying prediction to the current block based on the motion information and the video picture including the current block.
[0116] Method 400 includes, at operation 430, performing a conversion based on the prediction.
[0117] Figure 5 It is a flowchart of method 500 for video processing. Method 500 includes, at operation 510, for the conversion between the current block of a video and the bitstream representation of the video, determining motion information associated with the current block by using hash-based motion search over the region of the non-rectangular and non-square current block.
[0118] Method 500 includes, at operation 520, applying prediction to the current block based on the motion information and the video picture including the current block.
[0119] Method 500 includes, at operation 530, performing a conversion based on the prediction.
[0120] Figure 6 It is a flowchart of method 600 for video processing. Method 600 includes, at operation 610, for the conversion between the current block of a video and the bitstream representation of the video, determining motion information associated with the current block by using hash-based motion search over a fixed subset of the samples of the current block.
[0121] Method 600 includes, at operation 620, applying prediction to the current block based on the motion information and the video picture including the current block.
[0122] Method 600 includes, at operation 630, performing a conversion based on the prediction.
[0123] Figure 7 It is a flowchart of method 700 for video processing. Method 700 includes, at operation 710, as part of the conversion between the current block of a video and the bitstream representation of the video, performing hash-based motion search.
[0124] Method 700 includes, at operation 720, for each of one or more coding modes of the current block for the conversion, determining the rate-distortion cost of the skip mode according to the hash match found by the hash-based motion search and the quantization parameter (QP) of the reference frame being no greater than the QP of the current block.
[0125] Method 700 includes, at operation 730, applying prediction to a current block based on a rate-distortion cost and a video picture including the current block.
[0126] Method 700 includes, at operation 740, performing a transform based on the prediction.
[0127] Figure 8 is a flowchart of a method 800 for video processing. Method 800 includes, at operation 810, using hash-based motion search to determine motion information associated with a current block for transformation between the current block of a video and a bitstream representation of the video, the hash-based motion search being based on hash values of square sub-blocks of the current block, where the current block has a size of M×N, and where M and N are positive integers.
[0128] Method 800 includes, at operation 820, applying prediction to the current block based on the motion information and a video picture including the current block.
[0129] Method 800 includes, at operation 830, performing a transform based on the prediction.
[0130] Figure 9 is a flowchart of a method 900 for video processing. Method 900 includes, at operation 910, using hash-based motion search to determine motion information associated with a current block for transformation between the current block of a video and a bitstream representation of the video, the hash-based motion search including performing a K-pixel integer motion vector (MV) accuracy check for the hash-based motion search, where K is a positive integer.
[0131] Method 900 includes, at operation 920, applying prediction to the current block based on the motion information and a video picture including the current block.
[0132] Method 900 includes, at operation 930, performing a transform based on the prediction.
[0133] In some embodiments, the following technical solutions may be implemented:
[0134] A1. A video processing method, comprising: using hash-based motion search to determine motion information associated with a current block for transformation between the current block of a video and a bitstream representation of the video, where the size of the current block is M×N, where M and N are positive integers and M is not equal to N; applying prediction to the current block based on the motion information and a video picture including the current block; and performing a transform based on the prediction.
[0135] A2. The method according to solution A1, wherein M = K×N, and wherein K is a positive integer.
[0136] A3. The method according to solution A1, wherein N = K × M, and wherein K is a positive integer.
[0137] A4. The method according to solution A2 or A3, wherein K = 2, 4, 8, 16 or 32.
[0138] A5. The method according to solution A1, wherein the hash value of the current block is derived based on the hash value of the square blocks of the video.
[0139] A6. The method according to solution A5, wherein since it is determined that M is greater than N, the hash value of the current block is derived based on the hash values of M / N blocks of size N×N that form the current block.
[0140] A7. The method according to solution A5, wherein since it is determined that N is greater than M, the hash value of the current block is derived based on the hash values of N / M blocks of size M×M that form the current block.
[0141] A8. A method for video processing, comprising: for the conversion between the current block of the video and the bitstream representation of the video, determining motion information associated with the current block by using hash-based motion search on the region of the non-rectangular and non-square current block; applying prediction to the current block based on the motion information and the video picture including the current block; and performing the conversion based on the prediction.
[0142] A9. The method according to solution A8, wherein the region includes a triangular region, the triangular region includes sample points with coordinates (x, y), wherein x and y are relative to the starting point of the current block, wherein the size of the current block is M×N, and M and N are positive integers.
[0143] A10. The method according to solution A9, wherein x < k × y, k is a predefined value, and y = 0, 1,..., M - 1.
[0144] A11. The method according to solution A9, wherein x > k × y, k is a predefined value, and y = 0, 1,..., M - 1.
[0145] A12. The method according to solution A9, wherein x < (M - k × y), k is a predefined value, and y = 0, 1,..., M - 1.
[0146] A13. The method according to solution A9, wherein x > (M - k × y), k is a predefined value, and y = 0, 1,..., M - 1.
[0147] A14. A video processing method, comprising: for the conversion between a current block of a video and a bitstream representation of the video, determining motion information associated with the current block by using hash-based motion search on a fixed subset of samples of the current block; applying prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0148] A15. The method according to solution A14, wherein the fixed subset of samples includes a set of predefined coordinates relative to a starting point of the current block.
[0149] A16. The method according to any one of solutions A1 to A15, wherein the hash-based motion search includes adding a hash-based motion vector (MV) as a starting point during the MV estimation technique.
[0150] A17. The method according to solution A16, wherein the hash-based MV is checked for initialization of the best starting point in integer motion estimation.
[0151] A18. The method according to solution A16 or A17, wherein the fractional motion estimation operation is skipped due to the determination of the existence of the hash-based MV.
[0152] A19. The method according to any one of solutions A1 to A18, wherein the conversion generates the current block from the bitstream representation.
[0153] A20. The method according to any one of solutions A1 to A18, wherein the conversion generates the bitstream representation from the current block.
[0154] A21. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method according to any one of solutions A1 to A20.
[0155] A22. A computer program product stored on a non-transitory computer-readable medium, the computer program product including program code for performing the method according to any one of solutions A1 to A20.
[0156] In some embodiments, the following technical solutions can be implemented:
[0157] B1. A video processing method includes: performing hash-based motion search as part of a conversion between a current block of a video and a bitstream representation of the video; finding, based on the hash-based motion search, a hash match and a quantization parameter (QP) of a reference frame that is not greater than the QP of the current block, and for the conversion, determining a rate-distortion cost of a skip mode for each of one or more coding modes of the current block; applying prediction to the current block based on the rate-distortion cost and a video picture including the current block; and performing the conversion based on the prediction.
[0158] B2. The method according to solution B1, wherein the one or more coding modes include an ETM_MERGE_SKIP mode, an ETM_AFFINE mode, an ETM_MERGE_TRIANGLE mode, or an ETM_INTER_ME mode.
[0159] B3. The method according to solution B1 or B2, wherein the one or more coding modes do not include an advanced motion vector prediction (AMVP) mode or an ETM_INTRA mode.
[0160] B4. The method according to solution B1, further includes: after the determination, terminating a finer-grained block partition check.
[0161] B5. The method according to solution B4, wherein the rate-distortion cost of the skip mode for the ETM_HASH_INTER mode is determined to be lower than the rate-distortion cost of the skip mode for any other coding mode.
[0162] B6. The method according to any one of solutions B1 to B5, wherein the hash-based motion search and determination are performed based on the height or width of the current block.
[0163] B7. The method according to solution B6, wherein the size of the current block is M×N or larger.
[0164] B8. The method according to solution B6, wherein the size of the current block is 64×64.
[0165] B9. The method according to any one of solutions B1 to B8, wherein the conversion generates the current block from the bitstream representation.
[0166] B10. The method according to any one of solutions B1 to B8, wherein the conversion generates the bitstream representation from the current block.
[0167] B11. An apparatus in a video system includes a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method according to any one of solutions B1 to B10.
[0168] B12. A computer program product stored on a non - transitory computer - readable medium, the computer program product including program code for performing the method of any one of Solutions B1 to B10.
[0169] In some embodiments, the following technical solutions can be implemented:
[0170] C1. A video processing method, comprising: using hash - based motion search to determine motion information associated with a current block for conversion between the current block of a video and a bit - stream representation of the video, the hash - based motion search being based on hash values of square sub - blocks of the current block, wherein the size of the current block is M×N, and wherein M and N are positive integers; applying a prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0171] C2. The method according to Solution C1, wherein M is not equal to N.
[0172] C3. The method according to Solution C1, wherein M is equal to N.
[0173] C4. A video processing method, comprising: using hash - based motion search to determine motion information associated with a current block for conversion between the current block of a video and a bit - stream representation of the video, the hash - based motion search including performing a K - pixel integer motion vector (MV) accuracy check for the hash - based motion search, wherein K is a positive integer; applying a prediction to the current block based on the motion information and a video picture including the current block; and performing the conversion based on the prediction.
[0174] C5. The method according to Solution C4, wherein K = 4.
[0175] C6. The method according to any one of Solutions C1 to C5, wherein the size of the current block is N×N, where N is a positive integer less than or equal to a threshold.
[0176] C7. The method according to any one of Solutions C1 to C5, wherein the size of the current block is M×N, where M and N are positive integers and M is not equal to N, and wherein M or N is less than or equal to a threshold.
[0177] C8. The method according to Solution C6 or C7, wherein the threshold is 64.
[0178] C9. The method according to any one of Solutions C1 to C5, wherein the size of the current block is 8×4 or 4×8.
[0179] C10. The method according to any one of Solutions C1 to C5, wherein the hash - based motion search is applied to the luminance component of the video.
[0180] Method according to any one of Solutions C1 to C5, wherein hash-based motion search is applied to a base color component of a video.
[0181] Method according to any one of Solutions C1 to C11, wherein a transform generates a current block from a bitstream representation.
[0182] Method according to any one of Solutions C1 to C11, wherein a transform generates a bitstream representation from a current block.
[0183] An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method according to any one of Solutions C1 to C13.
[0184] A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the method according to any one of Solutions C1 to C13.
[0185] In some embodiments, the following technical solutions may be implemented:
[0186] A method for processing video, comprising: determining, by a processor, motion information regarding a first video block using hash-based motion search, the first video block having a size of M×N, wherein M is not equal to N; and performing further processing on the first video block using the motion information.
[0187] Method according to Solution D1, wherein M is 2N, 4N, 8N, 16N or 32N.
[0188] Method according to Solution D1, wherein N is 2M, 4M, 8M, 16M or 32M.
[0189] Method according to Solution D1, wherein determining the motion information using hash-based motion search includes deriving a hash value of the non-square video block using hash values of square video blocks that are sub-blocks of the non-square video block.
[0190] Method according to Solution D1, further comprising: determining that M is greater than N, and wherein determining the motion information using hash-based motion search includes deriving a hash value of the first video block from M / N (N, N) video block hash values.
[0191] D6. The method according to solution D1 further includes: determining that M is less than N, and wherein using hash-based motion search to determine motion information includes deriving the hash value of a first video block from M / N (N, N) video block hash values.
[0192] D7. A method for processing video includes: using, by a processor, hash-based motion search to determine motion information regarding a first video block having a non-rectangular and non-square shape; and performing further processing on the first video block using the motion information.
[0193] D8. The method according to solution D7, wherein the shape is triangular.
[0194] D9. The method according to solution D7, wherein the shape is a region defined by x < ky; x, y = 0, 1, 2,... M-1, where x and y are relative coordinates with respect to the starting point of the first video block, and where k is a predefined value.
[0195] D10. The method according to solution D7, wherein the shape is a region defined by x > ky, x, y = 0, 1, 2,... M-1, where x and y are relative coordinates with respect to the starting point of the first video block, and wherein k is a predefined value.
[0196] D11. The method according to solution D7, wherein the shape is defined as x < M-ky, x, y = 0, 1, 2,... M-1, where x and y are relative coordinates with respect to the starting point of the first video block, and wherein k is a predefined value.
[0197] D12. The method according to solution D7, wherein the shape is defined as x > M-ky, x, y = 0, 1, 2,... M-1, where x and y are relative coordinates with respect to the starting point of the first video block, and wherein k is a predefined value.
[0198] D13. The method according to solutions D1-D12, wherein the hash-based motion search is based on a fixed subset of the first video block.
[0199] D14. The method according to solution D13, wherein the fixed subset is defined by a set of predefined coordinates with respect to the starting point of the first video block.
[0200] D15. The method according to solutions D1-D14, wherein the hash-based motion search includes adding a hash-based motion vector (MV) as a starting point during the MV estimation technique.
[0201] D16. The method according to solution D15, wherein the hash-based MV is checked for initialization of the best starting point in integer motion estimation.
[0202] The method according to Solutions D1-D16 further includes: determining a hash-based motion search that results in a hash match; determining that a quantization parameter (QP) of a reference frame is greater than the QP of a first video block; and determining a rate-distortion cost of a skip mode based on the determination that the hash-based motion search that results in a hash match and the QP of the reference frame is greater than the QP of the first video block, wherein further processing of the first video block is performed based on the rate-distortion cost.
[0203] The method according to Solution D17, wherein determining the rate-distortion cost includes skipping advanced motion vector prediction (AMVP).
[0204] The method according to Solution D17, wherein determining the rate-distortion cost includes stopping a finer-grained block partition check.
[0205] The method according to Solution D17, wherein determining the rate-distortion cost is based on the size of the first video block.
[0206] The method according to Solution D20, wherein the size of the first video block is greater than or equal to M×N.
[0207] The method according to Solution D20, wherein the size of the first video block is 64×64 pixels.
[0208] The method according to Solution D1, wherein the hash pattern of the first video block is based on a sub-square hash value and the first video block has a rectangular shape.
[0209] The method according to Solution D1, wherein the hash pattern of the first video block is based on a sub-square hash value and the first video block has a square shape.
[0210] The method according to Solution D1 or D7, wherein the motion information includes a motion vector (MV).
[0211] The method according to Solution D1 or D7 further includes: performing a K-pixel integer motion vector (MV) accuracy check on the hash-based motion search.
[0212] The method according to Solution D26, wherein K is 4.
[0213] The method according to Solutions D1-D27, wherein the first video block is a partial size of an allowed block size.
[0214] The method according to Solutions D1-D27, wherein the first video block has a square shape and the hash-based motion search is applied to an N×N video block, wherein N is less than or equal to a threshold.
[0215] D30. The method according to solution D29, wherein the threshold is 64.
[0216] D31. The method according to solutions D1 - D27, wherein the first video block has a rectangular shape, and hash - based motion search is applied to an M×N video block, where M*N is less than or equal to the threshold.
[0217] D32. The method according to solution D31, wherein the threshold is 64.
[0218] D33. The method according to solutions D1 - D27, wherein the first video block is 4×8 or 8×4.
[0219] D34. The method according to solutions D1 - D27, wherein hash - based motion search is applied to color components.
[0220] D35. The method according to solutions D1 - D16, further comprising: determining the hash - based motion search that results in a hash match; determining that the quantization parameter (QP) of the reference frame is not greater than the QP of the first video block; determining the rate - distortion cost of the skip mode according to the hash - based motion search that results in a hash match and the QP of the reference frame being greater than the QP of the first video block, wherein further processing of the first video block is performed based on the rate - distortion cost.
[0221] D36. The method according to solution D35, wherein the skip mode includes skip modes for affine mode, triangular mode, or adaptive motion vector resolution (AMVR) mode.
[0222] D37. The method according to solution D35, wherein determining the rate - distortion cost includes stopping the check for finer - grained block partitioning.
[0223] D38. A video decoding device, comprising a processor configured to implement the method described in one or more of solutions D1 to D37.
[0224] D39. According to a video encoding device, comprising a processor configured to implement the method described in one or more of solutions D1 to D37.
[0225] D40. A computer program product having computer code stored thereon, which when executed by a processor causes the processor to implement the method described in any one of solutions D1 to D37.
[0226] Figure 10It is a block diagram of a video processing apparatus 1000. The apparatus 1000 can be used to implement one or more methods described herein. The apparatus 1000 can be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1000 can include one or more processors 1002, one or more memories 1004, and video processing hardware 1006. The (one or more) processors 1002 can be configured to implement one or more methods described in this document. The (one or more) memories 1004 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 1006 can be used to implement some of the techniques described in this document in hardware circuits.
[0227] In some embodiments, a video encoding method can be implemented using an apparatus implemented on a hardware platform as described with respect to Figure 10 the description.
[0228] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, the conversion from video blocks to the bitstream representation of the video will use the video processing tool or mode when the decision or determination to enable the video processing tool or mode is made. In another example, when a video processing tool or mode is enabled, the decoder processes the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the decision or determination.
[0229] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In an example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion from video blocks to the bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder processes the bitstream knowing that the bitstream has not been modified using the video processing tool or mode enabled based on the decision or determination.
[0230] Figure 11FIG. 0 is a block diagram illustrating an example video processing system 1100 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all components of system 1100. System 1100 may include an input 1102 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8- or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1102 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0231] System 1100 may include an encoding component 1104, which may implement the various encoding or coding methods described in this document. Encoding component 1104 may reduce the average bit rate of the video from input 1102 to the output of encoding component 1104 to produce an encoded representation of the video. Thus, encoding techniques are sometimes referred to as video compression or video transcoding techniques. As represented by component 1106, the output of encoding component 1104 may be stored or transmitted via the connected communication. The stored or transmitted bitstream (or encoded) representation of the video received at input 1102 may be used by component 1108 to generate pixel values or a displayable video, which is sent to display interface 1110. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "encoding" operations or tools, it should be understood that encoding tools or operations are used at the encoder and corresponding decoding tools or operations that reverse the encoding results will be performed by the decoder.
[0232] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0233] The disclosures and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or a multiprocessor or a group of computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0234] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including a compiled or interpreted language), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to be executed on one or more computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0235] The processing and logical flows described in this specification can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0236] For example, a processor suitable for executing a computer program includes general and special-purpose microprocessors, as well as any one or more of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes instructions and one or more storage devices that store instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as, for example, magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to one or more mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and the memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0237] Although this patent document contains many details, it should not be construed as limiting any subject or the scope of any claims, but rather as a description of features of particular embodiments of a particular technology. Certain features described in the context of separate embodiments of this patent document may also be implemented in combination in a single embodiment. Conversely, the various functions described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Moreover, although the above features may be described as acting in certain combinations, and even initially claimed as such, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0238] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood to mean that such operations must be performed in the particular order shown or in sequential order to achieve a desired result, or that all illustrated operations must be performed. Moreover, the separation of various system components described in the embodiments of this patent document should not be understood to be required in all embodiments.
[0239] Only some implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: For the conversion between the current block of a video and the bitstream of the video, motion information associated with the current block is determined by using hash-based motion search, where the hash-based motion search is performed on at least one of the following: the region of the current block that is non-rectangular and non-square, or a fixed subset of the samples of the current block; Apply prediction to the current block based on the motion information and the video picture including the current block; And Perform the conversion based on the prediction, wherein the method further includes: Determine that the hash-based motion search results in a hash match; Determine that the quantization parameter QP of the reference frame is greater than the QP of the current block of the video; and Based on (a) the determination that the hash-based motion search results in a hash match and (b) the determination that the QP of the reference frame is greater than the QP of the current block of the video, determine the rate-distortion cost of the skip mode, wherein further processing of the current block is based on the rate-distortion cost, wherein determining the rate-distortion cost is further based on the size of the current block of the video, and wherein determining the rate-distortion cost includes skipping advanced motion vector prediction AMVP, or stopping the check of finer-grained block partitioning.
2. The method according to claim 1, wherein, The region includes a triangular region, the triangular region includes samples with coordinates (x, y), where x and y are coordinates relative to the starting point of the current block, where the size of the current block is M×N, and M and N are positive integers.
3. The method according to claim 2, wherein, x < k×y, k is a predefined value, and y = 0, 1,..., M - 1.
4. The method according to claim 2, wherein, x > k×y, k is a predefined value, and y = 0, 1,..., M - 1.
5. The method according to claim 2, wherein, x < (M - k×y), k is a predefined value, and y = 0, 1,..., M - 1.
6. The method according to claim 2, wherein, x > (M - k×y), k is a predefined value, and y = 0, 1,..., M - 1.
7. The method according to claim 1, wherein, The fixed subset of samples includes a set of predefined coordinates relative to the starting point of the current block.
8. The method according to any one of claims 1 to 7, wherein, The hash-based motion search includes adding a hash-based motion vector MV as a starting point during the MV estimation technique.
9. The method according to claim 8, wherein the hash-based MV is checked for initialization of the best starting point in integer motion estimation.
10. The method according to claim 8, wherein the fractional motion estimation operation is skipped due to the determination of the existence of the hash-based MV.
11. The method according to claim 1, wherein, The conversion includes encoding the current block into the bitstream.
12. The method according to claim 1, wherein, The conversion includes decoding the current block from the bitstream.
13. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, When executed by the processor, the instructions cause the processor to implement the method according to any one of claims 1 to 12.
14. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Hash table construction and availability checking for hash-based block matching
CN105393537A