Fast Method, System and Storage Medium for Inter-Frame Prediction Based on CU Correlation

By adopting the fast mode decision-making and reference frame selection scheme of CU correlation in HEVC, the PU mode and reference frame selection are optimized, and the problem of high complexity of inter-frame prediction calculation of HEVC is solved, and efficient video encoding of real-time propagation system is realized.

CN113852811BActive Publication Date: 2025-07-18E SURFING VISION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110233573.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-03
Publication Date
2025-07-18
Estimated Expiration
2041-03-03

AI Technical Summary

Technical Problem

The computational complexity in the inter prediction process in the HEVC standard is too high, especially in the execution of motion estimation of multiple reference frames and rich inter-frame modes, resulting in unfriendly real-time sexual propagation systems.

Method used

Through fast mode decision-making and reference frame selection scheme based on coding unit (CU) correlation, the PU mode decision-making and reference frame number are optimized to reduce encoding complexity.

Benefits of technology

While maintaining the quality of video encoding, it significantly reduces the computational complexity of HEVC inter-frame prediction and is suitable for real-time propagation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113852811B_ABST
    Figure CN113852811B_ABST
Patent Text Reader

Abstract

The present invention provides a fast inter-frame prediction method, system and storage medium based on CU correlation. The present invention includes reading an encoded frame under a low-delay B-frame (LDB), low-delay P-frame (LDP) or random access (RA) coding configuration, and using the information correlation of parent and child coding units (CUs). When the best mode of the parent CU of the current depth coding unit (CU) is the Skip mode, the best reference frame of the parent CU is directly selected as the best reference frame for the current depth CU, thereby skipping unnecessary motion estimation processes and effectively reducing the inter-frame prediction coding time. Moreover, when the parent CU is encoded using the Skip mode, it is early determined whether to skip the remaining modes according to the best mode after the current CU is completed in the 2N×2N mode, thereby greatly reducing the computational complexity of HEVC inter-frame prediction while maintaining good video coding quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high - efficiency video coding, and particularly to the field of fast algorithms for inter - frame prediction in video coding. Background Art

[0002] The goal of video coding is to obtain the optimal output video quality under the constraint of bitrate. High Efficiency Video Coding (HEVC), namely H.265, is the latest international video coding standard at present. By adopting a flexible quadtree partitioning structure and rich intra - frame and inter - frame prediction modes, it greatly improves the coding efficiency. Compared with the previous generation video coding standard H.264 / AVC, its coding efficiency is doubled, but the computational complexity of the encoder also increases sharply.

[0003] Different from the fixed macro - block partitioning method in H.264 / AVC, in order to flexibly and efficiently represent video content with different textures in a video scene, HEVC introduces three structural concepts for block partitioning: Coding Unit (CU), Prediction Unit (PU), and Transform Unit (TU). The separation of these three blocks makes the transform, prediction, and entropy coding processes more flexible, and also makes the block partitioning more in line with the texture characteristics of the video, ensuring the optimization of coding performance.

[0004] Currently, the traditional fast algorithms for HEVC inter - frame prediction mainly predict the depth of the current Coding Tree Unit (CTU) based on the spatio - temporal correlation of video frames and information such as rate - distortion cost (RDC). However, in the current HEVC standard inter - frame prediction process, performing motion estimation on all reference frames introduces time consumption, resulting in a relatively complex inter - frame prediction complexity. In addition, when testing with the HM8.0 encoder on an ordinary PC, it is found that the mode selection of CTU consumes more than 2 / 3 of the overall coding time, so it is difficult to implement in an actual encoder, especially for some systems that require real - time transmission (such as video conferencing transmission systems and network live broadcast systems with HEVC as the coding standard).

[0005] Therefore, a solution is needed to solve the time consumption caused by performing motion estimation on multiple reference frames and performing rich inter - frame modes, and effectively reduce the computational complexity of inter - frame coding while maintaining video coding quality. Summary of the Invention

[0006] The present invention content is provided to introduce some concepts in a simplified form that will be further described in the following detailed description. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.

[0007] According to an embodiment of the present invention, a method for encoding a current frame in HEVC is described, including: (a) reading in the current frame under a low-delay P-frame (LDP), low-delay B-frame (LDB), or random access (RA) coding configuration; (b) performing coding tree unit (CTU) partitioning on the current frame; (c) if the current frame is not an I-frame, encoding the partitioned CTUs according to a fast mode decision scheme based on coding unit (CU) correlation and a fast reference frame selection scheme based on coding unit (CU) correlation; (d) if the currently encoded CTU is not the last CTU of the current frame, obtaining the next CTU and repeating steps (a) to (d); wherein the fast mode decision scheme based on coding unit (CU) correlation specifies that when performing inter-frame prediction mode encoding, if for the current CU with a depth not equal to 0, the best mode of its parent CU is the Skip mode, and after the current CU finishes Skip / Merge and Inter_2N×2N, the best mode is the Skip mode, then determine that the best mode of the current CU is the Skip mode and terminate the execution of the remaining prediction unit (PU) modes; otherwise, perform all PU modes for the current CU; wherein the fast reference frame selection scheme based on coding unit (CU) correlation specifies that if the parent CU of the current CU with a depth not equal to 0 performs inter-frame prediction mode encoding with the Skip mode as the best mode, then during the motion estimation process for each PU mode of the current CU, directly select the best reference frame of the parent CU as the best reference frame for the current mode, so as to perform motion estimation only on this best reference frame to select the best motion vector.

[0008] According to another embodiment of the present invention, a method for performing inter-frame prediction based on CU correlation is described, including: (a) obtaining a current CU with a current depth; (b) determining whether the current depth of the current CU is 0; (c) if the current depth is 0, performing all PU modes on the current CU, and determining and temporarily storing the best mode and the best reference frame index of the current CU; (d) if the current depth is not 0, determining whether the best mode of the parent CU of the current CU is the Skip mode; (e) if the best mode of the parent CU of the current CU is not the Skip mode, performing all PU modes on the current CU, determining the best mode and the best reference frame index of the current CU, and temporarily storing the best mode and the best reference frame index of the current CU when the current depth is not 3; (f) if the best mode of the parent CU of the current CU is the Skip mode, performing the Skip mode and the Inter_2N×2N mode on the current CU, and determining the best mode of the current CU; (g) if the best mode of the current CU is the Skip mode and the current depth is not 3, temporarily storing the best mode and the best reference frame index of the current CU; (h) if the best mode of the current CU is not the Skip mode, performing the remaining PU modes on the current CU, determining the best mode and the best reference frame index of the current CU, and temporarily storing the best mode and the best reference frame index of the current CU when the depth is not 3; (i) if the current depth is not 3, repeating the above steps (a)-(h) for the next depth.

[0009] According to still another embodiment of the present invention, a computer-readable storage medium is described. The computer-readable storage medium stores instructions executable by a processor, and the instructions executable by the processor are used to execute the above method when executed by the processor.

[0010] According to another embodiment of the present invention, a system for performing inter-frame prediction based on CU correlation is described, including: a processor; a memory, and the memory stores instructions that can execute the above method when executed by the processor.

[0011] By reading the following detailed description and referring to the associated drawings, these and other features and advantages will become apparent. It should be understood that the foregoing general description and the following detailed description are illustrative only and do not limit the various aspects claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to understand in detail the manner in which the above-described features of the present invention are used, a more specific description of the above briefly summarized content can be made with reference to the various embodiments, some of which are shown in the drawings. However, it should be noted that the drawings only show some typical aspects of the present invention and should not be considered to limit its scope, as the description may allow other equally effective aspects.

[0013] Figure 1 Fig. 100 shows a schematic diagram of the structure division between a CU and its corresponding coding tree CTU in the prior art;

[0014] Figure 2 Fig. 200 shows a flowchart of a method for encoding a current frame in HEVC according to an embodiment of the present invention;

[0015] Figure 3 Fig. 300 shows a flowchart of a method for performing mode selection by adopting a fast mode decision scheme based on CU correlation in inter-frame prediction according to an embodiment of the present invention;

[0016] Figure 4 Fig. 400 shows a flowchart of a method for selecting an optimal reference frame by adopting a fast reference frame selection scheme in inter-frame prediction according to an embodiment of the present invention; and

[0017] Figure 5 Fig. 500 shows a block diagram of an exemplary computing device according to an embodiment of the present invention. Detailed Embodiments

[0018] The present invention will be described in detail below with reference to the accompanying drawings, and the features of the present invention will be further manifested in the following specific description.

[0019] Introduction to the HEVC Standard:

[0020] The coding unit CU is the basic coding unit in the HEVC standard, and operations such as prediction, transformation, quantization, and entropy coding in the coding process are all completed based on the CU. HEVC applies a recursive structure of a quadtree for the division of the CU. Figure 1 Fig. 100 shows a schematic diagram of the structure division between a CU and its corresponding coding tree CTU in the prior art. As Figure 1 can be seen, a CTU is recursively divided according to the coding tree, and it can contain one or more CUs. For a CU with a depth of 0 and a size of 64×64, it is usually referred to as a coding tree unit CTU, which generally serves as the root node of the CU depth division. That is, the maximum coding unit size is 64×64 pixels, the minimum coding unit size is 8×8 pixels, and the large coding unit is recursively subdivided into 8×8 pixel sizes in a quadtree manner, with a maximum depth of 3.

[0021] The prediction unit PU is the basic unit for the prediction process. Based on the CU, it makes mode judgments according to intra-frame prediction or inter-frame prediction, and defines all prediction modes of the CU. Its maximum unit is the same size as the current CU. The PU modes mainly include intra-frame prediction mode (Intra Mode) and inter-frame prediction mode (Inter Mode). Among them, the intra-frame prediction mode includes two segmentation methods: 2N×2N and N×N, and N×N is only available when the current CU depth value is the minimum depth. The inter-frame prediction mode includes Merge mode, Skip mode, and general inter-frame (Inter) mode. The Merge mode can be used for all sizes of the PU, and Skip is a special case of Merge. It only occurs when the PU size is 2N×2N, the Merge mode is adopted, and the residual coding information is 0, indicating that the current coding mode is the Skip mode. There are 8 general inter-frame modes, mainly divided into two categories: symmetric segmentation and asymmetric segmentation. Among them, 2N×2N, 2N×N, N×2N, and N×N are 4 symmetric modes, and 2N×nU, 2N×nD, nL×2N, and nR×2N are 4 asymmetric modes. U, D, L, and R represent up, down, left, and right respectively, and the asymmetric division form is only used for CUs of sizes 32×32 and 16×16, and the N×N of the symmetric division form is only used for CUs of size 8×8. For example, 2N×nU and 2N×nD are divided in the ratio of 1:3 and 3:1 up and down respectively, and nL×2N and nR×2N are divided in the ratio of 1:3 and 3:1 left and right respectively.

[0022] Intra-frame prediction uses the pixel values of adjacent encoded blocks in the current frame to predict the pixel values of the current unencoded block, and encodes according to the difference between the predicted pixel values and the original values to effectively remove the spatial redundancy of the video. Inter-frame prediction mainly utilizes the similarity between consecutive images, and finds an optimal matching block in the encoded image through motion estimation (ME) and motion compensation (MC). The closer the pixel values of the matching block are to those of the original block, the more accurate the reconstructed pixel values will be. The task of motion estimation ME is to find the optimal corresponding block for the current encoded block in the encoded blocks and calculate the offset of the corresponding block, that is, the motion vector (MV, Motion Vector); while motion compensation MC is the process of obtaining the estimated value of the current frame according to the motion vector and the inter-frame prediction method. ME is a dynamic process, which involves many calculations, such as calculating differences, search algorithms, motion vector prediction (MVP, Motion Vector Predicting), etc.; MC is a static process, which estimates the corresponding block according to relevant information, such as MV, inter-frame prediction method, etc., and is equivalent to an index table.

[0023] The input video in the encoder is actually composed of a series of image sequences with strong correlations. In the HEVC standard, these video sequences are divided into different groups of pictures (GOPs), and the number of images in each GOP is determined by different profiles. Within each GOP, HEVC defines three types of frames: a frame that does not refer to other frames and is fully encoded is called an I-frame; a frame that is generated by referring to the previous I-frame and only encodes the differential part is called a P-frame; a frame that is encoded by referring to frames before and after is called a B-frame. The HEVC reference model (HM) encoder defines three different configurations according to the different types of encoded data: all intra (AI), low delay (LD), and random access (RA). Among them, the AI configuration is mainly used for intra-frame predictive coding, the LD configuration is mainly applied to real-time scenarios, and the RA configuration has the highest coding efficiency.

[0024] The LD configuration is further divided into the LDP (Low-Delay P) and LDB (Low-Delay B) configurations. In the LDP configuration, only the first frame is encoded as an I-frame, and all subsequent frames are encoded as P-frames. P-frames are only allowed to refer to reference frames that are earlier in the playback order, while B-frames refer to reference frames in both directions. Therefore, B-frames have a higher coding efficiency in the case of low delay. The RA configuration adopts a hierarchical B-frame structure, and all frames are numbered in the coding order. Because of the bidirectional B-frame hierarchical prediction structure, the RA configuration has a higher coding efficiency than other configuration methods. In the RA configuration, I-frames are inserted periodically to reduce the impact of transmission errors.

[0025] Problems encountered:

[0026] As described above, in HEVC, the CT is the basic coding unit. Each CTU can be divided into CUs of different sizes, and each CU can use PUs with different splitting modes for inter-frame prediction. Usually, the HEVC standard uses the rate-distortion RD value as the best mode judgment criterion.

[0027] The block partitioning of HEVC inter prediction adopts a recursive traversal method based on a quadtree structure. Each CU is recursively divided into 4 sub-CUs equally. PU mode prediction is performed at the CU layer, that is, different prediction modes are traversed respectively, including Merge, Skip, and Inter modes. Taking a CTU with a size of 64×64 and a maximum coding depth of 3 as an example, without considering the complexity of the Merge and Skip modes, only the complexity of the Inter mode is analyzed. When the coding depth is 0, 7 RD values need to be calculated; when the coding depth is 1, 28 RD values need to be calculated; when the coding depth is 2, 112 RD values need to be calculated; when the coding depth is 3, 256 RD values need to be calculated. Through the above analysis, a CTU needs to calculate 403 RD values to determine the best prediction mode.

[0028] In addition, HEVC adopts the technology of the reference picture set (RPS) to manage the decoded frames for use as references for subsequent pictures. HEVC supports the multi-reference picture technology. For example, 4 or 2 active reference pictures can be configured, which doubles the complexity of the motion estimation for each PU.

[0029] Based on the HEVC standard algorithm, the present invention optimizes the decision-making of the PU mode by utilizing the correlation between CUs, reduces the number of mode selections, optimizes the selection scheme of reference pictures, and adaptively reduces the number of reference pictures, which can effectively reduce the coding complexity while ensuring the video compression quality.

[0030] Figure 2 The flowchart of a method 200 for encoding a current picture in HEVC according to an embodiment of the present invention is shown. According to an embodiment of the present invention, the current picture may be a picture that is currently about to be encoded in a series of encoded pictures. In step 201, the current picture is read in under a low-delay P picture (LDP) or low-delay B picture (LDB) or random access (RA) coding configuration. In step 202, the current picture is partitioned into CTUs. In step 203, it is determined whether the current picture is an I picture. If so, go to step 204, and perform HEVC standard I picture encoding on the current picture, that is, perform intra prediction encoding on all CTUs, and the process ends. If not, in step 205, the partitioned CTUs are encoded. According to an embodiment of the present invention, the encoding of CTUs may adopt, for example, an inter prediction technique based on a fast mode decision scheme and a fast reference picture selection scheme according to an embodiment of the present invention, which will be described in detail in Figure 3 and Figure 4 below. In step 206, it is determined whether the currently encoded CTU is the last CTU of the current picture. If so, the process ends. If not, in step 207, the next CTU is obtained, and step 205 is performed on the next CTU until all CTUs are encoded.

[0031] Figure 3 FIG. 300 is a flowchart of a method for performing mode selection by adopting a fast mode decision scheme based on CU correlation in inter-frame prediction according to an embodiment of the present invention. This method is applied to Figure 2 step 205 in

[0032] In the HEVC standard, as previously introduced, a rich set of inter-frame prediction modes are introduced. Generally, the order of HEVC standard inter-frame prediction modes is Skip / Merge, Inter_2N×2N, Inter_N×2N, Inter_2N×N, and asymmetric partition modes to adapt to image blocks with different characteristics.

[0033] According to an embodiment of the present invention, the fast mode decision scheme mainly addresses the drawback of excessively high complexity caused by the rich inter-frame prediction modes in HEVC. Specifically, if the best mode of the parent CU of the current CU is the Skip mode (which usually indicates that the current image block has consistent motion characteristics or a simple background with surrounding image blocks), then after the current CU finishes executing the Skip / Merge and Inter_2N×2N modes, if the best mode is the Skip mode, the remaining modes (e.g., the remaining Inter_N×2N, Inter_2N×N, and asymmetric partition modes) are skipped; otherwise, all inter-frame prediction modes should be executed, and finally, the best mode is selected based on the rate-distortion cost (RDC).

[0034] According to an embodiment of the present invention, in the fast mode decision scheme: when performing inter-frame prediction mode encoding, if the current CU simultaneously satisfies condition 1 (i.e., for a CU with a depth not equal to 0, the best mode of the parent CU is the Skip mode) and condition 2 (i.e., after the current CU finishes executing Skip / Merge and Inter_2N×2N, the best mode is the Skip mode), then it is determined that the best mode of the current CU is the Skip mode, and the execution of the remaining modes is terminated. If conditions 1 and 2 cannot be satisfied simultaneously, the execution of the remaining modes at the current depth continues, and finally, the best mode of the CU at the current depth is determined. The following refers to Figure 3 to specifically describe the process of this fast mode decision.

[0035] In step 301, obtain the current CU with the current depth in the CTU.

[0036] In step 302, determine whether the current depth of the current CU is 0. If it is 0, then proceed to step 303. If the current depth is not equal to 0, then proceed to step 304.

[0037] In step 303, all PU modes are executed for the current CU, and the best mode and the best reference frame index of the current CU are determined. According to an embodiment of the present invention, for the current CU with a current depth less than 3, executing all PU modes for the current CU includes sequentially executing Skip / Merge, Inter_2N×2N, Inter_N×2N, Inter_2N×N, and an asymmetric partitioning mode for the current CU. For the current CU with a current depth of 3, executing all PU modes for the current CU includes sequentially executing Skip / Merge, Inter_2N×2N, Inter_N×2N, Inter_2N×N for the current CU. According to an embodiment of the present invention, the best mode is selected according to the rate-distortion cost (RDC) of each PU mode. According to an embodiment of the present invention, reference Figure 4 is used to describe how to determine the best reference frame index of the current CU.

[0038] In step 308, it is determined whether the current depth is 3. If so, the process ends. If not, step 309 is entered. That is, if the current CU is already a CU with a depth of 3, since there is no subsequent sub-CU of the current CU, there is no need to enter the staging step of step 309.

[0039] In step 309, the best mode and the best reference frame index of the current CU are staged. According to another embodiment of the present invention, the best reference frame index is stored as RFa (forward list) and RFb (backward list).

[0040] In step 304, it is determined whether the best mode of the parent CU of the current CU is the Skip mode. According to an embodiment of the present invention, for a CU with a depth of 1, its parent CU is a CU with a depth of 0. For a CU with a depth of 2, its parent CU is a CU with a depth of 1. For a CU with a depth of 3, its parent CU is a CU with a depth of 2. If the best mode of the parent CU of the current CU is the Skip mode, step 305 is entered; otherwise, step 303 is entered, that is, all PU modes are executed for the current CU.

[0041] In step 305, the Skip mode and the Inter_2N×2N mode are executed for the current CU, and the best mode of the current CU is determined.

[0042] In step 306, it is determined whether the best mode of the current CU is the Skip mode. If so, step 308 is entered, so that the remaining PU modes are no longer executed for the current CU, and the staging step of step 309 is entered when the depth of the current CU is not 3. If not, step 307 is entered.

[0043] In step 307, the remaining PU modes are executed for the current CU, and the best mode and the best reference frame index of the current CU are determined. According to an embodiment of the present invention, for the current CU with a current depth less than 3, executing the remaining PU modes for the current CU includes successively executing Inter_N×2N, Inter_2N×N, and an asymmetric partition mode. For the current CU with a current depth of 3, executing the remaining PU modes for the current CU includes successively executing Inter_N×2N and Inter_2N×N for the current CU. According to an embodiment of the present invention, the best mode is selected according to the rate-distortion cost (RDC) of each PU mode. According to an embodiment of the present invention, reference Figure 4 is used to describe how to determine the best reference frame index of the current CU.

[0044] After completing step 307, step 308 is entered. When the depth of the current CU is not 3, the best mode and the best reference frame index of the current CU are temporarily stored.

[0045] In step 310, it is determined whether the depth of the current CU is 3. If not, step 311 is entered; if so, the process ends.

[0046] In step 311, the depth of the current CU is incremented by 1, and the process returns to step 302 to repeat this process for the next depth. According to an embodiment of the present invention, a counter can be used to count the depth.

[0047] It can be seen that by using the fast mode decision of the present invention, when using the Skip mode to encode the parent CU, it is early determined whether to skip the remaining modes according to the best mode of the current CU after the completion of the Inter_2N×2N mode. This method can significantly reduce the computational complexity of HEVC inter-frame prediction while maintaining good video coding quality.

[0048] Figure 4 The flowchart of method 400 for selecting the best reference frame by adopting a fast reference frame selection scheme based on CU correlation in inter-frame prediction according to an embodiment of the present invention is shown. This fast reference frame selection scheme can be used for Figure 3 303 (i.e., executing all PU modes for the current CU), 305 (i.e., executing the Skip mode and the Inter_2N×2N mode for the current CU), and 307 (i.e., executing the remaining PU modes for the current CU).

[0049] In the HEVC standard, there are four reference frames in the single reference frame list for each video frame under the low-delay coding profile. For the low-delay P frame (LDP), it is unidirectional prediction (usually forward prediction), that is, each frame has a reference frame list; for the low-delay B frame (LDB), it is a bidirectional prediction frame, and each reference frame has two reference frame lists. In inter-frame prediction, for each prediction unit (PU) of the current CU at the current depth, each reference frame in the reference frame list will be traversed, and motion estimation will be performed on the corresponding reference frame to select the best matching block and obtain the motion vector. It can be seen that this will increase the complexity of motion estimation for each PU.

[0050] In the fast reference frame selection scheme of the present invention, the selection of reference frames is mainly for CUs with a depth not equal to 0. Specifically, if the parent CU of the current CU with a depth not equal to 0 is predicted and encoded in the Skip mode as the best mode, then during the motion estimation process for each PU mode of the current CU, the best reference frame of the parent CU is directly selected as the best reference frame for the current mode, so that motion estimation is only performed on this frame to select the best motion vector, thereby skipping unnecessary motion estimation processes and effectively reducing the inter-frame prediction coding time.

[0051] According to an embodiment of the present invention, for P-frame coding, the current CU performs motion estimation among the reference frames stored in its single reference frame list. If the parent CU is in the Skip mode, the current CU performs motion estimation on the reference frame corresponding to the reference frame index RFa and obtains the motion vector; if the parent CU is not in the Skip mode, motion estimation is performed on all frames in the reference frame list, and finally the motion vector is selected.

[0052] According to another embodiment of the present invention, for B-frame coding, the current CU performs motion estimation among the reference frames stored in its bidirectional reference frame list. If the parent CU is in the Skip mode, the current CU performs motion estimation on the reference frames corresponding to the reference frame indexes RFa and RFb respectively and obtains the forward reference and backward reference motion vectors respectively; if the parent CU is not in the Skip mode, motion estimation is performed in the bidirectional reference list and the motion vector is obtained.

[0053] The following references Figure 4 to specifically describe the process of the fast reference frame selection scheme.

[0054] In step 401, it is determined whether the depth of the current CU is 0. If so, the process ends. Otherwise, step 402 is entered.

[0055] In step 402, it is determined whether the best mode of the parent CU of the current CU is the Skip mode. If so, step 403 is entered; otherwise, step 404 is entered.

[0056] In step 403, directly obtain the best reference frame index of the parent CU as the best reference frame index of the current mode, obtain the best reference frame according to the best reference frame index, and perform motion estimation on the best reference frame to finally obtain the best motion vector.

[0057] In step 404, sequentially traverse all the reference frames in the reference frame list of the current frame, perform motion estimation on all the reference frames to select the best motion vector for each reference frame, and finally select the best reference frame from the reference frame list of the current frame according to the rate-distortion cost (RDC). Generally speaking, in HEVC, the rate-distortion cost (RDC) is used for prediction decision-making. The smaller the RDC value, the more the current prediction conforms to the current prediction standard.

[0058] In step 405, temporarily store the best reference frame index and the best motion vector.

[0059] Figure 5 FIG. 500 is a block diagram of an exemplary computing device according to an embodiment of the present invention. The computing device is an example of a hardware device applicable to various aspects of the present invention.

[0060] Refer to Figure 5 , now a computing device 500 will be described. The computing device is an example of a hardware device applicable to various aspects of the present invention. The computing device 500 can be any machine configured to perform processing and / or computing, and can be, but is not limited to, a workstation, a server, a desktop computer, a laptop computer, a tablet computer, a personal digital assistant, a smart phone, an in-vehicle computer, or any combination thereof. The foregoing various methods / devices / servers / client devices can be implemented in whole or at least in part by the computing device 500 or a similar device or system.

[0061] The computing device 500 may include components that can be connected or communicate via one or more interfaces and a bus 502. For example, the computing device 500 may include a bus 502, one or more processors 504, one or more input devices 506, and one or more output devices 508. The one or more processors 504 can be any type of processor and may include, but are not limited to, one or more general-purpose processors and / or one or more dedicated processors (e.g., specialized processing chips). The input device 506 can be any type of device capable of inputting information into the computing device and may include, but are not limited to, a mouse, a keyboard, a touch screen, a microphone, and / or a remote controller. The output device 508 can be any type of device capable of presenting information and may include, but are not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The computing device 500 may also include a non-transitory storage device 510 or be connected to the non-transitory storage device. The non-transitory storage device can be any storage device that is non-transitory and capable of implementing data storage, and the non-transitory storage device may include, but are not limited to, a disk drive, an optical storage device, a solid-state memory, a floppy disk, a flexible disk, a hard disk, a magnetic tape, or any other magnetic medium, an optical disk, or any other optical medium, a ROM (read-only memory), a RAM (random access memory), a cache memory, and / or any storage chip or cartridge, and / or any other medium from which a computer can read data, instructions, and / or code. The non-transitory storage device 510 can be separated from the interface. The non-transitory storage device 510 may have data / instructions / code for implementing the above methods and steps. The computing device 500 may also include a communication device 512. The communication device 512 can be any type of device or system capable of enabling communication with internal devices and / or communication with a network and may include, but are not limited to, a modem, a network card, an infrared communication device, a wireless communication device, and / or a chipset, such as a Bluetooth device, an IEEE 1302.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or a similar device.

[0062] The bus 502 may include, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0063] The computing device 500 may further include a working memory 514, which can be any type of working memory capable of storing instructions and / or data beneficial for the operation of the processor 504 and may include, but are not limited to, a random access memory and / or a read-only storage device.

[0064] The software components may be located in the working memory 514, and these software components include, but are not limited to, an operating system 516, one or more applications 518, drivers, and / or other data and code. The instructions for implementing the above methods and steps of the present invention may be included in the one or more applications 518, and the above method 200 of the present invention may be implemented by reading and executing the instructions of the one or more applications 518 by the processor 504.

[0065] The innovations of the present invention can be described in the general context of computer-readable storage media. A computer-readable storage medium is any available tangible medium that can be accessed within a computing environment. By way of example and not limitation, computer-readable storage media include non-transitory storage devices 510, memories 514, and any combination of the foregoing.

[0066] The innovations of the present invention can be described in the general context of computer-executable instructions, such as those computer-executable instructions included in program modules that are executed in a computing system on a target physical or virtual processor. Generally speaking, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. As described in the various embodiments, the functions of these program modules can be combined or split among these program modules. The computer-executable instructions for each program module can be executed in a local or distributed computing system.

[0067] The terms "system" and "device" are used interchangeably herein. Unless the context clearly indicates otherwise, the terms do not imply any limitation on the type of computing system or computing device. Generally speaking, a computing system or computing device can be local or distributed, and can include any combination of dedicated hardware and / or general hardware having software that implements the functions described herein.

[0068] For purposes of presentation, this detailed description uses terms such as "determine," "execute," "acquire," etc. to describe computer operations in a computing system. These terms are high-level abstractions of operations performed by a computer and should not be confused with actions performed by a human. The actual computer operations corresponding to these terms vary depending on the implementation. As used herein when describing coding selections, the term "optimal" (as in "optimal mode," "optimal reference frame") indicates: an option that is preferred in terms of distortion cost, bitrate cost, or some combination of distortion cost and bitrate cost as compared to other options. Any available distortion metric can be used for the distortion cost. Any available bitrate metric can be used for the bitrate cost. Other factors (such as algorithm coding complexity, algorithm decoding complexity, resource usage, and / or latency) may also affect the decision as to which options are "optimal."

[0069] It should also be recognized that changes can be made according to specific requirements. For example, customized hardware can also be used, and / or specific components can be implemented in hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. In addition, connections with other computing devices, such as network input / output devices, etc., can be adopted. For example, part or all of the disclosed methods and devices can be implemented by using the logic and algorithms according to the present invention through programming hardware (such as programmable logic circuits including field programmable gate arrays (FPGAs) and / or programmable logic arrays (PLAs)) with an assembly language or a hardware programming language (such as VERILOG, VHDL, C++).

[0070] Although the aspects of the present invention have been described with reference to the accompanying drawings so far, the above methods, systems, and devices are only examples, and the scope of the present invention is not limited to these aspects, but is defined only by the appended claims and their equivalents. Various components can be omitted or replaced by equivalent components. In addition, the steps can be implemented in an order different from the order described in the present invention. Moreover, various components can be combined in various ways. It is also important that, as technology develops, many of the components described can be replaced by equivalent components that appear later.

Claims

1. A method for encoding a current frame in HEVC, comprising: (a) Reading in the current frame under a Low-Delay P-frame (LDP) or Low-Delay B-frame (LDB) or Random Access (RA) coding configuration; (b) Performing Coding Tree Unit (CTU) partitioning on the current frame; (c) If the current frame is not an I-frame, encoding the partitioned CTUs according to a fast mode decision scheme based on Coding Unit (CU) correlation and a fast reference frame selection scheme based on CU correlation; (d) If the currently encoded CTU is not the last CTU of the current frame, obtaining the next CTU and repeating steps (a) to (d); wherein the fast mode decision scheme based on CU correlation specifies that when performing inter-frame prediction mode encoding, if for the current CU with a depth not equal to 0, the best mode of its parent CU is the Skip mode, and after the current CU executes Skip and Inter_2N×2N, the best mode is the Skip mode, then determine that the best mode of the current CU is the Skip mode and terminate the execution of the remaining Prediction Unit (PU) modes; otherwise, execute all PU modes for the current CU; wherein the fast reference frame selection scheme based on CU correlation specifies that if the parent CU of the current CU with a depth not equal to 0 performs inter-frame prediction mode encoding with the Skip mode as the best mode, then during the motion estimation process for each PU mode of the current CU, directly select the best reference frame of the parent CU as the best reference frame for the current mode, so as to perform motion estimation only on this best reference frame to select the best motion vector.

2. The method according to claim 1, wherein The method further comprises: If the current frame is an I-frame, performing HEVC standard I-frame encoding on the current frame to perform intra-frame prediction encoding on the partitioned CTUs.

3. A method for performing inter-frame prediction based on CU correlation, comprising: (a) Obtaining a current CU with a current depth; (b) Judging whether the current depth of the current CU is 0; (c) If the current depth is 0, executing all PU modes for the current CU and determining and temporarily storing the best mode and the best reference frame index of the current CU; (d) If the current depth is not 0, judging whether the best mode of the parent CU of the current CU is the Skip mode; (e) If the best mode of the parent CU of the current CU is not the Skip mode, executing all PU modes for the current CU, determining the best mode and the best reference frame index of the current CU, and temporarily storing the best mode and the best reference frame index of the current CU when the current depth is not 3; (f) If the best mode of the parent CU of the current CU is the Skip mode, executing the Skip mode and the Inter_2N×2N mode for the current CU and determining the best mode of the current CU; (g) If the best mode of the current CU is the Skip mode and the current depth is not 3, temporarily storing the best mode and the best reference frame index of the current CU; (h) If the best mode of the current CU is not the Skip mode, then perform the remaining PU modes on the current CU, determine the best mode and the best reference frame index of the current CU, and cache the best mode and the best reference frame index of the current CU when the depth is not 3; (i) If the current depth is not 3, repeat the above steps (a)-(h) for the next depth.

4. The method according to claim 3, wherein (j) Performing all the PU modes on the current CU further includes: For the current CU with a current depth less than 3, sequentially perform Skip / Merge, Inter_2N×2N, Inter_N×2N, Inter_2N×N, and asymmetric partitioning modes on the current CU; For the current CU with a current depth of 3, sequentially perform Skip / Merge, Inter_2N×2N, Inter_N×2N, Inter_2N×N on the current CU.

5. The method according to claim 3, wherein (k) Performing the remaining PU modes on the current CU further includes: For the current CU with a current depth less than 3, sequentially perform Inter_N×2N, Inter_2N×N, and asymmetric partitioning modes on the current CU; For the current CU with a current depth of 3, sequentially perform Inter_N×2N and Inter_2N×N on the current CU.

6. The method according to claim 3, wherein (l) Determining the best mode of the current CU further includes selecting the best mode based on the rate-distortion cost (RDC) of each PU mode.

7. The method according to claim 3, wherein (m) The best reference frame index is stored as a forward list and a backward list.

8. The method according to claim 3, wherein (n) Performing all the PU modes on the current CU or performing the remaining PU modes on the current CU further includes: If the current depth of the current CU is not 0 and the best mode of the parent CU of the current CU is the Skip mode, then directly select the best reference frame of the parent CU as the best reference frame of the current mode during the motion estimation process for each PU mode of the current CU.

9. A computer-readable storage medium storing processor-executable instructions that, when executed by a processor, are used to perform the method according to any one of claims 1-8.

10. A system for performing inter prediction based on CU correlation, comprising: A processor; A memory storing instructions that, when executed by the processor, can perform the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Low-complexity method for selecting HEVC coding multiple reference frames

    CN103813166A

  • HEVC interframe coding quick mode selection method

    CN105141954A