Encoding device, decoding device, and non-transitory machine-readable medium for coding video data
By determining motion vectors from specific positional relationships and incorporating relocated candidates, the method enhances the construction of candidate lists, addressing suboptimal prediction in complex video content and improving coding efficiency and compression.
Patent Information
- Application Number
- PCT/JP2024/046346
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-10
AI Technical Summary
Existing video coding methods struggle to construct optimal candidate lists for block prediction, particularly in dynamic or high-resolution video content, leading to increased residual errors and bitrate requirements due to suboptimal handling of complex motion patterns.
The method involves determining motion vectors for candidate lists based on various positional relationships of reference blocks, including top-left, top-right, center, and bottom-right positions, and incorporating relocated candidates to enhance prediction accuracy.
This approach improves the efficiency and compression performance of video coding by better capturing intricate motion patterns, reducing residual errors and bitrate demands.
Smart Images

Figure JP2024046346_10072025_PF_FP_ABST
Abstract
Description
ENCODING DEVICE, DECODING DEVICE, AND NON-TRANSITORY MACHINE-READABLE MEDIUM FOR CODING VIDEO DATA
[0001] The present disclosure is generally related to video coding and, more specifically, to techniques for constructing candidate lists for block prediction.
[0002] The present disclosure claims the benefit of and priority to U.S. Provisional Patent Application Serial No. 63 / 618,083, filed on January 5, 2024, entitled “PROPOSED INTER AND INTRA PREDICTION METHOD,” the content of which is hereby incorporated herein fully by reference in its entirety into the present disclosure for all purposes.
[0003] Prediction is a fundamental component of video coding, enabling efficient compression by reducing spatial and temporal redundancies in video sequences. It is categorized into two primary methods: intra prediction and inter prediction. Intra prediction utilizes spatial redundancies within a single frame by predicting target blocks based on neighboring blocks. Inter prediction, on the other hand, leverages temporal redundancies by predicting a target block in the current frame using reference blocks from other frames.
[0004] A key aspect of inter prediction is the construction of a candidate list, which consists of motion vectors that represent possible motion relationships between reference and target blocks. The candidate list serves as the basis for selecting the motion vector that provides the most accurate prediction for each block. However, the quality and completeness of the candidate list can significantly affect the efficiency of inter prediction. Suboptimal candidate lists may fail to capture complex motion patterns, resulting in higher residual errors and increased bitrate requirements. This challenge becomes more pronounced in dynamic video content or high-resolution sequences, where accurately predicting motion is particularly difficult.
[0005] Over the years, various methods have been proposed to construct candidate lists, often focusing on common patterns or simplifying assumptions about motion. While these approaches have achieved notable improvements, they may still fall short in scenarios with unconventional or intricate motion characteristics. There remains an opportunity to enhance the construction of candidate lists by incorporating strategies that better address such complexities, thereby further improving the coding efficiency and compression performance in video coding systems.Summery of Invention
[0006] The present disclosure is directed to an electronic device and a non-transitory machine-readable medium for encoding / decoding video data, aimed at the construction of candidate lists for block prediction, thereby enhancing prediction accuracy.
[0007] In a first aspect of the present disclosure, an electronic device for decoding video data is provided. The electronic device includes at least one processor, and at least one non-transitory computer-readable medium coupled to the at least one processor and storing one or more computer-executable instructions. The one or more computer-executable instructions, when executed by the at least one processor, cause the electronic device to: receive the video data; determine a block unit from an image frame according to the video data; determine a first motion vector that indicates a first reference block, among a plurality of reference blocks, in a reference frame relative to the block unit, based on at least one block vector of at least one of the plurality of reference blocks; include the first motion vector into a candidate list established for the block unit; determine a prediction of the block unit based on the candidate list; and reconstruct the block unit based on the prediction.
[0008] In an implementation of the first aspect, the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine a second reference block based on the at least one block vector of at least one other reference block, different from the first reference block and the second reference block, among the plurality of reference blocks. The first reference block is indicated by a second motion vector originated from the second reference block.
[0009] In another implementation of the first aspect, the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine the second motion vector based on at least one of a top-left position, a top-right position, a center position, a bottom-left position, and a bottom-right position of the second reference block.
[0010] In another implementation of the first aspect, determining the second motion vector includes: checking whether a third motion vector is stored corresponding to the center position of the second reference block; and determining the second motion vector based on the third motion vector when the third motion vector is stored corresponding to the center position of the second reference block.
[0011] In another implementation of the first aspect, the candidate list includes an inter merge candidate list, and determining the prediction of the block unit includes: selecting at least one motion vector from the inter merge candidate list; and determining, based on the at least one motion vector as selected, the prediction of the block unit.
[0012] In a second aspect of the present disclosure, an electronic device for encoding video data is provided. The electronic device includes at least one processor, and at least one non-transitory computer-readable medium coupled to the at least one processor and storing one or more computer-executable instructions. The one or more computer-executable instructions, when executed by the at least one processor, cause the electronic device to: receive the video data; determine a block unit from an image frame according to the video data; determine a first motion vector that indicates a first reference block, among a plurality of reference blocks, in a reference frame relative to the block unit, based on at least one block vector of at least one of the plurality of reference blocks; include the first motion vector into a candidate list established for the block unit; determine a prediction of the block unit based on the candidate list; and reconstruct the block unit based on the prediction.
[0013] In an implementation of the second aspect, the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine a second reference block based on the at least one block vector of at least one other reference block, different from the first reference block and the second reference block, among the plurality of reference blocks. The first reference block is indicated by a second motion vector originated from the second reference block.
[0014] In another implementation of the second aspect, the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine the second motion vector based on at least one of a top-left position, a top-right position, a center position, a bottom-left position, and a bottom-right position of the second reference block.
[0015] In another implementation of the second aspect, determining the second motion vector includes: checking whether a third motion vector is stored corresponding to the center position of the second reference block; and determining the second motion vector based on the third motion vector when the third motion vector is stored corresponding to the center position of the second reference block.
[0016] In another implementation of the second aspect, the candidate list includes an inter merge candidate list, and determining the prediction of the block unit includes: selecting at least one motion vector from the inter merge candidate list; and determining, based on the at least one motion vector as selected, the prediction of the block unit.
[0017] In a third aspect of the present disclosure, a non-transitory machine-readable medium of an electronic device storing one or more computer-executable instructions for decoding video data is provided. The one or more computer-executable instructions, when executed by at least one processor of the electronic device, causing the electronic device to: receive the video data; determine a block unit from an image frame according to the video data; determine a first motion vector that indicates a first reference block, among a plurality of reference blocks, in a reference frame relative to the block unit, based on at least one block vector of at least one of the plurality of reference blocks; include the first motion vector into a candidate list established for the block unit; determine a prediction of the block unit based on the candidate list; and reconstruct the block unit based on the prediction.
[0018] In an implementation of the third aspect, the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine a second reference block based on the at least one block vector of at least one other reference block, different from the first reference block and the second reference block, among the plurality of reference blocks. The first reference block is indicated by a second motion vector originated from the second reference block.
[0019] In another implementation of the third aspect, the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine the second motion vector based on at least one of a top-left position, a top-right position, a center position, a bottom-left position, and a bottom-right position of the second reference block.
[0020] In another implementation of the third aspect, determining the second motion vector includes: checking whether a third motion vector is stored corresponding to the center position of the second reference block; and determining the second motion vector based on the third motion vector when the third motion vector is stored corresponding to the center position of the second reference block.
[0021] In another implementation of the third aspect, the candidate list includes an inter merge candidate list, and determining the prediction of the block unit includes: selecting at least one motion vector from the inter merge candidate list; and determining, based on the at least one motion vector as selected, the prediction of the block unit.
[0022] Aspects of the present disclosure are best understood from the following detailed disclosure and the corresponding figures. Various features are not drawn to scale and dimensions of various features may be arbitrarily increased or reduced for clarity of discussion.
[0023] FIG. 1 is a block diagram illustrating a system having a first electronic device and a second electronic device for encoding and decoding video data, in accordance with one or more example implementations of this disclosure.
[0024] FIG. 2 is a block diagram illustrating a decoder module of the second electronic device illustrated in FIG. 1, in accordance with one or more example implementations of this disclosure.
[0025] FIG. 3 is a flowchart illustrating a method / process for decoding and / or encoding video data by an electronic device, in accordance with one or more example implementations of this disclosure.
[0026] FIG. 4 is a diagram illustrating multiple neighbor coded blocks of a block unit, in accordance with one or more example implementations of this disclosure.
[0027] FIG. 5 is a diagram illustrating a co-located coding unit, in accordance with one or more example implementations of this disclosure.
[0028] FIG. 6 is a diagram illustrating a derivation of a scaled motion vector, in accordance with one or more example implementations of this disclosure.
[0029] FIG. 7 is a diagram illustrating neighboring regions of a block unit, in accordance with one or more example implementations of this disclosure.
[0030] FIG. 8 is a diagram illustrating multiple non-adjacent coded blocks of a block unit, in accordance with one or more example implementations of this disclosure.
[0031] FIG. 9 is a diagram illustrating a template matching method, in accordance with one or more example implementations of this disclosure.
[0032] FIG. 10 is a diagram illustrating a determination of a motion vector for determining a relocated candidate, in accordance with one or more example implementations of this disclosure.
[0033] FIG. 11 is a diagram illustrating a determination of a motion vector for determining a relocated candidate, in accordance with one or more example implementations of this disclosure.
[0034] FIG. 12 is a diagram illustrating a determination of a block vector based on a block, in accordance with one or more example implementations of this disclosure.
[0035] FIG. 13 is a diagram illustrating a determination of a motion vector based on a block, in accordance with one or more example implementations of this disclosure.
[0036] FIG. 14 is a diagram illustrating a determination of a relocated candidate, in accordance with one or more example implementations of this disclosure.
[0037] FIG. 15 is a diagram illustrating a determination of a relocated candidate, in accordance with one or more example implementations of this disclosure.
[0038] FIG. 16 is a flowchart illustrating a method / process for constructing a candidate list, in accordance with one or more example implementations of this disclosure.
[0039] FIG. 17 is a diagram illustrating a template-based intra mode derivation mode, in accordance with one or more example implementations of this disclosure.
[0040] FIG. 18 is a diagram illustrating a decoder-side intra mode derivation mode, in accordance with one or more example implementations of this disclosure.
[0041] FIGS. 19A and 19B are diagrams illustrating a combined inter and intra prediction, in accordance with one or more example implementations of this disclosure.
[0042] FIG. 20 is a diagram illustrating a neighboring region, in accordance with one or more example implementations of this disclosure.
[0043] FIG. 21 is a diagram illustrating a regression-based prediction of a neighboring region, in accordance with one or more example implementations of this disclosure.
[0044] FIG. 22 is a block diagram illustrating an encoder module of the first electronic device illustrated in FIG. 1, in accordance with one or more example implementations of this disclosure.
[0045] The following disclosure contains specific information pertaining to implementations in the present disclosure. The figures and the corresponding detailed disclosure are directed to example implementations. However, the present disclosure is not limited to these example implementations. Other variations and implementations of the present disclosure will occur to those skilled in the art.
[0046] Unless noted otherwise, like or corresponding elements among the figures may be indicated by like or corresponding reference designators. The figures and illustrations in the present disclosure are generally not to scale and are not intended to correspond to actual relative dimensions.
[0047] For the purposes of consistency and ease of understanding, features are identified (although, in some examples, not illustrated) by reference designators in the exemplary figures. However, the features in different implementations may differ in other respects and shall not be narrowly confined to what is illustrated in the figures.
[0048] The present disclosure uses the phrases “in one implementation,” or “in some implementations,” which may refer to one or more of the same or different implementations. The term “coupled” is defined as connected, whether directly or indirectly through intervening components, and is not necessarily limited to physical connections. The term “comprising” means “including, but not necessarily limited to” and specifically indicates open-ended inclusion or membership in the so-described combination, group, series, and the equivalent.
[0049] For purposes of explanation and non-limitation, specific details, such as functional entities, techniques, protocols, and standards, are set forth for providing an understanding of the disclosed technology. Detailed disclosure of well-known methods, technologies, systems, and architectures are omitted so as not to obscure the present disclosure with unnecessary details.
[0050] Persons skilled in the art will recognize that any disclosed coding function(s) or algorithm(s) described in the present disclosure may be implemented by hardware, software, or a combination of software and hardware. Disclosed functions may correspond to modules that are software, hardware, firmware, or any combination thereof.
[0051] A software implementation may include a program having one or more computer-executable instructions stored on a computer-readable medium, such as memory or other types of storage devices. For example, one or more microprocessors or general-purpose computers with communication processing capability may be programmed with computer-executable instructions and perform the disclosed function(s) or algorithm(s).
[0052] The microprocessors or general-purpose computers may be formed of application-specific integrated circuits (ASICs), programmable logic arrays, and / or one or more digital signal processors (DSPs). Although some of the disclosed implementations are oriented to software installed and executing on computer hardware, alternative implementations implemented as firmware, as hardware, or as a combination of hardware and software are well within the scope of the present disclosure. The computer-readable medium includes, but is not limited to, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, compact disc read-only memory (CD ROM), magnetic cassettes, magnetic tape, magnetic disk storage, or any other equivalent medium capable of storing computer-executable instructions. The computer-readable medium may be a non-transitory computer-readable medium.
[0053] FIG. 1 is a block diagram illustrating a system 100 having a first electronic device and a second electronic device for encoding and decoding video data, in accordance with one or more example implementations of this disclosure.
[0054] The system 100 includes a first electronic device 110, a second electronic device 120, and a communication medium 130.
[0055] The first electronic device 110 may be a source device including any device configured to encode video data and transmit the encoded video data to the communication medium 130. The second electronic device 120 may be a destination device including any device configured to receive encoded video data via the communication medium 130 and decode the encoded video data.
[0056] The first electronic device 110 may communicate via wire, or wirelessly, with the second electronic device 120 via the communication medium 130. The first electronic device 110 may include a source module 112, an encoder module 114, and a first interface 116, among other components. The second electronic device 120 may include a display module 122, a decoder module 124, and a second interface 126, among other components. The first electronic device 110 may be a video encoder and the second electronic device 120 may be a video decoder.
[0057] The first electronic device 110 and / or the second electronic device 120 may be a mobile phone, a tablet, a desktop, a notebook, or other electronic devices. FIG. 1 illustrates one example of the first electronic device 110 and the second electronic device 120. The first electronic device 110 and second electronic device 120 may include greater or fewer components than illustrated or have a different configuration of the various illustrated components.
[0058] The source module 112 may include a video capture device to capture new video, a video archive to store previously captured video, and / or a video feed interface to receive the video from a video content provider. The source module 112 may generate computer graphics-based data, as the source video, or may generate a combination of live video, archived video, and computer-generated video, as the source video. The video capture device may include a charge-coupled device (CCD) image sensor, a complementary metal-oxide-semiconductor (CMOS) image sensor, or a camera.
[0059] The encoder module 114 and the decoder module 124 may each be implemented as any one of a variety of suitable encoder / decoder circuitry, such as one or more microprocessors, a central processing unit (CPU), a graphics processing unit (GPU), a system-on-a-chip (SoC), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combinations thereof. When implemented partially in software, a device may store the program having computer-executable instructions for the software in a suitable, non-transitory computer-readable medium and execute the stored computer-executable instructions using one or more processors to perform the disclosed methods. Each of the encoder module 114 and the decoder module 124 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in a device.
[0060] The first interface 116 and the second interface 126 may utilize customized protocols or follow existing standards or de facto standards including, but not limited to, Ethernet, IEEE 802.11 or IEEE 802.15 series, wireless USB, or telecommunication standards including, but not limited to, Global System for Mobile Communications (GSM), Code-Division Multiple Access 2000 (CDMA2000), Time Division Synchronous Code Division Multiple Access (TD-SCDMA), Worldwide Interoperability for Microwave Access (WiMAX), Third Generation Partnership Project Long-Term Evolution (3GPP-LTE), or Time-Division LTE (TD-LTE). The first interface 116 and the second interface 126 may each include any device configured to transmit a compliant video bitstream via the communication medium 130 and to receive the compliant video bitstream via the communication medium 130.
[0061] The first interface 116 and the second interface 126 may include a computer system interface that enables a compliant video bitstream to be stored on a storage device or to be received from the storage device. For example, the first interface 116 and the second interface 126 may include a chipset supporting Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, proprietary bus protocols, Universal Serial Bus (USB) protocols, Inter-Integrated Circuit (I2C) protocols, or any other logical and physical structure(s) that may be used to interconnect peer devices.
[0062] The display module 122 may include a display using liquid crystal display (LCD) technology, plasma display technology, organic light-emitting diode (OLED) display technology, or light-emitting polymer display (LPD) technology, with other display technologies used in some other implementations. The display module 122 may include a High-Definition display or an Ultra-High-Definition display.
[0063] FIG. 2 is a block diagram illustrating a decoder module 124 of the second electronic device 120 illustrated in FIG. 1, in accordance with one or more example implementations of this disclosure. The decoder module 124 may include an entropy decoder (e.g., an entropy decoding unit 2241), a prediction processor (e.g., a prediction processing unit 2242), an inverse quantization / inverse transform processor (e.g., an inverse quantization / inverse transform unit 2243), a summer (e.g., a summer 2244), a filter (e.g., a filtering unit 2245), and a decoded picture buffer (e.g., a decoded picture buffer 2246). The prediction processing unit 2242 further may include an intra prediction processor (e.g., an intra prediction unit 22421) and an inter prediction processor (e.g., an inter prediction unit 22422). The decoder module 124 receives a bitstream, decodes the bitstream, and outputs a decoded video.
[0064] The entropy decoding unit 2241 may receive the bitstream including multiple syntax elements from the second interface 126, as shown in FIG. 1, and perform a parsing operation on the bitstream to extract syntax elements from the bitstream. As part of the parsing operation, the entropy decoding unit 2241 may entropy decode the bitstream to generate quantized transform coefficients, quantization parameters, transform data, motion vectors, intra modes, partition information, and / or other syntax information.
[0065] The entropy decoding unit 2241 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique to generate the quantized transform coefficients. The entropy decoding unit 2241 may provide the quantized transform coefficients, the quantization parameters, and the transform data to the inverse quantization / inverse transform unit 2243 and provide the motion vectors, the intra modes, the partition information, and other syntax information to the prediction processing unit 2242.
[0066] The prediction processing unit 2242 may receive syntax elements, such as motion vectors, intra modes, partition information, and other syntax information, from the entropy decoding unit 2241. The prediction processing unit 2242 may receive the syntax elements including the partition information and divide image frames according to the partition information.
[0067] Each of the image frames may be divided into at least one image block according to the partition information. The at least one image block may include a luminance block for reconstructing multiple luminance samples and at least one chrominance block for reconstructing multiple chrominance samples. The luminance block and the at least one chrominance block may be further divided to generate macroblocks, coding tree units (CTUs), coding blocks (CBs), sub-divisions thereof, and / or other equivalent coding units.
[0068] During the decoding process, the prediction processing unit 2242 may receive predicted data including the intra mode or the motion vector for a current image block of a specific one of the image frames. The current image block may be the luminance block or one of the chrominance blocks in the specific image frame.
[0069] The intra prediction unit 22421 may perform intra-predictive coding of a current block unit relative to one or more neighboring blocks in the same frame as the current block unit based on syntax elements related to the intra mode in order to generate a predicted block. The intra mode may specify the location of reference samples selected from the neighboring blocks within the current frame. The intra prediction unit 22421 may reconstruct multiple chroma components of the current block unit based on multiple luma components of the current block unit when the multiple chroma components is reconstructed by the prediction processing unit 2242.
[0070] The intra prediction unit 22421 may reconstruct multiple chroma components of the current block unit based on the multiple luma components of the current block unit when the multiple luma components of the current block unit is reconstructed by the prediction processing unit 2242.
[0071] The inter prediction unit 22422 may perform inter-predictive coding of the current block unit relative to one or more blocks in one or more reference image blocks based on syntax elements related to the motion vector in order to generate the predicted block.
[0072] The inter prediction unit 22422 may receive the reference image block stored in the decoded picture buffer 2246 and reconstruct the current block unit based on the received reference image blocks.
[0073] The inverse quantization / inverse transform unit 2243 may apply inverse quantization and inverse transformation to reconstruct the residual block in the pixel domain. The inverse quantization / inverse transform unit 2243 may apply inverse quantization to the residual quantized transform coefficient to generate a residual transform coefficient and then apply inverse transformation to the residual transform coefficient to generate the residual block in the pixel domain.
[0074] The inverse transformation may be inversely applied by the transformation process, such as a discrete cosine transform (DCT), a discrete sine transform (DST), an adaptive multiple transform (AMT), a mode-dependent non-separable secondary transform (MDNSST), a Hypercube-Givens transform (HyGT), a signal-dependent transform, a Karhunen-Loeve transform (KLT), a wavelet transform, an integer transform, a sub-band transform, or a conceptually similar transform. The inverse transformation may convert the residual information from a transform domain, such as a frequency domain, back to the pixel domain, etc. The degree of inverse quantization may be modified by adjusting a quantization parameter.
[0075] The summer 2244 may add the reconstructed residual block to the predicted block provided by the prediction processing unit 2242 to produce a reconstructed block.
[0076] The filtering unit 2245 may include a deblocking filter, a sample adaptive offset (SAO) filter, a bilateral filter, and / or an adaptive loop filter (ALF) to remove the blocking artifacts from the reconstructed block. Additional filters (in loop or post loop) may also be used in addition to the deblocking filter, the SAO filter, the bilateral filter, and the ALF. Such filters (which are not explicitly illustrated for the brevity of description) may filter the output of the summer 2244. The filtering unit 2245 may output the decoded video to the display module 122 or other video receiving units after the filtering unit 2245 performs the filtering process for the reconstructed blocks of the specific image frame.
[0077] The decoded picture buffer 2246 may be a reference picture memory that stores the reference block to be used by the prediction processing unit 2242 in decoding the bitstream (e.g., in inter-coding modes). The decoded picture buffer 2246 may be formed by any one of a variety of memory devices, such as a dynamic random-access memory (DRAM), including synchronous DRAM (SDRAM), magneto-resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer 2246 may be on-chip along with other components of the decoder module 124 or may be off-chip relative to those components.
[0078] FIG. 3 is a flowchart illustrating a method / process 300 for decoding and / or encoding video data by an electronic device, in accordance with one or more example implementations of this disclosure. The method / process 300 is an example implementation, as there may be a variety of methods of decoding the video data.
[0079] The method / process 300 may be performed by an electronic device using the configurations illustrated in FIGS. 1 and 2, where various elements of these figures may be referenced to describe the method / process 300. Each block illustrated in FIG. 3 may represent one or more processes, methods, or subroutines performed by an electronic device.
[0080] The order in which the blocks appear in FIG. 3 is for illustration only, and may not be construed to limit the scope of the present disclosure, thus may be different from what is illustrated. Additional blocks may be added or fewer blocks may be utilized without departing from the scope of the present disclosure.
[0081] At block 310, the method / process 300 may start by receiving (e.g., by the decoder module 124) the video data. The video data received by the decoder module 124 may include a bitstream provided by the encoder module 114, which may include information of multiple image frames.
[0082] With reference to FIG. 1 and FIG. 2, the second electronic device 120 may receive the bitstream from an encoder, such as the first electronic device 110, or from other video providers, via the second interface 126. The second interface 126 may provide the bitstream to the decoder module 124.
[0083] The entropy decoding unit 2241 may decode the bitstream to determine multiple prediction indications and multiple partitioning indications for multiple video images. Then, the decoder module 124 may further reconstruct the multiple video images based on the prediction indications and the partitioning indications. The prediction indications and the partitioning indications may include multiple flags and multiple indices.
[0084] At block 320, the method / process 300 may determine (e.g., by the decoder module 124), a block unit from an image frame according to the video data. Specifically, the video data may include the bitstream received from the encoder, and a block unit may be determined from an image frame according to the bitstream.
[0085] With reference to FIG. 1 and FIG. 2, the decoder module 124 may determine the image frames based on the bitstream and may divide each image frame to determine the block units according to the partition indications in the bitstream. For example, the decoder module 124 may divide the image frames to generate multiple CTUs, and further divide one of the CTUs to determine the block units according to the partition indications based on any video coding standard.
[0086] In some implementations, the block unit may be a current block. For example, the current block may include at least one of a coding unit, a prediction unit, a macroblock, a luma block, and a chrome block.
[0087] At block 330, the method / process 300 may determine (e.g., by the decoder module 124), a prediction of the block unit.
[0088] INTER PREDICTION
[0089] In some implementations, the (inter) prediction of the block unit may be determined based on an inter prediction process using an inter prediction mode such as an (inter) merge mode.
[0090] In some implementations, the merge mode may allow a block to inherit motion information, such as motion vectors and reference indices, directly from neighboring blocks without the need to explicitly encode these parameters.
[0091] In some implementations, the merge mode may be achieved based on an (inter) merge candidate list, and the merge candidate list may include multiple (e.g., 15) (inter) merge candidates.
[0092] In some implementations, the merge candidate list may be established by sequentially including at least one adjacent candidate, at least one temporal candidate, at least one non-adjacent candidate, at least one history-based motion vector prediction (HMVP) candidate and at least one pair-wise average candidate, until the merge candidate list is full.
[0093] In some implementations, the at least one adjacent candidate may include at least one piece of neighbor motion information, and each piece of the neighbor motion information may include the motion vector of one of multiple neighbor coded blocks (e.g., adjacent coded blocks).
[0094] FIG. 4 is a diagram illustrating multiple neighbor coded blocks of a block unit 40, in accordance with one or more example implementations of this disclosure.
[0095] Referring FIG. 4, the neighbor coded blocks of the block unit 40 may include an above block 41, a left block 42, an above-right block 43, a bottom-left block 44, and an above-left block 45. The position of the top-left corner of the block unit 40 may be (x, y), the width of the block unit 40 may be W and the height of the block unit 40 is H, where W and H are positive integers. The above block 41 may be a block including a sample located at (x+W-1, y-1), the left block 42 may be a block including a sample located at (x-1, y+H-1), the above-right block 43 may be a block including a sample located at (x+W, y-1), the bottom-left block 44 may be a block including a sample located at (x-1, y+H), and the above-left block 45 may be a block including a sample located at (x-1, y-1).
[0096] In some implementations, the at least one temporal candidate may include a scaled motion vector derived based on at least one co-located coding unit (CU).
[0097] Specifically, the co-located CU may be a CU including a sample located at a specific position in a co-located picture. The co-located picture may be indicated by a reference list flag and a reference picture index. The reference list flag and the reference picture index may be signaled in the slice header. The reference list flag may indicate one of a first reference picture list (RPL) or a second RPL, and the reference picture index may indicate an index of the co-located picture in the first or second RPL indicated by the reference list flag.
[0098] FIG. 5 is a diagram illustrating a co-located coding unit, in accordance with one or more example implementations of this disclosure.
[0099] Referring to FIG. 5, the position of the top-left corner of the block unit 40 may be (x, y), the width of the block unit 40 may be W and the height of the block unit 40 is H, where W and H are positive integers. The specific position may be equal to (x+W, y+H). If the CU located at (x+W, y+H) is not available, the specific position may be equal to (x+(W / 2), y+(H / 2)). It should be noted that, a CU not available may indicate that the CU may not be coded by an intra prediction mode or the CU is outside of the current CTU row.
[0100] FIG. 6 is a diagram illustrating a derivation of a scaled motion vector, in accordance with one or more example implementations of this disclosure.
[0101] In some implementations, the scaled motion vector may be derived based on only one co-located CU. The motion vector of the co-located CU may be used to derive the scaled motion vector. For example, referring to FIG. 6, the scaled motion vector MVscalemay be derived based on the motion vector MVco-locatedof the co-located CU and multiple picture order count (POC) distances, using the following equation: MVscale= (tb / td) * MVco-located, where tb may be a POC distance between the current picture 62 and the current reference picture 63, and td may be a POC distance between the co-located picture 61 and the co-located reference picture 64.
[0102] In some implementations, the scaled motion vector may be derived based on two co-located CUs. The motion vectors of the two co-located CUs may be used to derive the scaled motion vector. The two co-located CU may be located at two co-located pictures, respectively. The two co-located pictures may be pictures that have minimum POC distances from the current picture. For example, in a case that the POC of the current picture is N, the POCs of the two co-located pictures may be N-1 and N+1, where N is an integer. One of the motion vectors of the two co-located CUs may be selected to derive the scaled motion vector based on a template cost. Specifically, each motion vector of the co-located CUs may have a template cost, which may be derived based on a neighboring region (which will be described below). More specifically, each motion vector of the co-located CUs may have a template cost, which may be derived based on a difference between a reconstructed neighboring region and a predicted neighboring region. More specifically, the reconstructed neighboring region may be the neighboring region with reconstructed samples, the predicted neighboring region may be the same neighboring region with predicted samples, and the difference between the reconstructed neighboring region and the predicted neighboring region may be calculated based on the reconstructed samples and the predicted samples. The predicted samples may be derived based on each motion vector of the co-located CUs, and the motion vector that yields the smallest template cost may be selected to derive the scaled motion vector.
[0103] FIG. 7 is a diagram illustrating neighboring regions of a block unit 40, in accordance with one or more example implementations of this disclosure.
[0104] Referring to FIG. 7, neighboring regions of the block unit 40 may include at least one of an above region 71, an above-left region 72, and a left region 73. In some implementations, the neighboring region of the block unit 40 may include all of the above region 71, the above-left region 72, and the left region 73. In some implementations, the neighboring region of the block unit 40 may include the above region 71 and the left region 73. In some implementations, the neighboring region of the block unit 40 may include the left region 73 and the above-left region 72. In some implementations, the neighboring region of the block unit 40 may include the left region 73. In some implementations, the neighboring region of the block unit 40 may include the above region 71 and the above-left region 72. In some implementations, the neighboring region of the block unit 40 may include the above region 71.
[0105] In some implementations, the at least one non-adjacent candidate may include at least one piece of non-adjacent motion information, and each piece of the non-adjacent motion information may include the motion vector of one of multiple non-adjacent coded blocks.
[0106] FIG. 8 is a diagram illustrating multiple non-adjacent coded blocks of a block unit 40, in accordance with one or more example implementations of this disclosure.
[0107] Referring to FIG. 8, blocks 801 to 805 may be the neighbor coded blocks of the block unit 40, and blocks 806 to 823 may be the non-adjacent coded blocks of the block unit 40. The distances between the non-adjacent coded blocks 806 to 823 and the block unit 40 may be determined based on the width and height of block unit 40.
[0108] In some implementations, the at least one HMVP candidate may be selected from a HMVP table. The HMVP table may include multiple pieces of motion information from previously coded blocks, each containing a motion vector. The number of motion information entries in the HMVP table may be, for example, five. The HMVP table may be reset upon encountering a new CTU row. When inserting new motion information into the HMVP table, a constrained first-in-first-out (FIFO) rule may be applied. Before insertion, a redundancy check may be performed to determine whether identical motion information already exists in the HMVP table. If an identical entry is found, it may be removed, all subsequent entries in the HMVP table may be shifted forward, and the identical motion information may be reinserted at the last entry. In some implementations, the last two entries in the HMVP table may then be selected as two HMVP candidates.
[0109] In some implementations, the at least one pair-wise average candidate may be generated by averaging predefined pair(s) of candidates from the merge candidate list, e.g., using the first two merge candidates. For example, the first merge candidate in the merge candidate list may be denoted as p0Cand, and the second merge candidate in the merge candidate list may be denoted as p1Cand. The averaged motion vector(s) may be calculated based on the availability of the motion vectors of the p0Cand and the p1Cand separately for each RPL. If both motion vectors are available in a given RPL, the two motion vectors are averaged, even if they point to different reference pictures, and the reference picture of the averaged motion vector is set to that of p0Cand. If only one motion vector is available, the available motion vector may be used directly. If no motion vector is available, the pair-wise average candidate may be deemed invalid.
[0110] FIG. 9 is a diagram illustrating a template matching method, in accordance with one or more example implementations of this disclosure.
[0111] In some implementations, the merge candidates may be refined by a template matching (TM) method. Specifically, the TM method may refine the motion vector of the merge candidates by finding the closest match between a neighboring region of the current block (e.g., block unit 40) and a neighboring region of the reference block. Referring to FIG. 9, a better motion vector may be, for example, searched around the initial motion vector (e.g., within the merge candidates) of the block unit 40 within a [-8, +8]-pel search range. The closest match between the neighboring region 91, 92 of the block unit 40 and the neighboring region 93, 94 of the reference block 90 may indicate that the smallest difference between the neighboring region 91, 92 of the block unit 40 and the neighboring region 93, 94 of the reference block 90.
[0112] In some implementations, the merge mode may be one of the Intra Block Copy (IBC) merge candidates in an IBC merge candidate list. The IBC merge candidate list may include multiple IBC merge candidates, which may include block vectors of adjacent candidates, non-adjacent candidates, HMVP candidates, and pairwise candidates, as described above.
[0113] In some implementations, when constructing / establishing an candidate list (e.g., an inter advanced motion vector prediction (AMVP) candidate list or an inter merge candidate list) for an inter prediction mode (e.g., an AMVP mode or a merge mode), a relocated candidate method may be applied to improve coding efficiency.
[0114] It should be noted that, neighboring block(s) of a block unit may include multiple adjacent block(s) and non-adjacent block(s) of the block unit, and neighboring candidate(s) in the candidate list (e.g., the inter AMVP candidate list or the inter merge candidate list) may include multiple adjacent candidate(s) and non-adjacent candidate(s).
[0115] In some implementations, by using the relocated candidate method, a relocated candidate may be determined and added into the candidate list when constructing the candidate list. The relocated candidate may include, for example, motion information. In some implementations, the relocated candidate (e.g., the motion information) may be determined based on at least one block vector of at least one neighboring block (e.g., coded by the IBC or IntraTMP mode) of the current block (e.g., block unit 40). In some implementations, the motion information may include at least one motion vector, at least one reference index, a reference list flag, or other related parameters.
[0116] FIG. 10 is a diagram illustrating a determination of a motion vector for determining a relocated candidate, in accordance with one or more example implementations of this disclosure.
[0117] Referring to FIG. 10, a motion vector MVnfor determining the relocated candidate may be determined based on at least one block vector of at least one reference block of the block unit 40. The at least one reference block may include a neighboring block B0of the block unit 40, and blocks B1to Bn, where n is an integer greater than or equal to 1. The block B1may be determined based on (e.g., the block vector BV0, 1or the motion vector of) the neighboring block B0, the block B2may be determined based on (e.g., the block vector BV1, 2or the motion vector of) the block B1, and so on, such that the block Bnmay be determined iteratively. The motion vector MVnmay be determined / originated from the block Bn.
[0118] For example, if the neighboring block B0is coded using the IBC or IntraTMP mode, B0, 1may represent a block vector used to predict the neighboring block B0, and B0, 1may point to the (reference) block B1. If the (reference) block B1is also coded using the IBC or IntraTMP mode, B1, 2may represent a block vector used to predict the (reference) block B1, and B1, 2may point to the (reference) block B2. Iteratively, the (reference) block Bnwhich is coded in an inter prediction mode may be found, and the motion information (e.g., including the motion vector MVn) of the block Bnmay be determined. The relocated candidate may then be determined based on the motion information and included in the candidate list.
[0119] FIG. 11 is a diagram illustrating a determination of a motion vector for determining a relocated candidate, in accordance with one or more example implementations of this disclosure.
[0120] Referring to FIG. 11, a motion vector MVnfor determining the relocated candidate may be determined based on at least one block vector of at least one reference block of the block unit 40. The at least one reference block may include blocks B’1to B’n, where n is an integer greater than or equal to 1. The block B’1may be determined based on the block unit 40 and the block vector BV0, 1or the motion vector of a neighboring block B0, the block B’2may be determined based on (e.g., the block vector BV1, 2or the motion vector of) the block B’1, the block B’3(not shown in FIG. 11) may be determined based on (e.g., the block vector or the motion vector of) the block B’2, and so on, such that the block B’nmay be determined iteratively. The motion vector MVnmay be determined / originated from the block B’n.
[0121] For example, if the neighboring block B0is coded using the IBC or IntraTMP mode, B0, 1may represent a block vector used to predict the neighboring block B0, and B0, 1may point to the block B1. In this example, starting from the block unit 40, the block vector B0, 1may be used to determine the (reference) block B’1. If the (reference) block B’1is also coded using the IBC or IntraTMP mode, B1, 2may represent a block vector used to predict the (reference) block B’1, and B1, 2may point to the (reference) block B’2. If the (reference) block B’2is also coded using the IBC or IntraTMP mode, a block vector or a motion vector of the (reference) block B’2may point to the (reference) block B’3(not shown in FIG. 11). Iteratively, the (reference) block B’nwhich is coded in an inter prediction mode may be found, and the motion information (e.g., including the motion vector MVn) of the block B’nmay be determined. The relocated candidate may then be determined based on the motion information and included in the candidate list.
[0122] FIG. 12 is a diagram illustrating a determination of a block vector based on a block, in accordance with one or more example implementations of this disclosure. It should be noted that, the determination of a block vector is exemplified and illustrated with reference to FIG. 12, however, a similar process may also be applied for determining a motion vector.
[0123] Referring to FIG. 12, in some implementations, when determining the block Bk+1or B’k+1based on the block vector BVk, k+1of the block Bkor B’k(e.g., k may be an integer selected from 1 to n-1), the top-left position 1201, the top-right position 1202, the center position 1203, the bottom-left position 1204, and the bottom-right position 1205 of the block Bkmay be checked if a block vector is stored for the corresponding position, to determine the block vector BVk, k+1.
[0124] For example, the top-left position 1201 of the block Bkmay be checked if any block vector is stored for the top-left position 1201 of the block Bk, and determine the block vector BVk, k+1based on the block vector stored for the top-left position 1201 or select the block vector BVk, k+1as the block vector stored for the top-left position 1201. For example, the top-right position 1202 of the block Bkmay be checked if any block vector is stored for the top-right position 1202 of the block Bk, and determine the block vector BVk, k+1based on the block vector stored for the top-right position 1202 or select the block vector BVk, k+1as the block vector stored for the top-right position 1202. For example, the center position 1203 of the block Bkmay be checked if any block vector is stored for the center position 1203 of the block Bk, and determine the block vector BVk, k+1based on the block vector stored for the center position 1203 or select the block vector BVk, k+1as the block vector stored for the center position 1203. For example, the bottom-left position 1204 of the block Bkmay be checked if any block vector is stored for the bottom-left position 1204 of the block Bk, and determine the block vector BVk, k+1based on the block vector stored for the bottom-left position 1204 or select the block vector BVk, k+1as the block vector stored for the bottom-left position 1204. For example, the bottom-right position 1205 of the block Bkmay be checked if any block vector is stored for the bottom-right position 1205 of the block Bk, and determine the block vector BVk, k+1based on the block vector stored for the bottom-right position 1205 or select the block vector BVk, k+1as the block vector stored for the bottom-right position 1205.
[0125] In some implementations, the top-left position 1201, the top-right position 1202, the center position 1203, the bottom-left position 1204, and the bottom-right position 1205 of the block Bkmay be checked in a pre-determined order until a valid block vector is found being stored for the corresponding position.
[0126] FIG. 13 is a diagram illustrating a determination of a motion vector based on a block, in accordance with one or more example implementations of this disclosure.
[0127] Referring to FIG. 13, in some implementations, when determining the motion vector MVnbased on the block Bnor B’n, the top-left position 1301, the top-right position 1302, the center position 1303, the bottom-left position 1304, and the bottom-right position 1305 of the block Bnmay be checked if a motion vector is stored for the corresponding position, to determine the motion vector MVn.
[0128] For example, the top-left position 1301 of the block Bnmay be checked if any motion vector is stored for the top-left position 1301 of the block Bn, and determine the motion vector MVnbased on the motion vector stored for the top-left position 1301 or select the motion vector MVnas the motion vector stored for the top-left position 1301. For example, the top-right position 1302 of the block Bnmay be checked if any motion vector is stored for the top-right position 1302 of the block Bn, and determine the motion vector MVnbased on the motion vector stored for the top-right position 1302 or select the motion vector MVnas the motion vector stored for the top-right position 1302. For example, the center position 1303 of the block Bnmay be checked if any motion vector is stored for the center position 1303 of the block Bn, and determine the motion vector MVnbased on the motion vector stored for the center position 1303 or select the motion vector MVnas the motion vector stored for the center position 1303. For example, the bottom-left position 1304 of the block Bnmay be checked if any motion vector is stored for the bottom-left position 1304 of the block Bn, and determine the motion vector MVnbased on the motion vector stored for the bottom-left position 1304 or select the motion vector MVnas the motion vector stored for the bottom-left position 1304. For example, the bottom-right position 1305 of the block Bnmay be checked if any motion vector is stored for the bottom-right position 1305 of the block Bn, and determine the motion vector MVnbased on the motion vector stored for the bottom-right position 1305 or select the motion vector MVnas the motion vector stored for the bottom-right position 1305. It should be noted that, the determination of motion vector MVnis exemplified and illustrated with reference to FIG. 13, however, a similar process may also be applied for determining the other motion vector(s) in the relocated candidate method.
[0129] In some implementations, the top-left position 1301, the top-right position 1302, the center position 1303, the bottom-left position 1304, and the bottom-right position 1305 of the block Bnmay be checked in a pre-determined order until a valid motion vector is found being stored for the corresponding position. In this case, the valid motion vector may be used for determining the relocated candidate. For example, if a motion vector is found being stored for the center position 1303 of the block Bn, the motion vector may be used for determining the relocated candidate.
[0130] In some implementations, the top-left position 1301, the top-right position 1302, the center position 1303, the bottom-left position 1304, and the bottom-right position 1305 of the block Bnmay be checked in a pre-determined order to find at least one valid motion vector stored for at least one corresponding position. In this case, each of the at least one valid motion vector may be selected as a motion vector of a relocated candidate and added into the candidate list (e.g., the inter AMVP / merge candidate list). For example, if all five motion vectors are found being stored for the top-left position 1301, the top-right position 1302, the center position 1303, the bottom-left position 1304, and the bottom-right position 1305 of the block Bn, each of the five motion vectors may be used for determining a relocated candidate.
[0131] FIG. 14 is a diagram illustrating a determination of a relocated candidate, in accordance with one or more example implementations of this disclosure.
[0132] Referring to FIG. 14, in some implementations, the motion vector MVnmay be determined based on at least one block vector of at least one reference block of the block unit 40, where the at least one reference block may include the block B0to Bn. The motion vector MVnmay be determined or selected as a motion vector of the relocated candidate, and the relocated candidate may be added into the candidate list. Relative to the block Bn, the motion vector MVnmay indicate a first reference block R0located in a reference frame 1401. Relative to the block unit 40, the motion vector MVnmay indicate a second reference block R1located in the reference frame 1401.
[0133] FIG. 15 is a diagram illustrating a determination of a relocated candidate, in accordance with one or more example implementations of this disclosure. It should be noted that, for the sake of brevity, n=2 is used as an example in FIG. 15. However, n is not limited to 2 in the present disclosure.
[0134] Referring to FIG. 15, in some implementations, the motion vector MV2may be determined based on at least one block vector of at least one reference block of the block unit 40, where the at least one reference block may include the block B0to B1. Specifically, the motion vector MV2may be determined based on the block vector BV0of a neighboring block B0of the block unit 40, and the block vector BV1of the block B1. Therefore, a (reference) block B2may be determined and the the motion vector MV2may be determined based on the motion vector stored for the block B2. The motion vector MV2originated from the block B2may indicate a first reference block R0located in a reference frame 1501. Based on the motion vector MV2and the first reference block R0located in the reference frame 1501, a motion vector MV’2, relative to or originated from the block unit 40, may be determined such that the motion vector MV’2may indicate the first reference block R0located in the reference frame 1501. The motion vector MV’2may be determined or selected as a motion vector of the relocated candidate, and the relocated candidate may be added into the candidate list.
[0135] In some implementations, the relocated candidate may be added to an (existing) candidate list. For example, the relocated candidate may be added to an inter merge candidate list, an inter AMVP candidate list, etc.
[0136] In some implementations, the relocated candidate may not be added to an (existing) candidate list (e.g., that is used for deriving the relocated candidate). For example, the relocated candidate, which is derived based on an (existing) candidate list, may not be added to the (existing) candidate list. For example, the relocated candidate may be used for establishing a separate candidate list for the inter prediction (e.g., using the (inter) merge mode). In some examples, the separate candidate list may include only the relocated candidate(s).
[0137] In some implementations, a flag may be signaled to indicate whether the relocated candidate is added to an (existing) candidate list (e.g., that is used for deriving the relocated candidate) or not. For example, a relocated flag may be signaled after a merge flag. When both the merge flag and the relocated flag are set to 1, the relocated candidate may be derived from an inter merge candidate list and subsequently added into the inter merge candidate list; when the merge flag is set to 1 and the relocated flag is set to 0, the relocated candidate is not added to the inter merge candidate list used for the inter merge mode; when the merge flag is set to 0, the relocated flag is not signaled.
[0138] In some implementations, a flag may be signaled to indicate whether the separate candidate list described above is used for the inter prediction (e.g., using the inter merge mode). For example, a relocated flag may be signaled. When the relocated flag is set to 1, the separate candidate list as described above is used for the inter prediction; when the relocated flag is set to 0, an existing candidate list (e.g., an inter merge candidate list without the relocated candidate) is used for the inter prediction.
[0139] For example, a relocated flag may be signaled to indicate whether the relocated candidate is used. The relocated candidate may be selected from a relocated candidate list, which may be derived based on a candidate list and include multiple relocated candidates. The candidate list may be, for example, an inter AMVP candidate list (e.g., as described in VVC / ECM) or an inter merge candidate list (e.g., as described in VVC / ECM). In some implementations, a relocated candidate may be used to derive subsequent relocated candidates.
[0140] Based on the relocated candidate method as described above, at least one relocated candidate may be determined and added in a candidate list (e.g., an inter AMVP candidate list or an inter merge candidate list).
[0141] FIG. 16 is a flowchart illustrating a method / process for constructing a candidate list, in accordance with one or more example implementations of this disclosure. The method / process 1600 is an example implementation, as there may be a variety of methods of constructing the candidate list. It should be noted that, additional blocks may be added for constructing the candidate list, without departing from the scope of the present disclosure.
[0142] At block 1610, the method / process 1600 may start by determining a first motion vector that indicates a first reference block, among multiple reference blocks, in a reference frame relative to the block unit 40, based on at least one block vector of at least one of the multiple reference blocks.
[0143] With reference to FIG. 15 (e.g., that takes n=2 as an example), a first motion vector (e.g., MV’n) that indicates a first reference block (e.g., R0) in a reference frame (e.g., 1501) relative to the block unit 40 may be determined, based on at least one block vector (e.g., at least one of BV0to BVn-1) of at least one reference block (e.g., at least one of B0to Bn-1) among multiple reference blocks (e.g., R0and B0to Bn). Specifically, a second reference block (e.g., Bn) may be determined based on the at least one block vector (e.g., at least one of BV0to BVn-1) of at least one other reference block (e.g., at least one of B0to Bn-1), different from the first reference block (e.g., R0) and the second reference block (e.g., Bn), among the multiple reference blocks (e.g., R0and B0to Bn). As shown in FIG. 15, the first reference block (e.g., R0) may be indicated by a second motion vector (e.g., MVn) originated from the second reference block (e.g., Bn). Details of each determination are described above and not repeated herein.
[0144] At block 1620, the method / process 1600 may include the first motion vector into a candidate list established for the block unit 40. The candidate list may include, for example, an inter AMVP candidate list and / or an inter merge candidate list.
[0145] With reference to FIG. 15 (e.g., that takes n=2 as an example), the first motion vector (MV’n) may be determined or selected as a motion vector of the relocated candidate, and the relocated candidate may be added into the candidate list.
[0146] Once the candidate list is completely constructed, the prediction may be determined (e.g., by the decoder module 124) using an inter prediction mode based on the constructed candidate list. In some implementations, the candidate list may be an inter merge candidate list, at least one candidate (e.g., each including a motion vector) may be selected from the inter merge candidate list, and the prediction of the block unit 40 may be determined based on the at least one candidate as selected, e.g., using the inter merge prediction mode.
[0147] INTRA PREDICTION
[0148] In some implementations, the (intra) prediction of the block unit (e.g., using the same reference numeral 40 for brevity) may be determined based on an intra prediction process using an intra prediction mode. The intra prediction mode may be, for example, a predefined mode or a derived mode.
[0149] In some implementations, the predefined mode may be a Planar mode, a direct current (DC) mode, or one of multiple angular modes. The number of the angular modes may be, for example, 65 or 129. In some implementations, the predefined mode may be the Planar mode in a case that the block unit 40 is a luma block. In some implementations, the predefined mode may be the DC mode in a case that the block unit 40 is a chroma block.
[0150] In some implementations, the derived mode may be a template-based intra mode derivation (TIMD) mode or a decoder-side intra mode derivation (DIMD) mode.
[0151] In some implementations, the TIMD mode may apply an intra TM method to calculate a cost between a first template, which includes multiple reconstructed samples, and a second template generated based on a candidate mode. The intra TM method may generate multiple TM costs for candidate modes in a candidate list (e.g., a Most Probable Mode (MPM) list). The candidate mode that yields the minimum TM cost may be selected as the TIMD mode.
[0152] In some implementations, a TIMD flag may be used to indicate whether the TIMD mode is selected as the final prediction mode for predicting or reconstructing the block unit 40. For example, when the TIMD flag is equal to 1, the TIMD mode is selected, and the intra-prediction mode index is not signaled.
[0153] FIG. 17 is a diagram illustrating a TIMD mode, in accordance with one or more example implementations of this disclosure.
[0154] In some implementations, the template may include samples neighboring the block unit 40. For example, the template may include samples in two rectangular regions: one located above the block unit 40 and the other located to the left of the block unit 40. These rectangular regions may be referred to as the template regions of the block unit 40. Referring to FIG. 17, the template regions 1701, 1702 of the block unit 40 and their associated reference lines 1703, 1704 are illustrated. The width of the block unit 40 may be denoted as W, and the height may be denoted as H. The above template region 1701 may have a width equal to W and a height denoted as K, while the left template region 1702 may have a height equal to H and a width denoted as K. Both template regions 1701, 1702 may include a set of reconstructed samples.
[0155] In some implementations, the predictions of the template regions 1701, 1702 may be determined based on a reference lines 1703. 1704 of the template regions 1701, 1702. In some implementations, the reference lines 1703, 1704 may include samples neighboring the template regions 1701, 1702.
[0156] In some implementations, the candidate modes in the candidate list for the intra TM method may include intra prediction modes such as the Planar, DC, and angular modes, where the number of angular modes may be 65 or 129. In some implementations, the candidate modes may be the modes in the Most Probable Mode (MPM) list. A (TM) cost between the template region (e.g., the template regions 1701, 1702) and its prediction from each candidate mode may be calculated using a specific metric, such as the Sum of Absolute Differences (SAD) or the Sum of Absolute Transformed Differences (SATD), and the candidate mode that yields the minimum (TM) cost may be selected as the TIMD mode.
[0157] In some implementations, the DIMD mode may select N intra prediction modes based on a Histogram of Gradient (HoG) generated using samples from a template region of the block unit 40, where N is a positive integer, e.g., two. The intra prediction modes may be angular modes in versatile video coding (VVC), where the angular mode indices may range from 2 to 66. After selecting N intra prediction modes, the predictors corresponding to the N intra predication modes may be computed, and the N predictors may be weighted and averaged to generate the final predictor for the block unit 40.
[0158] FIG. 18 is a diagram illustrating a DIMD mode, in accordance with one or more example implementations of this disclosure.
[0159] Referring to FIG. 18, the template region 1801 may include L neighboring reference lines of the block unit 40, where L is a positive integer, e.g., three. The HoG may be obtained by filtering the template region 1801 using a filter, e.g., a Sobel filter, where the horizontal and vertical filter kernels 1802 may be Gx(e.g., [(1, 0, -1), (2, 0, -2), (1, 0, -1)]) and Gy(e.g., [(1, 2, 1), (0, 0, 0), (-1, -2, -1)]). Specifically, the horizontal and vertical filter kernels 1802 may be centered on the pixels of the middle line of the template region 1801. For each position in the template region 1801, the (gradient) angle may be derived as arctan(Gx / Gy), and the (gradient) amplitude may be derived as |Gx| + |Gy|. The angles derived from the template region 1801 may be mapped to corresponding angular modes. After processing all positions in the template region 1801, the HoG may be generated. The relationship between the angles and the intra prediction mode indices (e.g., the mode indices of angular modes) may be represented in the form of lookup tables (LUTs), mathematical functions, predefined values, or a combination thereof. In the histogram, the x-axis may represent angular modes, while the y-axis may represent the corresponding amplitudes.
[0160] In some implementations, K angular modes corresponding to the K highest amplitudes may be selected from the HoG, where K is a positive integer. These K selected angular modes may be used to compute predictors, which may then be combined through weighted averaging to produce the final predictor for the block unit 40.
[0161] In some implementations, two angular modes with the highest two amplitudes may be selected from the HoG, and their corresponding predictors, along with one default predictor, are weighted and averaged to generate the final predictor for the block unit 40. The default predictor may be, for example, the Planar mode, and the weighted averaging process may also be referred to as “blending.”
[0162] In some implementations, the weighted averaging process may involve determining K weights corresponding to the K selected predictors. The weights may be calculated based on the ratio of the amplitudes of the K selected modes.
[0163] For example, if three (e.g., K = 3) selected intra prediction modes are denoted as M1, M2, M3, and their respective amplitudes are denoted as A1, A2, A3, the weights W1, W2, W3corresponding to M1, M2, M3may be calculated as follows: W1= (A1 / ( A1+ A2+ A3)) * W; W2= (A2 / ( A1+ A2+ A3)) * W; W3= (A3 / ( A1+ A2+ A3)) * W, where W may be the total weight to be distributed.
[0164] In some implementations, when two angular modes are selected from the HoG along with one default mode, the weights may be derived based on the ratio of the amplitudes and a constant factor. For example, if M1is the default mode and its corresponding weight W1is predefined as 22 / 64, when the total weight is 1, the weights W2and W3of the two angular modes may be calculated as follows: W2= (A2 / (A2+ A3)) * 42 / 64; W3= (A3 / (A2+ A3)) * 42 / 64.
[0165] COMBINED INTER AND INTRA PREDICTION
[0166] In some implementations, the (combined inter and intra) prediction of the block unit (e.g., using the same reference numeral 40 for brevity) may be determined based on both of the inter prediction process and the intra prediction process. The prediction determined based on both of the inter prediction process and the intra prediction process may also be referred to as a combined inter and intra prediction (CIIP). Specifically, the CIIP may combine the inter prediction signal and the intra prediction signal using weighted average.
[0167] In some implementations, the CIIP PCIIPof a block unit 40 may be derived as follows: PCIIP= ((4-wt)*Pinter+ wt * Pintra+2) >> 2, where Pintermay be the inter prediction signal, representing the prediction of the block unit 40 determined based on the inter prediction process; Pintramay be the intra prediction signal, representing the prediction of the block unit 40 determined based on the intra prediction process; and wt may be a (combined) weight.
[0168] In some implementations, the weight wt may be determined based on the coding modes of a top neighboring (e.g., adjacent) block and a left neighboring (e.g., adjacent) block as follows: 1) If the top neighboring block is available and coded based on the intra prediction process, setting a variable isIntraTop to 1; otherwise, setting isIntraTop to 0; 2) if the left neighboring block is available and coded based on the intra prediction process, setting a variable isIntraLeft to 1; otherwise, setting isIntraLeft to 0; 3) if (isIntraTop + isIntraLeft) is equal to 2, setting wt to 3; if (isIntraTop + isIntraLeft) is equal to 1, setting wt to 2; otherwise, setting wt to1.
[0169] FIGS. 19A and 19B are diagrams illustrating a combined inter and intra prediction, in accordance with one or more example implementations of this disclosure.
[0170] [Rectified under Rule 91, 28.01.2025]In some implementations, the block unit 40 may be divided into four sub-blocks based on the intra prediction mode, where each sub-block may be assigned different combined weights. The division strategy may be adopted when the intra prediction mode is an angular mode and may depend on whether the angular mode is classified as near-horizontal or near-vertical. Specifically, if the intra prediction mode is an angular mode and the angular mode falls within a near-horizontal range (e.g., 2 ≦ the angular mode < 34), the block unit 40 may be divided into four vertical sub-blocks with indexes 0 to 3, as shown in FIG. 19A. These sub-blocks may be created by partitioning the block unit 40 along vertical boundaries, resulting in sub-blocks that span the entire height of the block unit 40 but occupy distinct vertical sections of equal width. Conversely, if the intra prediction mode is an angular mode and the angular mode falls within a near-vertical range (e.g., 34 ≦ the angular mode ≦ 66), the block unit 40 may be divided into four horizontal sub-blocks with indexes 0 to 3, as shown in FIG. 19B. In this case, the block unit 40 may be partitioned along horizontal boundaries, resulting in sub-blocks that span the entire width of the block unit 40 but occupy distinct horizontal sections of equal height.
[0171] Each sub-block may be assigned a pair of combined weights, denoted as wIntra and wInter, which may correspond to the contributions of the intra prediction signal and the inter prediction signal, respectively. The combined weights for each sub-block may be predefined as shown in Table 1 below. The sub-block index, which determines the specific weights assigned to a sub-block, is illustrated in FIGS. 19A and 19B to aid in the assignment process.
[0172] The intra and inter prediction for each sub-block may be determined and the CIIP PCIIPfor the block unit 40 may be derived, based on the weighted average of the sub-blocks, as follows: PCIIP= (wIntra * Pintra+ wInter * Pinter+ 2) >> 3.
[0173] In some implementations, the combined weights of the inter predication signal and the intra prediction signal are not selected from a pre-defined set, as the pre-dined weight may not suitable for every combination of intra prediction signal and inter prediction signal or for every block size. For example, the CIIP of the block unit 40 may be determined based on both the inter prediction process and the intra prediction process, and the relationship between the CIIP and the inter and intra prediction signals may be determined based on a neighboring region using a regression method. In this case, the CIIP may be referred to as a regression CIIP or a regression-based CIIP.
[0174] FIG. 20 is a diagram illustrating a neighboring region, in accordance with one or more example implementations of this disclosure.
[0175] Referring to FIG. 20, the neighboring region of the block unit 40 may include one or more of the following regions: an above region 2001, a left region 2002, and an above-left region 2003. The above region 2001 may refer to a region immediately above the block unit 40, the left region 2002 may refer to a region immediately to the left of the block unit 40, and the above-left region 2003 may refer to a region diagonally above-left to the block unit.
[0176] Referring to FIG. 20, in some implementations, the width of the above region 2001 may be equal to a positive integer value, e.g., RW, and the height of the above region 2001 may be equal to a positive integer value, e.g., K. In some implementations, the value RW may be equal to the width of the block unit 40. In some implementations, the value RW may be equal to two times the width of the block unit 40. In some implementations, the value K may be equal to 1, 2, 3, 4, 8, 12, 16, or 32. In some implementations, the value of K may be equal to the height of the block unit.
[0177] Referring to FIG, 20, in some implementations, the height of the left region 2002 may be equal to a positive integer value, e.g., RH, and the width of the left region 2002 may be equal to a positive integer value, e.g., K. In some implementations, the value RH may be equal to the height of the block unit 40. In some implementations, the value RH may be equal to two times the height of the block unit 40. In some implementations, the value K may be equal to 1, 2, 3, 4, 8, 12, 16, or 32. In some implementations, the value K may be equal to the width of the block unit 40.
[0178] In some implementations, the neighboring region may include the above region 2001 and the left region 2002.
[0179] In some implementations, the neighboring region may include the left region 2002 and the region above-left region 2003.
[0180] In some implementations, the neighboring region may include the left region 2002 only. In some implementations, the neighboring region may include only the left region 2002 if the region above to the block unit 40 is not available (e.g., out of the picture boundary, outside of the current CTU row, etc.).
[0181] In some implementations, the neighboring region may include the above region 2001 and the above-left region 2003.
[0182] In some implementations, the neighboring region may include the above region 2001 only. In some implementations, the neighboring region may include only the above region if the region left to the block unit 40 is not available.
[0183] In some implementations, the regression-based CIIP may be determined based on multiple features and multiple weights (e.g., denoted as ci, i being an integer) corresponding to the features. The features include at least the intra prediction signal and the inter prediction signal, and the weights may be determined based on the neighboring region using the regression method.
[0184] In some implementations, the features include the intra prediction signal Pintraand the inter prediction signal Pinter, and the regression-based CIIP PREG_CIIPof a block unit 40 may be derived as follows: PREG_CIIP= c0Pintra+ c1Pinter.
[0185] In some implementations, the features include the intra prediction signal Pintra, the inter prediction signal Pinter, and a bias B, and the regression-based CIIP PREG_CIIPof a block unit 40 may be derived as follows: PREG_CIIP= c0Pintra+ c1Pinter+ c2B.
[0186] In some implementations, the bias may be determined based on the bit depth (e.g., denoted as bitDepth). For example, the bias may be derived as follows: B = 1 << (bitDepth - 1).
[0187] In some implementations, the features include the intra prediction signal Pintra, the inter prediction signal Pinter, and an intra mode index M, and the regression-based CIIP PREG_CIIPof a block unit 40 may be derived as follows: PREG_CIIP= c0Pintra+ c1Pinter+ c2M.
[0188] The intra mode index M may reflect the texture characteristics of the block unit 40. For example, if the intra mode index M is greater than or equal to 2 and less than 34, the texture of the block unit 40 may be horizontal-like. If the intra mode index M is greater than or equal to 34 and less than or equal to 66, the texture of the block unit 40 may be vertical-like.
[0189] In some implementations, the regression-based CIIP PREG_CIIPof a block unit 40 may be derived as follows: PREG_CIIP= c0Pintra+ c1Pinter+ c2M + c3B.
[0190] In some implementations, the regression-based CIIP PREG_CIIPof a block unit 40 may be associated with the position (x, y) of the predicted sample, and derived as follows: PREG_CIIP(x, y) = c0Pintra(x, y) + c1Pinter(x, y) + c2x + c3y, where x may represent the sample position along the x-axis within the block unit 40, and y may represent the sample position along the y-axis within the block unit 40. The regression-based CIIP prediction may be improved by incorporating the position-dependent terms. These position-dependent terms may allow the predicted samples of the regression-based CIIP to exhibit greater dependency on their spatial positions.
[0191] In some implementations, the regression-based CIIP PREG_CIIPof a block unit 40 may be derived as follows: PREG_CIIP(x, y) = c0Pintra(x, y) + c1Pinter(x, y) + c2x + c3y + c4B.
[0192] In some implementations, the regression-based CIIP PREG_CIIPmay be derived as follows: PREG_CIIP(x, y) = c0Pintra(x, y) + c1Pinter(x, y) + c2x + c3y + c4M.
[0193] In some implementations, the regression-based CIIP PREG_CIIPof a block unit 40 may be derived as follows: PREG_CIIP(x, y) = c0Pintra(x, y) + c1Pinter(x, y) + c2x + c3y + c4M + c5B.
[0194] In some implementations, the regression-based CIIP PREG_CIIPof a block unit 40 may be derived as follows: PREG_CIIP= Wintra(x, y)Pintra(x, y) + Winter(x, y)Pinter(x, y); Wintra(x, y) = c0x + c1y + c2B; Winter(x, y) = N - Wintra(x, y), where N may be a positive integer (e.g., 1, 2, 4, 8, 16, 32, 64, 128, etc.).
[0195] The weights cimay be derived by minimizing the difference between the prediction of the neighboring region and the reconstruction of the neighboring region. In some implementations, the difference may be calculated using the Mean Squared Error (MSE). In some implementations, the difference may be calculated using the Sum of Absolute Differences (SAD). In some implementations, the difference may be calculated using the Sum of Absolute Transformed Differences (SATD).
[0196] In some implementations, the reconstruction of the neighboring region may include multiple reconstructed samples within the neighboring region. The prediction of the neighboring region may be derived by combining an intra prediction of the neighboring region and an inter prediction of the neighboring region. The intra prediction of the neighboring region may be determined based on the intra prediction mode as described above, while the inter prediction of the neighboring region may be determined based on the inter prediction mode as described above.
[0197] FIG. 21 is a diagram illustrating a regression-based prediction of a neighboring region, in accordance with one or more example implementations of this disclosure.
[0198] Referring to FIG. 21, in some implementations, the regression-based prediction of the neighboring region 2101 of the block unit 40 may be derived as follows: PNEI_PRED(x, y) = c0Pintra(x, y) + c1Pinter(x, y) + c2x + c3y + c4B, where PNEI_PRED(x, y) may represent the regression-based prediction of each sample in the neighboring region 2101; Pintra(x, y) may represent the intra prediction of each sample in the neighboring region 2101; Pinter(x, y) may represent the inter prediction of each sample in the neighboring region 2101; x may represent the sample position along the x-axis within the neighboring region 2101; y may represent the sample position along the y-axis within the neighboring region 2101; and B may represent a bias. The weights c0, c1, c2, c3, c4may be derived by minimizing the difference between the regression-based prediction PNEI_PRED(x, y) and the reconstruction PNEI_RECof the neighboring region 2101.
[0199] Based on at least one prediction method described above, the prediction of the block unit 40 may be determined.
[0200] Referring back to FIG. 3, at block 340, the method / process 300 may reconstruct (e.g., by the decoder module 124) the block unit based on the prediction.
[0201] In some implementations, the decoder module 124 may add a plurality of residual components into the prediction block (e.g., the prediction of the block unit determined at block 330) to reconstruct the block unit. The residual components may be determined from the bitstream.
[0202] Referring back to FIG. 3, once the block unit is reconstructed, the method / process 300 may then end. By repeating the method / process 300, the multiple image frames included in the video data may be reconstructed.
[0203] FIG. 22 is a block diagram illustrating an encoder module 114 of the first electronic device 110 illustrated in FIG. 1, in accordance with one or more example implementations of this disclosure. The encoder module 114 may include a prediction processor (e.g., a prediction processing unit 2141), at least a first summer (e.g., a first summer 2142) and a second summer (e.g., a second summer 2145), a transform / quantization processor (e.g., a transform / quantization unit 2143), an inverse quantization / inverse transform processor (e.g., an inverse quantization / inverse transform unit 2144), a filter (e.g., a filtering unit 2146), a decoded picture buffer (e.g., a decoded picture buffer 2147), and an entropy encoder (e.g., an entropy encoding unit 2148). The prediction processing unit 2141 of the encoder module 114 may further include a partition processor (e.g., a partition unit 21411), an intra prediction processor (e.g., an intra prediction unit 21412), and an inter prediction processor (e.g., an inter prediction unit 21413).
[0204] The encoder module 114 may receive the source video and encode the source video to output a bitstream. The encoder module 114 may receive source video including multiple image frames and then divide the image frames according to a coding structure. Each of the image frames may be divided into at least one image block.
[0205] The at least one image block may include a luminance block having multiple luminance samples and at least one chrominance block having multiple chrominance samples. The luminance block and the at least one chrominance block may be further divided to generate macroblocks, CTUs, CBs, sub-divisions thereof, and / or other equivalent coding units.
[0206] The encoder module 114 may perform additional sub-divisions of the source video. It should be noted that the disclosed implementations are generally applicable to video coding regardless of how the source video is partitioned prior to and / or during the encoding.
[0207] During the encoding process, the prediction processing unit 2141 may receive a current image block of a specific one of the image frames. The current image block may be the luminance block or one of the chrominance blocks in the specific image frame.
[0208] The partition unit 21411 may divide the current image block into multiple block units. The intra prediction unit 21412 may perform intra-predictive coding of a current block unit relative to one or more neighboring blocks in the same frame as the current block unit in order to provide spatial prediction. The inter prediction unit 21413 may perform inter-predictive coding of the current block unit relative to one or more blocks in one or more reference image blocks to provide temporal prediction.
[0209] The prediction processing unit 2141 may select one of the coding results generated by the intra prediction unit 21412 and the inter prediction unit 21413 based on a mode selection method, such as a cost function. The mode selection method may be a rate-distortion optimization (RDO) process.
[0210] The prediction processing unit 2141 may determine the selected coding result and provide a predicted block corresponding to the selected coding result to the first summer 2142 for generating a residual block and to the second summer 2145 for reconstructing the encoded block unit. The prediction processing unit 2141 may further provide syntax elements, such as motion vectors, intra-mode indicators, partition information, and / or other syntax information, to the entropy encoding unit 2148.
[0211] The intra prediction unit 21412 may intra-predict the current block unit. The intra prediction unit 21412 may determine an intra prediction mode directed toward a reconstructed sample neighboring the current block unit in order to encode the current block unit.
[0212] The intra prediction unit 21412 may encode the current block unit using various intra prediction modes. The intra prediction unit 21412 of the prediction processing unit 2141 may select an appropriate intra prediction mode from the selected modes. The intra prediction unit 21412 may encode the current block unit using a cross-component prediction mode to predict one of the two chroma components of the current block unit based on the luma components of the current block unit. The intra prediction unit 21412 may predict a first one of the two chroma components of the current block unit based on the second of the two chroma components of the current block unit.
[0213] The inter prediction unit 21413 may inter-predict the current block unit as an alternative to the intra prediction performed by the intra prediction unit 21412. The inter prediction unit 21413 may perform motion estimation to estimate motion of the current block unit for generating a motion vector.
[0214] The motion vector may indicate a displacement of the current block unit within the current image block relative to a reference block unit within a reference image block. The inter prediction unit 21413 may receive at least one reference image block stored in the decoded picture buffer 2147 and estimate the motion based on the received reference image blocks to generate the motion vector.
[0215] The first summer 2142 may generate the residual block by subtracting the prediction block determined by the prediction processing unit 2141 from the original current block unit. The first summer 2142 may represent the component or components that perform this subtraction.
[0216] The transform / quantization unit 2143 may apply a transform to the residual block in order to generate a residual transform coefficient and then quantize the residual transform coefficients to further reduce the bit rate. The transform may be one of a DCT, DST, AMT, MDNSST, HyGT, signal-dependent transform, KLT, wavelet transform, integer transform, sub-band transform, and a conceptually similar transform.
[0217] The transform may convert the residual information from a pixel value domain to a transform domain, such as a frequency domain. The degree of quantization may be modified by adjusting a quantization parameter.
[0218] The transform / quantization unit 2143 may perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 2148 may perform the scan.
[0219] The entropy encoding unit 2148 may receive multiple syntax elements from the prediction processing unit 2141 and the transform / quantization unit 2143, including a quantization parameter, transform data, motion vectors, intra modes, partition information, and / or other syntax information. The entropy encoding unit 2148 may encode the syntax elements into the bitstream.
[0220] The entropy encoding unit 2148 may entropy encode the quantized transform coefficients by performing CAVLC, CABAC, SBAC, PIPE coding, or another entropy coding technique to generate an encoded bitstream. The encoded bitstream may be transmitted to another device (e.g., the second electronic device 120, as shown in FIG. 1) or archived for later transmission or retrieval.
[0221] The inverse quantization / inverse transform unit 2144 may apply inverse quantization and inverse transformation to reconstruct the residual block in the pixel domain for later use as a reference block. The second summer 2145 may add the reconstructed residual block to the prediction block provided by the prediction processing unit 2141 in order to produce a reconstructed block for storage in the decoded picture buffer 2147.
[0222] The filtering unit 2146 may include a deblocking filter, an SAO filter, a bilateral filter, and / or an ALF to remove blocking artifacts from the reconstructed block. Other filters (in loop or post loop) may be used in addition to the deblocking filter, the SAO filter, the bilateral filter, and the ALF. Such filters are not illustrated for brevity and may filter the output of the second summer 2145.
[0223] The decoded picture buffer 2147 may be a reference picture memory that stores the reference block to be used by the encoder module 114 to encode video, such as in intra-coding or inter-coding modes. The decoded picture buffer 2147 may include a variety of memory devices, such as DRAM (e.g., including SDRAM), MRAM, RRAM, or other types of memory devices. The decoded picture buffer 2147 may be on-chip with other components of the encoder module 114 or off-chip relative to those components.
[0224] The method / process 300 may be performed by the first electronic device 110 for encoding video data. The encoder module 114 may receive the video data. The video data received by the encoder module 114 may be a video. The encoder module 114 may determine a block unit from an image frame according to the video data. The encoder module 114 may divide the image frame to generate multiple CTUs, and further divide one of the CTUs to determine the block unit according to one of multiple partition schemes based on any video coding standard.
[0225] The encoder module 114 may establish a candidate list, such as an inter merge candidate list, based on the relocated candidate method as described above, for example. When establishing the candidate list, the encoder module 114 may determine a first motion vector that indicates a first reference block, among multiple reference blocks, in a reference frame relative to the block unit, based on at least one block vector of at least one of the multiple reference blocks. Afterwards, the encoder module 114 may include the first motion vector into a candidate list established for the block unit. Details for establishing the candidate list (e.g., based on the relocated candidate method) are described above and therefore are not repeated herein.
[0226] The encoder module 114 may use the method / process 300 to predict the block unit based on the candidate list (e.g., using an inter merge mode). The encoder module 114 may method / process 300 to further reconstruct the block unit based on the prediction, for example, to generate a reconstructed block including a plurality of reconstructed samples. The reconstructed samples of the block unit may be used as references for predicting a plurality of following blocks in the video data.
[0227] The disclosed implementations are to be considered in all respects as illustrative and not restrictive. It should also be understood that the present disclosure is not limited to the specific disclosed implementations, but that many rearrangements, modifications, and substitutions are possible without departing from the scope of the present disclosure.
Claims
1. An electronic device for decoding video data, the electronic device comprising: at least one processor; and at least one non-transitory computer-readable medium coupled to the at least one processor and storing one or more computer-executable instructions that, when executed by the at least one processor, cause the electronic device to: receive the video data; determine a block unit from an image frame according to the video data; determine a first motion vector that indicates a first reference block, among a plurality of reference blocks, in a reference frame relative to the block unit, based on at least one block vector of at least one of the plurality of reference blocks; include the first motion vector into a candidate list established for the block unit; determine a prediction of the block unit based on the candidate list; and reconstruct the block unit based on the prediction.
2. The electronic device of claim 1, wherein the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine a second reference block based on the at least one block vector of at least one other reference block, different from the first reference block and the second reference block, among the plurality of reference blocks, wherein the first reference block is indicated by a second motion vector originated from the second reference block.
3. The electronic device of claim 2, wherein the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine the second motion vector based on at least one of a top-left position, a top-right position, a center position, a bottom-left position, and a bottom-right position of the second reference block.
4. The electronic device of claim 3, wherein determining the second motion vector comprises: checking whether a third motion vector is stored corresponding to the center position of the second reference block; and determining the second motion vector based on the third motion vector when the third motion vector is stored corresponding to the center position of the second reference block.
5. The electronic device of claim 1, wherein the candidate list comprises an inter merge candidate list, and determining the prediction of the block unit based on the candidate list comprises: selecting at least one motion vector from the inter merge candidate list; and determining, based on the at least one motion vector as selected, the prediction of the block unit.
6. An electronic device for encoding video data, the electronic device comprising: at least one processor; and at least one non-transitory computer-readable medium coupled to the at least one processor and storing one or more computer-executable instructions that, when executed by the at least one processor, cause the electronic device to: receive the video data; determine a block unit from an image frame according to the video data; determine a first motion vector that indicates a first reference block, among a plurality of reference blocks, in a reference frame relative to the block unit, based on at least one block vector of at least one of the plurality of reference blocks; include the first motion vector into a candidate list established for the block unit; determine a prediction of the block unit based on the candidate list; and reconstruct the block unit based on the prediction.
7. The electronic device of claim 6, wherein the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine a second reference block based on the at least one block vector of at least one other reference block, different from the first reference block and the second reference block, among the plurality of reference blocks, wherein the first reference block is indicated by a second motion vector originated from the second reference block.
8. The electronic device of claim 7, wherein the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine the second motion vector based on at least one of a top-left position, a top-right position, a center position, a bottom-left position, and a bottom-right position of the second reference block.
9. The electronic device of claim 8, wherein determining the second motion vector comprises: checking whether a third motion vector is stored corresponding to the center position of the second reference block; and determining the second motion vector based on the third motion vector when the third motion vector is stored corresponding to the center position of the second reference block.
10. The electronic device of claim 6, wherein the candidate list comprises an inter merge candidate list, and determining the prediction of the block unit based on the candidate list comprises: selecting at least one motion vector from the inter merge candidate list; and determining, based on the at least one motion vector as selected, the prediction of the block unit.
11. A non-transitory machine-readable medium of an electronic device storing one or more computer-executable instructions for decoding video data, the one or more computer-executable instructions, when executed by at least one processor of the electronic device, causing the electronic device to: receive the video data; determine a block unit from an image frame according to the video data; determine a first motion vector that indicates a first reference block, among a plurality of reference blocks, in a reference frame relative to the block unit, based on at least one block vector of at least one of the plurality of reference blocks; include the first motion vector into a candidate list established for the block unit; determine a prediction of the block unit based on the candidate list; and reconstruct the block unit based on the prediction.
12. The non-transitory machine-readable medium of claim 11, wherein the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine a second reference block based on the at least one block vector of at least one other reference block, different from the first reference block and the second reference block, among the plurality of reference blocks, wherein the first reference block is indicated by a second motion vector originated from the second reference block.
13. The non-transitory machine-readable medium of claim 12, wherein the one or more computer-executable instructions, when executed by the at least one processor, further cause the electronic device to: determine the second motion vector based on at least one of a top-left position, a top-right position, a center position, a bottom-left position, and a bottom-right position of the second reference block.
14. The non-transitory machine-readable medium of claim 13, wherein determining the second motion vector comprises: checking whether a third motion vector is stored corresponding to the center position of the second reference block; and determining the second motion vector based on the third motion vector when the third motion vector is stored corresponding to the center position of the second reference block.
15. The non-transitory machine-readable medium of claim 11, wherein the candidate list comprises an inter merge candidate list, and determining the prediction of the block unit based on the candidate list comprises: selecting at least one motion vector from the inter merge candidate list; and determining, based on the at least one motion vector as selected, the prediction of the block unit.
Citation Information
Patent Citations
Methods and systems for intra block copy coding with block vector derivation
US20200404321A1