Construction of BV candidates using temporal candidates
By employing IBC and IntraTMP modes in video encoding, constructing a BV candidate list, and applying TM technology, the problem of insufficient utilization of intra-frame and inter-frame redundancy information in existing technologies is solved, achieving more efficient video encoding and decoding.
Patent Information
- Application Number
- CN202480026002.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-21
- Filing Date
- 2024-04-20
- Publication Date
- 2025-11-14
AI Technical Summary
Existing video encoding and decoding technologies struggle to effectively utilize redundant information within and between frames for efficient compression when processing video data, especially when constructing block vector candidates, resulting in low encoding efficiency.
The Intra-Block Copy (IBC) mode and Intra-Template Matching Prediction (IntraTMP) mode are adopted. The block vector (BV) candidate list is determined by format rules, and temporal BV candidates are constructed using motion information and BV information. Template matching (TM) technology is applied to select appropriate candidates for video encoding and decoding.
It improves the efficiency and quality of video encoding, reduces the amount of data, and enhances encoding efficiency.
Smart Images

Figure CN120958809A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 461,229, filed April 21, 2023, entitled “BV candidate construction by using temporal MV”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure describes aspects generally related to video encoding and decoding. Background Technology
[0003] The background description provided herein is for the purpose of presenting the general content of this disclosure. The extent of the work of the currently named inventors described in the background section and various aspects of this specification does not indicate that it was prior art at the time of filing of this application, nor is it expressly or implied that it was acknowledged as prior art of this application.
[0004] Image / video compression helps transfer image / video data between different devices, storage, and networks while minimizing quality loss. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which compresses images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image being reconstructed to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which compresses images based on temporal redundancy. For example, inter-frame prediction can predict samples in the current image from previously reconstructed images using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0005] This disclosure includes methods and apparatus for video encoding / decoding.
[0006] In one aspect, a method for processing visual media data includes processing a bitstream of visual media data according to format rules. The bitstream includes syntax elements that indicate the prediction of a current block in a current image based on either an intra-block copy (IBC) mode or an intra-template matching prediction (IntraTMP) mode.
[0007] The format rule specifies that the first time-bound BV candidate in the BV candidate list of the current block is determined based on one of the following: (i) motion information and (ii) block vector (BV) information associated with the first candidate position of the co-located block.
[0008] The format rule stipulates that when one of the information in (i) motion information and (ii) BV information is motion information and the motion information indicates: motion vector (MV) and reference image corresponding to the co-positioned image, the scaled MV is determined based on MV, the difference in first picture order count (POC) between the current image and the co-positioned image, and the difference in second POC between the reference image of the co-positioned image and the co-positioned image, wherein the first time BV candidate includes the scaled MV.
[0009] The format rules stipulate that the co-position block is located in the co-position image of the current image, a BV candidate list is constructed based on the first-time BV candidate, and the current sample is processed based on the BV candidate list using one of the IBC mode and IntraTMP mode.
[0010] In one example, the method for processing the visual media data includes generating a temporal candidate list comprising temporal BV candidates associated with each candidate position of a co-occurring block. Each temporal BV candidate is based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-occurring block, and the temporal BV candidates include a first temporal BV candidate. The method for processing the visual media data includes applying template-matching (TM) to determine the TM cost of each temporal BV candidate in the temporal candidate list. The BV candidate list may be constructed based on n1 temporal BV candidates, the TM cost of which is no greater than one or more TM costs of the remaining temporal BV candidates in the temporal candidate list.
[0011] In one aspect, a video coding method includes: determining a first temporal BV candidate in a BV candidate list for a current block in a current image based on one of (i) motion information and (ii) BV information associated with a first candidate location of a co-located block, wherein the co-located block is located in a co-located image of the current image; and predicting the current block according to one of an IBC mode and an IntraTMP mode. The method includes constructing a BV candidate list based on the first temporal BV candidate. The method includes encoding a current sample based on the BV candidate list using said one of the IBC mode and the IntraTMP mode.
[0012] In one aspect, the information in (i) motion information and (ii) BV information is motion information and the motion information indicates: MV, and a reference image corresponding to the co-positioned image. The scaled MV is determined based on the MV, a first POC difference between the current image and the co-positioned image, and a second POC difference between the reference image of the co-positioned image and the co-positioned image. In one example, the scaled MV is included as a first-time BV candidate.
[0013] In one aspect, when one of the motion information and (ii) BV information is motion information associated with a co-located block, a first-time BV candidate is determined by averaging multiple motion information from each candidate location associated with the co-located block.
[0014] In one aspect, a temporal candidate list is generated, comprising temporal BV candidates associated with each candidate position of the co-located block. The temporal BV candidates include a first temporal BV candidate. Each temporal BV candidate may be based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-located block. Based on the temporal candidate list, a BV candidate list is constructed.
[0015] In one example, TM is applied to determine the TM cost of each time BV candidate in the time candidate list, the time BV candidates in the time candidate list are reordered based on each TM cost, and a BV candidate list is constructed based on the reordered time candidate list, for example, by adding the first m reordered time candidates from the reordered time candidate list to the BV candidate list.
[0016] In one example, a Time Management (TM) is applied to determine the TM cost of each time BV candidate in the time candidate list. Based on this TM cost, n1 time BV candidates are selected. The TM cost of these n1 time BV candidates is no greater than the TM cost of each remaining time BV candidate in the time candidate list. The BV candidate list is constructed based on these n1 selected time BV candidates.
[0017] In one aspect, a TM is applied to determine the TM cost of each BV candidate in the initial BV candidate list. The BV candidates include first-time BV candidates. The BV candidate list is constructed by reordering the BV candidates in the initial BV candidate list based on their individual TM costs. In this example, the initial BV candidate list and the BV candidate list contain the same BV candidates that can be arranged in different orders.
[0018] According to one aspect of this disclosure, a video decoding apparatus includes processing circuitry. The processing circuitry is configured to receive encoded information indicating the prediction of a current block in a current image based on one of an intra-block copy (IBC) mode and an intra-template matching prediction (IntraTMP) mode; and to determine a first temporal BV candidate in a BV candidate list for the current block based on one of (i) motion information and (ii) BV information associated with a co-located block in a co-located image of the current image. The processing circuitry is configured to construct a BV candidate list based on the first temporal BV candidate. The processing circuitry is configured to reconstruct a current sample based on the BV candidate list using the IBC mode and the IntraTMP mode.
[0019] In one aspect, a first-time BV candidate is determined based on one of the following information: (i) motion information and (ii) BV information associated with the first candidate position of the co-located block.
[0020] In one aspect, the information in (i) motion information and (ii) BV information is motion information and the motion information indicates: MV, and a reference image corresponding to the co-positioned image. The scaled MV is determined based on the MV, a first POC difference between the current image and the co-positioned image, and a second POC difference between the reference image of the co-positioned image and the co-positioned image. In one example, the scaled MV is included as a first-time BV candidate.
[0021] In one aspect, when one of the motion information and (ii) BV information is motion information associated with a co-located block, a first-time BV candidate is determined by averaging multiple motion information from each candidate location associated with the co-located block.
[0022] In one aspect, a temporal candidate list is generated, comprising temporal BV candidates associated with each candidate position of a co-located block. The temporal BV candidates include a first temporal BV candidate. Each temporal BV candidate may be based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-located block. Based on the temporal candidate list, a BV candidate list is constructed.
[0023] In one example, TM is applied to determine the TM cost of each time BV candidate in the time candidate list, the time BV candidates in the time candidate list are reordered based on each TM cost, and a BV candidate list is constructed based on the reordered time candidate list, for example, by adding the first m reordered time candidates from the reordered time candidate list to the BV candidate list.
[0024] In one example, a Time Management (TM) is applied to determine the TM cost of each time BV candidate in the time candidate list. Based on this TM cost, n1 time BV candidates are selected. The TM cost of these n1 time BV candidates is no greater than the TM cost of each remaining time BV candidate in the time candidate list. The BV candidate list is constructed based on these n1 selected time BV candidates.
[0025] In one aspect, a TM is applied to determine the TM cost of each BV candidate in the initial BV candidate list. The BV candidates include first-time BV candidates. The BV candidate list is constructed by reordering the BV candidates in the initial BV candidate list based on their individual TM costs. In this example, the initial BV candidate list and the BV candidate list contain the same BV candidates that can be arranged in different orders.
[0026] In one aspect, a TM is applied to determine the TM cost of each BV candidate in the BV candidate list. The BV candidates include first-time BV candidates. n² BV candidates are selected based on their TM costs. The TM cost of these n² BV candidates is no greater than the TM cost of each of the remaining BV candidates in the BV candidate list. The current sample can be reconstructed based on the n² selected BV candidates using either the IBC mode or the IntraTMP mode.
[0027] In one example, the mode mentioned above, between IBC mode and IntraTMP mode, is IBC mode. In another example, the mode mentioned above, between IBC mode and IntraTMP mode, is IntraTMP mode.
[0028] This disclosure also provides a video encoding apparatus. The video encoding apparatus includes processing circuitry configured to implement any of the described video encoding methods.
[0029] This disclosure also provides a video decoding method. This method includes any method implemented by a video decoding apparatus.
[0030] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described video decoding / encoding methods. Attached Figure Description
[0031] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which: Figure 1 This is a schematic diagram of an example block diagram of a communication system (100).
[0032] Figure 2 This is a schematic diagram of an example decoder block diagram.
[0033] Figure 3 This is a schematic diagram of an example encoder block diagram.
[0034] Figure 4 An example of an intra-block copy (IBC) mode according to one aspect of this disclosure is shown.
[0035] Figure 5 An example of intra-frame picture patch compensation with a CTU-sized search range is shown according to one aspect of this disclosure.
[0036] Figure 6 An example of an IntraTMP (Intra-Template Matching Prediction) mode according to one aspect of this disclosure is shown.
[0037] Figure 7 An example of a candidate location for a time candidate (such as a time block vector (BV) candidate) according to one aspect of this disclosure is shown.
[0038] Figure 8 An example of scaling the motion vector (MV) of a time BV candidate according to one aspect of this disclosure is shown.
[0039] Figure 9 A flowchart outlining the decoding process is shown, based on some aspects of this disclosure.
[0040] Figure 10 A flowchart outlining the coding process is shown, which summarizes some aspects of this disclosure.
[0041] Figure 11 It is a schematic diagram of a computer system based on one aspect. Detailed Implementation
[0042] Figure 1 Block diagrams of some example video processing systems (100) are shown. This video processing system (100) illustrates a video encoder and video decoder in a streaming environment as examples of applications of the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, streaming media services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0043] The video processing system (100) includes an acquisition subsystem (113) that may include a video source (101) such as a digital camera, which creates an uncompressed video image stream (102). In one example, the video image stream (102) includes samples captured by a digital camera. The video image stream (102) is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data (104) (or encoded video bitstream). The video image stream (102) may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. Compared to the video image stream (102), the encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the lower data volume of the encoded video data (or encoded video bitstream), which can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 Client subsystems (106) and (108) in the streaming server (105) can access the streaming server to retrieve copies (107) and (109) of the encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and produces an output video picture stream (111) that can be displayed on a display (112) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (e.g., video stream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed topics are applicable within the context of the VVC standard.
[0044] It should be noted that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may include a video decoder (not shown), and electronic device (130) may also include a video encoder (not shown).
[0045] Figure 2An example block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., receiving circuitry). The video decoder (210) may be used in place of... Figure 1 The video decoder (110) in the example.
[0046] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream, to be decoded by the video decoder (210). In one aspect, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not indicated). The receiver (231) may separate the encoded video sequences from other data. To prevent network jitter, a buffer (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer (215) is part of the video decoder (210). In other cases, the buffer (215) may be located outside the video decoder (210) (not indicated). In other cases, an external buffer (not shown) may be configured for the video decoder (210) to prevent network jitter, for example, and another buffer (215) may be configured internally for, for example, handling broadcast timing. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer (215) may not be necessary, or it may be made smaller. Of course, for use on packet networks such as the Internet, a buffer (215) may also be required; this buffer may be relatively large and adaptive in size, and may be at least partially implemented in the operating system or a similar component (not shown) external to the video decoder (210).
[0047] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (210), and potential information for controlling display devices (212), such as displays that are not part of the electronic device (230) but may be coupled to it. Figure 2As shown in the diagram. The control information used for the display device may be a parameter set fragment (not shown) of supplemental enhancement information (SEI) messages or video usability information (VUI). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a subgroup parameter set of at least one subgroup of pixels in the subgroups of pixels in the encoded video sequence based on at least one parameter corresponding to a group. The subgroup may include a group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0048] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0049] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (220). For brevity, the flow of such subgroup control information between the parser (220) and the various units described below is not described.
[0050] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.
[0051] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) and control information from the parser (220), including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block containing sample values, which can be input into the aggregator (255).
[0052] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed image, but can use predictive information from a previously reconstructed portion of the current image. Such predictive information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from the current picture buffer (258) to generate a block with the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (255) adds the predictive information generated by the intra-picture prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0053] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (253) can access the reference image memory (257) to extract samples for prediction. After motion compensation of the extracted samples according to the block-related symbols (221), these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit (253) can obtain the predicted samples from the address in the reference image memory (257) under motion vector control, and the motion vector is available to the motion compensation prediction unit (253) in the form of symbols (221), which, for example, have X, Y and reference image components. Motion compensation may also include interpolation of sample values extracted from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0054] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also known as the encoded video stream), which are available as symbols (221) from the parser (220) to the loop filter unit (256). Video compression may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0055] The output of the loop filter unit (256) can be a sample stream, which can be output to a display device (212) and stored in a reference image memory (257) for subsequent inter-frame image prediction.
[0056] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (220)) are identified as reference images, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0057] The video decoder (210) performs decoding operations according to a predetermined video compression technique or standard (e.g., ITU-T Recommendation H.265). The encoded video sequence conforms to the syntax specified by the video compression technique or standard, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under that configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the limits defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the hypothetical reference decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.
[0058] In one aspect, the receiver (231) may receive supplementary (redundant) data along with the encoded video. This supplementary data may be a portion of the encoded video sequence. The supplementary data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The supplementary data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0059] Figure 3 An example block diagram of a video encoder (303) is shown. The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) can be used in place of Figure 1 The video encoder (103) in the example.
[0060] The video encoder (303) can be obtained from the video source (301) (not Figure 3 In one example, an electronic device (320) receives a video sample, the video source of which can capture video images that will be encoded by a video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0061] A video source (301) can provide a sequence of source video samples to be encoded by a video encoder (303) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc. used. The following focuses on describing the samples.
[0062] According to one aspect, the video encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (343) in real time or under any other required time constraints. Implementing an appropriate encoding rate is a function of the controller (350). In some aspects, the controller (350) controls and is functionally coupled to other functional units as described below. For simplicity, coupling is not indicated in the figures. Parameters set by the controller (350) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a particular system design.
[0063] In some respects, the video encoder (303) is configured to operate within an encoding loop. In one example, simply put, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (334) are also bit-accurately corresponding between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related techniques.
[0064] The operation of the “local” decoder (333) can be combined with, for example, the above-mentioned... Figure 2 The video decoder (210) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 2 When symbols are available and the entropy encoder (345) and parser (220) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (210), including the buffer (215) and parser (220), may not be fully implemented in the local decoder (333).
[0065] In one respect, any decoder technique other than parsing / entropy decoding, which exists in the decoder, must also exist in the corresponding encoder in essentially the same functional form. Therefore, the subject matter disclosed focuses on decoder operation. The description of encoder techniques can be simplified, as encoder techniques are inverses of the fully described decoder techniques. More detailed descriptions are only required in certain places, and are provided below.
[0066] In some examples, during operation, the source encoder (330) may perform motion-compensated predictive coding. This motion-compensated predictive coding predictively encodes the input image, referencing one or more previously encoded images from the video sequence designated as "reference images." In this way, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.
[0067] The local video decoder (333) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (330). The operation of the encoding engine (332) can be a lossy process. When the encoded video data can be decoded by the video decoder (333), Figure 3 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in the reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0068] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (335) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (335), it can be determined that the input image can have prediction references obtained from multiple reference images stored in the reference image memory (334).
[0069] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0070] The outputs of all the above functional units can be entropy encoded in the entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.
[0071] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0072] The controller (350) manages the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types: Intra-frame pictures (I-pictures) are pictures that can be encoded and decoded without using any other pictures in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, independent decoder refresh (IDR) pictures.
[0073] Predictive images (P-images) can be images that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses motion vectors and reference indices to predict sample values for each block.
[0074] Bidirectional predictive images (B-images) can be images that can be encoded and decoded using intra-frame or inter-frame prediction, which uses two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.
[0075] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined by the coding assignment of the corresponding images applied to the block. For example, a block of an I-image can be non-predictively coded, or it can be predictively coded with reference to already coded blocks of the same image (spatial prediction or intra-frame prediction). A pixel block of a P-image can be predictively coded with reference to a previously coded reference image via spatial prediction or temporal prediction. A block of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial prediction or temporal prediction.
[0076] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0077] In one aspect, the transmitter (340) may transmit additional data while transmitting encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0078] The captured video can be presented as multiple source images (video images) in a time-series format. Intra-frame image prediction (often simply called intra-prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, a specific image being encoded / decoded is segmented into blocks; this specific image being encoded / decoded is called the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference image.
[0079] In some respects, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image that precede the current image in the video in decoding order (but may be past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted using a combination of the first and second reference blocks.
[0080] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.
[0081] According to some aspects of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video picture sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU consists of three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Further, each CTU can be split into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, CUs are split into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking the luma prediction block as an example, a prediction block consists of a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and so on.
[0082] It should be noted that any suitable technology can be used to implement the video encoder (103) and video encoder (303), as well as the video decoder (110) and video decoder (210). In one aspect, one or more integrated circuits can be used to implement the video encoder (103) and video encoder (303), as well as the video decoder (110) and video decoder (210). In another aspect, one or more processors that execute software instructions can be used to implement the video encoder (103) and video encoder (303), as well as the video decoder (110) and video decoder (210).
[0083] In various examples, the current encoded frame (also referred to as the current frame or current picture) can be used as a reference region for block-based compensation, such as in intra-block copying, intra-template matching prediction (IntraTMP) mode, etc.
[0084] In one respect, block-based compensation from different images can be referred to as motion compensation. Similarly, block compensation can be performed from previously reconstructed regions within the same image, which can include intra-frame picture block compensation (also known as current picture referencing (CPR) or IBC mode). Figure 4 An example of intra-frame block compensation (e.g., IBC mode) according to one aspect of this disclosure is shown. The displacement vector indicating the offset between the current block (430) and the reference block (440) may be referred to as the block vector (BV) (450). The current block (430) and the reference block (440) are located in the current frame (400).
[0085] Unlike motion vectors (MV) in motion compensation, which can have arbitrary values (positive or negative in the x or y direction), BVs may be subject to certain restrictions to ensure that the reference block they point to is available and has been reconstructed. In one example, the reference... Figure 4 The current image (400) may include the region to be decoded (420) and the reconstructed region (410). In one example, the BV may be restricted to pointing to a reference block in the reconstructed region (410). In some examples, for parallel processing considerations, some reference regions located on tile boundaries or wavefront trapezoidal boundaries may be excluded.
[0086] The encoding of BV can be explicit or implicit. In explicit mode (known as advanced motion vector prediction (AMVP) mode in inter-frame coding), the difference between BV and the predicted BV value can be represented by a signal. In implicit mode, BV can be fully recovered from the predicted BV value in a manner similar to recovering MV in merge mode. In some implementations, the precision of BV can be limited to integer positions; in other systems, the precision of BV can be allowed to point to fractional positions.
[0087] A block-level flag (called the IBC flag) can be used to signal the use of intra-block copying at the block level. In one aspect, the IBC flag is signaled when the current block is not encoded in merge mode. In one example, the IBC flag can be signaled via a reference indexing method (e.g., by treating the currently decoded picture as the reference picture). In one example, such as in HEVC SCC, this reference picture (e.g., the currently decoded picture) is placed at the end of a list (e.g., a list of reference pictures). This particular reference picture (e.g., the currently decoded picture) can be managed along with other temporal reference pictures in the decoded picture buffer (DPB).
[0088] Intra-block copying can have several variations, such as considering it as a third mode distinct from intra-prediction or inter-prediction modes. By treating intra-block copying as a third mode, block vector prediction in merge and AMVP modes can be distinguished from regular inter-frame modes. In one example, the explicit mode described above could be called the IBC AMVP mode, and the implicit mode could be called the IBC merge mode. For example, a separate merge candidate list is defined for the IBC mode (e.g., the IBC merge mode), where all entries are BV. Similarly, in one example, the block vector prediction list in the IBC AMVP mode consists only of BV. In some examples, the general rules applied to these two lists include that, in terms of the candidate derivation process, these two lists can follow the same logic as the inter-merge candidate list used in inter-frame modes or the AMVP prediction list used in inter-frame modes. For example, the five spatially adjacent positions in an inter-merge mode, such as HEVC or VVC inter-merge mode, can be accessed to derive the merge candidate list for the IBC mode itself.
[0089] Figure 5 An example of intra-frame image patch compensation with a search range of one CTU size is shown according to one aspect of this disclosure, and in some examples, memory is reused to search a portion of the left CTU.
[0090] In some examples, such as VVC, the search range of the IBC pattern is limited to the current CTU. In one example, the effective memory requirement for storing reference samples for the IBC pattern is one CTU-sized sample. Considering that the existing reference sample memory is used to store reconstructed samples in the current 64×64 region, three additional 64×64-sized reference sample memories can be used. Therefore, a method can be used to extend the effective search range of the IBC pattern to a portion of the left CTU while the total memory requirement for storing reference pixels remains unchanged; for example, the total memory requirement is one CTU size, such as a total of four 64×64-sized reference sample memories. Figure 5 An example of this memory reuse mechanism is shown. Each vertical stripe block represents the current coding region (Curr), and the samples in each gray area are encoded samples. The crossed-out areas (marked with "X") cannot be used as references because these crossed-out areas may be replaced by the coding regions in the current CTU in the reference sample memory.
[0091] Figure 6 An example of an IntraTMP mode according to one aspect of this disclosure is shown. In one example, the IntraTMP mode is a special type of intra-prediction mode. In another example, the IntraTMP mode differs from the intra-prediction mode. (See also...) Figure 6 In the IntraTMP mode example, a prediction block (621) can be copied, such as the best prediction block from the reconstructed portion of the current frame. A template (620), such as an L-shaped template of the prediction block (621), can be matched with the current template (630) of the current block (631). For a predefined search range, the encoder can search for the template most similar to the current template (630) in the reconstructed portion of the current frame and can use the corresponding block (621) as the prediction block. In one example, the encoder then signals the use of IntraTMP mode and performs the same prediction operation on the decoder side.
[0092] refer to Figure 6 A predicted signal can be generated by matching the L-shaped causal neighboring blocks of the current block (631) with another block in a predefined search region. In one example, the predefined search region comprises or consists of the following parts: R1 (the current CTU), R2 (the upper-left CTU of the current CTU), R3 (the CTU above the current CTU), and R4 (the CTU to the left of the current CTU). In one example, the sum of absolute differences (SAD) is used as the cost function.
[0093] Within each region, the decoder can search for a template with the minimum cost (e.g., minimum SAD) relative to the current template, and can use the block corresponding to the minimum cost as the prediction block.
[0094] The dimensions of all regions (SearchRange_w, SearchRange_h) can be set to be proportional to the block size (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons. In one example, SearchRange_w = a × BlkW, SearchRange_h = a × BlkH, where "a" is a constant that balances gain and complexity. In one example, "a" equals 5.
[0095] In some examples, to speed up the template matching process, the search range of all search regions is downsampled, for example, by a downsampling factor of 2, thus reducing the template matching search to one-quarter of its original size. After finding the best match, refinement can be performed. This refinement can be accomplished by performing a second template matching search around the best match using the reduced range. In one example, the reduced range is defined as min(BlkW, BlkH) / 2.
[0096] The intra-template matching tool can be applied to CUs with a width and height both less than or equal to 64. The maximum CU size in IntraTMP mode can be configurable. For example, when decoder-side intramode derivation (DIMD) is not used for the current CU, IntraTMP mode can be notified at the CU level by a dedicated flag.
[0097] In some examples, such as some ECM applications, the BV candidates used in the BV candidate list in IBC modes (e.g., IBC skip mode, IBC merge mode, or IBC AMVP mode), IntraTMP mode, etc., can be derived from spatial candidates, hash-based black vector prediction (HBVP) candidates, pairwise average candidates (e.g., combining spatial and HBVP candidates), and some predefined BV candidates (e.g., (0,0)). Temporal BV can be used in some related methods. However, in some examples, temporal MV and temporal BV are not used to derive BV candidates from the BV candidate list.
[0098] This disclosure provides techniques, apparatus, and methods related to constructing BV candidates using temporal candidates (e.g., temporal MV, temporal BV, etc.) in a BV candidate list, for example. Using temporal candidates (e.g., temporal MV or temporal BV) to determine one or more BV candidates in a BV candidate list can improve the prediction accuracy of BV. In one example, using temporal candidates to determine one or more BV candidates in a BV candidate list can increase the diversity of the BV candidate list, thereby giving the encoder more BV candidate choices and resulting in more accurate BV candidates. In one aspect, the techniques, apparatus, and methods described in this disclosure can be used in IBC mode or other prediction modes that utilize the current image (also called the current coded frame) as a motion compensation reference region, such as IntraTMP mode.
[0099] The methods, aspects, and examples described in this disclosure can be used individually or in combination in any order. The term "IBC mode" may refer to the IBC mode or a variant thereof described in this disclosure. The term "IntraTMP mode" may refer to the IntraTMP mode or a variant thereof described in this disclosure. Furthermore, these methods, aspects, and examples can be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.
[0100] In one aspect, a prediction mode can be used to predict the current block in the same current image based on a reference block in a previously reconstructed region of the current image. The offset between the current block and the reference block can be called the Block Value (BV). The prediction mode can be an IBC mode or a variant of the IBC mode, an IntraTMP mode or a variant of the IntraTMP mode, etc. In various examples, to predict the current block, a candidate list (e.g., a BV candidate list) including one or more BV candidates can be constructed under the prediction mode (e.g., IBC mode, IntraTMP mode, etc.), and a BV candidate can be selected from the BV candidate list to determine the BV of the current block. The reference block can be determined based on the BV, and the current block can be predicted based on the reference block.
[0101] According to one aspect of this disclosure, the current block is predicted using the aforementioned prediction mode (e.g., one of the IBC mode and the IntraTMP mode). The first temporal BV candidate in the BV candidate list for the current block can be determined based on one of (i) motion information and (ii) BV information associated with a co-located block (also called a co-located block) in a co-located image (also called a co-located image) of the current image. Based on the first temporal BV candidate, a BV candidate list can be constructed. Using one of the IBC mode and the IntraTMP mode, the current sample can be reconstructed based on the BV candidate list.
[0102] In one example, the co-location image has already been reconstructed. One of the following information associated with the co-location block, (i) motion information and (ii) BV information, is used to reconstruct samples in regions within the co-location block or in regions adjacent to the co-location block.
[0103] In one aspect, motion information associated with a co-located block in a co-located image can be termed temporal motion information. Temporal motion information may include or indicate the time (MV) pointing to a reference block in a reference image, a reference index indicating the reference image, etc.
[0104] In one aspect, the BV information associated with a co-located block in a co-located image can be referred to as temporal BV information. Temporal BV information may include or indicate the temporal BV pointing to (e.g., in the co-located image) a reference block.
[0105] In one aspect, temporal motion information (e.g., including or indicating time MV) or temporal BV information (e.g., including or indicating time BV) can be used to derive BV candidates for the current block. BV candidates derived from temporal motion information or temporal BV information can be referred to as temporal BV candidates, such as the first temporal BV candidate described above. In some examples, temporal BV candidates can be derived by obtaining motion information (or temporal motion information) or BV information (or temporal BV information) from co-op blocks in a co-op picture (e.g., a co-op picture notified by a signal). In one example, the current picture is encoded in an inter-frame slice. The co-op picture can be a reference picture in a reference list of the current picture. The encoder can, for example, determine the co-op picture as the first reference picture closest to the current picture in a first reference list (e.g., L1) based on the picture order count (POC) difference.
[0106] In one aspect, inferred motion information or inferred BV information can be obtained from a specific location (also called a candidate location) of a co-location block in a co-location image (e.g., a co-location image notified by a signal). For example, the inferred motion information or inferred BV information is associated with a specific location of the co-location block. In one example, a first-time BV candidate is determined based on one of (i) motion information and (ii) BV information associated with a candidate location of the co-location block (e.g., a first candidate location).
[0107] Figure 7 An example of the candidate position of a time candidate (e.g., a time BV candidate) according to one aspect of this disclosure is shown. When time candidates are used in IBC skip mode, IBC merge mode, etc., time candidates such as time BV candidates can be referred to as time merge candidates.
[0108] For example, such as Figure 7As shown, the specific position or candidate position of the corresponding block (700) can be the center position (701) of the corresponding block (700), or the lower right position (702) of the corresponding block (700), etc. The center position (701) of the time candidate or the lower right position (702) of the time candidate (e.g., the time BV candidate) can be determined by... Figure 7 The corresponding gray blocks in the image are used to represent the region. In one example, one of the motion information and BV information used to reconstruct samples in block (711) (e.g., the region within the co-located block (700) represented by the gray blocks) is either the motion information or the BV information associated with the candidate location (701), and can be used to derive the temporal BV candidate.
[0109] As described above, the sibling block (700) is located in the sibling image of the current image. The sibling block (700) can be in the same position as the current block. For example, if the current block is located at (x0, y0) in the current image, then the sibling block can be located at (x0, y0) in the sibling image.
[0110] In one respect, the temporal motion information of co-located blocks in a co-located image can be scaled. Figure 8 An example of MV scaling for a time BV candidate (e.g., a time merging candidate) according to one aspect of this disclosure is shown. The current block (801) is located in the current picture (821). The co-located block (802) of the current block (801) is located in the co-located picture (822) of the current picture (821). A time BV candidate indicating a BV (812) can be derived based on the temporal motion information associated with the candidate position of the co-located block (802). The temporal motion information associated with the candidate position of the co-located block (802) may include an MV (811), which, for example, points from the co-located picture (822) to a reference picture (823) of the co-located picture (822). In one example, the derived motion information may be a scaled MV as described below. The scaled MV of the time BV candidate (e.g., a time merging candidate) may be as follows... Figure 8 The dashed lines in the diagram indicate the scaling factor (MV). In one example, the scaling MV can be determined (e.g., scaling) based on the MV (811) of the co-location block (802) (e.g., the co-location CU). The first MV (811) can be defined as the MV difference between the reference image (e.g., the co-location image (822)) and the current image (821). The second MV (822) can be defined as the MV difference between the reference image (823) and the co-location image (822) of the co-location image (or the co-location image (822)). In one example, the scaling MV is BV (812).
[0111] In one example, refer to Figure 8Based on the first POC difference t between the MV (811), the current image (821), and the co-position image (822), and the second POC difference T between the reference image (823) of the co-position image (822) and the co-position image (822), the scaled MV, such as the BV (812), can be determined. In one example, the first-time BV candidate of the current block (801) includes the scaled MV (812).
[0112] In one aspect, more than one location or multiple candidate locations can be used to derive BV candidates, such as temporal BV candidates. Based on multiple motion information derived from the corresponding multiple candidate locations, such as the average of multiple motion information derived from the corresponding multiple candidate locations, a derived BV candidate or a derived temporal BV candidate (e.g., a first temporal BV candidate) can be determined. In one example, a first temporal BV candidate is determined by averaging multiple motion information from various candidate locations associated with the co-location block.
[0113] In one example, the derived BV candidate or the derived temporal BV candidate (e.g., the first temporal BV candidate) can be the average of multiple BV information from various candidate locations associated with the co-block.
[0114] In one aspect, more than one candidate location can be used to generate a time candidate list. Since a time candidate list includes one or more time BV candidates, it can also be called a time BV candidate list. For example, a time candidate list can be used to construct a BV candidate list based on the order in which the BV candidate lists are constructed.
[0115] In one example, a time candidate list is generated, which includes time BV candidates associated with each candidate position of the co-location block. Each time BV candidate may be based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-location block. The time BV candidates include the first time BV candidate described above.
[0116] In one example, five candidate positions of the same block are associated with corresponding motion information and / or BV information, thus yielding five corresponding temporal BV candidates. The temporal BV candidate list can include five temporal BV candidates (e.g., candidates 1 to 5).
[0117] In one aspect, a template-matching (TM) process or TM can be applied to reorder the list of temporal candidates, for example, by using the TM costs of the individual temporal BV candidates. In one example, a reference block in the current image is determined based on each temporal BV candidate (e.g., BV). The TM cost is obtained based on the current template of the current block and the reference template of the reference block. The current template may include neighboring reconstructed samples of the current block. In one example, the current template may include upper reconstructed samples located above the current block and / or left reconstructed samples located to the left of the current block.
[0118] In one example, a Time Management (TM) is applied to determine the TM cost of each time BV candidate in the time candidate list. Based on the individual TM costs, the time BV candidates in the time candidate list can be reordered, and a BV candidate list can be constructed based on the reordered time candidate list. In one example, the first m time candidates from the reordered time BV candidates can be selected and added to the BV candidate list, where m is a positive integer less than or equal to the size of the time candidate list (e.g., the total number of time BV candidates in the time candidate list).
[0119] In one example, when the time BV candidate list includes five time BV candidates (e.g., candidates 1 through 5), five TM costs can be obtained. In one example, the time candidate list is reordered in ascending order based on the five TM costs. In one example, candidates 1 through 5 can be ordered as candidate 5, candidate 3, candidate 1, candidate 2, candidate 4. If m=2, then candidate 5 and candidate 3 are added to the BV candidate list.
[0120] In one aspect, the TM process can be applied to a time candidate list to select the best n1 candidates with one or more minimum TM costs. n is a positive number. In one example, n is less than or equal to the size of the time candidate list.
[0121] In one example, a time-based time management (TM) approach is applied to determine the TM cost of each time-based BV candidate in the time-candidate list. Based on the TM cost, n time-based BV candidates can be selected, where the TM cost of the n time-based BV candidates is no greater than the TM cost of each remaining time-based BV candidate in the time-candidate list. Based on the n selected time-based BV candidates, a BV candidate list can be constructed.
[0122] In one aspect, based on the order in which the BV candidate list is constructed, one or more temporal BV candidates derived from motion information (e.g., including MV) of co-located images can be used to construct the BV candidate list.
[0123] In one aspect, a TM process can be applied to determine the TM cost of each BV candidate in the initial BV candidate list, where the BV candidates include the first-time BV candidates. The BV candidate list can be constructed by reordering the BV candidates in the initial BV candidate list based on their respective TM costs.
[0124] In one example, a TM procedure can be applied to reorder a BV candidate list (e.g., an initial BV candidate list) using TM costs (e.g., in ascending order). In one example, the BV candidate list includes BV candidates, where each BV candidate may include one or more temporal BV candidates and / or one or more non-temporal BV candidates. The one or more non-temporal BV candidates may be derived from one or more spatially adjacent blocks in the current image, or the one or more non-temporal BV candidates may include one or more predefined BV candidates. In one example, the one or more non-temporal BV candidates include one or more spatial candidates, one or more HBVP candidates, one or more pairwise average candidates, one or more predefined BV candidates, etc. The TM cost of each BV candidate can be determined using the TM procedure described above. In one example, the BV candidate list (and the BV candidates) are reordered in ascending order based on the TM costs.
[0125] In one example, the TM process is applied to the BV candidate list to select the best n² candidates with one or more minimum TM costs. n² is a positive number. In one example, n² is less than or equal to the size of the BV candidate list. In another example, the TM process is applied to determine the TM cost of each BV candidate in the BV candidate list. The BV candidates include the first-time BV candidates. Based on the TM costs, n² BV candidates can be selected. The TM cost of these n² BV candidates is no greater than the TM cost of each remaining BV candidate in the BV candidate list. Using either the IBC mode or the IntraTMP mode, the current sample can be reconstructed based on the n² selected BV candidates.
[0126] In one aspect, when constructing a list of BV candidates that includes first-time BV candidates, the current block is encoded based on this list. In one example, a BV candidate is selected from the BV candidates to encode the current block. One of the BV candidates can be a first-time BV candidate. Based on the first-time BV candidate, the BV of the current block can be determined. In one example, the BV of the current block is either a BV from the first-time BV candidate or a scaled MV. Once the BV is determined, the current block can be reconstructed based on it.
[0127] Figure 9A flowchart outlining a process (900) according to one aspect of this disclosure is shown. The process (900) can be used in an apparatus such as a video decoder. In various aspects, the process (900) is executed by processing circuitry, such as processing circuitry that performs the functions of a video decoder (110), processing circuitry that performs the functions of a video decoder (210), etc. In some aspects, the process (900) is implemented in software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (900). The process begins at step (S901) and proceeds to step (S910).
[0128] At step (S910), encoded information is received. This encoded information indicates the prediction of the current block in the current image based on a prediction mode, wherein the current block is predicted based on a reference block in a previously reconstructed region of the same current image.
[0129] In one aspect, the prediction pattern is one of the IBC pattern and the IntraTMP pattern. In one example, one of the IBC pattern and the IntraTMP pattern is the IBC pattern. One of the IBC pattern and the IntraTMP pattern is the IntraTMP pattern.
[0130] In one example, the current image is encoded in an inter-frame slice.
[0131] At step (S920), a first time BV candidate in the BV candidate list of the current block is determined based on one of (i) motion information and (ii) BV information associated with the co-occurring block in the co-occurring image of the current image.
[0132] In one aspect, a first-time BV candidate is determined based on one of (i) motion information and (ii) BV information associated with the first candidate position of the co-located block, such as Figure 7 As shown.
[0133] In one respect, reference Figure 8 One of the information, (i) motion information and (ii) BV information, is motion information and the motion information indicates a reference image (e.g., reference image (823)) corresponding to the MV (e.g., MV (811)) and the co-positioned image (e.g., co-positioned image (822)). Based on a first POC difference (e.g., first POC difference t) between the MV (e.g., MV (811)), the current image (e.g., (821)) and the co-positioned image (e.g., (822)), and a second POC difference (e.g., second POC difference T) between the reference image (e.g., (823)) and the co-positioned image (e.g., (822)), a scaled MV (e.g., BV (812)) is determined. In one example, the scaled MV is included as a first-time BV candidate.
[0134] In one respect, reference Figure 7 When one of the motion information and (ii) BV information is motion information associated with a co-located block, a first-time BV candidate is determined by averaging multiple motion information of each candidate position (e.g., candidate positions (701) to (702)) associated with the co-located block (e.g., (700)).
[0135] At step (S930), a BV candidate list is constructed based on the first-time BV candidate. The BV candidate list for the current block may include BV candidates, such as one or more time-based BV candidates and one or more non-time-based BV candidates as described above. The one or more time-based BV candidates include the first-time BV candidate.
[0136] In one aspect, a temporal candidate list is generated, comprising temporal BV candidates associated with each candidate position of the co-located block. The temporal BV candidates include a first temporal BV candidate. Each temporal BV candidate may be based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-located block. Based on the temporal candidate list, a BV candidate list is constructed.
[0137] In one example, a TM is applied to determine the TM cost of each time BV candidate in the time candidate list, the time BV candidates in the time candidate list are reordered based on each TM cost, and a BV candidate list is constructed based on the reordered time candidate list, for example, by adding the first m reordered time candidates from the reordered time candidate list to the BV candidate list as described above.
[0138] In one example, a Time Management (TM) algorithm is applied to determine the TM cost of each time BV candidate in the time candidate list. Based on the TM cost, n1 time BV candidates are selected. The TM cost of these n1 time BV candidates is no greater than the TM cost of each remaining time BV candidate in the time candidate list. Based on these n1 selected time BV candidates, a BV candidate list is constructed.
[0139] In one aspect, a TM is applied to determine the TM cost of each BV candidate in the initial BV candidate list. The BV candidates include the first-time BV candidates. A new BV candidate list is constructed by reordering the BV candidates in the initial BV candidate list based on their individual TM costs. In this example, both the initial BV candidate list and the new BV candidate list contain the same BV candidates that can be arranged in different orders.
[0140] In one aspect, a time-mapping mechanism (TM) is applied to determine the TM cost of each BV candidate in the BV candidate list. The BV candidates include the first-time BV candidates. Based on the TM cost, n² BV candidates are selected. The TM cost of these n² BV candidates is no greater than the TM cost of each remaining BV candidate in the BV candidate list. Using either the IBC mode or the IntraTMP mode, the current sample can be reconstructed based on these n² selected BV candidates.
[0141] At step (S940), the current sample is reconstructed based on the BV candidate list using either the IBC mode or the IntraTMP mode.
[0142] Then, the process proceeds to step (S999) and ends.
[0143] The process (900) can be adjusted as appropriate. One or more steps in the process (900) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.
[0144] Figure 10 A flowchart outlining a process (1000) according to one aspect of this disclosure is shown. The process (1000) can be used with a video encoder. In various aspects, the process (1000) is executed by processing circuitry, such as processing circuitry executing the functions of a video encoder (103), processing circuitry executing the functions of a video encoder (303), etc. In some aspects, the process (1000) is implemented in software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (1000). The process begins at step (S1001) and proceeds to step (S1010).
[0145] At step (S1010), a first temporal BV candidate is determined in the block vector (BV) candidate list of the current block in the current image based on one of (i) motion information and (ii) BV information associated with the co-block, wherein the co-block is located in a co-image of the current image. The current block may be predicted according to one of the IBC mode and the IntraTMP mode. In one example, one of (i) motion information and (ii) BV information is associated with the first candidate position of the co-block.
[0146] In one respect, reference Figure 8One of the information, (i) motion information and (ii) BV information, is motion information and the motion information indicates a reference image (e.g., reference image (823)) corresponding to the MV (e.g., MV (811)) and the co-positioned image (e.g., co-positioned image (822)). Based on a first POC difference (e.g., first POC difference t) between the MV (e.g., MV (811)), the current image (e.g., (821)) and the co-positioned image (e.g., (822)), and a second POC difference (e.g., second POC difference T) between the reference image (e.g., (823)) and the co-positioned image (e.g., (822)), a scaled MV (e.g., BV (812)) is determined. In one example, the scaled MV is included as a first-time BV candidate.
[0147] In one respect, reference Figure 7 When one of the motion information and (ii) BV information is motion information associated with a co-located block, a first-time BV candidate is determined by averaging multiple motion information of each candidate position (e.g., candidate positions (701) to (702)) associated with the co-located block (e.g., (700)).
[0148] At step (S1020), a BV candidate list is constructed based on the first-time BV candidate. The BV candidate list for the current block may include BV candidates, such as one or more time-based BV candidates and one or more non-time-based BV candidates as described above. The one or more time-based BV candidates include the first-time BV candidate.
[0149] In one aspect, a temporal candidate list is generated, comprising temporal BV candidates associated with each candidate position of the co-located block. The temporal BV candidates include a first temporal BV candidate. Each temporal BV candidate may be based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-located block. Based on the temporal candidate list, a BV candidate list is constructed.
[0150] In one example, a TM is applied to determine the TM cost of each time BV candidate in the time candidate list, the time BV candidates in the time candidate list are reordered based on each TM cost, and a BV candidate list is constructed based on the reordered time candidate list, for example, by adding the first m reordered time candidates from the reordered time candidate list to the BV candidate list as described above.
[0151] In one example, a Time Management (TM) algorithm is applied to determine the TM cost of each time BV candidate in the time candidate list. Based on the TM cost, n1 time BV candidates are selected. The TM cost of these n1 time BV candidates is no greater than the TM cost of each remaining time BV candidate in the time candidate list. Based on these n1 selected time BV candidates, a BV candidate list is constructed.
[0152] In one aspect, a TM is applied to determine the TM cost of each BV candidate in the initial BV candidate list. The BV candidates include the first-time BV candidates. A new BV candidate list is constructed by reordering the BV candidates in the initial BV candidate list based on their individual TM costs. In this example, both the initial BV candidate list and the new BV candidate list contain the same BV candidates that can be arranged in different orders.
[0153] At step (S1030), the current sample is encoded based on the BV candidate list using either the IBC mode or the IntraTMP mode.
[0154] Then, the process proceeds to step (S1099) and ends.
[0155] Process (1000) can be adjusted as appropriate. One or more steps in process (1000) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.
[0156] In one aspect, a method for processing visual media data includes processing a bitstream of visual media data according to format rules. For example, the bitstream may be a bitstream decoded / encoded using any of the decoding and / or encoding methods described herein. The format rules may specify one or more constraints on the bitstream and / or one or more processes to be performed by the decoder and / or encoder.
[0157] In one example, the bitstream includes encoded information indicating the prediction of the current block in the current image based on either Intra-Block Copy (IBC) mode or Intra-Template Matching Prediction (IntraTMP) mode. The format rule specifies that first-temporal block vector (BV) candidates are determined for a BV candidate list of the current block based on either (i) motion information or (ii) BV information associated with co-located blocks in co-located images of the current image. The format rule specifies that a BV candidate list is constructed based on the first-temporal BV candidates. The format rule specifies that the current sample is processed according to the BV candidate list using either IBC mode or IntraTMP mode.
[0158] In one example, one of (i) motion information and (ii) BV information is associated with the first candidate position of the co-located block.
[0159] In one example, the format rule specifies that when one of (i) motion information and (ii) BV information is motion information and the motion information indicates: motion vector (MV) and a reference image corresponding to the co-image, the scaled MV is determined based on the MV, the first image order count (POC) difference between the current image and the co-image, and the second POC difference between the reference image of the co-image and the co-image, and the scaled MV is included in the BV candidate at the first moment.
[0160] In one example, the format rule specifies the generation of a time candidate list, which includes time BV candidates associated with each candidate position of a co-located block. Each time BV candidate is based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-located block, and the time BV candidates include a first time BV candidate. TM is applied to determine the TM cost of each time BV candidate in the time candidate list. Based on n1 time BV candidates, a BV candidate list is constructed, wherein the TM cost of the n1 time BV candidates is no greater than one or more TM costs of the remaining time BV candidates in the time candidate list.
[0161] The aspects, examples, and / or methods disclosed herein may be used individually or in combination in any order. For example, some aspects and / or examples executed by the decoder may be executed by the encoder, and vice versa. Each method (or aspect), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.
[0162] The above-described techniques can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 11 A computer system (1100) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0163] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode, etc.
[0164] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things (IoT) devices.
[0165] Figure 11 The components shown for the computer system (1100) are exemplary and are not intended to impose any limitation on the scope or functionality of computer software implementing aspects of this disclosure. Nor should the configuration of the components be construed as having any dependency or requirement relating to any or a combination of the components shown in the exemplary aspects of the computer system (1100).
[0166] The computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users, for example, through input such as: tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images acquired from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0167] Human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), and camera (1108).
[0168] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (1110), data gloves (not shown), or joystick (1105), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (1109), headphones (not depicted)), and visual output devices (e.g., screens (1110) including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touchscreen input functionality, each screen may or may not have tactile feedback functionality, some of which are capable of outputting two-dimensional visual output or more than three-dimensional output via means such as stereoscopic output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0169] The computer system (1100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) or similar media (1121) with CD / DVD, finger drives (1122), removable hard disk drives or solid-state drives (1123), conventional magnetic media such as magnetic tape and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0170] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0171] The computer system (1100) may also include an interface (1154) leading to one or more communication networks (1155). The network may be, for example, a wireless network, a wired network, or a fiber optic network. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, and so on. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port on the computer system (1100)) to connect to some general-purpose data port or peripheral bus (1149); other network interfaces are typically integrated into the core of the computer system (1100) by connecting to a system bus as described below (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system (1100) can use any of these networks to communicate with other entities. This communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as using a local area network (LAN) or wide area network (WAN) to connect to other computer systems. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.
[0172] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel (1140) of the computer system (1100).
[0173] The core (1140) may include one or more central processing units (CPU) (1141), graphics processing units (GPUs) (1142), dedicated programmable processing units in the form of field-programmable gate arrays (FPGAs) (1143), hardware accelerators (1144) for certain tasks, graphics adapters (1150), and so on. These devices, as well as read-only memory (ROM) (1145), random access memory (1146), and internal mass storage (1147) such as internal non-user-accessible hard disk drives, SSDs, etc., can be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of one or more physical plugs to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (1148) or via a peripheral bus (1149). In one example, a display (1110) may be connected to a graphics adapter (1150). Peripheral bus architectures include PCI, USB, etc.
[0174] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (1145) or RAM (1146). Transient data can also be stored in RAM (1146), while permanent data can be stored, for example, in internal mass storage (1147). Fast storage and retrieval of any storage device can be achieved by using a cache, which can be closely associated with one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc.
[0175] Computer-readable media may have computer code thereon that performs various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.
[0176] As a non-limiting example, a computer system (1100) having an architecture, particularly a kernel (1140), can provide functionality by having one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as some non-transitory memory of the kernel (1140), such as internal kernel mass storage (1147) or ROM (1145). Software implementing various aspects of this disclosure can be stored in such a device and executed by the kernel (1140). Depending on specific needs, the computer-readable media may include one or more storage devices or chips. The software can cause the kernel (1140), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1146) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system is enabled to provide functionality by means of hard-wired or otherwise embodied logic in the circuit (e.g., the accelerator (1144)), which may replace or operate with the software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0177] As used in this disclosure, "at least one of" or "one of" is intended to include any one or a combination of the elements. For example, at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include: A alone, B alone, C alone, or any combination thereof. One of A or B and one of A and B are intended to include: A or B or (A and B). Where applicable, such as when the elements are not mutually exclusive, "one of" does not exclude any combination of the listed elements.
[0178] Although examples of various aspects have been described in this disclosure, modifications, substitutions, and various alternative equivalents still exist within the scope of this disclosure. Therefore, it should be understood that those skilled in the art can design many systems and methods that, although not explicitly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.
Claims
1. A method for processing visual media data, the method comprising: The bitstream containing visual media data, including the current image, is processed according to format rules. The bitstream includes syntax elements that indicate the prediction of the current block in the current image according to one of Intra-Block Copy (IBC) mode and Intra-Template Match Prediction (IntraTMP) mode, and The format rules stipulate that: Based on one of (i) motion information and (ii) block vector (BV) information associated with the first candidate position of a co-located block in a co-located image, a first temporal BV candidate is determined in the BV candidate list of the current block. When one of the motion information (i) and the BV information (ii) is motion information and the motion information indicates: a motion vector (MV) and a reference image corresponding to the co-located image, a scaled MV is determined based on the MV, a first image sequence count (POC) difference between the current image and the co-located image, and a second POC difference between the reference image and the co-located image, wherein the first time-based BV candidate includes the scaled MV. The corresponding block is located in the corresponding image of the current image; Based on the first time-based BV candidates, the BV candidate list is constructed; and The current sample is processed based on the BV candidate list using one of the IBC mode and the IntraTMP mode.
2. The method according to claim 1, wherein, The formatting rules also specify: Generate a time candidate list, the time candidate list including time BV candidates associated with each candidate position of the co-position block, each time BV candidate based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-position block, the time BV candidate including the first time BV candidate; as well as Template matching (TM) is applied to determine the TM cost of each time BV candidate in the time candidate list, wherein the BV candidate list is constructed based on n1 time BV candidates, and the TM cost of the n1 time BV candidates is no greater than one or more TM costs of the remaining time BV candidates in the time candidate list.
3. A video encoding method, the method comprising: Based on one of (i) motion information and (ii) block vector (BV) information associated with the first candidate position of the co-op block, a first temporal BV candidate is determined in the BV candidate list of the current block in the current image, the co-op block being located in the co-op of the current image, the current block being predicted according to one of the intra-block copy (IBC) mode and the intra-template matching prediction (IntraTMP) mode; Based on the first time-based BV candidates, the BV candidate list is constructed; as well as The current sample is encoded based on the BV candidate list using one of the IBC mode and the IntraTMP mode.
4. The method according to claim 3, wherein, (i) the motion information and (ii) the BV information, one of which is motion information and indicates: a motion vector (MV) and a reference image corresponding to the co-position image; The method includes: determining a scaled MV based on the MV, a first image sequence count (POC) difference between the current image and the co-located image, and a second POC difference between the reference image of the co-located image and the co-located image; and The first time-based BV candidate includes the scaled MV.
5. The method according to claim 3, further comprising: When one of the motion information (i) and the BV information (ii) is the motion information associated with the first candidate position, the first time BV candidate is determined by averaging multiple motion information of each candidate position of the co-position block, wherein the multiple motion information includes the motion information associated with the first candidate position. as well as When one of the motion information (i) and the BV information (ii) is the BV information associated with the first candidate position, the first time BV candidate is determined by averaging multiple BV information at each candidate position of the co-position block, wherein the multiple BV information includes the BV information associated with the first candidate position.
6. The method according to claim 3, further comprising: Generate a time candidate list, the time candidate list including time BV candidates associated with each candidate position of the co-location block, wherein each time BV candidate is based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-location block, the time BV candidate including the first time BV candidate; and Based on the time candidate list, construct the BV candidate list.
7. The method according to claim 6, wherein, Constructing the BV candidate list includes: constructing the BV candidate list by applying template matching (TM) to the time candidate list.
8. The method according to claim 3, wherein, The method further includes: applying template matching (TM) to determine the TM cost of each BV candidate in the initial BV candidate list, wherein the BV candidates include the first-time BV candidates; and The construction includes: constructing the BV candidate list by reordering the BV candidates in the initial BV candidate list based on each TM cost.
9. A video decoding apparatus, the apparatus comprising: Processing circuit, the processing circuit being configured to: Receive encoded information, the encoded information indicating: predict the current block in the current image according to one of the Intra-Block Copy (IBC) mode and the Intra-Template Match Prediction (IntraTMP) mode; Based on one of (i) motion information and (ii) block vector (BV) information associated with a co-located block in a co-located image of the current image, a first-time BV candidate is determined in the BV candidate list of the current block; Based on the first time-based BV candidates, the BV candidate list is constructed; as well as The current sample is reconstructed based on the BV candidate list using one of the IBC mode and the IntraTMP mode.
10. The apparatus according to claim 9, wherein, The processing circuit is configured as follows: The first time BV candidate is determined based on one of the information (i) the motion information and (ii) the BV information associated with the first candidate position of the co-position block.
11. The apparatus according to claim 9 or 10, wherein, (i) the motion information and (ii) the BV information, one of which is motion information and indicates: a motion vector (MV) and a reference image corresponding to the co-position image; The processing circuit is configured to: determine the scaled MV based on the MV, a first image sequence count (POC) difference between the current image and the co-located image, and a second POC difference between the reference image and the co-located image; and The first time-based BV candidate includes the scaled MV.
12. The apparatus according to claim 9, wherein, The processing circuit is configured as follows: When one of the motion information (i) and the BV information (ii) is motion information associated with the co-position block, the first time BV candidate is determined by averaging multiple motion information of each candidate position associated with the co-position block; as well as When one of the motion information (i) and the BV information (ii) is the BV information associated with the co-position block, the first time BV candidate is determined by averaging multiple BV information at each candidate position associated with the co-position block.
13. The apparatus according to claim 9, wherein, The processing circuit is configured as follows: Generate a time candidate list, the time candidate list including time BV candidates associated with each candidate position of the co-position block, each time BV candidate being based on one of (i) motion information and (ii) BV information associated with the corresponding candidate position of the co-position block, the time BV candidate including the first time BV candidate; as well as Based on the time candidate list, construct the BV candidate list.
14. The apparatus according to claim 13, wherein, The processing circuit is configured as follows: The BV candidate list is constructed by applying template matching (TM) to the time candidate list.
15. The apparatus according to claim 9, wherein, The processing circuit is configured as follows: Template matching (TM) is applied to determine the TM cost of each BV candidate in the initial BV candidate list, which includes the first-time BV candidate; as well as The BV candidate list is constructed by reordering the BV candidates in the initial BV candidate list based on each TM cost.