Video coding and decoding method and device and storage medium

By using intra-block copying and string matching patterns in video encoding, and using information of spatial adjacent blocks and non-adjacent blocks for prediction, the problem of low coding efficiency in intra-block blocks in the prior art is solved, and more efficient video encoding is achieved.

CN120034661APending Publication Date: 2025-05-23TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510175563.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-06-01
Filing Date
2021-06-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the existing video encoding technology processes intra-frame picture blocks, it is difficult to effectively utilize the information of space neighboring blocks, resulting in low encoding efficiency.

Method used

Intra-block copy (IBC) mode and string matching mode are used to improve coding efficiency by using the block vectors and string offset vectors of spatially adjacent blocks and non-adjacent blocks in intra-picture blocks.

Benefits of technology

By effectively utilizing the information of spatial neighboring blocks and non-neighboring blocks, the encoding efficiency of intra-frame picture blocks is improved, the number of bit streams is reduced, and the performance of video encoding is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034661A_ABST
    Figure CN120034661A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes processing circuitry. A processing circuit decodes prediction information from a current block in an encoded video bitstream. The prediction information indicates an encoding mode in which the current block is encoded on the basis of reconstructed samples in the same picture as the current block. The processing circuitry determines a candidate vector for reconstructing at least a portion of a block having a predetermined spatial relationship with the current block, and then determines a displacement vector for reconstructing at least a portion of the current block from a candidate list including the candidate vector. Further, the processing circuitry reconstructs the at least a portion of the current block based on the displacement vector.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references

[0001] This application claims priority to U.S. Patent Application No. 17 / 335,880, filed on June 1, 2021, for "METHOD AND APPARATUS FOR VIDEO CODING," which claims priority to U.S. Provisional Application No. 63 / 093,520, filed on October 19, 2020, for "Spatial Displacement Vector Prediction for Intra-Picture Blocks and String Replication (Methods on Constraint of Chroma Quad Tree Split)." The entire disclosures of these prior applications are incorporated herein by reference in their entirety. Technical Field

[0002] This application describes embodiments that generally relate to video encoding and decoding. Background Art

[0003] The background description provided herein is intended to present the background of the present application as a whole. The extent to which the work of the presently named inventors described in the background section and various aspects of this specification is performed does not indicate that it is prior art at the time of filing this application, and it is never explicitly or implicitly admitted that it is prior art for this application.

[0004] Video encoding and decoding are enabled by inter-picture prediction techniques with motion compensation. An uncompressed digital video may include a series of pictures, each picture having spatial dimensions such as 1920×1080 luminance samples and associated chrominance samples. The series of pictures has a fixed or variable picture rate (also informally referred to as a frame rate), such as 60 pictures per second or 60Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution, 60Hz frame rate) with 8 bits per sample requires close to 1.5Gbit / s bandwidth. One hour of such video requires more than 600GB of storage space.

[0005] One goal of video encoding and decoding is to reduce redundant information in the input video signal through compression. Video compression can help reduce the bandwidth and / or storage space requirements mentioned above, in some cases by two or more orders of magnitude. Both lossless and lossy compression, as well as combinations of the two, can be used. Lossless compression refers to techniques that reconstruct an exact replica of the original signal from the compressed original signal. When lossy compression is used, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of distortion allowed depends on the application. For example, users of some consumer streaming applications may be able to tolerate higher distortion than users of television applications. The achievable compression ratio reflects that higher allowed / tolerant distortion results in higher compression ratios.

[0006] Video encoders and decoders may utilize techniques from several broad categories including, for example, motion compensation, transforms, quantization, and entropy coding and decoding.

[0007] Video codec techniques may include techniques known as intra-frame coding. In intra-frame coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all sample blocks are encoded in intra-frame mode, the picture may be an intra-frame picture. Intra-frame pictures and their derivatives (such as independent decoder refresh pictures) can be used to reset the decoder state, and can therefore be used as the first picture in an encoded video stream and video session, or as a still image. Samples of intra-frame blocks can be exposed to transformation, and transform coefficients can be quantized before entropy coding and decoding. Intra-frame prediction can be a technique for minimizing sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after transformation, and the smaller the AC coefficient, the fewer bits are required to represent the entropy coded block at a given quantization step size.

[0008] Conventional intra-frame codecs, such as those known from, for example, the MPEG-2 generation of codecs, do not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to use intra-frame prediction from, for example, surrounding sample data and / or metadata that is obtained during encoding / decoding of spatially adjacent data blocks and precedes the data blocks in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and not reference data from reference pictures.

[0009] There can be many different forms of intra-frame prediction. When more than one such technique can be used in a given video codec technique, the technique used can be encoded in the intra-frame prediction mode. In some cases, a mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be encoded separately or included in the mode codeword. The codeword to be used for a given mode, sub-mode, and / or parameter combination may have an impact on the coding efficiency gain through intra-frame prediction, and the same is true for the entropy coding and decoding technique that converts the codeword into a bitstream.

[0010] Some mode of intra prediction was introduced with H.264, improved in H.265, and further improved in newer codecs such as joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). The predictor block can be formed using neighboring sample values ​​belonging to samples that are already available. Sample values ​​of neighboring samples are copied into the predictor block according to the direction. A reference to the direction in use can be encoded in the bitstream, or it can be predicted itself.

[0011] Referring to FIG. 1A , a subset of nine known prediction directions from the 33 possible prediction directions of H.265 (corresponding to the 33 angular modes of the 35 intra modes) is depicted at the bottom right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is being predicted. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at an angle of 45 degrees to the horizontal direction at the upper right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at an angle of 22.5 degrees to the horizontal direction at the lower left.

[0012] Still referring to FIG. 1A , a square block (104) comprising 4×4 samples is shown at the upper left (indicated by the thick dashed line). The square block (104) comprises 16 samples, each of which is labeled with “S”, and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the X and Y dimensions. Since the block is a 4×4 size sample, S44 is located in the lower right corner. Reference samples following a similar numbering scheme are also shown. The reference samples are labeled with “R”, and their Y position (e.g., row index) and X position (e.g., column index) relative to the block (104). In H.264 and H.265, the prediction samples are adjacent to the block being reconstructed, so there is no need to use negative values.

[0013] Intra-picture prediction can be performed by copying reference sample values ​​from neighboring samples occupied by a signaled prediction direction. For example, assume that the encoded video bitstream includes signaling that indicates, for this block, a prediction direction consistent with arrow (102), i.e., predicting samples based on one or more prediction samples to the upper right at a 45 degree angle to the horizontal. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Sample S44 is predicted based on reference sample R08.

[0014] In some cases, for example by interpolation, the values ​​of multiple reference samples may be combined in order to calculate the reference sample, particularly when the direction is not divisible by 45 degrees.

[0015] As video coding techniques have evolved, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013) and JEM / VVC / BMS, and at the time of this application, up to 65 directions can be supported. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those possible directions using a small number of bits, accepting some penalty for less likely directions. Additionally, the direction itself can sometimes be predicted based on neighboring directions used in adjacent, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram ( 105 ) depicting 65 intra prediction directions according to JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of the intra-frame prediction direction bits representing directions in the coded video bitstream can vary depending on the video coding technique, and can range from simple direct mapping of prediction directions to intra-frame prediction modes to codewords, to complex adaptive schemes involving the most probable mode, and similar techniques, for example. However, in all cases, there may be certain directions in the video content that are statistically less likely to occur than certain other directions. Since the goal of video compression is to reduce redundancy, in well-working video coding techniques, those unlikely directions will be represented by a larger number of bits than more likely directions.

[0018] Motion compensation may be a lossy compression technique and may involve a technique where a block of sample data from a previously reconstructed picture or part of a reconstructed picture (reference picture) is used for prediction of a newly reconstructed picture or part of a picture after being spatially shifted in the direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, where the third dimension indicates the reference picture in use (the latter may indirectly be the temporal dimension).

[0019] In some video compression techniques, the MV applied to a region of sample data can be predicted based on other MVs, such as those MVs associated with another region of sample data that is spatially adjacent to the region being reconstructed and that precede the MV in decoding order. Doing so can greatly reduce the amount of data required to encode the MV, thereby eliminating redundant information and increasing the amount of compression. MV prediction can be done effectively, for example, when encoding an input video signal derived from a camera (called natural video), there is a statistical probability that regions larger than the region to which a single MV applies will move in a similar direction, so in some cases, similar motion vectors derived from MVs of neighboring regions can be used for prediction. This results in the MV found for a given region being similar or identical to the MV predicted from the surrounding MVs, and after entropy coding, it can be represented with fewer bits than the number of bits used when the MV is directly encoded. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself may be lossy, for example due to rounding errors when calculating predicted values ​​based on several surrounding MVs.

[0020] H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016) describes various MV prediction mechanisms. Among the various MV prediction mechanisms provided by H.265, this article describes a technique hereinafter referred to as "spatial merging".

[0021] Referring to Figure 2, the current block (201) includes samples found by the encoder during motion search, which can be predicted based on the previous block that has been spatially moved by the same size. The MV is not encoded directly, but is derived from metadata associated with one or more reference pictures, such as the most recent (in decoding order) reference picture, by using the MV associated with any of the five surrounding samples. The five surrounding samples are represented by A0, A1 and B0, B1, B2 (from 202 to 206). In H.265, MV prediction can use the prediction value of the same reference picture being used by the neighboring blocks. Summary of the invention

[0022] Various aspects of the present disclosure provide methods and devices for video encoding / decoding. In some examples, the device for video decoding includes a processing circuit. The processing circuit decodes prediction information of a current block from an encoded video code stream, the prediction information indicating a coding mode for encoding the current block based on reconstructed samples, the reconstructed samples being in the same picture as the current block; the processing circuit determines a candidate vector, the candidate vector being used to reconstruct at least a portion of a block having a predetermined spatial relationship with the current block; the processing circuit determines a displacement vector for reconstructing at least a portion of the current block from a candidate list including the candidate vector; further, the processing circuit reconstructs the at least a portion of the current block based on the displacement vector.

[0023] In some examples, the prediction information indicates an intra block copy (IBC) mode for encoding the current block, and the block is one of a spatially adjacent block and a spatially non-adjacent block encoded in a string matching mode.

[0024] In some examples, the prediction information indicates an intra block copy (IBC) mode used to encode the current block, and the blocks are spatially non-adjacent blocks encoded in one of the intra block copy (IBC) mode and a string matching mode.

[0025] In some examples, the prediction information indicates a string matching mode used to encode the current block, and the block is one of a spatially adjacent block and a spatially non-adjacent block encoded in one of the string matching mode and an intra block copy (IBC) mode.

[0026] In some examples, the prediction information indicates an intra block copy (IBC) mode, the block is encoded in a string matching mode, and the candidate vector satisfies the requirement of being a string offset vector of a last string in the block following a scan order.

[0027] In some examples, the prediction information indicates an intra block copy (IBC) mode for encoding the current block, the block is encoded in a string matching mode, and the candidate vector satisfies requirements associated with reconstruction of a non-rectangular portion of the block.

[0028] In some examples, the prediction information indicates a string matching mode used to encode the current block, the block is encoded in an intra block copy (IBC) mode, and the displacement vector satisfies requirements associated with reconstruction of a non-rectangular portion in the current block.

[0029] In some examples, the current block is encoded in an intra-block copy (IBC) mode, the block is encoded in a string matching mode, and a string offset vector of a string in the block is determined as the candidate vector in response to satisfying at least one of the following requirements: the block has more than one string offset vector with different values; the length of the string is not a multiple of the width of the block; the length of the string is not a multiple of the height of the block; the first sample and the last sample of the string are not aligned in the horizontal direction; the first sample and the last sample of the string are not aligned in the vertical direction; the first sample and the last sample of the string are not aligned in the horizontal direction or the vertical direction; and the first sample and the last sample of the string are not aligned in the horizontal direction and the vertical direction.

[0030] In some examples, the current block is encoded in a string matching mode, the block is encoded in an intra-block copy (IBC) mode, and a string offset vector of a string in the current block is determined based on the candidate vector in response to satisfying at least one of the following requirements: the current block has more than one string offset vector with different values; the length of the string is not a multiple of the width of the current block; the length of the string is not a multiple of the height of the current block; the first sample and the last sample of the string are not aligned in the horizontal direction; the first sample and the last sample of the string are not aligned in the vertical direction; the first sample and the last sample of the string are not aligned in the horizontal direction or the vertical direction; and the first sample and the last sample of the string are not aligned in the horizontal direction and the vertical direction.

[0031] Various aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by a computer for video decoding, causes the computer to perform any video encoding / decoding method. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Other features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0033] FIG. 1A is a diagram illustrating an exemplary subset of intra prediction modes.

[0034] FIG. 1B is a schematic diagram of exemplary intra prediction directions.

[0035] FIG2 is a schematic diagram of an exemplary current block and its surrounding spatial merging candidates.

[0036] Figure 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to an embodiment.

[0037] Figure 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to an embodiment.

[0038] Figure 5 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.

[0039] Figure 6 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.

[0040] Figure 7 A block diagram of an encoder according to another embodiment is shown.

[0041] Figure 8 A block diagram of a decoder according to another embodiment is shown.

[0042] Fig. 9 An example of intra block copy according to an embodiment of the present disclosure is shown.

[0043] Fig.10 An example of intra block copy according to an embodiment of the present disclosure is shown.

[0044] Fig.11 An example of intra block copy according to an embodiment of the present disclosure is shown.

[0045] Figures 12A-12D An example of intra block copy according to an embodiment of the present disclosure is shown.

[0046] Fig.13 An example of spatial classification for intra block copy block vector prediction of a current block according to an embodiment of the present disclosure is shown.

[0047] Fig.14 An example of a character string copy mode according to an embodiment of the present disclosure is shown.

[0048] Fig.15 A schematic diagram illustrating a picture during an encoding process according to some embodiments of the present disclosure is shown.

[0049] Fig.16 A flow chart outlining a process according to an embodiment of the present disclosure is shown.

[0050] Fig.17 A schematic diagram of a computer device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0051] Figure 3 3 is a simplified block diagram of a communication system (300) according to an embodiment disclosed in the present application. The communication system (300) includes a plurality of terminal devices, which can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and a terminal device (320) interconnected through the network (350). Figure 3 In an embodiment, the terminal device (310) and the terminal device (320) perform unidirectional data transmission. For example, the terminal device (310) may encode video data (e.g., a video picture stream collected by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data is transmitted in the form of one or more encoded video code streams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to restore the video data, and display the video picture based on the restored video data. Unidirectional data transmission is more common in applications such as media services.

[0052] In another embodiment, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which can occur, for example, during a video conference. For bidirectional data transmission, each of the terminal devices (330) and the terminal device (340) can encode video data (e.g., a video picture stream collected by the terminal device) for transmission to the other terminal device (330) and the terminal device (340) through the network (350). Each of the terminal devices (330) and the terminal device (340) can also receive the encoded video data transmitted by the other terminal device (330) and the terminal device (340), and can decode the encoded video data to restore the video data, and can display the video picture on an accessible display device based on the restored video data.

[0053] exist Figure 3In the embodiment of the present invention, the terminal device (310), the terminal device (320), the terminal device (330) and the terminal device (340) may be a server, a personal computer and a smart phone, but the principles disclosed in the present application may not be limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transmit encoded video data between the terminal devices (310), the terminal devices (320), the terminal devices (330) and the terminal devices (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit switching and / or packet switching channels. The network may include a telecommunications network, a local area network, a wide area network and / or the Internet. For the purpose of this discussion, unless explained below, the architecture and topology of the network (350) may be irrelevant to the operation disclosed in the present application.

[0054] As an example, Figure 4 The video encoder and video decoder are shown in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0055] The streaming system may include an acquisition subsystem (413), which may include a video source (401) such as a digital camera, which creates an uncompressed video picture stream (402). In an embodiment, the video picture stream (402) includes samples taken by a digital camera. Compared to the encoded video data (404) (or the encoded video bitstream), the video picture stream (402) is depicted as a thick line to emphasize the high data volume of the video picture stream, and the video picture stream (402) can be processed by an electronic device (420), which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the video picture stream (402), the encoded video data (404) (or the encoded video bitstream (404)) is depicted as a thin line to emphasize the lower amount of data of the encoded video data (404) (or the encoded video bitstream (404)), which can be stored on the streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4The client subsystem (406) and the client subsystem (408) in the video transmission system can access the streaming server (405) to retrieve the copy (407) and the copy (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in the electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or another presentation device (not depicted). In some streaming transmission systems, the encoded video data (404), the video data (407), and the video data (409) (e.g., a video bitstream) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265. In an embodiment, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application can be used in the context of the VVC standard.

[0056] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0057] Figure 5 is a block diagram of a video decoder (510) according to an embodiment disclosed in the present application. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 A video decoder (410) of an embodiment.

[0058] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be provided external to the video decoder (510) (not shown). In other cases, a buffer memory (not shown) is provided outside the video decoder (510) to prevent network jitter, for example, and another buffer memory (515) may be configured inside the video decoder (510) to handle broadcast timing, for example. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (515), or the buffer memory may be made smaller. Of course, in order to use on a service packet network such as the Internet, a buffer memory (515) may also be required, and the buffer memory may be relatively large and have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not shown) outside the video decoder (510).

[0059] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The types of symbols include information used to manage the operation of the video decoder (510) and potential information used to control a display device such as a display device (512) (e.g., a display screen) that is not part of the electronic device (530) but can be coupled to the electronic device (530), such as Figure 5As shown in . The control information for the display device may be a parameter set fragment (not indicated) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (520) may extract a subgroup parameter set of at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0060] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).

[0061] Depending on the type of coded video picture or portion of coded video picture (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors, the reconstruction of the symbol (521) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. For the sake of brevity, such subgroup control information flow between the parser (520) and the multiple units below is not described.

[0062] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In a practical embodiment operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0063] The first unit is a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) receives quantized transform coefficients as symbols (521) from the parser (520) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (551) can output a block including sample values, which can be input into an aggregator (555).

[0064] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates surrounding blocks of the same size and shape as the block being reconstructed using reconstructed information extracted from a current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0065] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (553) may access the reference picture memory (557) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (521), these samples may be added to the output of the scaler / inverse transform unit (551) (in this case referred to as residual samples or residual signals) by the aggregator (555) to generate output sample information. The acquisition of the predicted samples by the motion compensated prediction unit (553) from the address in the reference picture memory (557) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (553) in the form of the symbols (521), for example, including X, Y and reference picture components. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.

[0066] The output samples of the aggregator (555) may be employed by various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, as well as to previously reconstructed and loop filtered sample values.

[0067] The output of the loop filter unit (556) may be a sample stream that may be output to a display device (512) and stored in a reference picture memory (557) for subsequent inter-picture prediction.

[0068] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0069] The video decoder (510) may perform decoding operations according to a predetermined video compression technique, such as in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0070] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0071] Figure 6 6 is a block diagram of a video encoder (603) according to an embodiment disclosed in the present application. The video encoder (603) is provided in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace Figure 4 A video encoder (403) in an embodiment.

[0072] The video encoder (603) can be used to obtain the video source (601) (not Figure 6 In another embodiment, the video source (601) is a part of the electronic device (620) to receive video samples, and the video source can collect video images to be encoded by the video encoder (603). In another embodiment, the video source (601) is a part of the electronic device (620).

[0073] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0074] According to an embodiment, the video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, coupling is not indicated in the figure. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be used to have other suitable functions that are related to the video encoder (603) optimized for a certain system design.

[0075] In some embodiments, the video encoder (603) operates in an encoding loop. As a simple description, in embodiments, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (because in the video compression technology considered in this application, any compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, eg due to channel errors) is also used in some related techniques.

[0076] The operation of the "local" decoder (633) can be combined with the above Figure 4 The "remote" decoder described in detail for the video decoder (510) is identical. However, additional brief reference is made to Figure 5 , when symbols are available and the entropy encoder (645) and parser (520) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633).

[0077] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in substantially the same functional form. For this reason, the present application focuses on decoder operation. The description of encoder technology can be simplified because encoder technology is mutually inverse to the decoder technology described comprehensively. A more detailed description is only needed in certain areas and is provided below.

[0078] During operation, in some embodiments, the source encoder (630) may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (632) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.

[0079] The local video decoder (633) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data is available at the video decoder ( Figure 6 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0080] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as appropriate prediction references for the new picture. The predictor (635) may operate pixel-by-pixel based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (635), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (634).

[0081] The controller (650) may manage encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0082] The outputs of all the above functional units may be entropy encoded in an entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0083] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0084] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0085] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and features.

[0086] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0087] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.

[0088] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or the blocks may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0089] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0090] In an embodiment, the transmitter (640) may transmit additional data when transmitting the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0091] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an embodiment, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0092] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block may be predicted by a combination of the first reference block and the second reference block.

[0093] In addition, merge mode technology can be used in inter-picture prediction to improve coding efficiency.

[0094] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Furthermore, each CTU can be split into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0095] Figure 7 is a diagram of a video encoder (703) according to another embodiment disclosed in the present application. The video encoder (703) is used to receive a processed block (e.g., a prediction block) of sample values ​​in a current video picture in a video picture sequence, and encode the processed block into an encoded picture that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used to replace Figure 4A video encoder (403) in an embodiment.

[0096] In an HEVC embodiment, a video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion (RD) optimization to determine whether to use intra mode, inter mode, or bidirectional prediction mode to encode the processing block. When encoding the processing block in intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when encoding the processing block in inter mode or bidirectional prediction mode, the video encoder (703) may use inter prediction or bidirectional prediction techniques to encode the processing block into an encoded picture, respectively. In some video encoding techniques, the merge mode may be an inter-picture prediction submode, in which a motion vector is derived from one or more motion vector predictors without the aid of an encoded motion vector component external to the predictor. In some other video encoding techniques, there may be a motion vector component applicable to the subject block. In an embodiment, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the processing block mode.

[0097] exist Figure 7 In an embodiment of the present invention, the video encoder (703) includes: Figure 7 Shown are an inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together.

[0098] The inter-frame encoder (730) is used to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-frame prediction information (e.g., redundant information description according to an inter-frame coding technique, motion vectors, merge mode information), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some embodiments, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0099] The intra-frame encoder (722) is used to receive samples of a current block (e.g., a processing block), compare the block with an encoded block in the same picture in some cases, generate quantization coefficients after transformation, and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques). In an embodiment, the intra-frame encoder (722) also calculates an intra-frame prediction result (e.g., a predicted block) based on the intra-frame prediction information and a reference block in the same picture.

[0100] The general controller (721) is used to determine the general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information and add the intra prediction information to the bitstream; and when the mode is inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and add the inter prediction information to the bitstream.

[0101] The residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra-frame encoder (722) or the inter-frame encoder (730). The residual encoder (724) is used to operate based on the residual data to encode the residual data to generate a transform coefficient. In an embodiment, the residual encoder (724) is used to convert the residual data from the time domain to the frequency domain and generate a transform coefficient. The transform coefficient is then processed by quantization to obtain a quantized transform coefficient. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra-frame encoder (722) and the inter-frame encoder (730). For example, the inter-frame encoder (730) can generate a decoded block based on the decoded residual data and the inter-frame prediction information, and the intra-frame encoder (722) can generate a decoded block based on the decoded residual data and the intra-frame prediction information. The decoded blocks are appropriately processed to generate a decoded picture, and in some embodiments, the decoded picture may be buffered in a memory circuit (not shown) and used as a reference picture.

[0102] The entropy encoder (725) is used to format the code stream to produce an encoded block. The entropy encoder (725) generates various information according to a suitable standard such as the HEVC standard. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (such as intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the code stream. It should be noted that according to the disclosed subject matter, when the block is encoded in the inter-frame mode or the merge sub-mode of the bidirectional prediction mode, there is no residual information.

[0103] Figure 8FIG. 8 is a diagram of a video decoder (810) according to another embodiment disclosed in the present application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (810) is used to replace Figure 4 A video decoder (410) of an embodiment.

[0104] exist Figure 8 In an embodiment, the video decoder (810) includes Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874) and an intra-frame decoder (872) coupled together are shown in FIG.

[0105] The entropy decoder (871) may be used to reconstruct certain symbols from the encoded picture, which represent syntax elements constituting the encoded picture. Such symbols may include, for example, a mode for encoding the block (e.g., intra mode, inter mode, bidirectional prediction mode, a merged sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that may identify certain samples or metadata for prediction by the intra decoder (872) or the inter decoder (880), respectively, residual information in the form of, for example, quantized transform coefficients, and the like. In an embodiment, when the prediction mode is inter or bidirectional prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be inverse quantized and provided to the residual decoder (873).

[0106] The inter-frame decoder (880) is used to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0107] The intra-frame decoder (872) is used to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information.

[0108] The residual decoder (873) is used to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require some control information (to obtain the quantizer parameter QP), and this information can be provided by the entropy decoder (871) (the data path is not indicated because this is only low-volume control information).

[0109] The reconstruction module (874) is used to combine the residual output by the residual decoder (873) with the prediction result (which can be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which can be part of a reconstructed picture, which in turn can be part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations can be performed to improve visual quality.

[0110] It should be noted that the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using any suitable technology. In an embodiment, the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoder (403), video encoder (603) and video encoder (703) and video decoder (410), video decoder (510) and video decoder (810) may be implemented using one or more processors that execute software instructions.

[0111] Block-based compensation can be used for inter-frame prediction and intra-frame prediction. For inter-frame prediction, block-based compensation from different pictures is called motion compensation. Block-based compensation can also be done from previously reconstructed areas within the same picture, such as in intra-frame prediction. Block-based compensation from reconstructed areas within the same picture is called intra-picture block compensation, current picture referencing (CPR) or intrablock copy (IBC). The displacement vector indicating the offset between the current block and the reference block (also called the prediction block) in the same picture is called a block vector (BV, block vector), where the current block can be encoded / decoded based on the reference block. Unlike the motion vector in motion compensation, which can be any value (positive or negative in the x or y direction), BV has several constraints to ensure that the reference block is available and has been reconstructed. In addition, in some examples, some reference areas that are tile boundaries, strip boundaries, or wavefront trapezoid boundaries are excluded for parallel processing considerations.

[0112] The encoding of the block vector can be explicit or implicit. In the explicit mode, the BV difference between the block vector and its predictor is signaled. In the implicit mode, the block vector is recovered from the predictor (called the block vector predictor) without using the BV difference, in a similar way to the motion vector in the merge mode. The explicit mode may be referred to as the non-merged BV prediction mode. The implicit mode may be referred to as the merged BV prediction mode.

[0113] In some implementations, the resolution of block vectors is limited to integer positions. In other systems, block vectors are allowed to point to fractional positions.

[0114] In some examples, a block level flag (such as an IBC flag) may be used to signal the use of intra block copy at the block level. In an embodiment, the block level flag is signaled when the current block is explicitly encoded. In some examples, a reference index method may be used to signal the use of intra block copy at the block level. The current picture in decoding is then processed as a reference picture or a special reference picture. In an example, such a reference picture is placed at the last position of the reference picture list. Special reference pictures are also managed together with other temporal reference pictures in a buffer, such as a decoded picture buffer (DPB).

[0115] The IBC mode can vary. In the example, the IBC mode is treated as a third mode different from the intra prediction mode and the inter prediction mode. Accordingly, the BV prediction in the implicit mode (or merge mode) and the explicit mode is separate from the conventional inter mode. A separate merge candidate list can be defined for the IBC mode, wherein the entries in the separate merge candidate list are multiple BVs. Similarly, in the example, the BV prediction candidate list in the IBC explicit mode includes only multiple BVs. The general rule applied to these two lists (i.e., a separate merge candidate list and a BV prediction candidate list) is that, with respect to the candidate derivation process, these two lists can follow the same logic as the merge candidate list used in the conventional merge mode or the AMVP predictor list used in the conventional AMVP mode. For example, five spatially adjacent positions (e.g., A0, A1 and B0, B1, B2 in Figure 2) are accessed for the IBC mode, e.g., the HEVC or VVC inter-frame merge mode, to derive a separate merge candidate list for the IBC mode.

[0116] As described above, the BV of the current block in the reconstruction in the picture may have certain constraints, and therefore, the reference block of the current block is within the search range. The search range refers to a part of the picture from which the reference block can be selected. For example, the search range may be within certain parts of the reconstructed area in the picture. The size, position and / or shape of the search range, etc. may be constrained. Alternatively, the BV may be constrained. In the example, the BV is a two-dimensional vector including an x ​​component and a y component, and at least one of the x component and the y component may be constrained. Constraints may be specified with respect to the BV, the search range, or a combination of the BV and the search range. In various examples, when certain constraints are specified for the BV, the search range is constrained accordingly. Similarly, when certain constraints are specified for the search range, the BV is constrained accordingly.

[0117] Fig. 9 An example of intra-block copying according to an embodiment of the present disclosure is shown. The current picture (900) is reconstructed under decoding. The current picture (900) includes a reconstructed area (910) (gray area) and an area to be decoded (920) (white area). The current block (930) is reconstructed by the decoder. The current block (930) can be reconstructed from a reference block (940) in the reconstructed area (910). The position offset between the reference block (940) and the current block (930) is called a block vector (950) (or BV (950)). In Fig. 9 In the example of , the search range (960) is within the reconstructed region (910), the reference block (940) is within the search range (960), and the block vector (950) is constrained to point to the reference block (940) within the search range (960).

[0118] Various constraints may be applied to the BV and / or the search range. In an embodiment, the search range of the current block being reconstructed in the current CTB is constrained within the current CTB.

[0119] In an embodiment, the effective memory requirement for storing reference samples to be used in intra block copying is one CTB size. In the example, the CTB size is 128×128 samples. The current CTB includes the current region being reconstructed. The current region has a size of 64×64 samples. Since the reference memory can also store reconstructed samples in the current region, when the reference memory size is equal to the CTB size of 128×128 samples, the reference memory can store 3 more regions of 64×64 samples. Accordingly, the search range can include some parts of the previously reconstructed CTB, while the total memory requirement for storing reference samples remains unchanged (such as 1 CTB size of 128×128 samples or a total of 4 64×64 reference samples). In the example, the previously reconstructed CTB is the left neighbor of the current CTB, such as Fig.10 as shown in .

[0120] Fig.10An example of intra-block copying according to an embodiment of the present disclosure is shown. The current picture (1001) includes a current CTB (1015) being reconstructed and a previously reconstructed CTB (1010) which is a left neighbor of the current CTB (1015). The CTB in the current picture (1001) has a CTB size such as 128×128 samples and a CTB width such as 128 samples. The current CTB (1015) includes 4 regions (1016)-(1019), where the current region (1016) is being reconstructed. The current region (1016) includes multiple coding blocks (1021)-(1029). Similarly, the previously reconstructed CTB (1010) includes 4 regions (1011)-(1014). The coding blocks (1021)-(1025) are being reconstructed, the current block (1026) is being reconstructed, and the coding blocks (1026)-(1027) and regions (1017)-(1019) are to be reconstructed.

[0121] The current region (1016) has a co-located region (i.e., region (1011) in the previously reconstructed CTB (1010)). The relative position of the co-located region (1011) with respect to the previously reconstructed CTB (1010) may be the same as the relative position of the current region (1016) with respect to the current CTB (1015). Fig.10 In the illustrated example, the current region (1016) is the upper left region in the current CTB (1015), and therefore, the co-located region (1011) is also the upper left region in the previously reconstructed CTB (1010). Since the position of the previously reconstructed CTB (1010) is offset by the CTB width from the position of the current CTB (1015), the position of the co-located region (1011) is offset by the CTB width from the position of the current region (1016).

[0122] In an embodiment, the co-located region of the current region (1016) is in a previously reconstructed CTB, wherein the position of the previously reconstructed CTB is offset from the position of the current CTB (1015) by one or more CTB widths, and therefore, the position of the co-located region is also offset from the position of the current region (1016) by a corresponding one or more CTB widths. The position of the co-located region may be offset to the left or upward from the current region (1016), etc.

[0123] As described above, the size of the search range of the current block (1026) is constrained by the CTB size. Fig.10In the example of , the search range may include regions (1012)-(1014) in the previously reconstructed CTB (1010) and a portion of the reconstructed current region (1016) (such as coding blocks (1021)-(1025)). The search range further excludes the co-located region (1011), so that the size of the search range is within the CTB size range. Fig.10 , the reference block (1091) is located in the region (1014) of the previously reconstructed CTB (1010). The block vector (1020) indicates the offset between the current block (1026) and the corresponding reference block (1091). The reference block (1091) is within the search range.

[0124] Fig.10 The example illustrated in the figure can be appropriately adapted to other scenarios where the current region is located at another position in the current CTB (1015). In the example, when the current block is in the region (1017), the co-located region of the current block is the region (1012). Therefore, the search range can include the region (1013)-(1014), the region (1016), and a portion of the region (1017) that has been reconstructed. The search range further excludes the region (1011) and the co-located region (1012), so that the size of the search range is within the CTB size range. In the example, when the current block is in the region (1018), the co-located region of the current block is the region (1013). Therefore, the search range can include the region (1014), the region (1016)-(1017), and a portion of the region (1018) that has been reconstructed. The search range further excludes the region (1011)-(1012) and the co-located region (1013), so that the size of the search range is within the CTB size range. In the example, when the current block is in region (1019), the co-located region of the current block is region (1014). Therefore, the search range may include regions (1016)-(1018) and a portion of region (1019) that has already been reconstructed. The search range further excludes the previously reconstructed CTB (1010), so that the size of the search range is within the CTB size range.

[0125] In the above description, the reference block may be in a previously reconstructed CTB (1010) or a current CTB (1015).

[0126] In an embodiment, the search range may be specified as follows. In an example, the current picture is a luma picture, and the current CTB is a luma CTB including a plurality of luma samples, and BV(mvL) satisfies the following constraints of code stream consistency. In an example, BV(mvL) has a fractional resolution (e.g., 1 / 16 pixel resolution).

[0127] The constraint includes a first condition that a reference block of the current block has been reconstructed. When the reference block has a rectangular shape, an adjacent block availability check process (or a reference block availability check process) may be implemented to check whether to reconstruct an upper left sample and a lower right sample of the reference block. When the upper left sample and the lower right sample of the reference block are reconstructed, it is determined that the reference block is to be reconstructed.

[0128] For example, when the derivation process of reference block availability is called with the position of the upper left sample of the current block (xCurr, yCurr) (set to (xCb, yCb)) and the position of the upper left sample of the reference block (xCb+(mvL[0]>>4), yCb+(mvL[1]>>4)) as input, the output is equal to true when the upper left sample of the reference block is reconstructed, where the block vector mvL is a two-dimensional vector with an x-component mvL[0] and a y-component mvL[1]. When BV(mvL) has a fractional resolution such as 1 / 16 pixel resolution, the x-component mvL[0] and the y-component mvL[1] are shifted to have integer resolution, as indicated by mvL[0]>>4 and mvL[1]>>4, respectively.

[0129] Similarly, when the derivation process of block availability is called with the position of the top left sample of the current block (xCurr, yCurr) (set to (xCb, yCb)) and the position of the bottom right sample of the reference block (xCb+(mvL[0]>>4)+cbWidth-1, yCb+(mvL[1]>>4)+cbHeight-1) as input, the output is equal to true when the bottom right sample of the reference block is reconstructed. The parameters cbWidth and cbHeight represent the width and height of the reference block.

[0130] The constraint may also include at least one of the following second conditions: 1) the value of (mvL[0]>>4)+cbWidth is less than or equal to 0, indicating that the reference block is on the left side of the current block and does not overlap with the current block; 2) the value of (mvL[1]>>4)+cbHeight is less than or equal to 0, indicating that the reference block is above the current block and does not overlap with the current block.

[0131] The constraint may also include that the block vector mvL satisfies the following third condition: (yCb + ( mvL[1] >> 4) ) >> CtbLog2SizeY = yCb >> CtbLog2SizeY (1) (yCb + ( mvL[1] >> 4 + cbHeight - 1) >> CtbLog2SizeY = yCb >>CtbLog2Size (2) (xCb + ( mvL[0] >> 4)) >> CtbLog2SizeY >= (xCb >> CtbLog2SizeY) – 1 (3) (xCb+(mvL[0]>>4)+cbWidth-1)>>CtbLog2SizeY<=(xCb>>CtbLog2SizeY)(4) Wherein, the parameter CtbLog2SizeY represents the CTB width in log2 form. For example, when the CTB width is 128 samples, CtbLog2SizeY is 7. Equations (1)-(2) indicate that the CTB including the reference block is in the same CTB row as the current CTB (for example, when the reference block is in the previously reconstructed CTB (1010), the previously reconstructed CTB (1010) is in the same row as the current CTB (1015)). Equations (3)-(4) indicate that the CTB including the reference block is in the left CTB column of the current CTB or in the same CTB column as the current CTB. The third condition described by equations (1)-(4) indicates that the CTB including the reference block is the current CTB (such as the current CTB (1015)) or the left neighbor of the current CTB (such as the previously reconstructed CTB (1010)), similar to the reference block. Fig.10 Description.

[0132] The constraint may further include a fourth condition: when the reference block is in the left neighbor of the current CTB, the co-located region of the reference block is not reconstructed (i.e., no samples in the co-located region are reconstructed). Further, the co-located region of the reference block is in the current CTB. Fig.10 In the example of , the co-located region of the reference block (1091) is a region (1019) offset by a CTB width from the region (1014) where the reference block (1091) is located, and the region (1019) is not reconstructed. Therefore, the block vector (1020) and the reference block (1091) satisfy the fourth condition described above.

[0133] In an example, the fourth condition may be specified as follows: when (xCb+(mvL[0]>>4))>>CtbLog2SizeY isequal to (xCb>>CtbLog2SizeY)-1, the derivation process for reference block availability is called with the position (xCurr, yCurr) of the current block (set to (xCb, yCb)) and the positions (((xCb+(mvL[0]>>4)+CtbSizeY)>>(CtbLog2SizeY-1))<<(CtbLog2SizeY-1), ((yCb+(mvL[1]>>4))>>(CtbLog2SizeY-1))<<(CtbLog2SizeY-1)) as output, the output is equal to false to indicate that the co-located region has not been reconstructed, for example Fig.10 as shown in .

[0134] The constraints of the search range and / or block vector may include a suitable combination of the first condition, the second condition, the third condition and the fourth condition described above. In an example, the constraints include the first condition, the second condition, the third condition and the fourth condition, for example Fig.10 In the example, the first condition, the second condition, the third condition and / or the fourth condition may be modified, and the constraint includes the modified first condition, the second condition, the third condition and / or the fourth condition.

[0135] According to the fourth condition, when one of the coding blocks (1022)-(1029) is the current block, the reference block cannot be in the region (1011), and therefore, the search range of one of the coding blocks (1022)-(1029) excludes the region (1011). The reason for excluding the region (1011) is specified as follows: if the reference block is in the region (1011), the co-located region of the reference block is the region (1016), however, at least the samples in the coding block (1021) have been reconstructed, and therefore the fourth condition is violated. On the other hand, for a coding block to be reconstructed first in the current region, e.g. Fig.11 For a coding block (1121) in region (1116) in the coding block, the fourth condition does not prevent the reference block from being in region (1111) because the co-located region (1116) of the reference block has not yet been reconstructed.

[0136] Fig.11An example of intra-block copying according to an embodiment of the present disclosure is shown. The current picture (1101) includes a current CTB (1115) under reconstruction and a previously reconstructed CTB (1110) which is a left neighbor of the current CTB (1115). The CTB in the current picture (1101) has a CTB size and a CTB width. The current CTB (1115) includes 4 regions (1116)-(1119), wherein the current region (1116) is under reconstruction. The current region (1116) includes a plurality of coding blocks (1121)-(1129). Similarly, the previously reconstructed CTB (1110) includes 4 regions (1111)-(1114). The current block (1121) under reconstruction will first be reconstructed in the current region (1116), and the coding blocks (1122)-(1129) will be reconstructed. In the example, the CTB size is 128×128 samples, and each of the regions (1111)-(1114) and (1116)-(1119) is 64×64 samples. The reference memory size is equal to the CTB size and is 128×128 samples, and therefore, when the search range is defined by the reference memory size, the search range includes 3 regions and a portion of the additional region.

[0137] Similar to reference Fig.10 As described above, the current region (1116) has a co-located region (i.e., region (1111) in the previously reconstructed CTB (1110)). According to the fourth condition described above, the reference block of the current block (1121) may be in region (1111), and therefore, the search range may include regions (1111)-(1114). For example, when the reference block is in region (1111), the co-located region of the reference block is region (1116), wherein no samples in region (1116) are reconstructed before reconstructing the current block (1121). However, as shown in the reference block Fig.10 As described in the fourth condition, for example, after reconstructing the coding block (1121), the region (1111) is no longer available to be included in the search range for reconstructing the coding block (1122). Therefore, tight synchronization and timing control of the reference memory buffer will be used and may be challenging.

[0138] According to some embodiments, when the current block is first reconstructed in the current region of the current CTB, the search range may exclude the co-located region of the current region in a previously reconstructed CTB, wherein the current CTB and the previously reconstructed CTB are in the same current picture. The block vector may be determined so that the reference block is in the search range excluding the co-located region in the previously reconstructed CTB. In an embodiment, the search range includes coded blocks reconstructed after the co-located region and before the current block in decoding order.

[0139] In the following description, the CTB size may vary, and the maximum CTB size is set to be the same as the reference memory size. In an example, the reference memory size or the maximum CTB size is 128×128 samples. The description may be appropriately adapted to other reference memory sizes or maximum CTB sizes.

[0140] In an embodiment, the CTB size is equal to the reference memory size. The previously reconstructed CTB is a left neighbor of the current CTB, the position of the co-located region is offset from the position of the current region by the CTB width, and the coding block in the search range is in at least one of the current CTB and the previously reconstructed CTB.

[0141] FIG. 12A to FIG. 12D An example of intra-block copying according to an embodiment of the present disclosure is shown. FIG. 12A to FIG. 12D , the current picture (1201) includes a current CTB (1215) under reconstruction and a previously reconstructed CTB (1210) as a left neighbor of the current CTB (1215). The CTB in the current picture (1201) has a CTB size and a CTB width. The current CTB (1215) includes 4 regions (1216)-(1219). Similarly, the previously reconstructed CTB (1210) includes 4 regions (1211)-(1214). In an embodiment, the CTB size is the maximum CTB size and is equal to the reference memory size. In the example, the CTB size and the reference memory size are 128×128 samples, and therefore, each of the regions (1211)-(1214) and (1216)-(1219) has a size of 64×64 samples.

[0142] exist FIG. 12A to FIG. 12D In the illustrated example, the current CTB (1215) includes an upper left region, an upper right region, a lower left region, and a lower right region corresponding to regions (1216)-(1219), respectively. The previously reconstructed CTB (1210) includes an upper left region, an upper right region, a lower left region, and a lower right region corresponding to regions (1211)-(1214), respectively.

[0143] refer to Fig. 12A , the current region (1216) is being reconstructed. The current region (1216) may include a plurality of coding blocks (1221)-(1229). The current region (1216) has a co-located region in a previously reconstructed CTB (1210), namely, region (1211). The search range of one of the coding blocks (1221)-(1229) to be reconstructed may exclude the co-located region (1211). The search range may include regions (1212)-(1214) of the previously reconstructed CTB (1210) reconstructed in decoding order after the co-located region (1211) and before the current region (1216).

[0144] refer to Fig. 12A , the position of the co-located region (1211) is offset from the position of the current region (1216) by the CTB width, such as 128 samples. For example, the position of the co-located region (1211) is offset to the left by 128 samples from the position of the current region (1216).

[0145] Reference again Fig. 12A , when the current region (1216) is the upper left region (1215) of the current CTB, the co-located region (1211) is the upper left region (1210) of the previously reconstructed CTB, and the search region excludes the upper left region of the previously reconstructed CTB.

[0146] refer to Fig. 12B , the current region (1217) is under reconstruction. The current region (1217) may include multiple coding blocks (1241)-(1249). The current region (1217) has a co-located region (i.e., region (1212) in a previously reconstructed CTB (1210)). The search range of one of the multiple coding blocks (1241)-(1249) may exclude the co-located region (1212). The search range includes regions (1213)-(1214) of the previously reconstructed CTB (1210) and a region (1216) in the current CTB (1215) reconstructed after the co-located region (1212) and before the current region (1217). Due to the constraint of the reference memory size (i.e., the size of one CTB), the search range further excludes region (1211). Similarly, the position of the co-located region (1212) is offset from the position of the current region (1217) by a CTB width, such as 128 samples.

[0147] exist Fig. 12B In the example, the current region (1217) is the upper right region of the current CTB (1215), the co-located region (1212) is also the upper right region of the previously reconstructed CTB (1210), and the search region excludes the upper right region of the previously reconstructed CTB (1210).

[0148] refer to Fig. 12C, the current region (1218) is under reconstruction. The current region (1218) may include multiple coding blocks (1261)-(1269). The current region (1218) has a co-located region (i.e., region (1213)) in a previously reconstructed CTB (1210). The search range of one of the multiple coding blocks (1261)-(1269) may exclude the co-located region (1213). The search range includes region (1214) of the previously reconstructed CTB (1210) and region (1216)-(1217) in the current CTB (1215) reconstructed after the co-located region (1213) and before the current region (1218). Similarly, due to the constraint of the reference memory size, the search range further excludes region (1211)-(1212). The position of the co-located region (1213) is offset from the position of the current region (1218) by the CTB width, such as 128 samples. Fig. 12C In the example, when the current region (1218) is the lower left region of the current CTB (1215), the co-located region (1213) is also the lower left region of the previously reconstructed CTB (1210), and the search region excludes the lower left region of the previously reconstructed CTB (1210).

[0149] refer to Fig.12D , the current region (1219) is under reconstruction. The current region (1219) may include multiple coding blocks (1281)-(1289). The current region (1219) has a co-located region (i.e., region (1214)) in a previously reconstructed CTB (1210). The search range of one of the multiple coding blocks (1281)-(1289) may exclude the co-located region (1214). The search range includes regions (1216)-(1218) in the current CTB (1215) that are reconstructed in decoding order after the co-located region (1214) and before the current region (1219). Due to the constraint of the reference memory size, the search range excludes regions (1211)-(1213), and therefore, the search range excludes the previously reconstructed CTB (1210). Similarly, the position of the co-located region (1214) is offset from the position of the current region (1219) by the CTB width, such as 128 samples. Fig.12D In the example, when the current region (1219) is the lower right region of the current CTB (1215), the co-located region (1214) is also the lower right region of the previously reconstructed CTB (1210), and the search region excludes the lower right region of the previously reconstructed CTB (1210).

[0150] Referring back to FIG. 2 , the MVs associated with the five surrounding samples (or positions), represented as A0, A1 and B0, B1, B2 (202 to 206), respectively, may be referred to as spatial merge candidates. A candidate list (e.g., a merge candidate list) may be formed based on the spatial merge candidates. Any suitable order may be used to form the candidate list from the positions. In an example, the order may be A0, B0, B1, A1, and B2, with A0 being first and B2 being last. In an example, the order may be A1, B1, B0, A0, and B2, with A1 being first and B2 being last.

[0151] According to some embodiments, motion information of previously encoded blocks of a current block (e.g., a coding block (CB) or a current CU) may be stored in a history-based motion vector prediction (HMVP) buffer (e.g., a table) to provide motion vector prediction (MVP) candidates (also referred to as HMVP candidates) for the current block. The HMVP buffer may include one or more HMVP candidates and may be maintained during the encoding / decoding process. In an example, the HMVP candidates in the HMVP buffer correspond to motion information of previously encoded blocks. The HMVP buffer may be used in any suitable encoder and / or decoder. The HMVP candidate may be added to a merge candidate list after the HMVP spatial MVP and TMVP.

[0152] When a new CTU (or new CTB) row is encountered, the HMVP buffer may be reset (eg, cleared).When a non-subblock inter-coded block exists, the associated motion information may be added to the last entry of the HMVP buffer as a new HMVP candidate.

[0153] In an example, such as in VTM3, the buffer size (represented by S) of the HMVP buffer is set to 6, which indicates that up to 6 HMVP candidates can be added to the HMVP buffer. In some embodiments, the HMVP buffer can operate with a first-in-first-out (FIFO) rule, and therefore, for example, when the HMVP buffer is full, a piece of motion information (or HMVP candidate) first stored in the HMVP buffer will be removed from the HMVP buffer first. When a new HMVP candidate is inserted into the HMVP buffer, a constrained FIFO rule can be utilized, in which a redundancy check is first applied to determine whether there are identical or similar HMVP candidates in the HMVP buffer. If it is determined that the identical or similar HMVP candidate is in the HMVP buffer, the identical or similar HMVP candidate can be removed from the HMVP buffer, and the remaining HMVP candidates can be moved forward in the HMVP buffer.

[0154] HMVP candidates can be used in the merge candidate list construction process, such as in merge mode. The most recently stored HMVP candidates in the HMVP buffer can be checked in order and inserted into the merge candidate list after the TMVP candidate. Redundancy checks can be applied to HMVP candidates relative to spatial or temporal merge candidates in the merge candidate list. The description can be appropriately adapted to AMVP mode to construct an AMVP candidate list.

[0155] To reduce the number of redundant check operations, the following simplification can be used. (i) The number of HMVP candidates used to generate the merge candidate list can be set to (N<=4)? M: (8-N). N indicates the number of existing candidates in the merge candidate list, and M indicates the number of available HMVP candidates in the HMVP buffer. When the number of existing candidates in the merge candidate list (N) is less than or equal to 4, the number of HMVP candidates used to generate the merge candidate list is equal to M. Otherwise, the number of HMVP candidates used to generate the merge candidate list is equal to (8-N). (ii) When the total number of available merge candidates reaches the maximum allowed merge candidates minus 1, the merge candidate list construction process from the HMVP buffer is terminated.

[0156] When IBC mode is operated as a mode separate from inter prediction mode, a simplified BV derivation process for IBC mode can be used. A history-based block vector prediction buffer (referred to as an HBVP buffer) can be used to perform BV prediction. The HBVP buffer can be used to store BV information (e.g., BV) of previously encoded blocks of a current block (e.g., CB or CU) in a current picture. In an example, the HBVP buffer is a history buffer separate from other buffers (such as an HMVP buffer). The HBVP buffer can be a table.

[0157] The HBVP buffer can provide BV predictor (BVP, BV predictor) candidates (also referred to as HBVP candidates) for the current block. The HBVP buffer (e.g., a table) may include one or more HBVP candidates and may be maintained during the encoding / decoding process. In the example, the HBVP candidates in the HBVP buffer correspond to the BV information of the previously encoded blocks in the current picture. The HBVP buffer can be used for any suitable encoder and / or decoder. The HBVP candidate can be added to a merge candidate list, which is configured for BV prediction after the BV of the spatially neighboring blocks of the current block. The merge candidate list configured for BV prediction can be used for a merged BV prediction mode and / or a non-merged BV prediction mode.

[0158] The HBVP buffer may be reset (eg, flushed) when a new CTU (or new CTB) row is encountered.

[0159] In an example, such as in VVC, the buffer size of the HBVP buffer is set to 6, which indicates that up to 6 HBVP candidates can be added to the HBVP buffer. In some embodiments, the HBVP buffer can be operated with a FIFO rule, and therefore, for example, when the HBVP buffer is full, a piece of BV information (or HBVP candidate) first stored in the HBVP buffer will be removed from the HBVP buffer first. When a new HBVP candidate is inserted into the HBVP buffer, a constrained FIFO rule can be utilized, in which a redundancy check is first applied to determine whether the same or similar HBVP candidate exists in the HBVP buffer. If it is determined that the same or similar HBVP candidate is in the HBVP buffer, the same or similar HBVP candidate can be removed from the HBVP buffer, and the remaining HBVP candidates can be moved forward in the HBVP buffer.

[0160] HBVP candidates can be used in the merge candidate list construction process, such as in the merge BV prediction mode. The most recently stored HBVP candidates in the HBVP buffer can be checked in order and inserted into the merge candidate list after the spatial candidates. Redundancy checks can be applied to HBVP candidates relative to the spatial merge candidates in the merge candidate list.

[0161] In an embodiment, a HBVP buffer is established to store one or more BV information of one or more previously encoded blocks encoded in the IBC mode. The one or more BV information may include one or more BVs of one or more previously encoded blocks encoded in the IBC mode. Further, each of the one or more BV information may include side information (or additional information), such as a block size, a block location, etc. of the corresponding previously encoded block encoded in the IBC mode.

[0162] In class based history-based block vector prediction (also known as CBVP), for the current block, one or more BV information in the HBVP buffer that meets certain conditions can be classified into corresponding categories (also known as categories), and thus a CBVP buffer is formed. In the example, each BV information block in the HBVP buffer is used for a corresponding previously encoded block, such as a block encoded in IBC mode. The BV information of the previously encoded block may include BV, block size and / or block position, etc. The previously encoded block has a block width, a block height, and a block area. The block area may be the product of the block width and the block height. In the example, the block size is represented by the block area. The block position of the previously encoded block may be represented by the upper left corner (e.g., the upper left corner of a 4×4 area) or the upper left sample of the previously encoded block.

[0163] Fig.13An example (1310) of a spatial category for IBC BV prediction of a current block (e.g., CB, CU) according to an embodiment of the present disclosure is shown. The left area (1302) may be on the left side of the current block (1310). The BV information of a previously encoded block having a corresponding block position in the left area (1302) may be referred to as a left candidate or a left BV candidate. The top area (1303) may be above the current block (1310). The BV information of a previously encoded block having a corresponding block position in the top area (1303) may be referred to as a top candidate or a top BV candidate. The upper left area (1304) may be on the upper left of the current block (1310). The BV information of a previously encoded block having a corresponding block position in the upper left area (1304) may be referred to as an upper left candidate or an upper left BV candidate. The upper right area (1305) may be on the upper right of the current block (1310). The BV information of the previously encoded block with the corresponding block position in the upper right area (1305) may be referred to as an upper right candidate or an upper right BV candidate. The lower left area (1306) may be at the lower left of the current block (1310). The BV information of the previously encoded block with the corresponding block position in the lower left area (1306) may be referred to as a lower left candidate or a lower left BV candidate. Other types of spatial categories may also be defined and used in the CBVP buffer.

[0164] If the BV information of the previously encoded block satisfies the following conditions, the BV information may be classified into a corresponding category (or class). (i) Category 0: The block size (eg, block area) is greater than or equal to a threshold (eg, 64 pixels). (ii) Category 1: The occurrence rate (or frequency) of the BV is greater than or equal to 2. The occurrence rate of the BV may refer to the number of times the BV is used to predict a previously coded block. When a pruning process is used to form a CBVP buffer, when a BV is used multiple times to predict a previously coded block, the BV may be stored in one entry (rather than in multiple entries with the same BV). The occurrence rate of the BV may be recorded. (iii) Category 2: The block position is in the left region (1302), where a portion of the previously encoded block (e.g., the upper left corner of the 4×4 region) is to the left of the current block (1310). The previously encoded block may be in the left region (1302). Alternatively, the previously encoded block may span multiple regions including the left region (1302), where the block position is in the left region (1302). (iv) Category 3: The block position is in the top region (1303), where a portion of the previously encoded block (e.g., the upper left corner of the 4×4 region) is above the current block (1310). The previously encoded block may be within the top region (1303). Alternatively, the previously encoded block may span multiple regions including the top region (1303), where the block position is in the top region (1303). (v) Category 4: The block position is in the upper left region (1304), where a portion of the previously encoded block (e.g., the upper left corner of the 4×4 region) is at the upper left side (1310) of the current block. The previously encoded block may be within the upper left region (1304). Alternatively, the previously encoded block may span multiple regions including the upper left region (1304), where the block position is in the upper left region (1304). (vi) Category 5: The block position is in the upper right region (1305), where a portion of the previously encoded block (e.g., the upper left corner of the 4×4 region) is at the upper right side (1310) of the current block. The previously encoded block may be within the upper right region (1305). Alternatively, the previously encoded block may span multiple regions including the upper right region (1305), where the block position is in the upper right region (1305). (vii) Category 6: The block position is in the lower left region (1306), where a portion of the coded block (e.g., the upper left corner of the 4×4 region) is at the lower left side (1310) of the current block. The previously coded block may be in the lower left region (1306). Alternatively, the previously coded block may span multiple regions including the lower left region (1306), where the block position is in the lower left region (1306).

[0165] For each category (or class), the BV of the most recently encoded block can be derived as a BVP candidate. The CBVP buffer can be constructed by appending the BV predictor of each class in order from class 0 to class 6. The above description of CBVP can be appropriately adapted to include fewer classes or additional classes not described above. One or more of categories 0-6 can be modified. In the example, each entry in the HBVP buffer is classified into one of seven categories 0-6. An index can be signaled to indicate which category of categories 0-6 is selected. On the decoder side, the first entry in the selected category can be used to predict the BV of the current block.

[0166] Aspects of the present disclosure provide techniques for spatial displacement vector prediction for intra-picture blocks and string copy modes. String copy mode is also referred to as string matching or string prediction. String matching is similar to intra block copy (IBC) and can reconstruct reconstructed areas based on sample strings within the same picture. Further, string matching provides more flexibility regarding the shape of sample strings. For example, a block has a rectangular shape and a string can form a non-rectangular shape.

[0167] Fig.14An example of a string copy mode according to an embodiment of the present disclosure is shown. The current picture (1410) includes a reconstruction area (gray area) (1420) and an area under reconstruction (1421). The current block (1435) in the area (1421) is under reconstruction. The current block (1435) can be a CB, a CU, etc. The current block (1435) can include multiple strings, such as Fig.14 The strings (1430) and strings (1431) in the example. In the example, the current block (1435) is divided into a plurality of consecutive strings, wherein one string is followed by the next string along the scanning order. The scanning order can be any suitable scanning order, such as a raster scanning order or a traversal scanning order.

[0168] The reconstruction area (1420) may be used as a reference area for reconstructing the string (1430) and the string (1431).

[0169] For each of the multiple strings, a string offset vector (also referred to as a string vector (SV)) and the length of the string (referred to as the string length) may be signaled. The SV may be a displacement vector indicating a displacement offset between a string to be reconstructed and a reference string that is located in a reference region (1420) and has been reconstructed. The reference string may be used to reconstruct the string to be reconstructed. For example, SV0 is a displacement vector indicating a displacement offset between a string (1430) and a reference string (1400), and SV1 is a displacement vector indicating a displacement offset between a string (1431) and a reference string (1401). Thus, the SV may indicate where the corresponding reference string is located in the reference region (1420). The string length of the string indicates the number of samples in the string. Typically, the string to be reconstructed has the same length as the reference string.

[0170] refer to Fig.14 , the current block (1435) is an 8×8 CB including 64 samples. The current block (1435) is divided into a string (1430) and a string (1431) using a raster scan order. The string (1430) includes the first 29 samples of the current block (1435), and the string (1431) includes the remaining 35 samples of the current block (1435). The reference string (1400) used to reconstruct the string (1430) can be represented by the corresponding string offset vector SV0, and the reference string (1401) used to reconstruct the string (1431) can be represented by the corresponding string offset vector SV1.

[0171] In general, string size can refer to the length of the string or the number of samples in the string. Fig.14, string (1430) includes 29 samples, and thus the string size of string (1430) is 29. String (1431) includes 35 samples, and thus the string size of string (1431) is 35. A string location (or string position) may be represented by a sample position of a sample in a string (e.g., the first sample in decoding order).

[0172] The above description may be suitably adapted to reconstruct a current block comprising any suitable number of strings. Alternatively, in an example, when a sample in the current block does not have a matching sample in the reference region, an escape sample is signaled, and the value of the escape sample may be encoded directly without reference to the reconstructed sample in the reference region.

[0173] Various aspects of the present disclosure provide displacement vector prediction techniques for both block vector prediction of IBC mode and string offset vector prediction of string matching mode. Displacement vector prediction techniques can be used in skip mode, in direct / merge mode, or in vector prediction with differential coding. Various displacement vector prediction techniques can be used alone or in combination in any order. In the following description, the term "block" can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU.

[0174] Further, in the following description, the terms of spatially adjacent blocks and spatially non-adjacent blocks are used. A spatially adjacent block refers to a coded block that is adjacent (directly adjacent) to a current block, and the spatially adjacent block may share the same boundary line or the same corner point with the current block. In contrast, a spatially non-adjacent block refers to a previously coded block that is not a spatially adjacent block, and the spatially non-adjacent block does not have the same boundary line or the same corner point as the current block.

[0175] Fig.15 A diagram of a picture (1500) during an encoding process according to some embodiments of the present disclosure is shown. Fig.15 Some examples of a current block (1510) and spatially neighboring blocks, such as those shown by A0-E0, are shown. Spatially neighboring block A0 shares a portion of a boundary line (shown by (1521)) with the current block (1510). Spatially neighboring block B0 shares a portion of a boundary line (shown by (1522)) with the current block (1510). Spatially neighboring block C0 shares a corner point (shown by (1523)) with the current block (1510). Spatially neighboring block D0 shares a corner point (shown by (1524)) with the current block (1510). Spatially neighboring block E0 shares a corner point (shown by (1525)) with the current block (1510). It should be noted that other blocks, such as blocks positioned along a row above the current block (1510) between E0 and B0 or ​​blocks positioned along a column to the left of the current block (1510) between A0 and E0, may also be considered spatially neighboring positions.

[0176] Further, Fig.15 Also shown are spatially non-neighboring blocks such as shown by A1-E1, A2-E2, A3-E3, etc. that do not share any boundary portions or corner points with the current block (1510).

[0177] In some examples, the spatial neighboring blocks may be Fig.15 , and the larger blocks A0-B0, A1-B1, A2-B2, and A3-B3 shown in . For example, a block (1530) including A0 is encoded based on the displacement vector, and the block (1530) is a spatially adjacent block with respect to the current block (1510). In another example, a block (1540) including E1 is encoded based on the displacement vector, and the block (1540) does not have a shared boundary portion or corner with the current block (1510). The block (1540) is a spatially non-adjacent block with respect to the current block (1510).

[0178] According to an aspect of the present disclosure, when a current block is encoded in the IBC mode, a block vector or a string offset vector associated with a spatially adjacent block or a spatially non-adjacent block may be used to predict the current block.

[0179] In an embodiment, when the current block is encoded in IBC mode, spatial non-adjacent blocks encoded in IBC mode or string matching mode may be considered as candidates for predicting a block vector of the current block. In an example, when the current block is encoded in IBC mode and spatial non-adjacent blocks are encoded in IBC mode, block vectors of spatial non-adjacent blocks may be placed in a candidate list for predicting a block vector of the current block. In an example, when the current block is encoded in IBC mode and spatial non-adjacent blocks are encoded in string matching mode, string offset vectors of strings in spatial non-adjacent blocks may be placed in a candidate list for predicting a block vector of the current block.

[0180] In another embodiment, when the current block is encoded in IBC mode, the spatial non-adjacent blocks encoded in IBC mode can be considered as candidates for predicting the block vector of the current block encoded in IBC mode. Both the spatially adjacent blocks and the spatially non-adjacent blocks encoded in string matching mode are considered as candidates for predicting the block vector of the current block encoded in IBC mode. In an example, when the current block is encoded in IBC mode and the spatially non-adjacent blocks are encoded in IBC mode, the block vectors of the spatially non-adjacent blocks can be placed in a candidate list for predicting the block vector of the current block. In another example, when the current block is encoded in IBC mode and the spatially adjacent blocks are encoded in string matching mode, the string offset vectors of the strings in the spatially adjacent blocks can be placed in a candidate list for predicting the block vector of the current block. In another example, when the current block is encoded in IBC mode and the spatially non-adjacent blocks are encoded in string matching mode, the string offset vectors of the strings in the spatially non-adjacent blocks can be placed in a candidate list for predicting the block vector of the current block.

[0181] In some embodiments, when the current block is encoded in IBC mode, the string needs to meet certain requirements in order to use the string offset vector of the string to predict the block vector of the current block. In an embodiment, the string is required to be the last string in the block (spatial adjacent block or spatial non-adjacent block) containing the string. For example, the block includes multiple strings following a scan order, and then the last string of the multiple strings following the scan order can meet the requirements and the string offset vector of the string can be used (e.g., put into a candidate list) to predict the block vector of the current block.

[0182] In another embodiment, the length of the string is required to meet certain requirements. In an example, when the current block is encoded in IBC mode and the string in the block has a length that is not a multiple of the width of the block (using a horizontal scan order), the string offset vector of the string can be used (e.g., the string offset vector is placed in a candidate list) to predict the block vector of the current block. In another example, when the current block is encoded in IBC mode and the string in the block has a length that is not a multiple of the height of the block (using a vertical scan order), the string offset vector of the string can be used (e.g., the string offset vector is placed in a candidate list) to predict the block vector of the current block.

[0183] In another embodiment, the shape of the string is required to meet certain requirements. In an example, when the current block is encoded in IBC mode and the string in the block (e.g., spatially adjacent blocks, spatially non-adjacent blocks) does not form a rectangular shape in the block, the string offset vector of the string can be used (e.g., the string offset vector is put into a candidate list) to predict the block vector of the current block.

[0184] In some embodiments, when the current block is encoded in IBC mode, the block encoded in string matching mode needs to meet certain requirements in order to use the string offset vector of the string in the block to predict the block vector of the current block. In an embodiment, the block encoded in string matching mode needs to have more than one string offset vector (for example, there are at least two strings with different string offset vectors), then one of the string offset vectors can be used to predict the block vector of the current block.

[0185] In some embodiments, when the current block is encoded in IBC mode, the coordinates of the samples in the string need to meet certain requirements in order to use the string offset vector of the string to predict the block vector of the current block. For example, (x1, y1) represents the coordinates of the first sample in the string following the scan order, and (x2, y2) represents the coordinates of the last sample in the string following the scan order. For example, the first requirement may require that x1 is not equal to x2 (for example, the first sample and the last sample are not aligned in the horizontal direction); the second requirement may require that y1 is not equal to y2 (for example, the first sample and the last sample are not aligned in the vertical direction); the third requirement may require that x1 is not equal to x2 and y1 is not equal to y2; and the fourth requirement may require that x1 is not equal to x2 or y1 is not equal to y2. In an example, one of the four requirements can be selected to constrain the coordinates of the first and last samples in the string to allow the string offset vector of the string to be used to predict the block vector of the current block. In an example, when the selected requirement is not met, the string offset vector of the string is not considered as a candidate for predicting the block vector of the current block and cannot be put into the candidate list for predicting the block vector of the current block.

[0186] According to another aspect of the present disclosure, when a current block is encoded in a string matching mode, a block vector or a string offset vector associated with a spatially adjacent block or a spatially non-adjacent block may be used to predict the current block encoded in the string matching mode.

[0187] In an embodiment, when the current block is encoded in string matching mode, spatial non-adjacent blocks encoded in IBC mode or string matching mode may be considered as candidates for predicting a string offset vector of a string in the current block encoded in string matching mode. In an example, when the current block is encoded in string matching mode and the spatial non-adjacent blocks are encoded in IBC mode, the block vectors of the spatial non-adjacent blocks may be placed in a candidate list for predicting the string offset vector of the current block. In another example, when the current block is encoded in string matching mode and the spatial non-adjacent blocks are encoded in string matching mode, the string offset vectors of the strings in the spatial non-adjacent blocks may be placed in a candidate list for predicting the string offset vector of the current block.

[0188] In another embodiment, when the current block is encoded in string matching mode, the spatially neighboring blocks encoded in IBC mode or string matching mode are considered as candidates for predicting the string offset vector of the string in the current block. In an example, when the current block is encoded in string matching mode and the spatially neighboring blocks are encoded in IBC mode, the block vectors of the spatially neighboring blocks may be placed in a candidate list for predicting the string offset vector of the string in the current block. In another example, when the current block is encoded in string matching mode and the spatially neighboring blocks are encoded in string matching mode, the string offset vectors of the strings in the spatially neighboring blocks may be placed in a candidate list for predicting the string offset vector of the string in the current block.

[0189] In another embodiment, when the current block is encoded in string matching mode, the block vectors and string offset vectors associated with the spatially adjacent blocks and spatially non-adjacent blocks encoded in the IBC mode or the string matching mode may be considered as candidates for predicting the string offset vector of the string in the current block. In an example, when the current block is encoded in the string matching mode and the spatially adjacent blocks are encoded in the IBC mode, the block vectors of the spatially adjacent blocks may be placed in a candidate list for predicting the string offset vector of the string in the current block. In another example, when the current block is encoded in the string matching mode and the spatially adjacent blocks are encoded in the string matching mode, the string offset vectors of the strings in the spatially adjacent blocks may be placed in a candidate list for predicting the string offset vector of the string in the current block. In another example, when the current block is encoded in the string matching mode and the spatially non-adjacent blocks are encoded in the IBC mode, the block vectors of the spatially non-adjacent blocks may be placed in a candidate list for predicting the string offset vector of the string in the current block. In another example, when the current block is encoded in string matching mode and the spatial non-adjacent blocks are encoded in string matching mode, the string offset vectors of the strings in the spatial non-adjacent blocks can be put into a candidate list for predicting the string offset vectors of the strings in the current block.

[0190] In another embodiment, when a block vector is used to predict a string vector of a string in a current block, the current block needs to satisfy the requirement that the current block has more than one string offset vector (e.g., multiple string offset vectors of different values). Otherwise, the block vector cannot be considered as a candidate for predicting a string vector of a string in the current block.

[0191] In another embodiment, when the current block is encoded in string matching mode and has only one string offset vector, the constraint that the encoded BV should come from a spatially non-adjacent block can be imposed. In the example, the current block is encoded in string matching mode and has only one string offset vector, if a block vector is associated with a spatially non-adjacent block, the block vector can be put into a candidate list for predicting the string offset vector.

[0192] In another embodiment, the current block is encoded in a string matching mode. When a block vector of a block (e.g., a spatially adjacent block, a spatially non-adjacent block) is used to predict a string offset vector of a string in the current block, a constraint is imposed so that the length of the string is not a multiple of the width of the current block when the string follows a horizontal scan order. When the string follows a vertical scan order, similarly, a constraint may be imposed so that the length of the string is not a multiple of the height of the current block. In an example, when the current block is encoded in a string matching mode and the string in the current block has a length that is not a multiple of the width of the current block (using a horizontal scan order), then a block vector of a block (e.g., a spatially adjacent block or a spatially non-adjacent block) may be used (e.g., the block vector is placed in a candidate list) to predict a string offset vector of a string in the current block. In another example, when the current block is encoded in a string matching mode and the string in the current block has a length that is not a multiple of the height of the current block (using a vertical scan order), then a block vector of a block (e.g., a spatially adjacent block or a spatially non-adjacent block) may be used (e.g., the block vector is placed in a candidate list) to predict a string offset vector of a string in the current block.

[0193] In another embodiment, the current block is encoded in a string matching mode, and when a block vector of a block (e.g., a spatially adjacent block, a spatially non-adjacent block) is used to predict a string offset vector of a string in the current block, constraints are imposed on the coordinates of the samples in the string. For example, (x1, y1) represents the coordinates of the first sample in the string following the scan order, and (x2, y2) represents the coordinates of the last sample in the string following the scan order, and the coordinates of the first sample and the last sample need to meet certain requirements in order to use the block vector to predict the string offset vector of the string. For example, the first requirement may require that x1 is not equal to x2 (e.g., the first sample and the last sample are not aligned in the horizontal direction); the second requirement may require that y1 is not equal to y2 (e.g., the first sample and the last sample are not aligned in the vertical direction); the third requirement may require that x1 is not equal to x2 and y1 is not equal to y2; and the fourth requirement may require that x1 is not equal to x2 or y1 is not equal to y2. In the example, one of the four requirements is selected to constrain the coordinates of the first sample and the last sample in the string so as to allow the block vector to be used to predict the string offset vector of the string in the current block. In an example, when the selected requirement is not met, the block vector of the block is not considered as a candidate for predicting the string offset vector of the string in the current block and cannot be put into the candidate list for predicting the string offset vector of the string in the current block.

[0194] In another embodiment, the shape of the string is required to meet certain requirements. In an example, the current block is encoded in string matching mode, and the string offset vector of the string in the current block is predicted using the block vectors of the spatially adjacent blocks in the IBC mode, and the string is constrained not to form a rectangular shape in the current block (e.g., the string is required not to form a rectangular shape in the current block).

[0195] According to another aspect of the present disclosure, a block vector or a string offset vector from a spatially adjacent block or a spatially non-adjacent block can be combined with a history-based block vector or a history-based string offset vector in a candidate list. In some embodiments, a spatial candidate (e.g., a block vector of a block from a spatially adjacent block or a spatially non-adjacent block, a string offset vector from a spatially adjacent block or a spatially non-adjacent block) can be placed before a history-based candidate (e.g., a history-based block vector, a history-based string offset vector) in a candidate list for predicting a block vector or a string offset vector. In some embodiments, a spatial candidate (e.g., a block vector of a block from a spatially adjacent block or a spatially non-adjacent block, a string offset vector from a spatially adjacent block or a spatially non-adjacent block) can be placed after a history-based candidate (e.g., a history-based block vector, a history-based string offset vector) in a candidate list for predicting a block vector or a string offset vector.

[0196] In some embodiments, category-based prediction may be used for string vector prediction, for example in a manner similar to that used in category-based history-based block vector prediction (also referred to as CBVP). In some examples, additional information such as location information and size information associated with a new spatial candidate may be stored in a candidate list along with the vector information (block vector or string offset vector). The new spatial candidate may be a block vector from a spatially adjacent block, a block vector from a spatially non-adjacent block, a string offset vector from a spatially adjacent block, or a string offset vector from a spatially non-adjacent block. Using the location information and size information, the new spatial candidate may be classified into the correct predictor category.

[0197] According to another aspect of the present disclosure, in order to place several spatial candidates into a candidate list, spatial locations are accessed in a specific order (such as a predetermined order) to determine the spatial candidates. For example, consider predicting a string offset vector in a string matching mode with spatially non-adjacent blocks. In the example, the spatially non-adjacent blocks at locations A1-E1 can be accessed in the order of A1, B1, C1, D1, E1. In the example, the current block is encoded in string matching mode. When accessing a location, if the block including the location is encoded in IBC or string matching mode, the displacement vector (block vector or string offset vector) associated with the block can be used to predict the string offset vector in the current block, and can be placed in a candidate list to predict the string in the current block encoded in string matching mode.

[0198] In another example, the current block is encoded in IBC mode. When a location is accessed, if the block including the location is encoded in IBC or string matching mode, the displacement vector (block vector or string offset vector) associated with the block can be used to predict the block vector of the current block and can be put into a candidate list to predict the current block encoded in IBC mode.

[0199] Fig.16 A flowchart outlining a process (1600) according to an embodiment of the present disclosure is shown. The process (1600) may be used to reconstruct a block or string in a picture of an encoded video sequence. The process (1600) may be used for reconstruction of a block to generate a prediction block for the block in reconstruction. The term block in the present disclosure may be interpreted as a prediction block, CB, CU, etc. In various embodiments, the process (1600) is performed by a processing circuit, such as a processing circuit in a terminal device (310), (320), (330), and (340), a processing circuit that performs the function of a video encoder (403), a processing circuit that performs the function of a video decoder (410), a processing circuit that performs the function of a video decoder (510), and a processing circuit that performs the function of a video encoder (603). In some embodiments, the process (1600) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1600). The process starts at (S1601) and proceeds to (S1610).

[0200] At (S1610), prediction information of a current block from an encoded video code stream is decoded. The prediction information indicates a coding mode for encoding the current block based on reconstructed samples, the reconstructed samples being in the same picture as the current block.

[0201] It should be noted that the encoding mode may be an intra block copy (IBC) mode or may be a string matching mode.

[0202] At (S1620), a candidate vector is determined, the candidate vector being used to reconstruct at least a portion of a block having a predetermined spatial relationship with the current block.

[0203] It should be noted that the predetermined spatial relationship may include spatially adjacent blocks to the current block and / or spatially non-adjacent blocks to the current block.It should also be noted that the blocks may be encoded in IBC mode or string matching mode.

[0204] In some embodiments, the prediction information indicates an intra block copy (IBC) mode for encoding the current block, and the block is one of a spatially neighboring block and a spatially non-neighboring block encoded in a string matching mode.

[0205] In some embodiments, the prediction information indicates an intra block copy (IBC) mode for encoding the current block, and the block is a spatially non-adjacent block encoded in one of the IBC mode and the string matching mode.

[0206] In some embodiments, the prediction information indicates a string matching mode for encoding the current block, and the block is one of a spatially neighboring block and a spatially non-neighboring block encoded in one of a string matching mode and an intra block copy (IBC) mode.

[0207] In some embodiments, the prediction information indicates an intra block copy (IBC) mode, the block is encoded in a string matching mode, and the candidate vector satisfies the requirement of being a string offset vector of the last string in the block following a scan order.

[0208] In some embodiments, the prediction information indicates an intra block copy (IBC) mode for encoding the current block, the block is encoded in a string matching mode, and the candidate vector satisfies requirements associated with reconstruction of a non-rectangular portion of the block.

[0209] In some embodiments, the current block is encoded in an intra-block copy (IBC) mode, the block is encoded in a string matching mode, and a string offset vector of a string in the block is determined as a candidate vector in response to satisfying at least one of the following requirements: the block has more than one string offset vector with different values; the length of the string is not a multiple of the width of the block; the length of the string is not a multiple of the height of the block; the first sample and the last sample of the string are not aligned in the horizontal direction; the first sample and the last sample of the string are not aligned in the vertical direction; the first sample and the last sample of the string are not aligned in the horizontal direction or the vertical direction; and the first sample and the last sample of the string are not aligned in the horizontal direction and the vertical direction.

[0210] At (S1630), a displacement vector for reconstructing at least a portion of the current block is determined from a candidate list including candidate vectors.

[0211] The displacement vector can be a block vector or a string offset vector.

[0212] In some embodiments, the prediction information indicates a string matching mode for encoding a current block, the block is encoded in an intra block copy (IBC) mode, and the displacement vector satisfies requirements associated with reconstruction of a non-rectangular portion in the current block.

[0213] In some embodiments, the current block is encoded in a string matching mode, the block is encoded in an intra-block copy (IBC) mode, and then a string offset vector for a string in the current block is determined based on a candidate vector in response to satisfying at least one of the following requirements: the current block has more than one string offset vector with different values; the length of the string is not a multiple of the width of the current block; the length of the string is not a multiple of the height of the current block; the first sample and the last sample of the string are not aligned in the horizontal direction; the first sample and the last sample of the string are not aligned in the vertical direction; the first sample and the last sample of the string are not aligned in the horizontal direction or the vertical direction; and the first sample and the last sample of the string are not aligned in the horizontal direction and the vertical direction.

[0214] At (S1640), a portion of the current block is reconstructed based on the displacement vector.

[0215] Then, the process proceeds to (S1699) and terminates.

[0216] The process (1600) may be modified as appropriate. Steps in the process (1600) may be modified and / or omitted. Additional steps may be added. Any suitable implementation order may be used. For example, when it is determined that the current vector information is unique, the current vector information may be stored in the history buffer as described above. In some examples, a pruning process is used and one of the vector information in the history buffer is removed when the current vector information is stored in the history buffer.

[0217] The above techniques may be implemented as computer software via computer-readable instructions and physically stored in one or more computer-readable media. Fig.17 A computer device (1700) is shown that is suitable for implementing certain embodiments of the disclosed subject matter.

[0218] The computer software may be encoded in any suitable machine code or computer language, and may be assembled, compiled, linked, or the like to create a code comprising instructions, which may be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed by decoding, microcode, or the like.

[0219] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, IoT devices, etc.

[0220] Fig.17 The components shown for the computer device (1700) are exemplary in nature and are not intended to limit the scope of use or functionality of computer software implementing embodiments of the present application. The configuration of the components should not be interpreted as having any dependency or requirement on any component or combination of components shown in the exemplary embodiment of the computer device (1700).

[0221] The computer device (1700) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from one or more human users through tactile input (e.g., keyboard input, sliding, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0222] The human-computer interface input device may include one or more of the following (only one of which is drawn): keyboard (1701), mouse (1702), touchpad (1703), touch screen (1710), data gloves (not shown), joystick (1705), microphone (1706), scanner (1707), camera (1708).

[0223] The computer device (1700) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (1710), a data glove (not shown), or a joystick (1705), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1709), headphones (not shown)), visual output devices (e.g., screens (1710) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each of which has or does not have a touch screen input function, each of which has or does not have a tactile feedback function - some of which may output two-dimensional visual output or output of more than three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)) and printers (not shown).

[0224] The computer device (1700) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical disks (CD / DVD ROM / RW) (1720) with CD / DVD or similar media (1721), thumb drives (1722), removable hard disk drives or solid state drives (1723), traditional magnetic media such as tapes and floppy disks (not shown), special-purpose ROM / ASIC / PLD-based devices such as security software protectors (not shown), and the like.

[0225] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0226] The computer device (1700) may also include an interface (1755) to one or more communication networks (1754). The network may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. Examples of networks may include local area networks such as Ethernet, wireless local area networks, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), in-vehicle and industrial networks (including CANBus), and the like. Some networks typically require an external network interface adapter for connecting to some universal data port or peripheral bus (1749) (e.g., a USB port of the computer device (1700)); other systems are typically integrated into the core of the computer device (1700) by connecting to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smart phone computer system). By using any of these networks, the computer device (1700) can communicate with other entities. The communication can be one-way, for receiving only (e.g., wireless television), one-way for sending only (e.g., CAN bus to certain CAN bus devices), or two-way, such as to other computer systems via a local or wide area digital network. Each of the above networks and network interfaces can use certain protocols and protocol stacks.

[0227] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core (1740) of the computer device (1700).

[0228] The core (1740) may include one or more central processing units (CPUs) (1741), graphics processing units (GPUs) (1742), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1743), hardware accelerators for specific tasks (1744), etc. These devices, as well as read-only memory (ROM) (1745), random access memory (1746), internal mass storage (e.g., internal non-user accessible hard disk drives, solid-state drives, etc.) (1747), etc., may be connected via a system bus (1748). In some computer systems, the system bus (1748) may be accessed in the form of one or more physical plugs so that it may be expanded by additional central processing units, graphics processing units, etc. Peripheral devices may be directly attached to the core's system bus (1748) or connected via a peripheral bus (1749). The architecture of the peripheral bus includes PCI, USB, etc. In one example, the screen (1710) may be connected to a graphics adapter (1750). The architecture of the peripheral bus includes PCI, USB, etc.

[0229] The CPU (1741), GPU (1742), FPGA (1743) and accelerator (1744) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (1745) or RAM (1746). Transition data can also be stored in RAM (1746), while permanent data can be stored in, for example, internal mass storage (1747). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1741), GPUs (1742), mass storage (1747), ROM (1745), RAM (1746), etc.

[0230] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purpose of this application, or may be medium and code well known and available to those skilled in the art of computer software.

[0231] As an example and not a limitation, a computer system having an architecture (1700), in particular a core (1740), can be provided as a processor (including a CPU, a GPU, an FPGA, an accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with the above-mentioned user-accessible mass storage, as well as a specific memory of the core (1740) having non-volatility, such as a core internal mass storage (1747) or a ROM (1745). Software implementing various embodiments of the present application can be stored in such a device and executed by the core (1740). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the core (1740), in particular the processor therein (including a CPU, a GPU, an FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in the RAM (1746) and modifying such a data structure according to a software-defined process. Additionally or alternatively, the computer system may provide functionality hardwired in logic or otherwise contained in circuitry (e.g., accelerator (1744)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an integrated circuit (IC)) storing the executing software, circuitry containing the executing logic, or both. The present application includes any suitable combination of hardware and software. Appendix: Acronyms JEM: Joint Development Model VVC: Next Generation Video Coding BMS: Benchmark Collection MV: Motion Vector HEVC: High Efficiency Video Coding MPM: Most Probable Mode WAIP: Wide Angle Intra Prediction SEI: Supplemental Enhancement Information VUI: Video Availability Information GOPs: Group of Pictures TUs: Transformation Units PUs: prediction units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: prediction blocks HRD: Hypothesized Reference Decoder SDR: Standard Dynamic Range SNR: Signal to Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Field Programmable Gate Array IC: Integrated Circuit CU: Coding Unit PDPC: Position Dependent Prediction Combination ISP: Intra-frame sub-partitioning SPS: Sequence Parameter Set

[0232] Although the present application has described a number of exemplary embodiments, various changes, arrangements and various equivalent substitutions of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design a variety of systems and methods, which, although not explicitly shown or described herein, embody the principles of the present application and are therefore within the spirit and scope of the present application.

Claims

1. A video decoding method, It is characterized in that include: Obtain prediction information of a current block in an encoded video bitstream, the prediction information indicating a coding mode for encoding the current block based on reconstructed samples, the reconstructed samples and the current block being in the same picture; Determining, based on a coding mode of the current block and a coding mode of a block having a predetermined spatial relationship with the current block, a candidate vector indicating a displacement vector of a displacement offset between the current block and at least a portion of the block; Determining, from a candidate list including the candidate vector, a displacement vector for reconstructing at least a portion of the current block; as well as The at least a portion of the current block is reconstructed based on the displacement vector.

2. The method according to claim 1, It is characterized in that The prediction information indicates an intra block copy (IBC) mode for encoding the current block, and the block is one of a spatially adjacent block and a spatially non-adjacent block encoded in a string matching mode.

3. The method according to claim 1, It is characterized in that The prediction information indicates an intra block copy (IBC) mode used to encode the current block, and the blocks are spatial non-adjacent blocks encoded in one of the intra block copy (IBC) mode and a string matching mode.

4. The method according to claim 1, It is characterized in that The prediction information indicates a string matching mode for encoding the current block, and the block is one of a spatially adjacent block and a spatially non-adjacent block encoded in one of the string matching mode and an intra block copy (IBC) mode.

5. The method according to claim 1, It is characterized in that The prediction information indicates an intra block copy (IBC) mode, the block is encoded in a string matching mode, and the candidate vector meets the requirement of being a string offset vector of the last string in the block following a scanning order.

6. The method according to claim 1, It is characterized in that The prediction information indicates an intra block copy (IBC) mode for encoding the current block, the block is encoded in a string matching mode, and the candidate vector satisfies requirements associated with reconstruction of a string with a non-rectangular shape in the block.

7. The method according to claim 1, It is characterized in that The prediction information indicates a string matching mode for encoding the current block, the block is encoded in an intra block copy (IBC) mode, and the displacement vector meets requirements associated with reconstruction of a string with a non-rectangular shape in the current block.

8. The method according to claim 1, It is characterized in that The current block is encoded in an intra block copy (IBC) mode, the block is encoded in a string matching mode, and a string offset vector of a string in the block is determined as the candidate vector in response to satisfying at least one of the following requirements: more than one string offset vector having different values ​​for the block; The length of the string is not a multiple of the width of the block; the length of the string is not a multiple of the height of the block; The first sample and the last sample of the string are not aligned in the horizontal direction; The first sample and the last sample of the string are not vertically aligned; The first sample and the last sample of the string are not aligned in the horizontal direction or the vertical direction; as well as The first sample and the last sample of the string are not aligned in the horizontal direction and the vertical direction.

9. The method according to claim 1, It is characterized in that The current block is encoded in a string matching mode, the block is encoded in an intra block copy (IBC) mode, and a string offset vector of a string in the current block is determined based on the candidate vector in response to satisfying at least one of the following requirements: The current block has more than one string offset vector with different values; The length of the string is not a multiple of the width of the current block; The length of the string is not a multiple of the height of the current block; The first sample and the last sample of the string are not aligned in the horizontal direction; The first sample and the last sample of the string are not vertically aligned; The first sample and the last sample of the string are not aligned in the horizontal direction or the vertical direction; as well as The first sample and the last sample of the string are not aligned in the horizontal direction and the vertical direction.

10. The method according to any one of claims 1 to 9, It is characterized in that The candidate list includes the candidate vectors and location and size information of the blocks.

11. A video encoding method, It is characterized in that Used to generate video streams, including: Obtain prediction information of a current block, the prediction information indicating a coding mode for encoding the current block based on coded samples, the coded samples and the current block being in the same picture; Determining, based on a coding mode of the current block and a coding mode of a block having a predetermined spatial relationship with the current block, a candidate vector indicating a displacement vector of a displacement offset between the current block and at least a portion of the block; Determining a displacement vector for encoding at least a portion of the current block from a candidate list including the candidate vector; and The at least a portion of the current block is encoded based on the displacement vector.

12. A video decoding device, It is characterized in that comprising a processing circuit, the processing circuit being configured to: Obtain prediction information of a current block in an encoded video bitstream for decoding, the prediction information indicating a coding mode for encoding the current block based on reconstructed samples, the reconstructed samples and the current block being in the same picture; Determining, based on a coding mode of the current block and a coding mode of a block having a predetermined spatial relationship with the current block, a candidate vector indicating a displacement vector of a displacement offset between the current block and at least a portion of the block; Determining, from a candidate list including the candidate vector, a displacement vector for reconstructing at least a portion of the current block; as well as The at least a portion of the current block is reconstructed based on the displacement vector.

13. A video encoding device, It is characterized in that comprising a processing circuit, the processing circuit being configured to: Obtain prediction information of a current block, the prediction information indicating a coding mode for encoding the current block based on coded samples, the coded samples and the current block being in the same picture; Determining, based on a coding mode of the current block and a coding mode of a block having a predetermined spatial relationship with the current block, a candidate vector indicating a displacement vector of a displacement offset between the current block and at least a portion of the block; Determining, from a candidate list including the candidate vector, a displacement vector for encoding at least a portion of the current block; as well as The at least a portion of the current block is encoded based on the displacement vector.

14. A method for storing a video stream, It is characterized in that The video code stream is stored on a non-volatile computer-readable storage medium, wherein the video code stream is decoded according to the video decoding method according to any one of claims 1 to 10, or the video code stream is generated by the video encoding method according to claim 11.

15. A non-volatile computer readable medium storing instructions, It is characterized in that When the instructions are executed by a computer, the computer is caused to perform the method according to any one of claims 1 to 11.

16. A computer device, It is characterized in that include: A processor and a memory; the memory stores computer code, and when the computer code is executed by the processor, the processor executes the method according to any one of claims 1 to 11.