Method and apparatus for video decoding

By adopting the flip operation of string copy mode in video encoding, intra prediction is optimized, and the problems of increasing the number of intra prediction directions and redundancy of motion vectors are solved, and the compression rate and decoding efficiency of video encoding are improved.

CN115211111BActive Publication Date: 2025-07-22TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180018062.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-27
Filing Date
2021-09-01
Publication Date
2025-07-22
Estimated Expiration
2041-09-01

AI Technical Summary

Technical Problem

In the existing video encoding technology, in intra prediction, the increase in the number of intra prediction directions leads to a decrease in encoding efficiency, and the motion vector prediction is large, which affects the compression rate and decoding efficiency.

Method used

The string copy mode is used to process the string vector of the current block through the flip operation, including vertical flip, horizontal flip and combined flip, to generate a flip reference string to reconstruct the current string, and optimize the intra prediction process.

Benefits of technology

It improves the encoding efficiency of intra prediction, reduces the redundant information of motion vectors, and improves the video compression rate and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115211111B_ABST
    Figure CN115211111B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatuses for video decoding. A processing circuit may decode encoded information of a current block in a current picture according to an encoded video bitstream. The encoded information may indicate a string copy mode of the current block. The current block includes a current string. The processing circuit may determine whether to perform a flipping operation to predict the current string. Based on determining to perform the flipping operation to predict the current string, the processing circuit may determine an original reference string based on a string vector of the current string. The processing circuit may generate a flipped reference string by performing a flipping operation on the original reference string, and reconstruct the current string based on the flipped reference string.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure describes embodiments generally related to video decoding. Background Art

[0002] The background description provided herein is for the purpose of generally presenting the background of the present disclosure. Within the scope described in this background section, the works of the presently named inventors and aspects of this description that were not otherwise available as prior art at the time of filing are neither expressly nor implicitly admitted as prior art to the present disclosure.

[0003] Inter-frame image prediction with motion compensation can be used to perform video encoding and decoding. Uncompressed digital video can include a series of images, each having a spatial size of, for example, luminance samples and associated chrominance samples of 1920×1080. The series of images can have a fixed or variable image rate (also informally referred to as the frame rate), such as 60 images per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, a 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GB of storage space.

[0004] One purpose of video encoding and decoding can be to reduce redundancy in the input video signal through compression. Compression can help reduce the above-mentioned bandwidth and / or storage space requirements, and in some cases can reduce them by two orders of magnitude or more than two orders of magnitude. Lossless compression and lossy compression and their combinations can be employed. Lossless compression refers to techniques that can reconstruct an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application. In the case of video, lossy compression is widely used. The amount of tolerable distortion depends on the application; for example, users of certain consumer streaming applications can tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher admissible / tolerable distortion can result in a higher compression ratio.

[0005] Video encoders and decoders can utilize techniques from several broad categories, which include, for example, motion compensation, transformation, quantization, and entropy coding.

[0006] Video codec technology may include techniques referred to as intra coding. In intra coding, sample values are represented without reference to samples or other data from previously reconstructed reference images. In some video codecs, an image is spatially subdivided into sample blocks. When all sample blocks are coded in the intra mode, the image can be an intra image. Intra images and their derivatives (e.g., independent decoder refresh images) can be used to reset the decoder state and thus can be used as the first image in an encoded video bitstream and video session or as a still image. Samples of an intra block can be transformed, and the transform coefficients can be quantized before entropy coding. Intra prediction can be a technique to minimize sample values in the domain before transformation. In some cases, the smaller the DC value after transformation and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding for a given quantization step.

[0007] For example, traditional intra coding known from coding technologies such as MPEG-2 generations does not use intra prediction. However, some newer video compression technologies include techniques that attempt to use, for example, surrounding sample data and / or metadata obtained during coding and / or decoding of spatially adjacent data that is before the data block in decoding order. Such techniques are referred to hereinafter as "intra prediction" techniques. It should be noted that, at least in some cases, intra prediction uses only reference data from the currently being reconstructed image and not reference data from reference images.

[0008] Intra prediction can have many different forms. When more than one such technique can be used in a given video coding technology, the technique in use can be coded in an intra prediction mode. In some cases, a mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be coded separately or included in the mode codeword. Which codeword is used for a given combination of mode / sub-mode and / or parameters can affect the coding efficiency gain through intra prediction and thus can affect the entropy coding technique used to convert the codeword into a bitstream.

[0009] H.264 introduced a certain intra prediction mode, which was refined in H.265 and further refined in newer coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Adjacent sample values belonging to already available samples can be used to form a prediction block. The sample values of the adjacent samples are copied into the prediction block according to a direction. The reference to the direction used can be coded into the bitstream or can itself be predicted.

[0010] Reference Figure 1A, a subset of 9 prediction directions known from 33 possible prediction directions of H.265 (corresponding to 33 angular modes of 35 intra modes) is depicted in the lower right. The point (101) where the arrows converge represents the sample to be predicted. The arrows indicate the direction along which the predicted sample lies. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper right, at a 45-degree angle to the horizontal direction. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower left of sample (101), at a 22.5-degree angle to the horizontal direction.

[0011] Still referring to Figure 1A , a block of 4×4 samples (indicated by the dashed bold line) is depicted in the upper left. Block (104) includes 16 samples, each labeled with "S", its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is of size 4×4 samples, S44 is located in the lower right corner. Reference samples following a similar numbering scheme are also shown. The reference samples are labeled with "R", their Y position (e.g., row index) and X position (column index) relative to block (104). In H.264 and H.265, the predicted samples are all adjacent to the block being reconstructed; thus, negative values are not needed.

[0012] Intra image prediction can be achieved by copying the reference sample values from adjacent samples in a prediction direction suitable for signaling. For example, assume that the encoded video bitstream includes signaling that indicates, for the block, a prediction direction consistent with arrow (102), i.e., predicting the sample from one or more prediction samples in the upper right, at a 45-degree angle to the horizontal direction. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.

[0013] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample; especially when the directions cannot be evenly separated by 45 degrees.

[0014] As video coding technology develops, the number of possible directions increases. In H.264 (in 2003), nine different directions could be represented. In H.265 (in 2013), it increased to 33 directions, and JEM / VVC / BMS could support up to 65 directions when publicly available. Experiments have been conducted to identify the most likely directions, and some techniques in entropy coding are used to represent those possible directions with a small number of bits, accepting a certain cost for the less likely directions. Additionally, sometimes the direction itself can be predicted from adjacent directions used in already decoded adjacent blocks.

[0015] Figure 1B A schematic diagram (180) is shown, which depicts 65 intra prediction directions according to JEM to illustrate the increase in the number of prediction directions over time.

[0016] The mapping of the intra prediction direction bits representing the direction in the encoded video bitstream may vary depending on the video coding technology; for example, it can range from a simple direct mapping of the prediction direction to the intra prediction mode, mapping to codewords, mapping to complex adaptive schemes involving the most likely modes, and similar techniques. However, in all cases, there may be certain directions that are statistically less likely to occur in video content compared to some other directions. Since the goal of video compression is to reduce redundancy, in a well - functioning video coding technology, those less likely directions will be represented by more bits than the more likely directions.

[0017] Motion compensation can be a lossy compression technique and can involve the following technique: Blocks of sample data from a previously reconstructed image or a part thereof (reference image) are spatially offset along the direction indicated by a motion vector (hereinafter referred to as MV) and are then used to predict a newly reconstructed image or image part. In some cases, the reference image can be the same as the image currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference image being used (the latter can indirectly be the temporal dimension).

[0018] In some video compression techniques, an MV applicable to a certain region of sample data can be predicted based on other MVs, for example, an MV that is related to another region of sample data that is spatially adjacent to the region being reconstructed and whose decoding order is before that of the MV. Doing so can greatly reduce the amount of data required to encode the MV, thereby eliminating redundancy and increasing the compression ratio. MV prediction can work effectively. For example, when encoding an input video signal (referred to as natural video) obtained from a camera, there is the following statistical possibility: a region larger than the region applicable to a single MV moves along a similar direction. Therefore, in some cases, a similar motion vector derived from the MVs of adjacent regions can be used to predict the larger region. This makes the MV found for a given region similar or identical to the MV predicted based on the surrounding MVs. Furthermore, after entropy coding, the MV found for the given region can be represented using fewer bits than the number of bits used when directly encoding the MV. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, for example, due to rounding errors that occur when calculating the predicted value based on multiple surrounding MVs, the MV prediction itself can be lossy.

[0019] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms provided by H.265, the technique hereinafter referred to as "spatial merge" is described herein.

[0020] Reference Figure 2 , the current block (201) includes samples that have been found by the encoder during the motion search process. The samples can be predicted based on a previous block of the same size that has been spatially offset. Instead of directly encoding the MV, the MV can be derived from metadata associated with one or more reference images. For example, using the MV associated with any one of five surrounding samples labeled A0, A1 and B0, B1, B2 (corresponding to 202 to 206 respectively), the MV is derived from the nearest reference image (in decoding order). In H.265, MV prediction can use the predicted value from the same reference image that the adjacent block is using. Summary of the Invention

[0021] Aspects of the present disclosure provide methods and apparatuses for video encoding and / or decoding. In some examples, an apparatus for video decoding includes processing circuitry. The processing circuitry may decode encoded information of a current block in a current picture according to an encoded video bitstream. The encoded information may indicate a string copy mode of the current block. The current block may include a current string. The processing circuitry may determine whether to perform a flipping operation to predict the current string. Based on determining to perform the flipping operation to predict the current string, the processing circuitry may determine an original reference string based on a string vector of the current string. The processing circuitry may generate a flipped reference string by performing the flipping operation on the original reference string, and reconstruct the current string based on the flipped reference string.

[0022] In one embodiment, the flipping operation is one of the following operations: (i) a vertical flipping operation, (ii) a horizontal flipping operation, and (iii) a combined flipping operation. The processing circuitry may generate a flipped reference string by vertically flipping the original reference string based on the flipping operation being a vertical flipping operation. The processing circuitry may generate a flipped reference string by horizontally flipping the original reference string based on the flipping operation being a horizontal flipping operation. The processing circuitry may generate a flipped reference string by vertically and horizontally flipping the original reference string based on the flipping operation being a combined flipping operation.

[0023] In one example, the encoded information includes a string level flag of the current string located after the string length of the current string, and the string level flag indicates whether to perform a flipping operation to predict the current string.

[0024] In one embodiment, the current string is one of a plurality of strings included in the current block. The encoded information may include a block level flag of the current block. A first value of the block level flag may indicate that for each of the plurality of strings, a corresponding flipping operation is used to predict the string, and a second value of the block level flag may indicate that no flipping operation is performed on the plurality of strings. It may be determined to perform the flipping operation to predict the current string based on the block level flag having the first value. In one example, the processing circuitry may determine to prohibit performing the flipping operation on the current string based on the number of the plurality of strings in the current block being greater than a first threshold.

[0025] In one embodiment, the encoded information includes a string level flag of the current string. The processing circuitry may determine to perform the flipping operation to predict the current string based on the string level flag indicating to perform the flipping operation to predict the current string.

[0026] In one embodiment, the processing circuitry may determine to prohibit the flipping operation for the current string based on the string length of the current string being less than a second threshold. The string length of the current string may correspond to the number of samples in the current string.

[0027] In one embodiment, the processing circuitry may determine to prohibit performing the flipping operation on the current string based on the current string having a rectangular shape.

[0028] In one embodiment, the current string has a rectangular shape and includes samples that are not predicted using a flipping operation. The processing circuitry may determine to perform a flipping operation to predict a plurality of samples in the current string that are different from the samples that are not predicted using the flipping operation. In one example, the current string is the current block.

[0029] In one embodiment, the processing circuitry may dequantize the transform coefficients and inverse transform the transform coefficients into the residual of the current block.

[0030] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform methods for video decoding and / or encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0032] Figure 1A is a schematic illustration of an exemplary subset of intra prediction modes.

[0033] Figure 1B is an illustration of exemplary intra prediction directions.

[0034] Figure 2 is a schematic illustration of the current block and its surrounding spatial merge candidates in one example.

[0035] Figure 3 is a schematic illustration of a simplified block diagram of a communication system (300) according to one embodiment.

[0036] Figure 4 is a schematic illustration of a simplified block diagram of a communication system (400) according to one embodiment.

[0037] Figure 5 is a schematic illustration of a simplified block diagram of a decoder according to one embodiment.

[0038] Figure 6 is a schematic illustration of a simplified block diagram of an encoder according to one embodiment.

[0039] Figure 7 shows a block diagram of an encoder according to another embodiment.

[0040] Figure 8 shows a block diagram of a decoder according to another embodiment.

[0041] Figure 9 shows an example of intra block copy according to one embodiment of the present disclosure.

[0042] Figure 10 Shows an example of intra-block copy according to an embodiment of the present disclosure.

[0043] Figure 11 Shows an example of intra-block copy according to an embodiment of the present disclosure.

[0044] Figures 12A to 12D Shows an example of intra-block copy (IBC) according to an embodiment of the present disclosure.

[0045] Figure 13 Shows an example of a spatial category for IBC block vector prediction for a current block according to an embodiment of the present disclosure.

[0046] Figure 14 Shows an example of a string copy pattern according to an embodiment of the present disclosure.

[0047] Figure 15 Shows an example of a horizontal flip operation according to an embodiment of the present disclosure.

[0048] Figure 16 Shows an example of a vertical flip according to an embodiment of the present disclosure.

[0049] Figure 17 Shows an example of a combined flip operation according to an embodiment of the present disclosure.

[0050] Figure 18 Shows an example of a flip operation for a rectangular string according to an embodiment of the present disclosure.

[0051] Figure 19 Shows an example of a flip operation allowed for a current block according to an embodiment of the present disclosure.

[0052] Figure 20 Shows a flowchart outlining a process (2000) according to an embodiment of the present disclosure.

[0053] Figure 21 Is a schematic diagram of a computer system according to an embodiment. Detailed Description

[0054] Figure 3 Shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected by a network (350). In Figure 3In the example, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) may encode video data (such as a video image stream captured by the terminal device (310)) for transmission over the network (350) to another terminal device (320). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to recover the video image, and display the video image based on the recovered video data. Unidirectional data transmission may be common in applications such as media services.

[0055] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which may occur, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) may encode video data (such as a video image stream captured by the terminal device) for transmission over the network (350) to the other of the terminal devices (330) and (340). Each of the terminal devices (330) and (340) may also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), may decode the encoded video data to recover the video image, and may display the video image on an accessible display device based on the recovered video data.

[0056] In Figure 3 the example, the terminal devices (310), (320), (330), and (340) may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (350) represents any number of networks that convey encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise stated below, the architecture and topology of the network (350) may be immaterial to the operation of the present disclosure.

[0057] As an example of an application for the disclosed subject matter, Figure 4Shows the placement of a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0058] A streaming system may include an acquisition subsystem (413), and the acquisition subsystem (413) may include a video source (401) such as a digital camera. The video source (401) creates an uncompressed video image stream (402), for example. In one example, the video image stream (402) includes samples taken by the digital camera. The video image stream (402), depicted as a thick line to emphasize the high data volume, may be processed by an electronic device (420) compared to the encoded video data (404) (or encoded video bitstream). The electronic device (420) includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the lower data volume, may be stored on a streaming server (405) for future use compared to the video image stream (402). One or more streaming client subsystems, such as Figure 4 the client subsystem (406) and the client subsystem (408) in may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410) in an electronic device (430), for example. The video decoder (410) decodes an incoming copy (407) of the encoded video data and produces an output video image stream (411) that can be rendered on a display (412) such as a display screen or other rendering device (not depicted). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstream) may be encoded according to certain video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.

[0059] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0060] Figure 5A block diagram of a video decoder (510) according to an embodiment of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used in place of Figure 4 the video decoder (410) in the example of

[0061] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective consuming entities (not depicted). The receiver (531) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be located external to the video decoder (510) (not depicted). In other cases, a buffer memory (not depicted) may be provided external to the video decoder (510) to, for example, prevent network jitter, and another buffer memory (515) may be provided inside the video decoder (510) to, for example, handle presentation timing. When the receiver (531) receives data from a storage / forwarding device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be needed, or the buffer memory may be made smaller. For use on a service packet network such as the Internet, a buffer memory (515) may be needed, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least partially in the operating system or a similar element (not depicted) external to the video decoder (510).

[0062] The video decoder (510) may include a parser (520) to reconstruct symbols (521) according to the encoded video sequences. The categories of these symbols include information for managing the operation of the video decoder (510) and potential information for controlling a rendering device such as a rendering device (512) (e.g., a display screen), which is not an integral part of the electronic device (530) but may be coupled to the electronic device (530), as Figure 5As shown. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a video usability information (VUI) parameter set segment (not depicted). The parser (520) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be according to a video coding technology or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0063] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0064] Depending on the type of the encoded video image or a part thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different units. Which units are involved and how they are involved may be controlled by the parser (520) through subgroup control information parsed from the encoded video sequence. For clarity, such subgroup control information flows between the parser (520) and the multiple units below are not depicted.

[0065] In addition to the functional blocks already mentioned, the video decoder (510) may conceptually be subdivided into multiple functional units as described below. In practical implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the multiple functional units below.

[0066] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives the quantized transform coefficients as symbols (521) and control information from the parser (520), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output a block including sample values, and the sample values may be input into the aggregator (555).

[0067] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed image but may use prediction information from a previously reconstructed part of the current image. Such prediction information may be provided by the intra-image prediction unit (552). In some cases, the intra-image prediction unit (552) uses the surrounding reconstructed information extracted from the current image buffer (558) to generate a block having the same size and shape as the block being reconstructed. For example, the current image buffer (558) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (555) adds, on a per-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0068] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (553) may access the reference image memory (557) to extract samples for prediction. After motion-compensating the extracted samples according to the sign (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (which is referred to as the residual sample or residual signal in this case), thereby generating output sample information. The extraction of the prediction samples by the motion compensation prediction unit (553) from an address within the reference image memory (557) may be controlled by a motion vector, and the motion vector may be provided in the form of a sign (521) for use by the motion compensation prediction unit (553), and the sign (521) may have, for example, an X component, a Y component, and a reference image component. Motion compensation may also include interpolation of the sample values extracted from the reference image memory (557), a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.

[0069] The output samples of the aggregator (555) may be subject to various loop filtering techniques in the loop filter unit (556). The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and that may be available to the loop filter unit (556) as a sign (521) from the parser (520). However, the video compression technique may also respond to meta-information obtained during the decoding of a previous (in decoding order) part of the encoded image or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0070] The output of the loop filter unit (556) may be a sample stream that may be output to the rendering device (512) and stored in the reference image memory (557) for subsequent inter-image prediction.

[0071] Once fully reconstructed, some of the encoded images can be used as reference images for future prediction. For example, once the encoded image corresponding to the current image is fully reconstructed and the encoded image is identified as a reference image (e.g., by a parser (520)), the current image buffer (558) can become part of the reference image memory (557), and a new current image buffer can be reallocated before starting to reconstruct subsequent encoded images.

[0072] The video decoder (510) can perform decoding operations according to a predetermined video compression technique in a standard such as the ITU-T H.265 recommendation. The encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, a profile can select some tools from all the tools available in the video compression technique or standard as the only tools available under that profile. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference image size, etc. In some cases, the limits set by the level can be further defined by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.

[0073] In one embodiment, the receiver (531) can receive additional (redundant) data when receiving the encoded video. This additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (510) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0074] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present disclosure is shown. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of Figure 4 the video encoder (403) in the example of

[0075] The video encoder (603) can receive from a video source (601) (not Figure 6A part of the electronic device (620) in the example receives video samples. The video source (601) can acquire video images to be encoded by the video encoder (603). In another example, the video source (601) is a part of the electronic device (620).

[0076] The video source (601) can provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603). The digital video sample stream can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) can be a camera that acquires local image information as a video sequence. Video data can be provided as multiple individual images that are given motion when viewed in sequence. The images themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0077] According to one embodiment, the video encoder (603) can encode and compress the images of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to the other functional units. For clarity, the couplings are not depicted in the figure. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions that relate to optimizing the video encoder (603) for a certain system design.

[0078] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an overly simplified description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input image to be encoded and reference images) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to the way a (remote) decoder creates sample data (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference image memory (634). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the contents in the reference image memory (634) are also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference image samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference image synchronization principle (and the drift that occurs in cases where synchronization cannot be maintained, e.g., due to channel errors) is also used in some related technologies.

[0079] The operation of the "local" decoder (633) can be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 5 the video decoder (510). However, briefly referring additionally to Figure 5 , since the symbols are available and the entropy encoder (645) and the parser (520) are capable of losslessly encoding / decoding the symbols into the encoded video sequence, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0080] At this point, it can be observed that any decoder techniques other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of the encoder techniques can be simplified because the encoder techniques are reciprocal to the decoder techniques described comprehensively. More detailed descriptions are needed only in certain areas and are provided below.

[0081] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive coding, which predictively encodes an input image by referring to one or more previously encoded images in the video sequence designated as "reference images". In this way, the encoding engine (632) encodes the difference between a pixel block of the input image and a pixel block of the reference image, which can be selected as the prediction reference for the input image.

[0082] The local video decoder (633) may decode the encoded video data of an image that can be designated as a reference image based on the symbols created by the source encoder (630). The operation of the encoding engine (632) may advantageously be a lossy process. When the encoded video data can be decoded in a video decoder ( Figure 6 not shown), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The local video decoder (533) replicates the decoding process that may be performed by the video decoder on the reference image and may store the reconstructed reference image in the reference image cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference image that has common content (no transmission errors) with the reconstructed reference image that will be obtained by the remote video decoder.

[0083] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new image to be encoded, the predictor (635) may search the reference image memory (634) for sample data (as a candidate reference pixel block) or some metadata, such as reference image motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new image. The predictor (635) may operate on a per-pixel block basis of sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input image may have prediction references taken from multiple reference images stored in the reference image memory (634).

[0084] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0085] The outputs of all the above functional units may be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0086] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over a communication channel (660), which may be a hardware / software link to a storage device that can store the encoded video data. The transmitter (640) may merge the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0087] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding image. For example, any of the following image types can typically be assigned to an image:

[0088] An intra-frame image (I-image), which can be an image that can be encoded and decoded without using any other image in the sequence as a prediction source. Some video codecs allow different types of intra-frame images, including for example independent decoder refresh ("IDR") images. Those skilled in the art are aware of the variants of I-images and their corresponding applications and characteristics.

[0089] A predictive image (P-image), which can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, where the intra-frame prediction or inter-frame prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0090] A bi-predictive image (B-image), which can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, where the intra-frame prediction or inter-frame prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive images can use more than two reference images and associated metadata for reconstructing a single block.

[0091] The source image can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (encoded) blocks, which are determined by the encoding assignment applied to the corresponding image of the block. For example, blocks of an I-image can be non-prediction-encoded, or blocks of an I-image can be prediction-encoded (spatial prediction or intra-frame prediction) with reference to the encoded blocks of the same image. Pixel blocks of a P-image can be prediction-encoded by spatial prediction or by temporal prediction with reference to one previously encoded reference image. Blocks of a B-image can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference images.

[0092] The video encoder (603) can perform encoding operations according to a predetermined video encoding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including prediction encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard used.

[0093] In one embodiment, a transmitter (640) may send additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set segments, etc.

[0094] The captured video may be a plurality of source images (video pictures) in a time series. Intra-picture prediction (often simplified to intra-prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded is partitioned into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension identifying the reference picture.

[0095] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, e.g., a first reference picture and a second reference picture that are before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0096] In addition, merge mode techniques may be used in inter-picture prediction to improve coding efficiency.

[0097] According to some embodiments of the present disclosure, predictions such as inter - frame image prediction and intra - frame image prediction are performed on a block - by - block basis. For example, according to the HEVC standard, an image in a video image sequence is divided into coding tree units (CTUs) for compression. CTUs in an image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively split into one or more coding units (CUs) in a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In one example, each CU is analyzed to determine the prediction type for the CU, such as an inter - frame prediction type or an intra - frame prediction type. Depending on temporal and / or spatial predictability, a CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one embodiment, the prediction operation in encoding (encoding / decoding) is performed on a prediction - block basis. Taking the luminance prediction block as an example of a prediction block, the prediction block includes a matrix of values for pixels (e.g., luminance values), and the pixels are, for example, 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0098] Figure 7 FIG. shows a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video image in a video image sequence and encode the processing block into an encoded image that is part of an encoded video sequence. In one example, the video encoder (703) is used in place of Figure 4 the video encoder (403) in the example of

[0099] In the HEVC example, a video encoder (703) receives a matrix of sample values for processing a block, such as a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion optimization to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to best encode the processing block. When encoding the processing block in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded image; and when encoding the processing block in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction or bi-prediction techniques respectively to encode the processing block into an encoded image. In some video coding techniques, the merge mode may be an inter-picture prediction sub-mode, in which a motion vector is derived from one or more motion vector predictors without resorting to encoded motion vector components external to the predictor. In some other video coding techniques, there may be motion vector components applicable to the subject block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0100] In Figure 7 the example of, the video encoder (703) includes, as Figure 7 shown, an inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together.

[0101] The inter-frame encoder (730) is configured to receive samples of a current block (such as a processing block), compare the block with one or more reference blocks in a reference image (such as blocks in a previous image and a later image), generate inter-frame prediction information (such as redundancy information description, motion vector, merge mode information according to inter-frame coding techniques), and calculate an inter-frame prediction result (such as a predicted block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference image is a decoded reference image decoded based on encoded video information.

[0102] The intra-frame encoder (722) is configured to receive samples of a current block (such as a processing block), compare the block with encoded blocks in the same image in some cases, generate quantized coefficients after transformation, and also generate intra-frame prediction information in some cases (such as intra-frame prediction direction information according to one or more intra-frame coding techniques). In one example, the intra-frame encoder (722) also calculates an intra-frame prediction result (such as a predicted block) based on the intra-frame prediction information and a reference block in the same image.

[0103] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of a block and provides a control signal to the switch (726) based on the mode. For example, when the mode is an intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; and when the mode is an inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0104] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded image, and in some examples, the decoded image can be buffered in a memory circuit (not shown) and used as a reference image.

[0105] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information according to a suitable standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include the general control data, the selected prediction information (e.g., intra prediction information or inter prediction information), the residual information, and other suitable information in the bitstream. It should be noted that according to the disclosed subject matter, there is no residual information when encoding a block in the merge submode of the inter mode or the bi - directional prediction mode.

[0106] Figure 8A diagram showing a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed image. In one example, the video decoder (810) is used in place of Figure 4 the video decoder (410) in the example of

[0107] In Figure 8 the example of Figure 8 the video decoder (810) includes an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) coupled together as shown.

[0108] The entropy decoder (871) may be configured to reconstruct certain symbols based on the encoded image, and these symbols represent the syntax elements that make up the encoded image. Such symbols may include, for example, the mode for encoding a block (e.g., intra-frame mode, inter-frame mode, bi-prediction mode, merge sub-modes of the latter two, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can identify certain samples or metadata for use by the intra-frame decoder (872) or the inter-frame decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, etc. In one example, when the prediction mode is inter-frame or bi-prediction mode, the inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is intra-frame prediction type, the intra-frame prediction information is provided to the intra-frame decoder (872). The residual information may be inverse quantized and provided to the residual decoder (873).

[0109] The inter-frame decoder (880) is configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0110] The intra-frame decoder (872) is configured to receive the intra-frame prediction information and generate a prediction result based on the intra-frame prediction information.

[0111] The residual decoder (873) is configured to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (for including quantization parameter (QP)), and this information may be provided by the entropy decoder (871) (the data path is not depicted as this is only low-volume control information).

[0112] The reconstruction module (874) is configured to combine in the spatial domain the residual output by the residual decoder (873) with the prediction result (which may be output by an inter-frame prediction module or an intra-frame prediction module) to form a reconstructed block, which may be part of a reconstructed image, and the reconstructed image may in turn be part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations may be performed to improve the visual quality.

[0113] It should be noted that any suitable technique may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In one embodiment, one or more integrated circuits may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810).

[0114] Aspects of the present disclosure provide techniques for performing string matching by a flipping operation.

[0115] Block-based compensation can be used for both inter-frame prediction and intra-frame prediction. For inter-frame prediction, block-based compensation from different images is called motion compensation. For example, in intra-frame prediction, block-based compensation can also be performed from previously reconstructed regions within the same image. Block-based compensation from the reconstructed regions within the same image can be referred to as intra-image block compensation, current picture reference (CPR), or intra-block copy (IBC). The displacement vector indicating the offset between the current block and the reference block (also called the prediction block) within the same image is called the block vector (BV), where the current block can be encoded / decoded based on the reference block. The motion vector in motion compensation can be any value (positive or negative in the x or y direction). Different from the motion vector in motion compensation, the BV has several constraints to ensure that the reference block is available and reconstructed. Additionally, in some examples, for the consideration of parallel processing, some reference regions such as tile boundaries, slice boundaries, and / or wavefront trapezoid boundaries are excluded.

[0116] The encoding of the block vector can be explicit or implicit. In the explicit mode, the BV difference between the block vector and its predictor can be signaled. In the implicit mode, the block vector is recovered from the predictor (called the block vector predictor) in a manner similar to the motion vector in the merge mode, without having to use the BV difference. The explicit mode can be referred to as the non-merge BV prediction mode. The implicit mode can be referred to as the merge BV prediction mode.

[0117] In some implementations, the resolution of the block vector is restricted to integer positions. In other systems, the block vector is allowed to point to fractional positions.

[0118] In some examples, a block-level flag (e.g., the IBC flag) can be used to signal the use of intra-block copy at the block level. In one embodiment, when the current block is not explicitly encoded, the block-level flag is signaled. In some examples, a reference index method can be used to signal the use of intra-block copy at the block level. Then, the current image being decoded is treated as a reference image or a special reference image. In one example, this special reference image is placed at the last position in the reference image list. This special reference image is also managed together with other temporal reference images in a buffer such as the decoded picture buffer (DPB).

[0119] There can be variants of the IBC mode. In one example, the IBC mode is regarded as a third mode different from the intra prediction mode and the inter prediction mode. Thus, in the implicit mode (or merge mode) and the explicit mode, BV prediction is separated from the regular inter mode. A separate merge candidate list can be defined for the IBC mode, where the entries in the separate merge candidate list are BVs. Similarly, in one example, in the IBC explicit mode, the BV prediction candidate list only includes BVs. The general rule applied to the two lists (i.e., the separate merge candidate list and the BV prediction candidate list) is that in terms of the candidate derivation process, the two lists can follow the same logic as the merge candidate list used in the regular merge mode or the AMVP predictor list used in the regular AMVP mode. For example, five spatially adjacent positions (e.g., Figure 2 A0, A1 and B0, B1, B2 in

[0120] are accessed for the HEVC or VVC inter merge mode for the IBC mode, for example, to derive the separate merge candidate list for the IBC mode.

[0121] Figure 9Shows an example of intra block copy according to an embodiment of the present disclosure. The current image (900) is to be reconstructed under decoding. The current image (900) includes a reconstructed region (910) (gray region) and a region to be decoded (920) (white region). The current block (930) is reconstructed by the decoder. The current block (930) can be reconstructed from a reference block (940) in the reconstructed region (910). The position offset between the reference block (940) and the current block (930) is called the block vector (950) (or BV (950)). In Figure 9 the example, the search range (960) is located within the reconstructed region (910), the reference block (940) is located within the search range (960), and the block vector (950) is constrained to point to the reference block (940) within the search range (960).

[0122] Various constraints can be applied to the BV and / or the search range. In one embodiment, the search range of the current block being reconstructed in the current CTB is constrained to be within the current CTB.

[0123] In one embodiment, the effective memory requirement for storing the reference samples to be used in intra block copy is one CTB size. In one example, the CTB size is 128×128 samples. The current CTB includes the current region being reconstructed. The current region has a size of 64×64 samples. Since the reference memory can also store the reconstructed samples in the current region, when the reference memory size is equal to the CTB size of 128×128 samples, the reference memory can store an additional 3 regions of 64×64 samples. Therefore, the search range can include certain parts of the previously reconstructed CTB, while the total memory requirement for storing the reference samples remains the same (e.g., one CTB size of 128×128 samples, or a total of 4 reference samples of 64×64). In one example, the previously reconstructed CTB is the left neighbor of the current CTB, such as Figure 10 shown in.

[0124] Figure 10Shows an example of intra-block copy according to an embodiment of the present disclosure. The current image (1001) includes a currently reconstructed current CTB (1015) and a previously reconstructed CTB (1010) that is the left neighbor of the current CTB (1015). The CTBs in the current image (1001) have a CTB size (e.g., 128×128 samples) and a CTB width (e.g., 128 samples). The current CTB (1015) includes 4 regions (1016)-(1019), where the current region (1016) is being reconstructed. The current region (1016) includes a plurality of coding blocks (1021)-(1029). Similarly, the previously reconstructed CTB (1010) includes 4 regions (1011)-(1014). The coding blocks (1021)-(1025) have been reconstructed, the current block (1026) is being reconstructed, and the coding blocks (1026)-(1027) and regions (1017)-(1019) are to be reconstructed.

[0125] The current region (1016) has a collocated region (i.e., region (1011)) located in the previously reconstructed CTB (1010). The relative position of the collocated region (1011) with respect to the previously reconstructed CTB (1010) may be the same as the relative position of the current region (1016) with respect to the current CTB (1015). In Figure 10 the example shown, the current region (1016) is the upper left region in the current CTB (1015), and thus, the collocated region (1011) is also the upper left region in the previously reconstructed CTB (1010). Since the position of the previously reconstructed CTB (1010) is offset from the position of the current CTB (1015) by the CTB width, the position of the collocated region (1011) is offset from the position of the current region (1016) by the CTB width.

[0126] In one embodiment, the collocated region of the current region (1016) is located in the previously reconstructed CTB, where the position of the previously reconstructed CTB is offset from the position of the current CTB (1015) by one or more CTB widths, and thus, the position of the collocated region is also offset from the position of the current region (1016) by the corresponding one or more CTB widths. The position of the collocated region may be shifted leftward, upward, etc. from the current region (1016).

[0127] As described above, the size of the search range of the current block (1026) is constrained by the CTB size. In Figure 10In the example, the search range may include regions (1012)-(1014) in the previously reconstructed CTB (1010) and a part of the currently reconstructed region (1016), such as coding blocks (1021)-(1025). The search range further excludes the juxtaposed region (1011), such that the size of the search range is within the CTB size. Refer to Figure 10 , the reference block (1091) is located in the region (1014) of the previously reconstructed CTB (1010). The block vector (1020) indicates the offset between the current block (1026) and the corresponding reference block (1091). The reference block (1091) is within the search range.

[0128] Figure 10 The example shown may be suitable for other scenarios, in which the current region is located at another position in the current CTB (1015). In one example, when the current block is in the region (1017), the juxtaposed region of the current block is the region (1012). Thus, the search range may include regions (1013)-(1014), the region (1016), and a part of the reconstructed region (1017). The search range further excludes the region (1011) and the juxtaposed region (1012), such that the size of the search range is within the CTB size. In one example, when the current block is in the region (1018), the juxtaposed region of the current block is the region (1013). Thus, the search range may include the region (1014), regions (1016)-(1017), and a part of the reconstructed region (1018). The search range further excludes regions (1011)-(1012) and the juxtaposed region (1013), such that the size of the search range is within the CTB size. In one example, when the current block is in the region (1019), the juxtaposed region of the current block is the region (1014). Thus, the search range may include regions (1016)-(1018) and a part of the reconstructed region (1019). The search range further excludes the previously reconstructed CTB (1010), such that the size of the search range is within the CTB size.

[0129] In the above description, the reference block may be located in the previously reconstructed CTB (1010) or the current CTB (1015).

[0130] Figure 11Shows an example of intra block copy according to an embodiment of the present disclosure. The current image (1101) includes a currently reconstructed current CTB (1115) and a previously reconstructed CTB (1110) that is the left neighbor of the current CTB (1115). The CTBs in the current image (1101) have a CTB size and a CTB width. The current CTB (1115) includes four regions (1116)-(1119), where the current region (1116) is being reconstructed. The current region (1116) includes a plurality of coding blocks (1121)-(l129). Similarly, the previously reconstructed CTB (1110) includes four regions (1111)-(1114). The currently reconstructed current block (1121) is first reconstructed in the current region (1116), and the coding blocks (1122)-(1129) are to be reconstructed. In one example, the CTB size is 128×128 samples, and each of the regions (1111)-(l114) and (1116)-(1119) is 64×64 samples. The reference memory size is equal to the CTB size and is 128×128 samples. Thus, when the search range is bounded by the reference memory size, the search range includes three regions and a part of an additional region.

[0131] Similar to that described with reference to Figure 10 the current region (1116) has a collocated region (i.e., region (1111)) located in the previously reconstructed CTB (1110). The reference block of the current block (1121) may be located in the region (1111), and thus, the search range may include the regions (1111)-(1114). For example, when the reference block is located in the region (1111), the collocated region of the reference block is the region (1116), where the samples in the region (1116) are not reconstructed before the reconstruction of the current block (1121). However, as described with reference to Figure 10 after the reconstruction of the coding block (1121), for example, the region (1111) can no longer be included in the search range for reconstructing the coding block (1122). Therefore, tight synchronization and timing control of the reference storage buffer are required and may be challenging.

[0132] According to some embodiments, when the current block is first reconstructed in the current region of the current CTB, the search range may exclude the collocated region of the current region that is located in the previously reconstructed CTB, where the current CTB and the previously reconstructed CTB are located in the same current image. A block vector may be determined such that the reference block is located in the search range excluding the collocated region in the previously reconstructed CTB. In one embodiment, the search range includes coding blocks that are reconstructed in decoding order after the collocated region and before the current block.

[0133] In the following description, the CTB size may vary, and the maximum CTB size is set to be the same as the reference memory size. In one example, the reference memory size or the maximum CTB size is 128×128 samples. The description may be suitably applied to other reference memory sizes or maximum CTB sizes.

[0134] In one embodiment, the CTB size is equal to the reference memory size. The previously reconstructed CTB is the left neighbor of the current CTB, the position of the juxtaposed region is offset from the position of the current region by the CTB width, and the coded blocks in the search range are located in at least one of the current CTB and the previously reconstructed CTB.

[0135] Figures 12A to 12D An example of intra-block copy according to an embodiment of the present disclosure is shown. Refer to Figures 12A to 12D , the current image (1201) includes the current CTB (1215) being reconstructed and the previously reconstructed CTB (1210) that is the left neighbor of the current CTB (1215). The CTBs in the current image (1201) have a CTB size and a CTB width. The current CTB (1215) includes four regions (1216)-(1219). Similarly, the previously reconstructed CTB (1210) includes four regions (1211)-(1214). In one embodiment, the CTB size is the maximum CTB size and is equal to the reference memory size. In one example, the CTB size and the reference memory size are 128×128 samples. Thus, each of the regions (1211)-(1214) and (1216)-(1219) has a size of 64×64 samples.

[0136] In Figures 12A to 12D the example shown, the current CTB (1215) includes an upper left region, an upper right region, a lower left region, and a lower right region corresponding to the regions (1216)-(1219) respectively. The previously reconstructed CTB (1210) includes an upper left region, an upper right region, a lower left region, and a lower right region corresponding to the regions (1211)-(1214) respectively.

[0137] Refer to Figure 12A, the current region (1216) is being reconstructed. The current region (1216) may include a plurality of coding blocks (1221)-(1229). The current region (1216) has a collocated region (i.e., region (1211)) located in the previously reconstructed CTB (1210). The search range for one of the coding blocks (1221)-(1229) to be reconstructed may exclude the collocated region (1211). The search range may include regions (1212)-(1214) of the previously reconstructed CTB (1210), and regions (1212)-(1214) are reconstructed after the collocated region (1211) and before the current region (1216) in the decoding order.

[0138] Reference Figure 12A , the position of the collocated region (1211) is offset from the position of the current region (1216) by the CTB width (e.g., 128 samples). For example, the position of the collocated region (1211) is shifted 128 samples to the left from the position of the current region (1216).

[0139] Refer again to Figure 12A , when the current region (1216) is the upper left region of the current CTB (1215), the collocated region (1211) is the upper left region of the previously reconstructed CTB (1210), and the search region excludes the upper left region of the previously reconstructed CTB.

[0140] Reference Figure 12B , the current region (1217) is being reconstructed. The current region (1217) may include a plurality of coding blocks (1241)-(1249). The current region (1217) has a collocated region (i.e., region (1212)) located in the previously reconstructed CTB (1210). The search range for one of the coding blocks (1241)-(1249) may exclude the collocated region (1212). The search range includes regions (1213)-(1214) of the previously reconstructed CTB (1210) and region (1216) in the current CTB (1215), and these regions are reconstructed after the collocated region (1212) and before the current region (1217). Due to the constraint of the reference memory size (i.e., the size of one CTB), the search range further excludes region (1211). Similarly, the position of the collocated region (1212) is offset from the position of the current region (1217) by the CTB width (e.g., 128 samples).

[0141] In Figure 12B 's example, the current region (1217) is the upper right region of the current CTB (1215), the collocated region (1212) is also the upper right region of the previously reconstructed CTB (1210), and the search region excludes the upper right region of the previously reconstructed CTB (1210).

[0142] Reference Figure 12C , the current region (1218) is being reconstructed. The current region (1218) may include a plurality of coding blocks (1261)-(1269). The current region (1218) has a collocated region (i.e., region (1213)) located in the previously reconstructed CTB (1210). The search range for one of the plurality of coding blocks (1261)-(1269) may exclude the collocated region (1213). The search range includes the region (1214) of the previously reconstructed CTB (1210) and the regions (1216)-(1217) in the current CTB (1215) that were reconstructed after the collocated region (1213) and before the current region (1218). Similarly, due to the reference memory size constraint, the search range further excludes the regions (1211)-(1212). The position of the collocated region (1213) is offset from the position of the current region (1218) by the CTB width (e.g., 128 samples). In Figure 12C the example of, when the current region (1218) is the lower left region of the current CTB (1215), the collocated region (1213) is also the lower left region of the previously reconstructed CTB (1210), and the search region excludes the lower left region of the previously reconstructed CTB (1210).

[0143] Reference Figure 12D , the current region (1219) is being reconstructed. The current region (1219) may include a plurality of coding blocks (1281)-(1289). The current region (1219) has a collocated region (i.e., region (1214)) located in the previously reconstructed CTB (1210). The search range for one of the plurality of coding blocks (1281)-(1289) may exclude the collocated region (1214). The search range includes the regions (1216)-(1218) in the current CTB (1215) that were reconstructed after the collocated region (1214) and before the current region (1219) in decoding order. Due to the reference memory size constraint, the search range excludes the regions (1211)-(1213), so the search range excludes the previously reconstructed CTB (1210). Similarly, the position of the collocated region (1214) is offset from the position of the current region (1219) by the CTB width (e.g., 128 samples). In Figure 12D the example of, when the current region (1219) is the lower right region of the current CTB (1215), the collocated region (1214) is also the lower right region of the previously reconstructed CTB (1210), and the search region excludes the lower right region of the previously reconstructed CTB (1210).

[0144] Return reference Figure 2, the MVs associated with five positions (labeled A0, A1 and B0, B1, B2 (corresponding to 202 to 206 respectively)) can be referred to as spatial merge candidates. A candidate list (e.g., a merge candidate list) can be formed based on the spatial merge candidates. Any suitable order can be used to form the candidate list from multiple positions. In one example, the order can be A0, B0, B1, A1 and B2, where A0 is the first and B2 is the last. In one example, the order can be A1, B1, B0, A0 and B2, where A1 is the first and B2 is the last.

[0145] According to some embodiments, the motion information of the previously encoded block of the current block (e.g., the coding block (CB) or the current CU) can be stored in a history-based motion vector prediction (HMVP) buffer (e.g., a table) to provide motion vector prediction (MVP) candidates (also referred to as HMVP candidates) for the current block. The HMVP buffer can include one or more HMVP candidates and can be maintained during the encoding / decoding process. In one example, the HMVP candidates in the HMVP buffer correspond to the motion information of the previously encoded block. The HMVP buffer can be used in any suitable encoder and / or decoder. The HMVP candidates can be added to the merge candidate list, after the spatial MVP and the TMVP.

[0146] When a new CTU (or new CTB) row is encountered, the HMVP buffer can be reset (e.g., cleared). When there is a non-sub-block inter-coded block, the associated motion information can be added to the last entry of the HMVP buffer as a new HMVP candidate.

[0147] In one example, such as in VTM 3, the buffer size of the HMVP buffer (denoted by S) is set to 6, indicating that up to 6 HMVP candidates can be added to the HMVP buffer. In some embodiments, the HMVP buffer can operate according to the first-in-first-out (FIFO) rule. Therefore, the first motion information (or HMVP candidate) stored in the HMVP buffer is the first motion information (or HMVP candidate) to be removed from the HMVP buffer, e.g., when the HMVP buffer is full. When a new HMVP candidate is inserted into the HMVP buffer, a constrained FIFO rule can be utilized, where a redundancy check is first applied to determine whether there is an identical or similar HMVP candidate in the HMVP buffer. If it is determined that there is an identical or similar HMVP candidate in the HMVP buffer, the identical or similar HMVP candidate can be removed from the HMVP buffer, and the remaining HMVP candidates in the HMVP buffer can be moved forward.

[0148] The HMVP candidates can be used during the merged candidate list construction process, e.g., in the merge mode. The most recently stored HMVP candidates in the HMVP buffer can be checked in order and inserted into the merged candidate list after the TMVP candidates. Redundancy checking can be applied to the HMVP candidates with respect to the spatial or temporal merged candidates in the merged candidate list. The description can be suitably applied to the AMVP mode to construct the AMVP candidate list.

[0149] To reduce the number of redundancy checking operations, the following simplifications can be used.

[0150] (i) The number of HMVP candidates used to generate the merged candidate list can be set to (N <= 4)? M :

[0151] (8 - N). N indicates the number of existing candidates in the merged candidate list, and M indicates the number of available HMVP candidates in the HMVP buffer. When the number of existing candidates (N)

[0152] in the merged candidate list is less than or equal to 4, the number of HMVP candidates used to generate the merged candidate list is equal to M. Otherwise, the number of HMVP candidates used to generate the merged candidate list is equal to (8 - N).

[0153] (ii) When the total number of available merged candidates reaches the maximum allowed merged candidates minus 1, the merged candidate list construction process from the HMVP buffer is terminated.

[0154] When the IBC mode operates in a mode different from the inter prediction mode, a simplified BV derivation process for the IBC mode can be used. A history-based block vector prediction buffer (referred to as the HBVP buffer) can be used to perform BV prediction. The HBVP buffer can be used to store BV information (e.g., BV) of previously encoded blocks of the current block (e.g., CB or CU) in the current picture. In one example, the HBVP buffer is a history buffer different from other buffers (e.g., the HMVP buffer). The HBVP buffer can be a table.

[0155] The HBVP buffer can provide BV predictor (BVP) candidates (also referred to as HBVP candidates) for the current block. The HBVP buffer (e.g., a table) can include one or more HBVP candidates and can be maintained during the encoding / decoding process. In one example, the HBVP candidates in the HBVP buffer correspond to the BV information of previously encoded blocks in the current image. The HBVP buffer can be used in any suitable encoder and / or decoder. The HBVP candidates can be added to a merge candidate list configured for BV prediction, after the BV of one or more spatially adjacent blocks of the current block. The merge candidate list configured for BV prediction can be used for both the merged BV prediction mode and / or the non-merged BV prediction mode.

[0156] When a new CTU (or new CTB) row is encountered, the HBVP buffer can be reset (e.g., cleared).

[0157] In one example, such as in VVC, the buffer size of the HBVP buffer is set to 6, indicating that up to 6 HBVP candidates can be added to the HBVP buffer. In some embodiments, the HBVP buffer can operate according to the FIFO rule. Thus, the first BV information (or HBVP candidate) stored in the HBVP buffer is the first BV information (or HBVP candidate) removed from the HBVP buffer, e.g., when the HBVP buffer is full. When a new HBVP candidate is inserted into the HBVP buffer, a constrained FIFO rule can be utilized, where a redundancy check is first applied to determine whether there is an identical or similar HBVP candidate in the HBVP buffer. If it is determined that there is an identical or similar HBVP candidate in the HBVP buffer, the identical or similar HBVP candidate can be removed from the HBVP buffer, and the remaining HBVP candidates in the HBVP buffer can be moved forward.

[0158] For example, in the merged BV prediction mode, the HBVP candidates can be used in the merge candidate list construction process. The most recently stored HBVP candidates in the HBVP buffer can be checked in order, and the most recently stored HBVP candidates can be inserted into the merge candidate list, after the spatial candidates. A redundancy check can be applied to the HBVP candidates relative to the spatial merge candidates in the merge candidate list.

[0159] In one embodiment, the HBVP buffer is established to store one or more BV information of one or more previously encoded blocks encoded in the IBC mode. The one or more BV information can include one or more BVs of one or more previously encoded blocks encoded in the IBC mode. In addition, each of the one or more BV information can include the auxiliary information (or additional information) of the corresponding previously encoded block encoded in the IBC mode, such as block size, block position, etc.

[0160] In class-based, history-based block vector prediction (also referred to as CBVP), for a current block, one or more BV messages in the HBVP buffer that meet certain conditions can be classified into corresponding classes (also referred to as categories), thus forming the CBVP buffer. In one example, each BV message in the HBVP buffer is for a corresponding previously encoded block encoded, for example, by the IBC mode. The BV message of the previously encoded block can include BV, block size, block position, etc. The previously encoded block has a block width, a block height, and a block area. The block area can be the product of the block width and the block height. In one example, the block size is represented by the block area. The block position of the previously encoded block can be represented by the upper left corner of the previously encoded block (e.g., a 4×4 region in the upper left corner) or the upper left sample.

[0161] Figure 13 An example of a spatial category for IBCBV prediction for a current block (e.g., CB, CU) (1310) according to an embodiment of the present disclosure is shown. The left region (1302) can be located to the left of the current block (1310). The BV message of the previously encoded block having a corresponding block position in the left region (1302) can be referred to as a left candidate or a left BV candidate. The top region (1303) can be located above the current block (1310). The BV message of the previously encoded block having a corresponding block position in the top region (1303) can be referred to as a top candidate or a top BV candidate. The upper left region (1304) can be located above and to the left of the current block (1310). The BV message of the previously encoded block having a corresponding block position in the upper left region (1304) can be referred to as an upper left candidate or an upper left BV candidate. The upper right region (1305) can be located above and to the right of the current block (1310). The BV message of the previously encoded block having a corresponding block position in the upper right region (1305) can be referred to as an upper right candidate or an upper right BV candidate. The lower left region (1306) can be located below and to the left of the current block (1310). The BV message of the previously encoded block having a corresponding block position in the lower left region (1306) can be a lower left candidate or a lower left BV candidate. Other types of spatial categories can also be defined and used in the CBVP buffer.

[0162] If the BV message of the previously encoded block meets the following conditions, the BV message can be classified into the corresponding class (or category).

[0163] (i) Category 0: The block size (e.g., block area) is greater than or equal to a threshold (e.g., 64 pixels).

[0164] (ii) Category 1: The occurrence rate (or frequency) of BV is greater than or equal to 2. The occurrence rate of BV may refer to the number of times BV is used to predict a previous coded block. When using a pruning process to form the CBVP buffer, when BV is used multiple times in predicting a previous coded block, BV can be stored in one entry (instead of having the same BV in multiple entries). The occurrence rate of BV can be recorded.

[0165] (iii) Category 2: The block position is in the left region (1302), where a part of the previous coded block (e.g., the upper left 4×4 region) is to the left of the current block (1310). The previous coded block can be within the left region (1302). Alternatively, the previous coded block can span multiple regions including the left region (1302)

[0166] and the block position is in the left region (1302).

[0167] (iv) Category 3: The block position is in the top region (1303), where a part of the previous coded block (e.g., the upper left 4×4 region) is above the current block (1310). The previous coded block can be within the top region (1303). Alternatively, the previous coded block can span multiple regions including the top region (1303)

[0168] and the block position is in the top region (1303).

[0169] (v) Category 4: The block position is in the upper left region (1304), where a part of the previous coded block (e.g., the upper left 4×4 region) is in the upper left side of the current block (1310). The previous coded block can be within the upper left region (1304). Alternatively, the previous coded block can span multiple regions including the upper left region (1304)

[0170] and the block position is in the upper left region (1304).

[0171] (vi) Category 5: The block position is in the upper right region (1305), where a part of the previous coded block (e.g., the upper left 4×4 region) is in the upper right side of the current block (1310). The previous coded block can be within the upper right region (1305). Alternatively, the previous coded block can span multiple regions including the upper right region (1305)

[0172] and the block position is in the upper right region (1305).

[0173] (vii) Category 6: The block position is in the lower left region (1306), where a part of the coded block (e.g., the upper left 4×4 region) is on the lower left side of the current block (1310). The previously coded block can be within the lower left region (1306). Alternatively, the previously coded block can span multiple regions including the lower left region (1306)

[0174] where the block position is in the lower left region (1306).

[0175] For each category (or class), the BV of the most recently coded block can be derived as a BVP candidate. The CBVP buffer can be constructed by appending the BV predictors of each class in the order from category 0 to category 6. The description of CBVP above can be appropriately applied to include fewer categories or additional categories not described above. One or more of categories 0 - 6 can be modified. In one example, each entry in the HBVP buffer is classified into one of seven categories 0 - 6. An index can be signaled to indicate which one of categories 0 - 6 is selected. On the decoder side, the first entry in the selected category can be used to predict the BV of the current block.

[0176] Figure 14 An example of a string copy mode according to an embodiment of the present disclosure is shown. The string copy mode can also be referred to as a string matching mode, an intra - frame string copy mode, or string prediction. The current image (1410) includes a reconstructed region (gray region) (1420) and a region being reconstructed (1421). The current block (1435) in the region (1421) is being reconstructed. The current block (1435) can be a CB, a CU, etc. The current block (1435) can include multiple strings (e.g., strings (1430) and (1431)). In one example, the current block (1435) is divided into multiple consecutive strings, where following one string in the scan order is the next string. The scan order can be any suitable scan order, such as a raster scan order, a traversal scan order, or other predetermined scan order.

[0177] The reconstructed region (1420) can be used as a reference region to reconstruct strings (1430) and (1431).

[0178] For each of the multiple strings, a string offset vector (referred to as SV) and / or the length of the string (referred to as string length) can be signaled or inferred. The SV (e.g., SV0) can be a displacement vector indicating the displacement between the string to be reconstructed (e.g., string (1430)) and the corresponding reference string (e.g., reference string (1400)) located in the reconstructed reference region (1420). The reference string can be used to reconstruct the string to be reconstructed. Thus, the SV can indicate the position of the corresponding reference string in the reference region (1420). The string length can also correspond to the length of the reference string. Reference Figure 14, the current block (1435) is an 8×8 CB that includes 64 samples and is partitioned into two strings (e.g., strings (1430) and (1431)) using a raster scan order. String (1430) includes the first 29 samples of the current block (1435), and string (1431) includes the remaining 35 samples of the current block (1435). The reference string (1400) for reconstructing string (1430) may be indicated by the corresponding string vector SV0, and the reference string (1401) for reconstructing string (1431) may be indicated by the corresponding string vector SV1.

[0179] Generally, the string size (also referred to as the string length) may refer to the length of the string or the number of samples in the string. Refer to Figure 14 , string (1430) includes 29 samples, so the string size or string length of string (1430) is 29. String (1431) includes 35 samples, so the string size or string length of string (1431) is 35. The string position (or the position of the string) may be represented by the sample position of a sample in the string (e.g., the first sample in decoding order).

[0180] The above description may be suitably applied to reconstructing the current block that includes any suitable number of strings. Alternatively, in one example, when the samples in the current block do not have matching samples in the reference region, escape samples (or escape pixels) are signaled, and the values of the escape samples may be directly encoded without referring to the reconstructed samples in the reference region. In one example, the block includes multiple strings and one or more escape samples, where the multiple strings are reconstructed using a string copy mode, the one or more escape samples are directly encoded, and the one or more escape samples are not predicted using the string copy mode. The one or more escape samples may be located at any suitable position within the block. In one example, one or more escape samples in the block are located outside the multiple strings.

[0181] In one example, the strings in the block include one or more escape samples located within the string. Thus, the one or more escape samples in the string may be directly encoded, the one or more escape samples are not predicted using the string copy mode, and the remaining samples in the string are predicted using the string copy mode. The one or more escape samples may be located at any suitable position within the string.

[0182] In some examples, the block (e.g., CB) includes only one string and one or more escape samples, where the string is reconstructed using a string copy mode, the one or more escape samples are directly encoded, and the one or more escape samples are not predicted using the string copy mode.

[0183] Generally, the string shape (i.e., the shape of the string) may be any suitable shape, such as a non-rectangular shape (e.g., Figure 14 the shape of strings (1430)-(1431) in Figure 15the shape of the string (1532)-(l533) in, rectangular shape (e.g., Figure 18 the string (1833) in, etc.).

[0184] In some examples of string matching (or string copy mode), the samples in the current string to be reconstructed (or predicted) and the corresponding reference string have the same order along the scan direction. For example, reference Figure 15 , the string (1532) includes samples (1501)-(1506). The values of the samples (1501)-(1506) are represented by A, B, C, D, E, and F in the sequence. To apply the string copy mode to reconstruct the string (1532), identify a reference string in the reference region (1540) that has a pattern similar to or the same as the patterns A, B, C, D, E, and F arranged in the string (1532).

[0185] In some examples, there is no reference string in the reference region (1540) that has a pattern similar to or the same as the patterns A, B, C, D, E, and F arranged in the string (1532). For example, there may be no reference string having the same order along the scan direction. However, another reference string having a different pattern or different order along the scan direction can be identified, where the pattern or order of the current string can be obtained from the other reference string by performing one or more flipping operations on the other reference string.

[0186] Reference Figure 15 , a reference string (1534) having samples (1521)-(1526) can be identified in the reference region (1540). The samples (1521)-(1526) have values D', C', B', A', F', and E' respectively. In one example, the difference between value A and A' is less than a threshold, the difference between value B and B' is less than a threshold, the difference between value C and C' is less than a threshold, the difference between value D and D' is less than a threshold, the difference between value E and E' is less than a threshold, and the difference between value F and F' is less than a threshold. Allowing a more flexible prediction pattern or order (e.g., along the scan direction of the reference string, it is D', C', B', A', F', and E') in the string matching (or string copy mode) can improve efficiency. The flexible prediction pattern or order can include a pattern or order flipped from the pattern or order of the string (1532).

[0187] According to aspects of the present disclosure, the encoded information of the current block in the current image can be decoded according to the encoded video bitstream. The encoded information can indicate the string copy mode of the current block. The current block includes a current string.

[0188] It can be determined whether to perform a flipping operation to predict the current string. The flipping operation may include flipping an original reference string in a current image to generate a flipped reference string. Based on the determination to use the flipping operation to predict the current string, the original reference string in the current image can be determined based on the SV of the current string. The flipped reference string can be generated by flipping the original reference string, and the current string can be reconstructed based on the flipped reference string.

[0189] A method for applying a string matching pattern (or string copy pattern) by applying a flipping operation to a current string in a current block to be reconstructed or predicted is disclosed in the present disclosure. According to aspects of the present disclosure, the flipped reference string can be used to perform string matching prediction. The original reference string can undergo a flipping operation and then be used as a predictor to predict the current string to be predicted in the string copy pattern. A flipping operation can be performed on the original reference string to generate a flipped reference string, and subsequently the flipped reference string can be used as a predictor to predict the current string in the string copy pattern.

[0190] The flipping operation may refer to a flipping operation associated with a single flipping direction, such as Figure 15 the horizontal flipping operation shown, where the flipping direction is the horizontal flipping direction, or for another example Figure 16 the vertical flipping operation shown, where the flipping direction is the vertical flipping direction, and so on. In some examples, such as in a vertical flipping operation or a horizontal flipping operation, the flipped reference string is a mirror copy of the original reference string.

[0191] The flipping operation may refer to a combined flipping operation associated with multiple flipping directions (such as the horizontal flipping direction and the vertical flipping direction). In one example, as Figure 17 shown, the combined flipping operation includes a combination of multiple flipping operations (such as a vertical flipping operation and a horizontal flipping operation). The combination of multiple flipping operations can be performed in various orders. In one example, the combined flipping operation is performed in a single step.

[0192] In some embodiments, the flipping operation can be performed according to other predetermined conversion patterns and / or on a part of the original reference string. For example, the flipping operation can be performed on a subset of rows or columns of the original reference string.

[0193] As described above, the flipping operation can be of any suitable type (also referred to as the flipping type), such as a vertical flipping operation, a horizontal flipping operation, a combined flipping operation, etc. When more than one flipping type is available for the current string to be predicted, for example, a syntax element or an index can be assigned to the current string to indicate the flipping type used for the current string. The syntax element or index indicating the flipping type can be referred to as the flipping type index. The flipping type index can be signaled at any suitable syntax level. For example, for a string in a block, the flipping type index is signaled at the block level. In one example, the flipping type index indicates the flipping direction and is referred to as the flipping direction index. The flipping type index can be signaled for the current string in the current block to be predicted. In one example, when only one flipping type is available for the current string to be predicted, the flipping type index is not signaled. The flipping type index and / or the flipping type can be inferred as the available flipping types.

[0194] In one example, the flipping types available for the current string include two flipping operations, such as a vertical flipping operation and a horizontal flipping operation. Thus, the flipping type index can have 1 bit. For example, the flipping type index is a flag with 1 bit.

[0195] In one example, the flipping types available for the current string include more than two flipping operations, such as a vertical flipping operation, a horizontal flipping operation, and a combined flipping operation. Thus, the flipping type index can have multiple bits (e.g., two bits). In one example, the same flipping type index can be used to indicate the flipping type. Different values of the flipping type index can indicate different flipping types, such as a horizontal flipping operation, a vertical flipping operation, and a combined flipping operation.

[0196] Variable length coding can be used to indicate the flipping type. For example, a first value (e.g., "0") indicates a horizontal flipping operation, a second value (e.g., "01") indicates a vertical flipping operation, and a third value (e.g., "10") indicates a combined flipping operation, where the number of bits used for the first value, the second value, and the third value can be different. In other embodiments, various operations can be associated with different values.

[0197] In one embodiment, different flags can be provided for multiple flipping operations. For example, two different flags (e.g., two bits) can be used to indicate whether to perform each of a horizontal flipping operation and a vertical flipping operation. For example, a first flag (e.g., a first bit) indicates whether to perform a horizontal flipping operation, and a second flag (e.g., a second bit) indicates whether to perform a vertical flipping operation. In one example, if the values of the first flag and the second flag are both "1", it indicates performing the horizontal flipping operation and the vertical flipping operation respectively, and if the values of the first flag and the second flag are both "0", it indicates not performing the horizontal flipping operation and the vertical flipping operation respectively. Thus, the values of the first flag and the second flag being "00", "01", "10", and "11" indicate (i) no flipping operation, (ii) vertical flipping operation, (iii) horizontal flipping operation, and (iv) combined flipping operation for the current string respectively. The flags can be signaled at any suitable syntactic level, for example, for strings in a block, the flags are signaled at the block level.

[0198] As described above, when two flags (e.g., a first bit and a second bit) indicate performing a horizontal flipping operation and a vertical flipping operation, a combined flipping operation is performed.

[0199] If a first flag (e.g., a first bit) indicates whether to perform a vertical flipping operation, and a second flag (e.g., a second bit) indicates whether to perform a horizontal flipping operation, the above description can be applied appropriately.

[0200] In one embodiment, different flipping type indices can be used to indicate different flipping types, such as a horizontal flipping operation, a vertical flipping operation, and a combined flipping operation. In one example, a first index having, for example, 1 bit is used to indicate whether the flipping operation is a vertical flipping operation or a horizontal flipping operation, and a second index having, for example, 1 bit is used to indicate whether the flipping operation is a combined flipping operation. The different flipping type indices can be signaled at any suitable syntactic level, for example, for strings in a block, the different flipping type indices are signaled at the block level.

[0201] According to aspects of the present disclosure, the flipping operation can be one of (i) a vertical flipping operation, (ii) a horizontal flipping operation, and (iii) a combined flipping operation. Based on the flipping operation being a vertical flipping operation, a flipped reference string can be generated by vertically flipping an original reference string. Based on the flipping operation being a horizontal flipping operation, a flipped reference string can be generated by horizontally flipping an original reference string. Based on the flipping operation being a combined flipping operation, a flipped reference string can be generated by vertically and horizontally flipping an original reference string. For example, each flipping operation can correspond to a different type of flipping operation.

[0202] In one embodiment, the encoded information includes a string-level flag of the current string signaled after the string length of the current string. The string-level flag may indicate whether a flipping operation is performed to predict the current string.

[0203] In accordance with aspects of the present disclosure, a horizontally flipped (or mirrored) reference string may be used to perform string matching prediction in a string copy mode. The original reference string may undergo a horizontal flipping operation and then be used as a predictor to predict the current string in the string copy mode.

[0204] Figure 15 An example of a horizontal flipping operation in accordance with an embodiment of the present disclosure is shown. The current image (1550) includes a reconstructed reconstructed region (gray region) (also referred to as a reference region) (1540) and a region being reconstructed (1541). A current block (1531) in the region (1541) is being reconstructed. The current block (1531) may be a CB, CU, PB, PU, etc. The current block (1531) includes a plurality of strings (e.g., string (1532) and string (1533)). The reconstructed region (1540) may be used as a reference region to reconstruct strings (1532) and (1533).

[0205] To reconstruct the current string (e.g., string (1532)), a reference string (1534) (also referred to as an original reference string) within the reference region (1540) is determined based on an SV (e.g., SV3). The SV (e.g., SV3) may be a displacement vector indicating the displacement between the current string (e.g., string (1532)) and the corresponding reference string (1534) located in the reference region (1540). A flipping operation (e.g., a horizontal flipping operation) is performed on the original reference string (1534) to generate a flipped reference string (also referred to as a horizontally flipped reference string) (1535). Subsequently, the flipped reference string (1535) may be used as a predictor to predict the current string (1532) in the string copy mode.

[0206] Generally, the SV (e.g., SV3) of the current string to be predicted (e.g., string (1532)) may be a displacement vector between the current string (e.g., string (1532)) and the original reference string (e.g., string (1534)). The SV may point from a predetermined position in the current string to any suitable position of the original reference string, such as a predetermined position known to the encoder and decoder. In one embodiment, the SV may point from the sample position of a sample (e.g., the first sample or the starting sample to be reconstructed in scan order) in the current string to the sample position of the corresponding sample in the original reference string used to predict the sample in the current string. Refer to Figure 15, the current string is string (1532), and the sample in string (1532) is sample (1501) with value A. The original reference string is string (1534), and the corresponding sample in the original reference string (1534) for predicting sample (1501) is sample (1524) with value A'. SV3 points from the sample position of sample (1501) in string (1532) to the sample position of sample (1524) in the original reference string (1534).

[0207] In one embodiment, the flipping operation is a horizontal flipping operation, and SV can point from the leftmost sample in the topmost row of the current string to be predicted to the rightmost sample in the topmost row of the original reference string. Refer to Figure 15 , SV3 points from sample (1501) which is the leftmost sample in the topmost row of string (1532) to sample (1524) which is the rightmost sample in the topmost row of the original reference string (1534). In Figure 15 the example shown, due to the flipping operation, sample (1524) is used to predict sample (1501).

[0208] Generally, the flipping operation refers to an operation applied to the original reference string to generate a flipped reference string, and the anti-flipping operation refers to an operation applied to the current string to be predicted to determine the original reference string. The original reference string can be determined by applying the anti-flipping operation to the current string. The shape of the original reference string can be determined by applying the anti-flipping operation to the shape of the current string.

[0209] Refer to Figure 15 , the original reference string (1534) can be determined as follows. The sample at a predetermined position in the original reference string (such as the rightmost sample in the topmost row) can be determined based on SV and the sample at a predetermined position in the current string (such as the leftmost sample in the topmost row). For example, when SV3 points from sample (1501) which is the leftmost sample in the topmost row of string (1532) to sample (1524) which is the rightmost sample in the topmost row of the original reference string (1534), sample (1524) which is the rightmost sample in the topmost row of the original reference string (1534) is determined based on SV3 and sample (1501) of string (1532). Subsequently, the original reference string (1534) (or the shape of the original reference string (1534)) is determined by flipping the current string (1532) (or the shape of string (1532)).

[0210] A flipped reference string (1535) is generated by applying a horizontal flipping operation to an original reference string (1534). The flipped reference string (1535) can be used to predict a current string (1532). The flipped reference string (1535) can have the same shape (or pattern) as the shape (or pattern) of the string (1532). The original reference string (1534) and the string (1532) are mirror copies of each other.

[0211] The benefits of using a flipping operation (e.g., a horizontal flipping operation) in a string copy mode are described below. In Figure 15 the example shown, the string (1532) includes samples (1501)-(1506) having corresponding values A to F, respectively. The values A to F in the string (1532) are arranged in a first pattern that includes the values A to D from left to right in a first row (the topmost row) and the values E and F from left to right in a second row. The first pattern is associated with the string (1532). In one example, a pattern similar or identical to the first pattern is not recognized in a reference region (1540), and thus the string copy mode may not be applied to predict the string (1532), or the string copy mode may not be applied effectively to predict the string (1532) (e.g., the difference between the reference string and the string (1532) is relatively large).

[0212] On the other hand, the original reference string (1534) includes samples (1521)-(1526) having corresponding values D', C', B', A', F', and E', respectively, where the values A to F are respectively similar or identical to the values A' to F'. In one embodiment, the difference between one of the values A to F and a corresponding one of the values A' to F' (the corresponding value used to predict this one of the values A to F) is less than a threshold. For example, the difference between the values A and A' is less than the threshold. The values A' to F' in the original reference string (1534) are arranged in a second pattern that includes the values D', C', B', and A' from left to right in a first row and the values F' and E' from left to right in a second row. The second pattern is associated with the original reference string (1534).

[0213] The first pattern and the second pattern are different. For example, the sample (1524) used to predict the sample (1501) is the rightmost sample in the string (1534), while the sample (1501) is the leftmost sample in the string (1532). However, in the reference region (1540), the second pattern is similar or identical to a flipped copy (e.g., a mirror copy) of the first pattern. Thus, the original reference string (1534) can be flipped (i.e., horizontally flipped) to generate a flipped reference string (1535) associated with a third pattern that is similar or identical to the first pattern. The third pattern includes the values A', B', C', and D' from left to right in a first row and the values E' and F' from left to right in a second row in the flipped reference string (1535). The second pattern and the third pattern are flipped copies of each other. As Figure 15As shown, the samples (1521-1526) in the original reference string (1534) and the flipped copies (1521-1526) in the flipped reference string (1535) are symmetric about the vertical axis. That is, for each sample, the column where it is located changes after horizontal flipping, while the row remains unchanged. For example, the sample (1521) in the first column of the original reference string (1534) becomes located in the fourth column in the flipped reference string (1535) after horizontal flipping; the sample (1522) in the second column of the original reference string (1534) becomes located in the third column in the flipped reference string (1535) after horizontal flipping; the samples (1523, 1525) in the third column of the original reference string (1534) become located in the second column in the flipped reference string (1535) after horizontal flipping; the samples (1524, 1526) in the fourth column of the original reference string (1534) become located in the first column in the flipped reference string (1535) after horizontal flipping.

[0214] Therefore, in addition to using a pattern similar to or the same as the pattern of the current string (e.g., the first pattern) to predict the current string in the string copy mode, a pattern (e.g., the second pattern) similar to or the same as the flipped copy of the pattern of the current string (e.g., string (1532)) can be used in the string copy mode to predict the current string, thereby enabling a more flexible and efficient string copy mode.

[0215] In some examples, the current string to be predicted includes only one column. Therefore, the horizontally flipped reference string generated by performing a horizontal flipping operation on the original reference string having one column is the same as the original reference string. Thus, the horizontal flipping operation does not change the original reference string. In some examples, if the current string includes only one column, the horizontal flipping operation is prohibited or disabled. A string including only one column can occur in a vertical scan order. In one example, a string including only one column occurs only in a vertical scan order.

[0216] According to aspects of the present disclosure, a vertically flipped (or mirrored) reference string can be used to perform string matching prediction in the string copy mode. The original reference string can undergo a vertical flipping operation and then be used as a predictor to predict the current string in the string copy mode.

[0217] Figure 16 An example of vertical flipping according to an embodiment of the present disclosure is shown. The current image (1550) includes a reconstructed region (1540) and a region being reconstructed (1541). The current block (1531) is being reconstructed. The current block (1531) includes a plurality of strings (e.g., string (1532) and string (1533)). The reconstructed region (1540) can be used as a reference region to reconstruct strings (1532) and (1533).

[0218] To reconstruct the current string (e.g., string (1532)), a reference string (1536) (also referred to as the original reference string) within the reconstructed region (1540) is determined based on the SV (e.g., SV4). The SV4 can be a displacement vector indicating the displacement between the current string (e.g., string (1532)) and the corresponding reference string (1536) located in the reference region (1540). A flipping operation (e.g., a vertical flipping operation) is performed on the original reference string (1536) to generate a flipped reference string (also referred to as a vertically flipped reference string) (1537). Subsequently, the flipped reference string (1537) can be used as a predictor to predict the current string (1532) in the string copy mode.

[0219] Generally, the SV can point from a predetermined position in the current string to be predicted to any suitable position in the original reference string. Refer Figure 16 , the current string is string (1532). The original reference string is string (1536). The corresponding sample (1553) with the value "A" in the original reference string (1536) is used to predict the sample (1501) with the value "A" in the string (1532). The SV4 points from the sample position of the sample (1501) in the string (1532) to the sample position of the sample (1553) in the original reference string (1536).

[0220] In one embodiment, the flipping operation is a vertical flipping operation, and the SV can point from the leftmost sample in the topmost row of the current string to the leftmost sample in the bottommost row of the original reference string. Refer Figure 16 , the SV4 points from the sample (1501) which is the leftmost sample in the topmost row of the string (1532) to the sample (1553) which is the leftmost sample in the bottommost row of the original reference string (1536). In Figure 16 the example shown, due to the flipping operation, the sample (1553) is used to predict the sample (1501).

[0221] Refer Figure 16 , the original reference string (1536) can be determined as follows. When the SV4 points from the sample (1501) which is the leftmost sample in the topmost row of the string (1532) to the sample (1553) which is the leftmost sample in the bottommost row of the original reference string (1536), the sample (1553) which is the leftmost sample in the bottommost row of the original reference string (1536) is determined based on the SV4 and the sample (1501) of the string (1532). Subsequently, the original reference string (1536) (or the shape of the original reference string (1536)) is determined by vertically flipping the string (1532) (or the shape of the string (1532)).

[0222] A flipped reference string (1537) is generated by applying a vertical flip operation to an original reference string (1536). The flipped reference string (1537) can be used to predict a current string (1532). The flipped reference string (1537) can have the same shape (or pattern) as the shape (or pattern) of the string (1532). The original reference string (1536) and the string (1532) are mirror copies of each other.

[0223] The benefits of using a flip operation (e.g., a vertical flip operation) in the string copy mode are described below. In Figure 16 the example shown, the values A to F in the string (1532) are arranged in a first pattern as described in reference Figure 15 As described above, a pattern similar to or the same as the first pattern is not recognized in the reference region (1540), so the string copy mode may not be applied to predict the string (1532), or the string copy mode may not be applied to effectively predict the string (1532).

[0224] On the other hand, the original reference string (1536) includes samples (1551)-(1556) having corresponding values E”, F”, A”, B”, C” and D” respectively, where the values A to F are respectively similar to or the same as the values A” to F”. In one embodiment, the difference between one of the values A to F and one of the corresponding values A” to F” (the corresponding value used to predict this one of the values A to F) is less than a threshold. For example, the difference between the values A and A” is less than the threshold. The values A” to F” in the original reference string (1536) are arranged in a fourth pattern including the values E” and F” from left to right in the first row and the values A”, B”, C” and D” from left to right in the second row. The fourth pattern is associated with the original reference string (1536).

[0225] The first pattern and the fourth pattern are different. For example, the sample (1553) used to predict the sample (1501) is located in the bottommost row of the string (1536), while the sample (1501) is located in the topmost row of the string (1532). However, in the reference region (1540), the fourth pattern is similar to or the same as a flipped copy (e.g., a mirror copy) of the first pattern. Therefore, the original reference string (1536) can be flipped (i.e., vertically flipped) to generate a flipped reference string (1537) associated with a fifth pattern that is similar to or the same as the first pattern. The fifth pattern includes the values A”, B”, C” and D” from left to right in the first row and the values E” and F” from left to right in the second row in the flipped reference string (1537). The fourth pattern and the fifth pattern are flipped copies of each other. As Figure 16As shown, the samples (1551-1556) in the original reference string (1536) and the flipped copies (1551-1556) in the flipped reference string (1537) are symmetric about the horizontal axis. That is, for each sample, the row where it is located changes after vertical flipping, while the column remains unchanged. For example, the samples (1551, 1552) in the first row of the original reference string (1536) become located in the second row in the flipped reference string (1537) after vertical flipping; and the samples (1553, 1554, 1555, 1556) in the second row of the original reference string (1536) become located in the first row in the flipped reference string (1537) after vertical flipping. At the same time, after vertical flipping,

[0226] A pattern (e.g., the fourth pattern) that is similar or identical to the flipped copy of the pattern (e.g., the first pattern) of the current string (e.g., string (1532)) can be used in the string copy mode to predict the current string, thereby enabling a more flexible and efficient string copy mode.

[0227] In some examples, the current string to be predicted includes only one row. Therefore, the vertically flipped reference string generated by performing a vertical flipping operation on the original reference string with one row is the same as the original reference string. Thus, the vertical flipping operation does not change the original reference string. In some examples, if the current string includes only one row, the vertical flipping operation is prohibited or disabled. A string that includes only one row can appear in a horizontal scan order. In one example, a string that includes only one row appears only in a horizontal scan order.

[0228] According to aspects of the present disclosure, a combined flipping operation can be applied to a reference string to generate a flipped reference string, where the flipped reference string is used as a predictor to perform string matching prediction in the string copy mode. The combined flipping operation can include any suitable flipping operations performed sequentially in any suitable order. The combined flipping operation can also be performed in a single step of directly mapping the reference string to the flipped reference string.

[0229] In one example, the combined flipping operation includes, for example, the horizontal flipping operation described in reference Figure 15 and the vertical flipping operation described in reference Figure 16 The horizontal flipping operation and the vertical flipping operation can be combined to generate a flipped reference string, and the order of performing the combined operation can vary according to the embodiment.

[0230] Figure 17Shows an example of a combined flipping operation according to an embodiment of the present disclosure. The current image (1550) includes a reconstructed region (1540) and a region being reconstructed (1541). The current block (1531) in the region (1541) is being reconstructed. The current block (1531) includes a plurality of strings (e.g., string (1532) and string (1533)). The reconstructed region (1540) can be used as a reference region to reconstruct strings (1532) and (1533).

[0231] To reconstruct the current string (e.g., string (1532)), a reference string (1538) (also referred to as the original reference string) within the reconstructed region (1540) is determined based on the SV (e.g., SV5). SV5 can be a displacement vector indicating the displacement between the current string (e.g., string (1532)) and the corresponding reference string (1538) located in the reference region (1540). A combined flipping operation is performed on the original reference string (1538) to generate a flipped reference string (1539). Subsequently, the flipped reference string (1539) can be used as a predictor to predict the current string (1532) in the string copy mode.

[0232] As described above, the combined flipping operation can be performed on the original reference string (e.g., reference string (1538)) according to any suitable method and / or order to generate the flipped reference string (e.g., flipped reference string (1539)). The combined flipping operation can be performed on the original reference string (e.g., reference string (1538)) using one or more steps to generate the flipped reference string (e.g., flipped reference string (1539)).

[0233] In one embodiment, the combined flipping operation can be performed in two steps. In one example, the original reference string (1538) is flipped vertically to generate a first intermediate string, and then the first intermediate string is flipped horizontally to generate the flipped reference string (1539). In one example, the original reference string (1538) is flipped horizontally to generate a second intermediate string, and then the second intermediate string is flipped vertically to generate the flipped reference string (1539).

[0234] In one embodiment, the combined flipping operation can be performed in a single step. For example, the original reference string (1538) is directly flipped to generate the flipped reference string (1539).

[0235] Generally, the SV can point from a predetermined position in the current string to be predicted to any suitable position of the original reference string. Refer to Figure 17, the current string is string (1532). The original reference string is string (1538). The corresponding sample (1706) with value A''' in the original reference string (1538) is used to predict the sample (1501) with value A in string (1532). SV5 points from the sample position of sample (1501) in string (1532) to the sample position of sample (1706) in the original reference string (1538).

[0236] In one embodiment, the flipping operation is a combined flipping operation including a horizontal flipping operation and a vertical flipping operation, and the SV can point from the leftmost sample in the topmost row of the current string to the rightmost sample in the bottommost row of the original reference string. Refer to Figure 17 , SV5 points from sample (1501), which is the leftmost sample in the topmost row of string (1532), to sample (1706), which is the rightmost sample in the bottommost row of the original reference string (1538). In Figure 17 the example shown, due to the combined flipping operation, sample (1706) is used to predict sample (1501).

[0237] Refer to Figure 17 , the original reference string (1538) can be determined as follows. When SV5 points from sample (1501), which is the leftmost sample in the topmost row of string (1532), to sample (1706), which is the rightmost sample in the bottommost row of the original reference string (1538), sample (1706), which is the rightmost sample in the bottommost row of the original reference string (1538), is determined based on SV5 and sample (1501) of string (1532). Subsequently, the original reference string (1538) (or the shape of the original reference string (1538)) is determined by the reverse flipping operation of the combined flipping operation, such as vertically and horizontally flipping string (1532) (or the shape of string (1532)).

[0238] The benefits of using the combined flipping operation in the string copy mode are described below. In Figure 17 the example shown, the values A to F in string (1532) are arranged in the first pattern as described in reference Figure 15 . As described above, a pattern similar to or the same as the first pattern is not recognized in the reference area (1540), so the string copy mode may not be applied to predict string (1532), or the string copy mode may not be applied to effectively predict string (1532).

[0239] On the other hand, the original reference string (1538) includes samples (1701)-(1706) having corresponding values F''', E''', D''', C''', B''', and A''', respectively, where the values A to F are similar or identical to the values A''' to F'''. In one embodiment, the difference between one of the values A to F and a corresponding one of the values A''' to F''' (the corresponding value being used to predict this one of the values A to F) is less than a threshold. For example, the difference between value A and A''' is less than the threshold. The values A''' to F''' in the original reference string (1538) are arranged in a sixth pattern that includes values F''' and E''' from left to right in the first row and values D''', C''', B''', and A''' from left to right in the second row. The sixth pattern is associated with the original reference string (1538).

[0240] The first pattern and the sixth pattern are different. For example, the sample (1706) used to predict the sample (1501) is the rightmost sample in the bottom row of the string (1538), while the sample (1501) is the leftmost sample in the top row of the string (1532). However, in the reference region (1540), the sixth pattern is similar or identical to a flipped copy of the first pattern along the horizontal and vertical directions. Thus, the original reference string (1538) can be flipped to generate a flipped reference string (1539) associated with a seventh pattern that is similar or identical to the first pattern. The seventh pattern includes values A''', B''', C''', and D''' from left to right in the first row and values E''' and F''' from left to right in the second row in the flipped reference string (1539). The sixth pattern is a flipped copy of the seventh pattern along the horizontal and vertical directions.

[0241] Therefore, a pattern (such as the sixth pattern) that is similar or identical to a flipped copy of the pattern (such as the first pattern) of the current string (e.g., string (1532)) can be used in the string copy pattern to predict the current string, thereby enabling a more flexible and efficient string copy pattern.

[0242] In some examples, the current string to be predicted includes only one (e.g., one row or one column) sample. Thus, only one of the vertical flip operation and the horizontal flip operation can change the original reference string corresponding to the current string. In some examples, if the current string includes only one (e.g., one row or one column), the combined flip operation is prohibited or disabled.

[0243] In one example, the current string includes only one row, and the combined flip operation is prohibited for the current string. The horizontal flip operation is allowed for the current string. Additionally, the vertical flip operation is prohibited for the current string.

[0244] In one example, the current string includes only one column, and the combined flip operation is prohibited for the current string. The vertical flip operation is allowed for the current string. In addition, the horizontal flip operation is prohibited for the current string.

[0245] In accordance with aspects of the present disclosure, a syntax element or index (e.g., a flip type index) may be assigned to indicate a combined flip operation. In one example, the same flip type index may be used to indicate the flip type. Different values of the flip type index may indicate different flip types, such as a horizontal flip operation, a vertical flip operation, and a combined flip operation.

[0246] Variable length coding may be used to indicate the flip type. For example, a first value (e.g., "0") indicates a horizontal flip operation, a second value (e.g., "01") indicates a vertical flip operation, and a third value (e.g., "10") indicates a combined flip operation, where the number of bits used for the first value, the second value, and the third value may be different.

[0247] In one embodiment, different bits may be used to indicate whether each of the horizontal flip operation and the vertical flip operation is performed. For example, 1 bit (e.g., the first bit) indicates whether the horizontal flip operation is performed, and another bit (e.g., the second bit) indicates whether the vertical flip operation is performed. When both bits (e.g., the first bit and the second bit) indicate that the horizontal flip operation and the vertical flip operation are performed, the combined flip operation is performed.

[0248] In one embodiment, different flip type indexes may be used to indicate different flip types, such as a horizontal flip operation, a vertical flip operation, and a combined flip operation. In one example, a first index is used to indicate whether the flip operation is a vertical flip operation or a horizontal flip operation, and a second index is used to indicate whether the flip operation is a combined flip operation.

[0249] In accordance with aspects of the present disclosure, one or more methods may be used to indicate (e.g., signal) a flip operation for a current block to be predicted and / or a current string to be predicted in a string copy mode. Whenever applicable, the one or more methods may be used alone or in combination. The one or more methods may indicate (e.g., signal) whether the flip operation is used for the current block and / or the current string. The one or more methods may indicate (e.g., signal) the flip type for the current string in the string copy mode.

[0250] In one embodiment, the entire current block (e.g., the entire CB) may use a flipped reference string in the string copy mode. For example, the current block (e.g., block (1531)) includes at least one string (e.g., strings (1532)-(l533)). Each string in the at least one string is predicted using a corresponding flipped reference string. For example, string (1532) is predicted using flipped reference string (1535). Another flipped reference string is used to predict string (1533), similar to that described in reference Figures 15 - 17 as

[0251] The block-level flag of the current block can be used to indicate whether at least one flipped reference string (e.g., two different flipped reference strings) is used to predict at least one string (e.g., strings (1532)-(l533)) in the current block using the string copy mode. In some examples, the block-level flag of the current block is used to indicate the use of at least one flipped reference string to predict at least one string (e.g., strings (1532)-(1533)) in the current block (e.g., block (1531)) using the string copy mode.

[0252] In one example, the block-level flag of the current block indicates whether a corresponding flipped reference string is used to predict each string in at least one string in the current block. For example, the block-level flag of the current block indicates that a corresponding flipped reference string is used to predict each string in at least one string in the current block.

[0253] According to aspects of the present disclosure, the current string is one of the at least one string included in the current block. The encoded information may include the block-level flag of the current block. The block-level flag may indicate that at least one flipping operation is performed on at least one string to predict the current block. The block-level flag may indicate that a flipping operation is performed on each string in at least one string to predict the current block. Therefore, it can be determined to use at least one flipping operation to predict the current block based on the block-level flag indicating that a flipping operation is performed on each string in at least one string to predict the current block, where the at least one flipping operation includes a flipping operation performed on the original reference string.

[0254] According to aspects of the present disclosure, the current string is one of the multiple strings included in the current block. The encoded information may include the block-level flag of the current block. A first value of the block-level flag may indicate that a corresponding flipping operation is used to predict each string in the multiple strings, and a second value of the block-level flag may indicate that no flipping operation is performed on the multiple strings. In one example, the block-level flag has the first value, so a corresponding flipping operation is used to predict each string in the multiple strings. Therefore, based on the block-level flag having the first value, it is determined to perform a flipping operation to predict the current string.

[0255] According to aspects of the present disclosure, the current string is one of at least one string included in the current block. The number of at least one string in the current block is greater than a first threshold (also referred to as a quantity threshold), and thus it can be determined, based on the number of at least one string in the current block being greater than the first threshold, to prohibit at least one flipping operation for at least one string for the current block. The at least one flipping operation includes a flipping operation performed on an original reference string, and it is determined to prohibit this flipping operation. In one example, the at least one string corresponds to a plurality of strings, and the current string is one of the plurality of strings included in the current block. It can be determined, based on the number of the plurality of strings in the current block being greater than the first threshold, to prohibit the flipping operation for the current string.

[0256] After signaling the string lengths of the corresponding strings in the current block, a flag (e.g., a block-level flag) for indicating a flipping operation for the current block can be signaled. In some embodiments, the use of the flipping operation can depend on the number of strings in the current block. For example, when the number of at least one string (e.g., a plurality of strings) in the current block is greater than a first threshold (e.g., a quantity threshold), the flipping operation can be prohibited for the current block. In one example, the flag (e.g., a block-level flag) is not signaled. The flag (e.g., a block-level flag) can be inferred to indicate that no flipping operation is performed on the current block. The first threshold can be any suitable threshold, and any suitable method can be used to obtain it. In one example, the first threshold is a predetermined quantity threshold known to the encoder and / or decoder. In one example, for instance, the first threshold is calculated by the encoder and / or decoder based on the block size of the current block (e.g., block width, block height, block area). The first threshold can be signaled to the decoder.

[0257] In some examples, the current block includes a plurality of strings, and the flag of the current block can indicate whether a flipping operation is allowed for one or more of the plurality of strings. A first value of the flag can indicate that a flipping operation is allowed for one or more of the plurality of strings. A second value of the flag can indicate that a flipping operation is not allowed for the plurality of strings.

[0258] As described above, when more than one flipping type (e.g., more than one flipping direction or more than one flipping operation) is available for the current string to be predicted, a flipping type index can be used for the current string in the current block. The flipping type index can be signaled. In one example, the strings in the current block can have different flipping types, and for each string in the current block, the corresponding flipping type index is signaled.

[0259] Reference Figure 17, the current block (1531) includes strings (1532)-(1533). A block-level flag is used to indicate whether a flipping operation is used for the current block (1531). In one example, the block-level flag indicates that the flipping operation is used for the current block, so the flipping operation is used to predict strings (1532)-(1533). Two flipping types can be respectively used for strings (1532)-(1533). The first flipping type index of string (1532) is used to indicate the flipping type of string (1532). The second flipping type index of string (1533) is used to indicate the flipping type of string (1533).

[0260] Whether a flipped reference string is used for the current string in the current block predicted using the string copy mode can be indicated by a string-level flag (or string-level indication flag) of the current string in the current block. The corresponding string-level flag of each string in the current block can be used to indicate whether a flipping operation is used for the string.

[0261] According to aspects of the present disclosure, the encoded information may include a string-level flag of the current string, where the string-level flag may indicate that a flipping operation is used to predict the current string. Therefore, it can be determined to use the flipping operation to predict the current string based on the execution of the flipping operation indicated by the string-level flag to predict the current string.

[0262] In one embodiment, the use of the flipping operation may depend on the string length of the current string, for example when the number of samples in the current string is less than a second threshold (also referred to as a length threshold). It can be determined to prohibit the flipping operation from being used for the current string based on the string length of the current string being less than the second threshold.

[0263] When the string length of the current string is less than the length threshold (e.g., a predetermined length threshold), the flipping operation can be prohibited from being used for the current string to be predicted, and it is not necessary to signal relevant flags (e.g., string-level flags, flipping type indexes, etc.) indicating the flipping operation of the current string if the relevant flags exist.

[0264] As described above, when more than one flipping type (e.g., more than one flipping direction or more than one flipping operation) is available for a string in the current block using the flipping operation (or available for a string in the current block using the flipping operation), a flipping type index can be used for a string in the current block. The flipping type index can be signaled. In one example, the strings in the current block can have different flipping types, and for each string in the current block, the corresponding flipping type index is used (e.g., signaled).

[0265] Reference Figure 17, the current block (1531) includes strings (1532)-(1533). Two cascade flags are respectively used to indicate whether the flipping operation is used for the strings (1532)-(1533). In one example, the first cascade flag of the string (1532) indicates that the flipping operation is used for the string (1532), and the second cascade flag of the string (1533) indicates that the flipping operation is not used for the string (1533). More than one flipping type can be used for the string (1532), and the flipping type index of the string (1532) is used to indicate the flipping type of the string (1532).

[0266] Generally, if the flipped reference string generated by flipping the original reference string is different from the original reference string, the flipping operation can be effective. If the flipped reference string is the same as the original reference string, the flipping operation is invalid and may not be an option. Therefore, the flipping operation is not used. In some cases, it can be inferred that the flipping operation is not used, and it may not be necessary to signal the block-level flag or the cascade flag.

[0267] If the starting sample (or the first sample to be scanned) and the ending sample (or the last sample to be scanned) of the current string to be predicted are in two different rows of the current block to be predicted in the horizontal scan order, the flipping operation (e.g., horizontal flipping operation, vertical flipping operation or combined flipping operation) can be effective. If the starting sample (or the first sample to be scanned) and the ending sample (or the last sample to be scanned) of the current string to be predicted are in two different columns of the current block to be predicted in the vertical scan order, the flipping operation (e.g., horizontal flipping operation, vertical flipping operation or combined flipping operation) can be effective.

[0268] In some examples, when the string length of the current string to be predicted is greater than the length of one row (e.g., block width) or the length of one column (e.g., block height) of the current block, the flipping operation is effective.

[0269] According to aspects of the present disclosure, if a specific flipping operation (e.g., horizontal flipping operation or vertical flipping operation) results in a flipped reference string that is the same as the original reference string, the specific flipping operation may not be an option, and it can be inferred that the specific flipping operation is not used. For example, the specific flipping operation is not used. In one example, the specific flipping operation is not signaled. In one example, the current string to be predicted includes only one row, so the original reference string also includes only one row. The vertical flipping operation is not an option, and the vertical flipping operation is not used to vertically flip the original reference string, which would result in a flipped reference string that is the same as the original reference string. In one example, the current string to be predicted includes only one column, so the original reference string also includes only one column. The horizontal flipping operation is not an option, and the horizontal flipping operation is not used to horizontally flip the original reference string, which would result in a flipped reference string that is the same as the original reference string.

[0270] In some examples, the flip operation is valid and thus enabled. After signaling the string length of the current string, the flag (e.g., cascade flag) of the current string to be predicted can be signaled, where the flag indicates whether the flip operation is applied to the current string.

[0271] In one example, in the horizontal scan order, if the start sample (or the first sample to be scanned) and the end sample (or the last sample to be scanned) of the current string to be predicted are in the same row of the current block, then the current string has only one row of samples and the vertical flip operation cannot be an option. It is inferred that the vertical flip operation is not used. For a current string having only one row of samples, if only the vertical flip operation is considered by the decoder, for example, then there is no need to signal the flag (e.g., cascade flag) indicating whether the flip operation is applied to the current string. In one example, the flag (e.g., cascade flag) is not signaled, and it is inferred that the value of the flag indicates that the flip operation is not used for the current string.

[0272] On the other hand, if the decoder can consider both the vertical flip operation and another flip operation (e.g., horizontal flip operation), for example, then after selecting to perform the flip operation on a current string having only one row of samples, only the other flip operation (e.g., horizontal flip operation) is available. Therefore, there is no need to signal the flip type (e.g., flip direction), and it can be inferred that the flip type is the other flip type (e.g., horizontal flip operation). In one example, the flip type index (e.g., flip direction index) is not signaled, and it is inferred that the flip type index is a value indicating the other flip type.

[0273] Figure 18 An example of the flip operation for a rectangular string according to an embodiment of the present disclosure is shown. The current image (1850) includes a reconstructed region (also referred to as a reference region) (gray region) (1840) and a region being reconstructed (1841). The current block (1831) in the region (1841) is being reconstructed. The current block (1831) can be a CB, CU, PB, PU, etc. The current block (1831) includes a plurality of strings (e.g., string (1832) and string (1833)) and escape samples (1807)-(1808). String (1832) includes samples (1801)-(1806) and is a non-rectangular string. String (1833) includes samples (1809)-(1816) and is a rectangular string, where the shape of the string is a rectangular shape. The reconstructed region (1840) can be used as a reference region to reconstruct strings (1832) and (1833). The current string (e.g., string (1833)) is a rectangular string.

[0274] According to one aspect of the present disclosure, a flipping operation may be prohibited for a string of rectangles to be predicted (e.g., string (1833)). Based on the current string having a rectangular shape, the flipping operation may be prohibited for the current string. In one example, a string copy mode is used to predict the string of rectangles (1833) without a flipping operation.

[0275] According to one aspect of the present disclosure, when a first sample in a string of rectangles is derived by flipping a corresponding original reference string and at least one sample in the string of rectangles is not derived by flipping the corresponding original reference string, the flipping operation may be permitted for the string of rectangles to be predicted. At least one sample in the string of rectangles that is not derived by flipping the corresponding original reference string may be referred to as at least one non-flipped sample. For example, at least one non-flipped sample in the string of rectangles includes escape samples (or escape pixels), and the escape samples (or escape pixels) may be predicted without using the original reference string. As described above, the values of the escape samples may be directly encoded without referring to the corresponding reconstructed samples in the reference region. When performing a string copy mode with a flipping operation, the at least one encoded non-flipped sample may be arranged or rearranged.

[0276] In one embodiment, the current string has a rectangular shape and includes samples that are predicted without using a flipping operation. It may be determined to perform a flipping operation to predict a plurality of samples (e.g., first samples) in the current string that are different from the samples predicted without using the flipping operation. In one example, the current string is the current block.

[0277] In one example, a string copy mode with a flipping operation is used to predict a rectangular string (1833). The rectangular string (1833) includes a first set of samples (e.g., samples (1809)-(1812) and (1816)) predicted using a flipping operation (e.g., a horizontal flipping operation) and a second set of samples (e.g., samples (1813)-(1815)) predicted without using a flipping operation. More specifically, an original reference string (1834) within a reconstructed region (1840) is determined based on an SV (e.g., SV6). The original reference string (1834) includes samples (1821)-(1825). A flipping operation (e.g., a horizontal flipping operation) is performed on the original reference string (1834) to generate a flipped reference string (1835). Subsequently, the flipped reference string (1835) can be used as a predictor to predict the first set of samples in the current string (1833) in the string copy mode. For example, samples (1824), (1823), (1822), (1821), and (1825) in the flipped reference string (1835) are used to predict the first set of samples (1809)-(1812) and (1816) respectively. The second set of samples (1813)-(1815) is predicted without using a flipping operation and can be referred to as non-flipped samples. The second set of samples (1813)-(1815) can be escape samples that can be directly encoded without using the reconstructed samples in the reconstructed region (1840).

[0278] In one example, after predicting the first set of samples based on the flipped reference string (1835) and directly encoding the second set of samples (1813)-(1815), the encoded second set of samples (1813)-(1815) is arranged with the predicted first set of samples to form a predicted rectangular string (1833).

[0279] In one embodiment, the current block is a current string having a rectangular shape. It can be determined that a flipping operation is used for a subset of samples in the current string. At least one sample in the current string is predicted without using a flipping operation, and the subset of samples is different from at least one sample in the current string.

[0280] In one embodiment, the current block to be predicted includes only one string. In one example, the current block is a string, and the string has a rectangular shape. A flipping operation can be allowed for a subset of samples in the rectangular string, where the subset of samples is predicted using a string copy mode. At least one sample in the rectangular string is predicted without using a flipping operation and is referred to as a non-flipped sample in the rectangular string. At least one sample is different from the subset of samples. When performing a string copy mode with a flipping operation, the non-flipped samples can be arranged or rearranged. The non-flipped samples can include escape samples that can be directly encoded without referring to the corresponding reconstructed samples in a reference region.

[0281] Figure 19Shows an example of a flipping operation allowed for the current block (1931) according to an embodiment of the present disclosure. The current image (1950) includes a reconstructed region (also referred to as a reference region) (gray region) (1940) and a region being reconstructed (1941). The current block (1931) in the region (1941) is being reconstructed. The current block (1931) can be a CB, CU, PB, PU, etc. The current block (1931) may include samples (1901)-(1908).

[0282] In one example, the current block (1931) is a rectangular string (1931), where the rectangular string (1931) includes samples (1901)-(1908). The rectangular string (1931) can be reconstructed or predicted similar to the string (1833) described in the reference Figure 18 as described. The reference Figure 19 , allows a horizontal flipping operation for a subset of samples (such as samples (1902)-(1908)) in the current block (or current string) (1931), where the string copy mode is used to predict the subset of samples. The sample (1901) is not predicted using the flipping operation and is referred to as the non-flipping sample in the current block (1931). In one example, the sample (1901) is an escape sample. An original reference string (1934) including samples (1921)-(1927) is obtained based on the SV (e.g., SV7). A flipped reference string (1935) is generated by horizontally flipping the original reference string (1934). Subsequently, the samples (1902)-(1908) in the current string (1931) are predicted based on the flipped reference string (1935). The non-flipping sample (1901) in the current string (1931) can be directly encoded and then arranged at the upper left corner of the reconstructed string (1931).

[0283] In one example, the current block includes a current string to be predicted and at least one sample located outside the current string. The at least one sample located outside the current string can be located at any suitable position in the current block. A flipping operation can be allowed for the current string, and the flipping operation is performed on the current string, such as the reference Figures 15 - 17 as described for the string (1532) in. The at least one sample located outside the current string is not predicted using the flipping operation. When the string copy mode is performed on the current string, the at least one sample located outside the current string can be arranged. In one example, the at least one sample located outside the string includes an escape sample that can be directly encoded without referring to the corresponding reconstructed sample in the reference region.

[0284] In one example, the reference Figure 19, the current block (1931) includes a current string (1932) and samples (1901) located outside the current string (1932). A flip operation (e.g., a horizontal flip operation) is allowed for the current string (1932) that includes samples (1902)-(1908). As described above, the sample (1901) is not predicted using the flip operation and is referred to as a non-flipped sample in the current block (1931). An original reference string (1934) is obtained based on SV7. A flipped reference string (1935) is generated by horizontally flipping the original reference string (1934). Subsequently, the current string (1932) is predicted based on the flipped reference string (1935). The non-flipped sample (1901) can be directly encoded. Then it is arranged at the upper left corner of the reconstruction block (1931).

[0285] In some examples, the transformation process and the quantization process are skipped. Thus, after predicting the current block using the string copy mode, no transformation and quantization are performed on the residual of the current block. In some examples, after predicting the current block using the string copy mode, the transformation and quantization are performed on the residual of the current block.

[0286] In one embodiment, the residual of the current block predicted using the flip operation is transformed and quantized. A flag (e.g., a 1-bit flag) can be used to indicate whether to skip the transformation process. The flag can be signaled at the block level.

[0287] In one embodiment, the transformed coefficients in the current block can be dequantized. The residual of the current block can be generated by performing an inverse transformation on the transformed coefficients.

[0288] Figure 20 A flowchart outlining a process (2000) according to an embodiment of the present disclosure is shown. The process (2000) can be used to reconstruct blocks such as CB, PB, PU, CU, etc. In various embodiments, the process (2000) is performed by a processing circuit such as the processing circuits in the following terminal devices: terminal device (230), terminal device (320), terminal device (330), and terminal device (340), the processing circuit that performs the function of the video encoder (403), the processing circuit that performs the function of the video decoder (410), the processing circuit that performs the function of the video decoder (510), the processing circuit that performs the function of the video encoder (603), etc. In some embodiments, the process (2000) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (2000). The process starts at (S2001) and proceeds to (S2010).

[0289] At (S2010), the encoded information of the current block in the current image can be decoded according to the encoded video bitstream. The encoded information can indicate the string copy mode of the current block, and the current block includes the current string. The current block can include at least one string having the current string. In one example, the at least one string includes a plurality of strings. In one example, the at least one string includes only a single string as the current string. In one example, the current block includes escape samples. The escape samples can be located within and / or outside one or more of the at least one string. The shape of the current string can be a non-rectangular shape or a rectangular shape.

[0290] At (S2020), it can be determined whether to perform a flipping operation to predict the current string. The flipping operation can include flipping the original reference string in the current image to generate a flipped reference string.

[0291] In one embodiment, the encoded information includes a block-level flag of the current block. The current string can be one of the plurality of strings included in the current block. The first value of the block-level flag can indicate that for each of the plurality of strings, a corresponding flipping operation is used to predict the string, and the second value of the block-level flag can indicate that no flipping operation is performed on the plurality of strings. Whether to perform a flipping operation to predict the current string can be determined based on the block-level flag. It can be determined to perform a flipping operation to predict the current string based on the block-level flag having the first value.

[0292] It can be determined to prohibit the flipping operation for the plurality of strings in the current block based on the number of the plurality of strings in the current block being greater than a first threshold. Therefore, if the number of the plurality of strings in the current block is greater than the first threshold, it is determined to prohibit the flipping operation for the current string. In one example, when the number of the plurality of strings in the current block is greater than the first threshold, the block-level flag is not signaled, and it is inferred that the block-level flag indicates prohibiting the flipping operation for the current string.

[0293] In one embodiment, the encoded information includes a string-level flag of the current string. The string-level flag can indicate whether to perform a flipping operation to predict the current string. Whether to perform a flipping operation to predict the current string can be determined based on the string-level flag. If the string-level flag indicates performing a flipping operation to predict the current string, it is determined to perform a flipping operation to predict the current string. The encoded information can include the string-level flag of the current string signaled in the encoded information after the string length of the current string.

[0294] In one example, the string length of the current string is less than a second threshold, where the string length of the current string indicates the number of samples in the current string. Therefore, it is determined to prohibit the flipping operation for the current string. In one example, when the string length of the current string is less than the second threshold, the string-level flag is not signaled, and it is inferred that the string-level flag indicates prohibiting the flipping operation for the current string.

[0295] In one example, the current string has a rectangular shape and it is determined that a flip operation is prohibited for the current string.

[0296] In one example, the current string has a rectangular shape and includes samples (e.g., escape samples) that are not predicted using a flip operation. It may be determined to perform a flip operation to predict a plurality of samples in the current string that are different from the samples not predicted using a flip operation. In one example, the current string is the current block.

[0297] In one example, the current block is the current string having a rectangular shape. A flip operation may be determined for a subset of samples in the current string. At least one sample in the current string is not predicted using a flip operation, where the subset of samples is different from at least one sample in the current string.

[0298] Generally, it may be determined whether a flip operation is used to predict the current string based on various criteria, such as the block-level flag of the current block, the string-level flag of the current string, the number of at least one string in the current block, the string length of the current string, the shape of the current string, the shape of the current string, whether the current string includes escape samples, and / or the number of rows and / or columns in the current string.

[0299] If it is determined that the flip operation is used to predict the current string, the process (2000) proceeds to (S2030). Otherwise, if it is determined that the flip operation is not used to predict the current string, the process (2000) proceeds to (S2099) and ends.

[0300] As described above, the flip type of the flip operation can be any suitable type, such as (i) a vertical flip operation, (ii) a horizontal flip operation, or (iii) a combined flip operation. The combined flip operation may include a horizontal flip and a vertical flip of the original reference string. The combined flip operation can be performed in a single step or multiple steps.

[0301] For example, after signaling the string length of the current string, the index of the current string can be signaled in the encoded information. The index may indicate the flip type of the flip operation. The index may include a flip type index (e.g., a flip direction index), as described in Figures 15 - 19 as described.

[0302] At (S2030), based on the determination to perform a flip operation to predict the current string, the original reference string in the current image can be determined based on the string vector (SV) of the current string, as described in Figures 15 - 19As described. In one example, a sample (e.g., sample (1524)) in an original reference string (e.g., string (1534)) is determined based on an SV (e.g., SV3) and a corresponding sample (e.g., sample (1501)) in a current string (e.g., string (1532)). Subsequently, the shape of the original reference string (including the positions of the remaining samples (e.g., samples (1521)-(1523) and (1525)-(1526)) in the original reference string) can be determined by performing an inverse flip operation on the current string. The inverse flip operation is the inverse of the flip operation. Process (2000) proceeds to (S2040).

[0303] At (S2040), a flipped reference string can be generated by performing a flip operation on the original reference string, e.g., performing a flip operation on the original reference string based on a flip operation or flip type of the current string indicated by an index (e.g., a flip type index) to generate a flipped reference string.

[0304] The flipped reference string can be generated by vertically flipping the original reference string based on the flip operation or flip type being a vertical flip operation. The flipped reference string can be generated by horizontally flipping the original reference string based on the flip operation or flip type being a horizontal flip operation. The flipped reference string can be generated by vertically and horizontally flipping the original reference string based on the flip operation or flip type being a combined flip operation. Alternatively, the flipped reference string can be generated based on the original reference string in a single step based on the flip type being a combined flip operation. Process (2000) proceeds to (S2050).

[0305] At (S2050), the current string can be reconstructed based on the flipped reference string, as described in Figures 15 - 19 As described. Process (2000) proceeds to (S2099) and ends. In some examples, the current string can include samples that are not predicted using a flip operation (referred to as non-flipped samples) and samples that are predicted using a flip operation. Therefore, the non-flipped samples are not predicted or encoded based on the flipped reference string. For example, the non-flipped samples are directly encoded escape samples, as described in Figures 18 - 19 As described. The predicted samples predicted using the flip operation and the encoded non-flipped samples can be combined together to form a predicted current string, as described in Figures 18 - 19 As described.

[0306] Process (2000) can be appropriately adjusted. Steps in process (2000) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used. In one example, after predicting the current string based on the flipped reference string, the residual of the current block is transformed into transform coefficients, and the transform coefficients are quantized. Alternatively, the transformation and quantization are skipped, and the residual of the current block is considered to be zero.

[0307] In one example, at (S2020), if it is determined that the flipping operation is not used to predict the current string, the process (2000) may include the following steps: determining a reference string based on another SV, and using the reference string to predict the current string without any flipping operation.

[0308] Embodiments of the present disclosure may be used alone or in any order combination. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. Embodiments of the present disclosure may be applied to a luminance block or a chrominance block.

[0309] The above techniques may be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 21 A computer system (2100) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0310] The computer software may be encoded using any suitable machine code or computer language, and any suitable machine code or computer language may be subject to mechanisms such as assembly, compilation, linking, or the like to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0311] The instructions may be executed on various types of computers or their components, such as personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0312] Figure 21 The components of the shown computer system (2100) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (2100).

[0313] The computer system (2100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users, such as the following: tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., speech, clapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices may also be used to acquire certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, captured images from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0314] The human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard (2101), mouse (2102), touchpad (2103), touch screen (2110), data glove (not shown), joystick (2105), microphone (2106), scanner (2107), camera (2108).

[0315] The computer system (2100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of the touch screen (2110), data glove (not shown), or joystick (2105), but may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (2109), headphones (not depicted)), visual output devices (e.g., screens (2110) including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input function, each screen having or not having tactile feedback function, some of which are capable of outputting two-dimensional visual output or output beyond three dimensions through means such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).

[0316] The computer system (2100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2120) with media such as CD / DVD (2121), thumb drives (2122), removable hard disk drives or solid state drives (2123), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0317] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0318] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (2100)) attached to certain common data ports or peripheral buses (2149); as described below, other network interfaces are typically integrated into the kernel of the computer system (2100) by attaching to the system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (2100) may use any of these networks to communicate with other entities. Such communication may be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., CANbus connected to certain CANbus devices), or two-way, e.g., using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.

[0319] The above-mentioned human-machine interface devices, human-machine accessible storage devices, and network interfaces may be attached to the kernel (2140) of the computer system (2100).

[0320] The kernel (2140) may include one or more central processing units (CPUs) (2141), a graphics processing unit (GPU) (2142), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2143), a hardware accelerator (2144) for certain tasks, a graphics adapter (2150), etc. These devices, as well as a read-only memory (ROM) (2145), a random access memory (2146), internal mass storage such as an internal non-user accessible hard disk drive, SSD, etc. (2147) may be connected via a system bus (2148). In some computer systems, the system bus (2148) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripherals may be directly attached to the system bus (2148) of the kernel or attached to the system bus (2148) of the kernel via a peripheral bus (2149). In one example, a screen (2110) may be connected to the graphics adapter (2150). The architecture of the peripheral bus includes PCI, USB, etc.

[0321] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can execute certain instructions, which can be combined to form the above computer code. The computer code can be stored in the ROM (2145) or RAM (2146). Transitional data can also be stored in the RAM (2146), while permanent data can be stored, for example, in the internal mass storage (2147). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more of the following: one or more CPUs (2141), GPUs (2142), mass storage (2147), ROM (2145), RAM (2146), etc.

[0322] Computer code for performing various computer-implemented operations can be present on a computer-readable medium. The medium and the computer code can be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the type well-known and available to those skilled in the field of computer software.

[0323] By way of example and not limitation, a computer system having an architecture (2100), particularly a core (2140), can provide functionality due to software executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of the core (2140), such as the core internal mass storage (2147) or ROM (2145). The software implementing the various embodiments of the present disclosure can be stored in such devices and executed by the core (2140). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (2140), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute the specific processes or specific portions of the specific processes described herein, including defining data structures stored in the RAM (2146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hard-wired or otherwise embodied in a circuit (e.g., accelerator (2144)), which can replace the software or operate in conjunction with the software to execute the specific processes or specific portions of the specific processes described herein. In appropriate instances, portions that refer to software can include logic, and vice versa. In appropriate instances, portions that refer to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) that stores software for execution, a circuit that embodies logic for execution, or includes both. The present disclosure encompasses any suitable combination of hardware and software.

[0324] Appendix A: Acronyms

[0325] JEM: Joint Exploration Model

[0326] VVC: Versatile Video Coding

[0327] BMS: Benchmark Set

[0328] MV: Motion Vector

[0329] HEVC: High Efficiency Video Coding

[0330] SEI: Supplementary Enhancement Information

[0331] VUI: Video Usability Information

[0332] GOP: Group of Pictures

[0333] TU: Transform Unit

[0334] PU: Prediction Unit

[0335] CTU: Coding Tree Unit

[0336] CTB: Coding Tree Block

[0337] PB: Prediction Block

[0338] HRD: Hypothetical Reference Decoder

[0339] SNR: Signal-to-Noise Ratio

[0340] CPU: Central Processing Unit

[0341] GPU: Graphics Processing Unit

[0342] CRT: Cathode Ray Tube

[0343] LCD: Liquid Crystal Display

[0344] OLED: Organic Light-Emitting Diode

[0345] CD: Compact Disc

[0346] DVD: Digital Video Disc

[0347] ROM: Read-Only Memory

[0348] RAM: Random Access Memory

[0349] ASIC: Application-Specific Integrated Circuit

[0350] PLD: Programmable Logic Device

[0351] LAN: Local Area Network

[0352] GSM: Global System for Mobile Communications

[0353] LTE: Long Term Evolution

[0354] CANBus: Controller Area Network Bus

[0355] USB: Universal Serial Bus

[0356] PCI: Peripheral Component Interconnect

[0357] FPGA: Field Programmable Gate Array

[0358] SSD: Solid State Drive

[0359] IC: Integrated Circuit

[0360] CU: Coding Unit

[0361] Although the present disclosure has described multiple exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it is understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.

Claims

1. A method for video decoding, characterized in that, The method includes: Decoding the encoded information of a current block in a current image according to an encoded video bitstream, the encoded information indicating a string copy mode of the current block, the current block including a current string, the current string being one of a plurality of strings included in the current block, the encoded information including a block-level flag of the current block, a first value of the block-level flag indicating that for each of the plurality of strings, a corresponding flipping operation is used to predict the string, and a second value of the block-level flag indicating that no flipping operation is performed on the plurality of strings; Determining whether to perform a flipping operation to predict the current string includes: determining to perform the flipping operation to predict the current string based on the block-level flag having the first value; and Based on determining to perform the flipping operation to predict the current string, it further includes: Determining an original reference string based on a string vector of the current string, Generating a flipped reference string by performing the flipping operation on the original reference string, and Reconstructing the current string based on the flipped reference string.

2. The method according to claim 1, characterized in that The flipping operation is one of the following operations: (i) a vertical flipping operation, (ii) a horizontal flipping operation, and (iii) a combined flipping operation; And The generating a flipped reference string by performing the flipping operation on the original reference string includes: Based on the flipping operation being the vertical flipping operation, generating the flipped reference string by vertically flipping the original reference string, Based on the flipping operation being the horizontal flipping operation, generating the flipped reference string by horizontally flipping the original reference string, and Based on the flipping operation being the combined flipping operation, generating the flipped reference string by vertically and horizontally flipping the original reference string.

3. The method according to claim 1, characterized in that The method further includes: The encoded information includes a string-level flag of the current string located after the string length of the current string, the string-level flag indicating whether to perform the flipping operation to predict the current string.

4. The method according to claim 1, wherein The method further includes: The encoded information includes the string-level flag of the current string, and The determining whether to perform a flipping operation to predict the current string further includes: determining to perform the flipping operation to predict the current string based on the string-level flag indicating to perform the flipping operation to predict the current string.

5. The method according to claim 1, wherein The determining whether to perform a flipping operation to predict the current string further includes: determining to prohibit the flipping operation on the current string based on the number of the plurality of strings in the current block being greater than a first threshold.

6. The method according to claim 1, wherein The determining whether to perform a flipping operation to predict the current string further includes: determining to prohibit the flipping operation on the current string based on the string length of the current string being less than a second threshold, the string length of the current string corresponding to the number of samples in the current string.

7. The method according to claim 1, characterized in that, The determining whether to perform a flipping operation to predict the current string further includes: Determining to prohibit the flipping operation on the current string based on the current string having a rectangular shape.

8. The method according to claim 1, wherein The method further includes: The current string has a rectangular shape and includes samples that are not predicted using the flipping operation, and determining whether to perform a flipping operation to predict the current string further includes: determining to perform the flipping operation to predict a plurality of samples in the current string that are different from the samples that are not predicted using the flipping operation.

9. The method according to claim 8, wherein the current string is the current block.

10. The method according to any one of claims 1-9, characterized in that, The reconstruction includes: dequantizing the transform coefficients; and inversely transforming the transform coefficients into the residual of the current block.

11. A device for video decoding, characterized in that, The apparatus includes: a processing circuit configured to: decode the encoded information of a current block in a current image according to an encoded video bitstream, the encoded information indicating a string copy mode of the current block, the current block including a current string, the current string being one of a plurality of strings included in the current block, the encoded information including a block-level flag of the current block, a first value of the block-level flag indicating that for each of the plurality of strings, a corresponding flipping operation is used to predict the string, and a second value of the block-level flag indicating that no flipping operation is performed on the plurality of strings; determining whether to perform a flipping operation to predict the current string includes: determining to perform the flipping operation to predict the current string based on the block-level flag having the first value; and based on determining to perform the flipping operation to predict the current string, determine an original reference string based on a string vector of the current string, generate a flipped reference string by performing the flipping operation on the original reference string, and reconstruct the current string based on the flipped reference string.

12. A method for video encoding, characterized in that, The method includes: encode the encoding information of a current block in a current image according to a video bitstream to be encoded, the encoding information indicating a string copy mode of the current block, the current block including a current string, the current string being one of a plurality of strings included in the current block, the encoding information including a block-level flag of the current block, a first value of the block-level flag indicating that for each of the plurality of strings, a corresponding flipping operation is used to predict the string, and a second value of the block-level flag indicating that no flipping operation is performed on the plurality of strings; determining whether to perform a flipping operation to predict the current string includes: determining to perform the flipping operation to predict the current string based on the block-level flag having the first value; and based on determining to perform the flipping operation to predict the current string, further includes: determine an original reference string based on a string vector of the current string, generate a flipped reference string by performing the flipping operation on the original reference string, and encode the current string based on the flipped reference string.

13. An apparatus for video coding, characterized in that, The apparatus includes: a processing circuit configured to: Encode the encoding information of a current block in a current picture according to a video bitstream to be encoded, where the encoding information indicates a string copy mode of the current block, the current block includes a current string, the current string is one of a plurality of strings included in the current block, the encoding information includes a block-level flag of the current block, and a first value of the block-level flag indicates that for each of the plurality of strings, a corresponding flipping operation is used to predict the string, and a second value of the block-level flag indicates that no flipping operation is performed on the plurality of strings; Determining whether to perform a flipping operation to predict the current string includes: determining to perform the flipping operation to predict the current string based on the block-level flag having the first value; and Based on determining to perform the flipping operation to predict the current string, it further includes: Determining an original reference string based on a string vector of the current string, Generating a flipped reference string by performing the flipping operation on the original reference string, and Encoding the current string based on the flipped reference string.

14. A computer device, characterized in that, The computer device includes: a processor and a memory, and instructions are stored in the memory, and when the instructions are executed by the processor, the computer device executes the video decoding method according to any one of claims 1-10 and the video encoding method according to claim 12.

15. A computer-readable medium storing instructions, characterized in that, When the instructions are executed by a computer, the computer is caused to execute the video decoding method according to any one of claims 1-10 and the video encoding method according to claim 12.

16. A method for processing a video bitstream, characterized in that, The video bitstream is decoded based on the video decoding method according to any one of claims 1-10 or generated according to the video encoding method according to claim 12.

Citation Information

Patent Citations

  • Block flipping and skip mode in intra block copy prediction

    CN105247871A