Method and apparatus for video encoding / decoding
By adaptively selecting the resolution of block vectors during intra-frame image encoding and decoding, the method addresses inefficiencies in existing technologies, enhancing compression efficiency and computational performance.
Patent Information
- Application Number
- JP2023220187
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-07
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2039-03-06
AI Technical Summary
Existing video encoding and decoding technologies face challenges in efficiently encoding and decoding intra-frame images, particularly in adapting the resolution of block vectors to optimize compression efficiency.
The method involves a processing circuit that decodes prediction information from an encoded video bitstream, selects a resolution for the block vector difference from a set of candidate resolutions, and determines the block vector of the current block based on the selected resolution and a predicted value, ultimately reconstructing samples of the current block.
This approach enhances the encoding and decoding efficiency by adaptively selecting the resolution of block vectors, thereby improving the compression ratio and reducing computational complexity.
Smart Images

Figure 0007695333000001 
Figure 0007695333000002 
Figure 0007695333000003
Abstract
Description
Technical Field
[0001] “Cross - Reference to Related Applications” This disclosure claims priority to U.S. Provisional Application No. 62 / 639,862, filed on Mar. 7, 2018, entitled “Adaptive Methods for Motion and Block Vector Resolution in Image and Video Compression,” and U.S. Patent Application No. 16 / 182,788, filed on Nov. 7, 2018, entitled “Methods and Apparatus for Video Encoding / Decoding,” the entireties of which are incorporated herein by reference. [Technical Field] This disclosure generally describes embodiments related to video encoding / decoding.
Background Art
[0002] The description of the background art provided herein is intended to present the context of the disclosure as a whole. The degree of the work of the currently named inventors described in this background art section and in each aspect of this specification is not presented as prior art at the time of filing of this disclosure, nor is it expressly or implicitly admitted as prior art of this disclosure.
[0003] Video encoding and decoding can be performed using inter - frame image prediction with motion compensation. Uncompressed digital video can include a series of images, each image having a spatial dimension of, for example, 1920×1080 luminance samples and associated chrominance samples. This series of images can have, for example, 60 images per second or a fixed or variable image rate of 60 Hertz (Hz) (informally known as the frame rate). Uncompressed video has very high bit - rate requirements. For example, a 1080p60 4:2:0 video with 8 bits per sample (luminance sample resolution of 1920x1080 at a frame rate of 60 Hz) requires nearly 1.5 Gbit / s of bandwidth. Such a video requires more than 600 GB of storage space per hour.
[0004] One purpose of video encoding and decoding is to reduce redundant information in the input video signal by compression. Compression can help reduce the requirements for the above bandwidth or storage space, and in some cases, can reduce by more than two digits. Both lossless and lossy compression, as well as combinations of both, can be used. Lossless compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used in the expected application. In the case of video, lossy compression is widely used. The amount of allowable distortion depends on the application. For example, a user consuming a certain streaming application can tolerate higher distortion than a user of a television contribution application. The achievable compression ratio reflects the fact that higher allowed / tolerable distortion can produce a higher compression ratio.
[0005] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy encoding.
[0006] Video encoding / decoding techniques can include techniques known as intra-frame encoding. In intra-frame encoding, sample values are represented without reference to samples or other data from previously reconstructed reference images. In some video codecs, an image is spatially subdivided into sample blocks. If all sample blocks are encoded in the intra-frame mode, the image can be an intra-frame image. Intra-frame images and their derivatives, such as independent decoder refresh images, can be used to reset the state of the decoder and thus can be used as the first image or a still image in an encoded video bitstream and a video session. Samples of an intra-frame block are used for transformation, and the transformation coefficients can be quantized before entropy encoding. Intra-frame prediction can be a technique to minimize sample values in the pre-transform domain. In some cases, the smaller the DC value and AC coefficients after transformation, the fewer bits are required with a given quantization step size to represent the block after entropy encoding.
[0007] Conventional intra-frame encoding, such as that known from the MPEG-2 encoding technique, does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to obtain data blocks from surrounding sample data and / or metadata, where the surrounding sample data and / or metadata are obtained during the encoding / decoding of spatially adjacent blocks and before the decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. It should be noted that in at least some cases, intra-frame prediction uses only reference data from the currently reconstructed image without using reference data from a reference image.
[0008] There can be many different forms of intra prediction. In a given video coding technology, if two or more of such techniques can be used, the techniques in use can perform encoding in an intra prediction mode. In some cases, the mode may have sub - modes and / or parameters, and these modes may be encoded alone or may be included in a mode codeword. Which codeword to use for a given mode / sub - mode / parameter combination affects the encoding efficiency gain by intra prediction, so there are such cases also in the entropy coding technology used to convert the codeword into a bitstream.
[0009] A specific mode of intra prediction was introduced in H.264, improved in H.265, and further improved in newer encoding / decoding technologies such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). A prediction block can be formed using adjacent sample values that belong to already available samples. The sample values of the adjacent samples are copied into the prediction block according to a certain direction. The reference to the direction in use may be encoded in the bitstream or may itself be predicted.
[0010] Referring to FIG. 1, in the lower right, a subset of 9 prediction directions known from 35 predictable directions of H.265 is depicted. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.
[0011] Continuing to refer to FIG. 1, in the upper left, a square block (104) of 4×4 samples is depicted (shown by the thick dashed line). The square block (104) contains 16 samples, and each sample is labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample of block (104) in both the Y and X dimensions. Since this block is a 4×4 size sample, S44 is in the lower right. Additionally, reference samples following a similar numbering scheme are also shown. The reference samples are labeled with an "R", their Y position (e.g., row index) and X position (e.g., column index) relative to block (104). In both H.264 and H.265, since the predicted samples are adjacent to the block being reconstructed, there is no need to use negative values.
[0012] Intra-frame image prediction can function by copying the reference sample value from adjacent samples according to the predicted direction signaled. For example, assuming that the encoded video bitstream contains signaling, this signaling indicates, for this block, a predicted direction that coincides with arrow (102), i.e., the sample is predicted from one or more predicted samples in the upper right at a horizontal and 45-degree angle. In this case, samples S41, S32, S23, S14 are predicted from the same R05. And sample S44 is predicted from R08.
[0013] In some cases, in order to calculate the reference sample, especially when the direction is not evenly divisible by 45 degrees, for example, the values of multiple reference samples can be combined through interpolation.
[0014] With the development of video coding technology, the number of possible directions has already increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VC / BMS can support up to 65 directions at the time of disclosure. Experiments have been conducted to identify the most possible directions, and several techniques in entropy coding are used to represent those possible directions with fewer bits, at the expense of some for the less likely directions. Furthermore, the directions themselves may be predicted from adjacent directions used in adjacent already decoded blocks.
[0015] FIG. 2 shows a schematic diagram (201) depicting 65 intra-prediction directions by JEM to illustrate the number of prediction directions increasing over time.
[0016] The mapping from intra-prediction directions to the bits representing the directions in the encoded video bitstream can vary depending on the video coding technology and can be anything from a simple direct mapping to the prediction directions, to complex adaptive schemes including the intra-prediction mode, codewords, the most likely mode, and similar techniques. Those skilled in the art are readily proficient in those techniques. However, in all cases, in video content, there may be certain directions that are statistically less likely to occur than other specific directions. Since the purpose of video compression is redundancy reduction, those less likely directions are represented with more bits than the more likely directions in a properly functioning video coding technology. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0017] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, the apparatus includes a processing circuit for decoding video. The processing circuit decodes prediction information of a current block from an encoded video bitstream. The prediction information indicates an intra-block copy mode. The processing circuit selects a resolution of a block vector difference of the current block from a set of a plurality of candidate resolutions, and determines a block vector of the current block based on the selected resolution of the block vector difference and a predicted value of the block vector of the current block. Then, the processing circuit reconstructs at least one sample of the current block based on the block vector.
[0018] In one example, the processing circuit selects a resolution of the current block from two candidate resolutions based on a 1-bit flag.
[0019] According to aspects of the present disclosure, the processing circuit selects a first resolution of a first component of the block vector difference and selects a second resolution of a second component of the block vector difference. For example, the processing circuit selects the first resolution of the first component of the block vector difference based on a first flag and selects the second resolution of the second component of the block vector difference based on a second flag.
[0020] In some embodiments, the processing circuit selects the first resolution of the first component of the block vector difference based on a first component of the predicted value of the block vector and selects the second resolution of the second component of the block vector difference based on a second component of the predicted value of the block vector.
[0021] In some examples, when the first component of the predicted value of the block vector is smaller than a threshold, the processing circuit selects a default resolution as the first resolution from a set of two resolutions, and when the first component of the predicted value of the block vector is larger than the threshold, the processing circuit selects a smaller resolution as the first resolution from a set of two resolutions.
[0022] In one example, when the processing circuit uses a resolution different from the resolution for which the block vector prediction value is selected, instead of rounding the block vector prediction value to calculate the block vector, the processing circuit adds the block vector difference to the block vector prediction value. In another example, when the processing circuit has a resolution different from the resolution for which the block vector prediction value is selected, the processing circuit rounds the block vector prediction value of the current block to the selected resolution, and adds the block vector difference to the rounded block vector prediction value to calculate the block vector.
[0023] In some embodiments, the processing circuit modifies at least one of the block vector difference and the block vector prediction value to constrain the block vector to a valid region.
[0024] Aspects of the present disclosure provide a non-transitory computer-readable medium storing instructions that, when executed by a computer that decodes video, cause the computer to perform a method of video decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, where
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Best Mode for Carrying Out the Invention
[0026] FIG. 3 is a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) can encode video data (e.g., a video image stream captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video image, and display the video image based on the restored video data. Unidirectional data transmission is common in media serving applications and the like.
[0027] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data that may occur, for example, during a video conference. In the case of bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., a video image stream captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), and can decode the encoded video data to restore the video image, and based on the restored video data, display the video image on an accessible display device.
[0028] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may be shown as servers, personal computers, and smartphones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find applications having laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transmit encoded video data among the terminal devices (310), (320), (330), and (340), and includes wired (wired) and / or wireless communication networks. The communication network (350) can exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the present disclosure, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described herein below.
[0029] FIG. 4 shows the placement of a video encoder and a video decoder in a streaming environment as an example of an application to the disclosed subject matter. The disclosed subject matter is equally applicable to other video support applications, including, for example, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc., video conferencing, digital TV, and the like.
[0030] The streaming system can include a capture subsystem (413), which can include a video source (401), such as a digital camera, and create, for example, an uncompressed video image stream (402). In one example, the video image stream (402) includes samples taken by a digital camera. The video image stream (402), drawn with a thick line to emphasize the high data volume when compared to the encoded video data (404) (or encoded video bitstream), can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement various aspects of the disclosed subject matter, as will be described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), drawn with a thin line to emphasize the lower data volume when compared to the video image stream (402), can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data to generate an outgoing video image stream (411), which can be displayed on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video encoding / compression standard.Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally referred to as next-generation video coding or VVC (Versatile Video Coding). The disclosed subject matter can be used in the context of VVC.
[0031] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can similarly include a video encoder (not shown).
[0032] FIG. 5 shows a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.
[0033] The receiver (531) can receive one or more encoded video sequences decoded by the video decoder (510), and in the same or another embodiment, can receive one encoded video sequence at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence can be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) can receive the encoded video data along with other data, such as, for example, encoded audio data and / or auxiliary data streams, which can be transmitted to respective usage entities (not shown). The receiver (531) can separate the encoded video sequence from other data. To prevent network jitter, the buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as the "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be located outside the video decoder (510) (not shown). In still other cases, for example, to prevent network jitter, there may be a buffer memory outside the video decoder (not shown), and further, for example, to process the playback timing, there may be another buffer memory (515) inside the video decoder (510). If the receiver (531) receives data from a store-and-forward device or an isosynchronous network having sufficient bandwidth and controllability, the buffer memory (515) may not be necessary or may be small.For use in a best effort packet network such as the Internet, buffer memory (515) may be required and can be made relatively large, advantageously of an adaptable size, and can be implemented at least partially in an operating system or in a similar element external to video decoder (510) (not shown).
[0034] The video decoder (510) can include an analyzer (520) for reconstructing symbols (521) from the encoded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and potential information for controlling a rendering device (such as a display screen) (512) that is not an essential part of the electronic device (530) but can be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a visual user ability information (VUI) parameter set fragment (not shown). The analyzer (520) can perform analysis / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can follow video encoding techniques or standards and can follow various principles including variable length encoding, Huffman encoding, arithmetic encoding with or without context sensitivity, etc. The analyzer (520) can extract a set of at least one subgroup parameter of at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups can include a group of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The analyzer (520) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0035] The analyzer (520) can perform an entropy decoding / analysis operation on the video sequence received from the buffer memory (515) to create symbols (521).
[0036] The reconstruction of symbol (521) can be related to multiple different units depending on the type of the encoded video image or a part thereof (e.g., inter - frame image and intra - frame image, inter - frame block and intra - frame block) and other factors. Which unit it is related to and how it is related can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following multiple units is not shown for clarity.
[0037] In addition to the function blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In actual embodiments operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0038] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbol (521) from the parser (520), the quantized transform coefficients and control information including what transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).
[0039] In some cases, the output samples of the scaler / inverse transform unit (551) can belong to in-frame encoded blocks, i.e., blocks that do not use prediction information from the previously reconstructed image but can use prediction information from the previously reconstructed portion of the current image. Such prediction information may be provided by the in-frame image prediction unit (552). In some cases, the in-frame image prediction unit (552) uses the surrounding already reconstructed information extracted from the current image buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current image buffer (558) buffers, for example, the partially reconstructed current image and / or the fully reconstructed current image. The aggregator (555) adds, in some cases, sample-by-sample, the prediction information generated by the in-frame prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0040] In other cases, the output samples of the scaler / inverse transform unit (551) can belong to inter-frame encoded blocks and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference image memory (557) to extract the samples used for prediction. After the extracted samples are motion-compensated based on the symbols (521) associated with the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference image memory (557) when the motion compensation prediction unit (553) extracts the prediction samples can be controlled, for example, by the motion vectors available to the motion compensation prediction unit (553) in the form of symbols (521) that can have X, Y, and reference image components. Motion compensation can also include interpolation of the sample values extracted from the reference image memory (557), a motion vector prediction mechanism, etc., when an exact sub-sample motion vector is in use.
[0041] The output samples of the aggregator (555) may be employed by various loop filtering techniques in the loop filter unit (556). The video compression technology includes in-loop filter techniques that are controlled by parameters available to the loop filter unit (556) as symbols (521) from the parser (520) and that may be included in an encoded video sequence (also referred to as an encoded video bitstream), and may also respond to meta information obtained during the period of decoding a previous portion (in decoding order) of the encoded image or encoded video sequence and to previously reconstructed and loop filtered sample values.
[0042] The output of the loop filter unit (556) can be output to the rendering device (512) and can be a sample stream that can be stored in the reference image memory (557) for use in future inter-frame image prediction.
[0043] When a particular encoded image is fully reconstructed, it can be used as a reference image for future prediction. For example, when the encoded image corresponding to the current image is fully reconstructed and the encoded image is identified as a reference image (e.g., by the parser (520)), the current image buffer (558) can become part of the reference image memory (557), and a new current image buffer can be reallocated before disclosing the reconstruction of subsequent encoded images.
[0044] The video decoder (510) can perform a decoding operation according to a predetermined video compression technique in a standard such as, for example, ITU-T Rec. H.265. The encoded video sequence can comply with the syntax specified by the video compression technique or standard being used in the sense that the encoded video sequence complies with both the syntax of the video compression technique or standard and the profile as a document of the video compression technique or standard. Specifically, the profile can select some tools from all the tools available in the video compression technique or standard as the only tools available in that profile. It is also necessary for compliance that the complexity of the encoded video sequence is within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy restricts the maximum image size, the maximum frame rate, the maximum reconstructed sample rate (measured, for example, in mega samples per second), the maximum reference image size, etc. The restrictions set by the hierarchy can, in some cases, be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0045] In one embodiment, the receiver (531) can receive additional (redundant) data along with the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) extension layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0046] FIG. 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.
[0047] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that captures video images to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0048] The video source (601) can provide a source video sequence encoded by the video encoder (603) in the form of a digital video sample stream, and the digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit...), any color space (e.g., BT.601 Y CrCB, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device that stores previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual images that give motion when viewed in sequence. The images themselves may be configured as a spatial pixel array, where each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0049] According to one embodiment, the video encoder (603) can encode and compress the images of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is one function of the controller (650). In some embodiments, the controller (650) controls other functional units and is functionally coupled to other functional units as described below. This coupling is not shown for clarity. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, λ (lambda) value of rate distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.
[0050] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an overly simplified explanation, in one example, the encoding loop can include a source coder (630) (which is responsible for creating symbols such as a symbol stream based on, for example, an input image to be encoded and a reference image), and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs symbols in a similar manner as a (remote) decoder creates sample data (because in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference image memory (634). Since bit-exact results are obtained regardless of the position of the decoder (local or remote) due to the decoding of the symbol stream, the contents of the reference image memory (634) exactly correspond bit-by-bit between the local encoder and the remote encoder. In other words, the reference image samples that the prediction part of the encoder "saw" are exactly the same as the sample values that the decoder "sees" when using the prediction during the decoding period. This basic principle of reference image synchronization (and the drift that occurs, for example, when synchronization is not maintained due to channel errors) is also used in some related technologies.
[0051] The operation of the "local" decoder (633) may be the same as the operation of a "remote" decoder such as the video decoder (510) that has already been described in detail above in connection with FIG. 5. However, referring to FIG. 5 more briefly, since symbols are available and the encoding / decoding of symbols into the video sequence encoded by the entropy coder (645) and the parser (520) can be lossless, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully executable by the local decoder (633).
[0052] At this point, it has been observed that any decoder technology other than parsing / entropy decoding present in the decoder must necessarily be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technology can be omitted since it is the reverse of the decoder technology described comprehensively. More detailed description is needed only for specific areas and is provided below.
[0053] During operation, in some embodiments, the source coder (630) can perform motion-compensated predictive coding, which predictively encodes an input image by referring to one or more previously encoded images designated as "reference images" from a video sequence. In this way, the coding engine (632) encodes the difference between a pixel block of the input image and a pixel block of a reference image that can be selected as a prediction reference for the input image.
[0054] The local video decoder (633) can decode the encoded video data of an image that can be designated as a reference image based on the symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the encoded video data is decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) can copy the decoding process that can be performed by the video decoder for the reference image and store the reconstructed reference image in the reference image cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference image having the same content as the reconstructed reference image obtained by the remote video decoder (in the absence of transmission errors).
[0055] Predictor (635) can perform a prediction search on the encoding engine (632). That is, for a new image to be encoded, the predictor (635) can search the reference image memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference image motion vectors, block shapes, etc., that function as appropriate prediction references for the new image. The predictor (635) can operate on a pixel block-by-block basis based on the sample block to find an appropriate prediction reference. In some cases, the input image can have prediction references drawn from a plurality of reference images stored in the reference image memory (634) as determined by the search results obtained by the predictor (635).
[0056] The controller (650) can manage the encoding operation of the source coder (630), including, for example, setting parameters and subgroup parameters used for encoding video data.
[0057] The outputs of all the above functional units can be entropy encoded by the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0058] The transmitter (640) can buffer the encoded video sequence generated by the entropy coder (645) in preparation for transmission via a communication channel (660) that can be a hardware / software link to a storage device for storing the encoded video data. The transmitter (640) can merge the encoded video data from the video coder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0059] The controller (650) can manage the operation of the video encoder (603). During the encoding period, the controller (650) can assign a specific encoded image type to each encoded image, which may affect the encoding technique applicable to each image. For example, an image is often assigned as one of the following image types, that is, an intra-frame image (I-image) may be encoded and decoded without using any other image in the sequence as a prediction source.
[0060] Some video codecs allow different types of intra-frame images such as Independent Decoder Refresh (IDR) images. Those skilled in the art understand the variants of I-images and their applications and functions.
[0061] A predicted image (P-image) may be encoded and decoded using intra-frame prediction or inter-frame prediction that predicts the sample values of each block using at most one motion vector and a reference index.
[0062] A bi-directionally predicted image (B-image) may be encoded and decoded using intra-frame prediction or inter-frame prediction that predicts the sample values of each block using at most two motion vectors and a reference index. Similarly, multiple predicted images can use two or more reference images and related metadata for the reconstruction of a single block.
[0063] Source images can generally be spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be encoded block by block. These blocks can be encoded predictively by referring to other (already encoded) blocks as determined by the encoding assignment applied to each image of the block. For example, blocks of an I image may be encoded non-predictively or may be encoded predictively by referring to already encoded blocks of the same image (spatial prediction or intra-frame prediction). Pixel blocks of a P image may be encoded predictively via spatial prediction or via temporal prediction by referring to a previously encoded reference image. Blocks of a B image may be encoded predictively via spatial prediction or via temporal prediction by referring to one or two previously encoded reference images.
[0064] The video encoder (603) can perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T H.265. In that operation, the video encoder (603) can perform various compression operations including a predictive encoding operation that utilizes temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to a syntax specified by the video encoding technique or standard used.
[0065] In one embodiment, the transmitter (640) can transmit additional data along with the encoded video. The source coder (630) can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures or slices, supplementary enhancement information (SEI) messages, visual user utility information (VUI) parameter set fragments, and the like.
[0066] Video can be captured as a plurality of source images (video images) in a time series. Intra-frame image prediction (often abbreviated as intra-frame prediction) utilizes the spatial correlation in a given image, while inter-frame image prediction utilizes the correlation (temporal or other) between images. In one example, a particular image being encoded / decoded, called the current image, is divided into blocks. If a block of the current image is similar to a reference block in a reference image that has been previously encoded and is still buffered in the video, the block of the current image can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference image and can have a third dimension that identifies the reference image if multiple reference images are being used.
[0067] In some embodiments, bidirectional prediction techniques can be used for inter-frame image prediction. According to the bidirectional prediction technique, for example, two reference images such as a first and a second reference image that are both in front of the current image in the video in the order of decoding (although they may be in the past and future respectively in the order of display) are used. A block in the current image can be encoded by a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted by a combination of the first reference block and the second reference block.
[0068] Furthermore, to improve the encoding efficiency, the merge mode technique can be used in inter-frame image prediction.
[0069] According to some embodiments of the present disclosure, predictions such as inter-frame image prediction and intra-frame image prediction are performed in units of blocks. For example, according to the HEVC standard, images in a video image sequence are divided into coding tree units (CTUs) for compression, and the CTUs in an image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU may be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter-frame prediction type or an intra-frame prediction type. The CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Usually, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Using the luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0070] FIG. 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values in a current video image in a video image sequence and encode the processing block into an encoded image that is part of the encoded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.
[0071] In the example of HEVC, the video encoder (703) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples for example. The video encoder (703) determines whether to encode the processing block using an intra mode, an inter mode, or a bidirectional prediction mode, for example using rate distortion optimization. If the processing block is encoded in the intra mode, the video encoder (703) can encode the processing block into the encoded image using intra prediction techniques, and if the processing block is encoded in the inter mode or the bidirectional prediction mode, the video encoder (703) can encode the processing block into the encoded image using inter prediction or bidirectional prediction techniques respectively. In certain video encoding techniques, the merge mode can be an inter picture prediction submode in which the motion vector is derived from one or more motion vector prediction values when not utilizing the advantages of the encoded motion vector components other than the predicted value. In certain other video encoding techniques, there may be motion vector components applicable to the subject block. In one example, the video encoder (703) includes other components such as a mode decision module (not shown) for determining the mode of the processing block.
[0072] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general purpose controller (721), and an entropy encoder (725) that are coupled together as shown in FIG. 7.
[0073] The inter-frame encoder (730) receives samples of a current block (e.g., a processing block), compares the block with one or more reference blocks in a reference image (e.g., blocks in the previous and subsequent images), generates inter-frame prediction information (e.g., redundant information description by an inter-frame encoding technique, a motion vector, merge mode information), and is configured to calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference image is a decoded reference image that has been decoded based on the encoded video information.
[0074] The intra-frame encoder (722) receives samples of a current block (e.g., a processing block) and, in some cases, compares the block with blocks that have already been encoded in the same image, generates quantized coefficients after transformation, and, in some cases, is configured to generate intra-frame prediction information (e.g., intra-frame prediction direction information by one or more intra-frame encoding techniques). In one example, the intra-frame encoder (722) also calculates an intra-frame prediction result (e.g., a predicted block) based on the intra-frame prediction information and reference blocks in the same image.
[0075] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines the mode of a block and provides a control signal to the switch (726) based on that mode. For example, when the mode is intra-frame, the general-purpose controller (721) controls the switch (726) to select the intra-frame mode result used by the residual calculator (723), selects the intra-frame prediction information, and controls the entropy encoder (725) to include the intra-frame prediction information in the code stream. Also, when the mode is inter-frame mode, the general-purpose controller (721) controls the switch (726) to select the inter-frame prediction result used by the residual calculator (723), selects the inter-frame prediction information, and controls the entropy encoder (725) to include the inter-frame prediction information in the code stream.
[0076] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-frame encoder (722) or the inter-frame encoder (730). The residual encoder (724) is configured to operate based on the residual data to generate transform coefficients by encoding the residual data. In one example, the residual encoder (724) is configured to transform the residual data in the frequency domain and generate transform coefficients. Next, the transform coefficients undergo quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra-frame encoder (722) and the inter-frame encoder (730). For example, the inter-frame encoder (730) can generate a decoded block based on the decoded residual data and the inter-frame prediction information, and the intra-frame encoder (722) can generate a decoded block based on the decoded residual data and the intra-frame prediction information. The decoded block is appropriately processed to generate a decoded image, and in some examples, the decoded image can be buffered in a memory circuit (not shown) and used as a reference image.
[0077] The entropy encoder (725) is configured to format the bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other appropriate information in the bitstream. It should be noted that according to the disclosed subject matter, there is no residual information when encoding a block in the merge sub-mode of the inter-frame mode or the bi-directional prediction mode.
[0078] FIG. 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded image that is part of an encoded video sequence and decode the encoded image to generate a reconstructed image. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.
[0079] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) that are coupled together as shown in FIG. 8.
[0080] The entropy decoder (871) can be configured to reconstruct from the encoded image specific symbols that represent syntax elements that make up the encoded image. Such symbols include, for example, a mode for encoding a block (e.g., intra-frame, inter-frame, bi-directional prediction, two merge sub-modes of the latter, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information, etc.) that can identify specific samples or metadata used for prediction by the intra-frame decoder (872) or the inter-frame decoder (880), and residual information in the form of, for example, quantized transform coefficients. In one example, when the prediction mode is an inter-frame prediction mode or a bi-directional prediction mode, the inter-frame prediction information is provided to the inter-frame decoder (880). And when the prediction type is an intra-frame prediction type, the intra-frame prediction information is provided to the intra-frame decoder (872). The residual information can undergo inverse quantization and be provided to the residual decoder (873).
[0081] The inter-frame decoder (880) is configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.
[0082] The intra-frame decoder (872) is configured to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information.
[0083] The residual decoder (873) is configured to perform inverse quantization to extract the inverse quantized transform coefficients, process the inverse quantized transform coefficients, and convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantization parameter (QP)), and that information may be provided by the entropy decoder (871) (since this is only low-volume control information, the data path is not shown).
[0084] The reconstruction module (874) is configured to combine, in the spatial domain, the residual as the output from the residual decoder (873) and the prediction result (optionally, as the output from the inter-frame prediction module or the intra-frame prediction module) to form a reconstructed block, and the reconstructed block can be part of the reconstructed image, and then the reconstructed image can be part of the reconstructed video. Note that it can perform other appropriate operations such as deblocking operations to improve visual quality.
[0085] Note that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.
[0086] Aspects of the present disclosure provide techniques for in-frame image block compensation.
[0087] Block-based compensation from different images is called motion compensation. Similarly, block compensation can also be performed from previously reconstructed regions within the same image. Block-based compensation from reconstructed regions within the same image is called in-frame image block compensation or in-frame block copy. A shift vector indicating the offset between a current block and a reference block within the same image is called a block vector (or simply BV). Unlike motion vectors in motion compensation, which can be any value (positive or negative, in either the x or y direction), block vectors have several constraints, which ensure that the reference block is available and has already been reconstructed. Also, in some examples, some reference regions that are tile boundaries or wavefront ladder-shaped boundaries are excluded in consideration of parallel processing.
[0088] Encoding of the block vector may be either explicit or implicit. In the explicit mode, the difference between the block vector and its predictor is signaled, and in the implicit mode, the block vector is restored from a predictor (referred to as a block vector predictor) in a manner similar to the motion vector in the merge mode. In some implementations, the resolution of the block vector is limited to integer positions, and in other systems, the block vector is allowed to point to fractional positions.
[0089] In some examples, the use of in-frame block copy at the block level can be signaled using a reference index approach. The current image during decoding is treated as a reference image. In one example, such a reference image is placed at the last position in a list of reference images. This special reference image is also managed together with other temporal reference images in a buffer such as a decoded picture buffer (DPB).
[0090] There are also several variations of in-frame block copy, for example, inverted in-frame block copy (the reference block is inverted horizontally or vertically before being used to predict the current block), or line-based in-frame block copy (each compensation unit within the M×N coded block is a line of M×1 or 1×N).
[0091] FIG. 9 is a diagram showing an example of in-frame block copy according to an embodiment of the present disclosure. The current image (900) is being decoded. The current image (900) includes a reconstructed region (910) (gray region) and a region to be decoded (920) (white region). The current block (930) is being reconstructed by the decoder. The current block (930) can be reconstructed from a reference block (940) in the reconstructed region (910). The position offset between the reference block (940) and the current block (930) is called a block vector (950) (or BV(950)).
[0092] Traditionally, the motion vector resolution has been a fixed value such as 1 / 4 pixel (pixel) accuracy or 1 / 8 pixel accuracy in the main profiles of, for example, H.264 / AVC and HEVC. In HEVC SCC, the resolution of the motion vector can be selected to be 1 integer pixel or 1 / 4 pixel. The switching is done for each slice. In other words, the resolution of all motion vectors within a slice is the same.
[0093] In some subsequent developments, the resolution of the motion vector can be either 1 / 4 pixel, 1 integer pixel, or 4 integer pixels. In the example of 4 integer pixels, each unit represents 4 integer pixels. Therefore, there is a distance of 4 integer pixels between the symbols '0' and '1'. Furthermore, the adaptability occurs at the block level, that is, the motion vector can select a different resolution for each block.
[0094] Each aspect of the present disclosure provides an adaptive method for motion vector resolution and block vector resolution in image and video compression.
[0095] Block vectors are usually signaled with the signal at integer resolution. Therefore, when frame - in - block copy is extended to the full - frame range, all reconstructed regions of the current decoded image can be used as references. However, in the case of long - distance references, the cost of encoding the block - vector difference becomes high. Adaptive block - vector difference resolution can be used to improve the encoding of block - vector differences.
[0096] In some examples, multiple block - vector resolutions are used for the encoding of block vectors, and the block - vector resolution can be switched at the block level. Further, when more than two possible resolutions are used, a signaling flag that can contain more than 1 - bin (binary bit) is used to indicate which of the multiple block - vector resolutions is used for the block - vector difference of the current block. Possible resolutions include, but are not limited to, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel, 1 integer pixel, 2 integer pixels, 4 integer pixels, 8 integer pixels, etc. X integer pixels means that each minimum unit of the symbol represents X integer positions. For example, 2 integer pixels means that each minimum unit of the symbol represents 2 integer positions, 4 integer pixels means that each minimum unit of the symbol represents 4 integer positions, and 8 integer pixels means that each minimum unit of the symbol represents 8 integer positions.
[0097] In one embodiment, a set of resolutions can be used, and a flag can be signaled to indicate which resolution in the set is to be used. In one example, the block vector difference resolution set includes two resolutions such as 1 integer pixel and 2 integer pixels. Next, a 1-bin (1 binary) flag is signaled to select one resolution from the two possible resolutions. In another example, the block vector resolution set includes 1 integer pixel and 4 integer pixels. Next, a 1-bin flag is signaled to select one from the two possible resolutions.
[0098] In another example, the block vector difference resolution set includes 1 integer pixel, 2 integer pixels, and 4 integer pixels. A 1-bin flag is signaled to indicate whether the 1 integer pixel resolution is to be used. If the 1 integer pixel is not used, another 1-bin flag is signaled to indicate whether the 4 integer pixel resolution is selected. Different binarization embodiments can be derived in a similar way as showing resolution selection from three possible resolutions using 1-bin or 2-bin.
[0099] In another example, the block vector difference resolution set includes 1 integer pixel, 4 integer pixels, and 8 integer pixels. A 1-bin flag is signaled to indicate whether the 1 integer pixel resolution is to be used. If the 1 integer pixel is not used, another 1-bin flag is signaled to indicate whether the 4 integer pixel resolution is selected. Different binarization embodiments can be derived in a similar way as showing resolution selection from three possible resolutions using 1-bin or 2-bin.
[0100] In another example, the x and y components of the block vector difference can use different resolutions. For example, the x component of the block vector difference is at 1 integer pixel resolution, and the y component is at 4 integer pixel resolution. The resolution can be selected for each component by explicit signaling (one flag per component) or by inference. For example, the magnitude of the block vector prediction value of the current block can be used to estimate the resolution. In one example, if the magnitude of a component of the block vector prediction value is greater than a threshold, the block vector difference resolution for that component uses a larger step resolution, e.g., 4 integer pixel resolution, and otherwise, that component uses a smaller step resolution, e.g., 1 integer pixel resolution.
[0101] In another example, the x and y components of the block vector difference can use different resolutions. By default, the two components use a fixed resolution, such as 1 integer pixel resolution. Conditions for the components are set respectively to determine the resolution. For example, when conditions for some components are met, each component can be switched to a different resolution by inference. The components of the predicted value are also quantized to the corresponding resolution or maintained without changing the original resolution. For example, a condition for one component is related to the magnitude of the block vector prediction value of the current block. By evaluating the magnitude of the block vector prediction value of the current block in each component, the resolution of that component can be estimated. For example, if the magnitude of the component of the block vector prediction value is greater than a threshold, the block vector difference resolution of that component uses a larger step resolution, such as 4 integer pixel resolution, and otherwise, that component uses the default step resolution, such as 1 integer pixel resolution. In one specific example, the default resolution is 1 integer pixel. The block vector prediction value is set to (-21, -3), and the threshold for using 4 integer pixel resolution for each component is set to 20. The decoded block vector difference symbol is (2, 2). According to the rule, the x component uses 4 integer pixel resolution and the y component uses 1 integer pixel resolution. Then, the predicted value becomes (-20, -3) (-21 is rounded to -20), and the block vector difference becomes (8, 2) (due to 4 integer pixels for the x component and 1 integer pixel for the y component). In this example, the finally decoded block vector is (-12, -1).
[0102] Using a similar derivation as in the above example, a similar example can be derived. In one embodiment, a set of possible resolutions is formed and one or more signaling flags are used to indicate the selection of resolution for each block. In the above example, the binary bins of the signaling flags may be context encoded or bypass encoded. If context encoding is used, the block vector resolution of spatially adjacent blocks of the current block can be used for context modeling. Alternatively, the resolution of the last encoded block vector can be used.
[0103] According to another aspect of the present disclosure, the block vector prediction value is typically derived for the current block from one previously encoded block vector of an adjacent block. When used as a prediction value, the previously decoded block vector can have a resolution different from that of the current block. The present disclosure provides a way to solve this problem.
[0104] In one embodiment, the original resolution of the block vector prediction value remains unchanged. The decoded block vector difference is added to the block vector prediction value at its target resolution to obtain the finally decoded block vector. Thus, the finally decoded block vector will be the higher of the two resolutions (the original resolution of the block vector prediction value and the target resolution of the decoded block vector difference). For example, the block vector prediction value is a vector (-11, 0) with a resolution of 1 integer pixel. The decoded block vector difference is a vector (-4, 0) with a resolution of 4 integer pixels. The decoded block vector will be a vector (-15, 0) with a resolution of 1 integer pixel.
[0105] In another embodiment, the original resolution of the block vector prediction value is rounded to the target resolution of the current block. The decoded block vector difference is added to the block vector prediction value at that target resolution. Thus, the finally decoded block vector will be the lower of the two resolutions. For example, the block vector prediction value is a vector (-11, 0) with a 1 integer-pixel resolution. The decoded block vector difference is a vector (-4, 0) with a 4 integer-pixel resolution. The block vector prediction value is first rounded to a vector (-12, 0) with a 4 integer-pixel resolution before being added to the decoded block vector difference. The decoded block vector becomes a vector (-16, 0) with a 4 integer-pixel resolution.
[0106] Various rounding techniques can be used. In one example, the vector can be rounded to the nearest integer (corresponding to an integer symbol value) according to the difference value. For example, (-11, 0) with a 1 integer-pixel resolution is rounded to (-12, 0) with a 4 integer-pixel resolution, and (-13, 0) with a 1 integer-pixel resolution is also rounded to (-12, 0) with a 4 integer-pixel resolution.
[0107] In another example, a ceiling operation is applied to the vector so that it is rounded to the nearest integer (corresponding to an integer symbol value) that is not less than the current value. For example, (-11, 0) with a 1 integer-pixel resolution is rounded to (-8, 0) with a 4 integer-pixel resolution, and (-13, 0) with a 1 integer-pixel resolution is also rounded to (-12, 0) with a 4 integer-pixel resolution.
[0108] In another example, a flooring operation is applied to the vector such that it is rounded to the nearest integer that is not greater than the current value (corresponding to an integer symbol value). For example, (-11, 0) with a 1 integer pixel resolution is rounded to (-12, 0) with a 4 integer pixel resolution, and (-13, 0) with a 1 integer pixel resolution is also rounded to (-16, 0) with a 4 integer pixel resolution.
[0109] In another example, the vector is rounded towards zero. For example, (-11, 0) with a 1 integer pixel resolution is rounded to (-8, 0) with a 4 integer pixel resolution, and (11, -5) with a 1 integer pixel resolution is rounded to (8, -4) with a 4 integer pixel resolution.
[0110] Aspects of the present disclosure further provide techniques for handling boundary constraints having multiple resolutions. In one example, the block vector is constrained to point to a reference region where the use of in-frame image block compensation is permitted. The constraints can include the permitted reference region's picture / slice / tile and wavefront boundaries. When multiple resolutions are used, especially when a resolution greater than 1 integer pixel resolution is used, the boundary constraints need to be processed correctly.
[0111] In one embodiment, when the block vector (block vector prediction value + difference value) points to a location outside the reference region boundary, a clipping operation is performed to change one or two components of the block vector back to the edge of the boundary, thereby making the modified block vector valid. The clipping operation can be performed without considering the resolution of the decoded block vector. Thus, after the modification, the resolution of the block vector may be different from the resolution before the modification.
[0112] In another embodiment, when the block vector (block vector prediction value + difference value) points to a location outside the reference region boundary, a clipping operation is performed to change one or two components of the block vector to return to the edge of the boundary, whereby the changed block vector becomes valid. The clipping operation can be performed in consideration of the resolution of the decoded block vector. Therefore, after the change, the resolution of the block vector remains the same as before the change.
[0113] In another embodiment, when the block vector (block vector prediction value + difference value) points to a location outside the reference region boundary, the pixels outside the boundary can be displayed by horizontally or vertically expanding the pixels at the boundary.
[0114] In another embodiment, when the block vector (block vector prediction value + difference value) points to a location outside the reference region boundary, the resolution of the block vector difference is changed to the highest possible accuracy. For example, the block vector prediction value is the vector (0,0), and the block vector difference is the symbol (-5,0) with a 4-integer pixel resolution. The highest possible accuracy is 1 integer pixel. When the decoded block vector has a 4-integer pixel resolution, the block vector becomes (-20,0). If (-20,0) points to a location outside the left image boundary, the decoded block vector difference is changed to a 1-integer pixel resolution, whereby the decoded block vector becomes (-5,0), and in this case, it is a valid vector. This change can help move the block vector within the boundary.
[0115] According to another aspect of the present disclosure, the technique for processing multiple resolutions in block vectors for in-frame image block compensation can be similarly applied to motion vectors for inter-frame image block compensation. In some embodiments, multiple motion vector resolutions are used for encoding the motion vectors, and the motion vector resolution can be switched at the block level. Further, when more than two possible resolutions are used, a signaling flag that can exceed 1-bin is used to indicate which of the multiple motion vector resolutions is used for the motion vector difference of the current block. Possible resolutions include, but are not limited to, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel, 1 integer pixel, 2 integer pixels, 4 integer pixels, 8 integer pixels, etc. X integer pixels means that each minimum unit of the symbol represents X integer positions. For example, 2 integer pixels means that each minimum unit of the symbol represents 2 integer positions, 4 integer pixels means that each minimum unit of the symbol represents 4 integer positions, and 8 integer pixels means that each minimum unit of the symbol represents 8 integer positions.
[0116] In one example, the x and y components of the motion vector difference can use different resolutions. For example, the x component of the motion vector difference is at 1 integer pixel resolution, and the y component is at 4 integer pixel resolution. The resolution can be selected for each component by explicit signaling (one flag for each component) or inference. For example, the magnitude of the motion vector prediction value of the current block can be used to estimate the resolution. In one example, when the magnitude of the component of the motion vector prediction value is greater than a threshold, the motion vector difference resolution of that component uses a larger step resolution, such as 4 integer pixel resolution, and otherwise, that component uses a smaller step resolution, such as 1 integer pixel resolution.
[0117] In another example, the x and y components of the motion vector difference can use different resolutions. By default, the two components use a fixed resolution, such as 1 / 4 pixel resolution. Component conditions for each component are set to determine the resolution. For example, if the conditions for some components are met, each component can be switched to a different resolution by inference. The components of the predicted value are also quantized to the corresponding resolution or maintained without changing the original resolution. For example, the condition for one component is related to the magnitude of the motion vector prediction value of the current block. By evaluating the magnitude of the motion vector prediction value of the current block in each component, the resolution of the component can be estimated. For example, if the magnitude of the component of the motion vector prediction value is greater than a threshold, the motion vector difference resolution of the component uses a larger step resolution, such as 4 integer pixel resolution, and otherwise, the component uses the default step resolution, such as 1 / 4 pixel resolution. In one specific example, the default resolution is 1 / 4 pixel. The motion vector prediction value is set to (-20.75, -3.75), and the threshold for using 1 integer pixel resolution for each component is set to 5. The decoded motion vector difference symbol is (2, 2). According to the rule, the x component uses 1 integer pixel resolution and the y component uses 1 / 4 pixel resolution. Then, the motion vector prediction value becomes (-20, -3.75) (for example, -20.75 is rounded to -20 in the zero direction), and the block vector difference becomes (2, 0.5) (due to 1 integer pixel for the x component and 1 / 4 pixel for the y component). In this example, the finally decoded block vector is (-18, -3.25).
[0118] FIG. 10 shows a flowchart outlining a process (1000) according to an embodiment of the present disclosure. The process (1000) can be used for reconstructing blocks encoded in an intra mode and can generate a prediction block for a block being reconstructed. In various embodiments, the process (1000) is executed by processing circuits in terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of the video encoder (403), a processing circuit that executes the functions of the video decoder (410), a processing circuit that executes the functions of the video decoder (510), a processing circuit that executes the functions of the intra prediction module (552), a processing circuit that executes the functions of the video encoder (603), a processing circuit that executes the functions of the predictor (635), a processing circuit that executes the functions of the intra encoder (722), a processing circuit that executes the functions of the intra decoder (872), etc. In some embodiments, the process (1000) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1000). This process starts at (S1001) and proceeds to (S1010).
[0119] (S1010), the prediction information of the current block is decoded from the encoded video bitstream. The prediction information indicates an intra block copy mode.
[0120] (S1020), a resolution of the block vector difference of the current block is selected from a set consisting of a plurality of candidate resolutions. In one example, a resolution flag is received and the resolution is selected based on the flag. In another example, the resolution is determined based on inference.
[0121] (S1030), based on the selected resolution of the block vector difference and the predicted value of the block vector of the current block, the block vector of the current block is determined.
[0122] In (S1040), samples of the current block are reconstructed based on the determined block vector. Then, the process proceeds to (S1099) and ends.
[0123] The above technology is realized as computer software using computer-readable instructions and can also be physically stored in one or more computer-readable media. For example, FIG. 11 shows a computer system (1100) suitable for implementing a specific embodiment of the disclosed subject matter.
[0124] The computer software may be encoded using any suitable machine code or computer language, and code containing instructions may be created by mechanisms such as assembly, compilation, and linking. These instructions may be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or may be executed by interpretation, microcode, etc.
[0125] The instructions may be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, object network devices (internet of things devices), etc.
[0126] The components of the computer system (1100) shown in FIG. 11 are essentially exemplary, and are not intended to suggest any limitations regarding the scope or functionality of the use of the computer software for implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiment of the computer system (1100).
[0127] The computer system (1100) can include several human interface input devices. Such human interface input devices can respond to input by one or more users through tactile input (e.g., keystrokes, swipes, movements of a data glove, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures, etc.), and olfactory input (not shown). The human interface device can also be used to capture certain media that is not necessarily directly related to conscious human input, such as, for example, audio (e.g., voice, music, ambient sound, etc.), images (e.g., scanned images, photographic images obtained from a still image camera, etc.), video (e.g., 2D video, 3D video including stereoscopic video, etc.).
[0128] The human interface input device can include one or more of a keyboard (1101), a mouse (1102), a trackpad (1103), a touch screen (1110), a data glove (not shown), a joystick (1105), a microphone (1106), a scanner (1107), and a camera (1108) (only one of which is shown).
[0129] The computer system (1100) can also include several human interface output devices. Such human interface output devices can stimulate the senses of one or more users, for example, by tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (1110), a data glove (not shown) or a joystick (1105), which may be a tactile feedback device that does not act as an input device), audio output devices (e.g., speakers (1109), headphones (not shown)), visual output devices (e.g., a screen (1110) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each of which may or may not have a touch screen input function, and each of which may or may not have a tactile feedback function, and some of these can output two-dimensional visual output or three-dimensional or more visual output, for example, by stereographic output, virtual reality glasses (not shown), a holographic display and a smoke tank (not shown), and a printer (not shown)).
[0130] The computer system (1100) can include human-accessible storage devices and associated media such as optical media or similar media (1121) including a CD / DVD ROM / RW (1120) having a CD / DVD, a thumb drive (1122), a removable hard drive or a solid state drive (1123), legacy magnetic media such as tapes and floppy disks (not shown), special ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0131] One skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.
[0132] The computer system (1100) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be a local network, a wide area network, a metropolitan area network, a vehicle network and an industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include LANs such as Ethernet (registered trademark), Wi-Fi, cellular networks (such as GSM (registered trademark), 3G, 4G, 5G, LTE, etc.), TV cables or wireless wide area digital networks (including cable TV, satellite TV, terrestrial broadcast TV), vehicle and industrial networks (including CANBus), etc. Some networks generally require an external network interface adapter connected to some general-purpose data ports or peripheral buses (1149) (for example, the USB port of the computer system (1100)), and other systems are usually integrated into the core of the computer system system (1100) by connecting to the system bus as described below (for example, an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1100) can communicate with other entities. Such communication can be only unidirectional reception (for example, broadcast TV), only unidirectional transmission (for example, from Canbus to a specific Canbus device), or bidirectional, for example, communication to other computer systems using a local or wide area digital network. As described above, specific protocols and protocol stacks can be used for each of those networks and network interfaces.
[0133] The above human interface device, human-accessible memory device, and network interface can be connected to the core (1140) of the computer system (1100).
[0134] The core (1140) can include one or more central processing units (CPUs) (1141), a graphics processing unit (GPU) (1142), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1143), a hardware accelerator (1144) for specific tasks, and the like. These devices may be connected via a system bus (1148) together with a read-only memory (ROM) (1145), a random access memory (1146), an internal mass storage (1147) such as an internal non-user-accessible hard disk drive, SSD, etc. In some computer systems, in order to enable expansion by additional CPUs, GPUs, etc., one or more physical plugs can be accessed in the form of the system bus (1148). Peripheral devices may be directly connected to the system bus of the core (1148) or may be connected via a peripheral bus (1149). The architecture of the peripheral bus includes an external controller interface (PCI), a universal serial bus (USB), and the like.
[0135] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) can execute some instructions, and these instructions can be combined to form the above-mentioned computer code. The computer code can be stored in the ROM (1145) or RAM (1146). Also, temporary data can be stored in the RAM (1146), while permanent data can be stored, for example, in the internal mass storage (1147). By using high-speed storage that can be closely associated with one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc., high-speed storage and retrieval for any memory device become possible.
[0136] A computer-readable medium can have computer code for performing various computer-executed operations. The medium and the computer code may be specially designed and configured for the purposes of this disclosure, or they may be media and code known and available to those skilled in the computer software art.
[0137] By way of example and not limitation, a computer system having an architecture (1100), particularly a core (1140), can provide functionality as a processor (including a CPU, GPU, FPGA, accelerator, etc.) that executes software embodied on one or more tangible, computer-readable media. Such computer-readable media can be media related to the large-capacity memory accessible by the above user, and can also be a specific storage having a non-volatile core (1140), such as the on-core large-capacity storage (1147) or ROM (1145) inside the core. The software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (1140). The computer-readable media can include one or more memory devices or chips according to specific needs. This software includes defining a data structure stored in the RAM (1146) in the core (1140), specifically the processor (including a CPU, GPU, FPGA, etc.) therein, and modifying such a data structure according to the process defined by the software, and can execute a specific process or a specific part of a specific process described herein. Additionally or alternatively, the computer system can provide functionality as a result embodied in logic hardware or other circuitry (e.g., accelerator (1144)), which can operate instead of or together with the software to execute a specific process or a specific part of a specific process described herein. Where appropriate, references to software can include logic, and vice versa. Where appropriate, references to computer-readable media can include circuitry (such as an integrated circuit (IC)) that stores the executable software, circuitry that embodies the execution logic, or both. The present disclosure encompasses any suitable combination of hardware and software. Appendix A: Acronyms JEM: joint exploration model, collaborative development model VVC: Versatile Video Coding, General-Purpose Video Encoding BMS: Benchmark Set, Benchmark Set MV: Motion Vector, Motion Vector HEVC: High Efficiency Video Coding, High-Efficiency Video Encoding SEI: Supplementary Enhancement Information, Supplementary Enhancement Information VUI: Video Usability Information, Video User Usability Information GOPs: Groups of Pictures, Groups of Images TUs: Transform Units, Transform Units PUs: Prediction Units, Prediction Units CTUs: Coding Tree Units, Coding Tree Units CTBs: Coding Tree Blocks, Coding Tree Blocks PBs: Prediction Blocks, Prediction Blocks HRD: Hypothetical Reference Decoder, Hypothetical Reference Decoder SNR: Signal Noise Ratio, Signal-to-Noise Ratio CPUs: Central Processing Units, Central Processing Units GPUs: Graphics Processing Units, Graphics Processing Units CRT: Cathode Ray Tube, Cathode Ray Tube LCD: Liquid-Crystal Display, Liquid Crystal Display OLED: Organic Light-Emitting Diode, Organic Light-Emitting Diode CD: Compact Disc, Compact Disc DVD: Digital Video Disc, Digital Video Disc ROM: Read-Only Memory, Read-Only Memory RAM: Random Access Memory, Random Access Memory ASIC: Application-Specific Integrated Circuit, Application-Specific Integrated Circuit PLD: Programmable Logic Device, Programmable Logic Device LAN: Local Area Network, Local Area Network GSM: Global System for Mobile communications, Global System for Mobile Communications LTE: Long-Term Evolution, Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus, Universal Serial Bus PCI: Peripheral Component Interconnect, Peripheral Component Interconnect FPGA: Field Programmable Gate Areas, Field Programmable Gate Array SSD: solid-state drive, solid-state drive IC: Integrated Circuit, Integrated Circuit CU: Coding Unit, Coding Unit Although the present disclosure has described several exemplary embodiments, there are changes, arrangements, and various equivalent substitutions within the scope of the present disclosure. Therefore, those skilled in the art should understand that although not explicitly shown or described in this specification, various systems and methods that embody the principles of the present disclosure and are within the spirit and scope of the present disclosure can be designed.
Description of Symbols
[0138] 101 Sample 300 Communication System 310, 320 Terminal Devices 350 Network
Claims
1. A method for decoding video by a decoder, comprising: decoding prediction information of a current block from an encoded video bitstream, wherein the prediction information indicates whether to apply an in-frame block copy mode; selecting a resolution of a block vector difference of the current block from a set of multiple candidate resolutions, determining whether a first resolution included in the set is used based on a first syntax element, and when the first syntax element indicates that the first resolution is not used, selecting one of two resolutions included in the set as a second resolution based on a second syntax element, the first resolution and the two resolutions being three different resolutions, and at least one of the two resolutions being an integer pixel resolution; determining a block vector of the current block based on the selected resolution of the block vector difference and a block vector prediction value of the current block; reconstructing at least one sample of the current block based on the block vector. A method characterized by the above.
2. The method according to claim 1, wherein the first resolution is a 1 / 4 pixel resolution, and the second resolution is any one of 1 / 16, 1 / 8, 1 / 2, 1, 2, 4, 8 pixel resolutions.
3. When the block vector prediction value uses a resolution different from the selected resolution, further comprising adding the block vector difference to the block vector prediction value instead of rounding the block vector prediction value to calculate the block vector. The method according to claim 1 or 2, characterized by the above.
4. When the predicted block vector value has a resolution different from the selected resolution, rounding the predicted block vector value of the current block to the selected resolution; further including adding the block vector difference to the rounded predicted block vector value to calculate the block vector. The method according to any one of claims 1 to 3, characterized in that.
5. A program for causing a computer that decodes video to execute the method according to any one of claims 1 to 4.
6. An apparatus for decoding video, including a processing circuit, the processing circuit decoding prediction information of a current block from an encoded video bit stream, the prediction information indicating whether to apply an intra-block copy mode; selecting a resolution of a block vector difference of the current block from a set of a plurality of candidate resolutions, determining whether a first resolution included in the set is used based on a first syntax element, and when the first syntax element indicates that the first resolution is not used, selecting one of two resolutions included in the set as a second resolution based on a second syntax element, the first resolution and the two resolutions being three different resolutions, and at least one of the two resolutions being an integer pixel resolution; determining a block vector of the current block based on the resolution of the block vector difference selected in the selecting step and the predicted block vector value of the current block; and reconstructing at least one sample of the current block based on the block vector. The apparatus is characterized by that.
7. A method implemented by an encoder, Generating an encoded bitstream including a block of an image and prediction information of the block of the image and transmitting the encoded bitstream to a decoder including, wherein the prediction information indicates whether to apply an in-frame block copy mode at least one sample of the block of the image can be reconstructed based on a block vector of the block of the image the block vector of the block of the image is determined based on a resolution of a block vector difference of the block of the image and a predicted value of the block vector of the block of the image the resolution of the block vector difference of the block of the image is selected from a set of a plurality of candidate resolutions the encoded bitstream includes a first syntax element indicating whether a first resolution included in the set is used, and a second syntax element indicating that one of two resolutions included in the set is selected as a second resolution when the first syntax element indicates that the first resolution is not used wherein the first resolution and the two resolutions are three different resolutions, and at least one of the two resolutions is an integer pixel resolution Claim 8 A method performed by an encoder when an in-frame block copy mode is applied in video coding, determining a block vector difference of a current block of the video based on a block vector of the current block of the video and a predicted value of the block vector of the current block, and performing encoding of the video The resolution of the block vector difference is selected from a set of a plurality of candidate resolutions, and whether the first resolution included in the set is used is signaled to a decoder by a first syntax element. When the first syntax element indicates that the first resolution is not used, it is signaled to the decoder by a second syntax element that one of two resolutions included in the set is selected as a second resolution. The first resolution and the two resolutions are three different resolutions, and at least one of the two resolutions is an integer pixel resolution. Method.
9. A storage method, encoding a block of an image and prediction information of the block of the image to generate a bitstream; storing the bitstream in a storage medium and the prediction information indicates whether to apply an intra-frame block copy mode. From a set of a plurality of candidate resolutions, the resolution of the block vector difference of the block of the image is selected. Based on a first syntax element, it is determined whether the first resolution included in the set is used. When the first syntax element indicates that the first resolution is not used, based on a second syntax element, one of two resolutions included in the set is selected as a second resolution. The first resolution and the two resolutions are three different resolutions, and at least one of the two resolutions is an integer pixel resolution. Based on the resolution of the block vector difference selected in the selecting step and the block vector prediction value of the block of the image, the block vector of the block of the image is determined. At least one sample of the block of the image is reconstructed based on the block vector. A method characterized by this.
Citation Information
Patent Citations
Motion image coding apparatus and motion image decoding apparatus
JP2011229190A
Representing Motion Vectors in an Encoded Bitstream
US20150195527A1
Search region determination for inter coding within a particular picture of video data
US20160337661A1
Storage and signaling resolutions of motion vectors
US20160337662A1
Cited By
Method and apparatus for video coding / decoding
JP2025131763A