Encoder and decoder for drift-free padding and hashing of independently coded regions, encoding method and decoding method
By independently encoding boundary filtering of adjacent tiles during video encoding and decoding, and using reference blocks for filtering correction, the problem of insufficient parallel processing capability in the HEVC standard is solved, thereby improving the efficiency and quality of video encoding and decoding.
Patent Information
- Application Number
- CN202080050882.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-22
- Filing Date
- 2020-05-20
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-05-20
AI Technical Summary
Existing video encoders and decoders are insufficient in parallel processing capabilities, especially in the HEVC standard, where the image coding tree unit processing method fails to effectively support parallel processing, resulting in low efficiency.
By independently encoding the boundaries between adjacent tiles and performing filtering, and using reference blocks for decoding and encoding, the filter or filter kernel corrects the filtering influence range based on the boundary distance, thus achieving independent encoding and decoding of tiles.
It improves the parallel processing capabilities of video encoding and decoding, thereby enhancing processing efficiency and image reconstruction quality.
Smart Images

Figure CN114128264B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to video encoding and video decoding, and more particularly to encoders and decoders, encoding methods and decoding methods for drift-free padding and / or hashing of independent encoding regions. Background Technology
[0002] H.265 / HEVC (HEVC = High Efficiency Video Coding) is a video codec that provides tools for enhancing or even enabling parallel processing at the encoder and / or decoder. For example, HEVC supports subdividing an image into an array of independently coded tiles. Another concept supported by HEVC, related to WPP, allows for parallel processing of CTU rows or lines of images from left to right, such as in stripes, provided that some minimum CTU offset (CTU = Code Tree Unit) is observed when processing consecutive CTU rows. However, it would be advantageous to have an existing video codec that more efficiently supports the parallel processing capabilities of the video encoder and / or decoder.
[0003] The following describes an introduction to VCL partitioning (VCL = Video Coding Layer) according to existing technology.
[0004] Typically, in video coding, the encoding process for image samples requires small partitions, where samples are divided into rectangular regions for joint processing, such as prediction or transform coding. Therefore, images are divided into blocks of a specific size, which remains constant throughout the encoding of the video sequence. In the H.264 / AVC standard, fixed-size blocks of 16x16 samples are used, known as macroblocks (AVC = Advanced Video Coding).
[0005] In the existing HEVC standard (see [1]), there are coding tree blocks (CTBs) or coding tree units (CTUs) with a maximum size of 64x64 samples. In further descriptions of HEVC, the more common term CTU is used for such blocks.
[0006] CTUs are processed in raster scan order, starting from the CTU in the upper left corner and processing the CTUs in the image line by line until the CTU in the lower right corner.
[0007] Encoded CTU data is organized into containers called slices. Originally, in previous video coding standards, a slice referred to a segment comprising one or more consecutive CTUs of an image. Slices are used to divide encoded data into segments. Alternatively, a complete image can also be defined as a large segment; therefore, historically, the term slice remains applicable. In addition to the encoded image sample, a slice includes supplementary information related to the encoding process of the slice itself, which is placed in a so-called slice header.
[0008] According to existing technology, the VCL (Video Coding Layer) also includes techniques for segmentation and spatial partitioning. For example, such partitioning can be applied to video coding for various reasons, including processing load balancing in parallelization, CTU size matching in network transmission, error mitigation, etc. Summary of the Invention
[0009] The purpose of this invention is to provide improved concepts for video encoding and video decoding.
[0010] The objective of this invention is addressed by the subject matter of the independent claims.
[0011] A video decoder is provided for decoding an encoded video signal comprising encoded image data to reconstruct a plurality of images of a video, according to an embodiment. The video decoder includes an input interface configured to receive the encoded video signal, and a data decoder configured to reconstruct the plurality of images of the video by decoding the encoded image data. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of a first image, the data decoder is configured to filter across the boundary between the first tile and the second tile to obtain a first filtered tile, wherein the first tile and the second tile have been encoded independently relative to each other. The data decoder is configured to decode a current tile in a plurality of tiles of a second image in a plurality of images based on a reference block of the first filtered tile of the first image, wherein the reference block includes a first set of samples of the first filtered tile, and wherein the reference block does not include a second set of samples of the first filtered tile, wherein none of the first set of samples is affected by the filtering across the boundary between the first tile and the second tile, and wherein one or more of the second set of samples has been affected by the filtering across the boundary between the first tile and the second tile.
[0012] Furthermore, a video encoder is provided for encoding a plurality of images of a video by generating an encoded video signal according to an embodiment. Each of the plurality of images includes raw image data. The video encoder includes a data encoder configured to generate an encoded video signal including encoded image data, wherein the data encoder is configured to encode the plurality of images of the video into encoded image data, and an output interface configured to output the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of a first image of the plurality of images, a boundary exists between the first tile and the second tile. The data encoder is configured to encode the first tile and the second tile independently of each other. Furthermore, the data encoder is configured to encode the current tile in a plurality of tiles of the second image in the plurality of images based on a reference block of the first tile of the first image, wherein the filter defines filtering across the boundary between the first tile and the second tile, wherein the reference block includes a first set of samples of the first tile, and wherein the reference block does not include a second set of samples of the first tile, wherein none of the first set of samples will be affected by the filtering using the filter, and wherein one or more of the second set of samples will be affected by the filtering using the filter.
[0013] Furthermore, a method is provided according to an embodiment for decoding an encoded video signal including encoded image data to reconstruct a video from multiple images. The decoding method includes:
[0014] - Receive encoded video signals. And:
[0015] - Reconstruct multiple images from a video by decoding encoded image data.
[0016] Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of a first image, the method includes filtering across the boundary between the first tile and the second tile to obtain a first filtered tile, wherein the first tile and the second tile have been independently encoded. The method includes decoding a current tile in the plurality of tiles of a second image based on a reference block of the first filtered tile of the first image, wherein the reference block includes a first set of samples from the first filtered tile, and wherein the reference block does not include a second set of samples from the first filtered tile, wherein none of the first set of samples is affected by the filtering across the boundary between the first tile and the second tile, and wherein one or more of the second set of samples has been affected by the filtering across the boundary between the first tile and the second tile.
[0017] Furthermore, a method is provided according to embodiments for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes original image data. The method includes:
[0018] - Generate a encoded video signal comprising encoded image data, wherein generating the encoded video signal includes encoding multiple images of the video into encoded image data, and
[0019] - Output the encoded image data for each of the multiple images.
[0020] Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. For a first image of the plurality of images, a first tile and a second tile are two adjacent tiles among the plurality of tiles, a boundary exists between the first tile and the second tile. The method includes encoding the first tile and the second tile independently of each other. Furthermore, the method includes encoding a current tile of a plurality of tiles of a second image of the plurality of images according to a reference block of the first tile of the first image, wherein a filter defines filtering across the boundary between the first tile and the second tile, wherein the reference block includes a first set of samples from the first tile, and wherein the reference block does not include a second set of samples from the first tile, wherein none of the first set of samples is affected by the filtering using the filter, and wherein one or more of the second set of samples is affected by the filtering using the filter.
[0021] In addition, a computer program is provided for implementing one of the methods described above when executed on a computer or signal processor.
[0022] Furthermore, an encoded video signal is provided according to an embodiment for encoding a plurality of images comprising a plurality of tiles. Each of the plurality of tiles comprises a plurality of samples, wherein the encoded video signal comprises encoded image data encoding the plurality of images. The encoded video signal comprises encoding the plurality of images. For a first tile and a second tile of two adjacent tiles in a plurality of tiles of a first image in the plurality of images, a boundary exists between the first tile and the second tile. The first tile and the second tile are encoded independently of each other within the encoded video signal. The current tile of the plurality of tiles of a second image in the plurality of images is encoded according to a reference block of the first tile of the first image, wherein a filter defines filtering across the boundary between the first tile and the second tile, wherein the reference block comprises a first set of samples of the first tile, and wherein the reference block does not include a second set of samples of the first tile, wherein any of the first set of samples is not affected by the filtering using the filter, and wherein one or more of the second set of samples will be affected by the filtering using the filter.
[0023] In one embodiment, the encoded video signal may include, for example, an indication of an encoding mode that indicates that samples of the reference block used to decode the current tile will not be affected by the filtering across the boundary between the first and second tiles.
[0024] Furthermore, a system including the aforementioned video encoder and video decoder is provided. The video encoder is configured to generate an encoded video signal. The video decoder is configured to decode the encoded video signal to reconstruct an image of the video.
[0025] Furthermore, a video decoder is provided for decoding an encoded video signal comprising encoded image data to reconstruct a plurality of images of a video, according to an embodiment. The video decoder includes an input interface configured to receive the encoded video signal, and a data decoder configured to reconstruct the plurality of images of the video by decoding the encoded image data. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of blocks, and each of the plurality of blocks includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of one of the plurality of images, a boundary exists between the first tile and the second tile. The first tile and the second tile have been encoded independently relative to each other. The data decoder is configured to filter the first tile using a filter or a filter kernel, wherein the data decoder is configured to correct the reach of the filter or the filter kernel based on the block of the first tile to be filtered by the filter or the filter kernel and the distance between the boundary between the first tile and the second tile, the block being one of the plurality of blocks of the first tile.
[0026] Furthermore, a video encoder according to an embodiment is provided for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder includes a data encoder configured to generate an encoded video signal including encoded image data, wherein the data encoder is configured to encode the plurality of images of the video into encoded image data, and an output interface configured to output the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of blocks, wherein each of the plurality of blocks includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of a first image, a boundary exists between the first tile and the second tile. The data encoder is configured to encode the first tile and the second tile independently of each other. Furthermore, the data encoder is configured to filter the first tile using a filter or filter kernel, wherein the data encoder is configured to adjust the influence range of the filter or filter kernel based on the block to be filtered by the filter or filter kernel and the distance between the boundary between the first tile and the second tile, the block being one of the plurality of blocks of the first tile.
[0027] Furthermore, a method is provided according to an embodiment for decoding a encoded video signal including encoded image data to reconstruct a video from multiple images. The decoding method includes:
[0028] - Receive encoded video signals, and
[0029] - Reconstruct multiple images from a video by decoding encoded image data.
[0030] Each of the multiple images includes multiple tiles, each of the multiple tiles includes multiple blocks, and each of the multiple blocks includes multiple samples.
[0031] For a first patch and a second patch, two adjacent patches in a plurality of tiles of an image from multiple images, a boundary exists between the first patch and the second patch. The first patch and the second patch have been independently encoded relative to each other. The method includes filtering the first patch using a filter or filter kernel, wherein the method includes adjusting the influence range of the filter or filter kernel based on the block of the first patch to be filtered by the filter or filter kernel and the distance between the boundary between the first patch and the second patch, the block being one of a plurality of blocks of the first patch.
[0032] Furthermore, a method for encoding a plurality of images of a video by generating an encoded video signal, wherein each of the plurality of images includes original image data, wherein the method includes:
[0033] - Generate an encoded video signal comprising encoded image data, wherein generating the encoded video signal includes encoding multiple images of the video into encoded image data. And:
[0034] - Outputs the encoded image data for each of multiple images.
[0035] Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of blocks, and each of the plurality of blocks includes a plurality of samples. For a first image, two adjacent tiles, a first tile and a second tile, exist as a boundary between the first tile and the second tile. The method includes encoding the first tile and the second tile independently of each other. Furthermore, the method includes filtering the first tile using a filter or filter kernel, wherein the method includes adjusting the influence range of the filter or filter kernel based on the block to be filtered by the filter or filter kernel and the distance between the boundary between the first tile and the second tile, the block being one of the plurality of blocks of the first tile.
[0036] In addition, a computer program is provided for implementing one of the above methods when executed on a computer or signal processor according to the embodiments.
[0037] Furthermore, an encoded video signal is provided that encodes multiple images comprising multiple tiles according to an embodiment. Each of the multiple tiles includes multiple samples. The encoded video signal includes encoded image data encoding the multiple images. Furthermore, the encoded video signal includes the encoding of multiple images. Each of the multiple images includes multiple tiles, wherein each of the multiple tiles includes multiple blocks, and each of the multiple blocks includes multiple samples. For a first tile and a second tile of two adjacent tiles in the multiple tiles of a first image, a boundary exists between the first tile and the second tile. The first tile and the second tile are encoded independently relative to each other within the encoded video signal. The encoded video signal depends on filtering the first tile using a filter or filter kernel, wherein during filtering, the influence range of the filter or filter kernel has been corrected according to the block to be filtered by the filter or filter kernel and the distance between the boundary between the first tile and the second tile, the block being one of the multiple blocks of the first tile.
[0038] Furthermore, a system according to an embodiment is provided, comprising the aforementioned video encoder and video decoder. The video encoder is configured to generate an encoded video signal. The video decoder is configured to decode the encoded video signal to reconstruct an image of the video.
[0039] Furthermore, a video encoder according to an embodiment is provided for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder includes a data encoder configured to generate an encoded video signal including the encoded image data, wherein the data encoder is configured to encode the plurality of images of the video into encoded image data, and an output interface configured to output the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. The data encoder is configured to determine an independently encoded group of tiles, the independently encoded group of tiles including three or more tiles from a plurality of tiles of a reference image of the plurality of images. Furthermore, the data encoder is configured to encode the plurality of images based on a reference tile located within the reference image. Furthermore, the data encoder is configured to select a position for the reference tile within the reference image such that the reference tile is not both partially located within three of the three or more tiles of the independently encoded group of tiles and partially located within another tile of the plurality of tiles of the reference image that does not belong to the independently encoded group of tiles.
[0040] Furthermore, a method is provided according to embodiments for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes original image data. The method includes:
[0041] - Generate an encoded video signal comprising encoded image data, wherein generating the encoded video signal includes encoding multiple images of the video into encoded image data. And:
[0042] - Outputs the encoded image data for each of multiple images.
[0043] Each of a plurality of images comprises a plurality of tiles, wherein each of the plurality of tiles comprises a plurality of samples. The method includes determining an independently encoded group of tiles, the independently encoded group of tiles comprising three or more tiles from a plurality of tiles in a reference image of the plurality of images. Furthermore, the method includes encoding the plurality of images based on reference tiles located within the reference image. Additionally, the method includes selecting a location within the reference image for the reference tiles such that the reference tiles are not both partially located within three of the three or more tiles in the independently encoded group of tiles and partially located within another tile from the plurality of tiles in the reference image that does not belong to the independently encoded group of tiles.
[0044] In addition, a computer program is provided for implementing one of the above methods when executed on a computer or signal processor according to an embodiment.
[0045] Furthermore, an encoded video signal according to an embodiment is provided, the encoded video signal encoding a plurality of images comprising a plurality of tiles. Each of the plurality of tiles comprises a plurality of samples. The encoded video signal includes encoded image data encoding the plurality of images. The encoded video signal includes encoding of the plurality of images. Each of the plurality of images comprises a plurality of tiles, wherein each of the plurality of tiles comprises a plurality of samples. The encoded video signal includes independently encoded tile groups, the independently encoded tile groups comprising three or more tiles from a plurality of tiles of a reference image of the plurality of images. The plurality of images are encoded within a video data stream based on a reference tile located within the reference image. The reference tile is not located both partially within three tiles of the three or more tiles of the independently encoded tile group and partially within another tile of the plurality of tiles of the reference image that does not belong to the independently encoded tile group.
[0046] Furthermore, a system according to an embodiment is provided, including the video encoder and video decoder described above, for decoding an encoded video signal including encoded image data to reconstruct a plurality of images of a video. The video decoder includes an input interface configured to receive the encoded video signal, and a data decoder configured to reconstruct the plurality of images of the video by decoding the encoded image data. The video encoder is configured to generate the encoded video signal. The video decoder is configured to decode the encoded video signal to reconstruct images of the video.
[0047] A video encoder is provided according to embodiments for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder includes a data encoder configured to generate an encoded video signal including the encoded image data, wherein the data encoder is configured to encode the plurality of images of the video into encoded image data, and an output interface configured to output the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. The data encoder is configured to determine an independently encoded group of tiles comprising three or more tiles of a plurality of tiles of a reference image of the plurality of images. Furthermore, the data encoder is configured to encode the plurality of images based on a reference block located within the reference image, wherein the reference block is partially located within three of the three or more tiles of the independently encoded group of tiles and partially located within another tile of the plurality of tiles of the reference image that does not belong to the independently encoded group of tiles. Furthermore, the data encoder is configured to determine a plurality of reference samples located within a portion of the reference block that does not belong to the other block of the independently encoded tile group, based on one or more of a plurality of samples of the first of the three tiles in the independently encoded tile group and one or more of a plurality of samples of the second tile in the three tiles of the independently encoded tile group.
[0048] Furthermore, a video decoder is provided for decoding an encoded video signal including encoded image data to reconstruct a plurality of images of a video, according to an embodiment. The video decoder includes an input interface configured to receive the encoded video signal, and a data decoder configured to reconstruct the plurality of images of the video by decoding the encoded image data. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. The encoded video signal includes independently encoded groups of tiles, the independently encoded groups of tiles including three or more tiles from a plurality of tiles of a reference image of the plurality of images. The data decoder is configured to decode the plurality of images based on a reference block located within the reference image, wherein the reference block is partially located within three of the three or more tiles of the independently encoded group of tiles and partially located within another tile of the plurality of tiles of the reference image that does not belong to the independently encoded group of tiles. Furthermore, the data decoder is configured to determine a plurality of reference samples located within a portion of the reference block that does not belong to the other block of the independently encoded tile group, based on one or more of a plurality of samples of the first of the three tiles in the independently encoded tile group and one or more of a plurality of samples of the second tile in the three tiles of the independently encoded tile group.
[0049] Furthermore, a method is provided for encoding multiple images of a video by generating an encoded video signal. Each of the multiple images includes original image data. The method includes:
[0050] - Generate an encoded video signal comprising encoded image data, wherein generating the encoded video signal includes encoding multiple images of the video into encoded image data. And:
[0051] - Outputs the encoded image data for each of multiple images.
[0052] Each of a plurality of images comprises a plurality of tiles, wherein each of the plurality of tiles comprises a plurality of samples. The method includes determining an independently encoded group of tiles, the independently encoded group of tiles comprising three or more tiles from a plurality of tiles in a reference image of the plurality of images. Furthermore, the method includes encoding the plurality of images based on a reference block located within the reference image, wherein the reference block is partially located within three of the three or more tiles in the independently encoded group of tiles and partially located within another tile in the plurality of tiles of the reference image that does not belong to the independently encoded group of tiles. Furthermore, the method includes determining a plurality of reference samples located within a portion of the reference block that does not belong to the other tile in the independently encoded group of tiles, based on one or more of a plurality of samples from a first of the three tiles in the independently encoded group of tiles, and based on one or more of a plurality of samples from a second of the three tiles in the independently encoded group of tiles.
[0053] Furthermore, a method is provided according to an embodiment for decoding an encoded video signal including encoded image data to reconstruct a video from multiple images. The decoding method includes:
[0054] - Receive encoded video signals. And:
[0055] - Reconstruct multiple images from a video by decoding encoded image data.
[0056] Each of the multiple images comprises multiple tiles, and each of the multiple tiles comprises multiple samples. The encoded video signal comprises independent encoded tile groups, which include three or more tiles from multiple reference images.
[0057] The method includes decoding multiple images based on reference blocks located within a reference image, wherein the reference blocks are partially located within three or more tiles of an independently encoded tile group and partially located within another tile of a plurality of tiles of a reference image that does not belong to the independently encoded tile group. Furthermore, the method includes determining multiple reference samples based on one or more samples of a plurality of samples of the first of the three tiles of the independently encoded tile group, and based on one or more samples of a plurality of samples of the second of the three tiles of the independently encoded tile group, of a portion of the reference block located within the other tile of the independently encoded tile group.
[0058] In addition, a computer program is provided for implementing one of the above methods when executed on a computer or signal processor according to an embodiment.
[0059] Furthermore, according to an embodiment, an encoded video signal is provided that encodes a plurality of images comprising a plurality of tiles, each of which comprises a plurality of samples. The encoded video signal includes encoded image data encoding the plurality of images. Furthermore, the encoded video signal includes encoding of the plurality of images. Each of the plurality of images comprises a plurality of tiles, each of which comprises a plurality of samples. A group of three or more independently encoded tiles from a plurality of tiles in a reference image comprising the plurality of images is independently encoded within the encoded video signal. Based on a reference block located within the reference image, the plurality of images are encoded within the encoded video signal, wherein the reference block is partially located within three of the three or more tiles in the independently encoded tile group, and partially located within another tile of the plurality of tiles in the reference image that does not belong to the independently encoded tile group. Multiple reference samples located within a portion of the reference block that does not belong to the independently coded tile group can derive one or more samples from a plurality of samples of the first tile of the three tiles in the independently coded tile group and one or more samples from a plurality of samples of the second tile of the three tiles in the independently coded tile group.
[0060] Furthermore, a system according to an embodiment is provided, comprising the aforementioned video encoder and video decoder. The video encoder is configured to generate an encoded video signal. The video decoder is configured to decode the encoded video signal to reconstruct an image of the video.
[0061] Furthermore, a video encoder according to an embodiment is provided for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder includes a data encoder configured to generate an encoded video signal including the encoded image data, wherein the data encoder is configured to encode the plurality of images of the video into encoded image data, and an output interface configured to output the encoded image data of each of the plurality of images. The data encoder is configured to encode hash information within the encoded video signal. Furthermore, the data encoder is configured to generate hash information based on a current portion of a current image among the plurality of images, without depending on subsequent portions of the current image, wherein the current portion has a first position in the current image, and subsequent portions have a second position in the image different from the first position.
[0062] Furthermore, a video decoder according to an embodiment is provided for decoding an encoded video signal including encoded image data to reconstruct a plurality of images of a video. The video decoder includes an input interface configured to receive the encoded video signal, and a data decoder configured to reconstruct the plurality of images of the video by decoding the encoded image data. The data decoder is configured to analyze encoded hash information within the encoded video signal, wherein the hash information depends on a current portion of a current image of the plurality of images, but not on subsequent portions of the current image, wherein the current portion has a first position within the current image, and the subsequent portions have a second position in the image different from the first position.
[0063] Furthermore, a method for encoding multiple images of a video by generating an encoded video signal, wherein each of the multiple images includes original image data. The method includes:
[0064] - Generate an encoded video signal comprising encoded image data, wherein generating the encoded video signal includes encoding multiple images of the video into encoded image data. And:
[0065] - Output the encoded image data for each of the multiple images.
[0066] The method includes generating hash information based on a current portion of a current image among a plurality of images, independent of subsequent portions of the current image, wherein the current portion has a first position within the current image, and subsequent portions have a second position in the image different from the first position. Furthermore, the method includes encoding the hash information within an encoded video signal.
[0067] Furthermore, a method is provided according to an embodiment for decoding a encoded video signal including encoded image data to reconstruct a video from multiple images. The decoding method includes:
[0068] - Receive encoded video signals. And:
[0069] - Reconstruct multiple images from a video by decoding encoded image data.
[0070] The method includes analyzing encoded hash information within an encoded video signal, wherein the hash information depends on a current portion of a current image among multiple images, but not on subsequent portions of the current image, wherein the current portion has a first position within the current image, and subsequent portions have a second position in the image different from the first position.
[0071] In addition, a computer program according to an embodiment is provided for implementing one of the above methods when executed on a computer or signal processor.
[0072] Furthermore, an encoded video signal according to an embodiment is provided, the encoded video signal encoding a plurality of images comprising a plurality of tiles. Each of the plurality of tiles comprises a plurality of samples, wherein the encoded video signal comprises encoded image data encoding the plurality of images. The encoded video signal includes encoding of the plurality of images. The encoded video signal includes encoding of hash information, wherein the hash information depends on a current portion of a current image among the plurality of images, but not on subsequent portions of the current image, wherein the current portion has a first position in the current image, and subsequent portions have a second position in the image different from the first position.
[0073] Furthermore, a system according to an embodiment is provided, comprising the aforementioned video encoder and video decoder. The video encoder is configured to generate an encoded video signal. The video decoder is configured to decode the encoded video signal to reconstruct an image of the video.
[0074] Preferred embodiments are provided in the dependent claims. Attached Figure Description
[0075] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings, wherein:
[0076] Figure 1 A video encoder according to an embodiment is shown.
[0077] Figure 2 A video decoder according to an embodiment is shown.
[0078] Figure 3 A system according to an embodiment is shown.
[0079] Figure 4a The image shows a contaminated sample within a tile from the loop filter procedure.
[0080] Figure 4b The VVC tiling boundary extension of an independent region from a contaminated sample is shown.
[0081] Figure 5 The process of expanding the tile boundary according to an embodiment, which conforms to the influence range of the loop filter kernel, is illustrated.
[0082] Figure 6 The diagram shows the division of tiles and tile groups in the encoded image.
[0083] Figure 7 A reference block with prior art boundary padding is shown.
[0084] Figure 8 The boundary of a diagonally segmented concave block group according to an embodiment is shown.
[0085] Figure 9 A video encoder is shown.
[0086] Figure 10 The video decoder is shown.
[0087] Figure 11 This illustrates the relationship between the reconstructed signal (i.e., the reconstructed image) and the combination of the predicted residual signal and the predicted signal transmitted in the data stream. Detailed Implementation
[0088] The following description of the figures begins with the presentation of a description of the encoder and decoder of a block-based predictive codec for encoding images in video, in order to form an example of an encoding framework in which embodiments of the invention can be established. Relative to Figures 9 to 11 The corresponding encoders and decoders are described. The following provides a description of embodiments of the inventive concepts and how these concepts can be respectively incorporated into... Figure 9 and Figure 10 In the encoder and decoder, although Figures 1 to 3 The embodiments described below can also be used to form non-based Figure 9 and Figure 10 The encoder and decoder are the encoding framework operations of the encoder and decoder.
[0089] Figure 9 A video encoder is shown, an apparatus for predictively encoding an image 12 into a data stream 14 using transform-based residual coding, as exemplarily. This apparatus or encoder is indicated by reference numeral 10. Figure 10 The corresponding video decoder 20 is shown, i.e., the device 20 configured to also predictively decode the image 12' from the data stream 14 using transform-based residual decoding, wherein the apostrophe has been used to indicate that the image 12' reconstructed by the decoder 20 deviates from the image 12 originally encoded by the device 10 in terms of the coding loss introduced by the quantization of the predictive residual signal. Figure 9 and Figure 10 Transform-based prediction residual coding is used as an example, but embodiments of this application are not limited to this prediction residual coding. Regarding... Figure 9 and Figure 10 The same applies to other details described below.
[0090] Encoder 10 is configured to perform a spatial-to-spectral transformation on the prediction residual signal and encode the resulting prediction residual signal into data stream 14. Similarly, decoder 20 is configured to decode the prediction residual signal from data stream 14 and perform a spectral-to-spatial transformation on the resulting prediction residual signal.
[0091] Internally, encoder 10 may include a prediction residual signal former 22, which generates a prediction residual 24 to measure the deviation of the prediction signal 26 from the original signal, i.e., from the image 12. The prediction residual signal former 22 may, for example, be a subtractor that subtracts the prediction signal from the original signal, i.e., from the image 12. Encoder 10 then further includes a transformer 28 that subjectes the prediction residual signal 24 to a space-to-spectrum transformation to obtain a spectral-domain prediction residual signal 24', which is then quantized by a quantizer 32, also included in encoder 10. The quantized prediction residual signal 24' is thus encoded into bitstream 14. For this purpose, encoder 10 may optionally include an entropy encoder 34 to entropy encode the prediction residual signal transformed and quantized into data stream 14. Based on the prediction residual signal 24'' encoded into data stream 14 and decodeable from data stream 14, prediction signal 26 is generated by prediction stage 36 of encoder 10. For this purpose, prediction stage 36 may be internally configured, such as... Figure 9 As shown, a dequantizer 38 is included to dequantize the prediction residual signal 24” to obtain a spectral domain prediction residual signal 24”', which corresponds to the signal 24' except for the quantization loss. This is followed by an inverse transformer 40, which inversely transforms the latter's prediction residual signal 24”', i.e., a spectral-to-spatial transformation, to obtain a prediction residual signal 24”', which corresponds to the original prediction residual signal 24 except for the quantization loss. Then, a combiner 42 of prediction stage 36 recombines the prediction signal 26 and the prediction residual signal 24”', for example, by addition, to obtain a reconstructed signal 46, i.e., a reconstruction of the original signal 12. The reconstructed signal 46 can then correspond to signal 12'. The prediction module 44 of prediction stage 36 generates the prediction signal 26 based on signal 46 using, for example, spatial prediction (i.e., intra-image prediction) and / or temporal prediction (i.e., inter-image prediction).
[0092] Similarly, decoder 20, such as Figure 10 As shown, it can be internally composed of components corresponding to prediction stage 36 and interconnected in a manner corresponding to prediction stage 36. Specifically, the entropy decoder 50 of decoder 20 can perform entropy decoding on the quantized spectral domain prediction residual signal 24” from the data stream. Therefore, the dequantizer 52, inverse transformer 54, combiner 56, and prediction module 58, interconnected and cooperating in the manner described above with respect to prediction stage 36, recover the reconstructed signal based on the prediction residual signal 24”. Thus, as... Figure 10 As shown, the output of combiner 56 generates the reconstructed signal, i.e., image 12'.
[0093] Although not specifically described above, it is readily apparent that encoder 10 can set encoding parameters, including prediction modes and motion parameters, based on optimization schemes such as optimizing rates and distortion-related criteria, i.e., encoding costs. For example, modules 44 and 58 corresponding to encoder 10 and decoder 20 can support different prediction modes, such as intra-coding modes and inter-coding modes. The granularity of switching between these prediction mode types by encoder and decoder can correspond to subdividing images 12 and 12' into coded segments or coded blocks, respectively. For example, based on these coded segments, the image can be subdivided into blocks that are intra-coded and blocks that are inter-coded. As outlined in more detail below, intra-coded blocks are predicted based on the spatial composition of the corresponding block and the already encoded / decoded neighborhood. Multiple intra-coding modes can exist and be selected for corresponding intra-coding modes that include directional or angular intra-coding modes. Based on these modes, the corresponding segments are filled by extrapolating sample values of the neighborhood along a specific direction, which is specific to the corresponding oriented intra-coding mode, into the corresponding intra-coded segment. For example, the intra-coding mode may also include one or more other modes, such as a DC coding mode, according to which the prediction of the corresponding intra-coding block assigns DC values to all samples within the corresponding intra-coding segment, and / or an in-plane coding mode, according to which the prediction of the corresponding block is approximated or determined as a spatial distribution of sample values described by a two-dimensional linear function at the sample locations of the corresponding intra-coding block, driving the tilt and offset of the plane defined by the two-dimensional linear function on the basis of adjacent samples. In contrast, inter-coding blocks can be predicted, for example, in time. For inter-coding blocks, motion vectors can be signaled within the data stream, the motion vectors indicating the spatial displacement of a previously coded image portion of the video to which image 12 belongs, where the previously coded / decoded image is sampled to obtain the prediction signal for the corresponding inter-coding block. This means that, in addition to the residual signal encoding included in data stream 14, such as the entropy-coded transform coefficient level representing the quantized spectral domain prediction residual signal 24", data stream 14 may have already encoded therein coding mode parameters for assigning coding modes to various blocks, prediction parameters for some blocks, such as motion parameters for inter-coded segments, and optional further parameters, such as those for controlling and signaling to subdivide images 12 and 12' into segments respectively. Decoder 20 uses these parameters to subdivide the image in the same way as encoder, assign the same prediction modes to segments, and perform the same predictions to produce the same prediction signals.
[0094] Figure 11 This explains the relationship between the reconstructed signal, i.e., the reconstructed image 12', and the combination of the prediction residual signal 24"" signal notified by the signal in data stream 14, and the prediction signal 26. As mentioned above, this combination can be additive. The prediction signal 26 is in Figure 11The image region is shown as being subdivided into inner coding blocks, illustratively represented by shading, and inter coding blocks, illustratively represented without shading. The subdivision can be arbitrary, such as regularly subdividing the image region into rows and columns of square or non-square blocks, or subdividing the image 12 from a root block into multiple leaf blocks of variable size, such as quadtree subdivision, etc. Figure 11 The mixture is shown, in which the image region is first subdivided into rows and columns of root blocks, and then the root blocks are further subdivided into one or more leaf blocks according to recursive multi-subdivision.
[0095] Furthermore, data stream 14 may have an intra-coding mode encoded therein for intra-coding block 80, which assigns one of several supported intra-coding modes to the corresponding intra-coding block 80. For inter-coding block 82, data stream 14 may have one or more motion parameters encoded therein. In general, inter-coding block 82 is not limited to being time-coded. Alternatively, inter-coding block 82 may be any block predicted from a previously coded portion beyond the current image 12 itself, such as a previously coded image of the video to which image 12 belongs, or an image of another view, or a lower layer in the case where the encoder and decoder are scalable encoders and decoders, respectively.
[0096] Figure 11 The prediction residual signal 24” in the image is also shown as subdividing the image region into blocks 84. These blocks can be called transform blocks to distinguish them from coded blocks 80 and 82. In fact, Figure 11 The encoder 10 and decoder 20 can use two different subdivisions of images 12 and 12' to divide them into blocks, namely, one subdivision into coding blocks 80 and 82 respectively, and another subdivision into transform block 84. The two subdivisions can be the same, meaning that each coding block 80 and 82 can simultaneously form transform block 84, but... Figure 11 The illustration shows an extension of the subdivision of coded blocks 80 and 82, for example, by subdividing into transform blocks 84, such that any boundary between the two blocks 80 and 82 covers the boundary between the two blocks 84, or in other words, each block 80 and 82 either coincides with one of the transform blocks 84 or with a set of transform blocks 84. However, the subdivisions can also be determined or selected independently of each other, such that transform blocks 84 can alternatively span the block boundaries between blocks 80 and 82. With regard to subdivision into transform blocks 84, similar statements are therefore identical to those regarding subdivision into blocks 80 and 82, i.e., block 84 can be the result of a regular subdivision that separates the image region into blocks (with or without arrangement in rows and columns), the result of a recursive multi-tree subdivision of the image region, or a combination thereof, or any other type of block. Incidentally, it should be noted that blocks 80, 82, and 84 are not limited to squares, rectangles, or any other shape.
[0097] Figure 11This further illustrates that the combination of the prediction signal 26 and the prediction residual signal 24”” directly generates the reconstructed signal 12'. However, it should be noted that, according to alternative embodiments, more than one prediction signal 26 can be combined with the prediction residual signal 24”” to form image 12'.
[0098] exist Figure 11 In this context, transform block 84 should have the following meaning. Transformer 28 and inverse transformer 54 perform their transforms on a unit basis of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow skipping transforms so that for some transform blocks 84, the prediction residual signal is directly encoded in the spatial domain. However, according to the embodiments described below, encoder 10 and decoder 20 are configured in a way that they support several transforms. For example, the transforms supported by encoder 10 and decoder 20 may include:
[0099] o DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform
[0100] о DST-IV, where DST represents Discrete Sine Transform
[0101] о DCT-IV
[0102] о DST-VII
[0103] identity transformation (IT)
[0104] Naturally, while transformer 28 will support all forward transform versions of these transforms, decoder 20 or inverse transformer 54 will support their respective backward or inverse versions:
[0105] inverse DCT-II (or inverse DCT-III)
[0106] reverse DST-IV
[0107] reverse DCT-IV
[0108] о Inverse DST-VII
[0109] identity transformation (IT)
[0110] The following description provides further details about which transforms the encoder 10 and decoder 20 can support. In any case, it should be noted that the set of supported transforms may include only one transform, such as a spectrum-to-space or space-to-spectrum transform.
[0111] As mentioned above, Figures 9 to 11Specific examples of encoders and decoders according to this application have been presented as examples, in which the inventive concepts described further below can be implemented. In this regard, Figure 9 and Figure 10 The encoder and decoder can respectively represent possible implementations of the encoder and decoder described below. However, Figure 9 and Figure 10 This is merely an example. However, the encoder according to embodiments of this application can perform block-based encoding of image 12 using concepts outlined in more detail below, and differs from... Figure 9 Encoders, such as, for example, are not video encoders but still image encoders because they do not support inter-image prediction, or are different from... Figure 11 The example in the text demonstrates the subdivision into blocks 80. Similarly, the decoder according to embodiments of this application can perform block-based decoding of the image 12' from data stream 14 using the encoding concepts further outlined below, but can be combined with, for example... Figure 10 The difference between the decoder 20 and the others is that it is not a video decoder, but a still image decoder, and it also does not support intraprediction, or in a way that differs from the one mentioned above. Figure 11 The described method subdivides the image 12' into blocks and / or similarly does not derive prediction residuals from the data stream 14 in the transform domain, but derives them in the spatial domain.
[0112] Below, a general video encoder according to an embodiment is... Figure 1 As described in the description, the general video decoder according to the embodiments is in Figure 2 The general system described in the embodiments, and in the general system described in the embodiments, Figure 3 As described in the text.
[0113] Figure 1 A general video encoder 101 according to an embodiment is shown.
[0114] The video encoder 101 is configured to encode multiple images of a video by generating an encoded video signal, each of the multiple images including original image data.
[0115] The video encoder 101 includes a data encoder 110 configured to generate an encoded video signal comprising encoded image data, wherein the data encoder is configured to encode multiple images of the video into encoded image data.
[0116] In addition, the video encoder 101 includes an output interface 120 configured to output encoded image data of each of a plurality of images.
[0117] Figure 2 A general video decoder 151 according to an embodiment is shown.
[0118] The video decoder 151 is configured to decode an encoded video signal, including encoded image data, to reconstruct multiple images of the video.
[0119] The video decoder 151 includes an input interface 160 configured to receive encoded video signals.
[0120] In addition, the video decoder includes a data decoder 170, which is configured to reconstruct multiple images of the video by decoding the encoded image data.
[0121] Figure 3 A general system according to an embodiment is shown.
[0122] The system includes Figure 1 Video encoder 101 and Figure 2 The video decoder 151.
[0123] Video encoder 101 is configured to generate an encoded video signal. Video decoder 151 is configured to decode the encoded video signal to reconstruct a video image.
[0124] There are video applications where it is beneficial to divide the video into rectangular tiles / regions and encode them independently. For example, in a 360-degree video stream, the client's current viewing direction is used to select the resolution of individual regions (high resolution within the current viewport, and low resolution outside the current viewport as a replacement for changes in user orientation). These tiles / regions are then reassembled into a single bitstream and jointly decoded on the client, whereby each tile may have different neighboring tiles that were unavailable or nonexistent during encoding.
[0125] Other examples could be RoI (Region of Interest) encoding, where, for example, there is a region in the middle of the image that the viewer can choose from, for example, using a zoom-in operation (decoding only the RoI) or progressive decoder refresh (GDR), where the intra data (typically placed in a frame of the video sequence) is temporally distributed across several consecutive frames, for example, as columns of intra blocks, which slide across the image plane and locally reset the temporal prediction chain in the same way that the intra images reset the entire image plane. In the latter case, there are two regions in each image: one that was recently reset and another that may be affected by errors and error propagation.
[0126] For these use cases, and potentially others, limiting the predictive dependencies of images from different times is crucial so that (certain) regions / patterns can be encoded independently. However, this leads to several problems that this invention addresses. First, conventional boundary filtering cannot be used to mitigate the subjective quality effects of dividing an image plane into separate regions. Second, existing technologies do not describe how certain boundary geometries should be filtered. Third, standard image hashing signaling such as HEVC cannot be meaningfully used in the aforementioned use cases because value derivation involves the entire image plane.
[0127] The first aspect of the invention is claimed in claims 1 to 45.
[0128] The second aspect of the invention is claimed in claims 46 to 74.
[0129] The third aspect of the invention is claimed in claims 75 to 101.
[0130] The first aspect of the invention will now be described in detail below.
[0131] Specifically, the first aspect provides drift-free filtering.
[0132] A video decoder 151 according to an embodiment is provided for decoding an encoded video signal including encoded image data to reconstruct a plurality of images of a video. The video decoder 151 includes an input interface 160 configured to receive the encoded video signal, and a data decoder 170 configured to reconstruct the plurality of images of the video by decoding the encoded image data. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of a first image, the data decoder 170 is configured to filter across the boundary between the first tile and the second tile to obtain a first filtered tile, wherein the first tile and the second tile have been encoded independently relative to each other. The data decoder 170 is configured to decode a current tile in a plurality of tiles of a second image in the plurality of images based on a reference block of the first filtered tile of the first image, wherein the reference block includes a first set of samples of the first filtered tile, and wherein the reference block does not include a second set of samples of the first filtered tile, wherein none of the first set of samples is affected by the filtering across the boundary between the first tile and the second tile, and wherein one or more of the second set of samples has been affected by the filtering across the boundary between the first tile and the second tile.
[0133] For example, in some embodiments, two tiles may be considered adjacent, or may be considered adjacent tiles if they are adjacent / adjacent to each other in the images of multiple images, or if they are adjacent in location if they are adjacent / adjacent to each other in the map after being mapped, for example, by one or more mappings according to mapping rules into the map (e.g., to a projection map, or, for example, to a cube map, or, for example, to an isometric histogram). For example, tiles may be mapped, for example, by employing region packing. For example, tiles may be mapped to a projection map by a first mapping, and, for example, by a second mapping from the projection map to an isometric histogram.
[0134] In an embodiment, the data decoder 170 may be configured, for example, not to determine another reference block for decoding the current tile of the second image, wherein the other reference block includes one or more of the second set of samples of the first filtered tile, which has been affected by the filtering across the boundary between the first tile and the second tile.
[0135] According to an embodiment, the data decoder 170 may be configured, for example, to determine the reference block such that the reference block includes the first set of samples of the first filtered tile, and such that the reference block does not include the second set of samples of the first filtered tile, such that none of the first set of samples is affected by the filtering across the boundary between the first tile and the second tile, and such that one or more samples in the second set of samples have been affected by the filtering across the boundary between the first tile and the second tile.
[0136] In an embodiment, the data decoder 170 may be configured, for example, to determine the reference block based on the influence range of a filter or filter kernel, wherein the data decoder 170 may be configured, for example, to employ the filter or filter kernel for filtering across the boundary between the first and second blocks.
[0137] According to an embodiment, the data decoder 170 may be configured, for example, to determine the reference block based on filter information regarding the influence range of a filter or filter kernel. The filter information includes a horizontal filter kernel influence range indicating how many samples of the first patch within a horizontal row of the first patch are influenced by one or more samples of the second patch, which is filtered across the boundary between the first and second patches, wherein the data decoder 170 may be configured, for example, to determine the reference block based on the horizontal filter kernel influence range. And / or: the filter information includes a vertical filter kernel influence range indicating how many samples of the first patch within a vertical column of the first patch are influenced by one or more samples of the second patch, which is filtered across the boundary between the first and second patches, wherein the data decoder 170 may be configured, for example, to determine the reference block based on the vertical filter kernel influence range.
[0138] According to an embodiment, the data decoder 170 may be configured, for example, to determine the reference block based on the influence range of the vertical filter kernel by extrapolating samples from the first set of samples. And / or the data decoder 170 may be configured, for example, to determine the reference block based on the influence range of the horizontal filter kernel by extrapolating samples from the first set of samples.
[0139] In an embodiment, the data decoder 170 may be configured, for example, to determine the reference block based on the influence range of the vertical filter kernel by using vertical shearing with the first set of samples, and / or the data decoder 170 may be configured, for example, to determine the reference block based on the influence range of the horizontal filter kernel by using horizontal shearing with the first set of samples.
[0140] According to an embodiment, vertical shearing can be defined, for example, according to the following:
[0141] yInt i =Clip3(topTileBoundaryPosition+verticalFilterKernelReachInSamples,bottomTileBoundaryPosition–1–verticalFilterKernelReachInSamples,yInt L +i-3)
[0142] yInt L yInt represents one of a plurality of samples of the first patch located at position L in the vertical column of the first patch before vertical shearing.i Let represent one of the multiple samples of the first tile at position i in the vertical column of the first tile after vertical shearing; let represent the number of samples representing the influence range of the vertical filter kernel; let represent the top position of the multiple samples in the vertical column of the first tile; let represent the bottom position of the multiple samples in the vertical column of the first tile; and let represent the bottom position of the multiple samples in the vertical column of the first tile. The horizontal shearing can be defined, for example, according to the following:
[0143] xInt i =Clip3(leftTileBoundaryPosition+horizontalFilterKernelReachInSamples,rightTileBoundaryPosition–1–horizontalFilterKernelReachInSamples,xInt L +i-3)
[0144] xInt L xInt represents one of a plurality of samples of the first tile at position L in the horizontal row of the first tile before horizontal shearing. i This represents one of the multiple samples of the first tile at position i in the horizontal row after horizontal clipping. `horizontalFilterKernelReachInSamples` represents the number of samples representing the influence range of the horizontal filter kernel. `leftTileBoundaryPosition` represents the leftmost position of the multiple samples within the horizontal row of the first tile. `rightTileBoundaryPosition` represents the rightmost position of the multiple samples within the horizontal row of the first tile. `Clip3` is defined as follows:
[0145]
[0146] According to an embodiment, the data decoder 170 may be configured, for example, to determine the reference block by employing the horizontal shearing based on the influence range of the horizontal filter kernel and by employing the vertical shearing based on the influence range of the vertical filter kernel, wherein
[0147] verticalFilterKernelReachInSamples=horizontalFilterKernelReachInSamples.
[0148] In an embodiment, the data decoder 170 may be configured, for example, to determine the reference block by employing the horizontal shearing based on the influence range of the horizontal filter kernel and by employing the vertical shearing based on the influence range of the vertical filter kernel, wherein
[0149] verticalFilterKernelReachInSamples ≠ horizontalFilterKernelReachInSamples.
[0150] According to an embodiment, the data decoder 170 may be configured, for example, to filter the first tile using the filter or the filter kernel, wherein the data decoder 170 may be configured, for example, to adjust the influence range of the filter or the filter kernel based on the distance between the block of the first tile to be filtered by the filter or the filter kernel and the boundary between the first tile and the second tile.
[0151] In an embodiment, if the distance has a first distance value less than or equal to a threshold distance, the data decoder 170 may be configured, for example, to set the influence range of the filter or the filter kernel to a first size value. If the distance has a second distance value greater than the first distance value, and if the block and its neighboring blocks belong to the same reference image, the data decoder 170 may be configured, for example, to set the influence range of the filter or the filter kernel to a second size value greater than the first size value. The data decoder 170 may be configured, for example, to set the influence range of the filter or the filter kernel to the first size value if the distance has a second distance value greater than the first distance value and if the block and its neighboring blocks do not belong to the same reference image.
[0152] Furthermore, a video encoder 101 according to an embodiment is provided for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder 101 includes a data encoder 110 for generating an encoded video signal containing encoded image data, wherein the data encoder 110 is used to encode the plurality of images of the video into encoded image data, and an output interface 120 for outputting the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of a first image of the plurality of images, a boundary exists between the first tile and the second tile. The data encoder 110 is configured to encode the first tile and the second tile independently of each other. Furthermore, the data encoder 110 is configured to encode the current tile of a plurality of tiles in a plurality of images based on a reference block of the first tile of the first image, wherein the filter defines filtering across the boundary between the first tile and the second tile, wherein the reference block includes a first set of samples of the first tile, and wherein the reference block does not include a second set of samples of the first tile, wherein none of the first set of samples will be affected by the filtering using the filter, and wherein one or more of the second set of samples will be affected by the filtering using the filter.
[0153] For example, in some embodiments, two tiles may be considered adjacent, or may be considered adjacent tiles if they are adjacent / adjacent to each other in the images of multiple images, or if they are adjacent if they should be mapped to the decoder side such that they are adjacent / adjacent to each other in the mapping after being mapped to the decoder side, for example, by one or more mappings according to mapping rules to a map (e.g., a projection map, or, for example, a cube map, or, for example, an isometric histogram). For example, tiles may be mapped, for example, by employing region packing. For example, tiles may be mapped to a projection map, for example, by a first mapping, and, for example, by a second mapping from the projection map to an isometric histogram.
[0154] In an embodiment, the data encoder 110 may be configured, for example, to not encode the current tile based on another reference block, which includes one or more of the second set of samples of the first filtered tile that have been affected by the filtering across the boundary between the first tile and the second tile.
[0155] According to an embodiment, the data encoder 110 may be configured, for example, to determine the reference block such that the reference block includes the first set of samples of the first patch, and such that the reference block does not include the second set of samples of the first patch, such that none of the samples in the first set are affected by the filtering using the filter, and such that one or more samples in the second set are affected by the filtering using the filter.
[0156] In an embodiment, the data encoder 110 may be configured, for example, to determine the reference block based on the influence range of a filter or filter kernel, wherein the data encoder 110 may be configured, for example, to employ the filter or filter kernel for filtering across the boundary between the first and second blocks.
[0157] According to an embodiment, the data encoder 110 may be configured, for example, to determine the reference block based on filter information regarding the influence range of a filter or filter kernel. The filter information includes a horizontal filter kernel influence range indicating how many samples of the first patch within a horizontal row of the first patch are affected by one or more samples of the second patch, filtering across the boundary between the first and second patches, wherein the data encoder 110 may be configured, for example, to determine the reference block based on the horizontal filter kernel influence range. And / or the filter information includes a vertical filter kernel influence range indicating how many samples of the first patch within a vertical column of the first patch are affected by one or more samples of the second patch, filtering across the boundary between the first and second patches, wherein the data encoder 110 may be configured, for example, to determine the reference block based on the vertical filter kernel influence range.
[0158] According to an embodiment, the data encoder 110 may be configured, for example, to determine the reference block by extrapolating samples in the first set of samples based on the influence range of the vertical filter kernel. And / or, the data encoder 110 may be configured, for example, to determine the reference block by extrapolating samples in the first set of samples based on the influence range of the horizontal filter kernel.
[0159] In an embodiment, the data encoder 110 may be configured, for example, to determine the reference block based on the influence range of the vertical filter kernel by using vertical shearing with the first set of samples. And / or the data encoder 110 may be configured, for example, to determine the reference block based on the influence range of the horizontal filter kernel by using horizontal shearing with the first set of samples.
[0160] According to an embodiment, vertical shearing can be defined, for example, according to the following:
[0161] yInt i =Clip3(topTileBoundaryPosition+verticalFilterKernelReachInSamples,bottomTileBoundaryPosition–1–verticalFilterKernelReachInSamples,yInt L +i-3)
[0162] yInt L yInt represents one of multiple samples of the first patch at position L in the vertical column before vertical shearing. i The vertical filter kernel (verticalFilterKernelReachInSamples) indicates one of the samples in the first tile that is located at position i in the vertical column of the first tile after vertical shearing; the topTileBoundaryPosition indicates the topmost position of the samples in the vertical column of the first tile; and the bottomTileBoundaryPosition indicates the bottommost position of the samples in the vertical column of the first tile. The horizontal shearing can be defined, for example, according to the following:
[0163] xInt i =Clip3(leftTileBoundaryPosition+horizontalFilterKernelReachInSamples,rightTileBoundaryPosition–1–horizontalFilterKernelReachInSamples,xInt L +i-3)
[0164] xInt L xInt represents one of a plurality of samples of the first tile at position L in the horizontal row of the first tile before horizontal shearing. iThis represents one of the multiple samples of the first tile located at position i in the horizontal direction after horizontal clipping. `horizontalFilterKernelReachInSamples` represents the number of samples representing the influence range of the horizontal filter kernel. `leftTileBoundaryPosition` represents the leftmost position of the multiple samples within the horizontal row of the first tile. `rightTileBoundaryPosition` represents the rightmost position of the multiple samples within the horizontal row of the first tile. `Clip3` is defined as follows:
[0165]
[0166] In an embodiment, the data encoder 110 may be configured, for example, to determine the reference block by employing the horizontal shearing based on the influence range of the horizontal filter kernel and by employing the vertical shearing based on the influence range of the vertical filter kernel, wherein
[0167] verticalFilterKernelReachInSamples=horizontalFilterKernelReachInSamples.
[0168] According to an embodiment, the data encoder 110 may be configured, for example, to determine the reference block by employing the horizontal shearing based on the influence range of the horizontal filter kernel and by employing the vertical shearing based on the influence range of the vertical filter kernel, wherein...
[0169] verticalFilterKernelReachInSamples ≠ horizontalFilterKernelReachInSamples.
[0170] In an embodiment, the data encoder 110 may be configured, for example, to filter the first tile using the filter or the filter kernel. The data encoder 110 may be configured, for example, to adjust the influence range of the filter or the filter kernel based on the distance between the block of the first tile to be filtered by the filter or the filter kernel and the boundary between the first tile and the second tile.
[0171] According to an embodiment, if the distance has a first distance value less than or equal to a threshold distance, the data encoder 110 may be configured, for example, to set the influence range of the filter or the filter kernel to a first size value. If the distance has a second distance value greater than the first distance value and if the block and its neighboring blocks belong to the same reference image, the data encoder 110 may be configured, for example, to set the influence range of the filter or the filter kernel to a second size value greater than the first size value. Furthermore, if the distance has a second distance value greater than the first distance value and if the block and its neighboring blocks do not belong to the same reference image, the data encoder 110 may be configured, for example, to set the influence range of the filter or the filter kernel to the first size value.
[0172] Furthermore, a system including the aforementioned video encoder 101 and video decoder 151 is provided. The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is configured to decode the encoded video signal and reconstruct a video image.
[0173] Furthermore, a video decoder 151 is provided according to an embodiment for decoding an encoded video signal comprising encoded image data to reconstruct a plurality of images of a video. The video decoder 151 includes an input interface 160 configured to receive the encoded video signal, and a data decoder 170 configured to reconstruct the plurality of images of the video by decoding the encoded image data. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of blocks, and each of the plurality of blocks includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of one of the plurality of images, a boundary exists between the first tile and the second tile. The first tile and the second tile have been encoded independently relative to each other. The data decoder 170 is configured to filter the first tile using a filter or filter kernel, wherein the data decoder 170 is configured to adjust the influence range of the filter or filter kernel based on the distance between the block of the first tile to be filtered by the filter or filter kernel and the boundary between the first tile and the second tile, the block being one of the plurality of blocks of the first tile.
[0174] According to an embodiment, if the distance has a first distance value less than or equal to a threshold distance, the data decoder 170 may be configured, for example, to set the influence range of the filter or the filter kernel to a first size value. Furthermore, if the distance has a second distance value greater than the first size value and the block and its neighboring blocks belong to the same reference image, the data decoder 170 may be configured, for example, to set the influence range of the filter or the filter kernel to a second size value greater than the first size value. Additionally, if the distance has a second distance value greater than the first distance value and if the block and its neighboring blocks do not belong to the same reference image, the data decoder 170 may be configured, for example, to set the influence range of the filter or the filter kernel to the first size value.
[0175] According to an embodiment, the data decoder 170 may, for example, include a deblocking filter. For the block to be filtered of a first map block, the data decoder 170 may be configured, for example, to filter the first map block using the deblocking filter if, relative to the block to be filtered of the first map block, a second block within a second map block independently encoded from the first map block is within the filter influence range of the deblocking filter, and / or wherein, for the block to be filtered of the first map block, the data decoder (170) may be configured, for example, to set the deblocking filter strength of the deblocking filter based on whether, from the block to be filtered of the first map block, the second block within a second map block independently encoded from the first map block is within the filter influence range of the deblocking filter.
[0176] In an embodiment, the data decoder 170 may, for example, include a sample adaptive offset filter, wherein the sample adaptive offset filter may, for example, include an edge offset mode and a band offset mode. For the block to be filtered in the first tile, the data decoder 170 may, for example, be configured to activate the band offset mode and use the adaptive offset filter to filter the first tile if, from the block to be filtered in the first tile—that is, a third block within a second tile independently encoded relative to the first tile—is within the filter influence range of the sample adaptive offset filter in the edge offset mode.
[0177] According to an embodiment, the data decoder 170 may, for example, include an adaptive loop filter. For the block to be filtered within the first block, the data decoder 170 may, for example, be configured to deactivate the adaptive loop filter if, from the block to be filtered within the first block, a fourth block within a second block independently encoded relative to the first block, falls within the filter influence range of the adaptive loop filter.
[0178] Furthermore, a video encoder 101 according to an embodiment is provided for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder 101 includes a data encoder 110 configured to generate an encoded video signal including the encoded image data, wherein the data encoder 110 is configured to encode the plurality of images of the video into encoded image data, and an output interface 120 for outputting the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of blocks, wherein each of the plurality of blocks includes a plurality of samples. For a first tile and a second tile of two adjacent tiles in the plurality of tiles of a first image of the plurality of images, a boundary exists between the first tile and the second tile. The data encoder 110 is configured to encode the first tile and the second tile independently of each other. Furthermore, the data encoder 110 is configured to filter the first tile using a filter or filter kernel, wherein the data encoder 110 is configured to adjust the influence range of the filter or filter kernel based on the block to be filtered by the filter or filter kernel and the distance between the boundary between the first tile and the second tile, the block being one of a plurality of blocks of the first tile.
[0179] In an embodiment, if the distance has a first distance value less than or equal to a threshold distance, the data encoder 110 may be configured, for example, to set the influence range of the filter or the filter kernel to a first size value. The data encoder 110 may be configured, for example, to set the influence range of the filter or the filter kernel to a second size value greater than the first size value if the distance has a second distance value greater than the first distance value and if the block and its neighboring blocks belong to the same reference image. Furthermore, if the distance has a second distance value greater than the first distance value and if the block and its neighboring blocks do not belong to the same reference image, the data encoder 110 may be configured, for example, to set the influence range of the filter or the filter kernel to the first size value.
[0180] According to an embodiment, the data encoder 110 may, for example, include a deblocking filter. For the block to be filtered of a first map block, the data encoder 110 may be configured, for example, to filter the first map block using the deblocking filter if, relative to the block to be filtered of the first map block, a second block within a second map block independently encoded from the first map block is within the filter influence range of the deblocking filter, and / or wherein, for the block to be filtered of the first map block, the data decoder (170) may be configured, for example, to set the deblocking filter strength of the deblocking filter based on whether, from the block to be filtered of the first map block, the second block within a second map block independently encoded from the first map block is within the filter influence range of the deblocking filter.
[0181] In an embodiment, the data encoder 110 may, for example, include a sample adaptive offset filter, wherein the sample adaptive offset filter may, for example, include an edge offset mode and a band offset mode. For the block to be filtered in the first patch, the data encoder 110 may, for example, be configured to activate the band offset mode and use the adaptive offset filter to filter the first patch if, relative to the block to be filtered in the first patch, i.e., a third block within a second patch independently encoded from the first patch, it falls within the filter influence range of the sample adaptive offset filter in the edge offset mode.
[0182] According to an embodiment, the data encoder 110 may, for example, include an adaptive loop filter. For the block to be filtered in the first block, the data encoder 110 may be configured, for example, to deactivate the adaptive loop filter if, relative to the block to be filtered in the first block, a fourth block within a second block independently encoded from the first block is within the filter influence range of the adaptive loop filter.
[0183] Furthermore, a system according to an embodiment is provided, including the video encoder 101 and the video decoder 151 described above. The video encoder 101 is used to generate an encoded video signal. The video decoder 151 is used to decode the encoded video signal to reconstruct a video image.
[0184] When image regions are encoded independently, such as tiles in HEVC, subjective artifacts are visible. These unwanted artifacts can be mitigated by allowing loop filtering across tile boundaries. This is not a problem when the encoder and decoder can perform the same process; however, when tiles need to be independently decodeable and interchangeable, such as in the use case described above, this approach is only possible if the content is encoded using MCTS and the MV is constrained not to point to any samples affected by the filtering process. Otherwise, such features are prohibited if a region / tile boundary approach is used instead, because simply allowing loop filtering across tile boundaries will result in reconstructing incorrect sample values due to “incorrect” reference samples outside the tiles (given by the influence range of the loop filter kernel). In turn, using such contaminated samples, motion-compensated predictions will lead to further error propagation in the temporally following encoded image. In particular, the affected region expands rapidly when contaminated samples are used in the boundary-filling procedure envisioned in VVC independent tiles (VVC = Versatile Video Coding).
[0185] In 360-degree video streams, such as those using the MPEG OMAF standard (MPEG = Moving Picture Experts Group; OMAF = Omnidirectional Media Format), a particular approach to mitigating the aforementioned problem is to over-configure each individual region in the encoded image using some spare image samples, which can be omitted or mixed with spare image samples from spatially adjacent regions. However, this approach negatively impacts the sample budget available to a given level of decoder and is therefore undesirable, as the decoded samples are either discarded or mixed together after decoding.
[0186] The implementation aims to enable the loop filter to cross the boundaries of independent regions, but prevents motion-compensated predictions from using sample values influenced by potentially affected samples. One solution, similar to the encoder-side constraints used in HEVC, is to restrict motion compensation to adhere to independent patch boundaries plus an inward-pointing filter kernel, such that no reference contains contaminated samples, i.e., as shown below. Figure 4a The dashed line in the middle.
[0187] Figure 4a The contaminated sample is shown within a block from the loop filter process.
[0188] However, in VVC, an alternative approach is used to enable motion-compensated prediction within independently coded tiles, characterized by tile boundary extension. Here, sample values at tile boundaries are extrapolated perpendicularly to the tile boundaries, allowing motion estimation to reach this boundary extension. Now, when these tile boundary sample values are contaminated by the loop filter procedure, errors propagate into the boundary extension and thus into images such as... Figure 4b As shown.
[0189] Figure 4b The VVC tiling boundary extension of an independent region from a contaminated sample is shown.
[0190] Therefore, the embodiment aims to derive boundary filling not by using the tile boundary sample, which is the last sample within the tile, but by using the nearest sample within the tile that is unaffected by the loop filter procedure for vertical tile boundary expansion, also covering the sample values of contaminated samples within the tile, as follows: Figure 5 As shown.
[0191] Figure 5 The process of expanding the tile boundary according to an embodiment, which conforms to the influence range of the loop filter kernel, is illustrated.
[0192] According to the embodiment, prior art shearing for generating boundary-filled samples in the reference block is adapted. An example is given below, with equations for the horizontal and vertical components of the reference block sample positions conforming to the current VVC draft 5 specification v3 (while omitting motion vector wrapping):
[0193] yInt i =Clip3(0,picH-1,yInt) L +i-3)
[0194] xInt i =Clip3(0,picW-1,xInt) L +i-3)
[0195] It will be modified according to the embodiments to include further constants representing the influence range of the filter kernel, as follows:
[0196] yInt i =Clip3(topTileBoundaryPosition+verticalFilterKernelReachInSamples,bottomTileBoundaryPosition–1–verticalFilterKernelReachInSamples,yInt L +i-3)
[0197] xInti =Clip3(leftTileBoundaryPosition+horizontalFilterKernelReachInSamples,rightTileBoundaryPosition–1–horizontalFilterKernelReachInSamples,xInt L +i-3)
[0198] Clip3 is defined in [1] as follows:
[0199]
[0200] In an embodiment, verticalFilterKernelReachInSamples may be, for example, equal to horizontalFilterKernelReachInSamples.
[0201] In another embodiment, verticalFilterKernelReachInSamples may be different from horizontalFilterKernelReachInSamples, for example.
[0202] It is important to note that tile boundary extension is not part of the output image; it still contains contaminated samples.
[0203] As an alternative or supplement to the above concept, another embodiment is adopted. According to such an embodiment, the filterKernelReachInSamples (e.g., horizontalFilterKernelReachInSamples or verticalFilterKernelReachInSamples) of the filtering process are modified. Currently in VVC with respect to HEVC, all loop filters are either enabled or disabled across tiles. Still, to make the described region boundary extension more efficient, it is necessary to limit the number of samples affected by the filtering process. For example, the deblocking filter has two intensities, typically depending on checks such as whether the same reference image is used for two adjacent blocks. According to the embodiment, blocks at region boundaries can, for example, always be derived as blocks with this filter intensity such that a smaller number of samples are affected in the process, or the filter derivation process is independent / less dependent on the decoding background. Similarly, the influence range of each other filter (SAO and ALF) can be modified at the region boundaries (SAO = sample adaptive offset; ALF = adaptive loop filtering). Alternatively, filters can be disabled individually, rather than all at once as currently done, for example, only ALF is disabled.
[0204] According to an embodiment, the deblocking filter strength can be derived, for example, as follows: if at least one of two blocks at a block boundary belongs to a block to be independently encoded, the strength can be set to 1, for example.
[0205] In an embodiment, for SAO, if two blocks located in different tiles that are independently encoded are located on the boundary of the independent tile, then for example, the offset mode can always be used instead of the edge offset mode.
[0206] According to an embodiment, the combination of ALF, SAO, and deblocking can be disabled, for example, at individual tile boundaries.
[0207] The second aspect of the embodiments will now be described in detail below.
[0208] Specifically, the second aspect provides motion compensation predictions on the boundaries of concave patch groups (the inward-pointing boundaries of concave patch groups).
[0209] A video encoder 101 is provided according to an embodiment for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder 101 includes a data encoder 110 configured to generate an encoded video signal including the encoded image data, wherein the data encoder 110 is configured to encode the plurality of images of the video into encoded image data, and an output interface 120 for outputting the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. The data encoder 110 is configured to determine an independently encoded group of tiles, the group including three or more tiles from a plurality of tiles of a reference image of the plurality of images. Furthermore, the data encoder 110 is configured to encode the plurality of images based on reference tiles located within the reference image. Furthermore, the data encoder 110 is configured to select a location for the reference block within the reference image such that the reference block is not located partly within three of three or more tiles in the independently encoded tile group, and partly within another tile in a plurality of tiles in the reference image that do not belong to the independently encoded tile group.
[0210] In an embodiment, the three or more tiles may be arranged, for example, in a reference image such that they have an inwardly pointing boundary of a concave tile group relative to a plurality of tiles in the reference image that do not belong to an independently coded tile group.
[0211] According to an embodiment, the data encoder 110 may be configured, for example, to select a location for the reference block within a reference image such that the reference block is not located either partially within a first block of three or more blocks in the independently encoded tile group or partially within a second block of a plurality of tiles in the reference image that do not belong to the independently encoded tile group.
[0212] Furthermore, a system according to an embodiment is provided, including the aforementioned video encoder 101 and video decoder 151, for decoding an encoded video signal including encoded image data to reconstruct multiple images of a video. The video decoder 151 includes an input interface 160 configured to receive the encoded video signal, and a data decoder 170 configured to reconstruct multiple images of the video by decoding the encoded image data. The video encoder 101 is used to generate the encoded video signal. The video decoder 151 is used to decode the encoded video signal to reconstruct video images.
[0213] Furthermore, a video encoder 101 according to an embodiment is provided for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder 101 includes a data encoder 110 configured to generate an encoded video signal including the encoded image data, wherein the data encoder 110 is configured to encode the plurality of images of the video into encoded image data, and an output interface 120 for outputting the encoded image data of each of the plurality of images. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. The data encoder 110 is configured to determine an independently encoded group of tiles, the group including three or more tiles from a plurality of tiles of a reference image of the plurality of images. Furthermore, the data encoder 110 is configured to encode the plurality of images based on a reference block located within the reference image, wherein the reference block is partially located within three of the three or more tiles of the independently encoded group of tiles and is partially located within another tile of the plurality of tiles of the reference image that does not belong to the independently encoded group of tiles. Furthermore, the data encoder 110 is configured to determine a plurality of reference samples located within a portion of the reference block that does not belong to the independently encoded tile group, based on one or more of a plurality of samples of the first of the three tiles in the independently encoded tile group and based on one or more of a plurality of samples of the second tile that does not belong to the three tiles in the independently encoded tile group.
[0214] In an embodiment, the three or more tiles may be arranged, for example, in a reference image such that they have an inwardly pointing boundary of a concave tile group relative to a plurality of tiles in the reference image that do not belong to an independently coded tile group.
[0215] According to an embodiment, the data encoder 110 may be configured, for example, to determine, based on the separation of the reference block into a first sub-part and a second sub-part, a plurality of reference samples located within a portion of the reference block that does not belong to the other tile of the independently encoded tile group, such that the one or more samples of the first tile of the three tiles are used but the one or more samples of the second tile of the three tiles are not used to determine which of the plurality of reference samples located in the first sub-part, and such that the one or more samples of the second tile of the three tiles are used but the one or more samples of the first tile of the three tiles are not used to determine which of the plurality of reference samples located in the second sub-part.
[0216] In an embodiment, the separation of the portion of the reference block located within another tile that does not belong to the independently coded tile group into a first sub-part and a second sub-part can be, for example, diagonal separation of the reference block.
[0217] According to an embodiment, the data encoder 110 may be configured, for example, to determine the plurality of reference samples located within a portion of the reference block that does not belong to the other block of the independently encoded tile group by applying planar intraprediction using one or more samples of the first tile of the three tiles in the independently encoded tile group and using the one or more samples of the second tile of the three tiles or tree tiles.
[0218] Furthermore, a video decoder 151 is provided according to an embodiment for decoding an encoded video signal comprising encoded image data to reconstruct a plurality of images of a video. The video decoder 151 includes an input interface 160 configured to receive the encoded video signal, and a data decoder 170 configured to reconstruct the plurality of images of the video by decoding the encoded image data. Each of the plurality of images includes a plurality of tiles, wherein each of the plurality of tiles includes a plurality of samples. The encoded video signal includes independently encoded tile groups, the independently encoded tile groups including three or more tiles from a plurality of tiles of a reference image of the plurality of images. The data decoder 170 is configured to decode the plurality of images based on reference blocks located within the reference image, wherein the reference blocks are partially located within three of the three or more tiles of the independently encoded tile groups and partially located within another tile of the plurality of tiles of the reference image that does not belong to the independently encoded tile groups. Furthermore, the data decoder 170 is configured to determine, based on one or more samples of a first tile in the three tiles of the independently encoded tile group and one or more samples of a second tile in the three tiles of the independently encoded tile group, a plurality of reference samples located within a portion of the reference block that does not belong to the other tile in the independently encoded tile group.
[0219] In an embodiment, the three or more tiles are arranged in a reference image such that they have an inward pointing boundary relative to a plurality of tiles in the reference image that do not belong to an independently coded tile group.
[0220] According to an embodiment, the data decoder 170 may be configured, for example, to determine, based on the separation of the reference block into a first sub-part and a second sub-part, a plurality of reference samples located within a portion of the reference block that does not belong to the other tile of the independently encoded tile group, such that the one or more samples of the first tile of the three tiles are used but the one or more samples of the second tile of the three tiles are not used to determine which of the plurality of reference samples located in the first sub-part, and such that the one or more samples of the second tile of the three tiles are used but the one or more samples of the first tile of the three tiles are not used to determine which of the plurality of reference samples located in the second sub-part.
[0221] In an embodiment, the separation of a portion of a reference block located within another tile that does not belong to the independently coded tile group into a first sub-part and a second sub-part is a diagonal separation of the reference block.
[0222] According to an embodiment, the data decoder 170 may be configured, for example, to determine the plurality of reference samples located within a portion of the reference block that does not belong to the other block of the independently encoded tile group by applying in-plane predictions using one or more samples of the first tile of the three tiles in the independently encoded tile group and using the one or more samples of the second tile of the three tiles or tree tiles.
[0223] A system includes the aforementioned video encoder 101 and video decoder 151. The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is used to decode the encoded video signal and reconstruct a video image.
[0224] When tile groups are formed in raster scan order, as specified in VVC Draft 5 Specification v3, the boundary extension procedure can be performed on independently coded tile groups and needs to adapt to the following: Figure 6 The tile configuration in the image.
[0225] Figure 6 The diagram shows the division of tiles and tile groups in the encoded image.
[0226] As can be seen from the image, by clipping the sample positions, as defined in VVC Draft 5 Specification v3, it is insufficient to cover the tile boundaries, such as those marked with red crosses, when processing tile group 0.
[0227] The following section first presents what happens when the motion compensation prediction for a given reference block involves the top and left boundaries of tile 0 in tile group 0, specifically a close observation of the convex upper left boundary of tile 0 within tile group 0. According to existing techniques, three regions outside of tile 0 can be distinguished, as shown below. Figure 7 As shown.
[0228] Figure 7 A reference block with prior art boundary padding is shown.
[0229] The reference block portion located in the area marked "to be top" is filled with the vertical extrapolation of the corresponding top sample row of tile 0, while the reference block portion located in the area marked "to be left" is filled with the horizontal extrapolation of the corresponding left sample column of tile 0. As for the reference block portion located in the area marked "to be top-left," it is filled with a single sample value, that is, sampled at position 0,0 of tile 0, which is the top-left sample of tile 0.
[0230] Another example is... Figure 8 As shown, boundary extension will be applied to the L-shaped region. In this case, for the lower right corner of tile 0, i.e. Figure 8 The concave boundary of tile group 0 shown also needs to be extrapolated for motion compensation prediction. However, for convex tile group boundaries, similar to existing techniques for image boundaries, sample value extrapolation cannot be easily accomplished because the vertical extrapolation of the tile boundary in concave tile group boundaries produces two possible values at each sample location during boundary expansion.
[0231] Therefore, the embodiments involve limiting motion compensation prediction in the bitstream on the encoder side and disallowing a reference block that simultaneously contains samples from both patch 1 and patch 2, such as... Figure 8 The first reference block shown.
[0232] Figure 8 The boundary of a diagonally segmented concave block group according to an embodiment is shown.
[0233] Entry is permitted only if a sample from tile 1 or a sample from tile 2 is exclusively located within the reference tile. Figure 8 The boundary extension area of tile group 0 shown in the figure, as Figure 8 The second reference block is shown in the example. In this case, the vertical boundary fill of the rule is applied with respect to the boundaries of the adjacent tiles involved, i.e., tile 1 in the example.
[0234] As an alternative to the aforementioned bitstream constraints, the solution is to divide the boundary padding region diagonally, such as... Figure 8As shown by the dashed lines, the boundary filling region is filled in two parts by vertically extrapolating the sampled values from the boundaries of their respective adjacent tiles to the diagonal boundaries. In this alternative, a reference block containing samples from tiles 1 and 2 is also permitted, namely, an exemplary first reference block. In a further alternative to the diagonal division of the boundary filling region, the entire region is filled with sample values from tiles 1 and 2 according to an in-plane prediction pattern.
[0235] In a further alternative embodiment, if such an independent region exists, the motion compensation prediction from the region needs to be restricted to not pointing outside that region, unless it only crosses one boundary, i.e., neither the first reference block nor the second reference block is allowed.
[0236] The third aspect of the invention will now be described in detail below.
[0237] Specifically, the third aspect provides the decoded image hash for GDR.
[0238] A video encoder 101 is provided according to an embodiment for encoding a plurality of images of a video by generating an encoded video signal. Each of the plurality of images includes raw image data. The video encoder 101 includes a data encoder 110 configured to generate an encoded video signal containing the encoded image data, wherein the data encoder 110 is configured to encode the plurality of images of the video into encoded image data, and an output interface 120 for outputting the encoded image data of each of the plurality of images. The data encoder 110 is configured to encode hash information within the encoded video signal. Furthermore, the data encoder 110 is configured to generate a hash information image based on a current portion of a current image of the plurality of images, but not depending on subsequent portions of the current image, wherein the current portion has a first position in the current image, and subsequent portions have a second position within the image different from the first position.
[0239] According to an embodiment, the current image comprises multiple parts, the current part being one of the multiple parts, and a subsequent part being another of the multiple parts, each of which has a different position within the image. The data encoder 110 may, for example, be configured to encode the multiple parts within the encoded video signal in an encoding order, wherein the subsequent part immediately follows the current part in the encoding order, or wherein the subsequent part is interleaved with the current part and partially succeeds the current part in the encoding order.
[0240] In an embodiment, the data encoder 110 may be configured, for example, to encode the hash information that depends on the current portion.
[0241] According to an embodiment, the data encoder 110 may be configured, for example, to generate the hash information such that the hash information depends on the current portion and not on any other portion of the plurality of portions of the current image.
[0242] In an embodiment, the data encoder 110 may be configured, for example, to generate the hash information such that the hash information depends on the current portion and such that the hash information depends on one or more other portions of the plurality of portions preceding the current portion in the encoding order, but such that the hash information does not depend on any other portion of the plurality of portions of the current image following the current portion in the encoding order.
[0243] According to an embodiment, the data encoder 110 may be configured, for example, to generate the hash information such that the hash information depends on the current portion and the current portion of the current image, wherein the hash information depends on each other portion of the plurality of portions of the current image, the other portions preceding the current portion in the encoding order, or the other portions are interleaved with the current portion of the current image and partially preceding the current portion of the current image in the encoding order.
[0244] In an embodiment, the data encoder 110 may be configured, for example, to encode the current image among a plurality of images, such that the encoded video signal may be decoded, for example, by employing progressive decoding refresh.
[0245] For example, in one embodiment, the hash information may depend, for instance, on the refreshed area of the current image that is refreshed using progressive decoding refresh, rather than on another area of the current image that is not refreshed by progressive decoding refresh.
[0246] According to an embodiment, the data encoder 110 may be configured, for example, to generate hash information such that the hash information indicates one or more hash values based on the current portion of the current image rather than on the subsequent portions of the current image.
[0247] In an embodiment, the data encoder 110 may be configured, for example, to generate each of the one or more hash values based on a plurality of luminance samples of the current portion and / or a plurality of chrominance samples depending on the current portion.
[0248] According to an embodiment, the data encoder 110 may be configured, for example, to generate each of the one or more hash values as a message digest algorithm 5 value, or as a cyclic redundancy check value, or as a checksum, depending on the plurality of luminance samples of the current portion and / or depending on the plurality of chrominance samples of the current portion.
[0249] A video decoder 151 is provided according to an embodiment for decoding an encoded video signal including encoded image data to reconstruct images of a plurality of videos. The video decoder 151 includes an input interface 160 configured to receive the encoded video signal, and a data decoder 170 configured to reconstruct the plurality of images of the video by decoding the encoded image data. The data decoder 170 is configured to analyze encoded hash information within the encoded video signal, wherein the hash information depends on a current portion of a current image among the plurality of images, but not on subsequent portions of the current image, wherein the current portion has a first position within the current image, and subsequent portions have a second position within the image different from the first position.
[0250] According to an embodiment, the current image comprises multiple parts, the current part being one of the multiple parts, and a subsequent part being another of the multiple parts, each of which has a different position within the image. The multiple parts are encoded in an encoded video signal in an encoded order, wherein the subsequent part immediately follows the current part in the encoding order, or wherein the subsequent part is interleaved with the current part and partially follows the current part in the encoding order.
[0251] In an embodiment, the hash information, depending on the current portion, may be encoded, for example.
[0252] According to an embodiment, the hash information depends on the current portion, but not on any other portion of the plurality of portions of the current image.
[0253] In an embodiment, the hash information depends on the current portion and is such that the hash information depends on one or more other portions of the plurality of portions preceding the current portion in the encoding order, but is such that the hash information does not depend on any other portion of the plurality of portions of the current image following the current portion in the encoding order.
[0254] According to an embodiment, the hash information depends on the current portion, and wherein the hash information depends on the current portion of the current image, and wherein the hash information depends on each other portion of the plurality of portions of the current image, the other portions being located before the current portion in the encoding order, or the other portions being interleaved with the current portion of the current image and partially preceding the current portion of the current image in the encoding order.
[0255] In an embodiment, the data decoder 170 may be configured, for example, to use progressive decoding refresh to decode the encoded video signal to reconstruct the current image.
[0256] For example, in an embodiment, the hash information may depend, for instance, on the refresh area of the current image that is refreshed using progressive decoding refresh, rather than on another area of the current image that is not refreshed by progressive decoding refresh.
[0257] According to an embodiment, hash information indicates one or more values that depend on the current portion of the current image but not on the subsequent portions of the current image.
[0258] In an embodiment, each of the one or more hash values depends on a plurality of luminance samples and / or a plurality of chrominance samples of the current portion.
[0259] According to an embodiment, each of the one or more hash values is a Message Digest Algorithm 5 value, a Cyclic Redundancy Check (CDR) value, or a checksum that depends on multiple luminance samples and / or multiple chrominance samples of the current portion.
[0260] Furthermore, a system including the video encoder 101 and the video decoder 151 described in the embodiments is provided. The video encoder 101 is configured to generate an encoded video signal. The video decoder 151 is configured to decode the encoded video signal to reconstruct a video image.
[0261] A key tool for implementers is the control points carried in the encoded video bitstream to verify the correct operation of the decoder and the integrity of the bitstream. For example, there are methods that carry hashes of simple checksums of decoded sample values, such as MD5, CRC, or images, in SEI messages (SEI = supplemental enhancement information; MD5 = message-digest algorithm 5; CRC = cyclic redundancy check), with each SEI message associated with a single image. Therefore, it is possible to verify on the decoder side whether the decoded output matches the encoder's intended output without accessing the original material or the encoder. Mismatches can identify problems in the decoder implementation (during development) or bitstream corruption (in service) in case the decoder implementation has already been verified. For example, this can be used for error detection in systems, such as in session scenarios using RTP-based communication channels where the client must actively request the prediction chain to reset the IDR images to resolve decoding errors caused by bitstream corruption.
[0262] However, in GDR scenarios, the existing SEI message mechanism is insufficient to allow meaningful detection of decoder or bitstream integrity. This is because in GDR scenarios, an image can be considered correctly decoded even when only a portion is correctly reconstructed, i.e., the region whose temporal prediction chain was recently refreshed by scanning an inner column or row. Consider random access to a stream using a GDR-based encoding scheme. From the point the stream is decoded, many (partially) incorrect images will be decoded until the first fully correct image is decoded and can be displayed. Now, while the client can identify the first fully correct decoded image from the matching existing decoded image hash SEI message, the client will not be able to check the decoder or bitstream integrity within the decoded image until the first fully correct decoded image. Therefore, it is possible that bitstream corruption or decoder implementation errors prevent the client from ever obtaining a fully correct decoded image. Consequently, the point at which the client can take aversive measures (e.g., requesting different types of random access from the sending side to ensure bitstream integrity) will be significantly later than the present invention, which allows per-image detection for these situations. An example of this unfavorable existing detection mechanism is waiting for a predefined time threshold, such as multiple GDR cycles (which can be derived from the bitstream signaling) before identifying the bitstream as corrupted. Therefore, in GDR scenarios, it is crucial to distinguish encoder / decoder mismatches based on regions.
[0263] The embodiment provides bitstream signaling of the decoded image hash for the GDR refresh area so that the client can benefit from it as described above. In one embodiment, the area is defined in the SEI message by the location of the luminance sample in the encoded image, and a corresponding hash is provided for each area.
[0264] In another embodiment, a single hash for the GDR refresh region is provided in the bitstream, for example, via an SEI message. The region associated with the hash in a specific access unit (i.e., an image) changes over time and corresponds to all blocks whose prediction chains have been reset from the scanned inner columns or rows in the GDR cycle. That is, for the first image in the GDR period, for example, if there is one inner block column on the left side of the image, the region contains only the inner blocks of that specific column, while for each subsequent image, the region contains the inner block column and all blocks to the left, assuming the inner refresh columns are scanned from left to right, until the final image of the GDR cycle no longer contains unrefreshed blocks, and the hash is derived using all blocks of the encoded image. Exemplary syntax and semantics of such a message are given below.
[0265] Decoded GDR refresh zone hash SEI message syntax
[0266]
[0267] This message provides a hash value for the refresh area of each color component of the currently decoded image.
[0268] Note 1—The decoded GDR refresh area hash SEI message is a suffix SEI message and cannot be included in scalable nested SEI messages.
[0269] Before calculating the hash, the GDR refresh area of the decoded image data is arranged into one or three byte strings of length dataLen[cIdx] called pictureData[cIdx] by sequentially copying the corresponding sample value of each decoded image component into the byte string of pictureData.
[0270] The syntax elements hash_type, picture_md5[cIdx][i], picture_crc[cIdx], and picture_checksum[cIdx] are essentially equivalent to the content defined in HEVC for decoding the image hash SEI message, that is, deriving the hash from the data in the pictureData array.
[0271] When the above concept is combined with the first aspect of the present invention, namely the boundary expansion mechanism that complies with the influence range of the loop filter kernel, the region used for hash calculation also omits samples that may be contaminated by the influence range of the loop filter kernel.
[0272] In this embodiment, hash information can be sent at the beginning or end of the image.
[0273] Although certain aspects have already been described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent a description of corresponding blocks or items or features of the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices, such as microprocessors, programmable computers, or electronic circuits. In some embodiments, one or more of the most important method steps may be performed by such devices.
[0274] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware, software, or at least partially in hardware or at least partially in software. This implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, having stored electronically readable control signals thereon that cooperate (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium may be computer-readable.
[0275] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0276] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0277] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0278] In other words, embodiments of the method of the present invention are therefore computer programs having program code that, when run on a computer, performs one of the methods described herein.
[0279] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is generally tangible and / or non-transitory.
[0280] Therefore, a further embodiment of the method of the present invention is a data stream or signal sequence, which represents a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet.
[0281] Further embodiments include processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.
[0282] Further embodiments include a computer on which a computer program for performing one of the methods described herein is installed.
[0283] Further embodiments of the invention include an apparatus or system configured to transmit (e.g., electronically or optically) to a receiver a computer program for performing one of the methods described herein. For example, the receiver may be a computer, mobile device, storage device, etc. For example, the apparatus or system may include a file server for transmitting the computer program to the receiver.
[0284] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0285] The apparatus described herein can be implemented using hardware devices, computers, or a combination of hardware devices and computers.
[0286] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.
[0287] The above embodiments are merely illustrative of the principles of the invention. It should be understood that modifications and variations to the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the forthcoming patent claims, and not by the specific details presented through the description and explanation of the embodiments herein.
[0288] References
[0289] [1]ISO / IEC,ITU-T.High efficiency video coding.ITU-T RecommendationH.265|ISO / IEC23008 10(HEVC),edition 1,2013;edition 2,2014.
Claims
1. A video decoder for decoding an encoded video signal, wherein the video decoder comprises: An input interface is configured to receive the encoded video signal, the encoded video signal including a Supplemental Enhancement Information (SEI) message, the SEI message containing hash information, the hash information including a hash value of the decoded sample value of the color component of the refresh region of the progressive decoder refreshes the GDR image; as well as A data decoder is configured to reconstruct the refresh area of a GDR image by decoding the encoded video signal to obtain decoded sample values of the color components of the refresh area of the GDR image. The data decoder is configured to analyze the hash information to verify the decoded sample values of the color components of the refresh area of the GDR image.
2. The video decoder according to claim 1, The hash information includes the corresponding hash value for each color component of the refresh area of the GDR image.
3. The video decoder according to claim 2, The refresh area of a GDR image includes one luminance color component and two chrominance color components.
4. The video decoder according to claim 1, in, Analyzing the hash information encoded in the encoded video signal includes identifying corruption of the encoded video signal.
5. The video decoder according to claim 1, in, Analyzing the hash information encoded in the encoded video signal includes verifying the correct operation of the video decoder.
6. A video encoder for generating encoded video signals, the video encoder comprising: A data encoder is configured to generate the encoded video signal, the encoded video signal including the encoding of a progressive decoder refreshing a GDR image. The data encoder is configured to encode hash information into a Supplemental Enhancement Information (SEI) message within the encoded video signal. The hash information includes the hash value of the decoded sample value of the color components of the refresh area of the GDR image, and The hash information is configured for use by the decoder to verify the decoded sample values of the color components in the refresh area of the GDR image.
7. The video encoder according to claim 6, in, The hash information includes a luminance color component and a corresponding hash value for each of the two chrominance color components for the refresh area of the GDR image.
8. The video encoder according to claim 7, in, The data encoder is configured to generate each of the corresponding hash values as either a message digest algorithm 5 value, a cyclic redundancy check value, or a checksum.
9. A method for decoding an encoded video signal, the method comprising: The encoded video signal is received, the encoded video signal including a Supplemental Enhancement Information (SEI) message, the SEI message containing hash information, the hash information including a hash value of the decoded sample value of the color component of the refresh area of the progressive decoder refreshes the GDR image; The refresh area of the GDR image is reconstructed by decoding the encoded video signal to obtain the decoded sample values of the color components of the refresh area of the GDR image. as well as The hash information is analyzed to verify the decoded sample values of the color components in the refresh area of the GDR image.
10. The method according to claim 9, wherein, Analyzing the hash information encoded in the encoded video signal includes identifying corruption of the encoded video signal.
11. A method for generating an encoded video signal, the method comprising: Encode the GDR image refreshed by the progressive decoder; as well as The hash information is encoded into the Supplemental Enhancement Information (SEI) message within the encoded video signal. The hash information includes the hash values of decoded sample values of the color components for the refresh area of the GDR image. Furthermore, the hash information is configured for use by the decoder to verify the decoded sample values of the color components in the refresh area of the GDR image.
12. A non-transitory digital storage medium having a computer program stored thereon, the computer program being executed by a processor to perform the following steps: Receive an encoded video signal, the encoded video signal including a Supplemental Enhancement Information (SEI) message, the SEI message containing hash information, the hash information including a hash value of the decoded sample value of the color component of the refresh area of the progressive decoder refreshes the GDR image; The refresh area of the GDR image is reconstructed by decoding the encoded video signal to obtain the decoded sample values of the color components of the refresh area of the GDR image, and The hash information is analyzed to verify the decoded sample values of the color components in the refresh area of the GDR image.
Citation Information
Patent Citations
Gradual decoding refresh with temporal scalability support in video coding
CN104904216A
Hash table construction and availability checking for hash-based block matching
CN105393537A