Coding-loop-external smoothing of a boundary between two image regions
Patent Information
- Application Number
- EP2023757671
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-25
- Filing Date
- 2023-08-08
- Publication Date
- 2025-07-02
Smart Images

Figure 1.1
Abstract
Description
Description Title: Off-loop smoothing of a boundary between two image areas technical field
[0001] This disclosure falls within the domain of video compression.
[0002] More specifically, this disclosure relates to a process for processing at least one area of decoded image, a computer program, a recording medium, a digital signal, and a data processing circuit. Previous technique
[0003] Standardized video compression schemes have been based on the same principles since the first generation of MPEG standards, MPEG-2. In chronological order, the following standards are H.264 / AVC (2003), HEVC (2013), and VVC (2020). The AOM, VP9, and AV1 encoding schemes also follow the same concepts.
[0004] A video sequence to be encoded is divided into frames. Each frame is divided into fixed-size blocks, which can themselves be further subdivided. For a given frame, an encoder processes the blocks sequentially, from the top-left block to the bottom-right block. The encoder generates an output binary signal containing, for each frame, the result of the sequential processing of its constituent blocks.
[0005] The binary signal containing the compressed video sequence can then be broadcast and processed by a decoder whose operation, modeled on that of the encoder, considers the blocks sequentially in order to reconstruct the initial video sequence.
[0006] Reference is now made to Figure 1, which represents an example of an HEVC encoder built around a coding loop configured to perform different processing of a block of a source video sequence (100) supplied as input to the encoder.
[0007] One of these processes is a prediction of the provided block, using information that has already been encoded and decoded. The first image, called "Intra," is encoded using a spatial prediction (1 18) using only pixels reconstructed in the neighborhood of the block being processed. Subsequent images, called "Inter," can use a spatial prediction and, in addition, a temporal prediction (1 16) that exploits the previously encoded images using motion compensation (1 14) indicated by a motion vector, which generally allows for very efficient prediction. The images thus encoded and then decoded, and used for encoding future images, are grouped within a memory called the "Decoded Picture Buffer" (DPB) (1 12).
[0008] Another processing step performed within the coding loop is the encoding of the difference between the prediction result and the input block, or "pixel remnants." This encoding is performed after a transformation and quantization step (104). The quantization step is carried out for a given quantization parameter (QP) associated with each block and signaled in the bitstream. The QP represents a trade-off between the desired image quality after decoding and the desired degree of video compression. The higher the QP value, the less information there is about the pixel remnants in the encoded video sequence, and the higher the degree of video compression. Conversely, the lower the QP value, the more information there is about the pixel remnants, and the better the reconstruction quality at the decoder receiving the encoded video sequence.A quantization and inverse transformation step allows the pixel residues to be reconstructed.
[0009] Other processing steps performed within the coding loop involve successive filtering of the block being processed by different filters. The HEVC standard provides two filters named "Sample Adaptive Offset" (SAQ) (108) and "Deblocking Filter" (110), which can be translated into French as "décalage adaptatif d'exemplaires" and "filtre anti-blocs." These filters modify the reconstructed pixels of the block being processed without impacting the prediction of neighboring blocks within the same image, but they do impact the prediction of future blocks within subsequent images, since the images in the DPB are those post- filtering. In addition to these two filters, the VVC standard introduced an additional filter called "Adaptive Loop Filter" (ALF).
[0010] As already explained, the encoder thus generates, at output, a binary signal (124) comprising, for each image, the result of the sequential processing of the blocks that compose it.
[0011] Across these processing methods, various high-level image segmentation techniques have been introduced into standards to address different applications: for example, "Slices" and "Tiles" according to the HEVC standard, and sub-images or "SubPictures" according to the VVC standard. These segmentation techniques are described in nplcitl. Figure 2 shows an example of segmenting an image (200) according to the VVC standard into Tiles (delimited by thick solid lines) and Slices (delimited by thick dashed lines). In this example, the image is also divisible into blocks or "CTUs," marked with thin lines.
[0012] One of the main uses of high-level image slicing concerns applications involving a large number of pixels to be encoded. Examples include encoding video sequences with high image resolution (4K, 8K or 16K for example), video sequences with high frame rates (greater than 60 fps for example) or 360° video sequences, such as those used for virtual reality applications.
[0013] High-level image segmentation, such as into tiles, is used in HEVC to enable parallel processing across multiple encoding cores. This is achieved, for example, by processing one tile per encoding core, thus meeting the high computational demands with limited or even no data sharing between encoding cores. This parallel processing can utilize multiple threads or multiple cores of one or more data processing circuits, such as CPUs, ASICs, or FPGAs.
[0014] Generally, the compressed video stream is decoded by a decoder acting on a single decoding core and having no knowledge of the Parallelism is implemented during encoding. Decoding across multiple decoding cores is also possible.
[0015] In the HEVC standard, it is possible to implement processing within the encoding loop to improve visual quality between neighboring tiles. This involves sharing information between encoding cores for pixels at tile boundaries and specific boundary processing within each encoding core. Regarding filters, such as SAO and anti-blocking filters, normative parameters can be included in the encoded video sequence to indicate whether these filters have been implemented in the encoding loop to encode a given slice or tile. In the HEVC standard, these parameters are named "loop_filter_across_slice" and "loop_filter_across_tile". The concept of Motion Constrained Tile Set (MCTS), which relates to these specific modifications, is described in nplcit2.The MCTS is a set of measures taken at the encoder level to make the encoding / decoding of each tile independent of the encoding / decoding of other tiles. It thus enables parallelization, at least, of the prediction and reconstruction processes. Additionally, a parameter named "loop_filter_across_tile," which can take the values "0" or "1," indicates whether the parallelization can be extended to the border filtering processes (in the case of the value "0"), or whether, on the contrary, the border filtering processes require data related to the encoding / decoding of adjacent tiles (in the case of the value "1").
[0016] The VVC standard also introduced sub-pictures intended to replace MCTS. These sub-pictures are designed to be processed independently by their respective encoding cores. Several contributions, such as nplcit3, were made by standardization stakeholders to arrive at this design in the VVC standard.
[0017] Regardless of the high-level slicing method used, it is desirable to be able to smooth, at the decoder level, the boundaries between tiles, slices, or sub-images that have been processed by different encoding cores. Indeed, a lack of smoothing results in visually jarring boundary visibility after decoding.
[0018] With the current architecture of video compression standards, it is only possible to apply edge smoothing at the decoder level without risk of distortion if this smoothing has previously been applied at the encoder level. Therefore, the VVC standard requires that when edge smoothing is applied at the encoder level, a parameter called "sps_loop_filter_across_subpic_enabled_flag" signals in the bitstream that this edge smoothing should be reapplied at the decoder level.
[0019] Furthermore, when considering a boundary between two neighboring Tiles, Slices or sub-images whose encoding is carried out, for one, by a first encoding core and, for the other, by a second distinct encoding core, a smoothing of the boundary can only be implemented at the encoder level if pixels are transferred either between the first and second encoding cores, or from the first and second encoding cores to a third encoding core dedicated to the implementation of the smoothing of the boundary.
[0020] It follows from the above that, with the current architecture of video compression standards and in the case where encoding is implemented by parallelization on several encoding cores, applying edge smoothing induces significant constraints at the encoder level.
[0021] To avoid these constraints, it may be possible to indicate in the binary stream that boundary smoothing is to be performed at the decoder level, without however applying this smoothing at the encoder level.
[0022] Such processing is asymmetric, because it involves applying an anti-blocking filter at the decoder level to the boundaries of the sub-images, without a corresponding anti-blocking filter having also been applied to the boundaries at the encoder level to encode these same sub-images.
[0023] Such asymmetric processing between the encoder and the decoder induces drift, which is initially small and limited to the pixels on either side of the boundary between two sub-images. The more successive inter-predictions occur, the greater the drift becomes. This drift can be minimized by disabling the activation of SAO and ALF filters at both the encoder and decoder levels. CTU at the borders. Nevertheless, even with such provisions, such drift can generate significant visual artifacts upon decoding.
[0024] To illustrate this, a sequence of 60 images was encoded using the following procedure. Each image was first divided into two halves separated by a vertical boundary. The left image halves were encoded by a first encoding core, and in parallel, the right image halves were encoded by a separate second encoding core. No communication was established between the two encoding cores, and no block filter was implemented at the encoder level. Figure 3 shows the twentieth (300), fortieth (302), and sixtieth (304) images of the sequence, as obtained after decoding, aggregation of the decoded image halves, and implementation of a block filter at the decoder level only, within the decoding loop. The drift appears very significant from the fortieth image onwards and then continues to increase until it affects approximately three-quarters of the sixtieth image.
[0025] To preserve processing at the encoder and decoder, a solution has also been proposed in nplcit4 to allow filtering without impacting decoding due to the expectations of previous results. However, this approach is computationally expensive.
[0026] Finally, work related to nplcit5 and nplcitô was carried out within the framework of the JPEG 2000 image compression standard. The authors introduced the concept of "detiling," which involves, at the encoder level, decomposing the digital signals corresponding to groupings (into Tiles) of CTUs to be encoded into wavelets. The authors also plan to apply filtering at the encoder level in the wavelet-transformed domain. Summary
[0027] This disclosure improves the situation.
[0028] A method for processing at least one area of a decoded image is proposed, the method comprising: - at the output of a decoding loop that has decoded at least one area of the current image, processing of the decoded area of the current image using a module of boundary smoothing using metadata relating to at least one boundary between the current image area and a neighboring image area.
[0029] The term "image area" refers to any area, delimited by a closed line, within an image. When two closed lines, each delimiting an image area, share a common portion, this common portion forms a boundary between these image areas, which are then said to be adjacent.
[0030] By "decoding loop" we mean a set of logical instructions that allows, at a minimum, the following: - to receive, as input, an extract of the digital signal relating to the current image area, in encoded form, - decode the digital signal extract from, at least, information already decoded by the decoding loop, this already decoded information being, for example, related to one or more other areas of the same image and / or to one or more other images, - to provide, as output, the digital signal extract in decoded form, also called the "decoded current image area", and - to memorize the current decoded image area as already decoded information usable to predict image areas relative to future images.
[0031] By "metadata," we mean the digital data associated with the boundary between the current image area and a neighboring image area. This metadata is used, at least, for the processing implemented by the boundary smoothing module.
[0032] It is understood that, according to the proposed method, the processing using the edge smoothing module is performed outside the decoding loop, and that, as such, the result of this processing does not impact any future decoding of image areas by the decoding loop. Thus, the proposed method allows for post-decoding edge smoothing while maintaining symmetrical processing at the encoder and decoder levels. The proposed method therefore provides less edge visibility without generating visual artifacts, resulting in increased viewing comfort for the viewer.
[0033] Furthermore, the implementation of the proposed process at the output of the decoding loop requires no modification of existing algorithms or encoding devices. In particular, the proposed process is fully compatible with encoders implementing parallelization techniques using a plurality of encoding cores, with each encoding core responsible for encoding one image region from a set of image regions resulting from a high-level partitioning of the image.
[0034] In some examples, the current decoded image area includes an edge region corresponding to the boundary and a region far from the boundary, and the processing of the current decoded image area includes processing of the edge region and does not include processing of the region far from the boundary.
[0035] Thus, the processing of an image obtained by aggregating image regions from the decoder output can be limited to regions located at the boundaries between image regions, without impacting the image as a whole. In other words, it is possible to implement differentiated processing based on image region areas.
[0036] In some examples, the process further includes control of the processing of the current image area using a controller that uses first data relating to a decoding of the current image area and second data relating to a decoding of a neighboring image area.
[0037] Such a controller makes it possible to refine the processing implemented by the border smoothing module and in particular to take into account possible differences between the data relating to the decoding on either side of the border, for example possible differences between the quantization parameters (QP) associated with the current image area and the neighboring image area.
[0038] A computer program is also proposed, containing instructions for implementing the above process when this program is executed by a processor.
[0039] A non-transient recording medium readable by a computer is also proposed on which a program is recorded for the implementation of the above process when this program is executed by a processor.
[0040] A digital signal is also proposed, comprising at least one encoded current image area and metadata relating to at least one boundary between the current image area and a neighboring image area.
[0041] A digital signal is a signal that can be transmitted through a communication channel, stored in memory, or read by a processor. The encoded image area can be decoded by a suitable decoder.
[0042] Metadata can be used to process the current decoded image area, at least at the aforementioned boundary. In some examples, the metadata relates to at least one aspect of a smoothing process to be applied at the boundary.
[0043] Examples of aspects of smoothing to be applied at the border level include: - an activation of the smoothing function, or - an adjustment of a smoothing force, or - a choice of a smoothing function from a set of predetermined smoothing functions.
[0044] In general, the aspects considered may concern the delimitation of one or more image regions where smoothing should be implemented and / or a way of implementing smoothing in one or more regions among the delimited regions.
[0045] A data processing circuit is also proposed, including: - an input interface configured to receive the aforementioned digital signal, and - at least one output interface configured to provide the current encoded image area as input to a decoding loop, and to provide the metadata to a processing module intended as output of the decoding loop.
[0046] The data processing circuit can either be integrated into a single device or be made up of physical modules distributed across a plurality of devices placed in a communication network.
[0047] The decoding loop may or may not be part of the data processing circuit, and the same applies to the processing module. The processing module is configured, at a minimum, to process metadata—for example, to read, store, modify, relay, delete, and so on. In some examples, the processing module is configured to use metadata as a guide for implementing a smoothing operation on a boundary between adjacent image areas within the digital signal.
[0048] The output interface can also be configured not to provide metadata to the decoding loop. Brief description of the drawings
[0049] Other features, details, and advantages will become apparent upon reading the detailed description below and analyzing the attached drawings, on which: Fig. 1
[0050] [Fig. 1] schematically represents an example of an encoder according to the HEVC standard. Fig. 2
[0051] [Fig. 2] represents an example of image cropping according to the VVC standard. Fig. 3
[0052] [Fig. 3] represents three images in a sequence of decoded images, with an implementation of an anti-block filter at the decoder level only, in the decoding loop. Fig. 4
[0053] [Fig. 4] represents an encoding of an image area according to an example of implementation. Fig. 5
[0054] [Fig. 5] represents a decoding of an area of an encoded image according to an example of an embodiment. Fig. 6
[0055] [Fig. 6] represents an encoder comprising four independent encoding cores, according to an example embodiment. Fig. 7
[0056] [Fig. 7] represents a decoder configured to decode four image areas encoded according to an example embodiment. Description of the implementation methods
[0057] The proposed technique aims to resolve the problems described above by introducing normative smoothing of image area boundaries resulting from high-level image partitioning. This smoothing, like the classic filtering used in conventional video coding schemes, aims to attenuate potentially visible boundaries between image areas such as tiles, slices, or sub-images.
[0058] The smoothing achieved by the proposed technique is unique in that it is performed outside the encoding loop. Image areas stored in the DPB (Data Processing Box) and used as references for future image areas do not benefit from the proposed smoothing and the resulting visual attenuation of the boundary. This allows the smoothing of image area boundaries to be applied only on the decoder side. Since the reconstruction operation on the encoder side is limited to generating the same images that will be stored in the DPB on the decoder side, the smoothing of sub-image boundaries does not occur there.
[0059] A particular example of implementation is now described with reference to figure 4, which schematically represents an encoder.
[0060] A video sequence to be encoded (400) is provided as a digital input signal to the encoder. The video sequence to be encoded includes a A plurality of images. These images are assumed to have undergone high-level partitioning (not shown) into image areas. The encoder is configured to process the image areas contained in the video sequence sequentially. It can be configured that for a given image, the encoder will first process, for example, an image area located in the top left corner of the image, then a second image area, for example, adjacent to the first, and so on until it processes a final image area located, for example, in the bottom right corner of the image.
[0061] A particular feature of the proposed technique is that the video sequence to be encoded also includes, in addition to the image areas to be encoded, metadata, and / or that this metadata is made accessible by a separate digital signal associated with the video sequence to be encoded.
[0062] Metadata can relate to the entirety of the boundaries between image areas within a given image or several consecutive images. Alternatively, metadata can relate to one or more specific boundaries. Specific examples of metadata are presented later in relation to their intended use.
[0063] In the following description, we focus on the processing of a current image area supplied as input to the encoder. We also assume that the metadata is relative to at least one boundary between the current image area and at least one neighboring image area.
[0064] The processing of a current image region supplied as input to the encoder involves predicting the current image region based on previously encoded, decoded, and stored image regions (408). More precisely, the prediction of the current image region can be an intra-image prediction (414) implemented from one or more previously processed regions of the same image. The prediction of the current image region can also be an inter-image prediction (412) implemented based on an estimation (410) of a motion vector, itself implemented from regions of one or more previously processed images. Alternatively, it is also possible to implement in parallel an intra-image prediction (414) and a inter-prediction (412) as described previously and to provide the results of these predictions to a decision module (416). The prediction of the current image area is then a prediction according to a prediction mode chosen by the decision module (416).
[0065] The result of the prediction of the current image area is then compared to the current image area provided as input to the encoder, and the difference, called pixel residuals, undergoes a transformation and quantization (402). The current image area is then ready to be encoded. For this, an entropy encoder (418), i.e., a lossless encoder, for example of the CABAC or CAVLC type, is fed with all the information necessary to perform the encoding, namely: - an indication of the prediction method used to predict the current image area (i.e., intra, inter, or a prediction method chosen by the decision module), - prediction information according to the chosen mode, for example a motion vector in the case of an inter-prediction mode or an intra-prediction function in the case of the intra-mode, and - the result of the transformation and quantification of pixel residuals.
[0066] The same information is used to decode the current image area. Inverse quantization followed by inverse transformation (404) of the transformed pixel remnants is performed, allowing the pre-transformation pixel remnants to be reconstructed. The reconstructed pixel remnants and the result of the current image area prediction are then summed to reconstruct the current image area. Various filtering operations (406) can optionally be implemented. Finally, the reconstructed and optionally filtered current image area is stored in memory (408) and can then be used for future inter-image predictions when processing future image areas.
[0067] As described, the processing of the current image area includes the implementation of a coding loop defined as a sequence of processing steps, based on the processing of at least one previous image area and providing input for the processing of at least one area of the next image. Specifically, the current image area is first predicted based on previously processed image areas; therefore, processing of the current image area relies on the processing of at least one previous image area. From this prediction, the information necessary for encoding the current image area is determined. This information is used to determine the result of decoding the encoded current image area. Finally, this result is stored in memory for processing one or more subsequent image areas; thus, processing of the current image area provides input for the processing of at least one subsequent image area.
[0068] After processing the current image area, the entropy encoder (418) outputs a binary signal (420) containing the current image area in encoded form. More generally, after processing the video sequence (400) supplied as input to the encoder, the binary signal (420) output from the entropy encoder contains, in encoded form, all areas of all images in the video sequence (400).
[0069] A distinctive feature of the proposed technique is that the binary signal (420) is enriched with the aforementioned metadata, which, as specified, relates to at least one boundary between the current image area and at least one neighboring image area. In any case, this metadata is not processed by the coding loop; that is, it is not used within the coding loop in any of the processing steps that serve as the basis for determining the current encoded image area or any other encoded image area. Optionally, this metadata is not made accessible to the coding loop and is provided only to a post-processing module (not shown) which incorporates or joins it to the binary signal from the entropy encoder (418).
[0070] According to the proposed technique, this metadata serves as a signal from the encoder to a smoothing module located outside the decoding loop, in order to smooth the boundary between two adjacent image areas. The smoothing module is described later.
[0071] Referring to existing video coding standards, any signaling from an encoder-side entity to a decoder-side entity must be done via a specific message format carried by the binary signal containing the encoded video sequence. These messages are called "supplemental enhancement information" or SEI. The decoder's use of SEI is optional and intended to provide new functionalities conveyed by this information. It is possible to edit the standards' SEI sections after the fact. For example, the standard for VVC includes a separate document, nplcitô, relating to the SEI section of the standard.
[0072] By not being limited to existing video coding standards, signaling possibilities are expanded and are not restricted to SEIs (Signal Encoding Indicators). Signaling can thus be implemented at the image sequence level via one or more headers associated with the sequence, such as SPS, PPS, or VUI. Signaling can also be implemented at the image level via one or more headers associated with the image, such as "slice header" or "picture header." In certain examples where coding information is made accessible to the smoothing module, signaling can also be implemented at the level of image regions located on either side of a boundary between adjacent image areas.
[0073] Reference is now made, in an example of implementation, to figure 5 which schematically represents a decoder adapted to decode the aforementioned binary signal (420).
[0074] The decoder has an input interface configured to receive the binary signal. The encoded images, and more precisely the encoded image regions contained in the binary signal, are provided as input to an entropy decoder (502) and successively processed by the entropy decoder. For example, the implementation, by the entropy decoder, of the processing of the current encoded image region generates, at the output of the entropy decoder, two types of information, namely: the pixel remnants as they result from the transformation and quantization (402) at the encoder level, and the prediction method and prediction information used at the encoder level for predicting the current image area.
[0075] This information is processed by a decoding loop that functions similarly to the aforementioned encoding loop, relying on one or more previously decoded image areas to decode the current image area. After decoding, the current image area is stored in memory (408) and is in turn used to decode one or more subsequent image areas. The image areas stored in memory (408) form a sequence of decoded image areas that can be aggregated, frame by frame, to form a sequence of images—that is, a decoded video sequence.
[0076] A distinctive feature of the proposed technique is that the current decoded image area is transmitted to a boundary smoothing module (518), hereafter simply referred to as the "smoothing module," located outside the decoding loop. The smoothing module also receives metadata contained in the binary signal (420) and / or associated with the binary signal. This metadata can relate to any aspect of the processing implemented by the smoothing module. For example, the metadata can indicate whether smoothing is activated for all boundaries between image areas or, conversely, only for certain specific boundaries. In the latter case, the smoothing module implements boundary smoothing limited to those specific boundaries. For instance, the metadata can indicate a smoothing strength that is either common to all boundaries between image areas or differentiated by boundary.The smoothing module then implements the smoothing for each respective boundary, according to the smoothing strength indicated. For example, the metadata can indicate a smoothing function to be implemented from a set of predefined smoothing functions. A Gaussian filter is an example of a filter that can implement a smoothing function. The metadata can thus indicate, for one or more given boundaries, a particular smoothing function to be implemented and / or a specific parameter value used in the function's definition. smoothing. Considering as a smoothing function a function implemented by a Gaussian filter, the number of taps is an example of a suitable parameter.
[0077] The smoothing module (518) uses metadata relating to the smoothing of the boundary between the current image area and a neighboring image area to implement smoothing of this boundary. The smoothing module thus provides, as output, at least one decoded and smoothed current image area.
[0078] In the example shown, the smoothing module receives the decoded video sequence, which consists of a sequence of decoded images, one of which contains the current decoded image area. The smoothing module uses the metadata to process the video sequence and thus outputs a modified decoded video sequence (520), at least in that the current decoded image area is also smoothed.
[0079] Another distinctive feature of the proposed technique is that the smoothing result of one or more areas of one or more images, obtained from the smoothing module (518), is not fed into the decoding loop and is not used to decode subsequent image areas. The proposed technique thus achieves a visual rendering after decoding without visible boundaries that are bothersome to end users, while avoiding any transfer of pixels or other encoding information between sub-images on the encoder side. This is particularly beneficial in encoding schemes that utilize separate encoding cores, each processing a sub-image in parallel. Indeed, limitations in terms of connectors make data transfers between encoding cores expensive with these architectures.
[0080] As an illustration, Figure 6 shows an example of encoding an image contained in a video sequence (400) to be encoded, the image being divided into four sub-images (600, 602, 604, 606). The video sequence (400) to be encoded is associated with, or includes, metadata relating to at least one boundary between a given sub-image (600) and a neighboring sub-image (602). The image encoding is implemented in parallel by four distinct encoding cores (608, 610, 612, 614). The processing of a given image area (600) by a given encoding core (608) involves the implementation of a coding loop specific to that encoding core (608). This means, in particular, that each Each encoding core (608) has its own storage memory, configured to store the image areas encoded at that encoding core (608). Each encoding core operates independently, without any data sharing between encoding cores. The outputs of the four encoding cores are concatenated (616) to generate the binary signal (420) comprising the four sub-images in encoded form. This binary signal is also enriched with metadata associated with, or contained within, the video sequence (400) to be encoded. Unlike a method that would implement smoothing in the decoding loop on the decoder side but without implementing smoothing on the encoder side, the proposed technique has the advantage of being completely standardized and avoiding deviations such as those illustrated in Figure 3.
[0081] Reference is now made to Figure 7, which illustrates an example of a decoder adapted to decode encoded image regions (700, 702, 704, 706) of the same image contained in a binary signal (420). This decoding can be implemented either in parallel by a plurality of decoding cores (708, 710, 712, 714) or sequentially by a single decoding core.
[0082] In this example, the decoded image areas are aggregated by an aggregator (716) in order to reconstruct, in decoded form, the image contained in encoded form in the binary signal (420). It is after this aggregation that the boundaries of the sub-images are smoothed.
[0083] The smoothing of a boundary between two adjacent image areas located on either side of the boundary is implemented by the smoothing module (518) placed outside the decoding loop. As already described, the smoothing module relies, at least, on metadata for this purpose.
[0084] Optionally, the smoothing module (518) can be provided with access to the encoding information for these neighboring image regions. This encoding information can be provided to it, for example, by the decoding core(s) that decoded these image regions, either directly or indirectly. However, this is not always possible, particularly due to system constraints. For example, one can imagine that, as with the encoder, the sub-images are each decoded on a separate decoding core, and that, once a sub- image decoded by a decoding core from the corresponding part of the binary signal, the encoding information relating to this sub-image is no longer available.
[0085] It can be anticipated that when the smoothing module (518) has access to the coding information, it will use it to control the smoothing implementation. This allows, for example, the smoothing module to implement a standardized anti-blocking filter, since, according to current standards, such an anti-blocking filter uses information such as the transform size, the coding mode (inter / intra, in particular), the QP, or the motion vectors (in the case of inter-coding) to determine the length and power of the smoothing.
[0086] The smoothing module (518) can of course be configured to implement any other smoothing function that might use all or part of this encoding information to control the smoothing implementation. A simple example of an applicable smoothing function would be a Gaussian function whose standard deviation depends on the QP of the image areas on either side of the boundary to be smoothed.
[0087] Two specific examples of achievements are now described.
[0088] In a first specific implementation example, a source image is partitioned into several sub-images, each of which is encoded. The term "sub-image" here is equivalent to the term "sub-picture" as defined by the VVC video coding standard, described, for example, in nplcitô. The sub-images are indicated by headers as being completely independent, particularly with regard to smoothing. In other words, in the encoding loop, no normative smoothing is performed across the boundaries between sub-images. Several encoding cores are used to parallelize the encoding of the sub-images. The encoding of a sub-image is implemented completely independently on a single encoding core, without any pixel sharing between encoding cores. After the sub-images have been encoded, a binary signal, comprising the encoded sub-images, is generated. The binary signal also includes a SEI (Self-Encoded Image) containing a numerical value.
[0089] The binary signal is received at the decoder. The encoded sub-images are each decoded independently on a separate decoding core. At the output of the decoding cores, the decoded sub-images are aggregated into a single image. The encoding information for the blocks constituting the sub-images is not transmitted at the output of the decoding cores. After sub-image aggregation and before image reconstruction, out-of-loop smoothing of the boundaries between sub-images is performed. This smoothing does not use any encoding information for the blocks constituting the sub-images because this information is not accessible by the decoding cores. The smoothing implementation uses a simple five-tap Gaussian filter whose standard deviation, which determines the smoothing power, is fixed to the numerical value contained in the SEI (Self-Effective Index). The smoothing power is thus identical regardless of the smoothed boundary and the image.
[0090] In a second specific example, an image is partitioned into several "tiles," the latter term being defined in the HEVC encoding standard and described in nplcit7 in particular. The tiles are associated, via headers, with instructions not to implement any smoothing of the boundaries between tiles in the encoding loop.
[0091] The HEVC encoding of each tile is implemented independently on a separate encoding core. In this second specific implementation, constraints are imposed to prevent any sharing of pixels or information between encoding cores. These constraints ensure that the encoding / decoding of each tile is independent of the encoding / decoding of other tiles and fall under the concept of MCTS already described.
[0092] As previously explained, tile encoding involves tile prediction, which can include intra- and / or inter-prediction. Specifically, tile prediction can include a prediction of a motion vector, for example in inter- or "merge" encoding mode, both of which are defined in the HEVC standard.
[0093] A first example of a constraint concerns the prediction of the motion vector: it involves imposing that, when predicting a current tile, the prediction of the associated motion vector points to that current tile.
[0094] A second example of a constraint concerns the "merge" encoding mode, in which a list of candidates is provided to predict a motion vector for a current tile. The constraint is to provide only candidates whose motion vector prediction points to the current tile. Furthermore, a candidate labeled "TMVP" (temporal motion vector predictor) is typically provided, this TMVP candidate referring to a previously encoded image. However, it is possible that the TMVP candidate refers to a tile located in a different position within the previously encoded image than the current tile. In other words, in such a case, the current tile and the tile to which the TMVP candidate refers are processed by separate encoding cores.One possibility is then to prohibit the selection of this TMVP candidate to avoid any exchange of vectors, or more generally of information, between encoding cores.
[0095] Tile decoding is performed serially on a single decoding core. The decoded tiles are then aggregated into an image. After aggregation and before image rendering, out-of-loop smoothing of the tile boundaries is performed. This smoothing utilizes all the block encoding information available from the single decoding core. Out-of-loop smoothing is implemented by an anti-block filter whose operation is defined in the HEVC standard and relies on one or more SEI messages related to the block boundaries, and consequently, to the tile boundaries.
[0096] A toutes fins utiles, les documents non-brevet suivants sont cités : nplcitl : Y. -K. Wang et al., "The High-Level Syntax of the Versatile Video Coding (VVC) Standard," in IEEE Transactions on Circuits and Systems for Video Technology, vol. 31 , no. 10, pp. 3779-3800, Oct. 2021 , doi: 10.1109 / TCSVT.2021 .3070860 nplcit2 : WU, Yongjun, SULLIVAN, Gary J., ZHANG, Yifu, Motion-constrained tile set for region of interest coding, WO / 2014 / 168650 nplcit3 : Hendry, S. Hong, J. Chen, Y.-K. Wang (Huawei), JVET-N0109, CE12 / AHG12: Treating boundaries of independent tile groups as picture boundaries, Mars 2019 nplcit4 : S. Cho, H. Kim, H. Y. Kim and M. Kim, "Efficient In-Loop Filtering Across Tile Boundaries for Multi-Core HEVC Hardware Decoders With 4 K / 8 K-UHD Video Applications," in IEEE Transactions on Multimedia, vol. 17, no. 6, pp. 778-791 , June 2015, doi: 10.1 109 / TMM.2015.2418995 nplcit5 : Singh, S., Sharma, R. K., & Sharma, M. K. (2012). Post Processing Technique to Reduce Tile Boundary Artifacts in JPEG2000 Compressed Images. The International Journal of Multimedia & Its Applications, 4(1 ), 127 nplcit6 : Schwartz, E. L., Berkner, K., & Gormish, M. J. (1999). Optimal tile boundary artifact removal with CREW. In Proc, of Picture Coding Symposium, (pp. 285-288) nplcit7 : SO / IEC CD 23002-7 Supplemental enhancement information messages for coded video bitstreams nplcit8 : H.266 : Versatile video coding https: / / www.itu.int / rec / T-REC-H.266 nplcit9 : ITU-T H.265, High efficiency video coding https: / / www.itu.int / ITU-. T / recommendations / rec.aspx?rec=14107&lang=en nplcitl 0 : Y. He, M. Coban, M. Karczewicz, “AHG9 / AHG13: Film grain blending process for film grain characteristics SEI message”, JVET-Y0053, Jan. 2022.
Claims
Claims
1. Method for processing at least one decoded image area, the method comprising: at the output of a decoding loop having decoded at least one current image area, processing the decoded current image area, using a boundary smoothing module (518) using metadata relating at least to a boundary between the current image area and a neighboring image area, in which the metadata relates at least to a choice of a smoothing function to be applied at the boundary from a set of predetermined smoothing functions.
2. The method of claim 1, wherein: the decoded current image area comprises a border region corresponding to the boundary and a region distant from the boundary, the processing of the decoded current image area comprises processing of the border region and does not comprise processing of the region distant from the boundary.
3. A method according to claim 1 or 2, the method further comprising: controlling the processing of the decoded current image area using a controller using first data relating to a decoding of the current image area and second data relating to a decoding of a neighboring image area.
4. Computer program comprising instructions for implementing the method according to one of claims 1 to 3 when this program is executed by a processor.
5. Non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to one of claims 1 to 3 when this program is executed by a processor.
6. Digital signal comprising at least one encoded current image area and metadata relating at least to a boundary between the current image area and a neighboring image area, wherein the metadata relates at least to a choice of a smoothing function to be applied at the boundary from a set of predetermined smoothing functions.
7. A digital signal according to claim 6, wherein the metadata further relates to an activation of smoothing, and / or to a setting of a smoothing strength to be applied at the boundary.
8. Data processing circuit comprising: an input interface configured to receive a digital signal comprising at least one encoded current image area and metadata relating at least to a boundary between the current image area and a neighboring image area, in which the metadata relates at least to a choice of a smoothing function to be applied at the boundary from a set of predetermined smoothing functions, and at least one output interface configured to provide the encoded current image area as input to a decoding loop, and provide the metadata to a processing module (518) provided at the output of the decoding loop.
9. The data processing circuit of claim 8, wherein the output interface is configured not to provide the metadata to the decoding loop.