Method and apparatus for film grain modeling

By selecting regions for film grain modeling based on coding parameters during the video encoding process, the method effectively addresses the challenge of preserving and reconstructing film grain in video compression, achieving efficient and accurate film grain parameter extraction.

JP2025516240APending Publication Date: 2025-05-27INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024563820
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-05
Filing Date
2023-05-02
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing video compression technologies face challenges in efficiently compressing and reconstructing film grain, as traditional coding tools often remove film grain or require high bit rates to preserve it, contradicting the goal of bit-saving encoding.

Method used

A method and apparatus for selecting regions for film grain modeling by obtaining a reconstructed image from encoding, selecting regions based on coding parameters such as QP, CU size, and intra prediction mode, and determining film grain parameters from these selected regions and the input image.

Benefits of technology

This approach allows for fast and reliable film grain parameter extraction, improving the acceleration and accuracy of uniform region detection, thereby enhancing film grain preservation and reconstruction in video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516240000001_ABST
    Figure 2025516240000001_ABST
Patent Text Reader

Abstract

A method and apparatus are provided for selecting a region to be used for film grain modeling, where a reconstructed image is obtained from encoding an input image, at least one region of the reconstructed image is selected based on at least one coding parameter of the region when encoding the input image, and film grain parameters are determined from the selected at least one region and the input image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of European Patent Application No. 22305671.4, filed May 5, 2022, which is incorporated by reference in its entirety.

[0002] FIELD OF THEINVENTION The present embodiments relate generally to video compression, distribution and rendering, and more particularly to film grain modeling. The present embodiments relate to a method and apparatus for detecting uniform regions for use in film grain modeling. [Background technology]

[0003] To achieve high compression efficiency, image and video coding schemes usually employ prediction and transformation to exploit spatial and temporal redundancy in video content. In general, intra- or inter-prediction is used to exploit intra- or inter-picture correlation, and then the difference between the original block and the predicted block, often called the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to the entropy coding, quantization, transformation, and prediction.

[0004] Film grain is a specific type of noise that appears in video in a very pleasing and very characteristic way. In its essence, film grain is the result of the physical process of analog film stock exposure and development. Originally, the noise resulting from the photography process, the so-called film grain, is in the sensor noise that is naturally present and intrinsic to each analog camera. With the age of digital cameras, this sensor noise disappeared at the capture stage, but now it is added later to the content to recreate a cinematic look. The random nature of this noise makes it difficult to compress using traditional coding tools. Common parameters of encoding tools, such as those chosen for low bit rates, often remove film grain. High bit rates are needed to preserve and reconstruct film grain with sufficient quality, which goes against the encoding / decoding goal of saving bits while encoding the content. To overcome this encoder filtering problem, film grain modeling is usually performed before the encoding stage, and in the decoding stage, during the so-called compositing process, film grain is added to the reconstructed video using a model of the film grain. Summary of the Invention

[0005] According to one aspect, a method is provided for selecting a region to be used for film grain modeling, the method including obtaining a reconstructed image from encoding an input image, selecting at least one region of the reconstructed image, where the selection of the at least one region of the reconstructed image is based on at least one coding parameter of the region when encoding the input image, and determining film grain parameters from the selected at least one region and the input image.

[0006] According to another aspect, an apparatus is provided for selecting a region to be used for film grain modeling, comprising one or more processors operable to obtain a reconstructed image from encoding an input image, select at least one region of the reconstructed image based on at least one coding parameter of the region when encoding the input image, and determine film grain parameters from the selected at least one region and the input image.

[0007] One or more embodiments also provide a computer program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method for selecting a region to be used for film grain modeling according to any of the embodiments described herein. One or more of the embodiments also provide a computer-readable storage medium having stored thereon instructions for selecting a region to be used for film grain modeling according to the method described above. [Brief description of the drawings]

[0008] [Figure 1] 1 illustrates a block diagram of a system in which aspects of the present embodiments may be implemented. [Diagram 2] 1 illustrates a block diagram of one embodiment of a video encoder. [Diagram 3] 1 illustrates a block diagram of one embodiment of a video decoder. [Figure 4] Illustrate an example of block partitioning. [Diagram 5] 1 illustrates an example of a method for film grain encoding / decoding. [Figure 6] 1 illustrates an example of a method for encoding / decoding video content using film grain modeling. [Figure 7] 1 illustrates an example of a method for encoding / decoding video content using film grain modeling, according to one embodiment. [Figure 8]1 illustrates an example of a method for determining film grain parameters according to one embodiment. [Figure 9] 1 illustrates an example of a method for determining a region used to determine a film grain parameter, according to one embodiment. [Figure 10] Illustrates an example of partitioning-driven block selection for QP22(top) and QP(32) bottom. [Figure 11] 1 illustrates an example of intra-mode driven block selection based on planar mode (top) and DC mode (bottom). [Figure 12] 1 illustrates an example of residual-driven block selection based on code block flag values. [Figure 13] 1 illustrates an example of cost-driven block selection based on bpp cost below a given value. [Figure 14] 1 illustrates a block diagram of a system in which aspects of the present embodiments may be implemented, according to another embodiment. [Figure 15] 1 illustrates two remote devices communicating over a communications network, according to an example of the present principles. [Figure 16] 1 illustrates the syntax of a signal according to an example of the present principles; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] In this application, various aspects are described, including tools, features, embodiments, models, approaches, and the like. Many of these aspects are described in a specific and often descriptive manner to at least indicate their individual characteristics. However, this is for purposes of clarity of description and not to limit the applicability or scope of the aspects. In fact, all of the different aspects can be combined and substituted to provide further aspects. Moreover, these aspects can also be similarly combined and substituted with aspects described in previous applications.

[0010] The aspects described and contemplated in this application can be implemented in many different forms. Figures 1, 2, and 3 below provide some embodiments, but other embodiments are contemplated, and discussion of Figures 1, 2, and 3 is not intended to limit the scope of implementations. At least one of the above aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium having stored therein instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored therein a bitstream generated according to any of the described methods.

[0011] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably.

[0012] Various methods are described herein, each of which includes one or more steps or acts for achieving the described method. The order and / or use of specific steps and / or acts may be modified or combined, unless a specific order of steps or acts is required for proper operation of the method. Additionally, terms such as "first", "second", etc. may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decode" and "second decode". The use of such terms does not imply any ordering to the modified operations, unless specifically required. Thus, in this embodiment, the first decode need not be performed before the second decode, but may occur, for example, before, during, or during a period of overlap with the second decode.

[0013] Moreover, aspects of the present disclosure are not limited to VVC or HEVC, but may be applied, for example, to other standards and recommendations, whether existing or developed in the future, and to extensions of any such standards and recommendations (including VVC and HEVC).Unless otherwise indicated or technically precluded, aspects described in this application may be used individually or in combination.

[0014] FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device including various components described below and configured to perform one or more of the aspects described in the present application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected appliances, and servers. The elements of system 100, alone or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or separate components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 100 are distributed across multiple ICs and / or separate components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, through a communication bus or dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in the present application.

[0015] The system 100 includes at least one processor 110 configured to execute instructions loaded therein, for example, to implement various aspects described in the present application. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include non-volatile and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0016] The system 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded or decoded video, which may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of the system 100 or may be incorporated within the processor 110 as a combination of hardware and software known to those skilled in the art.

[0017] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described herein may be stored in the storage device 140 and then loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, expressions, operations, and operational logic.

[0018] In some embodiments, memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 also refers to ISO / IEC 13818, 13818-1 also known as H.222, and 13818-2 also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, H.265 and MPEG-H Part 2 are also known), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Experts Team (JVET)).

[0019] Inputs to the elements of system 100 may be provided through a variety of input devices, as shown in block 105. Such input devices may include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, an RF signal transmitted throughout a broadcast by a broadcaster, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples include composite video, not shown in FIG. 1.

[0020] In various embodiments, the input devices of block 105 have associated respective input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to (for example) as a channel, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs these various functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and refiltering RF signals transmitted over a wired (e.g., cable) medium to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, omit some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0021] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting the system 100 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or in the processor 110, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, for example, in a separate interface IC or in the processor 110, as desired. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and an encoder / decoder 130 operating in combination with memory and storage elements that process the data streams as desired for presentation on an output device.

[0022] The various elements of system 100 may be provided within an integrated housing in which the various elements may be interconnected and transmit data between each other using suitable connection arrangements 115, e.g., internal buses known in the art, including I2C buses, wiring, and printed circuit boards.

[0023] System 100 includes a communication interface 150 that enables communication with other devices over a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card, and communication channel 190 may be implemented in a wired and / or wireless medium, for example.

[0024] Data is streamed to the system 100 using a Wi-Fi network, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), in various embodiments. The Wi-Fi signal in these embodiments is received via a communication channel 190 adapted for Wi-Fi communication and the communication interface 150. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, a set-top box that delivers data via an HDMI connection of the input block 105 is used to provide the streamed data to the system 100. In yet other embodiments, an RF connection of the input block 105 is used to provide the streamed data to the system 100. As indicated above, various embodiments provide the data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0025] System 100 may provide output signals to various output devices, including a display 165, speakers 175, and other peripheral devices 185. Display 165 in various embodiments includes, for example, one or more of a touch screen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 165 may be for a television, a tablet, a laptop, a mobile phone, or other device. Additionally, display 165 may be integrated with other components (e.g., as in a smart phone) or may be separate (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 185 include one or more of a standalone digital video disc (or digital versatile disc) (both terms DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 to provide functionality based on the output of system 100. For example, a disc player performs the function of playing the output of system 100.

[0026] In various embodiments, control signals are communicated between system 100 and display 165, speaker 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that allow inter-device control with or without user intervention. Output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 may be integrated into a single unit with other components of system 100, for example, in an electronic device such as a television. In various embodiments, display interface 160 includes a display driver, for example, a timing controller (T Con) chip.

[0027] Display 165 and speakers 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments where display 165 and speakers 175 are external components, output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.

[0028] The embodiments may be performed by computer software implemented by the processor 110, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 120 may be of any type suitable for the technology environment, and may be implemented using any suitable data storage technology, such as, as non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memories, and removable memory devices. The processor 110 may be of any type suitable for the technology environment, and may include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a multi-core architecture-based processor.

[0029] 2 illustrates an encoder 200. Variations of this encoder 200 are contemplated, but for clarity, the following describes the encoder 200 without describing all possible variations.

[0030] In some embodiments, FIG. 2 also illustrates an encoder that employs improvements to the HEVC or VVC standards, or technologies similar to HEVC or VVC, such as encoders under development by the Joint Video Exploration Team (JVET).

[0031] Before being encoded, a video sequence may undergo encoding pre-processing (201), such as applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components to obtain a signal distribution that is more resistant to compression (e.g., using histogram equalization of color components) or resizing (ex: downscaling) the picture. Metadata may be associated with the pre-processing and attached to the bitstream.

[0032] In the encoder 200, a picture is coded by the encoder elements, as described below. The picture to be coded is partitioned (202) into units, e.g., CUs, and processed. Each unit is coded, e.g., using either intra mode or inter mode. When a unit is coded in intra mode, intra prediction (260) is performed. In inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder decides (205) whether to use one of intra mode or inter mode to code the unit, and indicates the intra or inter decision, e.g., by a prediction mode flag. The encoder may also mix (263) intra prediction and inter prediction results, or mix results from different intra / inter prediction methods. A prediction residual is calculated (210), e.g., by subtracting the prediction block from the original image block.

[0033] The motion refinement module (272) uses already available reference pictures to refine the motion field of a block without referring to the original block. The motion field for a region can be considered as a collection of motion vectors for all pixels that comprise the region. If the motion vectors are subblock-based, the motion field can also be represented as a collection of all subblock motion vectors in the region (all pixels in a subblock have the same motion vector, and the motion vectors can be different for each subblock). If a single motion vector is used for a region, the motion field for the region can also be represented by a single motion vector (the same motion vector for all pixels in the region).

[0034] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, as well as motion vectors and other syntax, are entropy coded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process.

[0035] The encoder decodes the coded block to provide a reference for further prediction. To decode the prediction residual, the quantized transform coefficients are dequantized (240) and inverse transformed (250). The decoded prediction residual and the predicted block are combined (255) to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / Sample Adaptive Offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (280).

[0036] 3 illustrates a block diagram of a video decoder 300. In the decoder 300, a bitstream is decoded by decoder elements, as described below. The video decoder 300 generally performs a decoding path that is the inverse of the encoding path described in FIG. 2. Additionally, the encoder 200 generally performs video decoding as part of the video data encoding.

[0037] In particular, the decoder's input includes a video bitstream, which may be generated by video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. Picture partition information indicates how the picture is partitioned. The decoder may then divide (335) the picture according to the decoded picture partitioning information. To decode the prediction residual, the transform coefficients are dequantized (340) and inverse transformed (350). The decoded prediction residual and the predicted block are combined (355) to reconstruct the image block.

[0038] The prediction block can be obtained (370) from intra prediction (360) or motion compensated prediction (i.e., inter prediction) (375). The decoder can mix (373) intra and inter prediction results, or mix results from multiple intra / inter prediction methods. Before motion compensation, the motion field can be refined (372) by using already available reference pictures. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (380).

[0039] The decoded picture may further undergo post-decoding processing (385), such as an inverse color conversion (e.g., from YCbCr 4:2:0 to RGB 4:4:4), or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (201), or a resizing (ex: upscaling) of the reconstructed picture. The post-decoding processing may use metadata derived in the pre-encoding processing and signaled in the bitstream.

[0040] Block Partitioning Structure In VVC, a picture is divided or partitioned into multiple non-overlapping coding tree units (CTUs). The CTU size in VVC can be set to a maximum of 128x128 for luma samples. Each coding tree unit (CTU) is treated as a coding unit or split into multiple coding units (CUs) by one or more recursive quad-tree (QTBT) partitions, followed by one or more recursive multi-type tree splits, such as a binary tree (BT) or a ternary tree (TT). The latter can be a horizontal binary tree split (SPLIT_BT_HOR), a vertical binary tree split (SPLIT_BT_VER), a horizontal ternary tree split (SPLIT_TT_HOR), or a vertical ternary tree split (SPLIT_TT_VER), as illustrated in FIG. 4. The CTU dual tree for intra-coded slices is described above with a block partitioning structure that separates the coding trees for luma and chroma. The splits (QTBT, BT, or TT) that are applied to each CTU and result in one or more CUs have square and rectangular shapes. They allow better adaptation to local picture content characteristics and features.

[0041] By increasing the CTU size up to 128x128 for luma, coding efficiency is improved, but at the expense of increased complexity. The size of each CU is related to the information present in this CU, or more precisely, to the rate distortion (RD) cost of the CU. The rate distortion (RD) cost of a candidate block can be calculated as follows: CostRD = D + Lambda(QP) × R where D is the distortion of the current block with coding parameters (e.g., L2, i.e., Euclidean distance), R is the associated rate (in number of coding bits), and Lambda(QP) is the Lagrangian parameter inferred from the quantization parameter QP. A high QP produces fewer but larger CUs than a small QP. Once partitioning is completed during the coding stage, several features are available for the intra frame and each CU, as follows: The intra prediction mode indicates which prediction mode is used, e.g., planar or angular prediction. The code block flag (CBF) indicates whether residual is added on top of the prediction. The CBF flag is coded per component. The bit cost per pixel (bpp) can be extracted by the encoder during the encoding stage. DCT transform coefficients can also be used, which are the values ​​of the transformed and quantized coefficients. The significance flag and the last xy can be used. These parameters signal which coding group in the transform unit contains the coefficient.

[0042] Film Grain Modeling Modern video compression and delivery systems can provide mechanisms for removing film grain before and / or during compression, and for users to add film grain back in an automatic, controllable manner through the use of parameter models. Figure 5 illustrates a block diagram of one example of film grain usage in a video coding framework. Typically, an encoder removes the original film grain from the video (500) and estimates the film grain model parameters based on uniform region selection (501) and analysis of the original content (502).

[0043] For film grain estimation, uniform regions in an image are detected. Common techniques used to detect uniform regions include applying some pre-filtering, such as noise removal or edge detection, from the encoding process. Since uniform region detection is an additional process that runs after pre-processing of frames to remove film grain of the input video and encoding of frames, it is an overhead to the encoding stage. It consumes CPU and memory.

[0044] In the framework illustrated in FIG. 5, content without film grain is encoded (503) and transmitted to a decoder, which decodes (504) the transmitted content and replaces (505) the removed film grain with a synthetic substitute that visually approximates the removed film grain.

[0045] In this manner, information about film grain is communicated as additional metadata, as illustrated in one embodiment according to Figure 6, which shows a block diagram of an example of film grain usage in a video coding framework according to another example. As explained above in conjunction with Figure 5, film grain preservation has two stages: modeling (region selection and analysis) and synthesis.

[0046] The embodiments presented herein focus on modeling, and more precisely on granularity analysis / noise extraction from the original content. As illustrated in Fig. 6, the grain analysis uses an original frame with grain (input video) and a corresponding filtered frame without film grain (filtered video). This filtered frame is obtained through a denoising step (600 pre-processing). From the filtered frame, flat regions are extracted (601) because they act as a mask for estimating (602) grain parameters in the grain layer, which are determined as the difference between the original frame and its filtered version. The filtered frame is coded (603).

[0047] Once estimated, the grain parameters (FG params) are transmitted along with the compressed video bitstream, for example in an SEI message (FG SEI). After decoding (604), the film grain is synthesized (605) from the film grain parameters and added back to the reconstructed video frames.

[0048] To determine the film grain parameters, noise removal and flat area detection are used in the grain analysis.

[0049] Many approaches address film grain removal, ranging from using a video encoder as a denoiser, to Wiener-type filters, nonlinear denoising with total variation minimisation, and multi-hypothesis motion compensated filtering. MCTF (Motion Compensated Temporal Filtering), a video filtering tool in VVC, is also an efficient denoiser. When denoising is applied, flat regions are estimated so that only flat / uniform / smooth regions of the picture are used for estimation, since edges and textures can affect the estimation of film grain intensity and patterns. To determine the smooth regions of a picture, a Canny edge detector is often applied to the denoised image at different scales, followed by a dilation operation.

[0050] It is noted that the above method relies on an additional process to extract the film grain parameters, which may not be suitable for very low complexity / latency encoders.

[0051] In the example illustrated in Figures 5 and 6, the workflow of the legacy solution can be summarized as follows: Film grain analysis and extraction is performed outside the codec workflow. Film grain parameters are sent to the decoder (typically using an SEI message). The compositor adds the film grain back into the reconstructed image for display.

[0052] According to one embodiment, a method and apparatus are provided for determining film grain parameters utilizing the results of an encoding process, which allows for fast and reliable film grain parameter extraction by selecting regions based on coding parameters obtained during the encoding of input frames.

[0053] In one embodiment, only intra-coded frames are considered for region selection.

[0054] In another embodiment, the regions are selected based on coding unit (CU) partitioning for intra frames during encoding.

[0055] In some embodiments, flat regions may be selected by analyzing relevant information for each CU from the CTU luminance, taking into account at least one or a combination of parameters, for example, QP, coding unit (CU) size, intra prediction mode (MPI, DC, plane, angle), code block flag (CBF) off or on, bits per pixel (BPP) cost, Last_xy_sig, DCT coefficient distribution.

[0056] In this way, uniform region detection for film grain determination is accelerated and improved.

[0057] Figure 7 illustrates an example of a method for encoding / decoding video content using film grain modeling, according to one embodiment. In this embodiment, steps 602-605 are similar to those described in Figure 6, except that encoding of the input frames (603) replaces the pre-processing of the input images (600) of Figure 6. According to the embodiment presented herein, the encoder is used as a denoiser. In other words, the encoding process of the encoder is used to remove the film grain of the input frames.

[0058] Also, in this embodiment, the region detection (700) is driven by the coding information provided in the encoder for encoding the input frame, thus avoiding an additional process for region detection.

[0059] In one embodiment, to detect flat / uniform regions, only intra frames are considered, where the partitioning and information for each CU is more relevant. In one embodiment, only the luma component is used for the analysis.

[0060] When encoding an input frame, each CTU is treated as one CU or split into multiple CUs by a recursive QT followed by a recursive binary-ternary tree (BTT), also called a multi-type tree (MTT). As explained above, the rate-distortion (RD) cost of a candidate block is driven by a quantization parameter QP. QP defines the quantization level.

[0061] At high bitrates and high PSNR, the QP value is low, but at low bitrates and low PSNR, the QP value is high.

[0062] For QP values ​​above a given QP value, e.g., QP≧22, the film grain is filtered and not preserved when encoding the input frame. Therefore, in this embodiment, the encoding can be considered as a denoiser, which is useful for estimating a filtered version of the input frame without the film grain. In this case, no external denoiser is required.

[0063] In another embodiment, when the QP value is higher than a given QP value, for example QP<22, the film grain is no longer removed from the input frame when encoding the input frame. Therefore, a pre-processing of the input frame is still required before encoding the input frame to filter and remove the film grain. This denoiser can be any denoiser outside the video coder tool, or an internal denoiser such as image filtering with or without temporal information, for example, the MCTF tool of VVC can be used. Therefore, in this embodiment, the method illustrated in FIG. 7 includes an additional pre-processing step of film grain removal. However, the region selection (700) is still based on at least one of the coding parameters obtained for the CU of the input frame when encoding the input frame.

[0064] 8 illustrates an example of a method for determining film grain parameters, according to one embodiment. At 800, a reconstructed image is obtained from encoding an input image. As described above, depending on the QP used to encode the input image, the input image may or may not have been processed prior to encoding to remove film grain. At 801, at least one region of the reconstructed image is selected based on at least one coding parameter of the region when encoding the input image. At 802, a film grain parameter is determined from the selected at least one region and the input image.

[0065] In the following, several embodiments of the region selection process are presented. The embodiments can be used alone or in combination.

[0066] Partitioning-Driven Block Selection The choice of block size depends heavily on the QP value. Higher and lower QP values ​​imply choosing larger and smaller block sizes, respectively. Lower QP values ​​preserve image details, while higher QP values ​​provide higher compression ratios at the expense of image quality loss. If the cost is not satisfied for a CU, this CU is split. The higher the QP, the larger the CU. Homogeneous regions should be localized in CUs with low cost and large surface.

[0067] 10 illustrates an example of quad-tree partitioning obtained for QP22 (up) and QP(32) down. It can be seen that for the same region in the frame, depending on the QP value, the quad-tree partitioning is not the same.

[0068] In this embodiment, all small CUs are discarded (i.e., not selected) and only large CUs are analyzed. Small CUs should be understood as CUs having a size smaller than a given CU size. Similarly, large CUs should be understood as CUs having a size larger than a given CU size. For example, to obtain a relevant model of film grain, a CU block size of 8x8 is too small. Either 32x32 and 16x16, or 64x64, 32x32 and 16x16 can be considered as a good choice.

[0069] When determining film grain from large CUs, if the block size is not 64x64, block scaling to 64x64 is applied. Experimental results show that at least four blocks of 64x64 are needed to have a good estimate of the film grain parameters.

[0070] Intra-mode driven block selection In this embodiment, the intra-mode coding information is used to check whether the CU is coded with orientation information. In VVC, intra-mode coding includes six most probable modes. Planar mode, which is always the Most Probable Mode (MPM), is coded first with a separate flag. DC mode and angular mode are coded using a list of the remaining five MPMs derived from the intra-modes of neighboring blocks. If an intra-mode is coded with angular mode, it means that the block includes coded orientation propagation.

[0071] In this embodiment, only CUs coded in at least DC mode or planar mode are selected, which corresponds to flat blocks and results in regions with a higher probability of being in a uniform region. Figure 11 illustrates an example of intra-mode driven block selection based on planar mode (mode 0, top of Figure 11) and DC mode (mode 1, bottom of Figure 11). Blocks selected based on this criterion are shown in dark gray in Figure 11. In these examples, 37% of CUs are in planar mode and 12% are in DC mode.

[0072] Depending on the intra prediction modes provided by the encoder and the video compression standard used to encode the video, other intra prediction modes can be used as long as they correspond to prediction modes for flat regions.

[0073] Residual-Driven Block Selection In another embodiment, the code block flag (CBF) is considered to select the CU. The CBF is used to indicate whether the encoding result of the CU contains a non-zero residual. If the CBF is equal to 0 (off), the residual coefficients of the CU do not need to be coded. This means that the coefficients are all 0 and the region can be considered to be well predicted. When combined with the above embodiment (intra mode), such a region with a CBF of 0 can be considered to have no details or to be uniform. Figure 12 illustrates an example of residual-driven block selection based on the code block flag value. The blocks selected based on this criterion are shown in dark gray in Figure 12. In this example, 10% of the CUs have a CBF of off value.

[0074] Cost-Driven Block Selection Another relevant information that can be used for block selection is the bit per pixel (BPP) cost per CU. The coding cost of textured regions is higher than that of uniform regions. In this embodiment, CUs with low BPP costs are selected, i.e., CUs with BPP costs below a given value are selected. For example, BPP<0.015 may be a given value for detecting uniform regions. Figure 13 illustrates an example of cost-driven block selection based on a BPP cost lower than 0.015. Blocks selected based on this criterion are shown in light gray in Figure 13.

[0075] Coefficient Distribution Driven Block Selection In this embodiment, the coefficient distribution of the transformed coefficients in a block is considered for region selection. The distribution of the DCT coefficients or energy of the first n DCT or transformed coefficients also gives some information about homogeneous regions. If most of the DCT coefficients are in low frequencies, it means that there is no or not much energy in high frequencies, that is, there is no detail in this CU, and the CU can be considered as a homogeneous region. The percentage of energy in the first n coefficients can be given by: Let X be the DCT coefficient of a CU, then an energy threshold is set, e.g., 90% of the value of the energy is considered, i.e., EnergyTh=0.9.

[0076] The total energy of the block is determined by summing the energies of all the coefficients, for example using L2.

[0077]

number

[0078] The sort() function sorts all the DCT coefficients of a CU from lowest frequency to highest frequency. The abs() function returns the absolute value of the coefficients. The normL2() function calculates the square root of the sum of the squares of the elements. This provides the energy of the Xsort vector.

[0079] Then a loop is performed over all the coefficients from low to high frequency, and for each coefficient c it is checked whether the sum of the energies of all the coefficients from the first coefficient to coefficient c relative to the total energy of the block is below an energy threshold.

[0080]

number

[0081] This function returns coeffIdx, i.e. the index of the DCT coefficient where 90% of the energy can be found in the previous lower DCT coefficient. If coeffIdx is low, it means that the energy is in the first / low frequencies, i.e. this CU is a homogeneous region. coeffIdx is considered low when it is below a given percentage of the number of coefficients in the block.

[0082] In a variant, the coefficient distribution can be determined using the last_xy_sig flag, which gives similar information. This flag corresponds to the position (x, y) of the last significant DCT coefficient in the CU. If (x, y) is close to the top-left position in the CU, the energy of the DCT coefficients is concentrated in low frequencies, i.e., the CU is a homogeneous region. As above, the position of the last significant coefficient is considered to be close to the top-left position in the CU when it is before the given position in the CU. The given position can depend on the size of the CU.

[0083] The above described embodiments for selecting the regions used to determine the film grain parameters may be used alone, or two or more embodiments may be used in combination.

[0084] FIG. 9 illustrates an example of a method for determining regions used to determine film grain parameters, according to one embodiment.

[0085] When encoding a CU of a frame, the partitioning of the frame is used to check the size of the CU at 900. In response to determining that the size of the CU is below a given size, for example, if the CU size is below 16×16, the CU is not selected and the process ends for this CU. Checking the size of the CU may include checking whether the CU is a square CU and, when the CU is square, checking whether the CU is below the given size.

[0086] In another variation, checking the size may also include whether the minimum dimension of the CU is below a given value when the CU is not square, i.e., when the CU is rectangular, for example, checking for a rectangular CU if the minimum dimension of the CU's width and height is below a given value, such as 16, 32, or 64.

[0087] If it is determined in 900 that the size of the CU is equal to or greater than the given size value, the intra-mode prediction used to code the CU is checked in 901. In response to determining that the intra-mode prediction is neither DC mode nor planar mode, the CU is not selected and the process ends for this CU. Otherwise, the CU is selected.

[0088] At 902, a check is made to see if residual is coded for the current CU, e.g., the value of the CBF is checked. In response to a determination that the CBF has a value of 0 (off, i.e., no residual is coded for the CU), the CU is selected and the process ends.

[0089] Otherwise (CBF value 1, ON), depending on the variant, other coding parameters may be checked for the CU.

[0090] At 903, the last_xy_flag is checked. If the last_xy_flag is close to the top-left position of the CU, the CU is selected. If not, the CU is not selected and the process ends.

[0091] In another variation, in 904, the energy of the N first coefficients of the CU is checked, and if it is determined that the energy of the N first coefficients of the CU is equal to or greater than a given amount of the energy of the entire CU (considering all coefficients of the CU), the CU is selected, where N is a given number of coefficients considered in low frequencies. N may depend on the CU size.

[0092] In another variation, at 905, the bit cost per pixel is checked for the CU, and if the BPP cost of the CU is determined to be below a given value, the CU is selected.

[0093] In another variation, the decisions at 903, 904, and 905 can be performed later if a CU is not selected from the previous decision.

[0094] 14 illustrates a block diagram of a system in which aspects of the present embodiment may be implemented, according to another embodiment. FIG 14 shows an embodiment of an apparatus 1400 for determining film grain parameters or for selecting regions for determining film grain parameters, according to any one of the embodiments described herein.

[0095] The device includes a processor 1410 and may be interconnected through at least one port to a memory 1420. Additionally, both the processor 1410 and the memory 1420 may also have one or more additional interconnects to external connections.

[0096] The processor 1410 is also configured to obtain a reconstructed image from the encoding of the input image according to any one of the embodiments described herein, select at least one region of the reconstructed image based on at least one coding parameter of the region when encoding the input image, and determine film grain parameters from the selected at least one region and the input image, For example, the processor 1410 is configured using a computer program product including code instructions implementing any one of the embodiments described herein.

[0097] In one embodiment illustrated in Fig. 15, in a transmission context between two remote devices A and B over a communication network NET, device A comprises a processor associated with memory RAM and ROM configured to implement a method for encoding video as described in conjunction with Figs. 1, 2, 7, and device B comprises a processor associated with memory RAM and ROM configured to implement a method for decoding video as described in conjunction with Figs. 1, 3, 7. Depending on the embodiment, device A is also configured to determine film grain parameters or to select regions for determining film grain parameters as described in conjunction with Figs. 4-13. In some embodiments, device B is also configured to synthesize film grain into the received video using the received film grain metadata.

[0098] According to one example, the network is a broadcast network adapted to broadcast / transmit the encoded video and the converted film grain metadata from device A to decoding devices including device B.

[0099] 16 shows an example of the syntax of a signal transmitted over a packet-based transmission protocol. Each transmission packet P includes a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may include video data and film grain metadata encoded according to any one of the embodiments described above.

[0100] Various implementations involve decoding. As used herein, "decoding" can encompass all or some of the processes performed on a received encoded sequence to produce a final output suitable for, e.g., a display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as, e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by a decoder in various implementations described herein, such as, e.g., decoding resampling filter coefficients, resampling a decoded picture, etc.

[0101] As a further example, in one embodiment, "decoding" refers to entropy decoding only, in another embodiment, "decoding" refers to differential decoding only, in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding, and in another embodiment, "decoding" refers to the entire reconstruction picture process including entropy decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process as a whole will be clear based on the context of the specific description and will be well understood by one of ordinary skill in the art.

[0102] Various implementations involve encoding. Similar to the above discussion regarding "decoding," "encoding" as used herein can encompass all or part of the processes performed on an input video sequence to produce, for example, an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as, for example, partitioning, differential encoding, transforming, quantizing, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by encoders in various implementations, such as those described herein, such as, for example, determining resampling filter coefficients, resampling decoded pictures, etc.

[0103] As a further example, in one embodiment, "encoding" refers to entropy encoding only, in another embodiment, "encoding" refers to differential encoding only, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process as a whole will be clear based on the context of the specific description and will be well understood by one of ordinary skill in the art.

[0104] Please note that syntaxes as used herein are descriptive terms, and therefore do not exclude the use of other syntax names.

[0105] This disclosure has described various information, such as, for example, syntax, that may be transmitted or stored. This information may be packaged or arranged in various manners, including, for example, common to video standards, such as placing the information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or slice header), or SEI message. Other manners are also available, including, for example, common to system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for purposes of session announcement and session invitation, e.g., as described in the RFCs and used in conjunction with RTP (Real-time Transport Protocol) transport. b. DASH MPD (Media Presentation Description) Descriptors, e.g., as used in DASH and transmitted over HTTP, Descriptors are associated with a representation or a collection of representations to provide additional characteristics to the content representation. c. RTP header extensions, for example as used during RTP streaming. d. ISO Base Media File Formats, such as those used in OMAF and in some specifications that use boxes, which are object-oriented building blocks defined by a unique type identifier and a length, also known as "atoms". e. HTTP Live Streaming (HLS) manifests transmitted over HTTP. A manifest can be associated with a version or collection of versions of content, for example to provide characteristics of the version or collection of versions.

[0106] Where a figure is presented as a flow diagram, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that the figure also provides a flow diagram of the corresponding method / process.

[0107] Some embodiments refer to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is usually considered, often due to computational complexity constraints. Rate-distortion optimization is usually formulated to minimize a rate-distortion function, which is a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, these approaches may be based on an extensive test of all encoding options, including all considered modes or coding parameter values, but with a full evaluation of their coding costs and associated distortion of the reconstructed signal after coding and decoding. To keep the encoding complexity down, more rapid approaches may also be used, especially with the calculation of approximate distortion based on the prediction or prediction residual signal rather than the reconstructed signal. A mixture of these two approaches may also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a full evaluation of both the coding cost and the associated distortion.

[0108] The implementations and aspects described herein may be implemented as, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed in the context of only a single implementation form (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. The method may be implemented in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors include, for example, communication devices such as computers, cellular phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.

[0109] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment", or "in one implementation" or "in an implementation" in various places throughout this application, as well as other variations, are not necessarily all referring to the same embodiment.

[0110] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0111] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, obtaining information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0112] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally involves in some manner, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0113] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first enumerated alternative (A), or the selection of only the second enumerated alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first enumerated alternative (A), or the selection of only the second enumerated alternative (B), or the selection of only the third enumerated alternative (C), or the selection of only the first and second enumerated alternatives (A and B), or the selection of only the first and third enumerated alternatives (A and C), or the selection of only the second and third enumerated alternatives (B and C), or the selection of all three alternatives (A and B and C). This can be expanded as many times as the items listed, as would be apparent to one of ordinary skill in this and related arts.

[0114] Also, as used herein, the term "signaling" specifically means to indicate something to a corresponding decoder. For example, in a particular embodiment, an encoder signals a particular one of a number of resampling filter coefficients. Thus, in an embodiment, the same parameters are used at both the encoder and decoder sides. Thus, for example, an encoder can transmit a particular parameter to a decoder (explicit signaling) so that the decoder can use the same particular parameter. In contrast, if the decoder already has the particular parameter as well as other parameters, a non-transmitting signaling (implicit signaling) can be used to simply allow the decoder to know and select the particular parameter. By avoiding transmitting any actual function, bit savings are realized in various embodiments. It will be understood that signaling can be achieved in various ways. For example, one or more syntaxes, flags, etc. are used in various embodiments to signal information to a corresponding decoder. Although the above relates to the verb form of the word "signal", the word "signal" may also be used as a noun in this specification.

[0115] As will be apparent to one of ordinary skill in the art, implementations can result in a variety of signals formatted to carry information that can be, for example, stored or transmitted. Information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bit stream of the described embodiments. For example, such a signal can be formatted as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. As is known, the signal can be transmitted over a variety of different wired or wireless links. The signal can be stored in a processor-readable medium.

[0116] A number of embodiments have been described above, the features of which may be provided alone or in any combination across the various claim categories and types.

Claims

1. 1. A method comprising: Obtaining a reconstructed image from an encoding of the input image; selecting at least one said region of the reconstructed image based on at least one coding parameter of the region when encoding the input image; determining film grain parameters from the selected at least one region and the input image.

2. 1. An apparatus comprising one or more processors, the one or more processors comprising: Obtaining a reconstructed image from the encoding of the input image; selecting at least one said region of the reconstructed image based on at least one coding parameter of the region when encoding the input image; An apparatus operable to determine film grain parameters from the selected at least one region and the input image.

3. The region is a coding unit, and the at least one coding parameter is the size of the coding unit used to encode the coding unit; an intra-prediction mode used to encode the coding unit; a code block flag indicating whether encoding of the coding unit provides a non-zero residual; the bit cost per pixel used to encode said coding unit; the position in said coding unit of the last significant coefficient; 3. The method of claim 1 or the apparatus of claim 2, wherein a given percentage of the energy of the coefficients of the coding unit is obtained by a coefficient located before the first coefficient, and an index of the first coefficient obtained by a coefficient located before the first coefficient.

4. The method or apparatus of claim 3 , wherein the at least one region is selected in response to determining that a size of the at least one region is greater than or equal to a given coding unit size.

5. The method or apparatus of claim 3 or 4, wherein the at least one region is selected in response to determining that the intra prediction mode is a DC mode or a planar mode.

6. The method or apparatus of any one of claims 3 to 5, wherein the at least one region is selected in response to a determination that the code block flag indicates that the encoding of the coding unit provides zero residual.

7. A method or apparatus according to any one of claims 3 to 6, wherein the at least one region is selected in response to determining that the bit cost per pixel is greater than or equal to a given value.

8. 8. The method or apparatus of claim 3, wherein the at least one region is selected in response to determining that a distance between a top left position in the coding unit and the position in the coding unit of the last significant coefficient is less than or equal to a given value.

9. 9. The method or apparatus of claim 3, wherein the at least one region is selected in response to determining that a distance between a top-left position in the coding unit and the position in the coding unit of a first coefficient at which a given percentage of the energy is contributed by the coefficients located between the top-left position and the first coefficient is less than or equal to a given value.

10. The method according to claim 1 or any one of claims 3 to 9 or the apparatus according to any one of claims 2 to 9, wherein the input image is coded as an intra picture.

11. The method according to any one of claims 1 or 3 to 10 or the apparatus according to any one of claims 2 to 10, wherein the at least one coding parameter of the region is obtained from an encoding of a luminance component of the input image.

12. The method of claim 1 or any one of claims 3 to 11 or the apparatus of any one of claims 2 to 11, wherein obtaining the reconstructed image comprises removing film grain from the input image.

13. 13. The method or apparatus of claim 12, wherein in response to determining that a value of a quantization parameter used to encode the input image is greater than or equal to a given value, removing film grain from the input image is performed by the encoding of the input image at the quantization parameter.

14. 13. The method or apparatus of claim 12, wherein in response to determining that a value of a quantization parameter used to encode the input image is less than a predetermined value, removing film grain from the input image is performed by denoising the input image using an external denoiser or image filtering tool of a video encoder.

15. A computer readable storage medium storing instructions for causing one or more processors to perform the method of any one of claims 1 or 3-14.

16. A computer program product comprising instructions, which when executed by one or more processors, cause the one or more processors to perform a method according to any one of claims 1 or 3 to 14.