Reversible production of rescaled grain-free and energy-aware video using an invertible framework

The invertible framework addresses the energy inefficiencies in processing high-resolution, grainy videos by converting them into low-resolution, grain-free, and energy-aware formats, achieving efficient energy use and maintaining image quality through reversible operations.

WO2025103888A1PCT designated stage expired Publication Date: 2025-05-22INTERDIGITAL CE PATENT HOLDINGS SAS

Patent Information

Application Number
PCT/EP2024/081627
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-16
Filing Date
2024-11-08
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing video processing techniques consume significant energy due to the need to handle high-resolution, grainy videos, and lack efficient methods for reversible conversion to low-resolution, grain-free, and energy-aware formats without compromising image quality.

Method used

An invertible framework is used to modify high-resolution, grainy videos into low-resolution, grain-free, and energy-aware videos through a forward pass, and then reconstruct the original video in a backward pass, optimizing operations for resolution reduction, film grain removal, and energy awareness simultaneously.

Benefits of technology

This approach reduces energy consumption in video processing and display by enabling efficient encoding, transmission, and decoding of modified videos while maintaining the quality of the original video, and allows for direct display of low-resolution, grain-free, and energy-aware videos to further reduce energy usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024081627_22052025_PF_FP_ABST
    Figure EP2024081627_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Apparatuses and methods are disclosed for reversible production of video. Techniques disclosed include obtaining an input image and modifying the input image using an invertible framework run in a forward pass to perform modification operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image. Also disclosed techniques for reconstructing the input image using the invertible framework run in a backward pass to map a version of the modified image into a reconstructed input image. The invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] REVERSIBLE PRODUCTION OF RESCALED GRAIN-FREE AND ENERGY- AWARE VIDEO USING AN INVERTIBLE FRAMEWORK CROSS REFERENCE TO RELATED APPLICATIONS [1] This application claims the benefit of European Application No. 23306984.8, filed on November 16, 2023, which is incorporated herein by reference in its entirety. BACKGROUND [2] Film grain, a sensor noise that was introduced into content captured by analog cameras, is no longer present in content captured by digital cameras. Nevertheless, film grain may be a desirable feature in video production as it allows for a creative expression, and so, many times, it is intentionally added into video to enhance viewers’ experience. Typically, grainy videos are distributed to viewers through a video chain that includes encoding, transmitting, and then decoding the video for display. Processing such videos in the video chain consumes energy that is affected by the video image resolution, its color distribution, as well as by the addition of the film grain to the video content. Therefore, techniques that modify grainy video content into a low-resolution, grain-free, and energy-aware content can reduce the energy consumed by the processing of such modified content in a video chain. To be beneficial, such techniques should be able to adequately reconstruct the grainy video so that it does not compromise the viewers’ experience and to do so with a feasible computational complexity. SUMMARY [3] Aspects disclosed in the present disclosure describe methods for reversible production of video. The methods comprise obtaining an input image and modifying the input image into a modified image using an invertible framework run in a forward pass. The modifying includes operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image. The methods further comprise obtaining a version of the modified image (e.g., a coded and decoded version of the modified image) and reconstructing the input image using the invertible framework run in a backward pass to map the version of the modified image into a reconstructed input image. The invertible framework is trained, according to one or more loss functions, to simultaneously optimize the operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image and the inversion of these operations. [4] Aspects disclosed in the present disclosure describe apparatuses for reversible production of video. The apparatuses comprise at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, cause the apparatuses to obtain an input image and to modify the input image into a modified image using an invertible framework run in a forward pass. The modifying includes operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image. The instructions further cause the apparatus to obtain a version of the modified image (e.g., a coded and decoded version of the modified image) and reconstruct the input image using the invertible framework run in a backward pass to map the version of the modified image into a reconstructed input image. The invertible framework is trained, according to one or more loss functions, to simultaneously optimize the operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image and the inversion of these operations. [5] Aspects disclosed in the present disclosure describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for reversible production of video. The methods comprise obtaining an input image and modifying the input image into a modified image using an invertible framework run in a forward pass. The modifying includes operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image. The methods further comprise obtaining a version of the modified image (e.g., a coded and decoded version of the modified image) and reconstructing the input image using the invertible framework run in a backward pass to map the version of the modified image into a reconstructed input image. The invertible framework is trained, according to one or more loss functions, to simultaneously optimize the operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image and the inversion of these operations. [6] This Summary is provided to introduce a selection of concepts in a simplified form that is further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to limitations that solve any or all disadvantages noted in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS [7] FIG. 1 is a block diagram of an example system, according to aspects of the present disclosure. [8] FIG. 2 is a block diagram of a system illustrating a film grain analysis and synthesis applied to video content processed in a video chain, according to aspects of the present disclosure. [9] FIG.3 is a block diagram illustrating modeling of a film grain, according to aspects of the present disclosure.

[0010] FIG. 4 is a block diagram of a system illustrating an invertible framework applied to video content processed in a video chain, according to aspects of the present disclosure.

[0011] FIG.5 is a block diagram illustrating an invertible framework, according to aspects of the present disclosure.

[0012] FIG. 6 is a flowchart of an example method for reversible production of video, according to aspects of the present disclosure. DETAILED DESCRIPTION

[0013] Aspects of the present disclosure describe techniques for reducing energy consumption associated with the processing of video content through a video chain (including encoding, transmitting, and decoding) and with the displaying of the video content. To that end, an invertible framework is used in a forward pass to modify an original video into a low-resolution, grain-free, and energy-aware video, to be further processed through the video chain; after which process the modified video may be displayed or the original video (or a version thereof) may be reconstructed by the invertible framework in a backward pass and then be displayed. Energy-aware video (or energy-aware images) are video (or images) that were modified to consume less energy when rendered on a display. According to aspects, the invertible framework is simultaneously trained to perform the modification and the reconstruction, including operations of downscaling and upscaling images, removing and restoring film grain, and converting images into energy-aware images and reversing this process. A system for processing and displaying content is generally described (in reference to FIG.1), existing arts are then described that are related to film grain analysis and synthesis (in reference to FIGS.2- 3) as well as to rescaling and energy-aware image generation, followed by description of the aspects of the present disclosure (in reference to FIGS.4-6).

[0014] FIG. 1 illustrates a block diagram of an example system 100. System 100 can be embodied as a device including the various components described below and can be configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, can be embodied in a single integrated circuit, multiple integrated circuits, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple integrated circuits and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 100 is configured to implement one or more of the aspects described in this application.

[0015] The system 100 includes at least one processor 110 that can be configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 can include embedded memory, input and output interfaces, and various other circuitries as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which can include non-volatile memory and / or volatile memory, including, for example, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 can be an internal storage device, an attached storage device, and / or a network accessible storage device, for example.

[0016] System 100 includes an encoder / decoder module 130 configured to process data to provide encoded video data or decoded video data. The encoder / decoder module 130 can include its own processor and memory. The encoder / decoder module 130 represents module(s) that can be included in a device to perform encoding and / or decoding functions. Additionally, the encoder / decoder module 130 can be implemented as a separate element of system 100 or can be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.

[0017] Program code that is to be loaded into processor 110 or into encoder / decoder 130 to perform the various aspects described in this application can be stored in a storage device 140 and subsequently loaded into memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 can store one or more of various items during the performance of the processes described in this application. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0018] In several embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing functions that are needed during encoding or decoding. In other embodiments, however, memory external to the processing device (where, for example, the processing device can be either the processor 110 or the encoder / decoder module 130) can be used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140 that may comprise, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations.

[0019] The input to the elements of system 100 can be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal (COMP), (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0020] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select, for example, a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements that perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs some of these functions, including, for example, down-converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to a baseband. In one set-top box embodiment, the RF portion and its associated input processing element receive an RF signal transmitted over a wired (for example, cable) medium, and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Added elements can include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.

[0021] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed- Solomon error correction, can be implemented, for example, within a separate input processing integrated circuit or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface integrated circuits or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.

[0022] Various elements of system 100 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using a suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.

[0023] The system 100 includes a communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 can include, but is not limited to, a modem or network card. The communication channel 190 can be implemented, for example, within a wired and / or a wireless medium.

[0024] Data are streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signal of these embodiments is received over the communication channel 190 and the communication interface 150 which are adapted for Wi- Fi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.

[0025] The system 100 can provide an output signal to various output devices, including a display device 165, an audio device (e.g., speaker(s)) 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display device 165, the audio device 175, or other peripheral devices 185 using signaling such as AV.link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to system 100 using the communication channel 190 via the communication interface 150. The display device 165 and the audio device 175 can be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.

[0026] The display device 165 and the audio device 175 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display device 165 and the audio device 175 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0027] Aspects described in this disclosure relate to reversible operations including: rescaling of images; analysis, removal, and synthesis of film grain in image content; and construction of energy-aware images. These operations may be applied to images of a video to be encoded into a bitstream and then transmitted to a decoder end where the decoded video may be displayed to a user. Thus, according to aspects, images of a given video are modified into low-resolution, grain-free, and energy-aware images with the goal of reducing the energy consumed by the further processing (e.g., encoding / transmission / decoding) of the modified images and with the goal of reducing the energy consumed by displaying these images. Related arts concerning film grain analysis and synthesis, rescaling of images, and construction of energy-aware images are next described.

[0028] Film grain may be a desirable feature in video production. This is because viewers may perceive film grain as a pleasant noise that enhances the natural appearance of the video content. Originally, film grain was unintentionally introduced into image content during the physical process of exposure and development of photographic film. However, the process of capturing images by digital sensors does not generate film grain into the captured content. Consequently, digitally captured images can be perceptively noiseless, clear, with pronounced edges and monotonous regions that can worsen the subjective experience of the viewer. Therefore, “re- graining” a video before its distribution may be desired to improve its visual appearance. This is especially common in the movie industry, where many video content creators utilize techniques for adding film grain to their produced video content to add texture and warmth, or sometimes to create a sense of nostalgia (e.g., to present events that occurred in a previous era).

[0029] When added to the video content, film grain can be adapted based on different characteristics of the content, such as the intensity and / or saturation of the pixel values, the presence of texture, and the type of that texture. Different grain patterns can be selected depending on the desired look that is to be created. The film grain can also be adapted based on the color component it is applied to, resulting in a colored film grain. The addition of film grain is, therefore, a process that can be content or image-region dependent which takes into account image characteristics (such as intensity, texture, saturation) as well as grain patterns.

[0030] The random nature of the film grain makes grainy image content difficult to compress using traditional coding tools. Common parameters of encoding tools, such as those chosen to generate low bit rates, can remove film grain from grainy image content. However, if preserving the film grain in the image content is desired, encoding at high bitrates is required, which limits the objective of efficiently representing content coded into a bitstream. To overcome this issue, the grainy image content can be analyzed to model the film grain and then to remove it before the image content is encoded. Then, after decoding the image content, the film grain can be synthesized using the model so that it may be added back into the decoded image content.

[0031] Supplemental Enhancement Information for film grain was introduced into the H264 and the AVC standards. (See, C. Gomila, A. Kobilansky, “SEI message for film grain encoding”, ISO / IEC JTC1 / SC29 / WG11, ITU-T SG16 Q.6 document JVT-H022, Geneva, CH, May 2003, and C. Gomila, “SEI message for film grain encoding: syntax and results”, ISO / IEC JTC1 / SC29 / WG11, ITU-T SG16 Q.6 document JVT-I013, Trondheim, NO, July 2003.) Since then, film grain modeling has become part of modern video coding standards. (See, McCarthy S., et al., “Illustration of the film grain characteristics SEI message in HEVC”, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, document JCTVC-AM0023-v1, teleconference, April 2020, and McCarthy S., et al., “AHG17: Illustration of the film grain characteristics SEI message for VVC”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, document JVET-R0359-v2, teleconference, April 2020.) In particular, the analysis and synthesis of film grain is available in the Versatile Video Coding (VVC) standard (also known as ITU-T H.266 and ISO / IEC 23090-3), as described in reference to FIG.2.

[0032] FIG.2 is a block diagram of a system 200 illustrating a film grain analysis and synthesis applied to video content processed in a video chain. The system 200 includes a pre-processor 220, a film grain estimator 230, a video chain 240, and a film grain synthesizer 250. As shown, a film grain analysis 215 may be applied to generate a grain-free video 225 and a film grain model 235 out of an input grainy video 210. The grain-free video 225 may be fed into the video chain 240 to be encoded 242, transmitted via a communication channel 244, and decoded 246 at a decoder’s end, generating a decoded grain-free video 252. A film grain synthesis 255 may then be applied to reconstruct the grainy video 210 out of the decoded grain-free video 252 using the film grain model 254, generating a reconstructed grainy video 260. The film grain analysis 215 and synthesis 255 are further described below.

[0033] The grain analysis begins, by the pre-processor 220, with the removal of the film grain from the grainy video 210 to generate therefrom the grain-free video 225. The removal of the film grain can be done by a denoising filter, for example. Then, the grainy video and the grain- free video are fed to the film grain estimator 230. The film grain estimator 230 is employed to determine the film grain model 235, computing the model’s parameters. For example, uniform regions can be segmented out of the grain-free video based on which a mask can be generated. The mask can then be used to estimate the model’s parameters from a grain layer, computed as the difference between the grainy video 210 and the grain-free video 225. In the state-of-the- art methods of grain estimation, the extraction of uniform regions is critical since edges and textures can affect the estimation of the film grain strength and pattern.

[0034] The grain synthesis 255 is performed by the film grain synthesizer 250. Based on the film grain model 254, the film grain synthesizer can, first, generate the grain layer and, second, add that grain layer to the decoded grain-free video 252 to generate therefrom the grainy video. Note that a film grain model 235 (estimated at the encoder end) has to be encoded into the bitstream (in-band or out-of-band with the coded video) to make the film grain model 254 available at the decoder end. Information on film grain can be communicated as metadata through, for instance, a SEI message specified by Versatile Supplemental Enhancement Information (VSEI) standard (also known as ITU-T Recommendation H.274 and ISO / IEC 23002-7).

[0035] FIG. 3 is a block diagram illustrating modeling of film grain 300. In the example of FIG.3, the film grain is defined in the DCT domain and is modeled based on an intensity range of the Y component it is applied to. For example, for a Y block size of 16x16 pixels, a 16x16 patch of film grain can be generated in the DCT domain as follows. Samples from a random Gaussian distribution (defined in the DCT domain) can be assigned to a part of the 16x16 patch. Such part can be determined, for example, by a vertical and a horizontal cutoff frequencies in the DCT domain (as illustrated in FIG. 3). In this case, the film grain model’s parameters are the variance σ2of the Gaussian distribution and the two cutoff frequencies. Note that the cutoff frequencies determine the film grain’s shape – e.g., cutoff frequencies that are equal will result in a circular film grain shape. As shown in FIG.3, the DC component is set to zero in order to obtain a centered film grain patch in the spatial domain. The value of the variance σ2of the Gaussian distribution depends on the mean value of the Y block to which the film grain patch is to be added, making the synthesis of the film grain region dependent.

[0036] Hence, a film grain can be synthesized based on a film grain model estimated according to the above described techniques (as illustrated in FIGS. 2 and 3) and then added to the decoded video content during a post processing stage. Further techniques include synthesizing film grain having a pattern that may be unique for the capturing camera and having an intensity level that depends on the video’s content. Different scaling factors can be used to control the intensity of the film grain, tuning it to the intensity of the content. Typically, the intensity level of the film grain is kept constant for a given block (e.g., of blocks of the same size) in a given frame. Moreover, depending on the desired artistic look, the film grain intensity can be controlled based on the texture of the content it is applied to (e.g., the more textured the regions, the higher the intensity of the applied film grain).

[0037] In the AOMedia Video 1 (AV1) standard, a film grain template is generated and applied to the video content. Therein, the film grain is modeled with an autoregressive process and represented by a piecewise linear function that includes the film grain strength for each of the Y, Cb, and Cr components. At the re-graining stage of the video decoding, a 64×64 luma film grain template is generated using the film grain model extracted from the bitstream. Blocks of size 32x32 are extracted from the 64x64 film grain template and added to respective 32×32 luma blocks to reconstruct the video content with the film grain. In this implementation, the film grain is added to the video image, using a block size that is fixed and not dynamically adapted to the image region.

[0038] Conventional techniques for film grain generation are applied to other digital content (outside of the context of video compression). Such techniques enable adding film grain to digital content without requiring prior analysis of that content. (See, e.g., P. Schallauer and R. M ̈orzinger, “Film grain synthesis and its application to re-graining,” in Image Quality and System Performance III, vol.6059, p.60590Z, International Society for Optics and Photonics, 2006, hereinafter “Schallauer”; Newson, A., at el., “A stochastic film grain model for resolution-independent rendering,” in Computer Graphics Forum, vol.36, pp.684–699, Wiley Online Library, 2017, hereinafter “Newson”). In Schallauer, a parametric model based on texture statistics is used to generate a film grain template. The film grain template is then decomposed into a steerable pyramid, using a linear, multi-scale, and multi-orientation image transform. Each scale and orientation of the pyramid is analyzed with respect to several statistical texture features, including minimum and maximum gray values and correlation of sub-bands. The synthesis starts with random noise which ensures high spatiotemporal variations. The algorithm produces film grain that matches the template very well, while the random noise-based approach inherently provides realistic spatial and temporal variations. However, the resulting synthesis is not content-adapted and is therefore uniformly applied to any image, whatever the underlying image characteristics (e.g., structures and textures) are. In Newson, a stochastic model is used to approximate the physical reality of film grain. To model the film grain, a Boolean inhomogeneous model is applied to mimic the analog photographic process as closely as possible. The model corresponds to uniformly distributed disks using a Poisson process of variable intensity which determines the amount of film grain with respect to the local image gray level. The film grain rendering algorithm is modeled with a single Gaussian kernel using a Monte Carlo simulation which performs simultaneous filtering and discretization of the film grain model. A wide range of grain sizes and intensities can be generated by varying the parameters of this model, among which are the average film grain radius and its standard deviation. Some larger values of these parameters accentuate the "grain" of the rendered result.

[0039] Deep-learning-based techniques have also been used for film grain generation. For example, using a deep-learning-based network, synthesis of film grain can be achieved by translating a given grain-free input image into a corresponding grainy output image. (See, Ameur Z., et al., “Deep-based film grain removal and synthesis,” arXiv preprint arXiv:2206.07411, 2022, hereinafter “Ameur”). The proposed model, in Ameur, is a U-Net with residual blocks trained with a discriminator in a conditional GAN (generative adversarial network) framework. The synthesis model is conditioned on both the input image and some expected grain intensity level, which allows the synthesis to be content-dependent with a controllable intensity level. The above-mentioned techniques for film grain generation all suffer from a major drawback: the properties of the synthesized film grain are limited by its size, shape, and intensity, resulting in the generation of limited grain types and, consequently, poor diversity.

[0040] Another deep learning framework is described in EP patent application 23305106.9. This framework leverages a style vector, used to gather film grain properties, that allows a better reconstruction at the decoder side. However, applying this framework requires transmitting the style vector in the bitstream and is limited by the large number of the deep network parameters. Other techniques related to film grain modeling and synthesis can be found in EP patent applications 21305329.1, 22306040.1, and 22305671.4.

[0041] As mentioned above, removing the film grain from the video allows for a more efficient representation of the video content in the bitstream. Furthermore, rescaling the video (into a lower resolution frame size) can reduce the computational complexity endured and the storage space required when processing the video in a video chain 240. Rescaling the video content also simplifies fitting the video to low-resolution displays. Yet, rescaling may come at the cost of maintaining adequate quality level of the rescaled video content. Additionally, rescaling may result in a loss of information, so that restoring the video to its original resolution may be challenging. To maximize the performance of the video restoration process, both the downscaling and the upscaling operations can be jointly learned. For example, an auto- encoder-based framework can be trained to learn the optimal low-resolution image that maximizes the performance of restoring the respective high-resolution image; that is, a downscaling method is devised that takes into account the upscaling process. Such a framework is trained in an unsupervised manner, with no assumptions made about how the original high- resolution image is downscaled, to learn the essential information for upscaling in an optimal way.

[0042] In a different paradigm, the downscaling and the upscaling operations are modeled using an invertible bijective transformation. (See, Xiao M., et al., “Invertible image rescaling,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part I 16, pages 126–144. Springer, 2020, thereinafter “Xiao”). In Xiao, in a forward pass, the framework performs the downscaling process by producing visually pleasing low-resolution images while capturing the distribution of the lost information using a latent variable that follows a specified distribution. Meanwhile, the upscaling process is made tractable such that by inversely passing a randomly drawn latent variable with the low- resolution image through the network in a backward pass, the high-resolution image is reconstructed.

[0043] The state of the art of generating energy-aware images can be classified into four categories. The first category includes methods that are based on histogram or look-up-tables which reduce the energy by, for example, manipulating the histogram. It could be a decrease, or a modification, of the number of empty bins with the goal of saving energy. The second category includes methods that rely on luminance-based or color-based transformation of the original images. The resulting colors or luminance values are chosen so that they consume less energy when displayed, and according to another criterion, for example, values that fulfill luminance and contrast-based Just-Noticeable-Difference (JND) models. The third category includes methods that require using side information (such as saliency maps, objective quality metrics, depth information, or gaze tracking) to apply, for example, region-based image modifications that reduce the energy requirements. The fourth category includes more recent methods that leverage the power of deep networks that are trained in a fully unsupervised manner, as there is no ground truth. Deep networks are then used to maximize the quality of the output image under a constraint of a power saving rate target. In this category, for example, an invertible energy-aware network can be used that produces invertible energy-aware images, where the original images can be recovered if required. (See, Olivier Le Meur and Claire- Hélène Demarty, “Invertible energy-aware images,” IEEE Signal Processing Letters, 2023; and EP patent application 23305294.3). Other energy-aware based applications are described in EP patent applications 21305604.7 and 22306719.0.

[0044] Invertible neural networks (INN) are used in the art to perform invertible processes, as described in the following examples. An INN can be used to produce invertible grayscale images, where the lost color information is encoded into a set of Gaussian distributed latent variables. (See, Zhao R., et al., “Invertible image decolorization,” IEEE Transactions on Image Processing, 30:6081–6095, 2021, hereinafter “Zhao”.) In Zhao, the original color version can be efficiently recovered by randomly re-sampling a new set of Gaussian distributed variables and using it as an input together with the synthetic grayscale, through reverse mapping. An invertible denoising network (InvDN) can also be used to transform a noisy input into a low- resolution clean image and a latent representation containing noise. (See, Liu Y., et al., “Invertible denoising network: A light solution for real noise removal,” in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 13365–13374, 2021, hereinafter “Liu”.) In Liu, to discard noise and restore the clean image, InvDN replaces the noisy latent representation with another one sampled from a prior distribution during reversion. In addition to denoising, an INN-based method can disentangle noise from the high-frequency image information. (See, Du W., et al., “Hierarchical disentangled representation for invertible image denoising and beyond,” arXiv preprint arXiv:2301.13358, 2023.)

[0045] Aspects of the present disclosure provide an invertible framework that when applied (in a forward pass) to a high-resolution, grainy, and not energy-aware video (namely, an original video) it modifies the video into a low-resolution, grain-free, and energy-aware video; and when applied (in a backward pass) to the modified video it reconstructs the original video (or a version thereof). Thus, according to aspects, the reconstructed video may be a grain-free version, an energy-aware version, or a grain-free and energy-aware version of the original video. The invertible framework is simultaneously trained to perform the tasks underlying the modification of the original video and its reconstruction. That is, the tasks of 1) downscaling and upscaling, 2) removing and restoring film grain, and 3) generating energy-aware content and reversing this process. Aspects described herein allow for efficient encoding, transmitting, and decoding of the modified video, without compromising the quality of the reconstructed original video at the decoder end. No additional information (e.g., metadata describing the removed film grain or fine image details) is required to be coded into the video bitstream to enable reconstruction of the original video based on the modified video. Furthermore, since the modified video (that is, the low-resolution, grain-free, and energy-aware video) is produced by the invertible framework with a viable image quality, according to aspects, it can be displayed (instead of the reconstructed original image) to reduce display energy consumption (i.e., reduce the energy consumed by the video rendering).

[0046] In contrast, existing methods independently carry out, by different modules, the processes of rescaling, removing / synthesizing film grain, and energy-aware conversion. Although, rescaling and denoising tasks have been proposed to be performed using the same INNs, noise and film grain do not share the same inherent characteristics, and so the task of denoising and film grain removal are not equivalent. Additionally, at least with respect to the task of film grain synthesis at the decoder side, existing methods require conveying the parameters of a film grain model in the bitstream to enable film grain synthesis with high fidelity to the original film grain that was removed from the video at the encoder side. Moreover, even with access to these model parameters, existing methods still suffer from discrepancies between the originally removed film grain and the synthesized film grain.

[0047] Hence, according to aspects, a deep-learning based invertible framework is employed that, through a multitasking process, produces a low-resolution, grain-free, and energy-aware image from a high-resolution, grainy, and not energy-aware image. These three tasks (rescaling, film grain removal / restoration, and energy-aware conversion) are concurrently realized with one light weight invertible framework. By generating content that can be represented with low bitrate (but still with acceptable quality), the invertible framework is capable of reducing the energy consumed by the encoding, transmission, decoding and displaying of the generated content. As the three tasks are concurrently realized by a light network, it also results in a reduction of the computation load compared to using three networks (each performing one task) with a large number of parameters, as in the current state of the art.

[0048] Because of the invertible property of the framework, only the low-resolution image produced by the framework (when run in a forward pass) should be submitted in the bitstream to enable reconstruction of the original image by the framework (when run in a backward pass) at the decoder end. This is accomplished by mapping the high-frequency information of the original image (including fine image details and film grain) into a set of random variables distributed according to a standardized Gaussian distribution. According to aspects, such mapping involves latent variables distributed according to a Gaussian distribution whose mean and variance are conditioned on the low-resolution image. Hence, it is not necessary to transmit any additional information (representing the high-frequency information) in addition to the low-resolution image (representing the low-frequency information) in the bitstream to reconstruct the original image at the decoder end. Furthermore, due to the disentanglement of the high-frequency information into information associated with the fine image details and into information associated with the film grain, the framework is capable of reconstructing also a grain-free version of the original image. The invertible framework is next described in reference to FIGS.4-6.

[0049] FIG.4 is a block diagram of a system 400 illustrating an invertible framework applied to video content processed in a video chain. The system 400 includes an invertible framework (that is run in a forward pass 420 and in a backward pass 460), a video chain 440 (including an encoder 442, a communication channel 444, and a decoder 446), and a display 480. The components of the system 400 can be implemented by respective components of the system 100 described in FIG.1. According to aspects, the invertible framework 420, 460 is trained to perform: 1) downscaling and upscaling, 2) removing and restoring film grain, and 3) generating energy-aware content and reversing this process. This is done with the goal of reducing the overall energy consumed in processing a given video in the video chain 440 – that is, reducing the energy consumed by the encoding 442, the transmission through a communication channel 444, and the decoding 446 of a video. This is also done with the goal of reducing the energy consumed by displaying a video, that can be achieved by directly displaying a low-resolution, grain-free, and energy-aware video (e.g., the decoded video 445, as illustrated by the dashed- arrow in FIG.4).

[0050] Thus, running in a forward pass 420, the invertible framework is trained to produce from a high-resolution, grainy, and not energy-aware video 410 (also referred to herein as the original or the input video) a low-resolution, grain-free, and energy-aware video 425 that is then fed into the encoder 442 of the video chain 440. Running in a backward pass 460, the invertible framework is trained to produce from the decoded low-resolution, grain-free, and energy-aware video 445, provided by the decoder 446 of the video chain 440, a reconstructed high-resolution, grainy, and not energy-aware video 470 (also referred to herein as the reconstructed video). According to aspects, running in a backward pass 460, the invertible framework is also trained to reconstruct a grain-free version of the original video (e.g., a high resolution, grain-free, and not energy-aware video), an energy-aware version of the original video (e.g., a high resolution, grainy, and energy-aware video), or both (e.g., a high resolution, grain-free, and energy-aware video). As the invertible framework 500 is trained so that the low- resolution, grain-free, and energy-aware video 425 corresponds to a minimum viable image quality, this video 425 or its decoded version 445 can be displayed directly, without going through the backward pass 460 to generate the reconstructed video 470. In this case, it enables further energy savings at the display. The invertible framework and its training are next described in reference to FIG.5.

[0051] FIG. 5 is a block diagram illustrating an invertible framework 500. Aspects of the framework 500 are described herein with respect to images (frames) of a video sequence. However, the aspects are applicable to images extracted from other contents, such as 360° video captured by camera(s) or generated by a computer. The invertible framework’s underlying model is trained through forward running 420 and backward running 460 of the framework 500. Thus, during the application of the trained model, by virtue of being invertible, video frames (images) can be pushed through the framework in both directions as indicated by the double-arrows. As shown in the example of FIG.5, the framework 500 includes three main components: a wavelet transformer 520, an invertible neural network (INN) 530 that can contain multiple INN units (e.g., L units 530.1-530.L), and a latent encoding unit 550, the operation of which is further described below.

[0052] The wavelet transformer 520 may be implemented by various types of wavelet transforms, such as the Haar wavelet transform. Employing the wavelet transformer 520 in a forward pass 420, an input image 510 can be decomposed into four low-resolution images for each color channel – one containing the low-frequency image information and three containing the high-frequency image information in the vertical, horizontal, and diagonal spatial directions. Formally, a wavelet transform decomposes the input image 510, including features ^^^^∈ ^ℝ^∗௪∗^, into one low-frequency image, including features ^^^^௪ ∈ ℝమ∗^మ∗^, and into three high- images, including features ^^ ∈ ℝଷ∗^మ∗^మ∗^; where respectively, the height and width of the input image 510, and ^^ represents the number of color channels. Generally, the low-frequency image ^^^^௪represents the overall structure and coarse features of the high-resolution image 510 (e.g., can be produced by average pooling), while the three high- frequency images ^^^^^^contain residuals in the vertical, horizontal, and diagonal directions respectively, which represent fine image details (e.g., edges and texture) and film grain. Employing the wavelet transformer 520 in a backward pass 460, the transformer can reconstruct the input image 510 based on a reconstructed version of the low-frequency image ^^^^௪and the three high-frequency images ^^^^^^provided by the invertible neural network 530. In an aspect, another rescaling decomposition can be used to replace the wavelet transformer.

[0053] The input image 510 can be represented in RGB format or in any other color format, such as YUV, Lab, and HSV formats, as well as a gray scale format. For example, when the input image 510 is an RGB image, in a forward pass, the output of the wavelet transformer 520 includes images rescaled by 2, including: three low-resolution images of the red, green, and blue channels (i.e., an RGB low-resolution image) that include the low-frequency component of the image 510 (i.e., ^^^^௪) and nine low-resolution gray scale images that contain the high- frequency component of the image 510 (i.e., ^^^^^^). In a backward pass 460, the wavelet transformer 520 reconstructs the RGB image based on a reconstructed version of the low- frequency component, ^^^^௪, and the reconstructed high-frequency component, ^^^^^^, provided by the invertible neural network 530.

[0054] In an aspect, the input image 510 may be decomposed into rescaled images of different dimensions. For example, the image 510 may be split into an RGB image of w / 2 width and h / 2 height (that is, the low-frequency component ^^^^௪) and three gray scale images of w / 2 width and h height (that is, the high-frequency component ^^^^^^). In another aspect, one or more convolutional layers can be added after the wavelet transformer 520 and / or in between pairs of INN units 530 to further refine the splitting between the low-frequency image information and the high-frequency image information.

[0055] Following the wavelet decomposition, the low-frequency component ^^^^௪and the high- frequency component ^^^^^^of the image 510 are fed (in the forward pass 420) into the invertible neural network 530, to be processed through the INN units. Each INN unit can be implemented by a coupling layer architecture. The forward mapping of an INN unit i (e.g., 530.i) can be expressed as follows: ℎ^ା^^ ൌ ℎ^^ ^ ^^൫ℎଶ^൯ (1) where ℎ^^is a vector containing high-frequency information (e.g., ℎ^ଶ ൌ ^^^^^^ ), where ^^,^^, ^^ are affinetransformations, and where the operators “^” and “⊙” denote point-wise multiplication, respectively. Thus, an additive transformation is applied to the low-frequency vector ℎ^^and an affine transformation is applied to the high-frequency vector ℎ^ଶ . The outputs of an INN unit 530.i in the forward pass 420 are thus vectors ℎ^^ା^and ℎଶ^ା^.

[0056] Consequently, the inverse operation (backward mapping) of the INN unit i can be expressed as follows: ℎ^^ ൌ ℎ^ା^^ െ ^^൫ℎଶ^൯ (3) The outputs of an INN unit in a backward pass 460, the outputs vectors ℎ^^ and ℎ^ଶ of an INN unit 530.1 represent the reconstructed low-frequency component ^^^^௪and the reconstructed high-frequency component ^^^^^^, respectively, that in turn are fed to the wavelet transformer 520 to reconstruct therefrom the high-resolution, grainy, and not energy-aware image 510. In an aspect, any number L ofINN units can be used (e.g., ^^ ൌ 8). In another aspect, other invertible architecture can be usedfor the INN units (e.g., residual networks). In yet another aspect, any number of combinations of wavelet transformer 520 and INN units 530 can be used in the invertible framework 500.

[0057] The invertible neural network 530 interfaces with a latent encoding unit 550, as illustrated in FIG. 5. In a forward pass 420, the latent encoding unit 550 encodes the high-frequency information provided by the last INN unit 530.L – that is, ^̃^ ≡ ℎ^ଶ – into a set ofGaussian distributed random variables ^^ ∽ ^^^0,1^. In the backward 460, the latentencoding unit 550 decodes a set of Gaussian distributed random variables ^^ ∽ ^^^0,1^ into ^̃^,recovering the high-frequency information. And so, no auxiliary information has to be delivered to the decoder end to reconstruct (in a backward pass 460) the high-resolution image 510 from the low-resolution image 570. However, to enable image-adaptivity, and, therefore, better recovery of the high-frequency information, according to aspects, a Gaussian distribution whose mean and variance are conditioned on the low-frequency information ℎ^^(that is, the low-resolution image 570) can be enforced on the high-frequency ^̃^ by the latent encoding unit 550. This can be accomplished by employing a forward mapping including a one-side affine coupling layer that normalizes ^̃^ into a standardized Gaussian distributed vector ^^, as follows: ^^ ൌ൫௭^ିథ ^^భ^^൯,^^^^^^ℎ ^^ ∽ ^^^0,1^ (5)where ^^, ^^ can be layers (as defined in Huang, G., at el., “Densely Connected Convolutional Networks,” CVPR 2017).

[0058] The reverse mapping, used in the backward pass 460, can be expressed as follows: ^̃^ ൌ ^^ ⊙ exp൫^^^ℎ^^^൯ ^ ^^^ℎ^^^,^^ℎ^^^^^^ ^̃^ ∽ ^^^^^^ℎ^^^, exp൫^^^ℎ^^^൯^ (6)Note a uniform distribution.

[0059] As shown in FIG. 5, the vector ^̃^ 540 can be disentangled into two parts: onerepresentative of film grain information ^^ீ(i.e., the removed grain) and one representative of image detail information ^^^(i.e., the down-scaled out image’s fine details). This allows the latent encoding unit 550 to independently learn ^^ீand ^^^. This disentanglement enables the reconstruction of a grain-free version of the input image 510 (e.g., if asked for by the viewer). Hence, by inversely running 460 the framework 500 described above, feeding the framework 500 with random samples ^^ 560 and with a law-resolution image (that may be an encoded and decoded version of the image 570 provided by the framework 500 in the forward pass), the original image 510 can be reconstructed, or a grain-free and / or energy-aware version of this image. This reconstruction can be done without the need of transmitting any auxiliary information (metadata that describes the high-frequency information) through the video chain 440 to the decoder end.

[0060] The underlying model of the invertible framework 500 can be optimized by training the framework based on an image category. Subsequently, several models can be learned based on respective image categories, such as outdoor images, images with persons, gaming images, user interface (UI) images, bright or dark images, and so on. In this manner, the framework’s production of low-resolution, grain-free, and energy-aware images and their reconstruction can be refined per image category.

[0061] Furthermore, to generate energy-aware images, the framework 500 can be trained separately for each chosen value of power reduction rate, namely, R (generating one set of model parameters per value of R). Alternatively, the training can be conditioned on R, thus generating one single model capable of reconstructing different images for different respective values of R. In an aspect, R can be fed as an input to the framework 500, for example, concatenated to the input image 510 or injected into any number of the INN units 530 of the framework. In another aspect, an additional invertible unit (e.g., an INN unit or a latent encoding unit) can be added to the framework 500 that is dedicated to the generation of energy- aware images. For example, such a dedicated invertible unit can be added to the front end of the framework 500 and can be fed with the high-resolution image 510 and a chosen value for R.

[0062] Next described are the objectives of the invertible framework 500 when run both in a forward pass 420 and a backward pass 460. These objectives can be formulated by loss functions according to which the underlying model of the framework 500 is optimized and its parameters are determined.

[0063] In the forward pass 420, the operation of the framework 500, as defined by the framework’s model, denoted by ^^, can be expressed as follows: ^^൫^^^൯ ൌ ^^^^^ / ோ, ^^^, (7)where ^^^denotes the input high-resolution, grainy, and not energy-aware image 510, where ^^^^ / ோdenotes the corresponding low-resolution, grain-free, and energy-aware image 570 with^^ ∈ ^0,1^ being the target energy reduction rate, and where ^^ denotes the random samples 560associated with the high-frequency information. To guide the model ^^ to generate a visually pleasant low-resolution, grain-free, and energy-aware image, we use a ground-truth image ^^^(created by down-sampling the original image 510 via, for example, a bicubic filter) to minimize the following loss function: ℒ^^^௪൫^^^^ / ோ,^^^൯ ൌ^ ே∑ே ^ୀ^ ฮ^^^^⁄ ோ ^^^^ െ ^^^^^^^ฮଶ, (8) where N denotes the be used (in addition or instead of the loss function shown in equation (8)) to maximize the similarity (or minimize the distance) between ^^^^ / ோand ^^^. For example, a structural similarity index measure (SSIM) loss function can be used: ℒௌௌூெ ൌ 1 െ ^^^^^^^^൫^^^^ / ோ,^^^൯. (9)

[0064] For ^^^^ / ோto be an energy-aware version of ^^^, the power consumed in processing ^^^^ / ோis defined to be reduced by ^^ compared to the power consumed in processing ^^^. Thus, in an aspect, the image power of ^^^^ / ோand ^^^, respectively, ^^^ and P, can be computed as the sum of the power of each led of an RGBW OLED display (following the power model proposed in Demarty CH, et al., “Display power modeling for energy consumption control,” ICIP 2023). In another aspect, denoting the luminance components of ^^^^ / ோand ^^^, respectively, by ^^^ and ^^, the image power of these components can be modeled, respectively, as: ^^^ ൌ ^ ∑ே ^^^ఊand ^^ ൌ^∑ே ^^ఊே^ୀ^ ^ே^ୀ^ ^, (10)where ^^ denotes the pixels considered in ^^^ and in ^^. The objective with respect to power consumption can then be expressed as: ℒ^^௪ ൌ ฮ^^^ െ ^1 െ ^^^^^ฮ^, (11)where ^1 െ ^^^^^ is the desired

[0065] Next, to enforce ^^ 560 to follow a standardized Gaussian distribution, the log-likelihood of the probability density function ^^^^^^ of the standardized Gaussian distribution is maximized as follows: ℒ^^^ ൌ െ log൫^^^^^^൯ ൌ െlog ^^ ವೋ exp ^െ ^ଶ ‖^^‖ଶ^^, (12) where ^^^is the

[0066] In the inverse pass 460, the operation of the framework 500 is defined by the inverse model ^^ି^as follows: ^^ି^^^^^^ / ோ, ^^) ൌ ^^^^, (13)where ^^^^is the reconstructed version of original input image ^^^510. Thus, the inverse model ^^ି^is a function of random samples ^^ 560 that are generated according to a standardized Gaussian distribution ^^^0,1^ and a function of the low-resolution, grain-free, and energy- aware image ^^^^ / ோ570.

[0067] Hence, in the backward pass 460, ^^ is decoded into ^̃^ by the latent encoding unit 550. To force the disentanglement between the film grain information and the image detail information, ^̃^ is segmented into ^^ீ(representing the film grain) and ^^^(representing the fineimage details) – that is, ^̃^ ൌ ^^^ீ , ^^^^ – and two respective loss functions are minimized, asfollows. To restore the high-resolution grainy image ^^^, ^̃^ ൌ ^^^ீ , ^^^^ is used and the followingreconstruction loss function is minimized to obtain ^^^^: ℒ^^^^^൫^^^^, ^^^൯ ൌ^ ே∑ே ^ୀ^ ฮ^^ି^^^^^^ୡ⁄ ୖ , z^^^^^^|௭^ୀ^௭ಸ,௭ವ^ െ ^^^^^^^ฮ^, (14) version of ^^^, we use a ground-truth image ^^^(e.g., created by smoothing the original image510). By setting ^^ீ to zero: ^̃^ ൌ ^0, ^^^^, the following reconstruction loss function is thenminimized to obtain ^^^^: ℒ^^^^^^^^^^ ,^^^^ ൌ^ ே∑ே ^ୀ^ ฮ^^ି^^^^^^^⁄ ோ , ^^^^^^^^|௭^ୀ^^,௭ವ^ െ ^^^^^^^ฮ^, (15)

[0068] Hence, according to aspects, the model of the framework 500 may be optimized in one or in multiple stages, using a combination of the loss functions described above. For example, optimization can be carried out with respect to the rescaling and the film grain removal and synthesis, by minimizing the following total loss function: ℒ௧^௧^^ ൌ ^^^ℒ^^^௪ ^ ^^ଶℒ^^^ ^ ^^ଷℒ^^^^^ ^ ^^ସℒ^^^^^ (16)Then, the model ℒ^^௪and ℒௌௌூெas follows: ℒ^^^^ ௧௨^^^^ ൌ ℒ௧^௧^^ ^ ^^ହℒ^^௪ ^ ^^^ℒௌௌூெ, (17)where ^^^, ^^ଶ, ^^ଷ, ^^ସ, ^^^^, ^^ଶ, ^^ଷ, ^^ସ, ^^ହ, ^^^^ ൌ ^40, 1, 1, 1, 1^^6, 1^^3^. of the power loss function, ℒ^^௪, from the image power as defined in equation (10) the luminance is retrieved from the RGB channels as follows: ^^ ൌ 0.299^^ ^ 0.587^^ ^ 0.114^^ (18)From this equation, it appears that a large reduction of the B (blue) channel is needed, compared to R (red) and G (green) channels, to significantly reduce the luminance pixel values. To avoid a color shift in the output of the model, ℒ^^^௪can be therefore adapted during the fine-tuning, to compute a weighted mean across the RGB channels with the same weights used in conversion. The same weighting is used while computing ℒ^^^௪during fine-tuning: ℒ^^^௪^^^^^,^^^^ ൌ^∑ே ^ୀ^ ೃ ே0.299ฮ^^^^⁄ ோ ^^^^ െ ^^^^^^^ฮଶ ^ 0.587ฮ^^^^⁄ ோ ^^^^ െ ^^^^^^^ฮଶ ^ 500. For example, constraints can be added that are related to the further processing, in the video chain 440, of the produced image ^^^^ / ோ425. Such constraints may include: ^ Energy-aware constraints that are related to a bitrate constraint, e.g., by minimizing the bitrate. ^ Robustness constraints to the encoding and decoding of the image ^^^^ / ோ. This can be done for example by including the encoding and the decoding steps in the framework and adding constraint(s) to minimize changes in ^^^^ / ோby the encoding and the decoding processes. ^ Robustness constraints for the reconstructed input image ^^^^, e.g., by an additional SSIM loss between ^^^^and ^^^.

[0071] The framework, during its operating and / or training phases, can be applied to the entire image (e.g., film grain removal and synthesis performed for the entire image). However, in further aspects, the framework can be selectively applied to region(s) of interest in the image. Thus, in an aspect, the input and the output of each module in the framework (e.g., 520, 530, and / or 550) can be modulated with respect to region(s) of interest indicated by a map. For example, a JND map (e.g., as computed in EP patent application 21305604.7) can be used to ensure that alterations on the input image 510 remain below a visibility threshold. In another example, a saliency map can be used to protect, e.g., during the training, visually important information. Thus, a map (such as the JND-based map or the saliency-based map) can be used as another input to the framework 500. Alternatively, the map can be applied to the computation of one or more of the loss functions used for the training of the framework – for example, the map’s pixels can be used as weights for respective elements of a loss function.

[0072] In an aspect, some further disentanglement between grayscale information and color information can be done, which allows processing the luma only while fine tuning for the energy-aware task. In this case the output of the framework is a disentangled low-resolution,grain-free, and energy-aware version of the luma component of the image, ^^^, corresponding toan energy reduced version of the luma component Y of the output image before the fine tuning. In this case, the RGB low-resolution, grain-free, and energy-aware color image is reconstructed based on luminance ratio as described below, for each pixel’s spatial location i: ^^^ ൌ ^^^^ ^^^ ൈ^^(20)where, ^^^ represent input pixels ^^^^ i. ^^^ and ^^^^ represent the original luminance and the reconstructed one, respectively.

[0073] FIG. 6 is a flowchart of an example method 600 for reversible production of video, according to aspects of the present disclosure. As illustrated in FIG.6, the method 600, in step 610, obtains an input image. Then, in step 620, the input image is modified using an invertible framework 500, running in a forward pass. According to aspects, the modification of the input image includes operations of reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image. The invertible framework 500 is trained, according to one or more loss functions (e.g., equations (16) and (17)), to simultaneously optimize these modification operations (including the reducing, the removing, and the converting operations) and the inversion of these operations. As explained above, the modified image, because it is a low-resolution, grain-free, and energy- aware image, consumes less power when processed in a video chain 630.

[0074] The method 600 may further operate on a version of the modified image to reconstruct the input image obtained in step 610. Thus, in step 640, a version of the modified image is obtained that may be a coded and then decoded version of the modified video provided by step 620. Then, in step 650, the input image is reconstructed using the invertible framework 500, running in a backward pass, to map the version of the modified image into a reconstructed input image. To that end, the high-frequency information of the input image can be disentangled into a first set including image detail information of the input image (e.g., ^^^in FIG.5) and into a second set including film grain information of the input image (e.g., ^^ீin FIG.5). Then, in a backward running of the invertible framework, the first set and the second set of the latent variables can be used to reconstruct the input image (that is, to reconstruct the grainy and not energy-aware input image), while the first set can be used (setting ^^ீto zero) to reconstruct a grain-free version of the input image.

[0075] Note that the image detail information of the input image may include information that is lost due to the downscaling operation and information that is lost due to the energy-aware conversion operation. Thus, in an aspect, the high-frequency information of the input image can be disentangled into: a first set including image detail information that is lost due to the downscaling operation; a second set including image detail information that is lost due to the energy-aware conversion operation; and a third set including the film grain information. In this aspect, in a backward running of the invertible framework: the first, the second, and the third sets can be used to reconstruct the input image (that is, to reconstruct the grainy and not energy- aware input image); the first and the second sets can be used to reconstruct a grain-free version of the input image; the first and the third sets can be used to reconstruct an energy-aware version of the input image; and the first set can be used to reconstruct a grain-free and energy- aware version of the input image.

[0076] As mentioned above, in an aspect, modified images (i.e., the low-resolution, grain-free, and energy-aware images provided by step 620 or obtained by step 640 from the video chain 630) can be directly displayed to a viewer, thus reducing the energy consumed by the display.

[0077] The modification of the input image (step 620), according to aspects described herein, may begin with employing a wavelet transformer 520 to transform the input image 510 (by a wavelet transform such as the Harr wavelet) to generate a low-frequency component of the input image and a high-frequency component of the input image. The low-frequency and the high-frequency components are then processed by one or more invertible neural network units 530 to generate the modified input image 570 and high-frequency information 540. The high- frequency information 540 is then mapped, by a latent encoding unit 550, into random variables 560 that are distributed according to a standard Gaussian distribution. In an aspect, the mapping enforces on the high-frequency information 540 a Gaussian distribution with a mean and a variance that are based on the modified image (e.g., see equation (5)).

[0078] According to further aspects, the invertible framework 500 is trained according to one or more loss functions. The one or more loss functions may include a combination of the following loss functions: a loss function that characterizes a difference between a down- sampled and grain-free version of the input image and a modified image (e.g., equation (8)); a loss function that enforces a standardized Gaussian distribution of random variables generated in a forward pass by the invertible framework (e.g., equation (12)); a loss function that characterizes a difference between an input image and a reconstructed grainy image (e.g., equation (14)); a loss function that characterizes a difference between a grain-free version of an input image and a reconstructed grain-free image (e.g., equation (15)); a structural similarity index measure loss function that characterizes a difference between a down-sampled and grain- free version of the input image and a modified image (e.g., equation (9)); and a loss function that characterizes a difference between a power measure of a down-sampled and grain-free version of an input image and a power measure of a modified image, constraining the power measure of the modified image by a power reduction rate (e.g., equation (11)).

[0079] The illustrations of the aspects described herein are intended to provide a general understanding of the structure, function, and operation of the various aspects. The illustrations are not intended to serve as a complete description of all of the elements and features of apparatuses and systems that utilize the structures or methods described herein. Many other aspects may be apparent to those of skill in the art upon reviewing the disclosure. Other aspects may be utilized and derived from the disclosure, such that structural and logical substitutions and changes may be made without departing from the scope of the disclosure. Accordingly, the disclosure and the figures are to be regarded as illustrative rather than restrictive.

[0080] The description of the aspects is provided to enable the making or use of the aspects. Various modifications to these aspects will be readily apparent, and the generic principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.

Claims

CLAIMS 1. A method, comprising: obtaining an input image; and modifying the input image into a modified image using an invertible framework run in a forward pass, the modifying includes modification operations of: reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

2. The method according to claim 1, further comprising: obtaining a version of the modified image; and reconstructing the input image using the invertible framework run in a backward pass to map the version of the modified image into a reconstructed input image.

3. The method according to claim 2, wherein: the modified image is a frame from a video to be processed in a video chain, wherein the video is encoded at an encoder end and decoded at a decoder end to generate a decoded video; and the obtained version of the modified image is a corresponding frame from the decoded video.

4. The method according to any one of the preceding claims, wherein the modifying of the input image comprises: transforming, by a wavelet transformer, the input image to generate a low-frequency component of the input image and a high-frequency component of the input image.

5. The method according to claim 4, wherein the modifying of the input image comprises: processing, by one or more invertible neural network units, the low-frequency component and the high-frequency component to generate the modified image and high- frequency information of the input image.

6. The method according to claim 5, wherein the modifying of the input image comprises: mapping, by a latent encoding unit, the high-frequency information into random variables that are distributed according to a standardized Gaussian distribution, wherein the mapping enforces on the high-frequency information a Gaussian distribution with a mean and a variance that are based on the modified image.

7. The method according to claim 5, wherein the high-frequency information is disentangled into: a first set including image detail information of the input image, and a second set including film grain information of the input image; and wherein the reconstructing of the input image comprises: using the first and the second sets to reconstruct the input image, or using the first set to reconstruct a grain-free version of the input image.

8. The method according to claim 5, wherein the high-frequency information is disentangled into: a first set including image detail information, of the input image, that is lost due to the reducing of the resolution of the input image, a second set including image detail information, of the input image, that is lost due to the converting of the input image into an energy-aware image, and a third set including film grain information of the input image; and wherein the reconstructing of the input image comprises: using the first, the second, and the third sets to reconstruct the input image, using the first and the second sets to reconstruct a grain-free version of the input image, using the first and the third sets to reconstruct an energy-aware version of the input image, or using the first set to reconstruct a grain-free and energy-aware version of the input image.

9. The method according to any one of the preceding claims, wherein the one ormore loss functions comprise: a loss function that characterizes a difference between a down-sampled and grain-free version of the input image and the modified image.

10. The method according to any one of the preceding claims, wherein the one or more loss functions comprise: a loss function that enforces a standardized Gaussian distribution on random variables outputted in a forward pass by the invertible framework.

11. The method according to any one of the preceding claims, wherein the one or more loss functions comprise: a loss function that characterizes a difference between the input image and the reconstructed image.

12. The method according to any one of the preceding claims, wherein the one or more loss functions comprise: a loss function that characterizes a difference between a grain-free version of the input image and a reconstructed grain-free version of the input image.

13. The method according to any one of the preceding claims, wherein the one or more loss functions comprise: a structural similarity index measure loss function that characterizes a difference between a down-sampled and grain-free version of the input image and the modified image.

14. The method according to any one of the preceding claims, wherein the one or more loss functions comprise: a loss function that characterizes a difference between a power measure of a down- sampled and grain-free version of the input image and a power measure of the modified image, constraining the power measure of the modified image by a power reduction rate.

15. A method, comprising: obtaining an input image, the input image is a modification of an original image that was modified using an invertible framework run in a forward pass, the modification includesmodification operations of: reducing the resolution of the original image, removing film grain from the original image, and converting the original image into an energy-aware image; and reconstructing the original image using the invertible framework run in a backward pass to map the input image into a reconstructed original image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

16. A method, comprising: training an invertible framework to modify an input image into a modified image when run in a forward pass and to reconstruct the input image from the modified image when run in a backward pass, the modification includes modification operations of: reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

17. An apparatus, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to to obtain an input image, and to modify the input image into a modified image using an invertible framework run in a forward pass, the modifying includes modification operations of: reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

18. The apparatus according to claim 17, wherein the instructions further cause the apparatus to: obtain a version of the modified image; and reconstruct the input image using the invertible framework run in a backward pass to map the version of the modified image into a reconstructed input image.

19. The apparatus according to claim 18, wherein the modified image is a frame from a video to be processed in a video chain, wherein the video is encoded at an encoder end and decoded at a decoder end to generate a decoded video; and the obtained version of the modified image is a corresponding frame from the decoded video.

20. An apparatus, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to: obtain an input image, the input image is a modification of an original image that was modified using an invertible framework run in a forward pass, the modification includes modification operations of: reducing the resolution of the original image, removing film grain from the original image, and converting the original image into an energy-aware image, and reconstruct the original image using the invertible framework run in a backward pass to map the input image into a reconstructed original image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

21. An apparatus, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to:train an invertible framework to modify an input image into a modified image when run in a forward pass and to reconstruct the input image from the modified image when run in a backward pass, the modification includes modification operations of: reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

22. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method, the method comprising: obtaining an input image; and modifying the input image into a modified image using an invertible framework run in a forward pass, the modifying includes modification operations of: reducing the resolution of the input image, removing film grain from the input image, and converting the input image into an energy-aware image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

23. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method, the method comprising: obtaining an input image, the input image is a modification of an original image that was modified using an invertible framework run in a forward pass, the modification includes modification operations of: reducing the resolution of the original image, removing film grain from the original image, and converting the original image into an energy-aware image; and reconstructing the original image using the invertible framework run in a backward pass to map the input image into a reconstructed original image, wherein the invertible framework is trained, according to one or more loss functions, to simultaneously optimize the modification operations and the inversion of these operations.

Citation Information

Patent Citations

  • EP22305671A

  • EP22306040A

  • EP22306719A

  • EP23305294A

  • EP21305604A

Cited By

  • Data encoding method, data decoding method, and data processing apparatus

    US12483273B2

  • Data encoding method, data decoding method, and data processing apparatus

    US20240235577A1