Apparatus and method for image encoding and decoding

Through the neural network-based image encoding and decoding method, the problem of efficient video data encoding and decoding in the prior art is solved, and efficient image data compression and decoding is realized, reducing bandwidth requirements and power consumption.

CN119946266APending Publication Date: 2025-05-06SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411151521.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2024-08-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the encoding and decoding problems of efficient video data, especially when network bandwidth is limited.

Method used

Using neural network-based image encoding and decoding methods, image blocks are transformed into potential representations through reversible and irreversible neural networks, and entropy encoding and decoding are performed based on these representations.

Benefits of technology

Efficient image data compression and decoding are realized, reducing the bandwidth requirement between IP and DRAM in the system on chip, reducing power consumption, and improving the quality of image data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946266A_ABST
    Figure CN119946266A_ABST
Patent Text Reader

Abstract

An apparatus and method for image encoding and decoding are provided. The image coding method comprises the following steps: transforming an image block into a first potential representation based on a reversible neural network; transforming the first potential representation into a second potential representation based on an irreversible neural network; estimating a probability distribution of the first potential representation; and performing entropy encoding on the first potential representation based on the probability distribution using an entropy encoder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is based on and claims the benefit of priority from Korean Patent Application No. 10-2023-0152118 filed in the Korean Intellectual Property Office on November 6, 2023, and Korean Patent Application No. 10-2024-0000749 filed in the Korean Intellectual Property Office on January 3, 2024, the disclosures of which are incorporated herein in their entirety by reference for all purposes. Technical Field

[0002] The present disclosure relates to encoding and decoding devices and encoding and decoding methods, and in particular to a device and method for image encoding and decoding based on a neural network. Background Art

[0003] In recent years, the Internet video market has continued to grow. However, since the amount of data contained in video is much larger than that of media such as voice, text, and photos, and the type or quality of service may be limited by network bandwidth, advanced video encoding technology is required. A related technology for image data compression is frame buffer compression (FBC), which can effectively utilize dynamic random access memory (DRAM) bandwidth during image data transmission between intellectual property (IP) in system-on-chip (SoC). Summary of the invention

[0004] According to one aspect of the present disclosure, there is provided an image encoding method, comprising: transforming an image block into a first latent representation based on a reversible neural network; transforming the first latent representation into a second latent representation based on an irreversible neural network; estimating a first probability distribution of the first latent representation based on the second latent representation; and performing entropy encoding on the first latent representation based on the first probability distribution.

[0005] The reversible neural network may include a regularized flow neural network comprising one or more coupled layers.

[0006] Parameters of one or more coupling layers may be obtained by training based on at least one of a neural network structure or a data distribution.

[0007] Estimating the first probability distribution may include dividing the first potential representation into a plurality of groups and estimating the first probability distribution for each of the plurality of groups.

[0008] The method may also include performing entropy encoding on the second latent representation based on the second probability distribution.

[0009] Transforming the first latent representation into a second latent representation may include transforming the first latent representation into the second latent representation using a super-a priori encoder.

[0010] Estimating the first probability distribution may include obtaining the super-prior as a probability distribution of the first latent representation using a super-prior decoder.

[0011] Estimating the first probability distribution may include: obtaining a super-prior based on an entropy decoding result of the second latent representation using a super-prior decoder; and estimating a first probability distribution of the first latent representation based on the super-prior using a context estimator.

[0012] The method may further include dividing the image block into a plurality of sub-blocks.

[0013] The method may also include performing operations within each of the plurality of sub-blocks based on one or more coupled layers of a regularized flow neural network.

[0014] The method may also include performing operations between the plurality of sub-blocks based on one or more coupling layers of the regularized flow neural network.

[0015] The operation between the plurality of sub-blocks may be performed by at least one of a super a priori encoder, a super a priori decoder, or a context estimator.

[0016] Transforming the image patch into the first latent representation may include transforming the image patch into the first latent representation by hierarchically using two or more reversible neural network based modules.

[0017] Transforming the image block into a first latent representation may include: inputting a portion of a third latent representation transformed by a reversible neural network-based module of a previous layer into a reversible neural network-based module of a next layer to transform a portion of the third latent representation into a fourth latent representation, and combining the remaining portion of the third latent representation with the fourth latent representation to transform the combined latent representation into the first latent representation.

[0018] According to another aspect of the present disclosure, there is provided an image decoding method, comprising: receiving a first bit stream of a first potential representation obtained by transforming an image block based on a reversible neural network; receiving a second bit stream of a second potential representation obtained by transforming the first potential representation based on an irreversible neural network; estimating a first probability distribution of the first potential representation based on the second bit stream; and reconstructing the image block based on the first bit stream and the first probability distribution.

[0019] The method may further include entropy decoding the first bitstream based on a first probability distribution of the first latent representation, wherein reconstructing the image block includes reconstructing the image block based on an entropy decoding result of the first bitstream and the first probability distribution.

[0020] The method may further include entropy decoding the second bitstream based on the second probability distribution.

[0021] Estimating the first probability distribution of the first latent representation may include obtaining a super-prior based on a decoding result of the second bitstream, and estimating the first probability distribution of the first latent representation based on the super-prior.

[0022] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing one or more instructions, and a processor configured to execute the one or more instructions to implement: a reversible neural network configured to transform an image block into a first latent representation; an irreversible neural network configured to transform the first latent representation into a second latent representation, and estimate a first probability distribution of the first latent representation based on the second latent representation; and an entropy encoder configured to perform entropy encoding on the first latent representation based on the first probability distribution.

[0023] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing one or more instructions, and a processor configured to execute the one or more instructions to implement: an irreversible neural network configured to: receive a first bit stream of a first latent representation obtained by transforming an image block, and a second bit stream of a second latent representation obtained by transforming the first latent representation, and estimate a probability distribution of the first latent representation; and a reversible neural network configured to reconstruct an image based on the first bit stream and the estimated probability distribution. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Embodiments of the present disclosure are shown in the accompanying drawings, and similar reference numerals in all drawings indicate corresponding parts in the various figures. The embodiments herein will be better understood from the following description with reference to the accompanying drawings, in which:

[0025] Figure 1 The present invention is a block diagram showing an image encoding device according to an embodiment of the present disclosure.

[0026] FIG. 2A to FIG. 2C The present invention is a block diagram showing an image encoding device according to another embodiment of the present disclosure.

[0027] Figure 3A and Figure 3B An example of a coupled layer in a regularized flow.

[0028] Figure 4 A diagram showing an example of applying a reversible neural network-based module in multiple layers.

[0029] Figure 5A and Figure 5B A block diagram showing a module based on an irreversible neural network according to an embodiment of the present disclosure.

[0030] Figure 6 The figure is a block diagram showing an image decoding apparatus according to an embodiment of the present disclosure.

[0031] Figure 7 The present invention is a block diagram showing an image decoding device according to another embodiment of the present disclosure.

[0032] Figure 8The figure is a flowchart showing an image encoding method according to an embodiment of the present disclosure.

[0033] Fig. 9 The figure is a flowchart showing an image decoding method according to an embodiment of the present disclosure.

[0034] Fig.10 FIG. 1 is a block diagram showing an electronic device according to an embodiment of the present disclosure.

[0035] Fig.11 is a block diagram illustrating an example of an image processing apparatus included in an electronic device. DETAILED DESCRIPTION

[0036] The embodiments herein and their various features and advantageous details are explained more fully with reference to the non-limiting embodiments shown in the accompanying drawings and described in detail in the following description. The description of known components and processing techniques is omitted so as not to unnecessarily confuse the embodiments herein. The examples used herein are only intended to facilitate understanding of the manner in which the embodiments herein can be implemented and to further enable those skilled in the art to implement the embodiments herein. Therefore, the examples should not be construed as limiting the scope of the embodiments herein. Throughout the drawings and the specific embodiments, unless otherwise described, the same reference numerals will be understood to refer to the same elements, features and structures.

[0037] The following detailed description is provided to help the reader obtain a comprehensive understanding of the methods, devices and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be apparent. For example, the order of operations described herein is merely an example and is not limited to those order of operations set forth herein, but is obviously changeable after understanding the disclosure of the present application, except for operations that must be performed in a certain order. In addition, for greater clarity and brevity, the description of known features may be omitted after understanding the disclosure of the present application.

[0038] The features described herein can be implemented in different forms and should not be interpreted as being limited to the examples described herein. Rather, the examples described herein are provided only to illustrate some of the many possible ways to implement the methods, devices and / or systems described herein, which will be apparent after understanding the disclosure of the present application. For example, one or more elements or components of the device described herein can be combined or separated without departing from the scope disclosed in the present application. As conventional in the art, as shown in the accompanying drawings, embodiments can be described and illustrated from the aspects of the blocks that perform the described functions. These blocks may be referred to herein as units or modules, etc., or names such as devices, logic, circuits, encoders, decoders, counters, comparators, generators, converters, etc., which can be physically implemented by analog and / or digital circuits including one or more logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, etc., and can also be implemented or driven by software and / or firmware (configured to perform the functions or operations described herein).

[0039] Throughout this disclosure, when a component is described as being "connected to" or "coupled to" another component, it may be directly "connected to" or "coupled to" the other component, or one or more other components may be interposed. Conversely, when an element is described as being "directly connected to" or "directly coupled to" another element, there are no other elements interposed. Similarly, similar expressions such as "between" and "immediately between" and "adjacent" and "directly adjacent" will also be interpreted in the same manner. As used herein, the term "and / or" includes any one of the items listed in the association and any combination of any two or more.

[0040] Although terms such as "first", "second", "third" and the like may be used herein to describe various components, assemblies, regions, layers or parts, these components, assemblies, regions, layers or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, component, region, layer or part from another component, component, region, layer or part. Therefore, without departing from the teachings of the examples described herein, the first component, component, region, layer or part mentioned in the examples may also be referred to as the second component, component, region, layer or part.

[0041] The terms used herein are only used to describe various examples and are not used to limit the present disclosure. Unless the context clearly indicates otherwise, the articles "a", "an" and "the" are also intended to include plural forms. The terms "comprise", "contain", "include", "include" and "have" indicate the presence of the described features, numbers, operations, components, elements and / or combinations thereof, but do not exclude the presence or addition of one or more other features, numbers, operations, components, elements and / or combinations thereof.

[0042] Figure 1 1 is a block diagram showing an image encoding apparatus 100 according to an embodiment of the present disclosure.

[0043] According to an embodiment, the image encoding device 100 may be included in various electronic devices, including but not limited to various image sending, receiving or processing devices, such as televisions, monitors, Internet of Things (IoT) devices, radar devices, smart phones, wearable devices, tablet PCs, netbooks, laptop computers, desktop computers, head-mounted displays (HMDs), autonomous driving vehicles, virtual reality (VR) devices, augmented reality (AR) devices, extended reality (XR) devices, vehicles, mobile robots, etc., as well as cloud computing devices, etc.

[0044] The image encoding device 100 can be used in an environment where operations are performed in units of blocks to randomly access data. According to an embodiment, the image encoding device 100 can perform lossless compression with a high compression ratio by using both lossless image compression and lossy image compression. For example, lossless compression with a high compression ratio can be applied to frame buffer compression (FBC). In this way, power consumption can be reduced by reducing the bandwidth between IP and DRAM in a system on chip (SoC), and the image encoding device 100 can be used for image data transmission between various devices and / or servers.

[0045] refer to Figure 1, the image encoding device 100 may include a first module 110, a second module 120, a first entropy encoder 131 and a second entropy encoder 141. According to an embodiment, the image encoding device 100 may include a memory storing one or more instructions, and a processor configured to execute one or more instructions to perform one or more operations of the image encoding device 100. For example, the memory may store a program code or an instruction set corresponding to the first module 110, the second module 120, the first entropy encoder 131 and the second entropy encoder 141. According to an embodiment, the processor may execute one or more instructions stored in the memory to implement the first module 110, the second module 120, the first entropy encoder 131 and the second entropy encoder 141. According to an embodiment, the processor may include one or more processors. According to an embodiment, the processor may include, but is not limited to, one or more logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical elements, etc.

[0046] The first module 110 may transform the image block IB used as input into a first potential representation. The first module 110 may include a reversible neural network that performs lossless compression, and may transform the image block IB into a first potential representation by using the reversible neural network. The reversible neural network may be a regularized flow neural network. The regularized flow neural network may include one or more coupling layers, and the one or more coupling layers may be modified in various ways. According to an embodiment, one or more parameters of one or more coupling layers may be obtained through training. For example, one or more parameters may include but are not limited to a division ratio, an element to be selected, and the like. One or more parameters of one or more coupling layers may be obtained based on a neural network structure, data distribution, and the like. According to an embodiment, one or more parameters of one or more coupling layers may be updated by retraining using image encoding and decoding results as training data.

[0047] Here, the term “latent representation” refers to the output of a neural network using an input image or motion information as input, and may be collectively referred to as latent features, latent vectors, etc.

[0048] The second module 120 can transform the first potential representation output by the first module 110 into a second potential representation, and can estimate the probability distribution of the first potential representation based on the transformed second potential representation. The probability distribution may include a mean μ and a standard deviation σ. The second module 120 can estimate the probability distribution by applying various probability models. The probability model may include, but is not limited to, a Laplace distribution model or a Gaussian distribution model.

[0049] The second module 120 may include an irreversible neural network that performs lossy compression. According to an embodiment, the irreversible neural network may include a hyper-a priori encoder, a hyper-a priori decoder, and / or a context estimator. However, the irreversible neural network is not limited thereto, and therefore, according to an embodiment, the irreversible neural network may include various other neural networks, such as a convolutional neural network (CNN), a recurrent neural network (RNN), a transformer-based neural network, etc., which may be used in combination as appropriate.

[0050] The first entropy encoder 131 can perform entropy encoding on the first potential representation based on the input probability distribution by using the first potential representation transformed by the first module 110 and the probability distribution estimated by the second module 120 as input. In addition, the first entropy encoder 131 can output a first bit stream of the first potential representation as a result of image encoding. Entropy encoding can be performed using arithmetic encoding techniques and arithmetic decoding techniques, but is not limited thereto.

[0051] The second entropy encoder 141 can perform entropy coding on the second potential representation based on the input probability distribution by using the second potential representation transformed by the second module 120 and the probability distribution obtained by training as input. For example, the second entropy encoder 141 can perform entropy coding on the second potential representation based on the input probability distribution by using the second potential representation transformed by the second module 120 and the reference probability distribution obtained by training as input. The reference probability distribution can be a predetermined probability distribution. In addition, the second entropy encoder 141 can output a second bit stream of the second potential representation as a result of image encoding. Entropy coding can be performed using arithmetic coding techniques and arithmetic decoding techniques, but is not limited thereto.

[0052] The first bit stream entropy encoded by the first entropy encoder 131 and the second bit stream entropy encoded by the second entropy encoder 141 can be transmitted to another IP in the SoC through DRAM, etc. However, the transmission of the bit stream is not limited thereto, and therefore, according to another embodiment, the bit stream can be transmitted to an external electronic device through wired and wireless communications.

[0053] FIG. 2A to FIG. 2C The present invention is a block diagram showing an image encoding device according to another embodiment of the present disclosure.

[0054] refer to Figure 2A , the image encoding device 200a may include a first module 110, a second module 120, a first entropy encoder 131, a second entropy encoder 141, and a second entropy decoder 142. A detailed description of the first module 110, the second module 120, the first entropy encoder 131, and the second entropy decoder 142 will be omitted.

[0055] The second entropy decoder 142 can perform entropy decoding using the second bit stream of the second potential representation generated by the second entropy encoder 141 and the probability distribution predefined by training as input to reconstruct the second potential representation. The reconstructed second potential representation can be input to the second module 120, and the second module 120 can estimate the probability distribution of the first potential representation based on the input second potential representation. Entropy decoding can be performed by arithmetic coding technology and arithmetic decoding technology, but is not limited thereto.

[0056] refer to Figure 2B , the image encoding device 200b may include a first module 110, a second module 120, a first entropy encoder 131, a first entropy decoder 132, a second entropy encoder 141, and a second entropy decoder 142. According to an embodiment, the second entropy decoder 142 may be omitted. A detailed description of the first module 110, the second module 120, the first entropy encoder 131, the second entropy encoder 141, and the second entropy decoder 142 will be omitted.

[0057] The second module 120 may divide the image block IB of the first potential representation into a plurality of sub-blocks, and may estimate the probability distribution of the first potential representation for each sub-block. According to an embodiment, the result of the entropy decoding performed by the first entropy decoder 132 for the previous sub-block may be input to the second module 120, and the second module 120 may estimate the probability distribution of the first potential representation of the current sub-block based on the result of the entropy decoding of the previous sub-block (the reconstructed first potential representation for the previous sub-block). For example, when estimating the probability distribution for the first sub-block, the second module 120 may estimate the probability distribution by using only the result of the entropy decoding performed by the second entropy decoder 142, and when estimating the probability distribution for the nth sub-block (n is an integer greater than 2), the second module 120 may estimate the probability distribution by using the result of the entropy decoding performed by the second entropy decoder 142 and the result of the entropy decoding performed by the first entropy decoder 132 for the first to n-1th sub-blocks.

[0058] The first entropy encoder 131 may perform entropy encoding on the first latent representation by using the first latent representation transformed by the first module 110 and the probability distribution estimated by the second module 120 as input. According to an embodiment, the first entropy encoder 131 may perform entropy encoding on the first latent representation for each sub-block based on the probability distribution of the first latent representation estimated by the second module 120 for each sub-block.

[0059] The first entropy decoder 132 can use the first bit stream of the first potential representation and the probability distribution estimated by the second module 120 as input to reconstruct the first potential representation by performing entropy decoding based on the input probability distribution. According to an embodiment, the first potential representation reconstructed for the sub-block can be input to the second module 120 for estimating the probability distribution of the first potential representation for subsequent sub-blocks. Entropy decoding can be performed by arithmetic coding technology and arithmetic decoding technology, but is not limited thereto.

[0060] refer to Figure 2C , the image encoding device 200c may include a first module 110, a second module 120, a first entropy encoder 131, a second entropy encoder 141, and a block partitioning module 150. According to an embodiment, the image encoding device 200c may further include a first entropy decoder 132 and / or a second entropy decoder 142. A detailed description of the first module 110, the second module 120, the first entropy encoder 131, the first entropy decoder 132, the second entropy encoder 141, and the second entropy decoder 142 will be omitted.

[0061] According to an embodiment, the block partitioning module 150 may partition the input image block into sub-blocks. For example, the block partitioning module 150 may partition an input image block of a predetermined size (e.g., 4x 32) into sub-blocks of a predetermined unit size (e.g., 4x 4) in consideration of the locality of the image. According to an embodiment, the unit size may be predefined in consideration of computing power, target decoding accuracy, etc.

[0062] The first module 110 can transform the image block into a first potential representation in units of sub-blocks using a reversible neural network. The reversible neural network may include one or more regularized stream neural networks. The regularized stream neural network may include one or more first coupling layers, which are configured to perform operations within the sub-block. In addition, the regularized stream neural network may also include one or more second coupling layers, which are configured to perform operations between sub-blocks between the first coupling layers or in the last layer. The first coupling layer and the second coupling layer may be included in a separate regularized stream neural network.

[0063] The second module 120 can use an irreversible neural network to transform the first potential representation into a second potential representation in units of sub-blocks, and can estimate the probability distribution of the first potential representation. The irreversible neural network may include a super-prior encoder, a super-prior decoder and / or a context estimator. The irreversible neural network can be configured to perform operations within a sub-block and / or between sub-blocks. The second module 120 can estimate the probability distribution of the first potential representation for the current sub-block based on the entropy decoding result of the previous sub-block (the first potential representation reconstructed for the previous sub-block).

[0064] The first entropy encoder 131 may perform entropy encoding on the first latent representation using the probability distribution estimated by the second module 120 as input, and may output a bitstream of the first latent representation. The first entropy decoder 132 may reconstruct the first latent representation by performing entropy decoding on the bitstream of the first latent representation using the probability distribution estimated by the second module 120 as input.

[0065] The second entropy encoder 141 may output a bitstream of the second potential representation by performing entropy encoding on the second potential representation using a probability distribution predefined by training. The second entropy decoder 142 may reconstruct the second potential representation by performing entropy decoding on the bitstream of the second potential representation using a probability distribution predefined by training.

[0066] Figure 3A and Figure 3B An example of a coupling layer in a regularized flow.

[0067] refer to Figure 3A , the coupling layer of the regularized flow neural network can be configured to divide the input image block x into the first image block x proportionally 1 and the second image block x 2 , for the first image block x 1 Processing is performed to combine the processed result t with the second image block x 2 Combine and convert the combined result x 2 +t and the first image block x 1 The first potential representation x' of the image block x is outputted by combining the first potential representation x' of the image block x. According to an embodiment, the ratio may be preset. According to an embodiment, the first image block x may be processed as follows 1 :The first image block x 1 Input to the neural network NN for inference by the neural network NN, then quantize the inference result Q, and add the result t to the second image block x 2 The regularized flow neural network may include multiple coupling layers to repeatedly perform the coupling process.

[0068] refer to Figure 3B , the image blocks are divided proportionally in the coupling layer of the regularized flow neural network ( Figure 3A ) can be modified to mask the parameters by inputting a trained predefined mask M. For example, the method may include masking the parameters of the coupling layer by inputting a mask M instead of dividing the input image block x into a first image block x 1 and the second image block x 2. According to an embodiment, the ratio may be preset. According to an embodiment, the mask M may contain information related to the parameters, including the ratio for dividing the image block, the elements to be selected or transferred in the image block, etc. The mask M may be determined by training in consideration of the structure of the neural network NN, the data distribution, etc. The mask M may be determined by training using learnable parameters and / or quantization of a straight-through estimator (STE). The mask M may be updated at predetermined intervals or when necessary by using the encoding and / or decoding results of the image block as training data for training. In this way, the influence of random seeds may be reduced, manually designed features may be reduced, and performance may be improved by optimization.

[0069] Figure 4 A diagram showing an example of applying a reversible neural network-based module in multiple layers.

[0070] refer to Figure 4 , the image encoding apparatus 100, 200a, 200b, and 200c may transform the image block x into the first potential representation z by hierarchically using the plurality of first modules 110. For example, the plurality of first modules 110 may include a first first module 111 and a second first module 112. The first first module 111 and the second first module 112 may include a regularized flow neural network 411. The image block x is input to the first first module 111 of the first layer to be transformed into the potential representation by the regularized flow neural network 411, and the transformed potential representation is divided into two parts z 1 and 1 The partitioned part y of the latent representation 1 Input to the second first module 112 of the second layer to be transformed into the second potential representation z through the regularized flow neural network 411 2 The transformed latent representation z 2 The part of the latent representation z transformed in the first layer 1 to generate the final latent representation. Figure 4 Two layers are shown, but the layers are not limited thereto, and the method performed in the first layer may be repeated in subsequent layers, extending to three or more layers.

[0071] Figure 5A and Figure 5B A block diagram showing a module based on an irreversible neural network according to an embodiment of the present disclosure.

[0072] refer to Figure 5A , the second module 510a according to an embodiment of the present disclosure may include a super a priori encoder 511 and a super a priori decoder 512. The second module 510a may be based on an irreversible neural network.

[0073] The super-a priori encoder 511 can use the potential representation ly output by the first module 110 as input and output the super-potential representation hlz. The entropy encoder and / or decoder 140 can use the probability distribution lpd as input to perform entropy encoding and / or decoding on the super-potential representation hlz. The probability distribution lpd can be a pre-trained probability distribution. The bitstream of the super-potential representation hlz is output by performing entropy encoding. According to an embodiment, the probability distribution is obtained by Gaussian modeling based on the trained parameters, but is not limited to this. The super-potential representation hlz reconstructed as a result of entropy decoding is input to the super-a priori decoder 512. The super-a priori decoder 512 can output the super-prior hp by decoding the super-potential representation hlz. According to an embodiment, the super-prior hp represents a feature vector used to express the potential representation ly as a probability distribution. The super-prior hp output by the super-a priori decoder 512 can be input to the first entropy encoder 131 and / or the first entropy decoder 132 as a probability distribution.

[0074] refer to Figure 5B According to another embodiment of the present disclosure, the second module 510b may include a super a priori encoder 511, a super a priori decoder 512, and a context estimator 513. The second module 510b may be based on an irreversible neural network.

[0075] The super-a priori encoder 511 can use the potential representation cly output by the first module 110 as input and output the super-potential representation hlz. The entropy encoder and / or decoder 140 can use the pre-trained probability distribution lpd as input to perform entropy encoding and / or decoding on the super-potential representation hlz. The bitstream of the super-potential representation hlz is output by performing entropy encoding. According to an embodiment, the probability distribution is obtained by Gaussian modeling based on the trained parameters, but is not limited to this. The super-potential representation hlz reconstructed as a result of entropy decoding is input to the super-prior decoder 512. The super-prior decoder 512 can output the super-prior hp by decoding the super-potential representation hlz. According to an embodiment, the super-prior hp represents a feature vector used to express the potential representation ly as a probability distribution. The super-prior hp output by the super-prior decoder 512 is input to the context estimator 513, and the context estimator 513 can use the super-prior hp to estimate the probability distribution of the potential representation cly. The context estimator 513 may divide the potential representation cly into a plurality of sub-blocks, and when estimating the probability distribution of the current potential representation cly, the context estimator 513 may estimate the probability distribution epd by using the potential representation ply entropy encoded and decoded for the previous sub-block. The estimated probability distribution epd may be input to the first entropy encoder 131 and / or the first entropy decoder 132 as a probability distribution.

[0076] Figure 6 The figure is a block diagram showing an image decoding apparatus according to an embodiment of the present disclosure.

[0077] refer to Figure 6 , the image decoding device 600 may include a first entropy decoder 610, a second module 620, a second entropy decoder 630, and a first module 640. According to an embodiment, the image decoding device 600 may include a memory storing one or more instructions, and a processor configured to execute one or more instructions to perform one or more operations of the image decoding device 600. For example, the memory may store program codes or instruction sets corresponding to the first entropy decoder 610, the second module 620, the second entropy decoder 630, and the first module 640. According to an embodiment, the processor may execute one or more instructions stored in the memory to implement the first entropy decoder 610, the second module 620, the second entropy decoder 630, and the first module 640. According to an embodiment, the processor may include one or more processors. According to an embodiment, the processor may include, but is not limited to, one or more logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, etc.

[0078] The first entropy decoder 610 can be decoded by using Figure 1 , Figure 2A , Figure 2B and Figure 2C The first bitstream BS1 generated by the image encoding apparatus 100, 200a, 200b and 200c is used as an input to perform entropy decoding to reconstruct the first latent representation. The first entropy decoder 610 can reconstruct the first latent representation based on the probability distribution estimated by the second module 620, and the second module 620 is based on an irreversible neural network. The second module 620 based on the irreversible neural network can estimate the probability distribution based on the second latent representation reconstructed by the second entropy decoder 630. The second entropy decoder 630 can be used to reconstruct the first latent representation by using Figure 1 , Figure 2A , Figure 2B and Figure 2C The second bit stream BS2 generated by the image encoding device 100, 200a, 200b and 200c is used as an input to perform entropy decoding to reconstruct the second potential representation. By using a reversible neural network, the first module 640 can reconstruct the image block using the first potential representation reconstructed by the first entropy decoder 610 as an input. The detailed description of the first entropy decoder 610, the second module 620, the second entropy decoder 630 and the first module 640 will be omitted, and they basically perform the same functions as the first entropy decoder 131, the second module 120, the second entropy decoder 142 and the first module 110 of the aforementioned image encoding device 100, 200a, 200b and 200c.

[0079] Figure 7 The present invention is a block diagram showing an image decoding device according to another embodiment of the present disclosure.

[0080] refer to Figure 7 , the image decoding device 700 may include a first entropy decoder 610, a second module 620, a second entropy decoder 630, a first module 640, and a block merging module 710. The first entropy decoder 610, the second module 620, the second entropy decoder 630, and the first module 640 are as described above, so their detailed description will be omitted. Figure 2C ) In the case where the image block is divided into a plurality of sub-blocks and image decoding is performed on each sub-block, the block merging module 710 can reconstruct the image block by merging the sub-blocks reconstructed by the first module 640.

[0081] Figure 8 The figure is a flowchart showing an image encoding method according to an embodiment of the present disclosure.

[0082] Figure 8 The method is an example of the image encoding method performed by the above-mentioned image encoding devices 100, 200a, 200b and 200c, and will be briefly described below to avoid redundancy.

[0083] According to an embodiment, in operation 810, the method may include transforming the image block into a first potential representation. For example, the image encoding device may transform the image block into the first potential representation by using a module based on a reversible neural network. For example, before being input to the module based on the reversible neural network, the image block is divided into a plurality of sub-blocks, and the module based on the reversible neural network may perform operations in units of sub-blocks to transform the image block into the first potential representation. The reversible neural network may include a regularized flow neural network. According to an embodiment, in the example case where the image block is divided into a plurality of sub-blocks, the coupling layer of the regularized flow neural network may be configured to perform operations within and / or between sub-blocks. In addition, for example, the parameters of the coupling layer (e.g., the division ratio, the elements to be selected, etc.) may be predefined by training in consideration of the neural network structure, data distribution, etc., and may be updated by retraining using the image encoding and decoding results as training data. For example, the image encoding device may transform the image block into the first potential representation by hierarchically using a plurality of modules based on a reversible neural network.

[0084] In operation 820, the method may include transforming the first potential representation into a second potential representation. For example, the image encoding device may transform the first potential representation into the second potential representation by using a module based on an irreversible neural network. The irreversible neural network may include a super prior encoder, a super prior decoder, a context estimator, a convolutional neural network (CNN), a recurrent neural network (RNN), a transformer-based neural network, etc., which may be used in appropriate combination.

[0085] In operation 830, the method may include performing entropy encoding on the second potential representation. For example, the image encoding device may perform entropy encoding on the second potential representation (e.g., transformed in operation 820) by using an entropy encoder. The image encoding device may output a bit stream of the second potential representation as a result of entropy encoding. Entropy encoding may be performed using arithmetic encoding techniques and arithmetic decoding techniques, but is not limited thereto.

[0086] In operation 840, the method may include estimating a probability distribution of a first potential representation. For example, the image encoding device may estimate the probability distribution of the first potential representation by using a module based on an irreversible neural network. For example, entropy decoding may be performed on the bitstream generated in operation 830, and the super prior output by the super prior decoder may be used as a probability distribution, wherein the super prior decoder uses the second potential representation reconstructed by entropy decoding as input. In another example, the super prior output by the super prior decoder is input to a context estimator, and the context estimator may estimate the probability distribution of the first potential representation by using the super prior. According to an embodiment, the probability distribution may be estimated by applying various probability models including a Gaussian-based model. In the case where the context estimator divides the first potential representation into a plurality of sub-blocks and estimates the probability distribution for each sub-block, the probability distribution of the first potential representation may be estimated for subsequent sub-blocks by using the entropy decoding result of the bitstream generated in operation 830 for the previous sub-block.

[0087] In operation 850, the method may include performing entropy encoding on the first potential representation. For example, by using an entropy encoder, the image encoding device may perform entropy encoding on the first potential representation based on the probability distribution estimated in operation 840. A bitstream of the first potential representation may be output as a result of entropy encoding. According to an embodiment, a bitstream of a subblock generated as a result of entropy encoding may be entropy decoded for estimating a probability distribution for a subsequent subblock.

[0088] Fig. 9 The figure is a flowchart showing an image decoding method according to an embodiment of the present disclosure.

[0089] Fig. 9 The method is Figure 6 or Figure 7 An example of an image decoding method performed by the image decoding device 600 or 700 will be briefly described below to avoid redundancy.

[0090] According to an embodiment, the method may include receiving a first bitstream and a second bitstream in operation 910. For example, the image decoding apparatus receives the first bitstream and the second bitstream.

[0091] In operation 920, the method may include performing entropy decoding on the second bitstream. For example, the image decoding device may perform entropy decoding on the second bitstream using an entropy decoder in operation 920. The entropy decoder may reconstruct the second potential representation by entropy decoding the second bitstream based on a probability distribution predetermined by training.

[0092] In operation 930, the method may include estimating a probability distribution of the first potential representation based on the second potential representation. For example, the image decoding device may estimate the probability distribution of the first potential representation by using the second potential representation reconstructed in operation 920 as an input and using a module based on an irreversible neural network. The module based on the irreversible neural network may include a super-prior decoder and / or a context estimator. The second potential representation reconstructed in operation 920 may be input to a super-prior decoder to output a super-prior, and the super-prior may be used as the probability distribution of the first potential representation, or the super-prior may be input to a context estimator to estimate the probability distribution.

[0093] In operation 940, the method may include performing entropy decoding on the first bitstream based on the probability distribution. For example, by using an entropy decoder, the image decoding apparatus may perform entropy decoding on the first bitstream using the probability distribution estimated in operation 930 as an input.

[0094] In operation 950, the method may include reconstructing the reconstructed first latent representation as an image block. For example, by using a module based on a reversible neural network, the image decoding device may reconstruct the reconstructed first latent representation as an image block. In the image encoding process, in the example case where the image block is divided into sub-blocks and image encoding is performed in units of sub-blocks, the reconstructed first latent representation in units of sub-blocks may be merged.

[0095] Fig.10 FIG. 1 is a block diagram showing an electronic device according to an embodiment of the present disclosure.

[0096] Electronic devices may include, for example, various image sending / receiving devices, such as televisions, monitors, Internet of Things (IoT) devices, radar devices, smart phones, wearable devices, tablet PCs, netbooks, laptops, desktop computers, head-mounted displays (HMDs), autonomous vehicles, virtual reality (VR) devices, augmented reality (AR) devices, extended reality (XR) devices, vehicles, mobile robots, etc., as well as cloud computing devices, etc.

[0097] refer to Fig.10 , the electronic device 1000 may include an image capturing device 1010 , an image processing device 1020 , a processor 1030 , a storage device 1040 , an output device 1050 , and a communication device 1060 .

[0098] The image capture device 1010 may include a device for capturing a still image or a moving image, such as a camera, etc., and may store the captured image in the storage device 1040 and transmit the image to the processor 1030. The image capture device 1010 may include a lens assembly having one or more lenses, an image sensor, an image signal processor, and / or a flash. The lens assembly included in the camera module may collect light emitted from an object to be imaged.

[0099] The image processing device 1020 may include the above-mentioned image encoding device and / or image decoding device. The image processing device 1020 can efficiently encode and / or decode the image based on the frame buffer compression (FBC) technology as described above, thereby reducing the DRAM bandwidth or power consumption required for data communication between the IP and DRAM in the SoC of the electronic device or data communication between electronic devices. In addition, by performing efficient image encoding, power consumption can be further reduced, and the battery time or thermal constraints of the electronic device can be improved.

[0100] The processor 1030 may include a main processor (e.g., one or more central processing units (CPUs) or application processors (APs), etc.), an intellectual property (IP) core, and an auxiliary processor (e.g., a graphics processing unit (GPU), an image signal processor (ISP), a sensor hub processor, or a communication processor (CP)), which may operate independently of the main processor or in conjunction with the main processor, etc. The processor 1030 may control components of the electronic device 1000 and process requests thereof.

[0101] The storage device 1040 may store data required to operate the components of the electronic device 1000 (e.g., images (still images or moving images captured by the image capture device), data processed by the processor 1030, neural networks used by the image processing device 1020, etc.), and instructions for executing functions. The storage device 1040 may include a computer-readable storage medium, such as a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a magnetic hard disk, an optical disk, a flash memory, an electrically programmable read-only memory (EPROM), or other types of computer-readable storage media known in the art.

[0102] The output device 1050 can output the image captured by the image capture device 1010 and / or the data processed by the processor 1030 in a visual / non-visual manner. The output device 1050 may include a sound output device, a display device (e.g., a display), an audio module, and / or a tactile module. The image processed by the image processing device 1020 can be displayed on the display to improve the user's experience of the image.

[0103] The communication device 1060 can support the establishment of a direct (e.g., wired) communication channel and / or a wireless communication channel between the electronic device 1000 and other electronic devices, servers, or sensor devices within the network environment, and perform communication via the established communication channel using various communication technologies. The communication device 1060 can transmit the image captured by the image capture device 1010, the bit stream output by the image processing device 1020 during the image encoding process, the image decoded during the image decoding process, and / or the data processed by the processor 1030 to other electronic devices. In addition, the communication device 1060 can receive the image to be processed from the cloud device or other electronic devices, and store the received image in the storage device 1040.

[0104] In addition, the electronic device 1000 may also include sensor devices for detecting various data (for example, an accelerometer, a gyroscope, a magnetic field sensor, a proximity sensor, an illumination sensor, a fingerprint sensor, etc.), input devices for receiving instructions and / or data from a user (for example, a microphone, a mouse, a keyboard and / or a digital pen (for example, a stylus, etc.), etc.).

[0105] Fig.11 is a block diagram illustrating an example of an image encoding and decoding process performed by an image processing apparatus included in an electronic device.

[0106] refer to Fig.11 , an example of image encoding and decoding processes performed by the image encoding device 1110 and the image decoding device 1120 of the image processing device 1100 will be described below.

[0107] In the image encoding process, the image block IB to be encoded is input to the first regularized stream neural network 1131 to be transformed into a potential representation. The transformed potential representation is input to the super prior encoder 1132 to be transformed into a super prior potential representation. The super prior potential representation is input to the second entropy encoder 1133 to output a bit stream of the super prior potential representation.

[0108] The bitstream of the super-a priori potential representation is input to the second entropy decoder 1134 to reconstruct the super-a priori potential representation. The second entropy encoder 1133 and / or the second entropy decoder 1134 can perform entropy coding and / or decoding based on the probability distribution generated by training in advance. The reconstructed super-a priori potential representation is input to the super-a priori decoder 1135, so that the super-a priori is output by the super-a priori decoder 1135, and the super-a priori is input to the context estimator 1136 to output the probability distribution. The output probability distribution can be input to the first entropy encoder 1137 and / or the first entropy decoder 1138. In addition, the potential representation transformed by the first regularized flow neural network 1131 is input to the first entropy encoder 1137 and entropy encoded, thereby outputting the bitstream of the potential representation. The output bitstream is entropy decoded by the first entropy decoder 1138 to reconstruct the potential representation, and the potential representation is input to the context estimator 1136 for estimating the probability distribution of the subsequent potential representation. Two bitstreams are output by performing image coding.

[0109] In the image decoding process, the image block is reconstructed by using the two bit streams output during the image encoding process as input. The bit stream output by the second entropy encoder 1133 is input to the second entropy decoder 1134 for entropy decoding to reconstruct the super prior potential representation. The reconstructed super prior potential representation is input to the super prior decoder 1135 to output the super prior, and the output super prior is input to the context estimator 1136 to estimate the probability distribution. The probability distribution estimated by the context estimator 1136 and the bit stream output by the first entropy encoder 1137 during the image encoding process are input to the first entropy decoder 1138, and entropy decoding is performed by the first entropy decoder 1138 to reconstruct the potential representation. The reconstructed potential representation is input to the second regularized flow neural network 1139 to reconstruct the image block DIB.

[0110] The present disclosure may be implemented as computer-readable codes written on a computer-readable recording medium. The computer-readable recording medium may be any type of recording device that stores data in a computer-readable manner.

[0111] Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, and carrier wave (e.g., data transmission via the Internet). The computer-readable recording medium can be distributed on multiple computer systems connected to the network so that computer-readable codes are written to and executed from the computer systems in a distributed manner. A programmer of ordinary skill in the art to which the present invention belongs can easily infer the functional programs, codes, and code segments required to implement the present invention.

[0112] The present disclosure has been described above with reference to preferred embodiments. However, it is obvious to those skilled in the art that various changes and modifications can be made without changing the technical concept and essential features of the present disclosure. Therefore, it is clear that the above embodiments are illustrative in all aspects and are not intended to limit the present disclosure.

Claims

1. An image encoding method, comprising: transforming the image patch into a first latent representation based on a reversible neural network; transforming the first latent representation into a second latent representation based on an irreversible neural network; estimating a first probability distribution of the first latent representation based on the second latent representation; and Entropy encoding is performed on the first latent representation based on the first probability distribution.

2. The method according to claim 1, wherein: The reversible neural network includes a regularized flow neural network including one or more coupled layers.

3. The method according to claim 2, wherein: The parameters of the one or more coupling layers are obtained by training based on at least one of a neural network structure or a data distribution.

4. The method according to claim 1, wherein: Estimating the first probability distribution includes: dividing the first latent representation into a plurality of groups, and The first probability distribution is estimated for each group in the plurality of groups.

5. The method according to claim 1, further comprising: Entropy encoding is performed on the second latent representation based on a second probability distribution.

6. The method according to claim 1, wherein: Transforming the first latent representation into the second latent representation includes: transforming the first latent representation into the second latent representation using a super-a priori encoder.

7. The method according to claim 1, wherein: Estimating the first probability distribution includes obtaining a super-prior as a probability distribution of the first latent representation using a super-prior decoder.

8. The method according to claim 1, wherein: Estimating the first probability distribution includes: obtaining a super-prior based on an entropy decoding result of the second latent representation using a super-prior decoder; and The first probability distribution of the first latent representation is estimated based on the super prior using a context estimator.

9. The method according to claim 1, further comprising: The image block is divided into a plurality of sub-blocks.

10. The method according to claim 9, further comprising: One or more coupled layers based on a regularized flow neural network perform operations within each of the plurality of sub-blocks.

11. The method according to claim 9, further comprising: One or more coupling layers based on a regularized flow neural network perform operations between the plurality of sub-blocks.

12. The method according to claim 11, wherein: The operations between the plurality of sub-blocks are performed by at least one of a super a priori encoder, a super a priori decoder, or a context estimator.

13. The method according to claim 1, wherein: Transforming the image block into the first latent representation includes transforming the image block into the first latent representation by hierarchically using two or more reversible neural network based modules.

14. The method according to claim 13, wherein: Transforming the image block into the first latent representation includes: inputting a portion of the third latent representation transformed by the reversible neural network-based module of the previous layer into the reversible neural network-based module of the next layer to transform the portion of the third latent representation into a fourth latent representation, and combining the remaining portion of the third latent representation with the fourth latent representation to transform the combined latent representation into the first latent representation.

15. An image decoding method, comprising: Receiving a first bitstream of a first latent representation, wherein the first latent representation is obtained by transforming an image block based on a reversible neural network; Receiving a second bitstream of a second latent representation, wherein the second latent representation is obtained by transforming the first latent representation based on an irreversible neural network; estimating a first probability distribution of the first latent representation based on the second bitstream; and The image block is reconstructed based on the first bitstream and the first probability distribution.

16. The method according to claim 15, further comprising: performing entropy decoding on the first bitstream based on the first probability distribution of the first latent representation, Wherein, reconstructing the image block includes: reconstructing the image block based on the entropy decoding result of the first bit stream and the first probability distribution.

17. The method according to claim 15, further comprising: The second bitstream is entropy decoded based on a second probability distribution.

18. The method according to claim 17, wherein: Estimating the first probability distribution of the first latent representation includes: obtaining a super prior based on a decoding result of the second bitstream, and estimating the first probability distribution of the first latent representation based on the super prior.

19. An electronic device comprising: a memory storing one or more instructions, and A processor configured to execute the one or more instructions to implement: a reversible neural network configured to transform the image patch into a first latent representation; Irreversible neural network, configured as: transforming the first latent representation into a second latent representation, and estimating a first probability distribution of the first latent representation based on the second latent representation; and An entropy encoder is configured to perform entropy encoding on the first latent representation based on the first probability distribution.

20. An electronic device, comprising: a memory storing one or more instructions, and A processor configured to execute the one or more instructions to implement: Irreversible neural network, configured as: receiving a first bitstream of a first latent representation obtained by transforming an image block and a second bitstream of a second latent representation obtained by transforming the first latent representation, and estimating a probability distribution of the first latent representation; and A reversible neural network is configured to reconstruct the image block based on the first bit stream and the estimated probability distribution.

Citation Information

Patent Citations

  • Pharmaceutical composition and method for preparing and using the same

    KR1020230152118A

  • Semiconductor devices and data storage systems including the same

    KR1020240000749A