Implicit neural representations driven by updated dictionaries for image and video compression
By decomposing the weights of the Implicit Neural Representation (INR) network into head and tail layers and encoding them using an updated dictionary, the high computational complexity of neural compression techniques is solved, achieving efficient signal compression and accurate reconstruction.
Patent Information
- Application Number
- CN202580012396.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-09
- Filing Date
- 2025-01-24
- Publication Date
- 2026-08-25
AI Technical Summary
Existing neural compression techniques have high computational complexity, making it difficult to effectively utilize global and local information for efficient signal compression.
An implicit neural representation (INR) network is used, which approximates the weights using a dictionary base and decomposes them into a head layer and a tail layer. The head layer is responsible for global information, and the tail layer is responsible for local information. Only the tail layer weights and the additional information from the head layer are transmitted, and the updated dictionary is used for encoding and decoding.
It reduces computational complexity, improves signal compression efficiency, and enhances the accuracy and flexibility of signal reconstruction.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit of European Patent Application No. 24305214.9, filed on February 9, 2024, which is incorporated herein by reference in its entirety. Technical Field
[0002] This embodiment generally relates to a method and apparatus for neural compression. Background Technology
[0003] Neural compression, or learning-based compression, applies neural networks and other machine learning methods to data compression. These techniques are currently being researched by MPEG, and a new ad hoc group within Working Group 4 focuses on compression based on implicit neural representations (INR). Typically, INR-based compression techniques have significantly lower computational complexity than end-to-end neural compression schemes. Summary of the Invention
[0004] In one implementation, the weights of the implicit neural representation (INR) network can be approximated by a dictionary base, which can be learned from a large dataset and is known to both the encoder and decoder. This dictionary can be updated. In an example, the updates can be obtained by the encoder and transmitted to the decoder.
[0005] In one implementation, the information content in a signal can be modeled as a combination of global and local information. Global information is common and shared across all natural signals and can be learned from a large dataset. However, local information is specific to each signal. Therefore, the INR network can be decomposed into a head layer and a tail layer, where the head layer handles global information and the tail layer handles local information. The weights of the head layer can be approximated by a dictionary base, which can be learned from a large dataset and is known to both the encoder and decoder. This dictionary can be updated. In the example, the update is obtained by the encoder and transmitted to the decoder. Therefore, only the weights of the tail layer are transmitted along with the update and some additional information about the head layer. Attached Figure Description
[0006] Figure 1 A block diagram of a system that can implement aspects of this embodiment is shown; Figure 2 A simple neural network for implicit neural representation (INR) is shown; Figure 3 This illustrates a typical process of encoding a signal using INR; Figure 4 A flowchart illustrating the encoding method based on the example is provided. Figure 5 A flowchart illustrating the decoding method based on the example is provided. Figure 6 A flowchart illustrating the encoding method based on the example is provided. Figure 7 A flowchart illustrating the decoding method based on the example is provided; and Figure 8 A flowchart illustrating a decoding method based on another example is provided. Detailed Implementation
[0007] This application describes various aspects, including tools, features, embodiments, models, solutions, etc. Many of these aspects are described in detail, and often in a manner that may sound restrictive, at least to demonstrate their respective characteristics. However, this is for clarity of description and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in prior applications.
[0008] The aspects described and envisioned in this application can be implemented in many different forms. The following... Figure 1 , Figure 2 and Figure 3 Some embodiments have been provided, but other embodiments are contemplated, and... Figure 1 , Figure 2 and Figure 3 The discussion does not limit the breadth of implementation methods. At least one aspect generally relates to video encoding and decoding, and at least another aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) storing instructions for encoding or decoding video data according to any described method, and / or a computer-readable storage medium storing a bitstream generated according to any described method.
[0009] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “encoded” and “coded” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.
[0010] This document describes various methods, each including one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, in various embodiments, terms such as "first," "second," etc., can be used to modify elements, components, steps, operations, etc., for example, "first decoding" and "second decoding." The use of such terms does not imply any requirement for the order of the modified operations unless a specific requirement exists. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or within an overlapping period of time with the second decoding.
[0011] For clarity, throughout the embodiments described herein, met conditions, unmet conditions, and configured condition parameters are described relative to a threshold (e.g., greater than or less than), a threshold value, and a configured threshold value. For example, met conditions may be described as being above a threshold value, while unmet conditions (e.g., performance criteria) may be described as being below a threshold value. The embodiments described herein are not limited to threshold-based conditions. Any other types of conditions and parameters (e.g., belonging to or not belonging to a range of values) may be applied to the embodiments described herein.
[0012] This aspect is not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations, whether pre-existing or developed in the future, as well as any extensions of such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically excluded, the aspects described in this application may be used alone or in combination.
[0013] Figure 1 A block diagram illustrating an example of a system that can implement various aspects and embodiments is shown. System 100 may be embodied as a device including various components described below and configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, networked home appliances, and servers. Elements of system 100 may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 may be communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in this application.
[0014] System 100 includes at least one processor 110 configured to execute instructions loaded thereon for implementing, for example, the aspects described herein. Processor 110 may include embedded memory, input / output interfaces, and a variety of other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.
[0015] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 130 may be implemented as a separate element of system 100, or may be incorporated within processor 110 as a combination of hardware and software, as is known to those skilled in the art.
[0016] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of a variety of items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions thereof, bitstreams, matrices, variables, and intermediate or final results of processing from equations, formulas, operations, and operational logic.
[0017] In some embodiments, the memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., the processing device may be processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external dynamic volatile memory (such as RAM) is used as working memory for video encoding and decoding operations, for example for MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, 13818-1 is also known as H.222, 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Coding, a new standard developed by the Joint Video Experts Group JVET).
[0018] Inputs to the components of system 100 can be provided via a variety of input devices as indicated in box 105. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or sets of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Other examples ( Figure 1 (Not shown in the image) Includes composite video.
[0019] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band); (ii) down-converting the selected signal; (iii) band-limiting it again to a narrower frequency band to select, for example, a signal band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements (e.g., frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers) to perform these functions. The RF section may include a tuner that performs multiple of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above (and other) components, remove some of them, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0020] Furthermore, USB and / or HDMI terminals may include corresponding interface processors for connecting system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that multiple aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed in a separate input processing IC or within processor 110. Similarly, aspects of USB or HDMI interface processing may be implemented as needed in a separate interface IC or within processor 110. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on the output device.
[0021] Various components of system 100 can be housed within an integrated housing. Within this integrated housing, various components can be interconnected and transmit data therebetween using suitable connection means 115 (e.g., internal buses known in the art, including I2C buses, wiring, and printed circuit boards).
[0022] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.
[0023] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11 (IEEE stands for Institute of Electrical and Electronics Engineers). In these embodiments, the Wi-Fi signal is received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that transmits data via an HDMI connection to input module 105 to provide streaming data to system 100. Still other embodiments use an RF connection to input module 105 to provide streaming data to system 100. As described above, various embodiments provide data in a non-streaming manner. Furthermore, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0024] System 100 can provide output signals to a variety of output devices, including a display 165, a speaker 175, and other peripheral devices 185. The display 165 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be used in a television, tablet computer, laptop computer, cellular phone (mobile phone), or other device. The display 165 can also be integrated with other components (e.g., in a smartphone) or separate (e.g., an external monitor for a laptop computer). In examples of various embodiments, other peripheral devices 185 include one or more of a standalone digital video disc (or digital universal optical disc) (DVR, used for both terms), an optical disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functionality based on the output of system 100. For example, an optical disc player performs the function of playing the output of system 100.
[0025] In various embodiments, signaling such as AV.Link, CEC, or other communication protocols is used to transmit control signals between system 100 and display 165, speaker 175, or other peripheral devices 185. These protocols enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 can be integrated into a single unit with other components of system 100, such as in an electronic device (e.g., a television). In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.
[0026] Display 165 and speaker 175 may alternatively be separated from one or more other components, for example, if the RF section of input 105 is part of a separate set-top box. In various embodiments where display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0027] These embodiments may be executed by computer software implemented by processor 110, or by hardware, or a combination of hardware and software. As a non-limiting example, these embodiments may be implemented by one or more integrated circuits. Memory 120 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical storage devices, magnetic storage devices, semiconductor-based storage devices, fixed memory, and removable memory, as a non-limiting example. Processor 110 may be of any type suitable for the technical environment and may include one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures, as a non-limiting example.
[0028] Figure 2 A simple neural network for implicit neural representation (INR) is shown. This neural network for INR can be called an INR network. For clarity, we use two-dimensional signals (such as images) for illustration, but INR can be used to represent signals of any dimension. INR parameterizes a signal (e.g., an image, a 3D scene) as a function (200) that takes coordinates (210) (e.g., image coordinates) as input and outputs possible approximations (220) of the signal at those coordinates (e.g., luminance and / or chrominance values). INR has recently been applied to images, 2D videos, or 3D objects, as well as other applications. In the case of images, the input (210) can be pixel coordinates. The INR outputs the color value of the input pixel (220). In the case of video, the output is similar, but the input can include frame indexes in addition to pixel coordinates. INR can be used to reconstruct a signal by calculating signal values for each required coordinate input.
[0029] INR networks (200) are typically neural networks composed of multiple neural layers (e.g., fully connected layers). Figure 2 In this network, there are four neural layers. The intermediate outputs are represented by circles. Each neural layer can be described as a function that first multiplies the input by a tensor, adds a vector called the bias, and then applies a nonlinear function to the result. In this document, we can also refer to a "neural layer" simply as a "layer." The shape (and other properties) of the tensor and the type of the nonlinear function are called the network's architecture. We will use the term "weights" to represent the values of the tensor and biases. The weights and, if applicable, the parameters of the nonlinear function are called the network's parameters. The architecture and parameters define the "model". We will use... To indicate by Parameterized INR function.
[0030] Figure 3 This illustrates a typical process for encoding a signal using an INR. This is achieved by acquiring (310) (e.g., learning or optimizing) the parameters of the INR network. This is accomplished by (or a subset thereof) and optionally encoding (320) parameters to create an output bitstream. Parameters It can be used to reconstruct signals. For signals of size... Image ,parameter Alternatively, (the selected subset) can be obtained, for example, by minimizing (e.g., optimizing) the following loss function: (1) (2) in It is distortion, its quantification is The predicted (e.g., reconstructed) image compared to the original image The differences between them It is the bit rate of the encoding parameters. yes and The trade-off parameters between them. It can be any differentiable distortion measure, such as the mean square error in equation (2). and It is an image Width and height. Other metrics, such as LPIPS (Learned Perceptual Image Patch Similarity), can also be used in this case. Weights Optimization can be performed using machine learning methods, such as batch / stochastic gradient descent. For each image... There exists a specific INR function. It is overfitted to a given image. .Depend on The quality of the reconstructed image depends on the size of the neural network. Since the weights serve as descriptors for the image, the larger the neural network, the higher the bit length. On the other hand, limiting the number of weights will reduce the bit length at the cost of distortion.
[0031] Image weights It can be encoded and transmitted to a receiver (e.g., a decoder) that is configured to receive the decoded weights. Reconstruct the image.
[0032] To reconstruct (e.g., decompress) the signal, evaluate at all relevant coordinates. These coordinates can be selected during decoding. For images or videos, a typical choice is all pixel coordinates. For example, for a 256×256 pixel image, these coordinates could be pairs of coordinates for all x∈{0,1,…,255} and y∈{0,1,…,255}. Other options are also available, such as upsampling, downsampling, or expanding the original image.
[0033] A signal or a portion of a signal (also known as a partition) can be better encoded by approximating some parameters of an INR network using dictionary approximation (e.g., a pre-known or trained dictionary). By doing so, the INR network can be encoded using the weights of the unapproximated portions and some additional information describing the approximated portions. In the example, the INR can be partitioned into a head layer and a tail layer, such as... In the example, only the weights of the tail layer can be transmitted. And some additional information about the head layer. This is an example of decomposition into a composite representation, and other types of decomposition can be used without loss of generality.
[0034] In this example, a single INR is used to encode the entire image. In other examples, the input can be partitioned, and different INRs are used for different parts / partitions.
[0035] Sparse dictionary learning is a representation learning method that aims to find a sparse representation of input data in the form of linear combinations of basic elements. These elements are called atoms, and they make up a dictionary. Dictionary learning algorithms learn a set of atoms (also called basis functions) from some training signal, such that the signal can then be approximated as a linear combination of only a few atoms.
[0036] To share redundant information between images or increase the capacity of the INR network, it is possible to learn [a method / mechanism]. A dictionary of atoms This allows the head layer (e.g., each head layer) to use a sparse linear combination of atoms from a dictionary. To approximate (represent). That is, (3) Sparse coefficients used in sparse linear combinations (Also known as the INR coefficient or approximate coefficient) is determined, for example, by optimizing the following loss function (e.g., approximating). , (4) in It's the weight of the top tier. It's a dictionary I learned. The goal is to determine (e.g., optimize) the sparse coefficients. To impose sparsity on the coefficient vector, the L1 norm is used, while... It is the trade-off between two terms in the equation.
[0037] Once the sparse coefficients are estimated, they can be transmitted to the decoder. On the decoder side, they can be obtained from the dictionary. An approximate head layer is obtained (e.g., computed) from the received sparse coefficients. Note that the dictionary can be known to both the encoder and decoder. In this case, the head layer size can be large enough to increase the representation capacity of the INR, since the dictionary will not be included in the bitstream.
[0038] In another example, certain layers of the INR network can be approximated using a dictionary approximation, thereby encoding the signal or a portion of the signal. This also allows the INR network to be encoded using the weights of the unapproximated portions and some additional information describing the approximated portions. As in the previous example, the INR can be divided into head and tail layers, such as... For approximations, it is possible to learn those with... A dictionary of INR functions This allows the head layer (e.g., each head layer) to be approximated (represented) using a sparse linear combination of the dictionary's INR functions. That is, .
[0039] Such an approximation can be achieved, for example, by optimizing the following costs: (5) in It is the spatial support for the INR approximation (the entire image, blocks, superpixels, etc.). This optimization problem can also be modified to optimize the dictionary. and / or weights It can also include additional losses, such as those imposed on some or all of the optimization parameters. or loss.
[0040] The dictionary used to encode (e.g., by approximating the INR parameters with a dictionary) a portion of the signal may not be optimal. This is often the case when the dictionary is computed on a different part of the signal, or learned on a dataset completely different from the part of the signal to be encoded. This can lead to suboptimal encoding of the current part of the signal. For example, when the signal is video, the dictionary might be learned on the coding unit of the first frame of the video (or another image segment) and reused for all subsequent frames. This dictionary may not be optimal for approximating the parameters of the INR representation of the coding unit in the second or subsequent frames.
[0041] In contrast, the following discloses an encoding method (and correspondingly a decoding method), in which a dictionary... The atoms (also known as basis functions) can be updated. In the example, an encoding method for updating dictionary atoms (dictionary functions) used to approximate INR parameters (e.g., for parts of the signal) is thus disclosed. These updates can be encoded in a bitstream representing the video and may be transmitted. A corresponding decoding method is also disclosed that uses the decoded updates to reconstruct the video signal (e.g., an image, video, or 3D scene).
[0042] Figure 4 A flowchart illustrating encoding method 400 based on the example is provided.
[0043] In S420, the signal 410 to be encoded can be divided into parts (also known as partitions). Video data can be divided into parts (also known as partitions). For example, the data can be divided into coding units. The partitioning of video data can be of various types, such as quadtree, superpixel, binary tree, ternary tree partitioning, or multiple tree types. The partitioning is known to the decoder; for example, it can be transmitted in the bitstream. This step is optional.
[0044] In S430, an INR is obtained (e.g., learned), that is, the INR parameters (e.g., weights) are obtained. For example, if the signal is divided in S420, then for each partition get.
[0045] In S440, obtain the dictionary. (e.g., a dictionary of atoms), obtained, for example, from memory or from previously encoded signals, or from previous frames of a video or 3D scene.
[0046] In S445, the dictionary is updated based on the INR parameters learned in S430. To do this, the dictionary update can be calculated. Updated dictionary Therefore Then, the updated dictionary. It can be used to use updated dictionaries The sparse linear combination of atoms is used to approximate (e.g., represent) the INR parameters. This update can, for example, be used to produce a vector (i.e., sparse coefficients). This is done to determine interesting characteristics (e.g., small size), for example by solving the following optimization problem:
[0047] In S450, sparse coefficients are obtained from the updated dictionary based on the INR parameters learned in S430. For example, the following optimization problem can be solved to achieve this. Approximate each parameter vector : .
[0048] In S460, all parameters of INR (e.g., sparsity coefficients) Dictionary update The partitions are encoded in bitstream 470. Instead of directly encoding the INR weights determined in S430, the sparse coefficients are encoded. Indeed, the INR weights determined in S430 are approximated by a sparse linear combination of dictionary atoms, where the weights of the linear combination are the sparse coefficients. These coefficients can also be quantized by the quantization function Q: This quantization can also be performed simultaneously with the optimization in step S450.
[0049] Figure 5 A flowchart of decoding method 500 based on the example is depicted.
[0050] In S520, parameters included in the decoded bitstream (e.g., dictionary updates) ,coefficient And possible partitions). If partitions are used, the signal partitions can be decoded / recalculated.
[0051] In S530, obtain the dictionary. (e.g., a dictionary of atoms), obtained, for example, from memory or from previously encoded signals, or from previous frames of a video or 3D scene.
[0052] In S540, the dictionary is updated based on the updates decoded in S520. The updated dictionary... Therefore .
[0053] In S550, an INR is obtained, that is, the INR parameters (e.g., weights) are obtained. For example, if the signal is divided in S420, then for each partition Obtained. The INR parameter is obtained using the updated dictionary. It is obtained through a sparse linear combination of atoms. The weights of the linear combination are the coefficients decoded by S520. Parameters The dictionary can be determined (e.g., calculated or reconstructed) as updated. A linear combination of atoms, for example, if the coefficients are quantized. , or if the coefficient When directly encoded in a bitstream, it is .
[0054] In S560, parameters can be used. Reconstruct signal data (e.g., video frames) from an INR network. As an example, inference is performed using an INR network for all coordinates in a corresponding partition.
[0055] In another example, the INR network An INR network can be decomposed into at least two parts, and the parameters of at least one part are approximated using a multi-scale dictionary approximation. By doing so, the INR network can be encoded using the weights of the unapproximated part and some additional information describing the approximated part. There are several possible methods to decompose an INR network into parts. For example, a part can consist of any subset of biases and / or weights and / or parameters of nonlinear functions and / or these elements. Such a subset can, for example, be defined as a subset of layers, such as the last... Layer, or last Layer bias, or subset of neurons. In the following text, we will use an INR divided into a head layer and a tail layer as an example, such as... In this model, the head layer is responsible for representing global information, while the tail layer represents local information. In this example, the decomposition into head and tail layers is determined by the user. Leveraging the composite nature of neural networks, the function is decomposed into a combination / composition of head (global information) and tail (local) layers. Typically, the head will have a larger capacity for versatility (e.g., from 6 to 10 layers), while the tail is a fitted portion with a limited number of layers (1 to 3). The user chooses the settings for what constitutes a head and what constitutes a tail, which may depend on the complexity of the signal to be encoded.
[0056] The weight of the top layer You can use a dictionary To approximate. A dictionary defined as the set of INR atoms (and corresponding INR functions). These INR atoms (and corresponding INR functions) are known to both the encoder and decoder. They can be randomly selected, learned on the first frame of the video sequence, or learned using a large dataset (e.g., an image database). These INR atoms (and corresponding INR functions) can be included in the bitstream earlier, thus being transmitted to the decoder. In another example, these INR atoms (and corresponding INR functions) can be normalized, in which case they are known to both the encoder and decoder without needing to be transmitted.
[0057] Figure 6 A flowchart illustrating an encoding method 600 based on an example is provided. This method can be used to encode video data, such as images, videos, or 3D scenes. The video data can be divided into portions (also known as partitions). For example, it can be divided into coding units. Video data can be partitioned in various ways, such as quadtrees, superpixels, binary trees, ternary trees, or multiple tree types. The partitioning is known to the decoder; for example, it can be transmitted in a bitstream.
[0058] In S602, INR parameters (e.g., weights) are obtained from the video data to be encoded (e.g., from an image). Indeed, for images There exists a specific INR function. It is overfitted to a given image. In the example, the INR function In the signal section to be encoded The INR function is overfitted. In the example, the INR function... It consists of a head layer and a tail layer, such as Therefore, it can be a part of the signal. (For example, for each part of the signal) obtain the INR parameters (e.g., train the INR) to thus provide parameters for that part. Generate parameter vector (corresponding to the first layer) and (Corresponding to the tail layer).
[0059] In S604, the dictionary is calculated. Update , including the dictionary It can be used to represent (e.g., approximate) parameters (also known as parameter vectors). Updating the dictionary allows for better encoding of the parameter vector. For example, calculating dictionaries. An update of one atom (e.g., each atom). Updated dictionary Therefore Then, the updated dictionary. It can be used to use updated dictionaries A sparse linear combination of atoms is used to approximate (e.g., represent) the head layer (e.g., the parameters of the head layer). This update could, for example, be used to generate vectors (i.e., sparse coefficients). This is done to determine interesting characteristics (e.g., small size), for example by solving the following optimization problem:
[0060] Other methods for updating the dictionary are also possible. In another example, it might be desirable to optimize the update by solving an optimization problem such as a rate-distortion optimization problem to minimize the size of the update:
[0061] in It controls the size of the update code. The importance of hyperparameters.
[0062] Solutions to such problems can be found, for example, using alternating minimization, by alternately fixing one variable and optimizing another. More specifically, given a dictionary... and updates The sparsity coefficients can be optimized using any sparse solver (such as Lasso). In the next step, the sparsity coefficient is fixed. And use (e.g., stochastic) gradient descent or coordinate descent to optimize dictionary updates. Continue (e.g., repeat) these alternating steps until convergence or a certain number of iterations are reached. Other algorithms, such as genetic algorithms, random sampling, or simulated annealing, can also be used.
[0063] In the example above, the optimization problem did not consider the tail layer. In another example, the parameters of the tail layer could also be considered. Dictionary updates can be calculated later in the case of optimization. Therefore, the optimization problem to be solved can be as follows:
[0064] The above process primarily describes the sequential scheme for updating the dictionary, calculating other parameters, and encoding. Besides those listed above, there are many possible variations of this process. For example, several elements can be optimized simultaneously. Dictionary Update ,parameter and / or parameters This can be achieved using any optimization algorithm, such as greedy search, gradient descent with a specific loss, genetic algorithms, machine learning algorithms, etc. This can be accomplished by solving any of the previously mentioned optimization problems and extracting the optimal values for these parameters.
[0065] In S606, sparse INR coefficients are obtained (e.g., sparse). ),in Approximate parameters The coefficients of linear combinations of atoms in the updated dictionary. Sparse coefficients ( The sparsity coefficients can be obtained directly from S604. Indeed, the sparsity coefficients ( The values can be determined and stored together with the update in S604, in which case they do not need to be calculated in S606. As mentioned above, several elements can be optimized simultaneously. Dictionary update ,parameter and / or parameters Any optimization algorithm can be used together for optimization.
[0066] In another example, sparse INR coefficients are calculated in S606 ( (by updating the dictionary) Approximate parameters of the head layer. In other words, based on the updated dictionary. To approximate the parameter vector This approximation can be computed using any readily available method, including those that can be used with the original dictionary. Methods for calculating approximations. For example, the following optimization problem can be solved to obtain an approximation. Approximate each parameter vector : .
[0067] In S608, dictionary updates ,parameter and parameters Encoded to create a bitstream, for example, for transmission or later use. This may involve using an entropy encoder and / or quantization. Dictionary update. ,parameter and tail layer parameters They can be encoded using any existing method. In particular, since they are parameters of a neural network, they can be encoded using existing codecs specifically designed for encoding such parameters, such as MPEG-NNC. (Dictionary update) This can be achieved using special features designed for encoding updates in such codecs, such as incremental neural network update data in MPEG-NNC. Parameters Entropy encoding can be performed directly, or quantization (e.g., using fixed-bit quantization) can be performed and then encoded in the bitstream, as shown below: ○ Find the maximum value:
[0068] ○ Normalize the value based on the maximum value:
[0069] ○ Use Each bit is quantized using fixed-bit quantization to obtain the symbol to be transmitted.
[0070] These parameters (i.e., dictionary update) ,parameter and parameters The encoding technique can be standardized and known to both the encoder and decoder. Alternatively, the encoder can be allowed to choose its own encoding technique. In the latter case, the chosen technique can be included in the bitstream and may be entropy-encoded. Therefore, the resulting bitstream can contain dictionary updates. ,coefficient And other parameters .
[0071] Figure 7 A flowchart of decoding method 700 based on the example is depicted.
[0072] In S702, the parameters included in the decoded bitstream (e.g., dictionary updates for all parts) INR coefficient And other parameters If partitioning is used, the signal partitions can be decoded / recalculated.
[0073] The following steps can be applied sequentially to each part of the signal to reconstruct the entire signal.
[0074] In S704, from the updated dictionary and INR coefficient Obtain (e.g., approximate or compute) the head layer. More precisely, obtain the parameters of the head layer. It is from the updated dictionary. and INR coefficient Parameters that are obtained (e.g., calculated). The dictionary can be determined (e.g., calculated or reconstructed) as updated. A linear combination of atoms, for example, if the coefficients are quantized. , or if When directly encoded in a bitstream, it is Updated dictionary The original dictionary known from the decoder and decoding updates Obtained.
[0075] In S706, it is possible to obtain parameters and The video data is reconstructed using an INR network composed of parameterized head and tail layers. As an example, the INR network is used to perform inference for all coordinates in the corresponding partition.
[0076] In another example, the dictionary can be composed of... Composed of several INR functions, such that the head layer (e.g., each head layer) is approximated (represented) using a sparse linear combination of the dictionary's INR functions. Update Modify the parameters of dictionary functions:
[0077] The encoding process is similar to Figure 6 However, the dictionary update calculation in S604 and the approximate coefficients obtained in S606 differ in the cost function of the optimization, as does the signal reconstruction. This difference stems from the different uses of the dictionary. In S604, such an update can be achieved, for example, by optimizing the following cost:
[0078] This objective may also involve terms aimed at minimizing the bitstream size required for encoding updates:
[0079] Optimization of the tail parameters can also be considered:
[0080] In S606, when the dictionary update has been calculated, the approximation coefficients... For example, the cost can be calculated for each part of the signal by optimizing the following:
[0081] These coefficients may have already been calculated in step S604, or may be adjusted using another loss from any loss used in step S604.
[0082] Figure 8 A flowchart of decoding method 800 based on the example is depicted.
[0083] In S802, the parameters included in the decoded bitstream (e.g., dictionary updates for all parts) ,coefficient And other parameters If partitioning is used, the signal partitions can be decoded / recalculated.
[0084] The following steps can be applied sequentially to each part of the signal to reconstruct the entire signal.
[0085] In S804, from the updated dictionary and INR coefficient Obtain (e.g., approximate or compute) the head layer. The head layer can be approximated as a linear combination of updated basis functions, for example, such as... Updated basis functions The original dictionary known from the decoder The original basis functions and the update of the decoder Obtained. In other words, the parameters of the updated basis functions. From the dictionary and updates Obtained (e.g., calculated).
[0086] In S806, it is possible to obtain the approximate head layer and the parameters. The parameterized tail layer of the INR network reconstructs the video data, with the approximation of the head layer. From the updated dictionary and decoding parameters Parameterization. As an example, inference is performed using an INR network for all coordinates in the corresponding partition.
[0087] Furthermore, this aspect is not limited to ECM, VVC, or HEVC, and can be applied to, for example, other standards and recommendations, as well as any extensions of such standards and recommendations. Unless otherwise stated or technically excluded, the aspects described in this application may be used alone or in combination.
[0088] Various numerical values are used in this application. The specific numerical values are for illustrative purposes, and the aspects described are not limited to these specific numerical values.
[0089] Note that the syntax elements used in this article (such as terms in equations and algorithms, signal labels / names, etc., for example, updates) ,parameter (etc.) are descriptive terms. Therefore, they do not preclude the use of other syntactic element names.
[0090] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by a decoder in various embodiments described herein, such as decoding dictionary updates and reconstructing the video signal from the dictionary update.
[0091] As a further example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding; and in yet another embodiment, "decoding" refers to the entire image reconstruction process including entropy decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to the broader decoding process will be determined based on the specific context of the description and is believed to be well understood by those skilled in the art.
[0092] Various implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding” as used herein can encompass all or part of the process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transforming, quantizing, and entropy coding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder according to various embodiments described herein, such as determining a dictionary update and updating the encoded video signal by the dictionary.
[0093] As a further example, in one embodiment, “encoding” refers only to entropy encoding; in another embodiment, “encoding” refers only to differential encoding; and in yet another embodiment, “encoding” refers to a combination of differential and entropy encoding. Whether the phrase “encoding process” is intended to specifically refer to a subset of operations or to refer to a broader encoding process will be determined based on the specific context of the description and is believed to be well understood by those skilled in the art.
[0094] This disclosure has described various information fragments that can be transmitted or stored, such as syntax. This information can be packaged or arranged in a variety of ways, including those common in video standards, such as placing the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers or stripe headers), or SEI messages. Other available methods include those common in system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (Session Description Protocol), a format for describing multimedia communication sessions for session announcement and session invitation, such as as described in the RFC, and used in conjunction with RTP (Real-Time Transport Protocol) transmission.
[0095] b. DASH MPD (Media Presentation Description) descriptors, such as those used in DASH and transmitted via HTTP, are associated with representations or sets of representations to provide additional features for content representation.
[0096] c. RTP header extensions, for example, used during RTP streaming.
[0097] d. ISO basic media file formats, such as those used in OMAF, use boxes, which are object-oriented building blocks defined by unique type identifiers and lengths, also known as "atoms" in some specifications.
[0098] e. An HLS (HTTP Live Streaming) manifest transmitted over HTTP. The manifest can be associated, for example, with a version or set of versions of the content to provide characteristics of that version or set of versions.
[0099] When the accompanying drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding devices. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding methods / processes.
[0100] Some implementations involve rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often under constraints of computational complexity. Rate-distortion optimization is generally formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist for solving the rate-distortion optimization problem. For example, these approaches can be based on extensive testing of all options, where their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding are fully evaluated. Faster schemes can also be used to save encoding complexity, particularly by computing approximate distortion. A combination of both approaches can also be used. Other schemes evaluate only a subset of possible options. More generally, many schemes employ any of a variety of techniques to perform optimization, but the optimization is not necessarily a comprehensive evaluation of encoding costs and associated distortion.
[0101] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), implementations of the discussed features can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in a processor, where processor refers generally to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0102] References to "an embodiment" or "an embodiment" or "an implementation" or "implementation" and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" appearing in various places in this application, as well as any other variations, do not necessarily refer to the same embodiment.
[0103] Furthermore, this application may refer to "determining" various information fragments. Determining information may include one or more of, for example, estimated information, calculated information, predicted information, or information retrieved from memory.
[0104] Furthermore, this application may refer to "accessing" various information fragments. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, one or more of these.
[0105] Furthermore, this application may refer to "receiving" various information fragments. Like "accessing," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) one or more of these. Moreover, "receiving" is generally referred to in some way during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0106] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of…”, for example in the cases of “A / B”, “A and / or B”, and “at least one of A and B”, is intended to include selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). This can be extended to any multiple items listed, as is clear to a person skilled in the art.
[0107] Furthermore, as used herein, the term "signal" specifically refers to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a specific dictionary update. Thus, in these embodiments, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already possesses that specific parameter along with other parameters, signaling (implicit signaling) can be used to simply inform the decoder and select that specific parameter without transmission. Bit savings are achieved in various embodiments by avoiding the transmission of any actual function. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. While the foregoing refers to the verb form of the term "signal," the term "signal" may also be used herein as a noun.
[0108] As will be apparent to those skilled in the art, implementations can generate various signals formatted to carry information, which may, for example, be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiments. Such a signal may, for example, be formatted as an electromagnetic wave (e.g., using the radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted via a variety of different wired or wireless links as known. The signal may be stored on a processor-readable medium.
[0109] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more features, devices, or aspects described herein, individually or in any combination across various claim classes and types. For example, features described herein may be implemented in a bitstream or signal including information generated as described herein. This information may allow a decoder to decode the bitstream, the encoder, bitstream, and / or decoder according to any of the described embodiments. For example, features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, features described herein may be implemented by a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, features described herein may be implemented by a television, set-top box, cellular phone, tablet computer, or other electronic device performing decoding. The television, set-top box, cellular phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image reconstructed from the residual of a video bitstream). The television, set-top box, cellular phone, tablet computer, or other electronic device may receive a signal including an encoded image and perform decoding.
[0110] Several embodiments have been described above. The features of these embodiments may be provided individually or in any combination across a variety of claim classes and types.
Claims
1. A decoding method, comprising: Decode the coefficients, dictionary updates, and parameters of the second layer set of the INR (Hidden Neural Representation) network that has been decomposed into a first layer set and a second layer set; The first layer set is obtained based on a linear combination of basis functions weighted by the coefficients, wherein the basis functions are the basis functions of the dictionary updated by the dictionary update; as well as The image or 3D scene is reconstructed based on the INR network using the first set of layers and the second set of layers.
2. The method of claim 1, wherein obtaining the first layer set based on a linear combination of basis functions weighted by the coefficients comprises: The parameters of the first layer set are obtained as a linear combination of atoms weighted by the coefficients, where the atoms are atoms of the dictionary updated by the dictionary update.
3. The method of claim 1 or 2, wherein the image or 3D scene comprises a plurality of partitions, and wherein the decoding of coefficients, dictionary updates and parameters of the second layer set, the acquisition of the first layer set and the reconstruction of the image or 3D scene are performed for each of the plurality of partitions.
4. The method of any one of claims 1 to 3, wherein the first layer set corresponds to global information of the image or 3D scene, and the second layer set corresponds to local information of the image or 3D scene.
5. A method for encoding video data representing an image or a 3D scene, comprising: Based on the video data, the parameters of the INR (Hidden Neural Representation) network, which is decomposed into a first-layer set and a second-layer set, are obtained; Update the dictionary of basis functions; Obtain the coefficients of a linear combination of the basis functions of the dictionary updated by the update, the linear combination being approximately used for the parameters of the first layer set; as well as The coefficients, the update, and the parameters of the second layer set are encoded.
6. The method of claim 5, wherein obtaining the coefficients of a linear combination of the basis functions of the dictionary updated by the update comprises: Obtain the coefficients of a linear combination of atoms in the dictionary updated by the update, the linear combination approximating the parameters used for the first-level set.
7. The method of claim 5 or 6, wherein the image or 3D scene comprises a plurality of partitions, and wherein the acquisition of the parameters of the INR network, the acquisition of the coefficients of the linear combination of atoms, and the encoding of the coefficients, the update, and the parameters of the second layer set are performed for each of the plurality of partitions.
8. The method of any one of claims 4 and 5, wherein the first layer set corresponds to global information of the image or 3D scene, and the second layer set corresponds to local information of the image or 3D scene.
9. A decoding apparatus comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform: Decode the coefficients, dictionary updates, and parameters of the second layer set of the INR (Hidden Neural Representation) network that has been decomposed into a first layer set and a second layer set; The first layer set is obtained based on a linear combination of basis functions weighted by the coefficients, wherein the basis functions are the basis functions of the dictionary updated by the dictionary update; as well as The image or 3D scene is reconstructed based on the INR network using the first set of layers and the second set of layers.
10. The decoding apparatus of claim 9, wherein obtaining the first layer set based on a linear combination of basis functions weighted by the coefficients comprises: The parameters of the first layer set are obtained as a linear combination of atoms weighted by the coefficients, where the atoms are atoms of the dictionary updated by the dictionary update.
11. The decoding apparatus of claim 9 or 10, wherein the image or 3D scene comprises a plurality of partitions, and wherein the decoding of coefficients, dictionary updates, and parameters of a second layer set, the acquisition of the first layer set, and the reconstruction of the image or 3D scene are performed for each of the plurality of partitions.
12. The decoding apparatus of claims 9 to 11, wherein the first layer set corresponds to global information of the image or 3D scene, and the second layer set corresponds to local information of the image or 3D scene.
13. An encoding apparatus for encoding video data representing an image or a 3D scene, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform: Based on the video data, the parameters of the INR (Hidden Neural Representation) network, which is decomposed into a first-layer set and a second-layer set, are obtained; Update the dictionary of basis functions; Obtain the coefficients of a linear combination of the basis functions of the dictionary updated by the update, the linear combination being approximately used for the parameters of the first layer set; as well as The coefficients, the update, and the parameters of the second layer set are encoded.
14. The encoding apparatus of claim 13, wherein obtaining the coefficients of a linear combination of the basis functions of the dictionary updated by the update comprises: Obtain the coefficients of a linear combination of atoms in the dictionary updated by the update, the linear combination approximating the parameters used for the first-level set.
15. The encoding apparatus of claim 13 or 14, wherein the image or 3D scene comprises a plurality of partitions, and wherein the acquisition of the parameters of the INR network, the acquisition of the coefficients of the linear combination of atoms, and the encoding of the coefficients, the update, and the parameters of the second layer set are performed for each of the plurality of partitions.
16. The encoding apparatus of any one of claims 13 to 15, wherein the first layer set corresponds to global information of the image or 3D scene, and the second layer set corresponds to local information of the image or 3D scene.
17. A signal comprising video data representing an image or a 3D scene, said video data being formed by performing the method as claimed in any one of claims 5 to 8.
18. A computer program comprising program code instructions for implementing the method according to any one of claims 1 to 8 when executed by a processor.
19. A computer-readable storage medium having stored thereon instructions for implementing the method as described in any one of claims 1 to 8.