Residual encoding and decoding using neural fields

By using optional remodeling technology to encode/decode residual image sequences in a bilayer neural network system, the problem of poor image quality and coding efficiency in the prior art is solved, and more efficient image processing and support for different visual experiences is achieved.

CN120239969APending Publication Date: 2025-07-01DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080647.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-25
Filing Date
2023-10-24
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, when using neural fields for image processing, it is difficult to effectively encode and decode residual image sequences, resulting in poor image quality and coding efficiency.

Method used

A framework for a bilayer neural network system is proposed to encode/decode residual image sequences using optional remodeling, the base layer provides baseline representation, and the enhancement layer provides enhanced information using trained neural fields.

Benefits of technology

The framework can provide better image quality and encoding efficiency, and can generate new frames that have not been encountered in the training data, supporting applications with different visual experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239969A_ABST
    Figure CN120239969A_ABST
Patent Text Reader

Abstract

Methods and apparatus for residual encoding and decoding using a neural field. According to an example embodiment, an image processing method implemented at an electronic encoder includes accessing a plurality of residual images or remodeled images, each of the residual images representing a difference between a respective reference image and a base image, the different respective reference images corresponding to respective values of one or more first parameters; training a neural field network to represent a plurality of residual images, the neural field network responsive to a variable input of the one or more first parameters and characterized by a set of second parameters, training producing fixed values of the second parameters; and transmitting the base image and the fixed value of the second parameter to a corresponding electronic decoder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 1. Cross - reference to related applications

[0002] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 419,099, filed on October 25, 2022, which is hereby incorporated by reference in its entirety. 2. Technical field

[0003] Each example embodiment relates to image processing, and more specifically but not exclusively to modeling image sequences using neural fields. 3. Background art

[0004] Visual computing involves synthesizing, estimating, manipulating, displaying, storing, and transmitting data about objects and scenes across space and time. For example, in computer graphics, visual computing is used to synthesize three - dimensional (3D) shapes and two - dimensional (2D) images, render novel views of various scenes, and animate human joint movements. In computer vision, visual computing is used to reconstruct 3D appearance, shape, pose, and deformation. In human - computer interaction, visual computing is used to enable interactive research on spatio - temporal data. In robotic applications, visual computing is used to support the planning and execution of robotic actions.

[0005] Advances in machine learning have spurred the use of methods employing coordinate - based neural networks to solve certain visual computing problems. Such neural networks (commonly referred to as "neural fields") parameterize the physical properties of scenes and objects across space and time. Example applications of neural fields include the synthesis of 3D shapes and images, human animation, and pose estimation. Other applications of neural fields are currently being actively developed. Summary of the invention

[0006] Disclosed herein are various embodiments of a framework for encoding / decoding a residual image sequence with optional reshaping in a two - layer neural network system, where a base layer provides a baseline representation and an enhancement layer provides enhancement information using a trained neural field. The optional reshaping can be used to reduce the size of the neural network. Compared to what is achieved by directly applying encoding to the original image, the disclosed framework can beneficially provide better quality and / or encoding efficiency. Additionally, the disclosed framework can also be used to generate new frames not encountered in the training data, thus supporting various applications with different visual experiences. Several example applications of the disclosed framework include, but are not limited to, motion blur at different time scales, different dynamic ranges, and faces of different ages.

[0007] According to an example embodiment, a method for residual encoding using a neural field is provided, the method comprising: accessing, via a processor, a plurality of residual images or reshaped images, each of the residual images representing a difference between a corresponding reference image and a base image, different corresponding reference images corresponding to respective values of one or more first parameters; training, via the processor, a neural field network to represent the plurality of residual images or reshaped images, the neural field network responsive to a variable input of the one or more first parameters and characterized by a set of second parameters, the training resulting in fixed values of the second parameters; and transmitting the base image and the fixed values of the second parameters to an electronic decoder.

[0008] According to another example embodiment, a method for residual decoding using a neural field is provided, the method comprising: accessing, via a processor, a base image and a set of fixed parameter values of a neural field network trained to represent a plurality of residual images, the neural field network responsive to a variable input of one or more first parameters and characterized by a set of second parameters, the fixed parameter values being values of the second parameters; testing, via the processor, the neural field network using a selected value of the variable input to generate an output residual image; and calculating, via the processor, a reconstructed image by combining the output residual image and the base image.

[0009] According to yet another example embodiment, an apparatus for residual encoding is provided, the apparatus comprising: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus to at least perform the following operations: access a plurality of residual images or reshaped images, each of the residual images representing a difference between a corresponding reference image and a base image, different corresponding reference images corresponding to respective values of one or more first parameters; train a neural field network to represent the plurality of residual images, the neural field network responsive to a variable input of the one or more first parameters and characterized by a set of second parameters, the training resulting in fixed values of the second parameters; and transmit the base image and the fixed values of the second parameters to an electronic decoder.

[0010] According to another example embodiment, a device for residual decoding is provided, the device including: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to, with the at least one processor, cause the device to at least perform the following operations: access a base image and a set of fixed parameter values of a neural field network trained to represent a plurality of residual images, the neural field network responsive to a variable input of one or more first parameters and characterized by a set of second parameters, the fixed parameter values being values of the second parameters; test the neural field network using a selected value of the variable input to generate an output residual image; and calculate a reconstructed image by summing the output residual image and the base image. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] By way of example, various other aspects, features, and advantages of the disclosed embodiments will become more fully apparent from the following detailed description and the accompanying drawings, in which:

[0012] Figure 1 is a block diagram illustrating a multi-layer perceptron (MLP) that can be used to implement a neural field according to an embodiment.

[0013] Figure 2 is a block diagram illustrating an encoder of an MLP using Figure 1 according to an embodiment.

[0014] Figure 3 is a block diagram illustrating a decoder corresponding to the encoder of Figure 2 according to an embodiment.

[0015] Figure 4 is a block diagram illustrating a part of motion blur image processing using an encoder of Figure 2 according to an embodiment.

[0016] Figure 5 is a block diagram illustrating a computing device according to an embodiment. DETAILED DESCRIPTION

[0017] The present disclosure and its various aspects can be embodied in various forms, including: hardware, devices, or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces, and application programming interfaces; and hardware-implemented methods, signal processing circuits, memory arrays, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. The following description is only intended to give a general idea of the various aspects of the present disclosure and does not limit the scope of the present disclosure in any way.

[0018] In the following description, many details such as device configuration, timing, operations, etc. are set forth to provide an understanding of one or more aspects of the present disclosure. It will be apparent to those skilled in the art that these specific details are merely exemplary and are not intended to limit the scope of the present application.

[0019] Neural field

[0020] In the neural field framework, field quantities are generated by sampling coordinates and feeding the sampled coordinates into a neural network. For example, a Neural Radiance Field (NeRF) is an implicit 3D scene representation that takes spatial positions (x, y, z) and viewing directions (θ, φ) as inputs and generates corresponding predicted color textures and volume densities as outputs. The corresponding neural network can be trained, for example, using a set of 2D images with known camera poses and related intrinsic information. After training, the neural network can be used to render any view of the 3D scene by: (i) querying the corresponding 3D positions and viewing directions of each pixel in the view, and (ii) performing volume rendering to construct the projected 2D image.

[0021] Figure 1 is a block diagram illustrating a multi-layer perceptron (MLP, 100) that can be used to implement a neural field according to an embodiment. The MLP (100) has three layers (1101 to 1103). The first layer (1101) is the input layer. The next layer (1102) is the hidden layer. The third layer (1103) is the output layer. Generally, the MLP (100) can have N hidden layers, where N is a positive integer. Thus, Figure 1 the specific example of the MLP (100) illustrated in corresponds to N = 1. In some specific examples, the number N ranges from 1 to 10.

[0022] The MLP (100) is a fully connected feedforward neural network. The "fully connected" property means that there are corresponding weighted connections between each neural network (NN) node (also referred to as a "processing element", "neuron", or "artificial neuron") in the previous layer and each NN node in the adjacent subsequent layer. An example NN node can scale, sum, and bias the incoming signal and use an activation function to produce an output signal that is a static non-linear function of the biased sum. Depending on the layer in which a particular NN node is located, the output of that node can become one of the outputs of the neural network or be sent via the corresponding connection(s) to one or more other NN nodes. The corresponding weights and / or biases applied by each NN node can be changed (e.g., optimized) during the training (learning) operation mode and are typically fixed (i.e., constant) during the testing (working) operation mode. Various embodiments disclosed hereinafter can employ or rely on one or more neural networks, such as the MLP (100).

[0023] The term w of the weight matrix W of MLP(100) ab is the weight of the connection between the a-th NN node in the subsequent layer and the b-th NN node in the previous layer. For a layer with multiple NN nodes, its output can be expressed as a function f of the weighted sum of the input values , where n o is the number of NN nodes in the subsequent layer, and n i is the number of NN nodes in the previous layer. The function f is the activation function mentioned above. Mathematically

[0024] O = f(WI + B) (1)

[0025] where the vector B is the bias vector. A non-limiting example of the activation function f is the ReLU (Rectified Linear Unit) function, which is expressed as follows:

[0026]

[0027] In other examples, other suitable activation functions can also be used.

[0028] In some examples, one or more of the final MLP outputs are configured to be within the range [0,1]. In such examples, the output layer (e.g., the output layer (1103)) is configured to apply the sigmoid function S(x) to map a wider range of values into the range [0,1]. The sigmoid function S(x) is expressed as follows:

[0029]

[0030] Since some neural networks (such as MLP(100)) may be inherently biased towards preferentially learning low-frequency functions, its input can be generated by mapping the initial low-dimensional input to a high-dimensional space using a series of trigonometric functions γ, so that the output data better fits the high-frequency components. In various examples, the series γ acting on the coordinate p is defined as follows:

[0031]

[0032] where 2L is the number of trigonometric components of γ; and {l0, l1,..., l L-1} are integers. In some examples, l k = k.

[0033] In Figure 1In the simplified non - restrictive example shown, the layers (1101, 1102, 1103) of the MLP (100) have two, three, and one NN node (102) respectively. In some other examples, the number of NN nodes (102) in the MLP layer can range from 1 to 256. Different hidden layers can have different corresponding numbers of NN nodes (102) or the same number of NN nodes (102). In a specific MLP example using the above - mentioned positional encoding, the input layer has 41 NN nodes (102), and each of the five hidden layers has 256 NN nodes (102).

[0034] Residual encoding

[0035] Figure 2 is a block diagram illustrating an electronic encoder (200) according to an embodiment. The following is a reference to Figure 3 describe the corresponding electronic decoder (300). Due to the use of residual images, the encoder (200) and decoder (300) can beneficially produce reconstructed images with better image quality than at least some traditional methods. Compared with calculations relying on integer values, the residual images can be represented by floating - point values for more accurate calculations. Additionally, the encoder (200) and decoder (300) can be used to accurately interpolate neural - field residuals according to neighboring reference residuals along a user - selected dimension of a related image sequence. Thus, the decoder (300) can be operated to accurately predict new residuals and then generate the corresponding new image by adding the new residuals to the base image. The use of the forward reshaping function and the backward reshaping function enables the encoder (200) and decoder (300) to further improve the efficiency of the neural - field model, thereby providing better reconstructed - image quality under a neural network of the same size or lower signal overhead under the same image quality.

[0036] In various examples, the encoder (200) receives a target (e.g., reference) image sequence (202) as input. In some examples, each image in the sequence (202) is normalized so that its value is in the range between 0 and 1. In a specific example, the sequence (202) is a video - frame sequence (I0, …, I T-1 )), where I0 indicates the first frame in time, and I T-1 indicates the last frame in time. In other examples, the sequence (202) is not necessarily a sequence in the time domain. In such examples, the index t representing the order of the target images in the sequence (202) corresponds to a user - defined domain or dimension, rather than time. In a specific example, the user - defined domain is the dynamic range of the content, where I0 represents the 100 - nit SDR version of the image, and I T-1represents the 4000 nits HDR version of the same image. In another specific example, I0 represents the 4000 nits HDR version of the image, and I T-1 represents the 100 nits SDR version of the same image. In this document, the acronyms SDR and HDR represent Standard Dynamic Range and High Dynamic Range respectively. Additional examples of the sequence (202) and various user-defined domains are described in the subsection titled "Example Use Cases" below.

[0037] One image in the sequence (202) is selected to be used as the base image (204), which is represented as I b . The encoder (200) operates to encode the base image (204) into a base layer bitstream (294), which is transmitted, for example, to the decoder (300). The encoder (200) further operates to calculate a residual image sequence (208) by subtracting (206) the base image (204) from each image in the sequence (202). A corresponding forward reshaped sequence (212) is generated by applying forward reshaping (210) to the sequence (208). The relevant parameters of the forward reshaping (210) are transmitted in the form of a metadata bitstream (296). The forward reshaped sequence (212) is used to train a suitable neural network (e.g., MLP (100)) in the neural field encoding block (214) to represent the sequence (212). Finally, the parameters of the trained neural network are compressed (216) and transmitted in the form of a neural network bitstream (298). Example operations performed by the decoder (200) according to various embodiments are described in more detail below. In some specific examples, the elements (204, 206, 208, 294) are absent or not used. In such examples, the forward reshaping (210) is directly applied to the sequence (202).

[0038] Residual image

[0039] The individual residual images of the sequence (208) are generated by subtracting (206) the base image (204) from each image in the sequence (202) (I0,..., I T-1 ), as follows:

[0040]

[0041] Thus, the pixel values of the residual image have signs, i.e., they can be positive or negative. In some examples, a rescaling operation is performed on the sequence (208), for example, when the MLP (100) has a sigmoid layer configured to operate according to Equation (3). In some examples, the rescaling operation is performed according to Equation (6):

[0042]

[0043] wherein, indicates the rescaled residual image. Thereafter, each pixel located at is represented as

[0044] reshaping

[0045] As already mentioned above, neural networks (such as MLP (100)) can tend to preferentially learn low-frequency functions. In various examples, the above-described positional encoding (e.g., see Equation (4)) can be used to reduce or eliminate such bias. One result of the positional encoding is an increase in the number of inputs to the input layer of the neural network (such as input layer (1101)). In some cases, this increase can be quite significant, which may translate into a corresponding increase in the size of the neural network. In at least some examples, forward reshaping (210) is used to mitigate this increase. More specifically, the forward reshaping (210) operation is used to change the signal frequency distribution such that the weights of the high-frequency components are reduced, thereby enabling a corresponding reduction in the number of triangular components of γ (also see Equation (4)). Various embodiments and components of the forward reshaping (210) are described in more detail below. In some specific examples, the forward reshaping (210) is optional and can be omitted.

[0046] Region of Interest Selection in the Spatial Domain

[0047] Generally, in the sequence (208), the active content is restricted to a region smaller than the full-size residual image. In other words, the residual image values outside this smaller region are constant, e.g., 0 (before rescaling) or 0.5 (after rescaling). To reduce the size of the image on which the neural network operates and avoid data bias for non-active pixels, a region of interest (ROI) is selected in at least some examples. Then, only the selected ROI is fed to the neural network for training in the neural network encoding block (214). Thus, in such examples, the corresponding neural network bitstream (298) only contains the coefficients and biases of the ROI. At the decoder (300), the ROI is reconstructed according to the ROI position transmitted to the decoder (300) via the metadata bitstream (296) and added to the base image.

[0048] In some examples, the boundaries of the ROI are determined as the tightest (e.g., the tightest) bounding box enclosing the ROI such that all pixels outside the ROI are non-active, e.g., having a value of 0.5 within the rescaled range. Equations (7) to (10) represent the Cartesian coordinates x0, x1, y0, y1 of the four boundary lines of the bounding box according to one possible example:

[0049]

[0050] In some specific examples, a small margin δ is added to expand the ROI, so as to make a clearer transition in the decoded image from the active region ROI encoded by the neural field to the inactive region from the base image (204). Equations (11) to (14) provide the corresponding mathematical formulas for redefining the ROI selection in the δ-margin example:

[0051] ROI * x0 = ROI * x0 - δ (11)

[0052] ROI * x1 = ROI * x1 + δ (12)

[0053] ROI * y0 = ROI * y0 - δ (13)

[0054] ROI * y1 = ROI * y1 + δ (14)

[0055] wherein, the ROI coordinates on the right side of Equations (11) to (14) are calculated using Equations (7) to (10) respectively. The two corners of the ROI determined in the manner indicated above (i.e., (ROI * x0 , ROI * y0 ) and (ROI * x1 , ROI * y1 )) define the rectangular ROI region and are transmitted in the form of the metadata bitstream (296). Thereafter, R t (x, y) is used to indicate the ROI image. H * and W * respectively indicate the height and width of the ROI.

[0056] Frame of Interest (FOI) Selection for User-Defined Axes

[0057] In some examples, similar to the ROI selection in the above spatial domain, individual frames with active residuals in the sequence (208) are selected in the user-defined domain or dimension. Equations (15) to (16) provide the formal definitions of the FOI that can be used for such examples:

[0058]

[0059] Two frame indices FOI determined in the manner indicated above * f0 and FOI * f1 define the FOI and are transmitted in the form of a metadata bitstream (296).

[0060] Spatial Domain Remodeling

[0061] In each relevant dimension (e.g., the x and y dimensions of a 2D image), the corresponding spatial axis is divided into M bins, where M is an integer greater than one. In some examples, M may be equal to the number of pixels in the corresponding dimension. Let us represent the number of pixels in the active region and the active frame as P. For a 2D image, the forward reshaping (210) operation is used to calculate the average local standard deviation (SD) of each bin (or overlapping bins) along the corresponding (x or y) dimension. Thereafter, the x SD and y SD calculated in this way are represented as σ h (x) and σ v (y). An example calculation corresponding to the x dimension is described below. Based on the provided example, a person of ordinary skill in the relevant art will easily understand how to implement the calculation corresponding to the y dimension without any undue experimentation.

[0062] In various examples, the calculation of the SD depends on a sliding window, and the size of the sliding window is represented as Δ. The lower limit x L and upper limit x H are expressed as follows:

[0063] x H = min(x + Δ, W * - 1) (17)

[0064] x L = max(x - Δ, 0)Δ, 0) (18)

[0065] Then, for each (x, y), the standard deviation is calculated as follows:

[0066]

[0067] where is the mean between [x L , x H . Then, for each x, the average SD of all y values in the corresponding range is calculated as follows:

[0068]

[0069] In a typical example, σ h (x) is non-uniform, i.e., it varies with x.

[0070] In an exemplary embodiment, the forward remapping (210) is configured to apply a mapping function f h (), which maps a uniform grid {x} to a non-uniform grid {x'}, such that the local SD in the new grid {x'} is more uniform than in the uniform grid {x}. Mathematically, the mapping function f h () satisfies the following criterion:

[0071] f h (σ h (x)) ≈ c h (21)

[0072] where is the sum of σ h (x) over the ROI x range; and is a constant.

[0073] In some examples, cumulative distribution function (CDF) matching is used to construct the mapping function f h (), as follows. First, the local SD is calculated:

[0074]

[0075] Since the target local SD is constant, the corresponding target function T h (x') is a linear function with a slope of 1 / (W * -1) and an offset of 0:

[0076]

[0077] For each point x, the mapping function f h () maps the point x to the point x' such that

[0078] T h (x') = C h (x) (24)

[0079] Based on equation (24), the forward remapping (210) operation is used to find the integer k such that T h (k - 1) ≤ C h (x) ≤ T h (k). Then, x' is found by interpolation based on two adjacent points, as follows:

[0080]

[0081] Then, the mapping function f h(x), where in each step the method represented by equations (21) to (25) is used. In some cases, the following example smoothing steps are used to apply additional smoothing to the mapping function:

[0082] Step 1: Copy

[0083] Step 2: Recalculate f h (x) as

[0084] Step 3: Clip the recalculated f h (x) to the valid range, i.e., f h (x) = clip3(f h (x), 0, 1) (28)

[0085] In this article, the function y = clip3(x,a,b) performs the following operations:

[0086] If x < a, then y = a;

[0087] Else If x > b, then y = b;

[0088] Else y = x;

[0089] End.

[0090] In various examples, the above operations are used to implement forward reshape (210) in the encoder (200) to achieve the following result: Given a uniform grid {x}, the mapping function f h (x) is applied to generate the corresponding non-uniform grid {x'}. In the neural field encoding (214), the non-uniform grid {x'} is used as the input for the above position encoding. The metadata of the mapping function f h (x) is transmitted to the decoder (300) in the form of a metadata bitstream (296). At the decoder (300), the mapping function metadata is extracted from the metadata bitstream (296), which enables the decoder (300) to reconstruct the mapping function f h (x). Thus, for each x of the uniform grid {x}, the decoder (300) can determine the corresponding x' of the non-uniform grid {x'}. In a typical example, the decoder (300) uses the x' determined in this way to query the MLP (100) to obtain the corresponding red, green, blue (RGB) pixel values. Then, the decoder (300) assigns the RGB value to the pixel x of the uniform grid {x}.

[0091] In various examples, similar operations are performed to implement reshape of the y-axis dimension in the forward reshape block (210). The corresponding mapping function f v() is constructed to satisfy equation (29):

[0092]

[0093] The reshaping function f v ()'s metadata is transmitted to the decoder (300) in the form of a metadata bitstream (296).

[0094] Remodeling of User-Defined Axes

[0095] In some examples, the above spatial reshaping algorithm is adapted to apply reshaping to user-defined dimensions. In a specific example, the encoder (200) is configured to measure SDσ along the time dimension t t (Δ), and apply substantially the same calculation to determine the reshaping function f in the time domain t (), where the function f t () satisfies equation (30):

[0096] f t (σ t (Δ))≈c t (30)

[0097] where c t is a constant. Similar to the case of the above spatial reshaping, the metadata of the reshaping function f t () is transmitted to the decoder (300) in the form of a metadata bitstream (296).

[0098] Codeword Domain Remodeling

[0099] In some examples, the above spatial reshaping algorithm is adapted to apply reshaping in the codeword domain. In such an example, the encoder (200) is configured to calculate the histogram σ in the codeword domain c (p). Then the corresponding reshaping function f is constructed using the above method c (), such that:

[0100] f c (σ c (p))≈c c (31)

[0101] where c c is a constant. Similarly, the metadata of the reshaping function f c () is transmitted to the decoder (300) in the form of a metadata bitstream (296).

[0102] In some examples, the forward reshaping (210) includes based on the above mapping functions (f h (), f v (), f t (), fc ()) Two or more forward reshape operations are performed. Such multiple reshape operations can be performed sequentially, for example. The order in which these reshape operations are applied to the sequence (208) is transmitted to the decoder (300) in the form of a metadata bitstream (296). Based on the transmitted order, the decoder (300) operates to apply the corresponding backward reshape operations in the reverse order. Note that changing the order of application of the individual mapping functions generally changes the final result. Therefore, the decoder (300) generally needs to follow the reverse order to obtain the best image reconstruction result.

[0103] Neural field encoding

[0104] As already pointed out above, positional encoding (see Equation (4)) is used to generate the input to the MLP (100) in the neural field encoding block (214). To control the MLP overhead, in at least some embodiments, for performance optimization purposes, it may be desirable to constrain the total number 2L of the triangular components of γ. The corresponding optimization problem can be stated as follows: How to select the values of the positional encoding to optimize the quality of the image computed by the decoder (300)? In this formula, does not need to be equal to k, and the value does not need to be consecutive integers.

[0105] Theoretically, the above optimization problem is a combinatorial problem with NP-hardness. Therefore, for relatively large L, finding the optimal solution may take a relatively long (e.g., impractical) time. To accelerate the optimization, in various examples, the neural field encoding block (214) is configured to use a greedy sub-optimal iterative algorithm. At the start of the algorithm, a frequency candidate list Ψ is initialized with L max different frequencies, where L max > L. The selected frequency list Ω is an empty set at initialization. Then, in each iteration of the iterative algorithm, for each frequency in the set Ψ, the algorithm operates to include the frequency into the frequency list and measure the corresponding distortion. For all frequencies in the set Ψ, the algorithm selects the frequency l * corresponding to the approximate minimum distortion. The value l * is placed into the selected frequency list Ω and removed from the frequency candidate list Ψ. The selection procedure is repeated until L frequencies are selected. The following pseudocode illustrates an example of the iterative frequency selection algorithm.

[0106]

[0107] Herein, the function D() is a function for calculating the distortion. The final frequency list Ω is transmitted to the decoder (300) in the form of a metadata bitstream (296).

[0108] Neural network compression

[0109] After the MLP (100) is trained in the neural field encoding block (214), the neural network compression (216) operation is used to compress various neural network parameters to generate a neural network bitstream (298). In different embodiments of the neural network compression (216), different suitable compression methods may be used. In a specific non-limiting example, the neural network compression (216) operates according to the standard of neural network compression and representation (NNCR) or Part 17 of the ISO / IEC 15938 standard, which is incorporated herein by reference in its entirety. The NNCR standard is promulgated by the ISO / IEC Moving Picture Experts Group (MPEG) and is specifically for the efficient compression and transmission of neural networks such as the MLP (100).

[0110] Residual decoding

[0111] Figure 3is a block diagram illustrating an electronic decoder (300) according to an embodiment. The corresponding electronic encoder (200) has been described above. The decoder (300) generally performs inverse encoding operations in the reverse order. The input to the decoder (300) is provided by the above-mentioned bitstreams (294, 296, 298). The decoder (300) applies neural network decompression (316) to the neural network bitstream (298) to obtain various neural network parameters. Then, the obtained neural network parameters are used to reconstruct the MLP (100) in the trained configuration previously calculated at the encoder (200). The decoder (300) uses the reconstructed MLP (100) and the relevant part of the metadata bitstream (296) in the neural field decoding block (314). This metadata part is used for spatial reshaping such that the reshaped coordinates are then used to query the reconstructed MLP (100). The user-specified input (315) provides one or more parameters such as time, luminance range, etc. Based on these parameters, the reconstructed MLP (100) operates in the test mode to render the corresponding reshaped residuals (312). The decoder (300) further operates to apply backward reshaping (310) to the reshaped residuals (312) to generate the corresponding residual image (308). Based on the metadata bitstream (296), the backward reshaping (310) is configured to perform a reshaping operation opposite to the reshaping operation of the forward reshaping (210). Summation (306) is used to combine the base image (204) received via the base layer bitstream (294) with the residual image (308) to generate the corresponding output image (302). Depending on the user-specified input (315), the image (302) can be one of the target (reference) images in the sequence (202) or a novel view generated using the reconstructed MLP (100). In the following description, several representative use cases of the encoder (200) and the decoder (300) are described to specifically illustrate several examples of user-defined dimensions and the corresponding user-specified input (315).

[0112] Example use case

[0113] Motion Blur at Different Time Scales

[0114] Figure 4 is a block diagram illustrating a part of motion blur image processing using the encoder (200) according to an embodiment. More specifically, Figure 4 illustrates the preprocessing (400) for generating the sequence (202) for the encoder (200). The preprocessing (400) uses a sequence of still images (4020, 4021,..., 402 K-1 ) to calculate the target motion blur (MB) images (2021, 2022,..., 202 K-1) sequence (202). In a representative example, the MB image at time dt is obtained by taking a still image (4020, 4021, ..., 402) from time 0 to time dt. K-1 ) is generated by taking the average value as follows:

[0115]

[0116] in, is the corresponding MB picture; and I k Is a still image (402 k ). When dt = 0, the MB image (2021) is the same as the first still image (4020), that is, The first still picture (4020) is also designated as the base picture (204). The encoder (200) is further operative to generate a base picture by extracting from each target MB picture (2021, 2022, ..., 202 K-1 ) by subtracting (206) the base image (204) from the base image to calculate a residual image sequence (208). The residual image sequence (208) thus calculated is then processed by the encoder (200), as described above with reference to Figure 2 After training using the sequence (208), the corresponding MLP (100) can be operated to predict the pixel value based on the pixel coordinate (x, y) and further based on the time dt. In this example, time dt is used as a user-specified variable input (315) at the decoder (300).

[0117] The value of dt may be an integer or a non-integer (e.g., a fraction) depending on the user specified input (315). When the value of dt is an integer value from the original value set, the decoder (300) operates the MLP (100) to generate the same MB images as the reference MB images (2021, 2022, ..., 202 K-1 ). When the value of dt is a fractional value from a range corresponding to the original set of values, the decoder (300) operates the MLP (100) to generate a reconstructed image (302) representing the corresponding new MB image.

[0118] Different Dynamic Ranges

[0119] In this particular use case, the parameter dt represents the dynamic range. In a particular example, the sequence (202) includes target images of the same scene with dynamic ranges of 100 nits, 162 nits, 260 nits, 413 nits, 652 nits, 1026 nits, 1612 nits, 2536 nits, and 4000 nits, respectively. In another particular example, a descending order of the dynamic range can also be used in the sequence (202). The sequence (208) of residual images is calculated using the sequence (202) and the base image (204), as described above with reference to Figure 2 as described. In various examples, the base image (204) is a target image with a dynamic range of 100 nits or a target image with a dynamic range of 4000 nits. After training using the sequence (208), the corresponding MLP (100) can operate at the decoder (300) to calculate an image with a different dynamic range based on a user-specified variable input (315). For example, the user-specified input (315) can be selected to have any dt value between 100 nits and 4000 nits. When the value of dt is a value selected from the original set of dynamic range values (i.e., one of 100 nits, 162 nits, 260 nits, 413 nits, 652 nits, 1026 nits, 1612 nits, 2536 nits, and 4000 nits), the decoder (300) operates the MLP (100) to generate a reconstructed image (302) corresponding to one of the reference images (202) used at the encoder (200). When the value of dt is not a value from the original set of dynamic range values, the decoder (300) operates the MLP (100) to generate a reconstructed image (302) with a corresponding new dynamic range according to the specified dt value.

[0120] Faces of Different Ages

[0121] In this particular use case, the parameter dt represents the age of the person in the target image (202). In a particular example, the sequence (202) includes 23 images of the same person with ages between 6 years and 60 years. The sequence (208) of residual images is calculated using the sequence (202) and the base image (204), as described above with reference to Figure 2As described. In one example, the base image (204) is a target image corresponding to the age of 6 years old. In another example, the base image (204) is a target image corresponding to the age of 60 years old. After training using the sequence (208), the corresponding MLP (100) can operate at the decoder (300) to calculate images corresponding to different ages based on the user-specified variable input (315). When the value of dt is a value selected from the original age set, the decoder (300) operates the MLP (100) to generate a reconstructed image (302) corresponding to one of the reference images (202) used at the encoder (200). When the value of dt is not a value from the original age set, the decoder (300) operates the MLP (100) to generate a reconstructed image (302) corresponding to the new age according to the specified value of dt.

[0122] Example Hardware

[0123] Figure 5 is a block diagram illustrating a computing device (500) according to an embodiment. The device (500) can be used, for example, to implement the encoder (200) or the decoder (300). The computing device (500) includes an input / output (I / O) device (510), an image processing engine (520), and a memory (530). The I / O device (510) can be used to enable the device (500) to receive various input signals (502) and output various output signals (504). For example, when the computing device (500) implements the encoder (200), the input signal (502) includes the sequence (202), and the output signal (504) includes bitstreams (294, 296, 298). When the computing device (500) implements the decoder (300), the input signal (502) includes bitstreams (294, 296, 298), and the output signal (504) includes the (multiple) reconstructed images (302).

[0124] The memory (530) may have a buffer for receiving image data and / or other related data. The image data may be in the form of, for example, one or more image files. Once the data is received, the memory (530) may provide a portion of the data to the image processing engine (520) for processing therein. The image processing engine (520) includes a processor (522) and a memory (524). The memory (524) may store program code therein, which when executed by the processor (522) enables the image processing engine (520) to perform image processing, including but not limited to image processing according to some or all of the above process flows (200, 300, 400). The program code may particularly include program code for simulating various neural networks (e.g., the above MLP (100)). Once the image processing engine (520) generates various above images by executing the corresponding portions of the code, the image processing engine (520) may perform its rendering process and provide the corresponding (multiple) visual images for viewing on a display. The visual images may be in the form of, for example, suitable image files output by the I / O device (510).

[0125] According to an example embodiment disclosed above, for example in the Summary of the Invention section and / or with reference to Figures 1 to 5 any one or any combination of some or all of those in, there is provided an apparatus for residual coding, the apparatus comprising: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to, together with the at least one processor, cause the apparatus to at least perform the following operations: accessing a plurality of residual images or reshaped images, each of the residual images representing a difference between a corresponding reference image and a base image, different corresponding reference images corresponding to respective values of one or more first parameters; training a neural field network to represent the plurality of residual images or reshaped images, the neural field network responsive to variable inputs of the one or more first parameters and characterized by a set of second parameters, the training resulting in fixed values of the second parameters; and transmitting the base image and the fixed values of the second parameters to an electronic decoder.

[0126] According to another example embodiment disclosed above, for example in the Summary of the Invention section and / or with reference to Figures 1 to 5Any one or any combination of some or all of the following provides an apparatus for residual decoding, the apparatus comprising: at least one processor; and at least one memory including program code; wherein the at least one memory and the program code are configured to, together with the at least one processor, cause the apparatus to at least perform the following operations: access a base image and a set of fixed parameter values of a neural field network trained to represent a plurality of residual images, the neural field network responsive to a variable input of one or more first parameters and characterized by a set of second parameters, the fixed parameter values being the values of the second parameters; use a selected value of the variable input to test the neural field network to generate an output residual image; and calculate a reconstructed image by summing the output residual image and the base image.

[0127] According to yet another example embodiment disclosed above, for example in the Summary of the Invention section and / or with reference to Figures 1 to 5 Any one or any combination of some or all of the following provides a method for residual encoding using a neural field, the method comprising: accessing, via a processor, a plurality of residual images or reshaped images, each of the residual images representing a difference between a corresponding reference image and a base image, different corresponding reference images corresponding to respective values of one or more first parameters; training, via the processor, a neural field network to represent the plurality of residual images or reshaped images, the neural field network responsive to a variable input of the one or more first parameters and characterized by a set of second parameters, the training resulting in fixed values of the second parameters; and transmitting the base image and the fixed values of the second parameters to an electronic decoder.

[0128] In some embodiments of the above method, the one or more first parameters are selected from the group consisting of time, dynamic range, and age.

[0129] In some embodiments of any of the above methods, the neural field network includes a multi-layer perceptron responsive to an input of coordinates or other field-type variables; and wherein the set of second parameters includes weights of connections between respective neural network nodes of the multi-layer perceptron and biases applied to signals by respective neural network nodes of the multi-layer perceptron.

[0130] In some embodiments of any of the above methods, the method further comprises: mapping, via the processor, a first quantity of inputs to a second quantity of inputs using a series of trigonometric functions, the second quantity being greater than the first quantity, the mapping being performed using a greedy algorithm; and applying, via the processor, the second quantity of inputs to the neural field network.

[0131] In some embodiments of any of the above methods, the method further comprises: reshaping, via the processor, the plurality of residual images to generate the first quantity of inputs, the reshaping being configured to reduce the relative weight of high-frequency signal components or to limit the size of the neural field network.

[0132] In some embodiments of any of the above methods, the method further comprises: selecting, via the processor, a region of interest in the residual image, the region of interest being smaller than a full image frame.

[0133] In some embodiments of any of the above methods, the reshaping comprises at least one of spatial domain reshaping, temporal domain reshaping, and codeword domain reshaping.

[0134] In some embodiments of any of the above methods, the method further comprises transmitting one or more parameters of the reshaping to the electronic decoder.

[0135] In some embodiments of any of the above methods, transmitting a fixed value of the second parameter to the electronic decoder comprises: generating a compressed data stream by applying compression to the fixed value of the second parameter via the processor; and transmitting the compressed data stream to the electronic decoder.

[0136] In some embodiments of any of the above methods, the neural field network is operable when a value specified by a variable input of the one or more first parameters is different from any of the respective values of the one or more first parameters.

[0137] According to another example embodiment disclosed above, for example in the Summary of the Invention section and / or in any one or any combination of some or all of Figures 1 to 5 there is provided a method for residual decoding using a neural field, the method comprising: accessing, via a processor, a base image and a set of fixed parameter values of a neural field network trained to represent a plurality of residual images, the neural field network responsive to a variable input of one or more first parameters and characterized by a set of second parameters, the fixed parameter values being values of the second parameters; testing, via the processor, the neural field network using a selected value of the variable input to generate an output residual image; and calculating, via the processor, a reconstructed image by combining the output residual image and the base image.

[0138] In some embodiments of the above method, the one or more first parameters are selected from the group consisting of time, dynamic range, and age.

[0139] In some embodiments of the above method, the neural field network includes a multi-layer perceptron; and wherein, the set of second parameters includes the weights of the connections between the respective neural network nodes of the multi-layer perceptron and the biases applied by the respective neural network nodes of the multi-layer perceptron to the signals.

[0140] In some embodiments of any of the above methods, the method further includes: receiving, from an electronic encoder, one or more reshaping parameters for reshaping, the reshaping being configured to increase the relative weight of high-frequency signal components; and applying, via the processor, the reshaping to the output of the neural field network to generate the output residual image.

[0141] In some embodiments of any of the above methods, the method further includes: selecting, via the processor, a region of interest of the output residual image, the region of interest being smaller than a full image frame.

[0142] In some embodiments of the above method, the reshaping includes at least one of spatial domain reshaping, time domain reshaping, and codeword domain reshaping.

[0143] In some embodiments of the above method, accessing the set of fixed parameter values includes: receiving, from an electronic encoder, a compressed data stream; and obtaining the fixed values of the second parameters by applying decompression to the compressed data stream via the processor.

[0144] In some embodiments of the above method, the neural field network is operable when the value specified by the variable input of the first parameter is different from any value of the first parameter used to train the neural field network.

[0145] Regarding the processes, systems, methods, heuristics, etc. described herein, it should be understood that although the steps of these processes, etc. have been described as being performed in a specific ordered sequence, these processes can be practiced using the described steps executed in an order different from the order described herein. Further, it should be understood that certain steps can be performed simultaneously, other steps can be added, or certain steps described herein can be omitted. In other words, the process descriptions herein are provided for the purpose of illustrating certain embodiments and should in no way be construed as limiting the claims.

[0146] Accordingly, it should be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided will be apparent to those of ordinary skill in the art upon reading the above description. The scope should not be determined with reference to the above description, but should be determined with reference to the appended claims and the full scope of equivalents to which those claims are entitled. It is expected and intended that the technology discussed herein will be developed in the future, and the disclosed systems and methods will be incorporated into such future embodiments. In summary, it should be understood that this application is capable of modification and change.

[0147] All terms used in the claims are intended to be given the broadest reasonable interpretation and ordinary meaning as understood by those who are knowledgeable in the technology described herein, unless an explicit contrary indication appears herein. In particular, the use of singular articles such as "a," "the," and "said" should be understood to recite one or more of the indicated elements unless the claim recites an explicit contrary limitation.

[0148] A summary of the disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. This summary is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. The method of the disclosure should not be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as reflected in the appended claims, the inventive subject matter lies in less than all of the features of a single disclosed embodiment. Thus, the appended claims are hereby incorporated into the detailed description, with each claim standing on its own as a separately claimed subject matter.

[0149] Although this disclosure includes references to illustrative embodiments, the specification is not to be construed as limiting. Various modifications to the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to those skilled in the art to which this disclosure pertains, are considered to be within the principles and scope of the disclosure, for example, as expressed in the claims.

[0150] Some embodiments may be implemented as circuit-based processes, including possible implementations on a single integrated circuit.

[0151] Some embodiments may be embodied in the form of methods and devices for practicing these methods. Some embodiments may also be embodied in the form of program code recorded in a tangible medium, such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy disk, a CD-ROM, a hard disk drive, or any other non-transitory machine-readable storage medium, wherein when the program code is loaded into a machine (such as a computer, etc.) and executed by it, the machine becomes a device for practicing the various embodiments described herein. Some embodiments may also be embodied in the form of program code, such as stored in a non-transitory machine-readable storage medium (including being loaded into a machine and / or executed by a machine), wherein when the program code is loaded into a machine (such as a computer or a processor, etc.) and executed by it, the machine becomes a device for practicing the various embodiments described herein. When implemented on a general-purpose processor, the program code segment is combined with the processor to provide a unique device that operates similarly to a specific logic circuit.

[0152] Unless expressly stated otherwise, each numerical value and range should be interpreted as being approximate as if the value or range were preceded by the word "about" or "approximately."

[0153] The use of figure numbers and / or reference numerals in the claims is intended to identify one or more possible embodiments of the claimed subject matter to facilitate the interpretation of the claims. Such use should not be construed as necessarily limiting the scope of these claims to the embodiments shown in the corresponding figures.

[0154] Although elements in the method claims (if any) are recited in a specific order with corresponding labels, these elements are not necessarily intended to be limited to being implemented in the specific order unless the claim recitation otherwise implies a specific order for implementing some or all of these elements.

[0155] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present disclosure. In this specification, the various appearances of the phrase "in one embodiment" are not necessarily all referring to the same embodiment, nor are separate or additional embodiments mutually exclusive of other embodiments. The same applies to the term "implementation".

[0156] Unless otherwise specified herein, the use of ordinal adjectives "first," "second," "third," etc. to refer to an object among multiple similar objects merely indicates that different instances of such similar objects are being referred to, and is not intended to imply that the similar objects so referred to must be in a corresponding order or sequence, whether temporally, spatially, in ranking, or in any other manner.

[0157] Unless otherwise specified herein, in addition to its plain meaning, the conjunction "if" may also or alternatively be construed to mean "when" or "while" or "in response to determining" or "in response to detecting", the interpretation of which may depend on the corresponding specific context. For example, the phrase "if it is determined that..." or "if [stated condition] is detected" may be construed to mean "after determining..." or "in response to determining..." or "after detecting [stated condition or event]" or "in response to detecting [stated condition or event]".

[0158] Also for the purposes of this description, the terms "coupled", "coupling", "being coupled", "connected", "linking" or "being connected" refer to any manner known in the art or later developed that allows energy to be transferred between two or more elements, and the insertion of one or more additional elements is contemplated, although this is not required. In contrast, the terms "directly coupled", "directly connected", etc. imply the absence of such additional elements.

[0159] The functions of the various elements shown in the figures, including any functional blocks labeled "processor" and / or "controller", can be provided by using dedicated hardware as well as hardware capable of executing software associated with appropriate software. When provided by a processor, the functions can be provided by a single dedicated processor, by a single shared processor, or by multiple individual processors (some of which may be shared). Additionally, the explicit use of the term "processor" or "controller" should not be construed as exclusively referring to hardware capable of executing software, and may implicitly include, but is not limited to, digital signal processor (DSP) hardware, network processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), read only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage devices. Other conventional and / or custom hardware may also be included. Similarly, any switches shown in the figures are merely conceptual. Their functions can be performed by the operation of program logic, by dedicated logic, by the interaction of program control and dedicated logic, or even manually, and the specific technique can be selected by the implementer, as more specifically understood from the context.

[0160] As used in this application, the terms "circuit" and "circuitry" can refer to one or more or all of the following: (a) a pure hardware circuit implementation (such as an implementation in a pure analog circuit and / or digital circuit); (b) a combination of hardware circuitry and software, such as (if applicable): (i) a combination of analog hardware circuitry and / or digital hardware circuitry with software / firmware, and (ii) any part of a hardware processor with software (including (a) digital signal processor(s)), software, and (a) memory(ies), which together cause a device such as a mobile phone or a server to perform various functions); and (c) (a) hardware circuitry(ies) and / or (a) processor(s), such as (a) microprocessor(s) or a part of (a) microprocessor(s), which require software (e.g., firmware) to operate but may not have software when not in operation. This definition of circuitry applies to all uses of the term in this application (including all claims). As a further example, as used in this application, the term "circuit" also covers an implementation of only one hardware circuit or one processor (or processors) or a part of a hardware circuit or processor and its accompanying software and / or firmware. The term "circuit" also covers, for example, a baseband integrated circuit or a processor integrated circuit for a mobile device, or a similar integrated circuit in a server, a cellular network device, or other computing or networking device, when applicable to a particular claim element.

[0161] Those of ordinary skill in the art will recognize that any block diagrams herein represent a conceptual view of illustrative circuitry embodying the principles of the present disclosure. Similarly, it will be recognized that any flowchart, flow diagram, state transition diagram, pseudocode, etc. represent various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or a processor, whether or not the computer or processor is explicitly shown.

[0162] The "Summary of the Invention" in this specification is intended to introduce some example embodiments, and additional embodiments are described in the "Detailed Description" and / or with reference to one or more of the drawings. The "Summary of the Invention" is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

1. A method for residual encoding using a neural field, the method comprising: Accessing, via a processor, a plurality of residual images or reshaped images, each of the residual images representing a difference between a corresponding reference image and a base image, different corresponding reference images corresponding to respective values of one or more first parameters; Training, via the processor, a neural field network to represent the plurality of residual images or reshaped images, the neural field network responsive to a variable input of the one or more first parameters and characterized by a set of second parameters, the training yielding fixed values of the second parameters; And Transmitting the base image and the fixed values of the second parameters to an electronic decoder.

2. The method according to claim 1, wherein The one or more first parameters are selected from the group consisting of time, dynamic range, and age.

3. The method according to claim 1 or 2, Among them, The neural field network includes a multi-layer perceptron responsive to an input of coordinates or other field-type variables; and Wherein the set of second parameters includes weights of connections between respective neural network nodes of the multi-layer perceptron and biases applied to signals by respective neural network nodes of the multi-layer perceptron.

4. The method according to any one of the preceding claims, further comprising: Mapping, via the processor, a first quantity of inputs to a second quantity of inputs using a series of trigonometric functions, the second quantity being greater than the first quantity, the mapping being performed using a greedy algorithm; And Applying, via the processor, the second quantity of inputs to the neural field network.

5. The method according to claim 4, further comprising reshaping, via the processor, the plurality of residual images to generate the first quantity of inputs, the reshaping being configured to reduce a relative weight of high-frequency signal components or to limit a size of the neural field network.

6. The method according to claim 5, further comprising selecting, via the processor, a region of interest in the residual images, the region of interest being smaller than a full image frame.

7. The method according to claim 5 or 6, wherein, The reshaping includes at least one of spatial-domain reshaping, time-domain reshaping, and codeword-domain reshaping.

8. The method according to any one of claims 5 to 7, further comprising transmitting one or more parameters of the reshaping to the electronic decoder.

9. The method according to any one of the preceding claims, wherein Transmitting the fixed values of the second parameters to the electronic decoder includes: Generating a compressed data stream by applying compression, via the processor, to the fixed values of the second parameters; and Transmitting the compressed data stream to the electronic decoder.

10. The method according to any one of the preceding claims, wherein, The neural field network is operative when a variable input of the one or more first parameters specifies a value different from any of the respective values of the one or more first parameters.

11. A method for residual decoding using a neural field, the method comprising: Accessing, via a processor, a base image and a set of fixed parameter values of a neural field network trained to represent a plurality of residual images, the neural field network responsive to a variable input of one or more first parameters and characterized by a set of second parameters, the fixed parameter values being values of the second parameters; Testing the neural field network using a selected value of the variable input via the processor to generate an output residual image; and Calculating a reconstructed image via the processor by combining the output residual image and the base image.

12. The method according to claim 11, wherein, The one or more first parameters are selected from the group consisting of time, dynamic range, and age.

13. The method according to claim 11 or 12, Among them, wherein the neural field network includes a multi-layer perceptron; and wherein the set of second parameters includes weights of connections between respective neural network nodes of the multi-layer perceptron and biases applied by respective neural network nodes of the multi-layer perceptron to signals.

14. The method according to any one of claims 11 to 13, further comprising: Receiving one or more reshaping parameters from an electronic encoder for reshaping, the reshaping being configured to increase a relative weight of high-frequency signal components; and Applying the reshaping to an output of the neural field network via the processor to generate the output residual image.

15. The method according to claim 14, further comprising selecting a region of interest of the output residual image via the processor, the region of interest being smaller than a full image frame.

16. The method according to claim 14 or 15, wherein The reshaping includes at least one of spatial domain reshaping, time domain reshaping, and codeword domain reshaping.

17. The method according to any one of claims 11 to 16, wherein, Accessing the set of fixed parameter values includes: Receiving a compressed data stream from an electronic encoder; and Obtaining the fixed values of the second parameters by applying decompression to the compressed data stream via the processor.

18. The method according to any one of claims 11 to 17, wherein The neural field network is operative when a variable input of the one or more first parameters specifies a value different from any value of the one or more first parameters used to train the neural field network.

19. An image processing apparatus for residual encoding, the apparatus comprising: At least one processor; and At least one memory including program code; wherein the at least one memory and the program code are configured to, using the at least one processor, cause the apparatus to at least perform the following operations: Accessing a plurality of residual images or reshaped images, each of the residual images representing a difference between a respective reference image and a base image, different respective reference images corresponding to respective values of one or more first parameters; Training a neural field network to represent the plurality of residual images or reshaped images, the neural field network responsive to a variable input of the one or more first parameters and characterized by a set of second parameters, the training resulting in fixed values of the second parameters; and Transmitting the base image and the fixed values of the second parameters to an electronic decoder.

20. An image processing apparatus for residual decoding, the apparatus comprising: At least one processor; and At least one memory including program code; wherein the at least one memory and the program code are configured to, using the at least one processor, cause the apparatus to at least perform the following operations: Access a base image and a set of fixed parameter values of a neural field network trained to represent multiple residual images, the neural field network responsive to variable inputs of one or more first parameters and characterized by a set of second parameters, the fixed parameter values being values of the second parameters; Test the neural field network using selected values of the variable inputs to generate an output residual image; and Compute a reconstructed image by summing the output residual image and the base image.