Signaling for implicit neural representation reconstruction

By carrying signal notification elements in the bitstream, the problem that INR network decoders cannot reconstruct images or videos is solved, achieving efficient signal reconstruction and meeting the requirements for high compression efficiency.

CN121359461APending Publication Date: 2026-01-16INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480039292.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-13
Filing Date
2024-06-10
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing image and video decoders cannot effectively reconstruct the output signal based solely on implicit neural representations (INRs), requiring additional information for the decoder to reconstruct the image or video.

Method used

By carrying signal notification (SEI syntax element) in the bitstream, information such as the type of INR network, input transformation, coordinate normalization, maximum upsampling size, and output format is provided so that the decoder can correctly reconstruct the signal.

Benefits of technology

It enables the efficient reconstruction of signals from images, videos, or 3D scenes using an INR network, improving the reconstruction capability of the decoder and meeting the requirements for high compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121359461A_ABST
    Figure CN121359461A_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding or decoding one or more syntax elements for using an INR decoder is provided. The one or more syntax elements are used to reconstruct at least a portion of the image using at least one implicit neural representation network representing the at least a portion of the image. The one or more syntax elements include at least one of an indication related to the implicit neural representation network, information for constructing implicit neural representation network inputs, and information related to implicit neural representation network outputs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to European application No. EP 23 30 59 37.7, filed on June 13, 2023, the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present embodiments generally relate to image, video and / or 3D scene compression using implicit neural representations (INR). The present embodiments relate to a method and apparatus for encoding, decoding, transmitting metadata used by a decoder to reconstruct a signal encoded using an INR network. BACKGROUND

[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to exploit spatial and temporal redundancy in video content. Emerging techniques leverage neural networks. Among them, implicit neural representations aim at parameterizing a function that takes coordinates as input and outputs signal values at these coordinates. INR can be used for example to compress images, videos or 3D objects or scenes. It can also be applied to any kind of signal. Methods are known for constructing INR networks for encoding 2D or 3D images. Methods are also known for compressing neural networks. However, any image / video decoder cannot reconstruct an image or video for display based on INR only. Additional information is needed for any image / video decoder to reconstruct the output signal. SUMMARY

[0004] According to one aspect, a method for signaling one or more syntax elements for using an INR decoder is provided. The one or more syntax elements are for reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing the at least a portion of the signal representing the scene. In some embodiments, the method further comprises encoding the INR network.

[0005] According to another aspect, an apparatus for signaling one or more syntax elements for using an INR decoder is provided. The apparatus comprises one or more processors operable for signaling one or more syntax elements, the one or more syntax elements for reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing the at least a portion of the signal representing the scene. In some embodiments, the apparatus is further operable for encoding the INR network.

[0006] According to another aspect, a method for decoding one or more syntax elements for use with an INR decoder is provided. The one or more syntax elements are for use in reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing at least a portion of the signal representing the scene. In some embodiments, the method further comprises decoding the INR network.

[0007] In some embodiments, the method further comprises reconstructing at least a portion of the signal representing the scene using the one or more syntax elements and the implicit neural representation network.

[0008] According to another aspect, an apparatus for decoding one or more syntax elements for use with an INR decoder is provided. The apparatus comprises one or more processors operable to decode one or more syntax elements for use in reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing at least a portion of the signal representing the scene. In some embodiments, the apparatus is further operable to decode the INR network. In some embodiments, the apparatus is further operable to reconstruct at least a portion of the signal representing the scene using the one or more syntax elements and the implicit neural representation network.

[0009] Further embodiments, which can be used individually or in combination, are described herein.

[0010] One or more embodiments also provide a computer program comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method for signaling / decoding one or more syntax elements for use with an INR decoder according to any of the embodiments described herein. One or more embodiments also provide a non-transitory computer-readable medium and / or computer-readable storage medium having stored thereon instructions for signaling / decoding one or more syntax elements for use with an INR decoder according to the methods described herein.

[0011] One or more embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the methods described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving a bitstream generated according to the methods described above. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 An example of a neural network for implicit neural representations is shown.

[0013] Figure 2An example of a method for encoding information used by a decoder to reconstruct at least a portion of a signal representing a scene from an encoded INR is shown, according to an embodiment.

[0014] Figure 3 An example of a method for decoding information used by a decoder to reconstruct at least a portion of a signal representing a scene from an encoded INR is shown, according to an embodiment.

[0015] Figure 4 An example of a method for reconstructing at least a portion of a signal representing a scene from an encoded INR is shown, according to an embodiment.

[0016] Figure 5 An example of a method for generating coordinates for use by a decoder to reconstruct at least a portion of a signal representing a scene from an encoded INR is shown, according to an embodiment.

[0017] Figure 6 An example of a method for performing INR network inference is shown, according to an embodiment.

[0018] Figure 7 An example of a method for reconstructing at least a portion of a signal representing a scene from an output of an INR network is shown, according to an embodiment.

[0019] Figure 8 A block diagram of a system in which aspects of the present embodiments can be implemented is shown.

[0020] Figure 9 A block diagram of a system in which aspects of the present embodiments can be implemented is shown, according to another embodiment.

[0021] Figure 10 Two remote devices communicating over a communication network according to an example of the present principles are shown.

[0022] Figure 11 A syntax of a signal according to an example of the present principles is shown. DETAILED DESCRIPTION

[0023] This application describes a number of aspects, including tools, features, embodiments, models, schemes, and the like. Many of these aspects are described in specific ways, and often in ways that can sound limiting at least to show individual features. However, this is for clarity of description, and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. Moreover, these aspects can also be combined and interchanged with aspects described in earlier applications.

[0024] The aspects described and contemplated in this application can be implemented in many different forms. The following detailed description Figures 1 to 11Some embodiments are provided, but other embodiments are contemplated, and the discussion is not limiting of the scope of implementation. Figures 1 to 11 The discussion of art herein is intended to be illustrative only and not limiting of the scope of implementation.

[0025] In this application, the terms “reconstruction” and “decoding” can be used interchangeably, the terms “pixel” and “sample” can be used interchangeably, and the terms “image,” “picture,” and “frame” can be used interchangeably.

[0026] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless otherwise specified, the order of the specific steps and / or acts can be modified and / or combined and / or used in combination with other steps and / or acts in various embodiments. Also, terms such as “first” and “second” can be used to modify an element, component, step, action, etc. in various embodiments, for example, “first decoding” and “second decoding.” Unless specifically required otherwise, the use of these terms does not limit the scope of the modification of the operations being modified. Thus, in this example, the first decoding does not necessarily need to be performed before the second decoding, and can occur before, during, or overlapping in time with the second decoding.

[0027] The aspects described in this application can be used individually or in combination, unless otherwise specified or technically impracticable.

[0028] At least one aspect generally relates to image, video, or 3D data encoding and decoding using implicit neural representations. More generally, at least one aspect described herein relates to using implicit neural representations to encode / decode any signal representing a scene.

[0029] At least another aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media (having instructions stored thereon for encoding or decoding a data signal representing a scene according to any of the described methods), and / or computer-readable storage media (having a bitstream generated according to any of the described methods stored thereon).

[0030] While not yet standardized, the MPEG group is standardizing new neural compression techniques. In the AhG within WG4, research groups from academia and some companies are working on compression techniques based on INR (INR stands for implicit neural representation), with an exponential growth in the number of papers.

[0031] Research on INR is mainly focused on 2D and video compression, but INR is also being researched for many other signals, in particular 3D scenes or objects. Moreover, the computational complexity of these schemes is much lower than end-to-end neural compression schemes.

[0032] Figure 1An example of a neural network used for Implicit Neural Representation (INR) is shown. This type of neural network for INR can be called an INR network. INR parameterizes the signal as a function 100, which takes coordinates 110 as input and outputs the signal values ​​120 at those coordinates. INR has recently been applied to applications such as images, videos, or 3D objects. In the case of images, the input 110 can be pixel coordinates. INR can output the color value of 120 input pixels. or Input coordinates can be modified through transformations before being used as input to a neural network. These transformations can be Fourier mappings, coordinate transformations, normalization, etc. INR can be used to reconstruct a signal by computing signal values ​​for each necessary coordinate input. It can be used to upsample a signal by generating outputs corresponding to the input coordinates of upsampled pixels, for example, generating the average coordinates between two consecutive pixels for a 2x upsampling.

[0033] Other types of input can be considered, such as when representing the scene as volume data, using point clouds, meshes, or any other suitable scene representation. In some variations, the scene can also be animated / dynamic, meaning it changes over time.

[0034] INR networks are typically neural networks, composed of multiple neural layers, such as fully connected layers. Figure 1 In this network, there are four layers. The intermediate outputs are represented by circles. Each neural layer can be described as a function: first, the input is multiplied by a tensor, added to a vector called the bias, and then a nonlinear function is applied to the result. The shape of the tensor (and other features) and the type of the nonlinear function are called the network's architecture. The values ​​of the tensor and bias are referred to in this paper as "weights." The weights and the parameters of the nonlinear function (if applicable) are called the network's parameters. Architecture and parameters define a "model". Symbols Used to indicate by Parameterized INR function.

[0035] The typical process of encoding a signal using INR is as follows. First, the weights of the INR network are optimized. (or a subset thereof) to reconstruct the signal. Next, these weights are optionally encoded to create an output bitstream. For a size of... Image For example, the weights can be optimized by minimizing the following loss function. :

[0036] , in It is quantified by the difference between the reconstructed image and the original image a distortion, the bitrate of the encoded parameters, is and a trade-off parameter between can be any differentiable distortion measure, for example the mean squared error in the second equation. and are the width and height of the image. Other measures can also be used in this case, for example LPIPS (Learned Perceptual Image Patch Similarity). The weight optimization is typically performed by a machine learning scheme, for example batch gradient descent.

[0037] To decompress the signal, evaluate at all relevant coordinates. These coordinates can be chosen at the time of decoding. For images or videos, a typical choice is all pixel coordinates. For example, for a 256x256 pixel image, these coordinates can be all pairs where all and all Other choices are also possible, for example upsampling, downsampling or extending the original image.

[0038] The scientific literature describes many schemes to build INR networks to encode 2D or 3D images. There are also many existing schemes to compress and encode neural networks, for example the MPEG Neural Network Compression (NNC) standard.

[0039] However, an INR network alone (optimized by the encoder and possibly encoded by NNC) is not sufficient to reconstruct the image at the decoder. In operation, this decoder needs additional information to correctly reconstruct the encoded input signal.

[0040] Some embodiments support the encoding of the information needed by the decoder to reconstruct the input signal (for example an image or video, or 3D data). In some variants, a bitstream is described that contains the INR network and the additional information needed to reconstruct the signal, including SEI syntax. It is also described how the decoder uses this information to reconstruct the signal.

[0041] Some embodiments describe mandatory information that the bitstream needs to carry, regardless of the INR method used.

[0042] Figure 2An example of a method 200 for encoding information used by a decoder to reconstruct at least a portion of a signal representing a scene from an encoded INR is shown, according to an embodiment. A bitstream containing an INR network and the information needed by the decoder to reconstruct the signal using the INR network is constructed, the INR network being optimized and encoded using an off-the-shelf scheme. At 201, the INR is encoded, for example using a NNC encoder. The INR represents at least a portion of a signal representing a scene. This can be a region of an image, a component of an image, a portion of a video, or data of a static or dynamic 3D scene. The weights of the INR are determined using an off-the-shelf scheme. The weights of the INR are encoded in the bitstream.

[0043] At 202, one or more syntax elements (SE) are also signaled in the bitstream. The one or more syntax elements are used to reconstruct at least a portion of the signal representing the scene at the decoder using the encoded INR.

[0044] Figure 3 An example of a method 300 for decoding information used by a decoder to reconstruct at least a portion of a signal representing a scene from an encoded INR is shown, according to an embodiment. For example, a bitstream encoded by the method is decoded. Figure 2 At 301, one or more syntax elements (SE) are decoded from the bitstream. At 302, the syntax elements are provided to the decoder for reconstructing at least a portion of the signal representing the scene using the syntax elements and an implicit neural representation network.

[0045] Some embodiments describing examples of the above syntax elements are provided below. In addition, some embodiments for reconstructing at least a portion of the signal representing the scene using the decoded syntax elements are also provided below.

[0046] As mentioned above, the bitstream should contain the signaling for some or all of the following elements.

[0047] In a variant, the type of the INR used to represent the at least a portion of the signal representing the scene is signaled. This signaling is related to the INR used, for example, a single network is used for the whole image or scene, one network is used for each channel or each group of channels (for example one network for Y and one network for UV), one network is used for each patch in the image or each sub-volume of data in the scene, etc.

[0048] In another variant, information related to the input of the INR network is signaled. These elements describe the information needed to build the input of the INR.

[0049] The dimension of the domain can be signaled, for example the image is a 2D signal. This describes the dimension of the output image or scene.

[0050] Input coordinate normalization can be signaled. Coordinate normalization refers to the range of coordinates that the INR expects. For example, an INR can expect input values in the range , the range or the range .

[0051] Input coordinate transformation can be signaled. Input coordinates are often transformed by a function before being fed to the neural network. In order to reconstruct the image, it is necessary to signal the transformation that must be used. If there are multiple transformations, their order can also be signaled.

[0052] The maximum upsample size can be signaled. This information can be optional. It can be meaningful to signal the maximum recommended upsample that is possible for this INR. The maximum upsample can be defined in various ways, for example, by a maximum upsample rate that will not increase distortion by a set amount, or by an upsample rate that makes the INR upsample better than the upsample algorithm available in the decoder.

[0053] In another variant, information related to the model is signaled. In practice, the model is often integerized. That is, the weights are quantized and represented as fixed-point values, and the floating-point operations are converted to their integer counterpart. For example, activation layers such as sin(x) can be converted to LUTs when the network is integerized. Therefore, it is necessary to precisely describe the range and bit-depth of the input / output.

[0054] In another variant, information related to the signal reconstruction is signaled.

[0055] For example, the output format can be signaled. This signaling indicates the format of the image encoding generated by the INR network, for example, RGB or YUV. Depending on this encoding, additional signaling can be needed. For example, when using the YUV format, signaling of chroma subsampling and / or phase can be needed.

[0056] The output range of the INR network can be signaled, for example, [0, 255] or [0, 1].

[0057] Assuming a quantized network, an example of a syntax to signal one or more of the information discussed above is shown in Table 1 below: Table 1 describes some of the syntax elements that can be needed to describe the input and output of the network.

[0058]

[0059] The semantics of the syntax elements shown in Table 1 are as follows: - inr_model_type: a label describing the type of INR scheme used to generate the INR output. For example, the following types can be defined: inr_model_type = 0: the model takes as input a vector of dimension d (specified by inr_input_dimension_count_minus1) and outputs an output. The input is uniformly sampled over the specified range in each dimension. inr_input_dimension_count_ minus1 inr_model_type = 1: the input is sampled at the center of each 32x32 patch using a 32x down-sampling function over the specified range in each dimension. Alternatively, an additional parameter can specify the down-sampling range.

[0060] inr_model_type = 2: the INR contains two neural networks, the first one is applied before the coordinate transformation and the second one is applied afterwards.

[0061] inr_model_type = 2: the INR contains two neural networks, the first one is applied before the coordinate transformation and the second one is applied afterwards.

[0062] Xxxx: other schemes can also be specified.

[0063] - inr_inference_type: a label describing the type of computation: integer or floating point. This can affect the syntax of some elements. For example, a value of 0 indicates integer computation, and a value of 1 indicates floating point computation.

[0064] - inr_input_dimension_count_minus1: an integer describing the number of input dimensions minus 1. For example, for an image INR model, the input has 2 dimensions (width and height), so inr_input_dimension_count_minus1 is 1. input_dimension_count_minus1

[0065] - inr_input_range_minus1[i]: for each dimension in the input, the input range minus 1. The number of valid values will be inr_input_range [i] = inr_input_range_minus1[i] + 1.

[0066] - inr_input_quantizer[i]: for each dimension in the input, an integer representing the fixed-point integer quantizer. For example, if inr_input_quantizer[i] = 0, then an input x on dimension d represents the value x. If inr_input_quantizer[i] = 3, then an input x on dimension d represents the value (x / 2^3).

[0067] ​- inr_input_zero_centered[i]: if the flag is 0 (false), the input on dimension d is in the range [0, inr_input_range[i]]; if the flag is 1 (true), the value is in the range [inr_input_range[i] / 2 - inr_input_range[i], inr_input_range[i] / 2]. Alternatively, the flag inr_input_zero_centered can be replaced by an explicit encoding of the offset of the zero value.

[0068] For example, a model intended to encode images of size WxH, sub-pixel precision of 1 / 8 pixel (i.e. the network can up-sample the original image by a factor of 8) is described by the following parameters: - inr_inference_type = 0 - inr_input_dimension_count_minus1 = 1 - inr_input_range_minus1[0] = W - 2 - inr_input_quantizer[0] = 3 - inr_input_zero_centered[0] = 0 - inr_input_range_minus1[1] = H - 2 - inr_input_quantizer[1] = 3 - inr_input_zero_centered[1] = 0 The coordinate generator for the input (to recover the original size of the image) would be: for (x = 0; x <= inr_input_range[0]; x++) { for (y = 0; y <= inr_input_range[1]; y++) { output(x, y) = infer(x 2^(inr_input_quantizer[0]), y 2^(inr_input_quantizer[1])) } } where infer(x, y) uses the input coordinates (x, y) to apply the model inference.

[0069] Another example of a model generating 3D point cloud data of voxel coordinate precision 1 / 256 for input coordinates in the range [-1,1]: -inr_inference_type = 0 -inr_input_dimension_count_minus1 = 2 -inr_input_range_minus1[0] = 1 -inr_input_quantizer[0] = 8 -inr_input_zero_centered[0] = 1 -inr_input_range_minus1[1] = 1 -inr_input_quantizer[1] = 8 -inr_input_zero_centered[1] = 1 -inr_input_range_minus1[2] = 1 -inr_input_quantizer[2] = 8 -inr_input_zero_centered[2] = 1 The coordinate generator for input (to recover the volume data) would be: fx = 2^inr_input_quantizer[0] fy = 2^inr_input_quantizer[1] fz = 2^inr_input_quantizer[2] startx = (inr_input_range[0] / 2 - inr_input_range[0]) fx = -256 endx = (inr_input_range[0] / 2) fx = 256 starty = (inr_input_range[1] / 2 - inr_input_range[1]) fy = -256 endy = (inr_input_range[1] / 2) fy = 256 startz = (inr_input_range[2] / 2 - inr_input_range[2]) fz = -256 endz = (inr_input_range[2] / 2) fz = 256 for(x=startx; x<=endx; x++) { for(y=starty; y<endy; y++) { for(z=startz; z<endz; z++) { output(x-startx, y-starty, z-startz) = infer(x, y, z) } } } - inr_input_transformation_count: Integer describing the number of successive transformations applied to the input before feeding into the INR network. For example, if one transformation is applied, the value is 0. If three transformations are applied, the value is 2.

[0070] - inr_input_transformation_type[i]: Label describing the type of transformation applied to the input. Type 0 is reserved for defining custom types. For example, the following types can be defined: inr_input_transformation_type[i] = 1: Hyperspherical coordinate transformation inr_input_transformation_type[i] = 2: Fourier mapping using a custom Fourier mapping matrix. This type value requires inr_inference_type to be floating-point. In this case, the custom matrix is described as follows: inr_Fourier_mapping_coefficient_count_minus1[i]: Integer describing the number of Fourier mapping coefficient transformations minus 1.

[0071] inr_Fourier_mapping_coefficient[i][j]: Value of the Fourier mapping coefficient.

[0072] inr_input_transformation_type[i] = 3: Normalization inr_input_transformation_type[i]=4: Fourier mapping using a first predefined Fourier mapping matrix inr_input_transformation_type[i]=5: Fourier mapping using a second predefined Fourier mapping matrix inr_input_transformation_type[i]=6: Fourier mapping using a predefined Fourier mapping LUT etc. - inr_output_type: describes the output type. Type 0 is reserved for defining a custom type. Other types can be predefined for example: inr_output_type=1: RGB output, each component coded in 8 bits in the range [0, 255], inr_output_type=2: RGB output, each component coded in 10 bits in the range [0, 1023], inr_output_type=3: YUV output, each component coded in 8 bits in the range [0, 255], inr_output_type=4: YUV output, each component coded in 10 bits in the range [0, 1023], inr_output_type=5: RGBA output, each component coded in 8 bits in the range [0, 255], where A represents the alpha value, etc.

[0073] When the type is 0, the output characteristics are described using the same logic as the input (number of components, range, quantizer, zero value offset).

[0074] For example, a model without any transformation and with an RGB output range in [0, 255] is described by the following parameters: - inr_input_transformation_count=0 - inr_output_type=1 A coordinate that must be mapped using a custom Fourier mapping with 3 coefficients a, b and c and then normalized and output in YUV format is described as follows: -inr_input_transformation_count = 2 -Inr_input_transformation_type[0] = 2 -inr_Fourier_mapping_coefficient_count_minus1[0] = 2 -inr_Fourier_mapping_coefficient[0][0] = a -inr_Fourier_mapping_coefficient[0][1] = b -inr_Fourier_mapping_coefficient[0][2] = c -Inr_input_transformation_type[1] = 3 -inr_output_type = 3 Note that the transformation can be represented as a LUT, typically used for integer extrapolation. For example, based on the previous example (reconstructing a 3D point set), using Fourier mapping, one signal decoding can be computed as follows: fx = 2^inr_input_quantizer[0] fy = 2^inr_input_quantizer[1] fz = 2^inr_input_quantizer[2] startx = (inr_input_range[0] / 2 - inr_input_range[0]) fx = -256 endx = (inr_input_range[0] / 2) fx = 256 starty = (inr_input_range[1] / 2 - inr_input_range[1]) fy = -256 endy = (inr_input_range[1] / 2) fy = 256 startz = (inr_input_range[2] / 2 - inr_input_range[2]) fz = -256 endz = (inr_input_range[2] / 2) fz=256 for(x=startx;x<=endx;x++) { for(y=starty;y<endy;y++) { for(z=startz;z<endz;z++) { mapped_coord = LUT_transform(x,y,z) output(x-startx,y-starty,z-startz)=infer(mapped_coord) } } } The LUT_transform function returns the mapped coordinate in the "mapped_coord" variable.

[0075] One scheme for constructing LUT_transform for Fourier mapping can be as follows. Assume a Fourier mapping coefficient set Fourier_mapping_coef and an input range [start, end]: LUT_table_s = initialize() LUT_table_c = initialize() for(x=start;x<=end;x++){ range = get_range(x) for(i=0;i<len( Fourier_mapping_coef );i++) LUT_table_s[x][i]=sin_int_rep(range, Fourier_mapping_coef [i]) LUT_table_c[x][i]=cos_int_rep(range, Fourier_mapping_coef [i]) } where len( Fourier_mapping_coef ) returns the length of the input set Fourier_mapping_coef , get_range(x) returns the real value range mapped to integer representation x, for example, if the uniform quantization step size is and x'=inv_quantization(x) is the real value mapped to x, usually sin_int_rep(range, coef) is a function that returns the integer representation of the value of the sine function on that interval. For example, the following method can be used:

[0076]

[0077] cos_int_rep(range, coef) can be defined similarly:

[0078]

[0079] Other multipliers can also be added inside the sin or cos, for example or multiples thereof. One can also take advantage of the fact that to avoid storing a table for the cosine values, but instead store a LUT_complementary_angle table that stores the integer representation of In this case, one can use for example LUT_table_c[x][i]:= LUT_table_s[LUT_complementary_angle table [x]][i].

[0080] LUT_transform(x,y,z) can then be implemented as follows: LUT_transform(x,y,z){ offset=0 n_coef = len( Fourier_mapping_coef ) input = (x,y,z) for(j=0;j<len(input);j++): val = input[j] for(i=0;i<n_coef;i++) output[offset+i]= LUT_table_s[val][i] output[offset+i+n_coef]= LUT_table_c[val][i] offset+=2 n_coef } In some embodiments, regarding Figure 2 and 3The mentioned syntax elements can also include network configuration information. For example, the neural network coding can be signaled. This element signals how the INR neural network is coded in the bitstream. For example, it can be a value associated with a specific format (such as NNC or Open Neural Network Exchange format) and / or a specific version of the format.

[0081] The inference engine configuration can be signaled. It can be meaningful to add signaling related to the inference engine configuration, for example, it can include the precision used for the operation, the memory required to store the INR network or the inference engine to be used.

[0082] Table 2 describes some syntax elements that can be needed to describe the network configuration:

[0083] The semantics of Table 2 are as follows: inr_mode_idc indicates the format used to encode the neural network. For example, a value equal to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bitstream, a value equal to 1 indicates that this SEI message contains a neural network encoded in a format identified by the tag URI inr_tag_uri.

[0084] inr_tag_uri contains a tag URI, the syntax and semantics of which are specified in IETF RFC 4151, for identifying the INR network or the updated format and related information encoded here.

[0085] inr_uri contains a URI, the syntax and semantics of which are specified in IETF Internet Standard 66, for identifying the neural network used as the INR network or the update relative to this network.

[0086] inr_complexity_info_present_flag specifies whether syntax elements indicating the INR network complexity are present. For example, a value equal to 1 specifies that one or more syntax elements indicating the INR network complexity are present, inr_complexity_info_present_flag equal to 0 specifies that no syntax elements indicating the INR network complexity are present.

[0087] inr_parameter_type_idc equal to 0 indicates that the neural network uses only integer parameters. inr_parameter_type_flag equal to 1 indicates that the neural network can use floating-point or integer parameters. inr_parameter_type_idc equal to 2 indicates that the neural network uses only binary parameters. inr_parameter_type_idc equal to 3 is reserved for future use.

[0088] inr_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicates that the neural network does not use parameters with bit length greater than 8, 16, 32, and 64, respectively. When inr_parameter_type_idc is present and inr_log2_parameter_bit_length_minus3 is not present, the neural network does not use parameters with bit length greater than 1.

[0089] inr_num_parameters_idc indicates the maximum number of neural network parameters for the INR network in powers of 2 (or another power of 2). inr_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of inr_num_parameters_idc shall be in the range of 0 to 63, inclusive.

[0090] If the value of inr_num_parameter_idc is greater than zero, the variable maxNumParameters is derived as follows: maxNumParameters = ( 2 « inr_num_parameters_idc ) 1 The number of neural network parameters for the post-processing filter shall be less than or equal to maxNumParameters.

[0091] inr_num_mac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per inference for the INR network is less than or equal to inr_num_mac_operations_idc. inr_num_mac_operations_idc equal to 0 indicates that the maximum number of multiply-accumulate operations for the network is unknown. The value of inr_num_mac_operations_idc shall be in the range of 0 to 2 32

[0092] ​inr_total_kilobyte_size greater than 0 indicates the total size (in kilobytes) required to store the uncompressed parameters for the neural network. The total size (in bits) is a sum equal to or greater than the number of bits used to store each parameter. inr_total_kilobyte_size is the total size (in bits) divided by 8000 rounded up. inr_total_kilobyte_size equal to 0 indicates that the total size required to store the parameters for the neural network is unknown. The value of inr_total_kilobyte_size shall be in the range of 0 to 2 32 1, inclusive.

[0093] inr_reserved_zero_bit_b shall be equal to 0 in bitstreams that conform to the syntax provided herein. Decoders shall ignore INR SEI messages for which inr_reserved_zero_bit_b is not equal to 0.

[0094] inr_payload_byte[ i ] contains the i-th byte of a conforming bitstream of ISO / IEC 15938-17. The sequence of bytes inr_payload_byte[ i ] shall be a complete conforming bitstream of ISO / IEC 15938-17 for all i values for which it exists.

[0095] The syntax elements provided above are examples only, additional syntax elements can be added and some can be removed. Furthermore, the order in which the syntax elements are presented is illustrative, one or more syntax elements can be presented in any other order.

[0096] Figure 4 An high level overview of an example method 400 that a decoder can rely on when decoding a bitstream 410 containing an encoded INR representing at least a portion of a signal characterizing a scene is shown. The embodiments described herein are for a signal characterizing an image. But the described embodiments are applicable to any other type of signal characterizing a scene. The bitstream also contains one or more syntax elements according to embodiments (e.g. described above in connection with Table 1 or 2).

[0097] The input bitstream 410 or a portion of this bitstream is fed to modules of a process. Module 420 extracts information from the bitstream, generates and outputs sets of input coordinates for the INR network. Module 430 performs INR inference. It is responsible for decoding the INR network and generating outputs of the INR network, in other words, associating outputs of the INR network with each set of input coordinates. Module 440 combines these outputs to reconstruct the image. This can involve transforming the outputs to match a requested output format. But this module is also responsible for correctly arranging the values into an image format.

[0098] Figure 5 One possible embodiment of the input generation module 520 is shown. In step 510, the original image size signaled in the bitstream is taken as input 515 and used to generate input coordinates corresponding to the original image. For example, these input coordinates are typically pairs of positions, where and range from 1 to the width or height of the image, respectively. Depending on the type of INR used, alternatives are possible, for example using coordinates associated with a pixel block. For example, for an INR outputting pixel block color values, the input is typically the coordinates of the block. For example, such an INR uses the syntax provided below to define it as inr_model_type = 1, and the input coordinates use this model type to subsample accordingly.

[0099] Step 530 takes as input the signaling 535 related to coordinate transformations in the bitstream and applies these transformations to the coordinates provided at 510.

[0100] Step 550 arranges the coordinates based on the INR type 555 signaled in the bitstream. This step can also be optional. This step ensures that the correct values are used for each input of the INR network. This module outputs the coordinates 560 to be used by the INR network.

[0101] Figure 6 One possible embodiment of the INR inference module 430 is shown. In step 610, the INR network encoding signaled in the bitstream 615 is taken as input to prepare the decoding. This piece of information is used to configure the INR decoder module 620. This module 620 takes as input the bitstream 625 of the INR network (i.e. the encoded INR) and outputs the decoded INR network.

[0102] Module 630 uses the inference engine configuration 635 and the decoded INR network to initialize the inference engine 640. This module 630 can involve steps such as reserving computation resources, loading the decoded model into the memory of the processor that will perform the inference, transforming the network into a format expected by the inference engine, adjusting the weights of the network for the inference engine, etc. If multiple networks are used in the INR, it can also involve defining in which networks which coordinates are used. In this case, the input 635 to module 630 can also include signaling of the INR type to use.

[0103] This initialized inference engine 640 takes as input the coordinates 560. These coordinates are fed to the INR network and an inference is performed to compute the values associated with these pixels. The inference can have several variations. For example, the inference can be performed on one set of coordinates at a time, while the inference is performed on multiple sets of coordinates, or the inference is performed in parallel on multiple processors. The module outputs the coordinates and the associated computed pixel values 660.

[0104] Figure 7 One possible embodiment of the inference output combination module 440 is shown. Module 710 receives the coordinates and the associated pixel values 660. It can also take as input the type of INR used and the output format of the network (716). This module is responsible for reordering the pixel values into a traditional image format, for example a matrix arranged in row-major or column-major order. It can also take as input the type of INR used if necessary. One example is when the input coordinates are associated with values of a pixel block, in order to know the order of the output, or when this step involves the combination of multiple INR network outputs. Another example is when different networks generate Y and UV values respectively.

[0105] Module 720 takes as input the output range of the network 726. Then a coordinate change is performed on the pixel values to obtain values that lie in the traditional range used for images, for example [0, 255] or [0, 1]. The final output 730 is the decoded picture.

[0106] This output can be further modified to obtain different image formats, for example from YUV to RGB or vice versa.

[0107] The above process is just one example. Depending on the INR type considered or signaled in the bitstream, additional steps can be required, or some steps can be omitted.

[0108] The order of some steps can also be different. For example, upsampling can be performed separately from the generation of the input range 420. Although a sequential process is described, some operations can be performed in parallel. For example, the network can be decoded 620 while the coordinates are prepared 420. Some steps can also be moved to different modules.

[0109] Figure 8An example block diagram of a system in which various aspects and embodiments can be implemented is shown. System 800 can be embodied as a device including various components described below and configured to perform one or more of the aspects described in the present disclosure. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, web servers, and the like. The elements of system 800, singly or in combination, can be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 800 are distributed across multiple ICs and / or discrete components. In various embodiments, system 800 is communicably coupled to other systems, or to other electronic devices, such as through a communications bus or through dedicated input and / or output ports. In various embodiments, system 800 is configured to implement one or more of the aspects described in the present disclosure.

[0110] System 800 includes at least one processor 810 configured to execute instructions loaded therein for implementing, for example, various aspects described in the present disclosure. Processor 810 can include embedded memory, input output interface, and various other circuitry as known in the art. System 800 includes at least one memory 820 (e.g., a volatile memory device and / or a non-volatile memory device). System 800 includes a storage device 840, which can include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disks, and / or optical disks. Storage device 840 can include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.

[0111] System 800 includes an encoder / decoder module 830, which is configured to, for example, process data to provide an INR representing at least a portion of a signal that characterizes a scene and / or an encoded INR representing at least a portion of a signal that characterizes a scene or a decoded INR representing at least a portion of a signal that characterizes a scene and / or at least a portion of the signal reconstructed from a decoded INR, and encoder / decoder module 830 can include its own processor and memory. Encoder / decoder module 830 represents a module that can be included in a device to perform encoding and / or decoding functions. As is known, a device can include one or both of encoding and decoding modules. Additionally, encoder / decoder module 830 can be implemented as a separate element of system 800, or can be incorporated as a combination of hardware and software within processor 810, as is known to those of skill in the art.

[0112] Program code to be loaded onto processor 810 or encoder / decoder 830 to perform aspects described in this application can be stored in storage device 840 and then loaded into memory 820 for execution by processor 810. In accordance with various embodiments, one or more of processor 810, memory 820, storage device 840, and encoder / decoder module 830 can store one or more of various items during the performance of processes described in this application. These stored items can include, but are not limited to, input images, video, 3D data, weights of INRs, decoded images, decoded video, decoded 3D data or portions of decoded data, bitstreams, matrices, variables, and intermediate or final results in equations, formulas, operations, and arithmetic logic processing.

[0113] In some embodiments, memory internal to processor 810 and / or internal to encoder / decoder module 830 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be processor 810 or encoder / decoder module 830) is used for one or more of these functions. The external memory can be memory 820 and / or storage device 840, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the operating system of the television.

[0114] Input to the elements of system 800 can be provided through various input devices as shown by block 805. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives RF signals transmitted, e.g., by a broadcaster over the air, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 8 Other examples not shown in FIG. 8 include composite video.

[0115] In various embodiments, the input devices of block 805 have respective associated input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a frequency band); (ii) downconverting the selected signal; (iii) band-limiting again to a still narrower frequency band to select (for example) a signal frequency band (which can be referred to as a channel in some embodiments); (iv) demodulating the downconverted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select the desired data stream. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs multiple ones of these functions, for example, including downconverting the received signal (s) to lower frequency (s) (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing elements receive an RF signal transmitted through a wired (for example, cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of those elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF portion includes an antenna.

[0116] In addition, the USB and / or HDMI terminals can include respective interface processors for connecting the system 800 to other electronic devices through USB and / or HDMI connections. It will be appreciated that various aspects of input processing (for example, Reed-Solomon error correction) can be implemented as desired within a separate input processing IC or within the processor 810, for example. Similarly, various aspects of USB or HDMI interface processing can be implemented as desired within a separate interface IC or within the processor 810. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including the processor 810 and the encoder / decoder 830, which operate in conjunction with memory and storage elements to process the data streams as desired for presentation on output devices.

[0117] Various elements of the system 800 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data to one another using suitable connection means 815, for example, internal buses (including an I2C bus), wiring, and printed circuit boards, as known in the art.

[0118] The system 800 includes a communication interface 850 that enables communication with other devices via a communication channel 890. The communication interface 850 can include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 890. The communication interface 850 can include, but is not limited to, a modem or network card, and the communication channel 890 can be implemented, for example, in wired and / or wireless media.

[0119] In various embodiments, data is streamed to the system 800 using a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)). The Wi-Fi signals of these embodiments are received through the communication channel 890 and the communication interface 850, which are adapted for Wi-Fi communication. The communication channel 890 of these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top traffic. Other embodiments use a set-top box to provide streaming data to the system 800 through the HDMI connection of the input block 805. Still other embodiments use the RF connection of the input block 805 to provide streaming data to the system 800. As mentioned above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0120] The system 800 can provide output signals to various output devices, including a display 865, speakers 875, and other peripheral devices 885. The display 865 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 865 can be used in a television, a tablet, a notebook computer, a cell phone (mobile phone), or other device. The display 865 can also be integrated with other components (e.g., as in a smartphone), or separate (e.g., an external display for a notebook computer). The other peripheral devices 885 include, in various example embodiments, one or more of a digital video recorder (or digital versatile recorder) (DVR, either term applies), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more of the peripheral devices 885 to provide functionality based on the output of the system 800. For example, the disk player performs the function of playing the output of the system 800.

[0121] In various embodiments, control signals between system 800 and display 865, speakers 875, or other peripheral devices 885 use signaling such as AV.Link, CEC, or other communications protocols that allow stream management between the devices with or without user intervention. Output devices can be communicatively coupled to system 800 through respective interfaces 860, 870, and 880 via dedicated connections. Alternatively, output devices can be connected to system 800 using communication channel 890 through communication interface 850. Display 865 and speakers 875 can be integrated in a single unit with other components of system 800 in an electronic device, for example, a television. In various embodiments, display interface 860 includes a display driver, for example, a timing controller (T Con) chip.

[0122] Display 865 and speakers 875 can be separate from one or more other components, for example, if the RF portion of input 805 is part of a standalone set-top box. In various embodiments where display 865 and speakers 875 are external components, output signals can be provided over dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0123] Embodiments can be performed by computer software implemented by processor 810, or by hardware, or by a combination of hardware and software. Embodiments can be implemented by one or more integrated circuits, as non-limiting examples. Memory 820 can be of any type suitable to the

[0124] Figure 9 A block diagram of a system that can implement aspects of the present embodiments is shown, in accordance with another embodiment. Figure 9 An embodiment of an apparatus 900 for encoding or decoding metadata used by an INR decoder according to any of the embodiments described herein is shown. The apparatus includes a processor 910 and can be interconnected with a memory 920 through at least one port. Both processor 910 and memory 920 can also have one or more additional interconnections to external connections.

[0125] The processor 910 is further configured to encode, using any of the embodiments described herein, at least one implicit neural representation network representing at least a portion of a signal representing a scene, and to signal one or more syntax elements for reconstructing the at least a portion of the signal representing the scene using the encoded implicit neural representation network.

[0126] In another variant, the processor 910 is configured to decode, using any of the embodiments described herein, one or more syntax elements for using an implicit neural representation network representing at least a portion of a signal representing a scene, and to reconstruct the at least a portion of the signal representing the scene using the one or more syntax elements and the implicit neural representation network. For example, the processor 910 is configured using a computer program product comprising code instructions implementing any of the embodiments described herein.

[0127] In one embodiment, as shown in Figure 10 the context of a transmission over a communication network NET between two remote devices A and B, device A comprises a processor associated with a memory RAM and ROM configured to implement a method for encoding one or more syntax elements for using an INR decoder as described in Figures 1 to 7 device B comprises a processor associated with a memory RAM and ROM configured to implement a method for decoding one or more syntax elements for using an INR decoder as described in Figures 1 to 7 According to an example, the network is a broadcast network adapted to broadcast / transmit the encoded INR and the one or more syntax elements from device A to a decoding device comprising device B. In some embodiments, the encoded INR and the one or more syntax elements are transmitted in the same signal. In other embodiments, the encoded INR and the one or more syntax elements are transmitted separately in different signals.

[0128] Figure 11 A syntax example of a signal transmitted over a packet-based transmission protocol is shown. Each transmitted packet P comprises a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD can comprise one or more syntax elements for using an INR decoder according to any of the embodiments described above.

[0129] In some embodiments, the one or more syntax elements comprise at least one of: an indication related to an implicit neural representation network, information for constructing an implicit neural representation network input, and information related to an implicit neural representation network output.

[0130] In one variant, the indications associated with the implicit neural representation network include at least one of the following: the model type of the implicit neural representation network, the computational type of the implicit neural representation network, the encoding format of the implicit neural representation network, an indication of a label URI identifying the format of the implicit neural representation network, an indication of a URI identifying the implicit neural representation network, an indication of the existence of complexity information associated with the implicit neural representation network, an indication of the type of parameters used by the implicit neural representation network, an indication of the bit length of the parameters used by the implicit neural representation network, an indication of the maximum number of parameters of the implicit neural representation network, an indication of the maximum number of multiply-accumulate operations per inference by the implicit neural representation network, and an indication of the size of the uncompressed parameters used to store the implicit neural representation network.

[0131] In another variant, the information used to construct the input of the implicit neural representation network includes at least one of the following: an indication of the number of input dimensions, an indication of the range of input dimensions, an indication of the input dimension quantizer, an indication of the zero-value offset within the range, an indication of the number of transformations applied to the input, an indication of the type of transformation applied to the input, and an indication of one or more parameters of the transformation.

[0132] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of the processes performed, for example, on a received encoded INR to produce a final output suitable for display. In various embodiments, such processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include, or alternatively include, processes performed by a decoder in various implementations described herein, such as entropy decoding of a sequence of binary symbols to reconstruct image, video, or 3D data.

[0133] It should be noted that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.

[0134] This application describes various information fragments, such as syntax, that can be transmitted or stored. This information can be packaged or arranged in various ways, including those common in image, video, or neural network standards, such as placing the information in SPS, PPS, NAL units, headers (e.g., NAL unit headers or fragment headers), or SEI messages. Other methods are also available, including those common in system-level or application-level standards, such as placing the information in one or more of the following: a. SDP (Session Description Protocol), a format for describing multimedia communication sessions, used for session announcements and invitations, such as those described in the RFC, and used in conjunction with RTP (Real-Time Transport Protocol) transmission.

[0135] b. DASH MPD (Media Presentation Description) descriptors, e.g., used in DASH and transported over HTTP, associated with a representation or set of representations to provide additional characteristics of the content representation.

[0136] c. RTP header extensions, e.g., used during RTP streaming.

[0137] d. ISO Base Media File Format, e.g., used in OMAF and using object-oriented building blocks (also referred to as "atoms" in some specifications) defined by a unique type identifier and a length.

[0138] e. HLS (HTTP Live Streaming) manifests transported over HTTP. A manifest, for example, can be associated with one version or set of versions of content to provide characteristics of that version or set of versions.

[0139] When a diagram is presented with a flowchart, it is to be understood that same can also be presented with a block diagram. Similarly, when a diagram is presented with a block diagram, it is to be understood that same can also be presented with a flowchart.

[0140] Implementations and aspects described herein can be implemented in a method or process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single implementation form (for example, as a method), implementation of the discussed features can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.

[0141] Reference to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", as well as other variants thereof, mean that a particular feature, structure, characteristic, and so forth described in connection with an embodiment is included in at least one embodiment. Therefore, the appearance of the phrase "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation", as well as any other variations thereof, throughout the application does not necessarily refer to the same embodiment.

[0142] Also, the application can refer to "determining" various pieces of information. Determining the information can include one or more of estimating the information, calculating the information, predicting the information, or retrieving the information from memory, for example.

[0143] Furthermore, the application can relate to "accessing" a variety of pieces of information. Accessing information can include one or more of, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0144] Furthermore, the application can relate to "receiving" a variety of pieces of information. Receiving, like "accessing," is meant to be a broad term. Receiving information can include one or more of, for example, accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" generally involves some action that is taken by a recipient with regard to information, such as storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0145] It should be understood that any "or" as used herein is intended to encompass "and / or," e.g., "A or B" is intended to mean "A or B or both A and B." Further, the use of "and / or" as used herein is intended to encompass any one item, any combination of items, or all of the items in a list. As an example, "A and / or B" is intended to mean "A or B or both A and B." As another example, "A, B, and / or C" is intended to mean "A or B or C or all of A and B and C." As a further example, "A, B, and / or C" is intended to mean "A or B or C or all of A and B and C." As a still further example, "A, B, and / or C" is intended to mean "A or B or C or all of A and B and C or any one of the individual items A and B and C or any combination of the items A and B and C." The use of "and / or" as used herein is merely an attempt to encompass all of the different combinations of items in a list. As a further example, "A, B, and / or C" is intended to mean "A or B or C or all of A and B and C or any one of the individual items A and B and C or any combination of the items A and B and C." The use of "and / or" as used herein is merely an attempt to encompass all of the different combinations of items in a list.

[0146] Furthermore, as used herein, the word "signal" means, among other things, indicating something to a corresponding decoder. In this manner, in embodiments, the same parameters are used at the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicitly signal) a particular parameter to a decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter along with other parameters, a signal can be employed (implicitly signal) to simply let the decoder know and select the particular parameter. By avoiding transmitting any actual function, bit savings are achieved in embodiments. It should be understood that signaling can be accomplished in a variety of ways. For example, in embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. While the foregoing relates to the verb form of the word "signal," the word "signal" can also be used as a noun herein.

[0147] As those skilled in the art will appreciate, the implementations can produce signals formatted to carry information that can be, for example, for transmission or for storage. The information can include instructions for performing a method, or data produced by one of the implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a portion of the radio frequency spectrum) or as a baseband signal. The formatting can include, for example, encoding the data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0148] A number of embodiments have been described. Features of the embodiments can be provided individually or in any combination.

Claims

1. A method comprising signaling one or more syntax elements for reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing the at least a portion of the signal representing the scene.

2. An apparatus comprising one or more processors, wherein the one or more processors are operable to signal one or more syntax elements for reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing the at least a portion of the signal representing the scene.

3. A method comprising decoding one or more syntax elements for reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing the at least a portion of the signal representing the scene.

4. The method of claim 3, further comprising reconstructing the at least a portion of the signal representing the scene using the one or more syntax elements and the implicit neural representation network.

5. An apparatus comprising one or more processors, wherein the one or more processors are operable to decode one or more syntax elements for reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing the at least a portion of the signal representing the scene.

6. The apparatus of claim 5, wherein the one or more processors are further operable to reconstruct the at least a portion of the signal representing the scene using the one or more syntax elements and the implicit neural representation network.

7. The method of claim 1, 3, or 4 or the apparatus of claim 2, 5, or 6, wherein the one or more syntax elements comprise at least one of: an indication related to the implicit neural representation network, information for constructing an input to the implicit neural representation network, and information related to an output of the implicit neural representation network.

8. The method or apparatus of claim 7, wherein the indication related to the implicit neural representation network comprises at least one of: a model type of the implicit neural representation network, a computation type of the implicit neural representation network, an encoding format of the implicit neural representation network, an indication of a label URI of a format identifying the implicit neural representation network, an indication of a URI identifying the implicit neural representation network, an indicator indicating whether there is complexity information related to the implicit neural representation network, an indication of a parameter type used by the implicit neural representation network, an indication related to a bit length of a parameter used by the implicit neural representation network, an indicator indicating a maximum number of parameters for the implicit neural representation network, an indication of a maximum number of multiply-accumulate operations per inference of the implicit neural representation network, and an indication of a size for storing uncompressed parameters of the implicit neural representation network. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 9. The method or apparatus of claim 7 or 8, wherein the information used to construct the input to the implicit neural representation network comprises at least one of: an indication of a number of dimensions of the input, an indication of a range of dimensions of the input, an indication of a quantizer of dimensions of the input, an indication of an offset of zero values within the range, an indication of a number of transformations applied to the input, an indication of a type of transformation applied to the input, and an indication of one or more parameters of the transformation.

10. The method of claim 4 or 7-9 or the apparatus of claim 6-9, wherein reconstructing the at least a portion of a signal representing a scene using the one or more syntax elements and the implicit neural representation network comprises: generating an input from the one or more syntax elements, providing the generated input to the implicit neural representation network.

11. The method of any of claims 4 and 7-10 or the apparatus of any of claims 6-10, wherein reconstructing the at least a portion of a signal representing a scene using the one or more syntax elements and the implicit neural representation network further comprises: performing an implicit neural representation network inference using the one or more syntax elements.

12. The method or apparatus of claim 11, wherein performing the implicit neural representation network inference using the one or more syntax elements comprises decoding the implicit neural representation network using the one or more syntax elements.

13. The method or apparatus of claim 11 or 12, wherein performing the implicit neural representation network inference comprises configuring an inference engine using the one or more syntax elements.

14. The method of any of claims 4 and 7-13 or the apparatus of any of claims 6-13, wherein reconstructing the at least a portion of a signal representing a scene using the one or more syntax elements and the implicit neural representation network further comprises: combining an output of the implicit neural representation network and an input to the implicit neural representation network using the one or more syntax elements.

15. The method or apparatus of claim 14, wherein combining an output of the implicit neural representation network and an input to the implicit neural representation network using the one or more syntax elements comprises arranging the output of the implicit neural representation network in a given format or within a given range of values.

16. The method of claim 1, further comprising encoding or the apparatus of claim 2, wherein the one or more processors are further operable to encode the at least one implicit neural representation network.

17. A computer program product comprising instructions for causing one or more processors to perform the method of any of claims 1, 3, 4, and 7-15.

18. A non-transitory computer readable medium storing executable program instructions to cause a computer executing the program instructions to perform the method of any of claims 1, 3, 4, and 7-15.

19. A bitstream comprising data representing one or more syntax elements for reconstructing at least a portion of a signal representing a scene using at least one implicit neural representation network, the at least one implicit neural representation network representing the at least a portion of a signal representing the scene.

20. The bitstream of claim 19, further comprising data encoding at least a portion of the implicit neural representation network or data for updating the at least one implicit neural representation network.

21. The bitstream of claim 19 or 20, wherein the one or more syntax elements comprise at least one of: an indication related to the implicit neural representation network, information for constructing an input to the implicit neural representation network, and information related to an output of the implicit neural representation network.

22. The bitstream of claim 21, wherein the indication related to the implicit neural representation network comprises at least one of: a model type of the implicit neural representation network, a computation type of the implicit neural representation network, an encoding format of the implicit neural representation network, an indication of a label URI of a format identifying the implicit neural representation network, an indication of a URI identifying the implicit neural representation network, an indicator indicating whether there is complexity information related to the implicit neural representation network, an indication of a parameter type used by the implicit neural representation network, an indication related to a bit length of a parameter used by the implicit neural representation network, an indicator indicating a maximum number of parameters for the implicit neural representation network, an indication of a maximum number of multiply-accumulate operations per inference of the implicit neural representation network, and an indication of a size of uncompressed parameters for storing the implicit neural representation network.

23. The bitstream of claim 21 or 22, wherein the information for constructing an input to the implicit neural representation network comprises at least one of: an indication of a number of dimensions of the input, an indication of a range of dimensions of the input, an indication of a quantizer of dimensions of the input, an indication of an offset of zero values within the range, an indication of a number of transforms applied to the input, an indication of a type of transform applied to the input, and an indication of one or more parameters of the transform.

24. A non-transitory computer readable medium storing the bitstream of any of claims 19 to 24.

25. An apparatus comprising: the apparatus of claim 3 or 4; and at least one of: (i) an antenna configured to receive or transmit a signal, the signal comprising the bitstream of any of claims 19 to 24; (ii) a band limiter configured to limit the signal to a frequency band containing the bitstream; and (iii) a display configured to display the at least a portion of a signal representing the scene.

26. The apparatus of claim 25, wherein the apparatus comprises at least one of: a television, a mobile phone, a tablet, and a set-top box. ​