Method and device for contextual coding and decoding of image sequences.

The method employs a reconstruction neural network to enhance the coding and decoding of image sequences by generating prediction images and updating correction information, addressing limitations in existing video compression techniques and achieving improved compression efficiency.

FR3156567A1Pending Publication Date: 2025-06-13ORANGE SA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2023013813
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing methods for coding and decoding image sequences, particularly in video compression, face limitations in efficiently correcting predicted signals, leading to suboptimal compression performance.

Method used

A method and device that utilize a reconstruction neural network to generate prediction images, encode correction information using these images to produce reconstruction maps, and update prediction and correction information based on coding performance measurements, thereby enhancing the coding and decoding process.

Benefits of technology

This approach allows for efficient compression of image sequences by effectively exploiting redundancy between prediction and correction signals, resulting in improved coding performance and simplified decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method and device for contextual coding and decoding of image sequences The invention relates to a coding method and a decoding method of a current image (Idv) of an image sequence (SV).The decoding comprises the following steps: – obtaining (E30) coded correction information (FCc) of said current image; – obtaining (E30, E31, E32, E33) at least one prediction image (IP / FMP) of said current image; – decoding (E35) said correction information (FCc) using said at least one prediction image to generate a set of feature maps, called reconstruction maps (FMR); – decoding (E34) at least part of a set of reconstruction parameters (WR) representative of at least one reconstruction neural network (MREC); – processing (E37) said reconstruction maps by said at least one reconstruction neural network to produce a decoded representation of said current image (Ivd) Figure for the abstract: Fig. 1.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for contextual coding and decoding of image sequences.

[0001] The invention relates to the general field of coding digital image sequences. It relates more particularly to the compression of digital videos.

[0002] Digital videos are generally subject to source coding aimed at compressing them in order to limit the resources required for their transmission and / or storage. There are many coding standards, such as the ITU / MPEG standards (H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc.) as well as their extensions (MVC, SVC, 3D-HEVC, etc.)

[0003] The encoding of an image is generally carried out by a prediction of the pixels using previously coded then decoded pixels present in the image being encoded, in which case we speak of “Intra prediction”, or of previously coded images, in which case we speak of “Inter prediction”.

[0004] In the case of inter-prediction, a classic approach is often used: after estimation and compensation of the motion at the encoder, a residual frame is calculated by subtracting the prediction frame from the original image. At the decoder, the prediction frame and the residual frame are added to obtain the reconstructed frame. This simple addition not being optimal, new approaches, notably neural, have been developed recently.

[0005] In this field, a conditional neural approach is proposed for example in the document "Conditional Residual Coding: A Remedy for Bottleneck Problems in Conditional Inter Frame Coding", by F. Brand et al (arXiv:2307.12864). The proposed conditional coding models the conditional distribution of an image as a function of its predictor. However ([Fig.3] of this document), the residual data and the predicted data are coded and decoded independently of each other. This limits the compression performance.

[0006] There is therefore a need for a solution for correcting a predicted signal in an image sequence more efficiently. Statement of the invention

[0007] The invention relates to a method for coding a current image at least from a sequence of images, comprising the following steps: - initialization of prediction information of the current image; - initialization of a set of parameters representative of a reconstruction neural network; - generation of at least one prediction image from the information of prediction; - encoding correction information using said at least one prediction image to generate a set of feature maps, called reconstruction maps; - processing said reconstruction feature maps by said at least one reconstruction neural network to produce a decoded representation of said current image; - updating at least one of said correction and / or prediction information and / or at least one parameter of said reconstruction neural network, based on a coding performance measurement; - a step of coding a binary stream comprising: - encoding of said correction information; - coding of said prediction information; - coding at least part of the set of parameters representative of the reconstruction neural network.

[0008] The invention also relates to a method for decoding a current image from a sequence of images, comprising the following steps: - obtaining coded correction information of said current image; - obtaining at least one prediction image of said current image; - decoding said correction information using said at least one prediction image to generate a set of feature maps, called reconstruction maps; - decoding at least part of a set of reconstruction parameters representative of at least one reconstruction neural network; - processing said reconstruction maps by said at least one reconstruction neural network to produce a decoded representation of said current image.

[0009] For the purposes of the invention, encoding, or "coding", means the operation which consists of representing a set of samples, or pixels, in a compact form carried for example by a digital binary train. Decoding means the operation which consists of processing a digital binary train to restore decoded samples.

[0010] By "sequence of images" is meant a plurality of ordered two-dimensional images, for example temporally in the case of a video. According to one example, the sequence corresponds to a scene. According to one example, the sequence corresponds to a set of predefined images, for example a fixed number, or, in the sense of the MPEG standards, a GOP (Group Of Pictures) comprising the images located between two images of the Intra-image type, also called "intra period". According to another example, the images can be views of the same scene represented in multi-views. According to another For example, the images can be a plurality of temporal and multi-view images (immersive video).

[0011] By "current image" we mean an image of the sequence. This image can cor respond to a time instant, or to several time instants, the data being for example grouped in an image of a size greater than that of the sequence.

[0012] By "prediction information" is meant any information useful for generating a prediction of the current image. Conventionally, this information can be inter prediction data (motion, reference images, etc.), intra prediction data (prediction direction), mode selection data (intra or inter mode or combined intra / inter or combination), one or more encoded prediction images, a latent representation of one or more prediction images in the form of encoded feature maps, etc. By "prediction image" is meant a prediction of the current image in the image domain or in the latent domain. In the first case, the image is made up of pixels, for example in (Y,U,V) or (R,G,B) format; in the second case, the image is made up of values ​​in the latent domain. It is also called a feature map.

[0013] By "correction information" is meant any information useful for correcting the prediction of the current image. This information may be correction coefficients, correction images, or a latent representation of correction images in the form of correction feature maps.

[0014] The "prediction images" are used to assist in the decoding of the correction information. The result of decoding the correction information using the prediction images is a representation of the current image to be coded or decoded in the latent domain, in the form of reconstruction feature maps. A "latent domain" is a representation space in which an image is composed of a set of feature maps, for example in two-dimensional form. It may also comprise a set of one-dimensional vectors, or a set of scalar values.

[0015] By "reconstruction neural network" is meant a neural network such as a convolutional neural network, a multi-layer perceptron, an LSTM (for "Long Short Term Memory" in English), etc. The neural network is defined for example by a plurality of layers of artificial neurons and by a set of activation, weighting and addition functions (for example, a layer can calculate y = f(Ax+b), where y and b are vectors of dimension N, x a vector of dimension M, A is a matrix of dimension MxN, and f is the activation function). Preferably the reconstruction neural network is convolutional, that is to say that it comprises at least one convolution layer. A plurality of such networks can be cascaded.

[0016] By "parameter of the neural network" is meant one of the values ​​which characterizes the neural network, for example a weight associated with one of the neurons (filter or convolution coefficient, weighting, bias, value affecting the operation of the non-linearity, etc.) At least part of this set is encoded by the encoder and decoded by the decoder. Another part can be obtained by other means, for example by reading from memory or by calculating parameters.

[0017] By “processing (of input values) by a reconstruction neural network” is meant the application of a function expressed by a neural network to the input values ​​to produce output values ​​representative of the samples of the current image to be decoded (reconstructed) (resp. encoded).

[0018] By "representation of the current image" we mean its representation in the image domain (representation in the form of pixels) or latent (representation in the form of feature maps).

[0019] By "performance measurement" is meant a measurement between at least one value of a sample to be coded and a decoded value of said sample. The measurement can evaluate for example a distortion, or a perceptual error. It can be carried out on a sample or a plurality of samples (for example, the current samples, or the current image, etc.) The measurement can also include a measurement of the flow rate, in particular associated with the coding of the neural network and / or the coding of the characteristic maps. The measurement can be a joint measurement between the flow rate and the distortion through their weighting. As is well known in the state of the art, the value of this measurement is generally minimized until a target or minimum value, or a predefined time, is reached.

[0020] By "sample" is meant a value taken from an image of the sequence. Sampling a signal produces a series of discrete values ​​called samples. In the case of an image signal, the sample is called a pixel, which can be, for example, a color pixel traditionally represented by a triplet of values, for example (R,G,B) or (Y,U,V), or by a single luminance value for a grayscale image. The position of the sample is identified by its abscissa (x) and ordinate (y) coordinates in the image.

[0021] Generally speaking, it is considered that the steps of a coding or decoding method should not be interpreted as being linked to a notion of temporal succession. In other words, the steps may be carried out in a different order from that indicated in the coding or decoding method, or even in parallel.

[0022] The coding method according to the invention carries out during its learning a construction of the coding parameters, from a sequence of input images, by training a reconstruction neural network on prediction and correction data sets, to obtain a faithful reconstruction of the image. Advantages The coding process encodes the reconstruction neural network. This can be optimized and adapted in complexity or quality, depending on the desired complexity and the desired ratio between bitrate and distortion. Similarly, the encoded and transmitted prediction and correction information can be optimized and adapted.

[0023] During training, or construction, or learning, the parameters of the neural network and the information to be coded are updated according to a performance measure, for example of the rate-distortion type. When the training is finished, that is to say that the performance measure obtained is satisfactory, the actual coding of the parameters of the reconstruction neural network as well as that of the prediction and reconstruction information can be carried out and the result of the coding stored or transmitted to the decoder. Advantageously, the training process therefore makes it possible to refine the parameters of the reconstruction neural network, as well as its input parameters, until an adequate representation in terms of performance is obtained, for example a desired balance between the rate generated and the distortion undergone by the input image being coded.Advantageously, the coding method according to the invention makes it possible to efficiently compress the signal.

[0024] Advantageously, the coding method is efficient since the prediction images are used to code (and decode) the correction images. The redundancy between the two signals is thus effectively exploited.

[0025] Advantageously, the decoding method is simple since it suffices to obtain the correction and prediction information and the reconstruction neural network to reconstruct a decoded version of the current image.

[0026] Advantageously, it is possible to design a transmission system which works image by image with low latency, each current image being decoded upon receipt of the parameters of the neural network and the associated prediction and correction information.

[0027] According to embodiments of the decoding method:

[0028] - The method comprises the steps of:

[0029] - decoding of motion information representative of the current image;

[0030] - obtaining at least one decoded reference image;

[0031] - compensation of said at least one reference image using the information of motion to produce said at least one prediction image.

[0032] Advantageously, according to this embodiment, the prediction image is obtained in the form of an image that is at least motion-compensated. The prediction image(s) (in the spatial domain or in the latent domain, in which case they may be called "feature maps") are then used to assist in the decoding of the in- correction training, in order to produce a latent representation of the image.

[0033] - the step of decoding said correction information comprises a sub-step of reconstruction of at least one value of said at least one reconstruction map, called current value, by entropic decoding of at least one piece of correction information, called correction data, as a function of at least one value of said at least one prediction image; advantageously according to this mode, the prediction images are used to assist the contextual decoding of the values ​​of the reconstruction maps. The redundancy between the two signals is thus effectively exploited.

[0034] - said at least one value of said at least one prediction image is located in a neighborhood of said current value. Advantageously, the decoding of the correction information is made particularly efficient by taking into account the neighborhood in the prediction information, which makes it possible to exploit the redundancies present in the data located at neighboring positions. This neighborhood takes, according to one embodiment, the form of a neighborhood vector. By "neighborhood vector", we mean a vector consisting of one or more elements, or data, constructed from the prediction images (pixel images or feature maps), and optionally from reconstruction maps already decoded. These data can be selected at coordinates close to the value to be reconstructed (within the sampling ratio, i.e. respecting the respective sizes of the prediction image and the image to be reconstructed). These data can be grouped in the neighborhood vector.The neighboring position may indicate a value in a prediction image (for example, the neighboring value of the one that is being reconstructed, at the same position in a prediction map). It may further indicate a value in the reconstruction map being processed (for example, the neighboring value at the top left of the one being processed). The neighborhood is associated with a context that allows a correction value to be decoded efficiently. According to another embodiment, the neighborhood is automatically defined by a convolution filter.

[0035] - the step of reconstructing at least one current value by decoding in- correction data tropic includes the following sub-steps: - creation of a neighborhood; - neighborhood processing to provide at least one statistical value; - decoding the correction data using said at least one statistical value

[0036] - the neighborhood processing step comprises the following sub-steps: - decoding at least part of a set of contextual processing parameters representative of at least one contextual processing neural network; - processing of the neighborhood by said at least one contextual processing neural network to produce said current value.

[0037] According to this embodiment, the neighborhood is the one that is applied to the input of the contextual processing neural network. According to one embodiment, it can be applied in the form of a neighborhood vector consisting of the values ​​selected in the neighborhood of the value being reconstructed. According to another embodiment, it is selected by a convolution kernel of a convolution layer. Thus, the correction feature maps are efficiently compressed by a contextual processing neural network specially trained to generate a value of the map according to its neighborhood by decoding the correction information in a contextual manner, the context being formed from the data of the prediction images. It is therefore able to represent them efficiently. It is also inexpensive to code.

[0038] - Said at least one reconstruction and / or processing neural network contextual has at least one convolution layer; thus, non-localized processing of the image can be performed, which improves the consistency of the generated image by limiting noise and enhancing the edges present in the image.

[0039] - Said at least one reconstruction and / or processing neural network contextual includes at least one MLP. An MLP (Multi Layer Perceptron) is a perceptron made up of several layers, an input layer adapted to the input format, optionally one or more hidden layers, and an output layer adapted to the output format of the output vector. Localized processing of the image, for example positional, can be carried out, which allows very simple and efficient coding of the image.

[0040] - Said at least one reconstruction and / or prediction neural network includes at least one attention module. This makes it possible to weight the characteristics of the image in order, for example, to adapt the allocation of the flow to the different regions and / or characteristics of the image according to their importance.

[0041] - At least part of the set of reconstruction and / or prediction parameters is obtained from data of said binary stream; advantageously according to this mode, the parameters of the neural network are transmitted in the stream. The reconstruction module can be transmitted in whole or in part in a quantized form and coded in a compact form using any quantizer and entropy coder accessible to those skilled in the art. A format similar to that of the MPEG-7 NNR standard can be used.

[0042] - At least part of the set of reconstruction and / or prediction parameters is decoded according to predetermined parameters; by "predetermined" we mean accessible to the decoder during decoding of the current image. Advantageously according to In this mode, part of the reconstruction parameters of the neural network are accessible to the decoder at the time when it decodes the current image without the need to decode information from the bitstream. According to one example, they may come from a previous decoding step (a previous image of the sequence, etc.). According to another example, they may be stored in a memory (storage memory, network, etc.). In particular, part of the contextual processing / reconstruction network may be stored in a quantized form and coded in a compact form using any quantizer and entropy coder accessible to those skilled in the art. A format similar to that of the MPEG-7 NNR standard may be used.

[0043] - At least part of the set of reconstruction and / or prediction parameters is decoded according to reference parameters; by "reference parameters" we mean parameters accessible to the decoder during the decoding of the current image and useful for decoding the other parameters. Advantageously, according to this mode, certain parameters of the neural network are decoded by taking into account reference parameters which may come from a prior decoding step, or stored in a memory (storage memory, network, etc.) accessible to the decoder. Thus, the storage space or the transmission rate on the network can be effectively reduced: certain parameters can be coded / decoded in a complementary manner (for example by updating a kernel, a layer, a bias, etc., of the network), others not be coded / decoded at all, because they are available in the reference parameters. The reference parameters can constitute a reference network. According to alternative embodiments, which can be combined with each other: 。 • the reference parameters can constitute a set of variants of a reconstruction network: type of convolution, attention modules, etc. Thus, only these parameters or their identification need to be coded, inserted into the flow and decoded by the decoder; • certain parameters are coded in a complementary manner to the parameters of the reference network. Thus, the decoder will simply have to decode these residual parameters and then add them, multiply them (or combine them in any other known way) with those of the reference neural network; • part of the parameters of the reference network is reused for the reconstruction network. For example, a complete layer can be copied from the reference network, the parameters of this layer are therefore neither encoded, nor transmitted, nor decoded; • a reference network indicator is transmitted in the stream. It is thus possible to indicate a network to be used during decoding from among a plurality of possible networks known to the decoder.

[0044] - Said at least one reconstruction and / or prediction network is selected from a plurality of reconstruction and / or synthesis networks, and the bitstream comprises an indicator indicating the selection of said at least one network of the plurality.

[0045] - Said at least one reconstruction and / or prediction network corresponds to a a plurality of cascaded reconstruction networks. Advantageously, the use of several cascaded networks makes it possible in particular to enhance filtering and improve the consistency of the generated image by limiting its noise and improving its contours. In addition, each network can be learned individually or frozen independently. According to alternative embodiments, which can be combined with each other: • the bitstream includes an indicator indicating the use of at least one of these networks. Thus, depending on the quality and complexity required, one or more networks can be used for decoding. The indicator indicates, for example, that network 1 is mandatory, but that networks 2 and 3, which are cascaded after network 1, are optional. The decoder can then choose whether or not to use them: if it uses them, it will obtain better quality at the expense of complexity, and vice versa. • At least two of the networks have common parameters. It is thus possible, advantageously, to transmit characteristic parameters of one network, which can be used for the others, which reduces the transmission cost. The network is thus shared through common parameters. • At least two of the networks have an identical structure and only their weights differ in part or completely. The same calculation elements are thus shared.

[0046] Correlatively, the invention also relates to a coding device and a decoding device.

[0047] The characteristics and advantages of the coding or decoding method apply in the same way to the coding or decoding device according to the invention and vice versa.

[0048] The invention also relates to a computer program on a recording medium, this program being capable of being implemented in a computer or a coding or decoding device according to the invention. This program comprises instructions adapted to the implementation of the corresponding method. This program can use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form. The program according to the invention can in particular be downloaded from a network such as the Internet.

[0049] The invention also relates to an information medium or a recording medium readable by a computer, and comprising computer program instructions mentioned above. The information or recording media may be any entity or device capable of storing the programs. For example, the media may include a storage medium, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording medium, for example a floppy disk or a hard disk, a DNA sequence, or a flash memory. On the other hand, the information or recording media may be transmissible media such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio link, by wireless optical link or by other means. Alternatively, each information or recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of a method according to the invention. Brief description of the drawings

[0050] Other characteristics and advantages of the present invention will emerge from the description given below, with reference to the appended drawings which illustrate exemplary embodiments thereof which are not limiting in nature.

[0051] [Fig-1] [Fig.l] schematically represents a decoding device used in the scope of the invention;

[0052] [Fig.2] [Fig.2] schematically represents a coding device used in the context of the invention;

[0053] [Fig.3] [Fig.3] illustrates a reconstruction feature map decoding module used in embodiments of the invention;

[0054] [Fig.4] [Fig.4] is a flowchart representing an example of a decoding method which can be implemented by the decoding device of [Fig.l];

[0055] [Fig.5] [Fig.5] is a flowchart representing a method for decoding reconstruction feature maps, which can be implemented by the decoding device of [Fig.l] during the decoding process of [Fig.4];

[0056] [Fig.6] [Fig.6] is a flowchart representing an example of a coding method which can be implemented by the coding device of [Fig.2];

[0057] [Fig.7] [Fig.7] illustrates a decoding method used in one embodiment of the invention. Description of the embodiments

[0058] The decoding device DECv of [Fig.l] decodes a subsequence of the sequence, proceeding image by image. At the end of the decoding, the V images of the sequence are decoded. A current subsequence to be decoded denoted Ivd comprises at least one current image to be decoded, each image comprising respectively a plurality of N samples. Subsequently, it is considered, without loss of generality, that the current subsequence comprises a single current image Iv. However, a current subsequence to be decoded may comprise, according to one embodiment, several images (which may for example be concatenated to form an image Iv of size greater than N).

[0059] The DECv decoder dedicated to the image receives as input the data necessary to decode it: - IPc encoded data corresponding to prediction information. This prediction information can be, for example: - motion information (motion field, reference frame numbers, etc.); - intra-image prediction information (prediction mode, prediction directions, etc.);

[0060] - signal distortion information, such as trans parameters geometric formations (translations, homotheties, etc.);

[0061] - prediction images;

[0062] - latent representations of prediction images;

[0063] - etc.

[0064] - FCc encoded data corresponding to coded correction information, generally, data aimed at correcting the prediction of the current image.

[0065] - the coded reconstruction parameters WRc of at least one neural network of reconstruction.

[0066] Other neural networks can be used, such as networks:

[0067] - contextual processing of latent reconstruction feature maps;

[0068] - synthesis of prediction information; - summary of correction information; - oversampling and / or processing of latent reconstruction feature maps;

[0069] - etc.

[0070] The DECV decoding module comprises, for a current image, a REF module for obtaining reference images, an IPD module for decoding prediction information, an FCD module for decoding correction information, an MREC module for reconstructing images from FMR reconstruction latent value maps, a NND module for decoding a neural network. According to one embodiment, DECV produces as output a current decoded image, denoted Ivd, comprising a plurality of decoded samples.

[0071] The IPD decoding module decodes IPc prediction information. According to one embodiment, it is a conventional decoder, for example of the JPEG type, or HEVC, etc. which produces at output at least one prediction image noted IP / FMP, where IP indicates an image in the pixel domain and FMP an image in the latent domain (a feature map). According to another example, it decodes quantized data corresponding to the images using an entropy decoder. According to one embodiment, it is a decoder that produces as output a motion field MV. The IPD module may comprise for this purpose a motion synthesis neural network.

[0072] The FCD decoding module decodes FCc correction information. For this purpose, it uses the IP or FMP prediction data, as will be detailed below. It generates as output one or more reconstruction feature maps in the latent domain. According to one embodiment, the FCD module may comprise for this purpose a contextual processing neural network, denoted ARM. According to one embodiment, the maps decoded by the FCD module, numbering NFR, are denoted FMR and are called reconstruction maps of the current image, or current reconstruction maps.

[0073] The optional REF module constitutes a set of NFR reference images, denoted IREF / FMREF, corresponding to one or more images previously coded then decoded from the sequence. Like the prediction images, they can take the form of images made up of pixels (in the image domain, pixelic) or feature maps (in the latent domain). They can be predetermined or their reference can be decoded in the stream in the form of an indicator. In one embodiment, a single reference image is selected. In one embodiment, several reference images are selected. In one embodiment, one or more reference maps are selected in the latent domain. In one embodiment, for example in the case where an intra image is decoded, no reference map is selected and the REF module is absent or it generates “neutral” images composed for example of constant values.

[0074] The optional WARP module performs motion compensation of the reference images from a decoded MV motion field, in order to obtain one or more compensated images, denoted IP if they are in the image domain, or FMP if they are in the latent domain. In the latter case, the prediction feature maps with the number of NFPs, are denoted FMP and are called prediction maps of the current image, or current prediction maps. The WARP module may be absent if the IPD decoder directly decodes the IP / FMP prediction images. In one embodiment, the WARP module produces the IP / FMP images with the data only from the IPD module, i.e. the optional REF module does not exist, or the WARP module does not take into account the reference images submitted to it.

[0075] The MREC module uses one or more IC / FMC reconstruction cards as input. and generates the reconstructed image. The MREC module may include an MLP. It may also include one or more convolution layers. A convolution layer may include convolutional elements, optionally including a residual structure or an attention module. An "attention module" means a layer of a neural network including an attention element. An attention element is an element that allows an attention mechanism to be applied to generate one or more masks that are used to weight (by multiplication or use of more complex functions) the characteristics of the image in order, for example, to adapt the allocation of the flow to the different regions and / or characteristics of the image according to their importance (contour, particular textures, etc.). The attention masks are applied to the inputs of the neural network (reconstruction maps). They are encoded and decoded as parameters of the neural network.Advantageously, it is therefore not necessary to use additional bits to encode the masks.

[0076] The coded information FCc, IPc and WRc is extracted from the BS stream which can be received on a communication network and / or stored in whole or in part in a memory accessible to the decoder.

[0077] The decoding device DEC can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be realized by the cooperation of the processor and computer program instructions stored in the aforementioned memory and designed to carry out the functionalities of the module concerned, in particular as described below, when these instructions are executed by the processor.

[0078] [Fig.2] schematically represents a coding device used in the context of the invention.

[0079] The coding device ENCv of [Fig.2] codes a current image or sub-sequence Iv of the sequence SV, proceeding image by image. At the end of the coding, the V images of the sequence are coded.

[0080] The ENCv coder receives as input a sequence of images including the image Iv and produces coded parameters as output. These parameters include:

[0081] - encoded data (IPc) corresponding to prediction information, intended for the generation of IP / FMP prediction images.

[0082] - FCc encoded data corresponding to correction information, intended to correct a prediction of the current image.

[0083] - the coded parameters WRc of at least one reconstruction neural network. As mentioned in the description of [Fig.l], other neural networks may be used, and in particular, as will be described later in one embodiment, the parameters of a contextual processing network of the character maps- latent reconstruction characteristics.

[0084] The ENCV coding module comprises, for a current image, an NNC module for coding neural network parameters, an FCC module for coding correction information, an IPC module for coding prediction information, an RD-OPT evaluation module, an INIT / MAJ initialization and update module, a REF module for generating and storing reference images.

[0085] Furthermore, in a conventional manner, the coding module comprises a decoder similar to the DECv decoder which was previously described, in order to reconstruct the current coded then decoded image, which is noted l'v.

[0086] During the process of training, or building, the coding, that is to say as long as the step of evaluating a performance is not satisfactory, the coding modules carry out a coding simulation, followed by a decoding, intended for the RD-OPT evaluation module. Subsequently, they carry out the actual coding of the data. In a known manner, the coding simulation can be identical to the actual coding, or carry out an approximation thereof.

[0087] The INIT / MAJ module is responsible for initializing and updating the values ​​of the images and the parameters of the neural network(s). The initialization data are denoted FCo, WRo and IPo. It updates the values ​​to be encoded in the current image, based on the results of the performance function. Once the prediction and correction information and the neural network(s) are stabilized, they can be encoded.

[0088] The RD-OPT module performs an evaluation and minimization of a coding performance. The evaluation function is for example of the rate-distortion type. The distortion can be evaluated between Iv and l'v, the rate can include the cumulative rate linked to the contributions of FCc, WRc and IPc, or only a sub-part. The minimization can be carried out by a gradient descent, or any other optimization method within the reach of a person skilled in the art.

[0089] The ENC coding device can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be realized by the cooperation of the processor and computer program instructions stored in the aforementioned memory and designed to carry out the functionalities of the module concerned, in particular as described below, when these instructions are executed by the processor.

[0090] [Fig.3] illustrates an example of a correction information decoding module used in decoding. A symmetrical module is used in encoding.

[0091] According to this embodiment, the FMD module comprises a contextual processing module, denoted MPS. The purpose of this module is to generate, from the information of coded correction and prediction images, a set of reconstruction maps corresponding to a representation of the image in the latent domain. More precisely, according to this embodiment, the module uses a context from the prediction images in the image (IP) or latent (FMP) domain to contextually decode the correction information. The result of such decoding provides a value (at least) of a reconstruction feature map, denoted Vn. A representation of a current neighborhood (for example a vector Cn from the prediction images and optionally from the reconstruction feature maps being decoded) is given as input to the module. The module provides statistical data as output which are used for the entropic decoding of the FCc value to provide a reconstruction value Vn in one of the reconstruction feature maps.In other words, the module models the law of appearance of a value Vn conditionally on its neighborhood.

[0092] As will be described later in support of [Fig.5], several embodiments are possible:

[0093] - use of an entropic (de)coder (ED) capable of predicting and then decoding a value depending on the context provided;

[0094] - use of an ARM neural network known as contextual processing, trained to predict a statistical value (e.g. probability, variance, etc.) based on context.

[0095] - etc.

[0096] In the illustration of [Fig.3], the decoding of the current value Vn located at the coordinates (xn, yn) in the current reconstruction feature map FMRi (where i represents the index of the current feature map, which can vary from 1 to NFR) uses contextual information from a neighborhood represented in gray. The neighborhood values, all available to the decoder, constitute the neighborhood vector Cn of the value Vn which can be used in one of the embodiments mentioned above and detailed in support of [Fig.5].

[0097] In this illustration, the decoding of a current FCc value to produce the value Vn being decoded located at the coordinates (xn, yn) in the current reconstruction feature map FMRi uses the contextual information of a prediction map FMPj previously obtained by the decoder: the values ​​located at the 9 coordinates (xn-k, yn-l), with k and 1 comprising the combinations of values ​​(-1,0,1) in the map FMPj are used to determine the neighborhood; according to this example, the (decoded) values ​​of the current map located at the 3 coordinates (xn-l, yn-l), (xn-l, yn), (xn, yn-l) are also used. This results in a vector Cn of 12 values.

[0098] The index j can be equal to the index i, that is to say that the feature maps of the same "rank" (same indices) are used in the prediction set and in the reconstruction. They can also be different, for example if we have a single IP image that is used to decode several FMRi cards.

[0099] In this embodiment, there are therefore as many FCc values ​​as Vn values ​​to be reconstructed. In other embodiments, there may be fewer FCc values ​​than Vn values ​​to be reconstructed.

[0100] In the illustrated embodiment, the neighborhood vector is extracted by a CTX module of the FMD decoding module, then it is applied to the input of the prediction MPS module. The MPS module may contain a contextual processing neural network ARM defined by its WAc parameters, in which case the neighborhood vector is applied for example to its inputs, and the network is used to estimate the statistical characteristics (p,o) or the probability (pr) of the FCc value to be decoded by the entropy decoder DE according to the context.

[0101] In another embodiment, a convolution kernel can be used to filter the values ​​of the FMPj map. This convolution kernel is for example that of a first convolution layer of the ARM neural network. This convolution kernel can only consider the available values ​​of the neighborhood vector, those which have been decoded.

[0102] [Fig.4] is a flowchart representing an example of a decoding method that can be implemented by the decoding device of [Fig.l].

[0103] The decoding described relates to a sub-sequence of images comprising at least one current image Iv of the sequence SV to be decoded.

[0104] During a step E30, the coded data stream is obtained. It can be received from a communication network, or read on a storage medium. In addition, certain information (for example, a reference reconstruction or contextual processing neural network) can be read in accessible memories of the decoder. The data obtained are at least the FCc correction information, IPc prediction information, and WRcd parameters of an MREC reconstruction network, optionally the WARc parameters of an ARM contextual processing network.

[0105] During a step E31, the prediction information IPc is decoded.

[0106] According to one embodiment, a pixel image IP is decoded. According to another embodiment, one or more FMP images are decoded in the latent domain (feature maps). Such an image can for example be decoded by a conventional technique within the reach of a person skilled in the art, for example a prediction followed by decoding by an entropic coder then dequantization, or it can come from a standard decoder (JPEG, HEVC, etc.)

[0107] According to one embodiment, a motion field is decoded. This motion field may be, for example, dense (at least one piece of motion information per pixel) or less precise (one piece of information then represents the motion information of several pixels). It can be decoded by a neural network, from decoded feature maps representative of the motion of the current image, or by any other known motion vector decoding technique. Conventionally, this motion field is used to compensate (warp) one or more IREF / FMREF reference images during a step E33 in order to obtain the motion-compensated IP / FMP prediction images. The IREF / FMREF reference images to be used are selected by the REF module in a predetermined or non-predetermined manner.

[0108] During a step E34, the reconstruction parameters WR of a reconstruction neural network of the MREC module are generated by decoding the WRc values. For this purpose, any known neural network decoding technique corresponding to that which was used by the encoder can be used, for example the neural network coding standard proposed by the MPEG-7 part 17, NNR standard.

[0109] According to one embodiment, the reconstruction parameters are received in the binary stream by the decoder. According to one embodiment, certain parameters of the reconstruction neural network are accessed in an accessible memory of the decoder. According to one embodiment, certain reference parameters of the reconstruction neural network are accessed in an accessible memory of the decoder. These different modes can be combined: for example, a part of the parameters is decoded from the stream, a predetermined part, obtained for example by copying certain parameters contained in memory, and a part inferred from a known model of a reference layer contained in memory.

[0110] According to one embodiment, an indicator received in the stream makes it possible to select one or more reconstruction networks from a plurality of networks accessible to the decoder.

[0111] During a step E35, the reconstruction characteristic maps are generated by decoding the FCc correction information. For this purpose, the coded FCc correction data are extracted from the stream, and decoded using the IP / FMP prediction images. Embodiments of this step are provided during the description of [Fig.5].

[0112] During a step E37, the reconstruction feature maps are then provided as input to a reconstruction neural network, for example a network composed of MLP and / or convolutional layers, to produce the final image Ivd. According to a variant represented in step E36, which will also be illustrated in [Fig.7], step E37 is preceded by a step E36 during which a feature vector is extracted from the set of feature maps and processed before being presented as input to the reconstruction neural network.

[0113] [Fig.5] is a flowchart representing a method of producing maps of reconstruction characteristics which can be implemented by the decoding device of [Fig.l] and by the decoding method illustrated in Figures 3 and 7.

[0114] These steps constitute sub-steps of step E35 described previously in support of [Fig.4]. Their purpose is to produce a current value Vnd of a reconstruction FMRi characteristic map by decoding the coded correction data FCc using values ​​of a neighborhood extracted from the prediction images FMPj / IPj.

[0115] During a sub-step E351, a neighborhood Cn is extracted, comprising values ​​neighboring the value Vn to be reconstructed in the reconstruction map. As illustrated in [Fig. 3], these neighboring values ​​are located in at least one prediction image, and optionally in an already reconstructed FMR map, or in the map being reconstructed. This neighborhood vector is made up of a number C of values, or data, corresponding to neighborhood values ​​(for example, C=12 in [Fig. 3]). These values ​​must be known to the decoder, they must therefore be located in a causal neighborhood of the value Vn. The neighborhood can also be selected automatically by a convolution kernel (for example, 9 values ​​can be presented as input to a convolution layer whose kernel is of size 3x3; the convolution kernel can also be non-rectangular, for example in the shape of a “gamma”, in order to consider only the upper and left neighborhood).

[0116] According to a first embodiment, these values ​​are used to determine during a step E352 the context of an entropy decoder to decode the current value during a step E354. This decoder can be a CAB AC (Context-Adaptive Binary Arithmetic Coding) type coder. This type of coder / decoder is well known to those skilled in the art. It is notably used in the H.265 / HEVC video compression standard. It is an arithmetic coder whose compression is lossless. During a first step, it decomposes all the non-binary symbols into binary symbols. Then, for each bit, the coder selects the most suitable probability model using a context to optimize the estimation of the probability. This context can be defined by information from the neighboring elements. Depending on the context, statistical information is generated (probability of 0 or 1).Binary arithmetic coding is then applied based on this statistical information to compress the resulting data. The reverse steps are used at the decoder.

[0117] As is known to those skilled in the art, there are several ways to use the neighborhood vector to produce context information. For example, one can count the number of neighboring values ​​other than zero, and associate a context with each number. Alternatively, one can perform comparisons between several neighboring values, and associate a given context with an order configuration between the neighboring values, for example by ranking neighboring values ​​in ascending order, and associating a context with each possible order.

[0118] In a second embodiment, a contextual processing neural network ARM is used in step E353 to predict statistical characteristics of the variable Vn to be decoded. The neighborhood is applied as input to the ARM network. According to one embodiment, the ARM network behaves like a function fv defined by a set of statistical parameters at output (mean, variance, median, etc.). These statistical parameters are used to decode the current value of FCc entropically. According to another embodiment, the contextual processing neural network ARM is used to produce the expected probability (pr) of the possible value of the current sample. The entropic decoding is then adapted to this probability (as is known for Huffman or arithmetic entropic coding) to decode a value FCc.

[0119] At the end of the decoding of step E354, the value decoded from the FCc stream using these statistics or this probability is injected into the reconstruction feature map at the position (xn, yn).

[0120] In another embodiment, not shown, the neighborhood is used to predict the current value Vn from an autoregressive model. In this mode, the past values ​​are constituted by the context, and the difference between the predicted variable and the actual value is entropically decoded during step E354 from the binary stream (FCc).

[0121] At the end of the process, all the current decoded values ​​V of the FMR reconstruction map being processed are decoded.

[0122] [Fig.6] is a flowchart representing an example of a coding method that can be implemented by the coding device of [Fig.2].

[0123] It is recalled that, during coding, the entire system is optimized during a learning phase which precedes the actual coding phase: the values ​​of the IP / FMP prediction images, the values ​​of the FCc correction parameters as well as the parameters of the MREC and optionally ARM neural networks are learned (in the sense of machine learning, for example by gradient descent) so as to optimize a performance measure, for example a cost function of the rate-distortion type noted D+ L*R, where D is the quadratic error of the reconstructed image Iv with respect to the original image, R is the sum of the coding rates of the characteristic images and optionally of the neural networks, and L a parameter set by the user for the RD-OPT module (Lagrange multiplier). In another example, D is calculated from a perceptual function such as SSIM (for Structural SIMilarity), or MSSSIM (for Multi-Scale Structural SIMilarity).

[0124] During a step E20, an input sequence SV to be coded, comprising at least one current image Iv comprising a plurality of N samples Pn, is provided as input to the method.

[0125] During a step E21, the prediction and correction information, as well as the parameters of the neural networks, are initialized. Subsequently, this data is optimized during the construction phase. According to one embodiment, the characteristic maps are initialized by predefined constant values. According to another embodiment, the characteristic maps are initialized by a set of random real values. According to one embodiment, the parameters of the neural network(s) are initialized by predefined values ​​known to give a satisfactory result (for example, following training on a corpus of image sequences). According to another embodiment, the parameters of the neural network are initialized by a set of random values.The correction information, images or prediction maps and parameters of the neural networks are subsequently updated, or refined, during a step E22, by the MAJ update module of the encoder during its learning.

[0126] During a step E23, a coding simulation is carried out for all the coding parameters: the prediction images are encoded by the IPC module, the correction information by the FCC module, the neural network(s) by the NNC module. The simulation may be identical to the actual coding but it may also be different (for example, simplified).

[0127] During a step E24, a decoding simulation is performed. To measure the distortion D, it is indeed necessary to simulate the coding then the decoding of at least part of the sequence of images, to obtain at least one pixel P'nd of at least one image l'v resulting from a simulation of coding then decoding of the samples of index n, then to measure the difference between these coded then decoded pixels and the pixels of the corresponding uncoded sequence. For this purpose, steps E31 to E37 of the decoding method are performed in order to reconstruct the current image or sub-sequence l'v decoded.

[0128] During a step E25, the performance measurement is evaluated. For this purpose, the coding simulation rates associated with the data to be encoded are calculated. According to one embodiment, the coding simulation rates associated with the parameters of the neural network(s) (WR coding of the MREC network and optionally WAR of the ARM network) are measured. According to one embodiment, the coding of the parameters of the neural networks is not simulated, because their influence is less important than that of the correction and prediction data.

[0129] As long as the cost function has not reached its minimum, the performance measurement is not satisfactory, and the method is repeated from step E22. Alternatively, the method can be interrupted after a predefined time or a number of predefined iterations, in order to control their complexity or duration.

[0130] During a step E26, if the cost function has reached its minimum or the time allocated to the encoding is deemed sufficient, the training stops. If a coded version corresponding to the last simulation of the parameters of the neural network and the correction and prediction information is available, the corresponding coding streams can be formed from it. According to another embodiment, the actual coding of the updated parameters of the neural network(s) and the values ​​of the characteristic maps is carried out at this step to produce the encoded parameters.

[0131] The encoded parameters may be stored, or transmitted, in a sub-stream or in different sub-streams which may be concatenated to produce a final BS stream. According to one embodiment, the sub-stream of the encoded parameters of the neural network is stored or transmitted before the sub-stream of the encoded prediction and correction information, so that it can be decoded before.

[0132] [Fig.7] illustrates an image decoding method used in one embodiment of the invention.

[0133] In this embodiment, the FMP prediction images are obtained in the form of feature maps, as described previously, from the IPc data of the BS stream, by conventional decoding, for example entropic decoding, possibly followed by warping of reference images.

[0134] In this embodiment, there are 4 current FMRi reconstruction maps. In a preferred embodiment, there are 7. In this embodiment, the first FMRI map has the same resolution as the image Iv to be decoded, and therefore comprises WxH values, where W represents the width of the image in pixels, and H its height. The second FMR2 map has half the resolution (in each dimension) of the FMRI map. Each additional map has half the resolution of the previous map. This structure makes it possible to reduce the number of values ​​of the characteristic maps, which facilitates decoding while minimizing the coding cost.

[0135] The FMR maps are made up of latent values. In the figure, a latent value Vn to be reconstructed is represented for the FMRI map. Such a value is obtained, as explained previously with reference to FIGS. 3 and 5, by decoding a value FCc extracted from the BS stream by an entropy decoder which takes into account a neighborhood in one or more of the IP / FMP images. In the embodiments for which the FCD module comprises an ARM contextual processing neural network, the WAR parameters of the ARM network are decoded from the BS stream.

[0136] In this embodiment, the reconstructed images are synthesized (at encoding and decoding) by an MREC neural network applied to FMS reconstruction feature maps processed by an XTR module. The XTR module optional may include one or more of the following steps / modules, shown in [Fig.7]:

[0137] - a step noted DQi (DQ1... DQ4) of dequantization of each card, the dequantization tification consisting for example of multiplication by a predefined or transmitted quantization step;

[0138] - a step noted SEi (SE1... SE4) of oversampling likely to put the FMRi map at the same resolution as the signal to be coded; according to the example shown, the FMR2 map is oversampled by a factor of 2 by the SE2 module in each dimension, according to any oversampling method within the reach of those skilled in the art. The FMR3 map is oversampled by a factor of 4 in each dimension, etc. According to one embodiment, it is possible to use an oversampling neural network.

[0139] - a step noted Ti (Tl... T4) of improvement processing, this step being able include convolutional filtering, a linear function, or even a neural network whose W1 parameters have been decoded in a classical manner (classical decoding or NNR) from the BS stream.

[0140] The steps thus described can be "mixed", in the sense that the oversampling can be based on a neural network. They can also be joint, that is to say applied to a set of maps and not to a single map.

[0141] According to this embodiment, these steps produce FMSi maps (FMS1... FMS4), which are at the same resolution as the signal to be coded, and therefore each comprise WxH values, where W represents the width of the image in pixels, and H its height. The transformed FMSi maps are processed by an MREC reconstruction neural network, typically a network comprising a Multi-Layer Perceptron, which associates with an input vector Zn consisting of the values ​​of the FMSi layers at the coordinates (xn,yn) of a pixel being processed an output vector (R,G,B) or (Y,U,V) at the coordinates (xn,yn) in a series of corresponding layers. These layers are for example the same number as the number of components of the output image (in figure 3). In this embodiment, the vector Zn is a 4-tuple (zb z2,z3, z4) consisting of the 4 values ​​of the FMS maps; located at the coordinates (x, y) of the current pixel Pn.The triplet (R, G, B) is inserted into each decoded image Iv of the subsequence v at the coordinates (xn, yn) in the corresponding color components.

[0142] According to another embodiment, the MREC module may comprise one or more convolution layers. In this case, the vector Zn is constructed by a convolution kernel moving on the FMS maps. It may comprise, for example, 9 values ​​in each FMS layer (for a convolution kernel of size 3x3).

[0143] The WR parameters of the MREC reconstruction network were decoded from the BS stream by conventional or NNR decoding. According to a variant not shown, additional correction feature maps are introduced. In such a variant, the vector Zn is of dimension greater than 4. Such an additional map typically comprises data that can assist the MLPC network in the task of image reconstruction. Thus, the added maps may be, but are not limited to, one or more of:

[0144] - a map including at each point the abscissa and / or the ordinate of this point;

[0145] - a map comprising at each point a positional coding;

[0146] - a map representing an image distinct from the images currently being processed, above capable of providing information about the images to be coded, for example a previously processed image or sequence of images;

[0147] - a map containing data representative of the time difference between the frames of the video being encoded. For example, if the first and last frames of the video are 8 frames apart, all samples in the map contain the value 8;

[0148] - a map representing a feature map of a distinct image of the images being processed, which may provide information on the images to be coded, for example a previously processed map;

[0149] - a card containing the value of an already decoded sample of the same card, by example the previous sample in the decoding order.

Claims

Claims

1. Method for coding a current image (Iv) at least from a sequence of images (SV) comprising the following steps: - initialization (E21) of current image prediction information (IP / FMP); - initialization (E21) of a set of parameters (WR) representative of a reconstruction neural network (MLPC); - generation (E23) of at least one prediction image from the prediction information; - coding (E23, E24) of correction information (FCc) using said at least one prediction image to generate a set of feature maps, called reconstruction maps; - processing (E23, E24) of said reconstruction feature maps by said at least one reconstruction neural network to produce a decoded representation of said current image; - updating (E22, E25) at least one of said correction and / or prediction information and / or at least one parameter of said reconstruction neural network, as a function of a coding performance measurement; - a step of coding (E26) a binary stream comprising: - coding of correction information (FCc); - coding of prediction information (IPc / FPc); - coding at least part of the set of parameters (WRc) representative of the reconstruction neural network (MLPC);

2. Method for decoding a current image (Idv) from an image sequence (SV), comprising the following steps: - obtaining (E30) coded correction information (FCc) of said current image; - obtaining (E30, E31, E32, E33) at least one prediction image (IP / FMP) of said current image; - decoding (E35) said correction information (FCc) using said at least one prediction image to generate a set of characteristic maps, called reconstruction maps (FMR); - decoding (E34) at least part of a set of reconstruction parameters (WR) representative of at least one reconstruction neural network (MREC); - processing (E37) of said reconstruction maps by said at least one a reconstruction neural network to produce a decoded representation of said current image (Ivd).

3. Decoding method according to claim 2, further comprising the following steps: - decoding (E31) of motion information (MV) representative of the current image; - obtaining (E32) of at least one decoded reference image (IREF, FMREF); - compensation (E33) of said at least one reference image (IREF, FMREF) using the motion information to produce said at least one prediction image (IP, FMP).

4. Decoding method according to one of claims 2 to 3, characterized in that the step of decoding (E35) said correction information comprises: - reconstruction of at least one value (Vn) of said at least one reconstruction map, called current value, by an entropic decoding of at least one data item of the correction information, called correction data item, as a function of at least one value of said at least one prediction image.

5. Decoding method according to the preceding claim, characterized in that said value at least of said at least one prediction image is located in a neighborhood of said current value.

6. Decoding method according to the preceding claim, characterized in that the step of reconstructing at least one current value (Vn) by an entropic decoding of the correction data comprises the following sub-steps: - constitution (E351) of a neighborhood; - processing (E352, E353) of the neighborhood to provide at least one statistical value; - decoding (E354) of the correction data using said at least one statistical value.

7. Decoding method according to the preceding claim, characterized in that the neighborhood processing step (E353) comprises the following sub-steps: - decoding (E34) of at least part of a set of contextual processing parameters (WAR) representative of at least one contextual processing neural network (ARM); - processing (E35) of the neighborhood by said at least one contextual processing neural network (ARM); contextual processing neurons (ARM) to produce said current value.

8. Decoding method according to claim 2 or 7, characterized in that the reconstruction or contextual processing neural network comprises a convolution layer.

9. Decoding method according to claim 2 or 7, characterized in that the reconstruction or contextual processing neural network comprises an MLP.

10. Decoding method according to claim 2 or 7, characterized in that the reconstruction or contextual processing neural network comprises an attention module.

11. Decoding method according to claim 2 or 7, characterized in that at least part of the set of reconstruction or contextual processing parameters is decoded according to predetermined parameters.

12. Decoding method according to claim 2 or 7, characterized in that at least part of the set of reconstruction or contextual processing parameters is decoded as a function of reference parameters.

13. Device (DEC, DECv) for decoding a current image (Idv) of a sequence of images (SV), said device being configured to implement, for at least one current image (Iv) of the sequence: - obtaining (FCD) coded correction information (FCc) of said current image; - obtaining (IPD, REF, WARP) of at least one prediction image (IP / FMP) of said current image; - decoding (FCD) of said correction information (FCc) using said at least one prediction image to generate a set of feature maps, called reconstruction maps (FMR); - decoding (NND) of at least part of a set of reconstruction parameters (WR) representative of at least one reconstruction neural network (MREC); - processing (MREC) of said reconstruction maps by said at least one reconstruction neural network to produce a decoded representation of said current image (Ivd).

14. Device for coding a sequence of images (SV), said device being configured to implement for at least one current image (Iv) of the sequence: - initialization (INIT / MAJ) of current image prediction information (IP / FMP); - initialization (INIT / MAJ) of a set of parameters (WR) representative of a reconstruction neural network (MLPC); - generation of at least one prediction image from the prediction information; - coding (FCC) of correction information (FCc) using said at least one prediction image to generate a set of feature maps, called reconstruction maps; - processing said reconstruction feature maps by said at least one reconstruction neural network to produce a decoded representation of said current image; - updating (INIT / MAJ) at least one of said correction and / or prediction information and / or at least one parameter of said reconstruction neural network, based on a coding performance measurement; - a step of coding a binary stream comprising: - coding (FCC) of correction information (FCc); - coding (IPC) of prediction information (IPc / FMPc); - coding (NNC) of at least part of the set of parameters (WRc) representative of the reconstruction neural network (MLPC).

15. A computer program comprising instructions for carrying out the steps of an encoding or decoding method according to claim 1 or 2 when said program is executed by a computer.

Citation Information

Patent Citations

  • Method and apparatus for content-adaptive online training in neural image compression

    US20220353521A1

  • Encoding with signaling of feature map data

    US20230336758A1