Method and apparatus for encoding and decoding image

By training synthetic neural networks and feature maps and utilizing the redundant information of images for efficient compression, the problem of high image coding complexity in existing technologies is solved, and efficient and high-fidelity image compression is achieved.

CN120641913APending Publication Date: 2025-09-12ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380093204.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-09
Filing Date
2023-12-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing image coding technologies find it difficult to strike a balance between compression efficiency and complexity. Although neural network-based methods have superior performance, they have high memory usage and computational complexity. Traditional methods are also difficult to adapt to diverse video formats and communication network requirements.

Method used

By training synthetic neural networks and feature maps, encoding parameters are constructed according to the input signal, the redundancy of the feature map is used for efficient compression, and the image values ​​are gradually predicted through the prediction neural network, achieving simple and efficient encoding and decoding.

Benefits of technology

It achieves the goal of reducing computational complexity while maintaining high compression efficiency, adapting to diverse video formats and communication network requirements, and providing high-fidelity image reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120641913A_ABST
    Figure CN120641913A_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for encoding and decoding a signal comprising a plurality of sample points. The decoding method comprises the following steps: decoding a first set of feature maps representing the signal. -decoding a set of parameters representing a neural network, referred to as a synthetic neural network,-for at least one sample point of the signal to be decoded, referred to as a current sample point, associated with a position in the signal to be decoded:-decoding, by means of the synthetic neural network defined by the decoded parameters, the current sample point of the signal to be decoded on the basis of said position of the feature vector; said feature vectors are constructed from the feature maps in said first set, thereby providing a vector representing decoded values of the current sample point,-decoding / encoding said first set of feature maps, comprising, for at least one value of one of said feature maps, referred to as the current value, entropy encoding said value based on values of at least its neighborhood.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the general field of encoding one-dimensional or multi-dimensional signals. More particularly, the present invention relates to compressing digital images. Background Art

[0002] Digital images are typically source-coded to achieve compression, thereby limiting the resources required for their transmission and / or storage. Numerous coding standards exist. For still images, there is the JPEG family of standards, while for moving images or video, there are standards developed by the ITU / MPEG organizations (H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc.) and their extensions (MVC, SVC, 3D-HEVC, etc.).

[0003] An image is typically encoded by dividing it into a number of rectangular blocks and encoding these blocks of pixels in a given processing sequence. In existing video compression techniques, the processing of the blocks typically involves block pixel prediction, which is performed using previously encoded and then decoded pixels present in the image being encoded (in this case, called "intra-frame prediction") or using previously encoded images (in this case, called "inter-frame prediction"). This exploitation of any spatial and / or temporal redundancy avoids the need to transmit or store the pixel values ​​of each pixel block by representing at least some of the blocks using a residual (which represents the difference between the predicted values ​​of the pixels in the block and the actual values ​​of the pixels in the predicted block).

[0004] As video formats continue to evolve to achieve higher compression rates and accommodate a diverse range of expected formats and communication networks, the prediction possibilities continue to grow, and conventional encoding and decoding algorithms become very complex.

[0005] In addition to these conventional methods proposed by compression standards (JPEG, MPEG, ITU), methods based on artificial intelligence (especially neural network intelligence) are also emerging.

[0006] Some of these neural network methods can be viewed as simple extensions of the competition concept in previous compression techniques (such as prediction mode competition, video coding transformation, etc.).

[0007] Other approaches use the concept of "autoencoders." Autoencoders are learning algorithms based on artificial neural networks that allow for the construction of new representations of a dataset. The autoencoder architecture consists of two parts: an encoder and a decoder. The encoder consists of multiple layers of neurons that process the data to construct new representations, called "encoded" representations, also known as "latent representations." The decoder's neural network layers, in turn, receive these representations and filter them to attempt to reconstruct the original data. The difference between the reconstructed and original data allows for the measurement of any error introduced by the autoencoder. Training involves modifying the autoencoder's parameters to reduce the reconstruction error measured at each sample point in the dataset. While these autoencoder-based systems offer superior performance, they come at the expense of significantly increased memory usage and complexity compared to conventional approaches, such as those proposed by compression standards. Such systems can have millions of parameters, and decoding a single pixel can require up to a million MAC (multiply-add) operations. This makes these decoders significantly more complex than conventional decoders, potentially hindering the application of learning-based compression.

[0008] A simple neural network-based image coding technique was recently described in the paper "Compression with Implicit Neural Representations" (arXiv:2103.03123) by Emilien Dupont et al. The proposed encoding technique involves adapting a neural network to an image, quantizing the network's weights, and transmitting the quantized weights. During decoding, the neural network is evaluated at each pixel location to reconstruct the image. However, this technique still suffers from compression inefficiencies.

[0009] Therefore, there is a need for a solution for simply and efficiently encoding / compressing an image or a sequence of images. Summary of the Invention

[0010] The invention is directed to a coding method according to claim 1 and a decoding method according to claim 8 .

[0011] Within the meaning of the present invention, the term "encoding" is understood to mean an operation involving representing a set of samples in a compact form, such as for transmission via a digital bit stream. Decoding is understood to mean an operation involving processing a digital bit stream to restore decoded samples.

[0012] The term "sample" of a signal should be understood to mean a value sampled from the signal. Sampling a signal produces a sequence of discrete values, called samples. In the case of image signals, a sample is called a pixel, which can be, for example, a color pixel traditionally represented by a triplet of values, such as (R, G, B) or (Y, U, V). The location of a sample is identified by its coordinates on the horizontal (x) and vertical (y) axes in the image.

[0013] The expression "a signal comprising a plurality of samples" should be understood to mean a signal having one dimension (audio, sound), two dimensions (images), or more than two dimensions (stereoscopic vision, multi-view images, images associated with depth maps, video, etc.). Depending on the dimensionality, the samples have one, two, or several coordinates in the signal. For image signals, the position of a sample is identified by its x and y coordinates.

[0014] The term "characteristic map" should be understood to mean an abstract representation of a signal, comprising a plurality of potentially discrete variable data, also known as values, real numbers or integers. As is well known, these maps are also known as "latent representations".

[0015] The expression "feature map transformation" should be understood to mean the application of a mathematical operation to transform the values ​​of a first map into the values ​​of a second map. The first map used for encoding (referred to as the first set of maps) can be any type of map. The second map (referred to as the transformed map or map in the second set) has the same resolution as the input signal, meaning it contains a number of values ​​equal to the number of samples (N) contained in the (correspondingly decoded) input signal. Transformation can involve, for example, interpolation, upsampling, filtering, quantization, Fourier transforms, and so on.

[0016] The expression "a feature vector of data constructed from a feature map according to a position" should be understood to mean a vector consisting of one or more (preferably discrete) elements or data, where the data is constructed from a feature map at a position determined by the position of the sample being processed in the signal. This feature vector is an input to the synthesis neural network. For example, in the case of a one-dimensional audio signal, such a vector can be constructed from a plurality of values ​​sampled from each of the feature maps at the same coordinates as the sample to be encoded. For an image, such a vector can be constructed from a plurality of values ​​sampled from each of the feature maps at the same x and y coordinates as the sample to be encoded (or respectively decoded). Once these values ​​have been sampled from the feature maps, they can be processed (e.g., by quantization, filtering, interpolation, etc.) to form a feature vector, which is then input to the synthesis neural network.

[0017] The term "synthetic neural network" should be understood to mean a neural network, such as a convolutional neural network, a multilayer perceptron, an LSTM (long short-term memory), etc. A neural network is defined, for example, by multiple layers of artificial neurons and a set of activation functions, weighting functions, and summation functions (for example, a layer may compute y = f (Ax+b), where y and b are N-dimensional vectors, x is an M-dimensional vector, A is an M × N-dimensional matrix, and f is an activation function).

[0018] The term “parameter of a neural network” is understood to mean one of the values ​​characterizing a neural network, for example a weight associated with one of the neurons (filter coefficients, weights, biases, values ​​influencing nonlinear operations), etc.

[0019] The expression "processing using a synthetic neural network" should be understood to mean applying a function expressed by the synthetic neural network to an input feature vector in order to produce an output vector representing the samples to be encoded (or respectively decoded). The output vector may include one or more data representing the samples.

[0020] The term "performance measure" is understood to mean a measure between at least one value of a sample to be encoded and the decoded value of said sample. This measure may assess, for example, distortion or perceptual error. This measure may be performed for one sample or for multiple samples (e.g., the current sample or the current image, etc.). This measure may also include a bitrate measure, in particular a bitrate measure associated with the encoding of the synthetic neural network and / or the encoding of the feature maps in the first group. This measure may be a joint measure of bitrate and distortion achieved by weighting the bitrate and distortion. As is well known in the art, the value of this measure is typically minimized until a target value is reached.

[0021] The term "construction step" is understood to mean a step aimed at constructing the parameters representing the image before actually encoding them. The construction sub-step may be repeated as many times as necessary to obtain an acceptable performance measure.

[0022] Generally, the steps of the encoding or decoding method should not be interpreted as being associated with the concept of a chronological order. In other words, these steps can be performed in a different order than indicated in the independent encoding or decoding claims, or even performed simultaneously.

[0023] The encoding method according to the present invention constructs encoding parameters based on an input signal (e.g., an image) by training a neural network (referred to as a synthetic neural network) on feature vectors associated with the locations of the samples to be encoded. These feature vectors are constructed from feature maps, which can have the same resolution as the input signal or a lower resolution. During training or construction, the parameters of the neural network and the values ​​of the feature maps are updated based on performance measurements (e.g., bitrate distortion type). When training is complete (i.e., the performance measurements obtained are satisfactory), the actual encoding of the parameters of the synthetic neural network and the values ​​of the feature maps can be performed and stored or transmitted to a decoder.

[0024] Advantageously, the training process allows for refining the parameters of the synthetic neural network and / or the values ​​of the feature maps until a representation that satisfies performance requirements is achieved, for example, until a desired balance between the generated bit rate and the distortion experienced by the input signal is achieved. The training of the values ​​of the feature maps and the training of the parameters of the synthetic neural network can be joint training. Advantageously, the encoding method according to the present invention allows for efficient compression of signals.

[0025] Advantageously, the decoding method is simple because only the feature maps and the synthetic neural network need to be decoded to reconstruct the decoded version of the signal (e.g., image).

[0026] Advantageously, the encoding of the feature map becomes particularly efficient by taking into account the coding neighborhood, thereby allowing to exploit the redundancy present in the map.

[0027] Advantageously, such a synthetic neural network can have a very simple structure with very few parameters.

[0028] Alternatively, decoding may be performed step-by-step, sample-by-sample.

[0029] According to an embodiment of the encoding or decoding method:

[0030] - the encoding method comprises the following sub-steps of encoding said current value of one of said characteristic maps:

[0031] - constructing a neighborhood vector based on the feature maps in the first group; and

[0032] - processing the neighborhood vector using an artificial neural network, called a predictive neural network, defined by a set of parameters, so as to provide a prediction of the current value;

[0033] - updating at least one parameter of the prediction network based on the coding performance measurement;

[0034] - Encoding the set of parameters of the prediction network.

[0035] The decoding method comprises the following sub-steps of decoding said current value of one of said feature maps:

[0036] - decoding a set of parameters representing a neural network called a predictive neural network;

[0037] - constructing a neighborhood vector based on the feature maps in the first group; and

[0038] - Processing the vector using a predictive neural network to provide a prediction of the current value.

[0039] Advantageously, according to this encoding or decoding mode, the feature map is efficiently compressed by a prediction neural network that can predict the value of the map based on its neighborhood. The term "neighborhood vector" should be understood to mean a vector consisting of one or more elements or data that are constructed from the feature map in a position close to the current sample position (which is also the position of the current value in the feature map). The neighboring position can indicate the value in the map being processed (for example, the neighboring value of the upper left corner of the map being processed) or the value in another feature map (for example, the neighboring value of the same position in the previous map). The neighborhood vector is an input to the prediction neural network. The term "prediction" should be understood to mean at least one data used to estimate the current value of the feature map, such as a probability, a statistical value, etc. The prediction neural network trained on the maps of the image can efficiently represent these maps. In addition, the encoding cost is low.

[0040] The method comprises the step of transforming said first set of feature maps so as to obtain a second set of feature maps having the resolution of the input signal, the method being characterized in that said feature vector is constructed from feature maps of said second set.

[0041] Advantageously, according to this embodiment, the feature maps are divided into two groups, one of which is reserved for extracting feature vectors and the other for encoding. Thus, two methods with different objectives can be separated: the maps in the first group to be encoded (and correspondingly decoded) must be compressed as efficiently as possible, while the maps in the second group must facilitate the process of extracting and constructing feature vectors.

[0042] According to one variant, at least one of the feature maps in the first set has a lower resolution than the resolution of the signal to be encoded (respectively, to be decoded), and the transformation operation involves upsampling. Advantageously, according to this embodiment, compression of the feature maps is more efficient because at least one of the feature maps in the first set to be encoded (respectively, to be decoded) contains fewer values ​​than would be the case if it had the resolution of the signal. For example, in the case of a digital image, one of the feature maps in the first set may have a 1 / 2 resolution, i.e., it contains half the number of x and y values ​​as the number of samples contained in the input signal (i.e., a total of 1 / 4 the number of values ​​of a feature map with the resolution of the signal). In contrast, a feature map in the second set (corresponding to the transformation of this map in the first set) has the same resolution as the signal. Therefore, in this case, the transformation includes at least one upsampling operation to obtain in the transformed map the same number of values ​​as the number of samples contained in the input signal (respectively, to be decoded).

[0043] - At least one of said feature maps of the first group has the same resolution as the resolution of the signal to be encoded (respectively to be decoded).

[0044] Advantageously, according to this embodiment, at least one of the feature maps has the same resolution as the input signal to be encoded (and respectively decoded), thereby achieving high fidelity and preserving the details of the original resolution of the signal. According to one embodiment, in this case, the transformation leaves the number of values ​​of the transformed feature map unchanged; the transformation can be simplified to an identity transformation (without any manipulation of the values ​​of the maps of the first group) or include filtering, quantization, Fourier transform operations, etc.

[0045] During the encoding step, if the feature maps include, for example, floating point values ​​or real values, quantization is crucial for the correct operation of the system. They need to be quantized before being encoded and / or provided as input to the synthesis and / or prediction neural network. In contrast, for decoding, depending on the embodiment, an inverse quantization operation is not necessary.

[0046] - Constructing said feature vector comprises the sub-step of extracting the value of said at least one feature map in the same position as the position of the current sample in the signal to be encoded (respectively to be decoded).

[0047] Advantageously, the elements of the feature vector can be constructed by extracting values ​​from feature maps in the first or second group at the same locations as the samples in the signal (input signal for encoding, signal to be decoded for decoding). This approach is easy to implement. For example, if J input feature maps with the same resolution as the signal are available, simply extracting the values ​​from these maps at the coordinates of the current sample (at the same horizontal and vertical coordinates in the feature maps) allows for the direct construction of a feature vector of J elements.

[0048] - The construction of the feature vector includes the following sub-steps:

[0049] - extracting a plurality of values ​​of the feature map in the first group according to the position of the current sample point;

[0050] - processing said extracted values ​​in order to obtain a feature vector.

[0051] Advantageously, according to this embodiment, feature vectors are extracted from feature maps before processing, which can be arbitrary and in particular have a resolution lower than that of the signal to be encoded (or respectively decoded). This processing can correspond to, for example, quantization, scaling, filtering, etc. of the extracted data. For encoding, if the feature maps include floating-point values ​​or real values, quantization is crucial for the correct operation of the system. They need to be quantized before being encoded and / or provided as input to the synthesis and / or prediction neural network. In contrast, for decoding, depending on the embodiment, an inverse quantization operation is not necessary.

[0052] - The method comprises the steps of constructing a third set of feature maps, and further constructing a feature vector from said feature maps.

[0053] Advantageously, these third sets of additional maps, constructed identically in the encoder and decoder, are neither stored in nor transmitted to the decoder, nor decoded in the decoder. Thus, these additional maps allow the use of additional data to improve compression without reducing the bit rate. This additional data may include, for example, coordinates, causal data available in the maps of the first or second sets, data about other images already processed by the encoder or decoder, etc.

[0054] Relatedly, another object of the invention is an encoding device and a decoding device.

[0055] The features and advantages of the encoding or decoding method are also applicable to the encoding or decoding device according to the present invention, and vice versa.

[0056] Another object of the present invention is a computer program on a storage medium, wherein the program can be implemented in a computer or control device according to the present invention. The program comprises instructions designed to implement the corresponding method. The program can use any programming language and can be in the form of source code, object code, or an intermediate code between source code and object code, such as in a partially compiled format, or in any other desired format.

[0057] The present invention also relates to a computer-readable information medium or storage medium comprising instructions for the aforementioned computer program. The information or storage medium may be any entity or device capable of storing a program. For example, the medium may include a storage device such as a ROM (e.g., a CD-ROM or a microelectronic circuit ROM), or even a magnetic storage device (e.g., a floppy disk, a hard disk, a DNA sequence, or a flash memory). Furthermore, the information or storage medium may be a transmissible medium such as an electrical or optical signal, which may be routed via an electrical or optical cable, a radio link, a wireless optical link, or other means.

[0058] The program according to the invention can in particular be downloaded via the Internet.

[0059] Alternatively, each information medium or storage medium may be an integrated circuit incorporating the program, the circuit being designed to execute or being used to execute the method according to the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Other characteristics and advantages of the invention will become apparent from the following description given with reference to the accompanying drawings, which illustrate embodiments of the invention that are by no means limiting.

[0061]

Figure 1

[0062]

Figure 2

[0063]

Figure 3

[0064]

Figure 4

[0065]

Figure 5

[0066]

Figure 6

[0067]

Figure 7

[0068]

Figure 8

[0069]

Figure 9

[0070]

Figure 10

[0071]

Figure 11

[0072]

Figure 12

[0073] Figure 1 An encoding device ENC is schematically shown.

[0074] The encoding device ENC includes a feature map generation module GEN, an additional feature map generation module FME, a transformation module SE, a data extraction module XTR, a processing and quantization module TT, a module MLP corresponding to a synthetic artificial neural network, a neural network encoding module NNC capable of encoding a synthetic neural network and an optional prediction neural network, a feature map encoding module FMC, a coding performance evaluation module EVAL, and an update module MAJ.

[0075] The encoding device ENC may be implemented by an electronic device comprising a processor and a memory (not shown); each of the aforementioned modules may then be generated via the interaction of the processor with computer program instructions stored in the aforementioned memory and designed, in particular as described below, to perform the functions of the module in question when these instructions are executed by the processor.

[0076] Figure 1 The encoding device ENC in the encoding device receives a series of samples to be encoded (denoted as P n , for example, a time-continuous sound sample) or a set of image data (denoted as I(Pn ))as input. In the second case, the image signal I(P n ) can represent a two-dimensional image, or multiple two-dimensional images (video, color components, stereoscopic components, multi-view components, etc.). n represents a sample n of an input signal comprising N samples. In one embodiment, the signal is a color image signal represented by at least one two-dimensional representation (such as a pixel matrix), where each pixel has a red component (R), a green component (G), and a blue component (B), or, as a variant, a luminance component (Y) and at least one chrominance component (U, V). The position of each pixel is defined by its x and y coordinates (x and y) in the image. In one embodiment, the image is grayscale and represented by a two-dimensional representation (such as a pixel matrix), where each pixel has either a grayscale component or a luminance component. In this case, the vector representing the pixel is reduced to a single component.

[0077] As will be referred to below Figures 3 to 7 Described in more detail:

[0078] The feature map generation module GEN is configured to generate multiple (M) feature maps, denoted as FM i The optional module FME can generate one or more additional maps (number L) that are neither encoded nor transmitted and are represented as FME l .

[0079] In one embodiment, the module SE processes the first set of feature maps FM i Transform to generate a second set of feature maps FMS with the same resolution as the input signal i .

[0080] The optional FME module can quantify the FM graphs from this set of M i The data extracted from or the vector Z formed by these data n . It should be noted that quantization of a value refers to matching that value to a member of a discrete set of possible code symbols. For example, this set of possible code symbols may consist of integer values, and the quantization system simply rounds the actual value to an integer value. According to another example, quantization involves multiplication by a given value and then rounding. Next, the module SE transforms the values ​​of at least one feature map, for example by upsampling, interpolation, filtering, etc. At the end of the transformation, the transformed feature maps of the second set have the same resolution as the images of the input sequence. Advantageously, according to this embodiment, the encoded feature maps may have a lower resolution than the resolution of the images to be encoded, while the maps of the second set used for constructing the feature vector have the same resolution as the image sequence, thereby facilitating the extraction of the values.

[0081] In one embodiment, the module SE is not present, in which case the values ​​to be used to construct the feature vector are extracted from the first set of feature maps.

[0082] The module XTR is for the current sample point P to be coded n , extract the feature map FM according to its coordinates in the input signal i (and / or FMSi and / or FME l , according to one of the previously described embodiments). For example, when it is intended to calculate the coordinates (x n ,y n ) at the sample point P n When encoding, the module XTR extracts the n ,y n ) determines the value in the graph at the position.

[0083] In one embodiment, the extracted values ​​form a vector Z n . Z n is a J-tuple, i.e. it contains J elements or data z i . The vector Z with index n n Refers to pixel P' n The eigenvector of .

[0084] In one embodiment, the optional module TT processes the extracted values ​​to generate a vector Z n Module TT can quantize the data extracted from the set of feature maps. The processing can include other operations such as filtering, scaling, etc. In particular, if module SE is not used and if the feature maps in the first set have a lower resolution than the images in the sequence, module TT can take into account the coordinates of the values ​​in the map with lower resolution.

[0085] It should be noted that at least one of the modules SE or TT has to quantize the feature maps.

[0086] The module MLP consists of K parameters W k The synthetic neural network defined here is able to process the vector Z n (or J-tuple) as input to generate the sample point P to be encoded n As output, the second vector of is taken as an output. According to one embodiment, the synthetic neural network is an MLP (Multi-layer Perceptron) consisting of an input layer adapted to the input format (J-tuples), optionally one or more hidden layers, and an output layer adapted to the output format of the output vector (typically a vector containing A elements). According to one embodiment, A is equal to 3, and the output vector is the encoded and then decoded pixel P' n The (R, G, B) triplet.

[0087] The module NNC is used to synthesize the neural network (especially its parameters W k ) is encoded. Optionally, the module NNC encodes the prediction neural network ARM (especially its parameters O b ) is encoded. During the encoding training or construction process, that is, as long as the step of evaluating the performance is still not satisfactory, the module NNC simulates the encoding and then the decoding, the result of which is sent to the evaluation module. Subsequently, it evaluates the synthetic neural network W k The actual encoding of the parameters of the prediction neural network ARM is performed, and the encoded parameters are denoted as Wc k and Oc b In a known manner, the coding simulation can be identical to or similar to the actual coding.

[0088] Module FMC to FM i , i.e., the values ​​of the feature maps in the first group (excluding any additional maps FME optionally upsampled by SE) l and the second group of graphs). During the encoding training or construction process, i.e. as long as the step of evaluating the performance is still not satisfactory, the module FMC simulates encoding and then decoding, the result of which is sent to the evaluation module. Subsequently, the execution of the graph FM i The actual encoding of the value of . The encoded graph is represented as FMc i In a known manner, the encoding simulation can be identical to or similar to the actual encoding. If necessary, the encoding module quantizes the latent representation of the values ​​of the maps in the first group using a quantizer to generate an ordered set of quantized values. The encoding module then compresses the quantized data using a coding that takes into account the neighborhood of the values ​​of the feature map to be encoded. As will be described below, the FMC module can include a predictive neural network (ARM).

[0089] The module EVAL performs evaluation and minimization of coding performance. For example, the evaluation function is of the bit rate distortion type. Minimization can be performed via gradient descent or any other method within the capabilities of those skilled in the art.

[0090] The module MAJ updates the graph FM to be encoded according to the results of the performance function i The value of .

[0091] Figure 2 A decoding device DEC is schematically shown.

[0092] Figure 2 The decoding device DEC receives as input a first set of encoded data organized into M feature maps FMc i (also called layer FM) and the encoded parameters Wc of the synthetic neural network MLP' k, and optionally the encoded parameters Oc of the synthetic neural network ARM' b .

[0093] The decoding device DEC comprises a neural network decoding module NND capable of decoding a synthetic neural network MLP' and optionally a prediction neural network ARM', a feature map decoding module FMD, a data extraction module XTR', an inverse transformation module SE', a processing and inverse quantization module TT', a module MLP' corresponding to the synthetic neural network, an additional feature map generation module FME'. According to one embodiment, it outputs a decoded image, denoted as I(Pd n ), including multiple decoded sample points Pd n .

[0094] The graph (number M) decoded by the module FMD is denoted as FMd i The parameters of the synthetic neural network (MLP') decoded by the module NND are denoted as Wd k The parameters of the prediction neural network (ARM') decoded by the module NND are represented as Od b .

[0095] The decoder module FME' can also generate one or more additional graphs, which are represented by FME' l , and the number is L, with the additional graph FME generated by the encoder l The same number.

[0096] In one embodiment, the module SE′ processes the first set of decoded feature maps FMd i Transformed to generate a second set of feature maps with the same resolution as the signal to be decoded, denoted as FMS' i . The module SE' optionally performs an inverse quantization corresponding to the quantization performed on the encoder. If the quantizer Q of the encoder only rounds the actual values ​​it receives, then there is no need to perform inverse quantization. If the neural network is able to take into account the quantization of its input data, then there is no need to perform inverse quantization. Otherwise, the decoder performs the inverse operation of the quantizer Q. The module SE' then performs a transformation on the values ​​of the feature maps similar to that performed by the encoder (including, for example, upsampling, interpolation, filtering, etc.). Once the transformation is completed, the transformed feature maps in the second group have the same resolution as the images of the sequence to be decoded.

[0097] In one embodiment, the module SE' is not present, in which case the values ​​to be used to construct the feature vector are extracted from the first set of feature maps.

[0098] Module XTR' with Figure 1 The module is the same as the module XTR. This module is for the sample point Pd to be decoded n, extract M feature maps FMd according to their coordinates in the signal to be decoded i (and / or FMS' i and / or FME' l , according to one of the previously described embodiments). In one embodiment, J = M. In one embodiment, J = M + L.

[0099] In one embodiment, the extracted values ​​form a vector Zd n . Zd n It is a J-tuple, that is, it contains J elements or data. i .

[0100] In one embodiment, the optional module TT' processes the extracted values ​​to generate a vector Zd n The module TT' may perform inverse quantization on the data extracted from the set of feature maps. The processing may include other operations similar to those performed by the encoder, such as filtering, scaling, etc.

[0101] The module MLP' is composed of K parameters Wd k The neural network defined is called a synthetic neural network and is able to process the vector Zd n (or J-tuple) as input to generate a representation of the sample point P to be decoded n The second vector (usually a vector containing A elements) is output. According to one embodiment, K = 3, and the output vector is the decoded pixel Pd n The (R, G, B) triplet of . Module MLP' has the same structure as module MLP, and if the parameter W k If the encoding is lossless, its parameters are the same, or if the encoding is lossy, its parameters are different.

[0102] When all the sample points P of the signal n When all have been decoded, the reconstructed signal I(Pd n ), for example, the image I includes N vectors Pd n N decoded samples in the form of .

[0103] The decoding device DEC may be implemented by an electronic device comprising a processor and a memory (not shown); each of the aforementioned modules may then be generated via the interaction of the processor with computer program instructions stored in the aforementioned memory and designed, in particular as described below, to perform the functions of the module in question when these instructions are executed by the processor.

[0104] Figure 3 An example of a synthetic artificial neural network for encoding and decoding in the context of an embodiment of the present invention is presented.

[0105] The synthetic artificial neural network MLP for encoding and the synthetic artificial neural network MLP' for decoding are defined by the same structure (eg, comprising multiple layers of artificial neurons) and a set of weights and activation functions respectively associated with the artificial neurons of the networks in question.

[0106] The vector representation of the current sample point (from the feature map FM on the encoder i / FMS i and FME l Or from FMd on the decoder i / FMS' i and FME' l The vector Z obtained n or Zd n ) is applied to the input (i.e., input layer) of the synthetic artificial neural network MLP or MLP'. According to one embodiment, the synthetic artificial neural network outputs a vector having color components (R, G, B) of color pixels forming an image.

[0107] All of these reconstructed pixels are stitched together into a (2D, 3D) image, forming the decoded or reconstructed image.

[0108] On the encoder, a synthetic artificial neural network MLP is trained on the image to minimize the input representation I(P n ) and its output representation I(P' n ), while also minimizing the amount of data to be encoded. In this sense, the module EVAL performs a performance measurement.

[0109] Once the encoder is trained, the network's parameters are either losslessly encoded (in which case the neural network MLP' is identical to the MLP) or lossily encoded (in which case the network MLP' may be slightly different from the MLP).

[0110] Figure 4 An example of a predictive artificial neural network for encoding (ARM) and decoding (ARM') feature maps in the context of an embodiment of the present invention is presented.

[0111] The predictive artificial neural network ARM for encoding and the predictive artificial neural network ARM' for decoding are defined by the same structure (eg comprising multiple layers of artificial neurons) and a set of weights and activation functions respectively associated with the artificial neurons of the network in question.

[0112] The vector representation of the current sample point (from the feature map FM on the encoder i / FMS i and FME l Or from FMd on the decoderi / FMS' i and FME' l The obtained vector C n or Cd n ) is applied to predict the input (i.e., the input layer) of the artificial neural network ARM (for encoding) or ARM' (for decoding).

[0113] The prediction neural network is represented as a function that outputs a prediction of the current value of the feature map being processed, which can be in the form of a predicted value or probability data.

[0114] According to one embodiment, at the encoder, the network implements the function f ѱ , this function is to encode the graph FMi i The current value V n The current value of provides the expected mean and / or variance (µ, σ). These statistics are used to perform entropy coding of the value. For example, if the function produces a mean, then the mean is subtracted from the current value and only the difference is entropy coded, and the mean is considered a prediction of the current value. Alternatively, if the function produces a mean and a variance, then the mean is subtracted from the current value and the difference is encoded using an entropy coding appropriate to the variance, for example, by quantizing the variance to a set of predetermined variances and associating a type of entropy coding with each quantized variance value. At the encoder, the network implements the function f ѱ , this function is to decode the image FMd i Current value Vd n The current value of provides an expected mean and / or variance. These statistics are used to perform entropy decoding of the value. For example, if the function produces a mean, the current value is decoded by the entropy decoder and the mean is added to the current value. Alternatively, if the function produces a mean and a variance, the current value is decoded by the decoder using entropy decoding appropriate to the variance, for example, by quantizing the variance into a set of predetermined variances and associating a type of entropy decoding with each quantized variance value.

[0115] According to another embodiment, the neural network can generate an expected probability (pr) for each possible value of the current sample. In this case, entropy coding or decoding will adapt to this probability (such as the well-known Huffman or arithmetic entropy coding).

[0116] At the encoder, the prediction artificial neural network ARM is trained on the image to minimize the amount of data to be encoded. In this sense, the module EVAL performs a performance measurement. It should be noted that the overall performance measurement involves minimizing the number of encoded and then decoded images I(P') while minimizing the decoding rate. n ) and the input image I(P n). According to one embodiment, the feature map is losslessly encoded via entropy coding. In this case, encoding the feature map affects the bit rate but does not affect the distortion of the encoded image. According to another embodiment, if the feature map is lossily encoded, encoding the feature map affects not only the bit rate but also the distortion.

[0117] Once the training is completed, the network's B parameters Oc b is either losslessly encoded (in which case the neural network ARM' is identical to ARM) or lossily encoded (in which case the network ARM' may be slightly different from ARM).

[0118] Figure 5 It shows that Figure 1 Flowchart of an example of an encoding method implemented by an encoding device.

[0119] According to this embodiment, the signal is a two-dimensional image, so each sample point to be encoded has a coordinate (x n ,y n ) pixel P n .

[0120] Coding is done in two main stages:

[0121] In the first phase, called the construction phase, learning is performed so that for the input signal I(P n ) Determine the FM i The value and parameter W k and optionally O b , to optimize the total cost function. For example, via gradient descent, the parameters of the synthetic neural network MLP and the feature map FM are then updated i The cost function can be of the bitrate-distortion type, or of the bitrate, or distortion, or perceptual type. To measure the bitrate R, it is necessary to simulate the image FM. i The encoding of , then the associated encoding bit rate (the size of stream B1) needs to be measured. According to one embodiment, the parameter W is not simulated. k and / or O b The encoding of , because their influence is smaller than that of the feature map. According to one embodiment, the parameter W is also simulated. k and / or O b To measure the distortion D, it is necessary to simulate the encoding and then decoding of at least a portion of the image I in order to obtain at least one pixel P' resulting from the simulation of the encoding and then decoding. n , then, we need to measure the input image I(P n) with the coded and then decoded image I(P' n ) between the corresponding parts.

[0122] Next, during the second phase, called the encoding phase, the graph FM i and parameter W k and optionally O b Encoded to produce the encoded value FMc i and WC k (and optionally Oc b ), and then transmitted or stored. They form the input signal I(P n ) is a compressed representation of .

[0123] The steps of a method according to one embodiment of the present invention will now be described.

[0124] During step E20, a plurality (N) of sample points P will be included n The signal to be coded I(P n ) is passed as input to this method.

[0125] During a step E21 , a first set of M graphs FM is initialized i Subsequently, the parameters W of the synthetic neural network MLP must be optimized during the construction phase k Hetu FM i and optionally predict the parameters of the neural network O b .

[0126] According to one embodiment, FIG. FM i With the input signal I(P n ) and therefore each picture contains the same resolution as the sample points P to be coded n The number of values ​​is the same as the number N.

[0127] According to one embodiment, FIG. FM i The resolution is less than or equal to the input signal I(P n ) and therefore for at least one of the maps the number N' of values ​​to be coded is less than N. According to one variant, the first map FM i has the resolution of the image, and each subsequent image has the Figure 1 Half resolution.

[0128] According to one embodiment, a plurality of graphs FM i With the same resolution, it is smaller than the input signal I(P n ) resolution.

[0129] According to one embodiment, FIGFM iTransformed to provide a second set of transformed feature maps FMS i In this embodiment, the feature vectors are preferably extracted from the transformed graphs in the second group, rather than directly from the graphs in the first group. Thus, in this embodiment, the feature vectors are indirectly extracted from the graphs in the first group. The graphs in the second group are not encoded; they are only used to construct the feature vectors.

[0130] According to one embodiment, the map FM is initialized using predefined constant values. i .

[0131] According to another embodiment, the feature map is initialized using a set of random real values.

[0132] According to one embodiment, FME generates one or more graphs l And added to the first set, these one or more maps form another set of L additional feature maps. They are used to construct the feature vector but are not stored or transmitted.

[0133] Subsequently, during a step E22 , the updating module MAJ of the encoder updates or refines the feature maps FM of the first group during its learning phase i .

[0134] During a step E23, the map FMi of the first group is processed by the module FMC of the coder. i During the construction phase, this operation is a simulation of encoding. During the encoding phase, this operation is the actual encoding, and the encoded values ​​form the stream B1. The simulation can be the same as the actual encoding, but can also be different (e.g. simplified). For this encoding, a technique is used to predict the value of the feature map via its neighborhood, e.g. by using the reference Figure 9 In one embodiment, the structure and parameters of the prediction neural network are b For example, they are initialized during the first iteration of this step. These parameters are then updated or refined during subsequent iterations of the method during the construction phase.

[0135] In one embodiment, FIGFM i The encoding is performed in the order (FM1, FM2, ..., FM4), with the variables of each graph encoded in a predefined order (e.g., lexicographic order). Each graph undergoes entropy encoding. Entropy encoding produces a compressed stream B1, the bit rate of which is subsequently measured during step E29.

[0136] During a step E24, according to one embodiment, the module SE processes the M maps FM of the first group i Transformed to generate a second set of images FMS with the image resolution of the input sequence i .

[0137] According to one embodiment, M graphs FMS are generated i .

[0138] According to one embodiment, each map FM i Transformed into graph FMS i .

[0139] According to one embodiment, at least one map FM i has a lower resolution than the resolution of the image of the signal to be coded, and the transform operation includes upsampling so that the transformed image FMS i Includes the same number of samples as the image in the sequence. Upsampling involves adding values ​​to the image FMS i , so as to achieve the resolution of the images of the input sequence. This operation can be simple (by nearest neighbor replication) or include interpolation (linear, polynomial, filtering, etc.).

[0140] During a step E25, the transformed map FM is extracted by the module XTR i , or optionally FMS i , and optionally additional diagrams FME i The extraction is based on the sample point P of the input signal. n The coordinates (x n ,y n This extraction can also be performed depending on the resolution of the graph in question.

[0141] According to one embodiment, the feature vector Z n directly from this extraction.

[0142] The samples to be coded are processed in order from n=1 to n=N, for example.

[0143] According to one embodiment, during a step E26, the module TT determines the coordinates (x n ,y n ) for each sample point P n , according to Figure FM i or FMS i and optionally FME l Extract the values ​​to construct the feature vector Z n If necessary, this processing may involve processing the i Extract the value or construct the vector Z n This processing may include other operations such as filtering, scaling, applying any function (preferably a monotonic function), etc.

[0144] In one embodiment, Z n Included with input graph FM i or FMSi (and optionally FME l ) is the same number of values ​​as the number of values ​​in the median. In this case, J = M(+L).

[0145] In one embodiment, Z n is located at the current pixel P n The coordinates (x n ,y n ) at Figure FM i or FMS i (and optionally FME l ) values ​​form a J-tuple (z1, z2, ..., z J ), such as referring to Figure 6 What is shown.

[0146] In one embodiment, Z n It is from Figure FM i (and optionally FME l ) is constructed from the values ​​sampled in the graph, where the sampling coordinates of these values ​​vary depending on the graph. For example, if the graph FM i (and / or FME l ) Since they have been downsampled and have different resolutions, the coordinates are adapted (by scaling) to match the resolution of each figure.

[0147] In one embodiment, Z n is a J-tuple constructed from values ​​that are retrieved from the graph FM by applying a process to one or more values ​​of the graph (e.g., filtering the neighbors of the target value in the graph). i (and optionally FME l ) is obtained by sampling. For example, for a graph FM with the same resolution as the input signal i The current sample point P in n , we can extract the n ,y n )、(x n -1,y n )、(x n ,y n -1) and (x n -1,y n -1) and process these values ​​(filtering, averaging, interpolation, etc.) to obtain the value of FM i or FME l Related vector Z n The final value of element i (z i According to another example, in the image FM having half the resolution of the input signal i In the example, we can consider the coordinates (x n / 2,y n / 2)、(x n / 2-1,y n / 2)、(x n / 2,y n / 2-1) and (x n / 2-1,y n / 2-1) and process these values ​​(filtering, averaging, interpolation, etc.) to obtain the value of FM in this figure. i or FME l Related vector Z n The final value of element i (z i ).

[0148] During step E27, the vector Z n Processed by the synthetic neural network MLP to generate a representation of the sample point to be encoded P n The vector is output, that is, according to one embodiment, (sample point P n The sample point P' is encoded and then decoded n The (R, G, B) triplet.

[0149] The structure and parameters W of the synthetic neural network k For example, they are initialized during the first iteration of this step. These parameters are then updated or refined during subsequent iterations of the method during the construction phase.

[0150] According to one embodiment, the parameters of the synthetic neural network and / or the predictive neural network are initialized with predefined values ​​that are known to produce satisfactory results (e.g., after training on a library of images).

[0151] According to another embodiment, the parameters of the synthetic neural network and / or the predictive neural network are initialized with a set of random values.

[0152] During a step E28 , the parameters W of the synthetic neural network MLP are adjusted k and the parameters O of the prediction neural network ARM b (if present) are quantized and encoded. During the construction phase, this operation is a simulation of encoding. During the encoding phase, this operation is the actual encoding, and the encoded values ​​form the stream B2. The simulation can be identical to the actual encoding, but can also be different (e.g., simplified). For this purpose, any known technique can be used, such as the Neural Network Coding standard proposed in Part 17 of the MPEG-7 standard, also known as Neural Network Representation or NNR. It should be noted that in this case, the coding pair weights W need to be chosen k and optionally O b The amount of degradation caused.

[0153] During a step E29 , the performance measures are evaluated.

[0154] To this end, the measurements are made on the feature maps in the first group (by comparing the FM i is encoded to simulate the flow B1) and optionally with the parameters of the neural network (via the parameters W k and optionally O b The encoding is performed to simulate the encoding simulation bit rate associated with stream B2).

[0155] According to one embodiment, the cost function is of the bitrate-distortion type, expressed as (D+L*R), where D is, for example, the squared error measured between the input signal and the decoded signal (or the error measured on a subset of the signal's samples). According to another example, D is calculated from a perceptual function such as SSIM (Structural Similarity) or MSSSIM (Multi-Scale Structural Similarity). According to one embodiment, R is the simulated bitrate of stream B1; according to another embodiment, R is the total bitrate used to encode the image, i.e., the sum of the simulated bitrates of B1 and B2. L is a parameter that adjusts the bitrate-distortion tradeoff. Other cost functions may also be used.

[0156] As long as the cost function has not reached its minimum, the performance measure is not satisfactory and the method is repeated starting from step E22. This minimization can be performed via a known mechanism such as gradient descent, wherein the parameters are updated by updating the values ​​of the feature maps during step E22 and updating the parameters of the network or networks during steps E23, E27.

[0157] During step EF, if the cost function has reached its minimum, the training stops. k ) and feature maps (FM i ) is available, the streams B1 and B2 can be formed from this. According to another embodiment, during this step the updated parameters of the synthetic neural network (W k ) and feature maps (FM i ) values, and optionally predicting the neural network ( b ) to produce the coded parameters Wc that form the streams B1 and B2 k (Optionally b ) and FMc i Streams B1 and B2 may be concatenated to produce the final stream. According to one embodiment, stream B2 of encoded parameters of one or more neural networks is stored or transmitted before stream B1 so as to be able to be decoded before stream B1.

[0158] Figure 6 A schematic diagram showing a decoding method used in an embodiment of the present invention is shown.

[0159] In this embodiment, there are 4 generated graphs FM i In the preferred embodiment, there are 7 graphs.

[0160] The first map FM1 has the same n ) and therefore contain W × H variables, where W represents the width of the image (in pixels) and H represents its height. The second map FM2 has half the resolution (in each dimension) of map FM1. The resolution of each additional map is half that of the previous one. This structure allows the number of variables in the feature map to be reduced, thereby facilitating encoding and learning while minimizing encoding costs.

[0161] According to the reference Figure 6 In the presented method, module SE upsamples the image FM2 by a factor of 2 in each dimension, upsamples the image FM3 by a factor of 4 in each dimension, and upsamples the image FM4 by a factor of 8 in each dimension.

[0162] The obtained graph FMS i With the same image I(P n ) and therefore each image consists of W × H values, where W represents the width of the image in pixels and H represents its height (N = W × H).

[0163] According to this embodiment, the layer FMS i Quantization is performed by module SE.

[0164] Other types of structures are possible, for example, using a drop rate other than half between graphs (such as one quarter, one third, etc.).

[0165] In the variant shown with dashed lines, there are five feature maps: an additional map FME0 has been introduced, which is neither encoded nor transmitted. This additional map typically contains data that can assist the network MLP in performing the task of reconstructing the signal. Thus, the additional map can be one or more maps from the following non-limiting list:

[0166] - A plot that contains at each point the x-coordinate of that point.

[0167] - A plot that contains at each point the y-coordinate of that point.

[0168] - A diagram containing a positional code at each point (e.g. as described on the following website: https: / / skosmos.loterre.fr / P66 / fr / page / -K0D65X2X-X).

[0169] - A map representing an image that is different from the one being processed and that can provide information about the image to be encoded, for example, previously processed images if the current image forms part of a sequence of images to be encoded (such as a video, a set of medical images, a multi-view representation, etc.).

[0170] - A map representing a feature map of an image different from the one being processed and which can provide information about the image to be encoded, for example, a previously processed map if the current image forms part of a sequence of images to be encoded (such as a video, a set of medical images, a multi-view representation, etc.).

[0171] - a picture containing the values ​​of an already decoded sample of the same picture (i.e. the previous sample in decoding order).

[0172] In this embodiment, vector Z n is obtained from the pixel at the current pixel P n The coordinates (x n ,y n ) at the FMS i The extracted values ​​form a 4-tuple (z1...z4). From the graph FMS i The vector Z of the (quantized) extracted values n Processed by the synthetic neural network MLP to output a second vector, for example, the output vector is the encoded and then decoded pixel P' n The (R, G, B) triplet is inserted into the decoded image I(P' n ) in the color components (R', G', B') position (x n ,y n ) place.

[0173] In another embodiment (not shown), the image is directly retrieved from the layer FM at a position recalculated according to the size of the image. i Extract vector Z n , and then optionally after extraction, the extracted values ​​are processed and quantized.

[0174] According to the variant shown as the dotted line, the vector Z n It is a 5-tuple (z0...z4), where the value z0 is extracted from the attached map fme0.

[0175] Figure 6 It shows that Figure 2 A flowchart of an example of a decoding method performed by a decoding device.

[0176] During a step E30, the streams B1 and B2 are extracted from the coded stream. These streams respectively contain a first set of images FMc i and parameter Wc k , optional parameter Ocb The encoded representation of .

[0177] During step E31, by i Decode to generate M graphs FMd i For this decoding, a technique is used to predict the value of a feature map via its neighborhood, e.g. Figure 10 In one embodiment, FIGFMd i Decoding is performed in the order of (FMd1, FMd2, ..., FMd4), and the values ​​of each map are decoded in a predefined order (e.g., lexicographic order).

[0178] According to the embodiment as described for the encoder:

[0179] - Figure FMd i With the signal to be reconstructed I(Pd n ) same resolution, i.e., the graphs contain N values.

[0180] - Figure FMd i The resolution is less than or equal to the signal to be reconstructed I(Pd n ) resolution.

[0181] - Multiple FMd i has the same resolution, which is lower than the resolution of the signal.

[0182] During a step E32, according to one embodiment, one or more graphs FME' are generated l And added to the first set, this one or more maps form another set of L additional feature maps. These maps are not decoded, but are generated by the decoder in the same way as they were generated in the encoder. These maps usually include data that can assist the network MLP's task of reconstructing the signal. Figure 6 The non-limiting list of possible additional feature maps described for the encoder applies in this case as well.

[0183] During a step E33, according to one embodiment, the module SE' processes the M maps FMd of the first group. i Transformed to generate a second set of images FMS' with the resolution of the input image i This step is similar to the reference Figure 5 Step E24 described for the encoder is similar and these embodiments are applicable. In particular:

[0184] According to one embodiment, M graphs FMS' are generated. i .

[0185] According to one embodiment, each map FMd i Transformed into graph FMSi .

[0186] According to one embodiment, at least one map FMd i has a lower resolution than the image resolution of the image to be coded, and the transform operation includes upsampling, so that the transformed image FMS' i Contains the same number of samples as the input image. Upsampling involves adding values ​​to the image FMS' i , so that the resolution of the input image is achieved. This operation can be simple (nearest neighbor copying) or include interpolation (linear, polynomial, filtering, etc.).

[0187] If necessary, the transform may optionally include inverse quantization performed on the extracted values. However, inverse quantization is not mandatory.

[0188] During a step E34, the transformed map FMd is extracted by the module XTR' i , or alternatively FMS' i , and optionally additional diagrams FME' i The extraction is based on the sample point P of the input signal. n The coordinates (x n ,y n ). This extraction can also be performed depending on the resolution of the graph in question. This step is similar to the reference Figure 5 Step E25 described for the encoder is similar and these embodiments are applicable. In particular:

[0189] According to one embodiment, the feature vector Zd n directly from this extraction.

[0190] In one embodiment, Zd n is located at the current pixel P n The coordinates (x n ,y n ) at the FMdi or FMS' i (and optionally FME' l ) values ​​form a J-tuple (z1, z2, ..., z J ), such as referring to Figure 8 What is shown.

[0191] The samples to be decoded are processed in order from n=1 to n=N, for example.

[0192] According to one embodiment, during step E35, for the coordinates to be decoded (x n ,y n ) for each sample point Pd n , module TT' is based on the coordinate (x n ,yn ) According to Figure FMd from the first group i Or from the second group of Figure FMS' i Extract the values ​​and optionally append them to the graph in FME' i Extract the values ​​to construct the vector Zd n This step is similar to the reference Figure 5 The step E26 described for the encoder is similar and the described embodiments are applicable. If necessary, the extraction may include the extraction of the value or the vector Zd formed. n Perform dequantization.

[0193] During a step E36 , the parameters Wd of the synthetic neural network MLP′ are k and optionally predict the parameters Od of the neural network b is the value Wc of convection B2 k and Oc b The synthetic neural network MLP' is similar to the synthetic neural network MLP, i.e., it has the same structure and includes the same parameters, except for the encoding, which can be lossy or lossless. Similarly, the prediction neural network ARM' (if used to decode the feature map) is similar to the prediction neural network ARM, i.e., it has the same structure and the same parameters, except for the encoding, which can be lossy or lossless.

[0194] According to one embodiment, stream B2 is decoded before stream B1 in order to be able to obtain the synthesis neural network MLP' and the optional prediction neural network ARM' before starting to decode the samples.

[0195] During a step E37 , the vector Zd is processed by the synthetic neural network MLP′ n To generate the sample point Pd to be decoded n The second vector is outputted, according to one embodiment, at the position (x) of the color component (Rd, Gd, Bd) n ,y n ) is injected into the decoded image I(Pd n ). This step is the same as reference Figure 5 Step E27 described for the encoder is similar.

[0196] When all samples of the signal have been processed, we obtain an image corresponding to, for example, I(Pd n ) corresponds to the decoded signal.

[0197] Figure 8 A schematic diagram showing a decoding method used in an embodiment of the present invention is shown.

[0198] In this embodiment, there are 4 decoded maps FMd i In the preferred embodiment, there are 7 graphs.

[0199] In this embodiment, the first map FMd1 has the same resolution as image I and therefore contains W × H variables, where W is the width of the image (in pixels) and H is its height. The second map FMd2 has half the resolution (in each dimension) of FMd1. The resolution of each additional map is half that of the previous one. This structure allows the number of variables in the feature map to be reduced, thereby facilitating decoding while minimizing decoding cost.

[0200] Graph FMd2 is upsampled by a factor of 2 in each dimension, graph FMd3 is upsampled by a factor of 4 in each dimension, and graph FMd4 is upsampled by a factor of 8 in each dimension, according to any upsampling method within the capabilities of a person skilled in the art.

[0201] Figure FMS' i has the same resolution as the image to be decoded, and therefore each map contains W × H variables, where W is the width of the image in pixels and H is its height.

[0202] In this embodiment, the vector Zd n is obtained from the pixel at the current pixel Pd n The coordinates (x n ,y n ) at the FMS' i The value of the vector Zd forms a 4-tuple (z1...z4). n is optionally dequantized and then processed by the synthetic neural network MLP' to generate a representation of the sample to be decoded Pd n The (R, G, B) triplet is output. The (R, G, B) triplet is inserted into the decoded image I(Pd n ) in the color components (Rd, Gd, Bd) (x n ,y n ) place.

[0203] According to the variant shown in dashed lines, there are 5 graphs: an additional graph FME'0 has been introduced. In this embodiment, the vector Zd n It is a 5-tuple.

[0204] Figure 9 It means that Figure 1 Coding equipment and Figure 5 Flowchart of a method for encoding a feature map implemented by an encoding method.

[0205] These steps form the reference above Figure 5The sub-steps of step E23 described above are intended to use the neighborhood values ​​for the feature maps FM in the first group being processed. i The current value V n to encode.

[0206] During sub-step E231 , a neighborhood vector (C n ), which includes the value V n The adjacent values. Figure 11 and Figure 12 As shown, these neighboring values ​​can be located in the same graph and / or in multiple (M) graphs FM i The neighborhood vector is formed by C values ​​or data corresponding to the neighborhood values ​​(for example, C = 10). The encoder and decoder must know these values; therefore, they must be located at the value V n in the causal neighborhood of .

[0207] According to a first embodiment, these values ​​are used to determine the context of the entropy encoder used to encode the current value during step E234. This encoder can be a CABAC (Context-Adaptive Binary Arithmetic Coding) encoder. This type of encoder is well known to those skilled in the art. It is primarily used in the H.265 / HEVC video compression standard. It is a lossless arithmetic encoder that decomposes all non-binary symbols into binary symbols. For each bit, the encoder selects the most appropriate probability model and uses context to optimize the probability estimate. This context can be defined by information about neighboring elements. Arithmetic coding is then applied to compress the resulting data. Those skilled in the art are aware of various ways to generate context information using neighborhood vectors. For example, the number of non-zero neighboring values ​​can be counted, and a context can be associated with each number. Alternatively, several neighboring values ​​can be compared, and a given context can be associated with an ordering configuration between the neighboring values, for example, by sorting the neighboring values ​​in ascending order and associating a context with each possible order.

[0208] In a second embodiment, during step E232, the neighborhood is used to predict the current value according to an autoregressive model. It should be noted that the autoregressive model predicts a series of samples using past values. In this embodiment, the past values ​​are formed by the context, and during step E234, the difference between the predicted variable and the actual value is quantized and then entropy coded.

[0209] In the third embodiment, as shown in FIG. Figure 12As shown, during step E233, a prediction neural network ARM is used to predict the statistical characteristics of the variable to be encoded. The neighborhood vector is provided as input to the network ARM so as to output a prediction of the current value. According to one embodiment, the network ARM is represented by a function f ѱ , as referenced Figure 4 As mentioned above, it provides a set of statistical parameters (mean value, variance, median value, etc.) for entropy encoding of the current value. The function of this module ARM is to i All the values ​​to be encoded V n The current value is best predicted to reduce the bit rate required to encode the feature map. According to another embodiment, a prediction neural network ARM is used to generate an expected probability (pr) of the possible value of the current sample. Entropy coding will adapt to this probability (such as the well-known Huffman or arithmetic entropy coding).

[0210] In the fourth embodiment, each feature map is divided into blocks of a predetermined size, and encoding of each block includes transformation (eg, DCT (Discrete Cosine Transform), Haar Transform, etc.), and the transformed value is encoded by entropy coding.

[0211] In the fifth embodiment, each feature map is divided into blocks having a predetermined size, and each block is represented by a product code of a lattice vector quantization type.

[0212] When this method is completed, the graph FM being processed i The current encoded value Vc n is encoded.

[0213] Figure 10 It means that Figure 2 Decoding equipment and Figure 7 Flowchart of a method for decoding a feature map implemented by a decoding method.

[0214] These steps form the reference above Figure 7 The purpose of these steps is to use the neighborhood values ​​to represent the feature maps FMd in the first group being processed. i Current value Vd n to decode.

[0215] During sub-step E311 , a neighborhood vector (C n ), which includes the value Vd n This step is similar to the step E231 described above and the same embodiments are applicable. The neighborhood vector is composed of the neighboring vectors located in the same graph and / or in multiple (M) graphs FMd i The neighborhood values ​​in different graphs correspond to C values ​​or data (for example, C = 10). nThese values ​​within the causal neighborhood of are known to the decoder.

[0216] According to a first embodiment, these values ​​are used to determine the context of the entropy decoder used to decode the current value during step E314. This decoding is similar to that used by encoders, such as CABAC encoders. The use of neighborhoods to generate context information is similar to that selected by the encoder. For example, the number of non-zero neighboring values ​​can be counted, and a context can be associated with each number. Alternatively, several neighboring values ​​can be compared, and a given context can be associated with an order configuration between the neighboring values, for example, by sorting the neighboring values ​​in ascending order and associating a context with each possible order.

[0217] In a second embodiment, during step E312 , the neighborhood is used to predict the current value according to an autoregressive model. In this embodiment, past values ​​are formed from the context and the difference between the predicted variable and the actual value is quantized and then entropy coded during step E314 .

[0218] In the third embodiment, as shown in FIG. Figure 12 As shown, during step E313, a prediction neural network ARM' is used to predict the statistical characteristics of the variable to be decoded. The neighborhood vector is provided as input to the network ARM' so as to output a prediction of the current value. According to one embodiment, the network ARM is represented by a function f ѱ , as referenced Figure 4 As described, it is defined by a set of statistical parameters (mean, variance, median, etc.) used to entropy encode the current value. According to another embodiment, a prediction neural network ARM' is used to generate the expected probability (pr) of the possible value of the current sample. Entropy decoding will adapt to this probability (such as the well-known Huffman or arithmetic entropy coding). If the encoding is lossless, the network ARM' is the same as the network ARM.

[0219] In a fourth embodiment, each feature map is divided into blocks of a predetermined size, and decoding each block includes entropy decoding the values ​​and then inverse transforming them (e.g., inverse DCT (discrete cosine transform), inverse Haar transform, etc.).

[0220] In the fifth embodiment, each feature map is divided into blocks having a predetermined size, and each block is decoded by a lattice vector quantization type product code to generate a decoded block.

[0221] When this method is completed, the graph being processed FMd i The current decoded value Vd n is decoded.

[0222] Figure 11A schematic diagram of a method for encoding or decoding a feature map according to an embodiment is shown.

[0223] In this illustration, the coordinates (x n ,y n ) in the current value V n (Correspondingly Vd n ) uses the context information of its own map and the previous map FM2 (or FMd2) respectively. The coordinates (x n ,y n -1)、(x n ,y n -2)、(x n -1,y n -1)、(x n -1,y n )、(x n -1,y n +1)、(x n -2,y n ) and the coordinates (x n -1,y n -1)、(x n -1,y n )、(x n -1,y n +1)、(x n ,y n -1), (x n ,y n )、(x n ,y n +1)、(x n +1,y n -1)、(x n +1,y n )、(x n +1,y n +1) is used to determine the neighborhood in which to encode (and decode, respectively) the current value. These values, available in both the encoder and decoder, constitute the variable V n (Correspondingly Vd n )’s neighborhood vector C n (Correspondingly Cd n ), the neighborhood vector can be used to combine the above Figure 9 (Correspondingly Figure 10 ) in any of the embodiments described.

[0224] Figure 12A schematic diagram showing another method for encoding or decoding a feature map according to an embodiment is shown.

[0225] In this figure, the current feature map FM i (Correspondingly FMd i ) coordinates (x n ,y n ) at the current value V n (Correspondingly Vd n ) uses the context information of its own graph. The values ​​shown in gray are used to determine the neighborhood in which the current value is encoded (respectively decoded). These values, available in both the encoder and decoder, constitute the value V n (Correspondingly Vd n )’s neighborhood vector C n (Correspondingly Cd n ), the neighborhood vector can be used to combine the above Figure 9 (Correspondingly Figure 10 ) in any of the embodiments described.

[0226] In the embodiment shown, the neighborhood vectors are extracted by the module CTX (respectively CTX') of the encoding module FMC (respectively FMD) and then applied to the input of the prediction neural network ARM (respectively ARM') for predicting the statistical characteristics (µ, σ) or probability (pr) of the value to be encoded (respectively decoded) by the entropy encoder CE (respectively DE).

Claims

1. A method for performing a multi-sample (P n ) of the signal to be coded (I(P n )) is a method for encoding, the method comprising the following steps: - Build step, which includes the following sub-steps: - Construct (E21, E22) to represent the signal (FM i ) of the first set of feature maps (FM i ); - For the position in the signal to be encoded (x n ,y n ) is associated with at least one sample point of the signal to be encoded, called the current sample point (P n ): - According to the current sample point (P n ) of the position (x n ,y n ) according to the feature map (FM) in the first group i ) Construct (E25) feature vector (Z n ); - Use a set of parameters (W k ) defines an artificial neural network (MLP) called a synthetic neural network to process the feature vector (Z) of (E27) n ), in order to provide a vector (P') representing the decoded value of the current sample. n ); - updating (E22, E27) at least one value of one of said feature maps of said first group and / or at least one parameter of said network as a function of a coding performance measure; - For the first set of feature maps (FM i ) is encoded (E23, E27, EF), comprising at least one value of one of said characteristic maps, called the current value (V n ), entropy encoding the value based on values ​​in at least its neighborhood; - the set of parameters (W) of the synthetic neural network k ) to perform encoding steps.

2. The encoding method according to claim 1, wherein The current value of one of the characteristic maps (V n ) The encoding step includes the following sub-steps: - According to the feature map (FM) in the first group i ) Construct (E231) neighborhood vector (C n ); - Use a set of parameters (O b ) defines an artificial neural network called predictive neural network (ARM) to process the neighborhood vector (C n ) in order to provide the current value (V n ) predictions; - updating (E22, E23, E28) at least one parameter of said prediction network according to the coding performance measure; - the set of parameters of the prediction network (O b ) to perform encoding steps.

3. The encoding method according to claim 1 or 2, wherein: The method comprises transforming the first set of feature maps (FM i ) step (E24) to obtain a second set of feature maps (FMS) having the resolution of the image of the input sequence i ); And the eigenvector (Z n ) is obtained from the feature map (FM) in the first group i ) The transformed feature maps (FMS) in the second group are obtained i ) constructed.

4. The encoding method according to claim 3, wherein: The feature map (FM) in the first group i ) has a lower resolution than the resolution of the signal to be encoded, and is characterized in that the transformation operation involves upsampling.

5. The encoding method according to claim 1, wherein: The eigenvector (Z n ) consists of the following sub-steps: - According to the current sample point (P n ) of the position (x n ,y n ) extract the feature map (FM) in the first group i ) multiple values; - processing said extracted values ​​in order to obtain a feature vector.

6. The encoding method according to any one of the preceding claims, characterized in that The method involves constructing another set of feature maps (FME l ) steps (E21, E22), and is characterized in that the feature vector is also constructed based on the feature map in the other group.

7. A method for detecting a plurality of sample points (Pd n ) is a method for decoding a signal to be decoded, the method comprising the following steps: - The first set of feature maps (FMd i ) is decoded (E31), comprising at least one value for one of said characteristic maps, called the current value (V n ), entropy decoding the value based on values ​​in at least its neighborhood; - A set of parameters (Wd k ) to decode (E35); - For the position (x n ,y n ) is associated with at least one sample point of the signal to be decoded, called the current sample point (Pd n ): - According to the position of the current sample point (x n ,y n ) According to the feature map (FMd) in the first group i ) Construct (E34) feature vector (Zd n ); as well as: - Using the decoded parameters (Wd k ) defines a synthetic neural network (MLP') to process the vector (Zd n ), in order to provide a representation of the current sample point (Pd n ) is a vector of decoded values.

8. The decoding method according to claim 7, wherein: The current value of one of the characteristic maps (V n ) The decoding step includes the following sub-steps: - A set of parameters (Od k ) to decode (E36); - According to the feature map (FM) in the first group i ) Construct (E311) neighborhood vector (C n ); - Process (E313) the neighborhood vector (C) using the prediction neural network (ARM') n ) to provide the current value (V n ) predictions.

9. The decoding method according to claim 7 or 8, wherein: The method comprises transforming the first set of decoded feature maps (FMd i ) step (E33) to obtain a second set of feature maps (FMS') with the resolution of the input signal i ); And the eigenvector (Zd n ) is obtained from the decoded feature map (FMd i ) The transformed feature maps (FMS') in the second group are obtained i ) constructed.

10. The decoding method according to claim 9, wherein: The characteristic map (FMd in the first group i ) has a lower resolution than the resolution of the signal to be decoded, and is characterized in that the transformation operation comprises upsampling.

11. The decoding method according to any one of claims 7 to 10, characterized in that: The characteristic vector (Zd n ) consists of extracting the current sample point (Pd n ) is the same position (x n ,y n ) in the at least one feature map (FMd i , FME' i ) value of the sub-step (E34).

12. The decoding method according to claim 7, wherein: The characteristic vector (Zd n ) consists of the following sub-steps: - According to the current sample point (P n ) of the position (x n ,y n ) extract the feature map (FMd) in the first group i ) multiple values; - Process (E35, TT') the extracted values ​​to obtain the feature vector.

13. The decoding method according to any one of claims 7 to 12, characterized in that: The method involves constructing another set of feature maps (FME' l ) step (E32), and is characterized in that the feature vector is also constructed based on the feature map in the other group.

14. A method for performing a multi-sample n ) of the signal to be coded (I(P n )) is a device for encoding, characterized in that The device is configured to perform the following operations: - Construct (GEN, MAJ) to represent the signal (FM i ) of the first set of feature maps (FM i ); - For the position in the signal to be encoded (x n ,y n ) is associated with at least one sample point of the signal to be encoded, called the current sample point (P n ): - According to the current sample point (P n ) of the position (x n ,y n ) according to the feature map (FM) in the first group i ) Construct (XTR) feature vector (Z n ); - Use a set of parameters (W k ) defines an artificial neural network (MLP) called a synthetic neural network to process the feature vector (Z n ), in order to provide a decoded value P' representing the current sample point n The vector (S n ); - updating (MAJ, NNC) at least one value of one of said feature maps of said first group and / or at least one parameter of said synthesis network according to a coding performance measure; - For the first set of feature maps (FM i ) is encoded (FMC), comprising at least one value of one of the characteristic maps, called the current value (V n ), entropy encoding the value based on values ​​in at least its neighborhood; - the set of parameters (W) of the synthetic neural network k ) is performed in the encoding step (NNC).

15. A method for performing a multi-sample n ) of a device for decoding a signal to be decoded, characterized in that The device is configured to perform the following operations: - The first set of feature maps (FMd i ) is decoded (FMD), including at least one value for one of the feature maps, called the current value (V n ), entropy decoding the value based on values ​​in at least its neighborhood; - A set of parameters (Wd k ) for decoding (NND); - For the position (x n ,y n ) is associated with at least one sample point of the signal to be decoded, called the current sample point (Pd n ): - According to the position of the current sample point (x n ,y n ) According to the feature map (FMd) in the first group i ) Construct (XTR') feature vector (Zd n ); - Using the decoded parameters (Wd k ) defines a synthetic neural network (MLP') to process (MLP') the feature vector (Zd n ), in order to provide a representation of the current sample point (Pd n ) is a vector of decoded values.

16. A computer program comprising instructions for executing the steps of the encoding method according to claim 1 or the decoding method according to claim 7 when said program is executed by a computer.