Method and device for encoding and decoding images
The method and device facilitate independent decoding of signal regions by encoding latent value maps and masks, addressing limitations in existing encoding methods and enhancing parallel processing and memory efficiency.
Patent Information
- Application Number
- FR2024006994
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2026-01-02
AI Technical Summary
Existing image encoding methods, both classical and neural, fail to allow for independently decodable areas in the coded signal, limiting interaction with semantic content and memory requirements at the decoder, as well as parallel decoding possibilities.
A method and device that encode a region of a signal by obtaining latent value maps, constructing characteristic vectors, processing them through a synthetic neural network, and encoding masks and parameters, enabling independent decoding of segmented areas.
Enables independent decoding of signal regions, reducing memory requirements and allowing parallel processing, while improving interaction with semantic content.
Smart Images

Figure 00000040_0000 
Figure 00000041_0000 
Figure 00000042_0000
Abstract
Description
Title of the invention: Method and device for encoding and decoding images. Prior art.
[0001] The invention relates to the general field of coding one-dimensional or multidimensional signals. It relates more particularly to the compression of digital images or videos.
[0002] Digital videos are generally subject to source coding aimed at compressing them in order to limit the resources required for their transmission and / or storage. Numerous coding standards exist, such as the ITU / MPEG standards (H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc.) and their extensions (MVC, SVC, 3D-HEVC, etc.). In these approaches, image encoding is generally performed by predicting pixels using previously encoded and then decoded pixels present in the image being encoded, in which case it is called "Intra prediction," or previously encoded images, in which case it is called "Inter prediction."
[0003] In addition to these classic approaches, approaches based on artificial intelligence, and in particular neural networks, tend to develop.
[0004] Some neural approaches, starting from an input signal, for example an image, train a so-called synthetic neural network on characteristic vectors associated with a position of a sample of the input signal to be encoded. These characteristic vectors are constructed from feature maps that may be at the resolution of the input signal, or at a lower resolution. During training, or construction, the parameters of the neural network and the values of the feature maps are updated according to a performance measure, for example, a data rate-distortion type. When training is complete, i.e., when the performance measure obtained is satisfactory, the actual encoding of the synthetic neural network parameters and the values of the feature maps can be performed and stored or transmitted to the decoder.The decoding of the coded signal is then carried out by applying the synthesis neural network to the feature maps.
[0005] One drawback of the classical and neural approaches described above is that they do not allow the creation of independently decodable areas in the coded signal, which limits not only the ability to interact with the semantic content of the coded signal but also the possibility of limiting the memory required at the decoder level or of parallelizing the decoding of this coded signal.
[0006] There is therefore a need for a solution to improve upon the classical and neural approaches described above. Summary of the invention
[0007] The invention relates to a method for encoding a region of a signal, called the encoding region, said encoding region comprising a plurality of samples to be encoded, said encoding method comprising the following steps: - obtaining a first group of at least one latent value map representative of said signal, - obtaining a mask in said signal from said zone to be coded, - for at least one sample of said area to be coded, called the current sample, associated with a position in said signal to be encoded: • the construction of a characteristic vector from said latent value maps of said first group, as a function of said position of said current sample, • the processing of said characteristic vector by an artificial neural network, called a synthetic neural network, said synthetic neural network being defined by a set of parameters and comprising at least one synthetic neural layer to obtain, at the output of said synthetic neural network, a vector representing a decoded value of the current sample, • the updating of at least one value of one of said latent value maps of said first group and / or of at least one parameter of said synthetic neural network, as a function of a coding performance measure, - obtaining a second group of latent value maps representative of at least the said area to be encoded from said mask, - the encoding of the second group of latent value maps, - the encoding of said mask, and - the coding of said set of parameters of said synthetic neural network. said processing comprising the application of said at least one neural synthesis layer to transform at least one group of input latent maps, referred to as input maps, into at least one group of output latent maps, referred to as output maps, said application comprising: - obtaining a characteristic region of said zone to be decoded in said entry maps from said mask, said characteristic region comprising the points associated with said mask, - for at least one point in said area to be decoded, called the point to be decoded: • obtaining a point associated with said point to be decoded in at least one of the latent input maps, • the evaluation of a proximity criterion of said associated point in relation to a boundary of said characteristic region, • the construction of an input vector for said neural synthesis layer from a neighborhood of said associated point and said evaluation of said proximity criterion, and • the processing of said input vector by said neural synthesis layer to obtain a point of at least one of the output maps,
[0008] The invention also relates to a method for decoding an area, said area to be decoded, of a signal comprising at least two areas, said area to be decoded comprising a plurality of samples to be decoded, said decoding method comprising the following steps: - the decoding of a group of latent value maps representative of said area to be decoded of said signal, - obtaining a mask of said area to be decoded in said signal, - the decoding of a set of parameters representative of a neural network, called a synthetic neural network comprising at least one synthetic neural layer, - the processing of said group of latent maps decoded by said synthetic neural network to produce at the output of said synthetic neural network at least said area to be decoded, said processing comprising the application of said at least one synthetic neural layer to transform at least one group of input latent maps, said input maps, into at least one group of output latent maps, said output maps, said application comprising obtaining a characteristic region of said area to be decoded in said input maps from said mask, said characteristic region comprising the points associated with said area to be decoded and for at least one point of said area to be decoded, said decode point, the processing of a point associated with said decode point in at least one of the input latent maps by said neural layer as a function of its proximity to a boundary of said characteristic region.
[0009] For the purposes of this invention, encoding, or "coding", means the operation of representing a set of samples in a compact form. for example by a digital binary stream. Decoding is understood to be the operation of processing a digital binary stream to produce decoded samples.
[0010] By "sample" of the signal, we mean a value taken from the signal. Sampling the signal produces a sequence of discrete values called samples. In the case of an image signal, the sample is called a pixel, which can be, for example, a color pixel traditionally represented by a triplet of values, for example (R, G, B) or (Y, U, V). The position of the sample is located by its abscissa (x) and ordinate (y) coordinates in the image.
[0011] By "signal comprising a plurality of samples" is meant a signal with one (audio, sound), two (image) or more than two (stereoscopic image, multiscopic image, image associated with a depth map, video, etc.) dimensions. Depending on this dimensionality, the sample has one, two, or more coordinates in the signal. In the case of an image signal, the position of the sample is located by its abscissa (x) and ordinate (y) coordinates.
[0012] By "feature maps" or equivalently by "latent value maps" is meant an abstract representation of the signal comprising a plurality of variable data, discrete or not, which are also called values, for example real or integer numbers.
[0013] By "characteristic data vector constructed from feature maps as a function of a position" is meant a vector consisting of one or more elements, or data, preferably discrete, the data being constructed from the feature maps at a position determined by that of the sample being processed in the signal. This characteristic vector is the one that is applied to the input of the synthesis neural network. For example, in the case of a one-dimensional audio signal, such a vector can be constructed from a plurality of values taken from each of the feature maps at the same coordinate as the sample to be encoded or in a neighborhood thereof. In the case of an image, such a vector can be constructed from a plurality of values taken from each of the feature maps at the same x- and y-coordinates as the sample to be encoded (respectively,to decode) or in a neighborhood of it. Once these values are taken from the feature maps, they can be processed to form the feature vector, before entering the synthesis neural network, for example by quantization, filtering, interpolation, etc.
[0014] By "synthetic neural network," we mean a neural network such as a convolutional neural network, a multilayer perceptron, an LSTM (for "Long Short Term Memory"), etc. The neural network is defined, for example, by a plurality of layers of artificial neurons and by a set of functions activation, weighting and addition (for example, a layer can compute y = f (Ax+b), where y and b are N-dimensional vectors, x is an M-dimensional vector, A is an MxN-dimensional matrix, and f is the activation function).
[0015] By "neural network parameter" we mean one of the values that characterizes the neural network, for example a weight associated with one of the neurons (filter coefficient, weighting, bias, value affecting the functioning of non-linearity, etc.)
[0016] By "processing by a synthetic neural network" is meant the application of a function expressed by a synthetic neural network to the input characteristic vector to produce an output vector representative of the sample to be encoded (resp. decoded). This output vector may contain one or more data points representative of the sample.
[0017] By "performance measurement," we mean a measurement between at least one value of a sample to be encoded and a decoded value of said sample. The measurement may, for example, assess distortion or perceptual error. It may be performed on one or more samples (for example, a current sample, or the current image, etc.). The measurement may also include a measurement of throughput, particularly associated with the encoding of the synthetic neural network and / or the encoding of the feature maps of the first group. The measurement may be a joint measurement of throughput and distortion through their weighting. As is well known in the prior art, the value of this measurement is generally minimized until a target value is reached.
[0018] Generally speaking, the steps of an encoding or decoding process should not be interpreted as being linked to a notion of temporal succession. In other words, the steps may be carried out in a different order than that indicated in the independent encoding or decoding claim, or even in parallel.
[0019] The coding method according to the invention encodes a region of a signal from a representation of that signal in the form of feature maps. These feature maps are segmented into characteristic regions whose values are subsequently encoded entropically, independently of one another. Furthermore, a synthetic neural network is trained on all of these regions, taking into account the segmentation into characteristic regions. Thus, it is possible to obtain a coded representation of a region of the original signal that can subsequently be decoded by the synthetic neural network independently of any other part of that signal.
[0020] The decoding process (and symmetrically the encoding process) may further include one or more of the following optional features, taken individually or in any technically possible combination.
[0021] According to a first characteristic, the processing step of said associated point (Pan) comprises: - obtaining a point associated with said point to be decoded in at least one of the latent input maps, - the evaluation of a proximity criterion of said associated point in relation to a boundary of said characteristic region, - the construction of an input vector for said neural synthesis layer from a neighborhood of said associated point and said evaluation of said proximity criterion, and - the processing of said input vector by said neural synthesis layer to obtain a point of at least one of the output maps.
[0022] According to another feature, the neighborhood of the associated point is independent of the evaluation of the proximity criterion.
[0023] According to another feature, the neighborhood of the associated point is selected based on the evaluation of the proximity criterion.
[0024] According to another feature, the neighborhood of the associated point is selected based on the associated point.
[0025] According to another feature, during the construction of the input vector, a component of the input vector is associated with a point in the neighborhood, the value of the component being equal to the value of the associated point if the associated point is a point in the characteristic region and to a replacement value otherwise.
[0026] According to another feature, the replacement value is dependent on the points of the characteristic region or on the points associated with the neighborhood belonging to the characteristic region.
[0027] According to another feature, the replacement value does not depend on the points of the feature region.
[0028] According to another characteristic, the proximity criterion is a distance, for example a Euclidean distance.
[0029] According to another characteristic, the neuronal synthesis layer is a convolutional neuronal layer.
[0030] Correspondingly, the invention also relates to a coding device for a signal comprising a plurality of samples to be coded, characterized in that said coding device is configured to implement: - obtaining a first group of at least one latent value map representative of said signal, - obtaining a mask in said signal from said zone to be coded, - for at least one sample of said area to be coded, called the current sample, associated with a position in said signal to be encoded: • the construction of a characteristic vector from said latent value maps of said first group, as a function of said position of said current sample, • the processing of said characteristic vector by an artificial neural network, called a synthetic neural network, said synthetic neural network being defined by a set of parameters and comprising at least one synthetic neural layer to obtain, at the output of said synthetic neural network, a vector representing a decoded value of the current sample, • the updating of at least one value of one of said latent value maps of said first group and / or of at least one parameter of said synthetic neural network, as a function of a coding performance measure, - obtaining a second group of latent value maps representative of at least the said area to be encoded from said mask, - the encoding of the second group of latent value maps, - the encoding of said mask, and - the coding of said set of parameters of said synthetic neural network. said processing comprising the application of said at least one neural synthesis layer to transform at least one group of input latent maps, referred to as input maps, into at least one group of output latent maps, referred to as output maps, said application comprising: - obtaining a characteristic region of said zone to be decoded in said entry maps from said mask, said characteristic region comprising the points associated with said mask, - for at least one point in said area to be decoded, called the point to be decoded: • obtaining a point associated with said point to be decoded in at least one of the latent input maps, • the evaluation of a proximity criterion of said associated point in relation to a boundary of said characteristic region, • the construction of an input vector for said neural synthesis layer from a neighborhood of said associated point and said evaluation of said proximity criterion, and • the processing of said input vector by said neural synthesis layer to obtain a point on at least one of the output maps,
[0031] The invention also relates to a device for decoding an area of a signal comprising at least two areas, said area to be decoded, said area to be decoded comprising a plurality of samples to be decoded characterized in that the decoding device is configured to implement: - the decoding of a group of latent value maps representative of said area to be decoded of said signal, - obtaining a mask of said area to be decoded in said signal, - the decoding of a set of parameters representative of a neural network, called a synthetic neural network comprising at least one synthetic neural layer, - the processing of said group of latent maps decoded by said synthetic neural network to produce at the output of said synthetic neural network at least said area to be decoded, said processing comprising the application of said at least one synthetic neural layer to transform at least one group of input latent maps, said input maps, into at least one group of output latent maps, said output maps, said application comprising obtaining a characteristic region of said area to be decoded in said input maps from said mask, said characteristic region comprising the points associated with said area to be decoded and for at least one point of said area to be decoded, said decode point, the processing of a point associated with said decode point in at least one of the input latent maps by said neural layer as a function of its proximity to a boundary of said characteristic region.
[0032] The characteristics and advantages of the coding or decoding process apply in the same way to the coding or decoding device according to the invention and vice versa.
[0033] The invention also relates to a computer program on a recording medium, this program being capable of being implemented in a computer or an encoding or decoding device according to the invention. This program comprises instructions adapted to the implementation of the corresponding method. This program may use any programming language and may be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0034] The invention also relates to an information or recording medium readable by a computer, and comprising computer program instructions as mentioned above. The information or recording media can be any entity or device capable of storing programs. For example, the media can comprise a storage means, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a floppy disk or a hard drive, a DNA sequence, or flash memory. On the other hand, information or recording media can be transmissible media such as an electrical or optical signal, which can be transmitted via an electrical or optical cable, by radio link, by wireless optical link or by other means.
[0035] The program according to the invention can in particular be downloaded onto an Internet-type network.
[0036] Alternatively, each information or recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of a process according to the invention. Brief description of the figures
[0037] The invention will be better understood with the aid of the following description, given solely by way of example and made with reference to the accompanying drawings in which: - Figure [1] schematically represents a coding device according to one embodiment of the invention, - Figure [Fig. 2] schematically represents a decoding device according to one embodiment of the invention, - Figure [Fig. 3] illustrates an example of a synthetic artificial neural network used in the context of the invention, - [Fig.4] schematically illustrates an example of processing performed by a neuronal synthesis layer of the neuronal network of [Fig.3], - [Fig.5] schematically illustrates a second example of processing carried out by a neuronal synthesis layer of the neuronal network of [Fig.3], - [Fig.6] schematically illustrates a third example of processing carried out by a neuronal synthesis layer of the neuronal network of [Fig.3], - [Fig.7] is a logic diagram representing an example of a coding process that can be implemented by the coding device of [Fig.1], - Figure 8 illustrates a coding method used in one embodiment of the invention. - [Fig.9] is a flowchart representing an example of a decoding process that can be implemented by the decoding device of [Fig.2], - Figure 10 illustrates a decoding method used in one embodiment of the invention. - [Fig.1 1] is a logic diagram representing a method for encoding feature maps that can be implemented by the encoding device of [Fig.1] and by the encoding process of [Fig.7], - [Fig. 12] is a logic diagram representing a method for decoding feature cards which can be implemented by the decoding device of [Fig.2] and by the decoding process of [Fig.9]. Detailed description of the invention
[0038] Figure 1 schematically represents, according to a first embodiment, an ENC coding device for at least one zone of a signal (I(Pn)). In the example described here, the signal I(Pn) is segmented into a set of J zones, and all zones are coded independently. Alternatively, only one zone or only a few zones may be coded independently.
[0039] This ENC coding device includes a SEG segmentation module, an SEGC encoding module of the segmentation provided by the SEG segmentation module, a GEN module for generating feature maps, an SE module for transforming feature maps, an XTR module for extracting data from feature maps, a TT module for processing and quantification, an NNSYN module corresponding to a synthetic artificial neural network, an NNC module for coding a neural network capable of encoding the synthetic neural network, an FMC module for coding feature maps, an EVAL module for evaluating coding performance, and an MAJ module for updating.
[0040] The ENC coding device can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be realized by the cooperation of the processor and computer program instructions stored in the aforementioned memory and designed to perform the functionalities of the module concerned, in particular as described below, when these instructions are executed by the processor.
[0041] The ENC coding device of [Fig. 1] receives as input a signal consisting of a succession of samples to be coded, denoted Pn, for example, a temporal sequence of sound samples, or a set of image data denoted I(Pn). In this second case, the image signal I(Pn) can represent a two-dimensional image, or a plurality of two-dimensional images (video, color components, stereoscopic components, multiscopic components, etc.). Pn designates a sample n of the input signal comprising N samples. In one embodiment, the signal is a color image signal represented by means of at least one two-dimensional representation, such as a pixel matrix, each pixel comprising a The image consists of a red component (R), a green component (G), a blue component (B), or, alternatively, a luminance component (Y) and at least one chrominance component (U, V). The location of each pixel is defined by its x-coordinate (abscissa) and y-coordinate (ordinate) in the image. In one embodiment, the grayscale image is represented by a two-dimensional representation, such as a pixel matrix, with each pixel having a grayscale component, or luminance. In this case, the pixel's representative vector is reduced to a single component.
[0042] The SEG module performs a segmentation Sg of the sequence of samples to be encoded Pn into J (greater than or equal to 2) zones Zo'. This segmentation operation allows the samples to be grouped into J different homogeneous sets according to one or more predefined criteria. For example, in the case of a temporal sequence of sound samples, the segmentation operation makes it possible to obtain sequences of units corresponding to silences, noises, phonemes, words, etc. Similarly, in the case of an image signal I(Pn), the segmentation operation makes it possible to group the pixels Pn of the image signal into J homogeneous zones according to criteria, notably intensity or spatial criteria. For example, the segmentation operation can make it possible to identify in the image signal I(Pn) two different zones (J=2), one corresponding to the background of the image and the other to the foreground of the image.Segmentation Sg is, for example, represented as a set of J masks MZo', each allowing the identification of a ZoJ zone, each mask being associated with a value different from that associated with the other masks.
[0043] The SEGC module performs lossless encoding of the Sg segmentation. This encoding can be achieved by encoding the segmentation map corresponding to the J MZo' masks (for example, using the JPEG-LS algorithm defined by the international standard ISO / IEC 14495-1) or, alternatively, by independently encoding each MZoj mask, or, in yet another alternative, by independently encoding the contours of the J MZoj masks. The encoded segmentation is denoted Sgc.
[0044] The NNC module performs a coding simulation, followed by a decoding, for the evaluation module
[0045] The GEN feature map generation module is configured to generate a plurality of M feature maps, denoted FM;.
[0046] In one embodiment, the SE module performs a transformation of the first group of FM characteristic maps; to generate a second group of FMS characteristic maps; at the same resolution as the input signal.
[0047] The optional SE module can perform a quantification of the data extracted from this set of M FM cards; It should be noted that the quantification of a value refers to mapping that value to a member of a set A discrete set of possible code symbols. For example, the set of possible code symbols might consist of integer values, and the quantization system performs a simple rounding from a real value to an integer. In another example, quantization involves multiplication by a given value followed by rounding. Then, the SE module performs a transformation of the values in at least one of the feature maps, such as oversampling, interpolation, filtering, etc. After the transformation, a transformed feature map from the second group has the same resolution as the images in the input sequence. Advantageously, according to this method, the feature maps that are coded can be of lower resolution than the images to be coded, while the maps of the second group, which are used to construct the feature vectors, are at the same resolution as the image sequence, which facilitates the extraction of values.
[0048] In one embodiment, the SE module is absent; in this case, the values that will be used to construct the characteristic vector are extracted from the first group of characteristic maps.
[0049] The XTR module performs value extraction from the FMSi (or FM; depending on one of the embodiments described above) feature maps for a current sample Pn to be encoded, based on its coordinates in the input signal and optionally the segmentation Sg performed by the SEG module. For example, if the goal is to encode the sample Pn at coordinates (xn, yn) in a Zo' region of an input image, the XTR module performs value extraction from the maps at positions determined by the coordinates (xn, yn) and by the MZo' mask of the Zo* region.
[0050] In one embodiment, the extracted values constitute the vector Zn. Zn is an L-tuple, that is, it contains L elements, or data z;. For example, in one embodiment, L=M, meaning, for example, that only one value is extracted for each feature map FM;. In another embodiment, L>M, meaning, for example, that several values are extracted for at least one feature map FM;. The dimension L of the vector depends on the topology of the NNSYN synthesis neural network and, more particularly, on the topology of the input layer of this NNSYN synthesis neural network.
[0051] The vector Zn of index n refers to the characteristic vector of the pixel P'n.
[0052] In one embodiment, the optional TT module processes the extracted values to generate the vector Zn. The TT module can quantify the data extracted from the set of feature maps. The processing may include other operations, such as filtering, scaling, etc. In particular, if the SE module is not used and if the feature maps of the first group have resolutions lower than those of the sequence images, The TT module can take into account the coordinates of values in lower resolution maps.
[0053] It should be noted that at least one of the SE or TT modules must perform a quantification of the characteristic maps.
[0054] The NNSYN module is a synthetic neural network defined by K parameters Wk, capable of processing the input vector Zn, or L-Tuple, to generate as output a second vector representative of the sample Pn to be coded.
[0055] An example of a synthetic neural network is presented later with reference to [Fig.3].
[0056] It should be noted that the behavior of the synthesis neural network may depend on the Sg segmentation as will be presented in more detail later with reference to Figures 4 to 6.
[0057] The NNC module performs the coding of the synthetic neural network, specifically its parameters Wk. During the coding training, or construction, process—that is, until the performance evaluation stage is satisfactory—the NNC module performs a coding simulation, followed by decoding, for the evaluation module. Subsequently, it performs the actual coding of the synthetic neural network parameters Wk. The coded parameters are denoted Wck. As is known, the coding simulation can be identical to the actual coding, or it can approximate it.
[0058] The FMC module performs the encoding of the FM maps; that is, the values of the feature maps of the first group (excluding the maps of the second group, which optionally result from oversampling by the SE module). During the encoding training, or construction, process—that is, as long as the performance evaluation step is not satisfactory—the FMC module performs a coding simulation, followed by decoding, for the evaluation module. Subsequently, it performs the actual encoding of the FM map values.
[0059] The FMC module determines, taking into account the Sg segmentation performed by the SEG module, an SgL segmentation of each of the FM characteristic maps; in J zones ZoL^. Thus, the FMC module uses the Sg segmentation defined in the domain of the signal to be encoded to obtain a segmentation in the latent domain of the FM characteristic maps. Obtaining this segmentation in the latent domain depends on the transformation transform used to go from the domain of the signal to be encoded to the latent domain. Thus, the SgL segmentation in the latent domain results from a calculation that simultaneously takes into account the location of the Sg segmentation zones in the signal to be encoded and the transformation transform. For example, if the latent domain and the signal to be encoded have the same resolution, then the segmentation in the domain The latent region is identical to the segmentation in the domain of the signal to be encoded. As another example, if a latent region has a resolution lower than that of the signal to be encoded, then the boundary between two regions of the segmentation in the domain of that latent region lies between the points in the latent region whose co-located samples in the signal to be encoded belong to two different regions of the Sg segmentation. Each of the FM maps is represented as a set of J characteristic regions ZoL', each of the characteristic regions ZoL' corresponding in the FM map to region j of the SgL segmentation. It is possible that some characteristic regions ZoL' may be empty. The encoding of the values of the FM maps is performed taking into account the SgL segmentation. More precisely, each of the J characteristic regions ZoL' is encoded successively, for example, in the order of their indexing.Thus, each of the characteristic zones ZoL^ is coded independently of the other characteristic zones. The coded data of each coded zone ZoL' in each of the characteristic maps FM; is noted. Each zone ZoL' is an entropic coding. Thus, the coding of the M FM; maps corresponds to the independent coding of the J characteristic zones of the M FMi maps.
[0060] In a first embodiment, the FMC module only encodes the values of the points in the characteristic zone ZoLj, excluding any other points in the FM; characteristic maps. Thus, the encoding of a characteristic zone ZoLj comprises the successive encoding, for each FM; map, of the values of the points in the characteristic zone ZoLj in the FM; map.
[0061] Alternatively, the FMC module creates a secondary feature map LM2J for each feature zone ZoLj. The value of a point on the secondary feature map pjyf2^ is equal to the value of the point on the FM map; if this point belongs to the feature zone ZoLj, it is a predetermined value, for example, zero, otherwise. The FMC module then encodes the entire feature map to obtain different encoded data for each feature zone ZoLj.
[0062] In a known manner, the coding simulation can be identical to the actual coding, or perform an approximation of it. The coding module quantifies, if necessary, the latent representation of the values of the cards in the first group using a quantifier to generate an ordered collection of quantized values. Then the coding module compresses the quantized data, using a coding that takes into account the neighborhood of a value to be coded in the feature card.
[0063] The EVAL module performs an evaluation and minimization of coding performance. The evaluation function is, for example, of the rate-distortion type. The minimization can be performed by gradient descent, or any other method within the reach of a person skilled in the art.
[0064] The MAJ module performs an update of the values of the FM cards; and / or the parameters of the neural network to be encoded, according to the results of the performance function.
[0065] Fig. 2 schematically represents a decoding device DEC of a decoding zone Zodj of a signal, called the decoding zone, said decoding zone Zodj comprising a plurality of samples Pdn to be decoded.
[0066] This DEC decoding device includes an NND module for decoding neural network(s) capable of decoding the NNSYN' synthesis neural network, a SEGD module for decoding a segmentation, an FMD module for decoding feature maps, an XTR' module for data extraction, an SE' module for inverse transformation, and a TT' module for inverse processing and quantization.
[0067] The DEC decoding device produces at output a decoded image comprising at least the decoded area, denoted Zodj (Pdn), comprising a plurality of decoded samples Pdn.
[0068] The DEC decoding device of Figure 2 receives as input the Sgc-encoded segmentation and a group of pMci-encoded data
[0069] The DEC decoding device of [Fig.2] also receives as input the Wck encoded parameters of the NNSYN' synthesis neural network.
[0070] The parameters of the NNSYN' synthesis neural network decoded by the NND module are noted Wdk.
[0071] The SEGD module decodes the MZo* mask from the Sgc encoded data.
[0072] According to embodiments as described for the encoder: - The segmentation map corresponding to the J MZo* masks is decoded from the Sgc data and the MZo' mask is extracted from the segmentation map. - The MZo' mask is decoded directly when it has been previously encoded in the Sgc encoded data independently of other masks.
[0073] The FMD module constructs M decoded maps using the mask MZo' and the encoded data pMcf. Depending on the data encoding implementation of the pjyfcj data implemented by the ENC encoder, and starting from the mask MZo* in the domain of the signal to be decoded, the FMD module determines the corresponding characteristic area ZoLj in the latent domain, and then the FMD module decodes the pj^J data. to obtain the point values of this characteristic zone ZoLj in each of the pMdJ maps.
[0074] In one embodiment, the SE' module performs a transformation of the first set of decoded feature maps to generate a second set of feature maps at the same resolution as the signal to be decoded, denoted pMS'i. The SE' module optionally performs inverse quantization corresponding to the quantization performed by the encoder. Inverse quantization is not necessary if the encoder's quantizer Q has simply rounded the real values submitted to it. Inverse quantization is also unnecessary if the neural network is capable of handling quantization of its input data. Otherwise, the decoder performs the inverse operation of the quantizer Q. Then, the SE' module performs a transformation of the feature map values, including, for example, oversampling, interpolation, filtering, etc., similar to that performed by the encoder.Following the transformation, a transformed feature map of the second group has the same resolution as the images of the sequence to be decoded.
[0075] In one embodiment, the SE' module is absent; in this case, the values that will be used to construct the characteristic vector are extracted from the first group of characteristic maps.
[0076] The XTR' module is identical to the XTR module of Figure 1. It performs an extraction of values from the M characteristic maps (or pJVIS'- sc'on ' one of the embodiments described above), for a sample Pdn to be decoded, as a function of its coordinates in the signal to be decoded and the mask MZo'.
[0077] In one embodiment, the extracted values constitute the vector Zdn. Zdn is an L-tuple, that is to say, it comprises L elements, or zd data;
[0078] In one embodiment, the optional TT' module processes the extracted values to generate the Zdn vector. The TT' module can perform inverse quantization of the data extracted from the feature set. The processing may include other operations, such as filtering, scaling, etc., similar to those performed by the encoder.
[0079] The NNSYN' module is a so-called synthetic neural network, defined by K parameters Wdk, capable of processing the input vector Zdn, or L-tuple, to generate as output a second vector representative of the sample Pdn to be decoded, generally a vector containing A elements. In one embodiment, K=3 and the output vector is the triplet (R, G, B) of the decoded Pdn pixel. The NNSYN' module is of identical structure to the NNSYN module, and its parameters are either identical if the encoding of its parameters Wk, or different if the encoding is done with losses.
[0080] When all the Pdn samples of the signal have been decoded, we have a reconstructed signal Zodj(Pdn) of the Zo' zone.
[0081] Furthermore, the DEC decoding device can successively be implemented to decode all the Zodj areas in order to reconstruct all the samples to be decoded of the signal, i.e. of the image I(Pdn).
[0082] The DEC decoding device can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be implemented through the cooperation of the processor and computer program instructions stored in the aforementioned memory and designed to perform the functionalities of the module concerned, in particular as described below, when these instructions are executed by the processor. The DEC device can also comprise a plurality of processors, the processors being dedicated to the parallel decoding of image areas.
[0083] Figure 3 illustrates an example of a synthetic artificial neural network used for encoding and decoding in the context of embodiments of the invention.
[0084] The synthetic artificial neural network used for encoding, NNSYN, and the synthetic artificial neural network used for decoding, NNSYN', are defined by an identical structure, comprising for example a plurality of layers of artificial neurons, and by a set of weights and activation functions associated respectively with the artificial neurons of the network concerned.
[0085] The synthetic neural network is, according to one embodiment, an MLP, or Multi Layer Perceptron, followed by one or more convolutional neuron layers Ci, ... Cn, each of the convolutional neuron layers being associated with a filtering mask of predefined size, for example of size 3x3.
[0086] The MLP consists of an input layer adapted to the input format (the L-tuple), optionally one or more hidden layer(s), and an output layer providing an intermediate output vector Vsn, generally a vector comprising A' elements.
[0087] Thus, a vector representation of a current sample (a vector Zn or Zdn from the FM / FMS characteristic maps, at the encoder or FMd^FMS'- at the decoder) is applied to the input (i.e. on the input layer) of the MLP which produces at the output the intermediate output vector Vsn.
[0088] The concatenation of all these intermediate output vectors constitutes intermediate latent value maps which are then processed successively by the convolutional layer(s) of neurons to provide as output a set of output vectors also comprising A elements.
[0089] According to one embodiment, A is equal to 3, the intermediate output vector is a triplet and the output vector is the triplet (R, G, B) of the pixel P'n encoded and then decoded.
[0090] The concatenation of all these triplets in an image constitutes the reconstructed signal I(Pdn), according to an example an image I, or Zodj (Pdn), according to an example the decoded area Zodj.
[0091] At the encoder, the NNSYN synthetic artificial neural network is trained on the image so as to minimize the differences between the input representation of the current image I(Pn) and its output representation I(P'n), while also minimizing the amount of data to be encoded. The EVAL module performs a performance measurement in this regard.
[0092] Once the encoder training is complete, the network parameters are encoded, either losslessly, in which case the NNSYN' neural network is identical to NNSYN, or with losses, in which case the NNSYN' network may be slightly different from NNSYN.
[0093] With reference to [Fig.4] and [Fig.5], we will now present the application of a convolutional neural layer Ck of the NNSYN / NNSYN' synthesis neural network to a group of input latent value maps CCEi, CCE2, CCE3, to obtain at the output of the convolutional neural layer Ck a group of output latent value maps CCSi, CCS2, CCS3.
[0094] In the example described here, the convolutional neural layer Ck is associated with a convolutional kernel of predefined size, for example of size 3x3. This convolutional neural layer Ck transforms three input latent value maps CCEh CCE2, CCE3 into three output latent value maps CCSi, CCS2, CCS3.
[0095] The application of this convolutional neural layer Ck to the three input latent value maps CCEi, CCE2, CCE3 includes, in a step (1), obtaining a characteristic region Rj comprising all the points of the maps CCEi, CCE2, CCE3 co-located with samples belonging to the mask MZo' of the Zodj area. The points of this characteristic region Rj are called determined points.
[0096] Then, in a step (2), for a point Psn of the CCSi, CCS2, CCS3 maps colocalized with a sample Ptn belonging to the mask MZo' of the area to be decoded Zodj, an input vector Vcon of the convolutional neural layer Ck is constructed.
[0097] This input vector Vcon depends on the Pan point of the CCEi, CCE2, CCE3 maps, which is co-located with the Psn point of the CCSi, CCS2, CCS3 maps, and on contextual information about this Pan point. This contextual information is represented by a neighborhood of the Pan point, this neighborhood including, by definition, the Pan point itself. In In the example presented here, the neighborhood used includes 9 points associated with a 3x3 square mask centered on the point Pan.
[0098] This contextual information is also conditioned by the characteristic region Rj in the input maps CCEi, CCE2, and CCE3. More precisely, the contextual information depends on the location of point Pan relative to the boundary of area Rj. This location is obtained by evaluating a proximity criterion Cp(Pan) of point Pan relative to the boundary of characteristic region Rj. In the example described here, the proximity criterion Cp(Pan) is the Euclidean distance between point Pan and the boundary of characteristic region Rj. Alternatively, the proximity criterion Cp(Pan) could be a Manhattan, Minkowski, or Chebyshev distance. In another variant, the distance could simply be measured in pixels, for example, by counting the number of pixels along the two directions of a coordinate system associated with the current map.In yet another variant, the proximity criterion Cp(Pan) may also be a quasi-distance or a gap or any other relevant measure indicating the proximity of point Pan to the boundary of the characteristic region Rj.
[0099] Two cases must then be distinguished: - In the first case, illustrated in [Fig.4], the proximity criterion Cp(Pan) is greater than (possibly equal to) a distance d (for example predefined) and the neighborhood values of the input maps CCEi, CCE2, CCE3 are used to form the input vector Vcon. - In the second case, illustrated in [Fig. 5], the proximity criterion Cp(Pan) is less than this distance d (in the example described here, the point Pan is located at zero distance from the boundary of the characteristic region Rj). In this case, the input vector Vcon is defined component by component. The value of a component of the input vector Vcon is equal to the value of the associated point on the input map if that point belongs to the characteristic region Rj, and to a replacement value otherwise.
[0100] During a step (3), the input vector is processed by the convolutional neural layer Ck to generate as output the point Psn of the CCSi, CCS2, CCS3 maps.
[0101] In [Fig. 4] and [Fig. 5], the values shown in gray are part of the characteristic region Rj and are used to determine components of the input vector Vcon. The values shown in white ([Fig. 5]) do not belong to the characteristic region Rj but to another region Rx and, since they cannot be used to define the input vector Vcon, are replaced by a replacement value Rp.
[0102] In one embodiment, the replacement value is a predetermined constant value, for example equal to 0.
[0103] In another embodiment, the replacement value is a function of the values of the neighborhood points belonging to the characteristic region Rj; for example, the replacement value is equal to the value of point Pan. In another embodiment, a set of replacement values is calculated based on the values of the neighborhood points belonging to the characteristic region Rj. A neural network can be implemented to calculate a given number of replacement values based on the neighborhood points of the characteristic region Rj.
[0104] In the preceding embodiments, the definition of the neighborhood of point Pan is independent of the proximity criterion Cp(Pan). In another embodiment, and as shown in [Fig. 6], the definition of the neighborhood of point Pan may depend on the proximity criterion Cp(Pan).
[0105] Two cases must again be distinguished: - In the first case, the proximity criterion Cp(Pan) is greater than (possibly equal to) the distance d and the values associated with a first neighborhood of the point Pan (for example identical to that defined previously in connection with [Fig.4]) of the input maps CCEi, CCE2, CCE3 are used to form the input vector Vcon. - In the second case, the proximity criterion Cp(Pan) is less than this distance d. In this case, a second neighborhood of the point Pan is used. Preferably, to respect the topology of the convolutional neural layer Ck, the number of points in the first neighborhood is equal to the number of points in the second neighborhood. In the example shown in [Fig. 6], the second neighborhood is obtained, for example, from the first neighborhood by adapting the shape of the latter to the boundary of the characteristic region Rj so that the points of the second neighborhood all belong to the characteristic region Rj.
[0106] Once determined, the input vector is processed, in a step (3), by the convolutional neural layer Ck to generate as output the PSn point of the CCSi, CCS2, ccs3 maps
[0107] Fig. 7 is a logic diagram representing an example of a coding process that can be implemented by the coding device of Fig. 1 when the NNSYN synthesis neural network is, for example, that shown in Fig. 3.
[0108] According to this embodiment, the signal is a two-dimensional image, each sample to be coded is therefore a pixel Pn with coordinates (xn, yn).
[0109] The encoding takes place in three main phases:
[0110] In a first phase, called the segmentation phase, the segmentation Sg of the input signal I(Pn) into J (greater than or equal to 2) zones Zo' is carried out.
[0111] According to a first example of segmentation, the image is divided into J regular zones, for example of identical size and shape (except possibly at the edges of the image). For example, the image can be divided into rectangular zones whose boundaries are horizontal and vertical. Such a division corresponds to the concept of "tiles" implemented by encoding standards such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC.
[0112] According to a second example of segmentation, the image is divided into J zones corresponding to samples (or blocks) traversed in a lexicographical order. Such a division corresponds to the concept of "slices" implemented by coding standards such as H.264 / AVC, H.265 / HEVC and H.266 / VVC.
[0113] According to a third example of segmentation, the image is divided into semantic zones, for example, a background and a foreground. This can be done manually by an operator. Segmentation can also be automatic or semi-automatic, depending on the segmentation algorithm used. It should be noted that there are no restrictions related to the type of segmentation algorithm used.
[0114] Advantageously, such a division into zones allows, during the subsequent decoding of the signal I(Pdn), the decoding of the zones to be decoded Zodj to be parallelized by distributing the decoding load equally on each of the decoders and / or processors.
[0115] In a second phase, called the construction phase, a learning process is performed to determine, for an input signal I(Pn), the values of the FM maps and the Wk parameters to optimize an overall cost function. The learning is, for example, performed by gradient descent, followed by updating the parameters of the NNSYN synthesis neural network and the values of the FM feature maps. As is known in the prior art, the cost function can be of the rate-distortion type, or rate-distortion type, or perceptual type. To measure the rate R, it is necessary to simulate the encoding of the J ZoLj zones of the FM maps and then measure the associated encoding rate (the size of the B2 flow). According to one embodiment, the encoding of the Wk parameters is not simulated because their influence is less significant than that of the feature maps.According to one embodiment, the encoding of the parameters Wk is also simulated and the associated throughput (the size of the stream Bl) is measured. To measure the distortion D, it is necessary to simulate the encoding and then the decoding of at least a part of the image I, to obtain at least one pixel P'n resulting from a simulation of encoding and then decoding, and then to measure the difference between this part. of the input image I(Pn) and a corresponding part of the encoded and then decoded image I(P'n).
[0116] Then, during a third phase, called the coding phase, the Sg segmentation, the data of each ZolJet zone and its Wk parameters are encoded to produce the coded values Sgc and Wck before transmission or storage. These constitute the compressed representation of the input signal I(Pn), this compressed signal being able to be decoded zone by zone.
[0117] We will now describe the steps of a coding process according to one embodiment of the invention.
[0118] During a step E20, a signal I(Pn) to be coded, comprising a plurality of N samples Pn, is provided as input to the process.
[0119] During a step E21, the segmentation Sg of the input signal I (Pn) into J (greater than or equal to 2) zones ZoJ is carried out.
[0120] During step E22, the M FM maps of the first group are initialized. Subsequently, the Wk parameters of the NNSYN synthesis neural network and the values of the FM maps must be optimized during the construction phase.
[0121] According to one embodiment, the FM cards; are of the same resolution as the input signal I (Pn) and therefore each have the same number of values N as there are samples Pn to be coded.
[0122] According to one embodiment, the FM cards; have a resolution less than or equal to that of the input signal I(Pn) and therefore include, for at least one of them, a number N' of values to be coded less than N. According to a variant, the first FM card; has the resolution of the images and each subsequent card has half the resolution of the previous one.
[0123] According to one embodiment, several FM cards; are of the same resolution, lower than that of the input signal I(Pn).
[0124] According to one embodiment, the FM maps are transformed to provide a second group of transformed feature maps FMS. In this embodiment, the feature vectors are preferably extracted from the transformed maps of the second group, and not directly from the maps of the first group. Thus, in this embodiment, the feature vectors are indirectly extracted from the maps of the first group. The maps of the second group are not coded; they serve only for the construction of the feature vectors.
[0125] According to one embodiment, the FM cards are initialized with predefined constant values.
[0126] According to another embodiment, the feature cards are initialized by a set of random real values.
[0127] The FM feature cards; of the first group are subsequently updated, or refined, during an E23 step, by the encoder update module during its learning.
[0128] During step E24, the J zones zo]J ^es FM maps of the first group are encoded by the FMC module of the encoder. During the construction phase, this operation is a coding simulation. During the encoding phase, this operation is an actual coding, and the encoded values constitute the B2 stream. The simulation may be identical to the actual coding, but it may also be different (for example, simplified). For this coding, a technique for predicting a feature map value by its neighborhood is used, as will be described, for example, in support of [Fig. 11].
[0129] In one embodiment, the J zones ZoLj of the FM maps are encoded in the indexing order of the J zones and in a predefined order of the associated feature maps (i.e., fm{, FM1, FM2, FM3, FM4, FM5, ...) and the values of each map, for example, lexicographic. Each zone ZoLj of the FM maps undergoes entropy coding.
[0130] Entropic coding of all zones of all FM maps; produces a compressed stream B2 whose throughput is subsequently measured during a step E29.
[0131] During a step E25, according to one embodiment, the M cards of the first group FM; are transformed by the SE module to generate cards of the second group FMSi at the resolution of the images of the input sequence.
[0132] According to one embodiment, M FMS cards are generated.
[0133] According to one embodiment, each FM card; is transformed into an FMS card;.
[0134] According to one embodiment, at least one FM card; is of lower resolution to that of the images in the sequence to be encoded, and the transformation operation includes oversampling so that the transformed FMSi map contains the same number of samples as the images in the sequence. Oversampling consists of adding values to the FMS maps to achieve the resolution of the images in the input sequence. It can be simple (by nearest neighbor replication) or involve interpolation (linear, polynomial, filtered, etc.).
[0135] During step E26, and taking into account the segmentation Sg of the input signal I(Pn) into J zones Zo', values are extracted by the XTR module from the FM cards; or optionally FMSi transformed. This extraction is performed based on the coordinates (xn, yn) of the sample Pn of the input signal and optionally the mask MZo'. It can also be performed based on the resolution of the card considered.
[0136] According to one embodiment, the characteristic Zn vector results directly from this extraction.
[0137] The samples to be coded are, for example, processed sequentially, from n=1 to n=N.
[0138] According to one embodiment, during a step E27, the characteristic vector Zn is constructed by the TT module from the values extracted from the FM or FMS maps for each sample Pn with coordinates (xn, yn) of the input signal. The processing may include quantization of the values extracted from the FM maps or of the resulting vector Zn, if necessary. The processing may include other operations, such as filtering, scaling, the application of any function, preferably monotonic, etc.
[0139] In one embodiment, Zn is an M-tuple (zb z2,..., Zj), consisting of the values of the FM; or FMS; maps located at the coordinates (xn, yn) of the current pixel Pn, as will be illustrated in support of [Fig.8].
[0140] In one embodiment, Zn is an M-tuple constructed from values taken from the FM maps; at coordinates that may be different for each map. For example, if the FM maps; are at different resolutions because they have been downsampled, the coordinates are adjusted (by scaling) to match the resolution of each map.
[0141] During a step E28, the vector Zn is processed by the NNSYN synthesis neural network to generate as output a vector representative of the sample Pn to be coded, according to an embodiment the triplet (R, G, B) of the sample P'n (the sample Pn coded then decoded).
[0142] The structure and parameters Wk of the synthesis neural network are initialized, for example, during the first iteration of this step. These parameters are subsequently updated, or refined, during the construction phase, in subsequent iterations of the process.
[0143] According to one embodiment, the parameters of the synthesis neural network and / or the prediction neural network are initialized by predefined values known to give a satisfactory result (for example, following training on a corpus of images).
[0144] According to another embodiment, the parameters of the synthesis neural network and / or the prediction neural network are initialized by a set of random values.
[0145] During step E29, the parameters Wk of the NNSYN synthesis neural network are quantized and encoded. During the construction phase, this operation is a coding simulation. During the encoding phase, this operation is the actual encoding, and the encoded values constitute the stream Bl. The simulation may be identical to the actual encoding, but it may also be different (for example, simplified). Any known technique can be used for this purpose, for example, the network coding standard. of neurons proposed by the MPEG-7 Part 17 standard, also called NNR (Neural Network Representation). Note that in this case, the amount of degradation that the encoding introduces to the Wk weights must be chosen.
[0146] During an E30 step, a performance measure is evaluated.
[0147] To this end, the coding simulation rates associated with the feature maps of the first group (simulation of the B2 flow by coding the J zones pj^jj of the FM maps;) and optionally with the parameters of the neural network(s) (simulation of the B1 flow by coding the Wk parameters) are measured.
[0148] In one embodiment, the cost function is of the rate-distortion type, denoted (D+L*R), where D, for example, is the root mean square error measured between the input signal and the decoded signal (or the error measured on a subset of the signal samples). In another example, D is calculated from a perceptual function such as SSIM (for Structural SIMilarity) or MSSSIM (for Multi-scale Structural SIMilarity). In one embodiment, R is the simulated rate of stream B1; in another embodiment, R is the total rate used to encode this image, i.e., the sum of the simulated rates of B1 and B2. L is a parameter that controls the rate-distortion trade-off. Other cost functions are possible.
[0149] As long as the cost function has not reached its minimum, or a maximum number of iterations of the cost function minimization algorithm has been reached, the performance measurement is not satisfactory, and the process is repeated from step E23. This minimization can be performed by a mechanism known as gradient descent with parameter updates during step E23 for the feature map values and E29 for the network(s) parameters.
[0150] During an EF step, if the cost function has reached its minimum, or if a maximum number of iterations of the cost function minimization algorithm has been reached, training stops. If a coded version corresponding to the last simulation of the synthetic neural network parameters (Wk) and feature maps (pj^i) is available, streams B1 and B2 can be formed from it. In another embodiment, the actual coding of the updated synthetic neural network parameters (Wk) and feature map values (FM;) is performed at this step to produce the encoded parameters Wcket pjypd that constitute streams B1 and B2. Furthermore, at this step, the coding of the segmentation Sg is performed to produce an encoded segmentation of the J masks MZoc' that constitutes stream B3.
[0151] The B2 stream comprises sub-streams corresponding to each area of the feature maps. When encoding each latent area, it is possible, in a The preferred implementation method is to place location information for sub-flows within the overall flow at an identifiable point in the flow. This location information can be: - a flow pointer (which indicates the starting address of each sub-flow in the overall flow), or - a marker (a series of bits otherwise prohibited, which allows traversing the stream to find the beginning of each sub-stream), or - any other means of identifying a sub-part of a coded stream.
[0152] Streams B1 and B2 can be multiplexed and / or concatenated to produce a final stream. In one embodiment, stream B3 of the coded segmentation and stream B1 of the coded parameters of the neural network(s) are stored or transmitted before stream B2, so that they can be decoded before stream B2.
[0153] It should be noted that a single NNSYN neural network is used for the coding of each Zo' area.
[0154] Figure 8 illustrates a coding method used in one embodiment of the invention.
[0155] In this embodiment, there are 4 FM cards generated. In a preferred mode, there are 7.
[0156] The first FMi map has the same resolution as the image I(Pn), and therefore contains WxH variables, where W represents the width of the image in pixels, and H its height. The second FM2 map has half the resolution (in each dimension) of the FMi map. Each additional map has half the resolution of the previous map. This structure reduces the number of variables in the feature maps, which facilitates coding and learning while minimizing the coding cost.
[0157] The FM2 map is oversampled by the SE module by a factor of 2 in each dimension, according to a method illustrated in Figure 6. The FM3 map is oversampled by a factor of 4 in each dimension, and the FM4 map by a factor of 8 in each dimension. The FMi map is not affected by the oversampling. (FMS^FMj).
[0158] The resulting FMS maps are of the same resolution as the image I(Pn), and therefore each have WxH values, where W represents the width of the image in pixels and H its height (N=WxH).
[0159] Other types of structure are possible, for example one can use a different reduction rate of one half between the cards (one quarter, or one third, etc.).
[0160] In this embodiment, the vector Zn is a 4-tuple (zi...z4) consisting of the values extracted from the FMS maps located at the coordinates (xn, yn) of the current pixel Pn. The vector Zn consisting of the extracted (quantized) values from the FMS maps is processed by The NNSYN synthesis neural network generates a second output vector; in this example, the output vector is the triplet (R, G, B) of the encoded and then decoded pixel P'n. This triplet is inserted into the decoded image I(P'n) at the positions (xn, yn) of the color components (R', G', B').
[0161] In another embodiment, not shown, the Znest vector is extracted directly from the FM layers;, at positions recalculated according to the size of the maps, then the extracted values are optionally processed and quantified after extraction.
[0162] The [Fig.9] is a logic diagram representing an example of a decoding process for a Zodj decoding zone which can be implemented by the DEC decoding device of the [Fig.2] when the synthesis neural network NNSYN' is for example that presented in the [Fig.3].
[0163] During a step F20, the Bl stream, a portion of the B2 stream (that corresponding to the feature maps of the Zodj zone), and the B3 stream are extracted from the encoded stream. They contain, respectively, Wck parameters, an Sgc segmentation, and encoded representations of the M maps of the first group representative of the Zodj zone to be decoded.
[0164] During an F21 step, an MZoj segmentation mask is generated by decoding the values of the encoded segmentation.
[0165] During a step F22, the M pjypjl cards are generated by decoding the FMc-- values. In one embodiment, the pxqd' cards are decoded in the order (FMdJp FMdb • • • FMdp' ct 'cs values of each card in a predefined order, for example lexicographical, possibly taking into account (depending on the coding technique implemented by the ENC encoder) the MZoj mask.
[0166] According to embodiments as described for the encoder: - The p cards have the same resolution as the I(Pdn) signal to be reconstructed, that is to say they contain N=WxH values. - The pM^p maps have a resolution less than or equal to that of the I(Pdn) signal to be reconstructed. - Several p^jC[J cards have the same resolution, lower than the resolution of the signal.
[0167] In step F23, according to one embodiment, the M maps of the first group FMd- are transformed by the module SE' to generate maps of the second group FMS'- at the resolution of the input images. This step is similar to step E25 which has been described for the encoder in support of [Fig. 5], and the embodiments apply. In particular:
[0168] According to one embodiment, M p^g'J cards are generated.
[0169] According to one embodiment, each card pj^i is transformed into a card 1 1
[0170] According to one embodiment, at least one map pM(ÿ) has a resolution lower than that of the images in the image to be encoded, and the transformation operation includes oversampling so that the transformed map has the same number of samples as the input image. The oversampling consists of adding values to the maps pMS'j to achieve the resolution of the input image. It can be simple (by nearest neighbor replication) or involve interpolation (linear, polynomial, filtered, etc.).
[0171] The transformation may optionally include inverse quantization of the extracted values, if necessary. However, inverse quantization is not mandatory.
[0172] During step F24, values are extracted by the XTR' module from the transformed PjVfji or, optionally, pjyJS'j maps. This extraction is performed based on the coordinates (xn, yn) of a sample Pn of the input signal. It can also be performed based on the resolution of the map in question. This step is similar to step E26, which was described for the encoder in support of [Fig. 7], and the embodiments apply.
[0173] In one embodiment, Zdn is an M-tuple (zb z2,..., Zj), consisting of the values of the pj^yj or FMS'- maps located at the coordinates (xn, yn) of a current pixel Pdn, as will be illustrated in support of [Fig. 10].
[0174] The samples to be decoded are, for example, processed in sequential order relative to the MZof mask
[0175] According to one embodiment, in step F35, a vector Zdn is constructed by the module TT' from the values extracted from the pMjJ maps of the first group or the pjÿJS'j maps of the second group, for each sample Pdn of coordinates (xn, yn) to be decoded, as a function of the coordinates (xn, yn). This step is similar to step E27, which was described for the encoder supporting [Fig. 7], and the described embodiments apply. The extraction may include inverse quantization of the extracted values or of the constructed vector Zdn, if necessary.
[0176] During step F26, the Wdk parameters of the NNSYN' synthesis neural network are generated by decoding the Wck values of the Bl stream. Any known decoding technique corresponding to the encoding technique that has been used by the encoder. The NNSYN' synthesis neural network is similar to the NNSYN synthesis network, that is to say, it has the same structure and the same parameters, except for the encoding, which can be done with or without loss.
[0177] According to one embodiment, stream B1 is decoded before streams B2 and B3, so that the NNSYN' synthesis neural network is available before decoding the samples. Similarly, stream B3 is decoded before stream B2 so that the mask of the area to be decoded is available before decoding the samples.
[0178] During step F27, the vector Zdn is processed by the NNSYN' synthesis neural network to generate as output a second vector representing the sample Pdn to be decoded, according to an embodiment a triplet which is injected into the image of the decoded area Zodj (Pdn) at the positions (xn, yn) of the color components (Rd, Gd, Bd). This step is similar to step E28 which was described for the encoder in support of [Fig. 5].
[0179] When all the samples of the signal have been processed, the decoded signal corresponding for example to the Zodj(Pdn) zone is available.
[0180] It should be noted that only one neural network NNSYN' is used during decoding regardless of the Zodj area being decoded.
[0181] Fig. 10 illustrates a method for decoding a Zodj zone used in an embodiment of the invention.
[0182] In this embodiment, there are 4 decoded cards. In a Their preferred mode of operation is 7 in number.
[0183] In this embodiment, the decoded cards are representative 1 only from the Zodj zone. In other words, only the data from the ZoLj zone (corresponding in the latent domain to the Zodj zone) were decoded.
[0184] In this embodiment, the first FMjj map has the same resolution as the image I, and therefore comprises WxH variables, where W represents the width of the image in pixels, and H its height. The second FM^ map has half the resolution (in each dimension) of the FMdh map. Each additional map has half the resolution of the preceding map. This structure reduces the number of variables in the feature maps, thus facilitating decoding while minimizing the coding cost.
[0185] Map pMd1 is oversampled by a factor of 2 in each dimension, using any oversampling method available to a person skilled in the art. Map FMd^ is oversampled by a factor of 4 in each dimension, and map F]\p|i by a factor of 8 in each dimension.
[0186] The pM$'i cards have the same resolution as the image to be decoded, and therefore have WxH values, where W represents the width of the image in pixels, and H its height.
[0187] In this embodiment, the vector Zdn is a 4-tuple (zi...z4) consisting of the values of the pMS'i maps located at the coordinates (xn, yn) of the current pixel Pdn. The vector Zdn is optionally dequantized and then processed by the NNSYN' synthesis neural network to generate as output the triplet (R, G, B) representative of the Pdn sample to be decoded. The triplet (R, G, B) is inserted into the decoded image I(Pdn) at the coordinates (xn, yn) in the color components (Rd, Gd, Bd).
[0188] The [Fig. 11] is a flowchart representing a method of coding feature cards which can be implemented by the coding device of [Fig.1] and by the coding process of [Fig.5].
[0189] These steps constitute sub-steps of step E30 described previously with support of figure 7. They aim to encode a current value Vnd' of a point of a characteristic zone 2qL- of a feature map pMjJ of the first group being processed using values from the neighborhood.
[0190] In a substep E301, a neighborhood vector Cn is established, comprising values close to the value Vn. These neighboring values may be located in the same map and / or in a different map from the plurality M of maps. This neighborhood vector consists of a number C of values, or data, corresponding to neighborhood values (for example, C=10). These values must be known to the encoder and the decoder; therefore, they must be located in a causal neighborhood of the value Vn.
[0191] According to a first embodiment, these values are used to determine the context of an entropy encoder to encode the current value during an E303 step. This encoder can be a CAB AC (Context-adaptive binary arithmetic coding) encoder. This type of encoder is well known to those skilled in the art. It is notably used in the H.265 / HEVC video compression standard. It is an arithmetic encoder with lossless compression. It decomposes all non-binary symbols into binary symbols. Then, for each bit, the encoder selects the most suitable probability model and uses a context to optimize the probability estimation. This context can be defined by information from neighboring elements. Arithmetic coding is then applied to compress the resulting data. As is known to those skilled in the art, there are several ways to use the neighborhood vector to produce context information.For example, one can count the number of neighboring values other than zero, and associate a context with each number. Alternatively, one can perform comparisons between several neighboring values, and associate a given context with an ordering configuration between them. neighboring values, for example by ranking neighboring values in ascending order, and associating a context with each possible order.
[0192] In a second embodiment, the neighborhood is used to predict, during a step E302, the current value from an autoregressive model. It is recalled that an autoregressive model predicts a sample from a series based on its past values. In this embodiment, the past values are constituted by the context, and the difference between the predicted variable and the actual value is quantified and then entropically coded during step E303.
[0193] At the end of the process, the current coded value Vcnde of the card being processed is coded.
[0194] The [Fig. 12] is a logic diagram representing a method for decoding feature cards which can be implemented by the decoding device of [Fig.2] and by the decoding process of [Fig.7]
[0195] These steps constitute sub-steps of step F22 described previously in support of figure 7. They aim to decode a current value of a point of a characteristic zone 2oL^ a feature map pj^jî of the first group being processed using values from the neighborhood.
[0196] In a substep F221, a neighborhood vector Cdn is established, comprising values close to the value Vdn. This step is similar to the previously described step E301 and the same embodiments apply. This neighborhood vector consists of a number C of values, or data, corresponding to neighborhood values (for example, C=10) located in the same map and / or in a different map of the plurality M of maps of a point of a characteristic zone ZoL- of a feature map. These values, being in a causal neighborhood of the value Vdn, are known to the decoder.
[0197] According to a first embodiment, these values are used to determine the context of an entropy decoder for decoding the current value during an F223 step. This decoding is similar to that used in the encoder, for example, CAB AC. The use of the neighborhood to produce context information is similar to that chosen in the encoder. For example, one can count the number of non-zero neighbor values and associate a context with each number. Alternatively, one can perform comparisons between several neighbor values and associate a given context with an ordering configuration among the neighbor values, for example, by ranking the neighbor values in ascending order and associating a context with each possible order.
[0198] In a second embodiment, the neighborhood is used to predict the current value from an autoregressive model during a step F222. In this embodiment, The values passed are constituted by the context, and the difference between the predicted variable and the actual value is quantified and then entropically coded during step F223.
[0199] At the end of the process, the current decoded value Vdnde of the FMd card being processed is decoded.
[0200] It should also be noted that the invention is not limited to the embodiments described above. It will indeed be apparent to those skilled in the art that various modifications can be made to the embodiments described above, in light of the information just disclosed to them.
[0201] For example, NNSYN / NNSYN' synthesis neural networks can be recurrent neural networks.
[0202] In another example, the synthetic neural networks can consist of one or more convolutional neural networks, followed by an MLP and then followed by one or more convolutional neural networks, each of the convolutional neural networks being associated with a convolution kernel of predefined size, for example, size 3x3. In these examples, the acquisition of the Zn / Zdn vectors is adapted to the topology of the NNSYN / NNSYN' synthetic neural networks in a similar way to the adaptation of the input vector Vcon shown with reference to Figures 4 to 6.
[0203] In the detailed presentation of the invention given above, the terms used shall not be interpreted as limiting the invention to the embodiments set forth in this description, but shall be interpreted as including all equivalents which can be foreseen by a person skilled in the art by applying their general knowledge to the implementation of the teaching which has just been disclosed to them.
Claims
Demands
1. A method for encoding a region (Zo^) of a signal (I(Pn)), called the region to be encoded, said region to be encoded comprising a plurality of samples (Pn) to be encoded, said encoding method comprising the following steps: - obtaining a first group of at least one latent value map (FM;) representative of said signal (I(Pn)), - obtaining a mask (MZoJ) in said signal (I(Pn)) of said zone to be coded (Zo*), - for at least one sample of said zone to be coded (Zo^, called current sample (Pn), associated with a position (xn, yn) in said signal (I(Pn)) to be coded: • the construction of a characteristic vector (Zn ) from said latent value maps (FM; ) of said first group, as a function of said position (xn, yn) of said current sample (Pn), • the processing of said characteristic vector (Zn) by an artificial neural network, called a synthetic neural network (NNSYN), said synthetic neural network being defined by a set of parameters (Wk) and comprising at least one synthetic neural layer (ck) to obtain, at the output of said synthetic neural network, a vector (P'n) representative of a decoded value of the current sample (Pn), • the updating of at least one value of one of said latent value maps of said first group and / or of at least one parameter of said synthetic neural network, as a function of a coding performance measure, - obtaining a second group of latent value maps representing at least the said zone to be coded (Zo*) from said mask (MZoJ), - the coding of the second group of latent value maps (FMcp' - the encoding of said mask (MZoJ), and
2. - the coding of said set of parameters (Wk) of said synthetic neural network. said processing comprising the application of said at least one neural synthesis layer to transform at least one group of latent input maps (LIMs), referred to as input maps, into at least one group of latent output maps (LIMs), referred to as output maps, said application comprising: - obtaining a characteristic region (Rj) of said zone to be decoded in said input maps from said mask (MZoj), said characteristic region (Rj) comprising the points associated with said mask, - for at least one point (Ptn) of said area to be decoded, called point to be decoded: • obtaining an associated point (Pan) to the said point to be decoded (Ptn) in at least one of the latent input maps (CCE), • the evaluation of a proximity criterion of said associated point (Pan) with respect to a boundary of said characteristic region, • the construction of an input vector (Vcon) of said neural synthesis layer from a neighborhood of said associated point (Pan) and said evaluation of said proximity criterion, and • the processing of said input vector by said neural synthesis layer to obtain a point (Psn) of at least one of the output maps, Decoding method of an area (Zodj), said area to be decoded, of a signal comprising at least two areas, said area to be decoded comprising a plurality of samples (Pdn) to be decoded, said decoding method comprising the following steps: - the decoding (F22) of a group of latent value maps (FMdp) representative of said zone to be decoded (Zodj) of said signal, - obtaining (F20) a mask (MZoJ) of said area to be decoded in said signal (I(Pn)),
3. - the decoding (F26) of a set of parameters (Wdk) representative of a neural network (NNSYN'), called a synthetic neural network comprising at least one synthetic neural layer, - the processing (F27) of said group of decoded latent maps (FMd^ P31) by said synthesis neural network (NNSYN') to produce at the output of said synthesis neural network at least said decoding area (Zodj), said processing comprising the application of said at least one synthesis neural layer to transform at least one group of input latent maps (CCE), said input maps, into at least one group of output latent maps (CCS), said output maps, said application comprising obtaining a characteristic region (Rj) of said decoding area in said input maps from said mask (MZoJ), said characteristic region (Rj) comprising the points associated with said decoding area (Zodj) and for at least one point (Ptn) of said decoding area, said decoding point,the processing of an associated point (Pan) to be decoded in at least one of the latent input maps (CCE) by said neural layer as a function of its proximity to a boundary of said characteristic region. A method for decoding a zone (Zodj) of a signal according to the preceding claim, wherein the processing step of said associated point (Pan) comprises: - obtaining an associated point (Pan) to be decoded in at least one of the latent input maps (CCE), - the evaluation of a proximity criterion of said associated point (Pan) in relation to a boundary of said characteristic region, - the construction of an input vector (Vcon) of said neural synthesis layer from a neighborhood of said associated point (Pan) and said evaluation of said proximity criterion, and - the processing of said input vector by said neural synthesis layer to obtain a point (Psn) of at least one of the output maps.
4. Method for decoding an area (Zodj) of a signal according to the preceding claim wherein said neighborhood of said associated point (Pan) is independent of said evaluation of the proximity criterion.
5. Method for decoding an area (Zodj) of a signal according to claim 3 wherein said neighborhood of said associated point (Pan) is selected according to said evaluation of the proximity criterion.
6. A method for decoding a zone (Zodj) of a signal according to any one of claims 3 to 5, wherein said neighborhood of said associated point (Pan) is selected based on said associated point (Pa
7. nj- Method of decoding an area (Zodj) of a signal according to any one of claims 3 to 6 wherein during the construction of the input vector (Vcon), a component of the input vector (Vcon) is associated with a point of said neighborhood, the value of said component being equal to the value of said associated point if said associated point is a point of said characteristic region (Rj) and to a replacement value otherwise.
8. Method of decoding a zone (Zodj) of a signal according to the preceding claim wherein said replacement value is dependent on the points of said characteristic region (Rj) or of the points associated with said neighborhood belonging to said characteristic region (Rj).
9. Decoding method according to claim 7 wherein said replacement value does not depend on the points of said characteristic region (Rj).
10. Method for decoding an area (Zodj) of a signal according to any one of claims 3 to 9 wherein said proximity criterion is a distance.
11. Method for decoding a zone (Zodj) of a signal according to any one of claims 2 to 9 wherein said synthesis neural layer is a convolutional neural layer.
12. A signal coding device (NCD) for a signal (I(Pn)) comprising a plurality of samples (Pn) to be coded, characterized in that said coding device is configured to implement: - obtaining a first group of at least one latent value map (FM;) representative of said signal (I(Pn)), - obtaining a mask (MZoJ) in said signal (I(Pn)) of said zone to be coded (Zo1), - for at least one sample of said zone to be coded (Zo*), called current sample (Pn), associated with a position (xn, yn) in said signal (I(Pn)) to be coded: • the construction of a characteristic vector (Zn ) from said latent value maps (FM; ) of said first group, as a function of said position (xn, yn) of said current sample (Pn), • the processing of said characteristic vector (Zn) by an artificial neural network, called a synthetic neural network (NNSYN), said synthetic neural network being defined by a set of parameters (Wk) and comprising at least one synthetic neural layer (ck) to obtain, at the output of said synthetic neural network, a vector (P'n) representative of a decoded value of the current sample (Pn), • the updating of at least one value of one of said latent value maps of said first group and / or of at least one parameter of said synthetic neural network, as a function of a coding performance measure, - obtaining a second group of latent value maps representing at least the said zone to be coded (Zo*) from said mask (MZoJ), - the encoding of the second group of latent value maps <FMcp’ - the encoding of said mask (MZ©*), and - the coding of said set of parameters (Wk) of said synthetic neural network.
13. said processing comprising the application of said at least one neural synthesis layer to transform at least one group of latent input maps (LIMs), referred to as input maps, into at least one group of latent output maps (LIMs), referred to as output maps, said application comprising: - obtaining a characteristic region (Rj) of said zone to be decoded in said input maps from said mask (MZoj), said characteristic region (Rj) comprising the points associated with said mask, - for at least one point (Ptn) of said area to be decoded, called point to be decoded: • obtaining an associated point (Pan) to the said point to be decoded (Ptn) in at least one of the latent input maps (CCE), • the evaluation of a proximity criterion of said associated point (Pan) with respect to a boundary of said characteristic region, • the construction of an input vector (Vcon) of said neural synthesis layer from a neighborhood of said associated point (Pan) and said evaluation of said proximity criterion, and • the processing of said input vector by said neural synthesis layer to obtain a point (Psn) of at least one of the output maps, Decoding device (DEC) of an area (Zodj) of a signal comprising at least two areas, said area to be decoded, said area to be decoded comprising a plurality of samples (Pdn) to be decoded, characterized in that the decoding device is configured to implement: - decoding a group of latent value maps (FMdp representative of said zone to be decoded (Zodj) of said signal, - obtaining a mask (MZo1) of said zone to be decoded in said signal (I(Pn)), - the decoding of a set of parameters (Wdk) representative of a neural network (NNSYN'), called
14. synthetic neural network comprising at least one synthetic neural layer, - the processing of said group of decoded latent maps (FMdp P31 redistributes synthetic neural network (NNSYN') to produce at the output of said synthetic neural network at least said decoded area (Zodj), said processing comprising the application of said at least one synthetic neural layer to transform at least one group of input latent maps (CCE), said input maps, into at least one group of output latent maps (CCS), said output maps, said application comprising obtaining a characteristic region (Rj) of said decoded area in said input maps from said mask (MZoJ), said characteristic region (Rj) comprising the points associated with said decoded area (Zodj) and for at least one point (Ptn) of said decoded area, said decoded point,the processing of an associated point (Pan) to be decoded in at least one of the latent input maps (CCE) by said neural layer as a function of its proximity to a boundary of said characteristic region. Computer program comprising instructions for carrying out the steps of an encoding process according to claim 1 or a decoding process according to any one of claims 2 to 11 when said program is executed by a computer.
Citation Information
Patent Citations
Coding concept allowing efficient multi-view / layer coding
EP2984839B1
Method and device for encoding and decoding images.
FR3143245A1
Method and apparatus for decoding with signaling of feature map data
US20230353764A1
AU2016259446A1