Method and device for encoding and decoding images
By constructing characteristic maps and applying a dilation rule to create independently decodable regions, the method improves interaction with semantic content and enables parallel decoding of images.
Patent Information
- Application Number
- FR2024006996
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2026-01-02
AI Technical Summary
Existing image encoding methods, both classical and neural, do not allow for the creation of independently decodable areas in the coded signal, limiting interaction with semantic content and parallel decoding possibilities.
A method and device for encoding and decoding images by constructing characteristic maps, applying a dilation rule to create a dilated characteristic zone, and encoding a mask, enabling independent decoding of regions.
This approach enhances the ability to interact with the semantic content of the coded signal and allows for parallel decoding processes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for encoding and decoding images. Prior art.
[0001] The invention relates to the general field of coding one-dimensional or multidimensional signals. It relates more particularly to the compression of digital images or videos.
[0002] Digital videos are generally subject to source coding aimed at compressing them in order to limit the resources required for their transmission and / or storage. Numerous coding standards exist, such as the ITU / MPEG standards (H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc.) and their extensions (MVC, SVC, 3D-HEVC, etc.). In these approaches, image encoding is generally performed by predicting pixels using previously encoded and then decoded pixels present in the image being encoded, in which case it is called "Intra prediction," or previously encoded images, in which case it is called "Inter prediction."
[0003] In addition to these classic approaches, approaches based on artificial intelligence, and in particular neural networks, tend to develop.
[0004] Some neural approaches, starting from an input signal, for example an image, train a so-called synthetic neural network on characteristic vectors associated with a position of a sample of the input signal to be encoded. These characteristic vectors are constructed from feature maps that may be at the resolution of the input signal, or at a lower resolution. During training, or construction, the parameters of the neural network and the values of the feature maps are updated according to a performance measure, for example, a data rate-distortion type. When training is complete, i.e., when the performance measure obtained is satisfactory, the actual encoding of the synthetic neural network parameters and the values of the feature maps can be performed and stored or transmitted to the decoder.The decoding of the coded signal is then carried out by applying the synthesis neural network to the feature maps.
[0005] One drawback of the classical and neural approaches described above is that they do not allow the creation of independently decodable areas in the coded signal, which limits not only the ability to interact with the semantic content of the coded signal but also the possibility of limiting the memory required at the decoder level or of parallelizing the decoding of this coded signal.
[0006] There is therefore a need for a solution to improve upon the classical and neural approaches described above. Summary of the invention
[0007] The invention relates to a method for encoding a region of a signal, called the encoding region, said encoding region comprising a plurality of samples to be encoded, said encoding method comprising the following steps: - a step of constructing a group of at least one characteristic map representative of said signal, - a step of obtaining a mask of said area to be coded and a rule for expanding a region of the group of at least one feature map, - a step of obtaining a feature area of the group of at least one feature map as a function of said mask, - a step of determining a dilated characteristic zone within said group of at least one feature map by applying said dilation rule to said characteristic zone, - a step of encoding the value of the points of said dilated characteristic zone, and - a step of encoding said mask.
[0008] The invention also relates to a method for decoding a region of a signal, called the region to be decoded, said region to be decoded comprising a plurality of samples to be decoded, said decoding method comprising the following steps: - a step of decoding a mask of said area to be decoded, - a step of obtaining a dilation rule for a region of the group of minus a characteristics card, - a step of determining a dilated characteristic area within said group of at least one feature map according to said dilation rule and said mask, - a step of decoding the value of the points in said dilated characteristic zone, and - a synthesis step of said area to be decoded from said decoded values and said mask.
[0009] For the purposes of this invention, encoding, or "coding," means the operation of representing a set of samples in a compact form, for example, using a digital binary stream. Decoding means the operation of processing a digital binary stream to produce decoded samples.
[0010] By "sample" of the signal, we mean a value taken from the signal. Sampling the signal produces a sequence of discrete values called samples. In the case of an image signal, the sample is called a pixel, which can be, for example, a color pixel traditionally represented by a triplet of values, for example (R, G, B) or (Y, U, V). The position of the sample is located by its abscissa (x) and ordinate (y) coordinates in the image.
[0011] By "signal comprising a plurality of samples" is meant a signal with one (audio, sound), two (image) or more than two (stereoscopic image, multiscopic image, image associated with a depth map, video, etc.) dimensions. Depending on this dimensionality, the sample has one, two, or more coordinates in the signal. In the case of an image signal, the position of the sample is located by its abscissa (x) and ordinate (y) coordinates.
[0012] By "feature maps" is meant an abstract representation of the signal comprising a plurality of variable scalar data, discrete or not, also called values, for example real or integer numbers, signed or unsigned. These maps are also known as "latent representations" or "representations in the transformed domain".
[0013] “By region expansion rule”, we mean a procedure allowing to Transforming a region involves increasing its size (one-dimensional region), area (two-dimensional region), volume (three-dimensional region), or space (higher-dimensional region). A rule for expanding a region is, for example, the Minkowski sum, which defines the morphological expansion of that region by a structuring element.
[0014] Generally speaking, the steps of an encoding or decoding process should not be interpreted as being linked to a notion of temporal succession. In other words, the steps may be carried out in a different order than that indicated in the independent encoding or decoding claim, or even in parallel.
[0015] The coding method according to the invention encodes a region of a signal from a representation of that signal in the form of characteristic maps by identifying in these characteristic maps the set of values necessary to subsequently decode said region. In other words, encoding this set of values, defined by an expanded characteristic region, makes it possible to obtain a coded representation of the region that can subsequently be decoded independently of any other part of the signal.
[0016] Furthermore, this dilated characteristic area can be determined using a mask of the area to be encoded in the signal domain and depends predictably of the type of encoding and / or decoding used to encode and / or decode the feature cards.
[0017] Thus, such a coding / decoding process improves the ability to interact with the semantic content of the coded signal and offers the possibility of parallelizing the decoding process of the entire signal.
[0018] According to one embodiment of the coding process, the step of obtaining a group of at least one feature map of the coding process comprises, for at least one sample, called the current sample, of the signal to be coded, associated with a position in the signal to be coded, a step of constructing a feature vector from said group of at least one feature map, as a function of said position of said current sample and a step of processing said feature vector by an artificial neural network, called a synthesis neural network defined by a set of parameters, to provide a vector representative of a decoded value of the current sample, and a step of updating at least one value of the group of said at least one feature map and / or of at least one parameter of said network, as a function of a coding performance measure.
[0019] By "characteristic data vector constructed from feature maps as a function of a position" is meant a vector consisting of one or more elements, or data, preferably discrete, the data being constructed from the feature maps at a position determined by that of the sample being processed in the signal. This characteristic vector is the one applied to the input of the synthesis neural network. For example, in the case of a one-dimensional audio signal, such a vector can be constructed from a plurality of values taken from each of the feature maps at the same coordinate as the sample to be encoded or in a neighborhood thereof. In the case of an image, such a vector can be constructed from a plurality of values taken from each of the feature maps at the same x- and y-coordinates as the sample to be encoded (or decoded).Once these values are taken from the feature maps, they can be processed to form the feature vector, before entering the synthesis neural network, for example by quantization, filtering, interpolation, etc.
[0020] By "synthetic neural network", we mean a neural network such as a convolutional neural network, a multilayer perceptron, an LSTM (for "Long Short Term Memory"), etc. The neural network is defined for example by a plurality of layers of artificial neurons and by a set of activation, weighting and addition functions (for example, a layer can compute y = f (Ax+b), where y and b are vectors of dimension N, x a vector of dimension M, A is a matrix of dimension MxN, and f is the activation function).
[0021] By "neural network parameter" we mean one of the values that characterizes the neural network, for example a weight associated with one of the neurons (filter coefficient, weighting, bias, value affecting the functioning of non-linearity, etc.)
[0022] By "processing by a synthetic neural network" is meant the application of a function expressed by a synthetic neural network to the input characteristic vector to produce an output vector representative of the sample to be encoded (resp. decoded). This output vector may contain one or more data points representative of the sample.
[0023] By "performance measure," we mean a measurement between at least one value of a sample to be encoded and a decoded value of said sample. The measurement may, for example, assess distortion or perceptual error. It may be performed on one or more samples (for example, a current sample, or the current image, etc.). The measurement may also include a measurement of throughput, particularly associated with the encoding of the synthetic neural network and / or the encoding of the feature maps of the first group. The measurement may be a joint measurement of throughput and distortion through their weighting. As is well known in the prior art, the value of this measurement is generally minimized until a target value is reached.
[0024] By "construction step" is meant a step which aims to construct the representative parameters of the image, before their actual encoding. The construction substeps can be repeated as many times as necessary to obtain an acceptable performance measurement.
[0025] According to one embodiment of the decoding process, the dilation rule of a region of the group of at least one feature map includes the dilation of the region by a structuring element.
[0026] According to one embodiment of the decoding process, the dilation by a structuring element is a morphological dilation.
[0027] According to one embodiment of the decoding process, the structuring element comprises at least one maximum distance, the expanded region comprising said region and the points of said region at least one feature map of said group located at a distance from an edge of said region in at least one direction outwards from said region less than and / or equal to said maximum distance.
[0028] According to one embodiment of the decoding process, said dilation rule is defined by a number of parameters equal to twice the dimension of said signal, each direction of one of the dimensions of said signal being associated with a parameter, the dilated region being obtained by adding to said region, in at least one direction of a dimension of said signal, a number of points equal to said parameter associated with said at least one direction of a dimension of said signal.
[0029] According to one embodiment of the decoding process, the characteristic values are representative of the area to be decoded in a transformed domain, for example associated with a direct discrete cosine transform or a wavelet transform, and the synthesis step includes the application to the characteristic values of an inverse transform associated with the direct transform.
[0030] According to one embodiment of the decoding process, the characteristic values are representative of the area to be decoded in a latent domain and the decoding process further includes a step of decoding the parameters of a neural network, called a synthesis neural network and the synthesis step includes, for at least one sample, called the current sample, of the signal to be decoded, associated with a position in the signal to be decoded: - a step of constructing a characteristic vector from said characteristic values as a function of said position of the current sample, and - a processing step of the characteristic vector by the synthesis neural network defined by the decoded parameters to provide a decoded value of the current sample.
[0031] Correspondingly, the invention also relates to a device for encoding a region of a signal, called the encoding region, said encoding region comprising a plurality of samples to be encoded, characterized in that said encoding device is configured to implement: - a step of obtaining a group of at least one characteristic map representative of said signal, - a step of obtaining a mask in said signal of said area to be coded and a dilation rule for a region of the group of at least one feature map, - a step of obtaining a characteristic area of the group from at least one feature map based on said mask, - a step of determining a dilated characteristic zone within said group of at least one feature map by applying said dilation rule to said characteristic zone, - a step of encoding the value of the points of said dilated characteristic zone, and - a step of encoding said mask.
[0032] Correspondingly, the invention also relates to a device for decoding a region of a signal, called the region to be decoded, said region to be decoded comprising a plurality of samples to be decoded, characterized in that said decoding device is configured to implement: - a step of decoding a mask of said area to be decoded, - a step of obtaining a dilation rule for a region of the group of minus a characteristics card, - a step of determining a dilated characteristic area within said group of at least one feature map according to said dilation rule and said mask, - a step of decoding the value of the points in said dilated characteristic zone, and - a synthesis step of said area to be decoded from said decoded values and said mask.
[0033] The characteristics and advantages of the coding or decoding process apply in the same way to the coding or decoding device according to the invention and vice versa.
[0034] The invention also relates to a computer program on a recording medium, this program being capable of being implemented in a computer or an encoding or decoding device according to the invention. This program includes instructions adapted to the implementation of the corresponding method. This program can use any programming language and be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0035] The invention also relates to a computer-readable information or recording medium comprising the computer program instructions mentioned above. The information or recording medium may be any entity or device capable of storing programs. For example, the medium may include a storage means, such as a ROM, for example a CD-ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a floppy disk or a hard drive, a DNA sequence, or flash memory. Furthermore, the information or recording medium may be a transmissible medium such as an electrical or optical signal, which may be transmitted via an electrical or optical cable, by radio link, by wireless optical link, or by other means.
[0036] The program according to the invention can in particular be downloaded onto an Internet-type network.
[0037] Alternatively, each information or recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of a process according to the invention. Brief description of the figures
[0038] The invention will be better understood with the aid of the following description, given solely by way of example and made with reference to the accompanying drawings in which: - [Fig. 1] schematically represents a coding device according to a first embodiment of the invention, - Figure [Fig. 2] schematically represents a decoding device according to the first embodiment of the invention, - Figure [Fig. 3] illustrates an example of a synthetic artificial neural network used in the context of the invention, - [Fig.4] schematically represents an example of determining a characteristic area dilated by the encoding device of [Fig.1] or by the decoding device of [Fig.2], - [Fig.5] is a logic diagram representing an example of a coding process that can be implemented by the coding device of [Fig.1], - [Fig.6] illustrates a coding method used by the coding device of [Fig.1], - [Fig.7] is a flowchart representing an example of a decoding process that can be implemented by the decoding device of [Fig.2], - [Fig.8] illustrates a decoding process used by the decoding device of [Fig.2], - [Fig.9] is a flowchart representing a method for encoding feature cards that can be implemented by the encoding device of [Fig.1] and by the encoding process of [Fig.5], - [Fig. 10] is a logic diagram representing a method for decoding feature cards that can be implemented by the decoding device of [Fig. 2] and by the decoding process of [Fig. 7], - Figure
[11] schematically represents a transcoding device according to a first embodiment of the invention, - Figure 12 schematically represents a coding device according to a second embodiment of the invention, - Figure 13 schematically represents a decoding device according to the second embodiment of the invention, - [Fig. 14] schematically represents an example of determining a characteristic area dilated by the encoding device of [Fig. 12] or by the decoding device of [Fig. 13], - [Fig. 15] is a flowchart representing an example of a coding process that can be implemented by the coding device of [Fig. 12], and - [Fig. 16] is a logic diagram representing an example of a decoding process that can be implemented by the decoding device of [Fig. 13]. Detailed description of the invention
[0039] Figure 1 schematically represents, according to a first embodiment, an ENC coding device for at least one zone of a signal (I(Pn)). In the example described here, the signal I(Pn) is segmented into a set of J zones, and all zones are coded independently. Alternatively, only one zone or only a few zones may be coded independently.
[0040] This ENC coding device includes a GEN module for generating feature maps, an SE module for transformation, an XTR module for data extraction, a TT module for processing and quantification, an NNSYN module corresponding to a synthetic artificial neural network, an NNC module for coding a neural network capable of coding the synthetic neural network, an FMC module for coding feature maps, an EVAL module for evaluating coding performance, and an MAJ module for updating.
[0041] The ENC coding device also includes an FMD decoding module for feature maps, a SEG module for obtaining a segmentation into zones, a CSEG module for coding these zones, an OES module for obtaining a structuring element, an optional CES module for coding this structuring element, a DZCD module for determining dilated characteristic zones in the feature maps and a CE module for coding the value of the points of the dilated characteristic zones.
[0042] The ENC coding device can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be realized by the cooperation of the processor and computer program instructions stored in the aforementioned memory and designed to perform the functionalities of the module concerned, in particular as described below, when these instructions are executed by the processor.
[0043] The ENC coding device of [Fig. 1] receives as input a succession of samples to be coded, denoted Pn, for example a temporal sequence of sound samples, or a set of image data denoted I(Pn). In this second case, the image signal I(Pn) can represent a two-dimensional image, or a plurality of two-dimensional images (video, color components, stereoscopic components, multiscopic components, etc.). Pn designates a sample n of the input signal comprising N samples. In one embodiment, the signal is a color image signal. represented by means of at least one two-dimensional representation, such as a pixel matrix, each pixel having a red (R), green (G), blue (B) component, or, alternatively, a luminance (Y) component and at least one chrominance (U, V) component. The location of each pixel is defined by its x and y coordinates in the image. In one embodiment, the grayscale image is represented by means of a two-dimensional representation, such as a pixel matrix, each pixel having a grayscale component, or luminance component.
[0044] The GEN feature map generation module is configured to generate a plurality of M feature maps denoted FM;.
[0045] In one embodiment, the SE module performs a transformation of the first group of FM characteristic maps; to generate a second group of FMS characteristic maps; at the same resolution as the input signal.
[0046] The optional SE module can perform quantization of the data extracted from this set of M FM maps, or of the vector Zn constructed from this data. Recall that quantizing a value refers to mapping that value to a member of a discrete set of possible code symbols. For example, the set of possible code symbols may consist of integer values, and the quantization system performs a simple rounding of a real value to an integer. In another example, quantization consists of multiplication by a given value followed by rounding. The SE module then performs a transformation of the values of at least one of the feature maps, for example, oversampling, interpolation, filtering, etc. After the transformation, a transformed feature map from the second group has the same resolution as the images in the input sequence.Advantageously, in this method, the feature maps that are coded can be of lower resolution than the images to be coded, while the maps of the second group, which are used to construct the feature vectors, are at the same resolution as the image sequence, which facilitates the extraction of values.
[0047] In one embodiment, the SE module is absent; in this case, the values that will be used to construct the characteristic vector are extracted from the first group of characteristic maps.
[0048] The XTR module performs value extraction from the FM (and / or FMSi) feature maps for a current sample Pn to be encoded, based on its coordinates in the input signal. For example, if one seeks to encode the sample Pn at the coordinates (xn, yn) of an input image, the XTR module performs a extracting values from maps at positions imposed by coordinates (xn, yn).
[0049] In one embodiment, the extracted values constitute the vector Zn. Zn is an L-tuple, that is, it contains L elements, or data z;. For example, in one embodiment, L=M, meaning, for instance, that only one value is extracted for each feature map FM;. In another embodiment, L>M, meaning, for instance, that several values are extracted for at least one feature map FM;. The dimension L of the vector depends on the topology of the NNSYN synthesis neural network and, more particularly, on the topology of the input layer of this NNSYN synthesis neural network.
[0050] The vector Zn of index n refers to the characteristic vector of the pixel P'n.
[0051] In one embodiment, the optional TT module processes the extracted values to generate the vector Zn. The TT module can quantify the data extracted from the set of feature maps. The processing may include other operations, such as filtering, scaling, etc. In particular, if the SE module is not used and if the feature maps in the first group have lower resolutions than the images in the sequence, the TT module can take into account the coordinates of the values in the lower-resolution maps.
[0052] It should be noted that at least one of the SE or TT modules must perform a quantification of the characteristic maps.
[0053] The NNSYN module is a synthetic neural network defined by K parameters Wk, capable of processing the input vector Zn, or L-Tuple, to generate as output a second vector representative of the sample Pn to be coded.
[0054] An example of a synthetic neural network is presented later with reference to [Fig.3].
[0055] The NNC module performs the coding of the synthetic neural network, specifically its parameters Wk. During the coding training or construction process—that is, as long as the performance evaluation step is not satisfactory—the NNC module performs a coding simulation, followed by decoding, for the evaluation module. Subsequently, it performs the actual coding of the synthetic neural network parameters Wk. The coded parameters are denoted Wck. As is known, the coding simulation can be identical to the actual coding, or it can approximate it.
[0056] The FMC module performs the encoding of the FM maps; that is, the values of the characteristic maps of the first group (excluding the maps of the second group, which optionally result from oversampling by the SE module). During the The encoding training, or construction, process—that is, as long as the performance evaluation step is not satisfactory—requires the FMC module to perform a coding simulation, followed by decoding, for the evaluation module. Subsequently, it performs the actual encoding of the FM; chart values. The encoded charts are denoted FMc;. As is known, the coding simulation can be identical to the actual encoding, or it can approximate it. The encoding module quantifies, if necessary, the latent representation of the values in the charts of the first group using a quantifier to generate an ordered collection of quantized values. Then, the encoding module compresses the quantized data, using an encoding that takes into account the proximity of a value to be encoded on the feature chart.
[0057] The EVAL module performs an evaluation and minimization of coding performance. The evaluation function is, for example, of the rate-distortion type. Minimization can be performed by gradient descent, or any other method within the grasp of a person skilled in the art.
[0058] The MAJ module performs an update of the values of the FM cards; to be encoded and / or the parameters of the Wk synthesis neural network, according to the results of the performance function.
[0059] When the evaluation function is minimized, or a predefined number of iterations have been performed, the coded values FMc and Wck constitute the compressed representation of the input signal I(Pn). At this stage, it is not possible to decode a specific area of the coded signal from this compressed representation without decoding the entire coded signal.
[0060] To achieve this functionality, the FMD module performs the decoding of the FMC-encoded values. The cards decoded by the FMD module, numbering M, are denoted FMdi.
[0061] The SEG module performs a segmentation into J (greater than or equal to 2) zones ZoJ of the sequence of samples to be encoded Pn. This segmentation operation allows the samples to be grouped into J different homogeneous sets according to one or more predefined criteria. For example, in the case of a temporal sequence of sound samples, the segmentation operation makes it possible to obtain sequences of units corresponding to silences, noises, phonemes, words, etc. Similarly, in the case of an image signal I(Pn), the segmentation operation makes it possible to group the pixels Pn of the image signal into J homogeneous zones according to criteria, notably intensity or spatial criteria. For example, the segmentation operation can make it possible to identify in the image signal I(Pn) two different zones (J=2), one corresponding to the background of the image and the other to the foreground of the image.Segmentation is, for example, represented as a set of J MZo' masks. allowing each to identify a Zo^ zone, each of the masks being associated with a value different from that associated with the other masks.
[0062] The CSEG module performs effective lossless encoding of the segmentation. This encoding can be achieved by encoding the segmentation map corresponding to the J MZo' masks (for example, using the JPEG-LS algorithm defined by the international standard ISO / IEC 14495-1) or, alternatively, by independently encoding each MZo' mask, or, in yet another alternative, by independently encoding the contours of the J MZo' masks. The encoded segmentation is denoted MZoc'.
[0063] The OES module obtains a dilation rule for a region of the FMd feature map group. This dilation rule is specified, for example, by means of a structuring element ES that depends on the topology of the NNSYN synthesis neural network. It should be noted that this structuring element ES can be determined deterministically based on the topology of the NNSYN synthesis neural network. This structuring element ES can be defined globally for all FMd maps or individually for each FMd map.
[0064] The optional CES module performs the encoding, for example losslessly, of the structure element ES. The encoded structure element is denoted ESc. Such an ESc element can, for example, be an index allowing a structure element to be referenced from a list of structure elements known to the decoder. In other words, the CES module performs the encoding of the dilation rule by encoding the structure element ES. Alternatively, no encoding of the dilation rule is performed.
[0065] From this structuring element ES, for each zone Zo', a dilated characteristic zone ZCD' of the FMd maps is determined. This dilated characteristic zone ZCD* includes at least all the points of the FMd maps necessary to decode the sample values of a given zone Zo*.
[0066] The determination of the dilated characteristic zones ZCDj in the FMd characteristic maps; for the Zo' zones is carried out by the DZCD module.
[0067] The CE module performs the entropic coding for each zone Zo* of the value of the points of the dilated characteristic zones ZCD* in the form of different coded data for each zone Zo'.
[0068] In a first embodiment, the CE module only codes the values of the points in the expanded characteristic zone ZCD* to the exclusion of any other point in the FMd characteristic maps;.
[0069] Alternatively, the CE module creates a secondary feature map pi\42dJ for each dilated feature zone ZCD'. The value of a point on the secondary map The pM2d characteristic value is equal to the value of the point on the FMd map; if this point belongs to the expanded characteristic zone ZCD', it is set to a predetermined value, for example, zero; otherwise, the CE module encodes the entire FM2d characteristic map to obtain different encoded data for each ZoJ zone.
[0070] Thus, the coded values Wck and MZoCj (and optionally ESc) constitute the compressed representation of the Zo' region of the input signal I(Pn). These coded values are represented as binary streams Bl, B2j, and B3j. Thus, according to this example, the binary stream Bl comprises the coded values Wck, the binary stream B2j the coded values gcJ, and the binary stream B3j the coded values MZoCj (and optionally ESc).
[0071] Thus, thanks to the invention, each zone Zo* is coded independently of the other zones, which allows their subsequent decoding to be carried out independently and possibly in parallel.
[0072] It should be noted that the ENC coding device for an area of a signal I(Pn) described above includes the successive application of a coding device for the signal I(Pn) in its entirety in order to produce a stream B0 comprising the coded values FMc; and of a transcoding device for this coded signal in its entirety into a coding of one or more area(s) of the signal I(Pn), the said area(s) being decodable independently of each other.
[0073] According to this embodiment, the transcoding device includes the FMD decoding module for feature maps, the SEG module for obtaining a segmentation into zones, the CSEG module for encoding these zones, the OES module for obtaining a structuring element, the CES module for encoding this structuring element, the DZCD module for determining dilated characteristic zones in the feature maps and the CE module for encoding the value of the points of the dilated characteristic zones.
[0074] Fig. 2 schematically represents a DEC decoding device for a Zodj zone of a signal, called the decoding zone, said decoding zone Zodj comprising a plurality of Pdn samples to be decoded.
[0075] The DEC decoding device of Figure 2 receives as input data streams Bl, B2j and B3j. The Bl stream includes the encoded parameters Wck of the synthesis neural network NNSYN'. The B2j stream includes a group of encoded data gcJ corresponding to the M feature maps pi\.'îdJ representative of the area to be decoded Zodj and the B3j stream includes the encoded parameters of a mask MZoc1 of said area to be decoded Zodj (and optionally the encoded parameters of a structuring element ESc).
[0076] This decoding device DEC includes an NND module for decoding neural network(s) capable of decoding the synthesis neural network NNSYN', a DSEG module capable of decoding a mask MZo' of said area to be decoded Zodj, a DES module capable of decoding a structuring element ES, a DZCD module capable of determining a dilated characteristic area ZCD* in the M feature maps, a DE module for decoding feature maps from the characteristic area ZCD* and the encoded data group, an XTR' module for data extraction, an SE' module for inverse transformation, a TT' module for inverse processing and quantization, an NNSYN' module corresponding to a synthesis neural network and an EXTR module for extracting samples belonging to the Zodj(Pn) area.
[0077] According to one embodiment, the DEC decoder produces at output a decoded area of the image Zodj(Pdn) comprising a plurality of decoded samples Pdn.
[0078] The DES module obtains a structuring element ES. This can be achieved by decoding the encoded parameters ESc if these are available in the B3j stream. Alternatively, the DES module deduces the structuring element ES from the topology of the synthesis neural network NNSYN' after decoding it by the NND module.
[0079] The DESG module performs the decoding of the Mzo' mask from the MZoc' encoded parameters.
[0080] The DZCD module performs the determination of the dilated characteristic zone ZCD* in the characteristic maps -
[0081] The DE module decodes the values of the points in the dilated characteristic zone ZCD' from the coded data pp*ct of the dilated characteristic zone ZCD' in order to obtain the characteristic maps pi\4(]f
[0082] The parameters of the NNSYN' synthesis neural network decoded by the NND module are noted Wdk.
[0083] In one embodiment, the SE' module performs a transformation of the first group of decoded feature maps pMjj to generate a second group of feature maps with the same resolution as the signal to be decoded, denoted pfçfS'j. The SE' module optionally performs inverse quantization corresponding to the quantization performed by the encoder. Inverse quantization is not necessary if the encoder's quantizer Q has simply rounded the real values submitted to it. Inverse quantization is also unnecessary if the neural network is capable of handling quantization of its input data. Otherwise, the decoder performs the inverse operation of the quantizer Q. Then the SE' module performs a transformation of the feature map values, including, for example, oversampling, interpolation, filtering, etc., similar to that performed by the encoder. After the transformation, a transformed feature map from the second group has the same resolution as the images in the sequence to be decoded.
[0084] In one embodiment, the SE' module is absent; in this case, the values that will be used to construct the characteristic vector are extracted from the first group of characteristic maps.
[0085] The XTR' module is identical to the XTR module of Figure 1. It performs an extraction of values from the M characteristic maps (and / or sc'on ' one of the embodiments described above), for a sample Pdn to be decoded as a function of its coordinates in the signal to be decoded.
[0086] In one embodiment, the extracted values constitute the vector Zdn. Zdn is an L-tuple, that is to say, it comprises L elements, or data zd;.
[0087] In one embodiment, the optional TT' module processes the extracted values to generate the Zdn vector. The TT' module can perform inverse quantization of the data extracted from the feature set. The processing may include other operations, such as filtering, scaling, etc., similar to those performed by the encoder.
[0088] The NNSYN' module is a so-called synthetic neural network, defined by K parameters Wdk, capable of processing the input vector Zdn, or L-tuple, to generate as output a second vector representing the sample Pdn to be decoded, generally a vector containing A elements. In one embodiment, A=3 and the output vector is the triplet (R, G, B) of the decoded pixel Pdn. The NNSYN' module has the same structure as the NNSYN module, and its parameters are either identical if the encoding of its parameters Wk is lossless, or different if the encoding is lossy.
[0089] The DEC decoding device can successively be implemented to decode all the Zodj areas in order to reconstruct all the samples to be decoded of the signal, i.e. of the image I(Pdn).
[0090] Alternatively, a plurality of DEC decoding devices can be implemented in parallel, so as to decode the different areas of the image in parallel. Once all the areas are decoded, the image can be reconstructed by compositing the different areas.
[0091] The DEC decoding device can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be implemented through the cooperation of the The processor and computer program instructions stored in the aforementioned memory are designed to perform the functions of the module in question, particularly as described below, when these instructions are executed by the processor. The DEC device may also include multiple processors, with each processor dedicated to the parallel decoding of image areas.
[0092] Figure 3 illustrates an example of a synthetic artificial neural network used for encoding and decoding in the context of embodiments of the invention.
[0093] The synthetic artificial neural network used for encoding, NNSYN, and the synthetic artificial neural network used for decoding, NNSYN', are defined by an identical structure, comprising for example a plurality of layers of artificial neurons, and by a set of weights and activation functions associated respectively with the artificial neurons of the network concerned.
[0094] The synthetic neural network is, according to one embodiment, an MLP, or Multi Layer Perceptron, followed by one or more convolutional neural network(s) Ci, ... Cn, each of the convolutional neural networks being associated with a filtering mask of predefined size, for example of size 3x3.
[0095] The MLP consists of an input layer adapted to the input format (the L-tuple), optionally one or more hidden layers, and an output layer providing an intermediate output vector Vsn, for example a vector containing A' elements. The intermediate output vector is then processed successively by the convolutional neural network(s) to provide an output vector also containing A elements.
[0096] Thus, a vector representation of a current sample (a vector Zn or Zdn from the FM / FMS feature maps, at the encoder, or FMd^FMS'- at the decoder) is applied to the input (i.e., to the input layer) of the NNSYN or NNSYN' synthetic artificial neural network, which produces the output vector. In one embodiment, A is equal to 3, and the output vector is the triplet (R, G, B) of the pixel P'n encoded and then decoded, or of the pixel Pdn decoded by the decoder.
[0097] The concatenation of all these reconstructed pixels in an image constitutes the reconstructed signal Zodj (Pdn), according to an example an image I containing only the Zodj area.
[0098] At the encoder, the NNSYN synthetic artificial neural network is trained on the image so as to minimize the differences between the input representation of the current image I(Pn) and its output representation I(P'n), while also minimizing the amount of data to be encoded. The EVAL module performs a performance measurement in this regard.
[0099] Once the encoder training is complete, the network parameters are encoded, either losslessly, in which case the NNSYN' neural network is identical to NNSYN, either with losses, in which case the NNSYN' network may be slightly different from NNSYN.
[0100] With reference to [Fig. 4], we will now describe an example of determining a dilated characteristic area ZCD1 associated with an area Zo1 in the case of an image signal I(Pn) segmented into three areas Zo1, Zo2, Zo3 identified by their masks MZo1, MZo2 and MZo3. [Fig. 4] thus presents two steps (1)-(2) for determining the dilated characteristic area corresponding to the area Zo1.
[0101] In the example described, the dilated characteristic area is determined in four input maps FMdB, FMd2, FMd3, and FMd4 of the NNSYN / NNSYN' synthesis network. It is further assumed in this example that the NNSYN / NNSYN' synthesis network has only one convolutional neural network Cid, whose convolutional kernel is a 3x3 convolutional kernel.
[0102] This determination includes, in a step (1), obtaining a region R1 comprising all the points of the FMdB FMd2, FMd3 FMd4 maps co-located with samples belonging to the MZo1 mask of the Zo1 zone.
[0103] When the FMdB FMd2, FMd3 FMd4 feature maps have the same resolution as the I(Pn) image, as is the case in the example shown, the boundary of this region is in each FMdB FMd2, FMd3 FMd4 feature map identical to the boundary of the MZo1 mask.
[0104] This region is then expanded, in a step (2), by a structuring element ES which may depend on the topology of the synthesis neural network NNSYN / NNSYN'.
[0105] Thus, in the example presented, decoding a sample Pdn[ typically requires access not only to the values of the points on the FMdB, FMd2, FMd3, and FMd4 maps co-located with this sample Pdn[, but also to the values of the points on the FMdB, FMd2, FMd3, and FMd4 maps co-located with the Pdn2 samples belonging to the neighborhood used during the filtering by the neural network Cb of this point Pdn[. In other words, when the sample Pdn[ is on the edge of the Zo1 area, its decoding requires access to points on the FMdB, FMd2, FMd3, and FMd4 maps co-located with a Pdn2 sample belonging to an area Zo2 different from the Zo1 area. The set of points of the FMdi, FMd2, FMd3 FMd4 maps needed to decode the Zo1 zone defines the dilated characteristic zone ZCD1 of these FMdb FMd2, FMd3 FMd4 maps.
[0106] In the example described here, the structuring element ES can be defined for each map FMdi, FMd2, FMd3, FMd4 as a 3x3 square corresponding to the convolution kernel used within the neural network Ci. In this example, the area ZCD1 is thus obtained in each map FMdB, FMd2, FMd3, FMd4 by a morphological dilation of the trace of the region R1 in each map FMdi, FMd2, FMd3 FMd4 by a square structuring element of size 3x3, i.e. the size of the convolution kernel.
[0107] Alternatively, the structuring element ES can be defined in the case of an I(Pn) image by four values indicating the number of points from the FMdb, FMd2, FMd3, and FMd4 maps to be added to the boundary of the R1 region in both the horizontal and vertical directions. In the example given, all four values are equal to one. It should be noted that, depending on the value chosen, the expanded characteristic area ZCD1 of the FMdh, FMd2, FMd3, and FMd4 maps may, depending on the shape of the Zo1 contour, include more points than the set of points from the FMdb, FMd2, FMd3, and FMd4 maps strictly necessary to decode the Zo1 area.
[0108] In another variant, the structuring element can be defined by a scale factor indicating proportionally to the size of the zone Zo1 the number of points of the maps FMdi, FMd2, FMd3 FMd4 to be added to the border of the region R1 to obtain the dilated characteristic zone ZCD1.
[0109] In the preceding example, the FMdB, FMd2, FMd3, and FMd4 feature maps have the same resolution as the I(Pn) image. However, it should be noted that the determination of the dilated feature area associated with the Zo1 area is similar when the FMdB, FMd2, FMd3, and FMd4 feature maps have different resolutions than the I(Pn) image. Simply put, the determination of the boundary of the R1 region and the structuring element takes into account the transformation function from the resolution of the I(Pn) image to that of the FMdB, FMd2, FMd3, and FMd4 feature maps. For example, if the FMdB, FMd2, FMd3, and FMd4 feature maps have a lower resolution, the boundary of the R1 region is located between the points whose co-located samples in the I(Pn) image are on either side of the Zo1 area boundary.
[0110] The [Fig.5] is a logic diagram representing an example of a method for encoding at least one area of a signal (I(Pn)) which can be implemented by the encoding device of the [Fig.1].
[0111] According to this embodiment, the signal is a two-dimensional image, each sample to be coded is therefore a pixel Pn with coordinates (xn, yn).
[0112] The encoding takes place in several phases.
[0113] In a first phase, called the construction phase, a learning process is performed to determine, for an input signal I(Pn), the values of the FM maps and the Wk parameters to optimize an overall cost function. The learning is, for example, performed by gradient descent, followed by an update of the parameters of the NNSYN synthesis neural network and the values of the feature maps. FM; As is known in the prior art, the cost function can be of the rate-distortion type, or rate-distortion type, or perceptual type. To measure the rate R, it is necessary to simulate the encoding of the FM; maps, and then measure the associated encoding rate (the size of the stream B0). In one embodiment, the encoding of the parameters Wk is not simulated because their influence is less significant than that of the feature maps. In another embodiment, the encoding of the parameters Wk is also simulated, and the associated rate (the size of the stream Bl) is measured. To measure the distortion D, it is necessary to simulate the encoding and then the decoding of at least a portion of the image I, to obtain at least one pixel P'n resulting from an encoding and then decoding simulation, and then measure the difference between this portion of the input image I (Pn) and a corresponding portion of the encoded and then decoded image I (P'n).
[0114] Then, during a second phase, called the coding phase, the parameters Wk, the FMd-' 'cs cards, the MZûj masks, and optionally the ES structuring element are encoded to produce the coded values Wck, ESc, and MZoCj before transmission or storage. They constitute the compressed representation of the Zo' area of the input signal I(Pn).
[0115] Naturally, the first phase and the second phase correspond to two independent processes which can be carried out on different devices.
[0116] We will now describe the steps of a method for encoding a Zo* zone of an input signal I(Pn) according to the first embodiment of the invention.
[0117] During a step E20, a signal I (Pn) to be coded, comprising a plurality of N samples Pn, is provided as input to the process.
[0118] During step E21, the M FM maps of the first group and the Wk parameters of the synthesis neural network are initialized. Subsequently, the Wk parameters of the NNSYN synthesis neural network and the values of the FM maps must be optimized during the construction phase.
[0119] According to one embodiment, the FM cards; are of the same resolution as the input signal I(Pn) and therefore each have the same number of values N as there are samples Pn to be coded.
[0120] According to one embodiment, the FM cards; have a resolution less than or equal to that of the input signal I (Pn) and therefore include, for at least one of them, a number N' of values to be coded less than N. According to a variant, the first FMi card is at the resolution of the images and each subsequent card is at half the resolution of the previous one.
[0121] According to one embodiment, several FMi cards have the same resolution, lower than that of the input signal I (Pn).
[0122] According to one embodiment, the FM maps are transformed to provide a second group of transformed FMSi feature maps. In this embodiment, the feature vectors are preferably extracted from the transformed maps of the second group, and not directly from the maps of the first group. Thus, in this embodiment, the feature vectors are indirectly extracted from the maps of the first group. The maps of the second group are not coded; they serve only for the construction of the feature vectors.
[0123] According to one embodiment, the FM cards are initialized with predefined constant values.
[0124] According to another embodiment, the feature cards are initialized by a set of random real values.
[0125] The FM feature cards; of the first group are subsequently updated, or refined, during an E22 step, by the encoder update module during its learning.
[0126] During step E23, the FM cards of the first group are encoded by the FMC module of the encoder. During the build phase, this operation is a coding simulation. During the encoding phase, this operation is the actual encoding, and the encoded values constitute the B0 stream. The simulation may be identical to the actual encoding, but it may also be different (for example, simplified). For this encoding, a technique for predicting a feature card value based on its neighborhood is used, as will be described, for example, in support of [Fig. 9]. These parameters are subsequently updated, or refined, during the build phase, in later iterations of the process.
[0127] In one embodiment, the FM cards are coded in the order (FM1, FM2, ..., FM4), and the variables of each card in a predefined order, for example, lexicographical. Each card undergoes entropy coding. The entropy coding produces a compressed stream B0 whose throughput is subsequently measured during a step E29.
[0128] During a step E24, according to one embodiment, the M cards of the first group FM; are transformed by the module SE to generate cards of the second group FMSi at the resolution of the images of the input sequence.
[0129] According to one embodiment, M FMS cards are generated.
[0130] According to one embodiment, each FM card; is transformed into an FMS card;.
[0131] According to one embodiment, at least one FM card; is of lower resolution to that of the images in the sequence to be encoded, and the transformation operation includes oversampling so that the transformed FMSi map has the same number of samples as the images in the sequence. Oversampling consists of adding values in FMS maps; to achieve the resolution of the images in the input sequence. It can be simple (by nearest neighbor replication) or involve interpolation (linear, polynomial, filtered, etc.).
[0132] During step E25, values are extracted by the XTR module from the FM cards; or possibly FMS cards; transformed. This extraction is performed according to the coordinates (xn, yn) of the sample Pn of the input signal. It can also be performed according to the resolution of the card considered.
[0133] According to one embodiment, the characteristic Zn vector results directly from this extraction.
[0134] The samples to be coded are, for example, processed sequentially, from n=1 to n=N.
[0135] According to one embodiment, during a step E26, the characteristic vector Zn is constructed by the TT module from the values extracted from the FM or FMS maps for each sample Pn with coordinates (xn, yn) of the input signal. The processing may include quantization of the values extracted from the FM maps or of the resulting vector Zn, if necessary. The processing may include other operations, such as filtering, scaling, the application of any function, preferably monotonic, etc.
[0136] In one embodiment, Zn has as many values as there are FMi or FMS cards; in input. In this case, L=M.
[0137] In one embodiment, Zn is an L-tuple (zb z2,..., Zj), consisting of the values of the FM; or FMS; maps located at the coordinates (xn, yn) of the current pixel Pn, as will be illustrated in support of [Fig.6].
[0138] In one embodiment, Zn is an L-tuple constructed from values taken from the FMi maps at coordinates that may be different for each map. For example, if the FMi maps are at different resolutions because they have been undersampled, the coordinates are adjusted (by scaling) to match the resolution of each map.
[0139] In one embodiment, Zn is an L-tuple constructed from values taken from FM maps; by applying processing to one or more values from the maps, for example, filtering out values close to the target value in a map. For example, for a current sample Pn in an FM map; which is at the same resolution as the input signal, one can extract the values located at the coordinates (xn, yn), (xn-l, yn), (xn, yn-l), and (xn-l, yn-l) and apply processing to these values (filtering, averaging, interpolation, etc.) to obtain the final value (¾) of the i-th element of the vector Zn relative to that FM map;. According to another example, in an FM map; which is at half the resolution of the input signal, we can consider the values located at the coordinates (xn / 2, yn / 2), (xn / 2-l, yn / 2), (xn / 2, yn / 2-l) and (xn / 2-l, yn / 2-l) and apply processing to these values (filtering, averaging, interpolation, etc.) to obtain the final value (¾) of element i of the vector Zn relative to this FM map;.
[0140] During a step E27, the vector Zn is processed by the NNSYN synthesis neural network to generate as output a vector representative of the sample Pn to be coded, according to an embodiment the triplet (R, G, B) of the sample P'n (the sample Pn coded then decoded).
[0141] The structure and parameters Wk of the synthesis neural network are initialized, for example, during the first iteration of this step. These parameters are subsequently updated, or refined, during the construction phase, in subsequent iterations of the process.
[0142] According to one embodiment, the parameters of the synthesis neural network are initialized by predefined values known to give a satisfactory result (for example, following training on a corpus of images).
[0143] According to another embodiment, the parameters of the synthesis neural network are initialized by a set of random values.
[0144] During step E28, the Wk parameters of the NNSYN synthesis neural network are quantized and encoded. During the construction phase, this operation is a coding simulation. During the encoding phase, this operation is the actual encoding, and the encoded values constitute the Bl stream. The simulation may be identical to the actual encoding, but it may also be different (for example, simplified). Any known technique may be used for this purpose, for example, the neural network coding standard proposed by the MPEG-7 Part 17 standard, also called NNR (Neural Network Representation). Note that in this case, the amount of degradation that the encoding introduces to the Wk weights must be chosen.
[0145] During an E29 step, a performance measure is evaluated.
[0146] To this end, the coding simulation rates associated with the feature maps of the first group (simulation of the B0 flow by coding the FM maps;) and optionally with the parameters of the neural network (simulation of the Bl flow by coding the Wk parameters) are measured.
[0147] In one embodiment, the cost function is of the rate-distortion type, denoted (D+L*R), where D, for example, is the squared error measured between the input signal and the decoded signal (or the error measured on a subset of the signal samples). In another example, D is calculated from a perceptual function such as SSIM (for Structural SIMilarity) or MSSSIM (for Multi-scale Structural SIMilarity). In one embodiment, R is the simulated rate of the B0 stream. In another embodiment, R is the total rate used to encode this image, i.e. Let L be the sum of the simulated flow rates of BO and Bl. L is a parameter that regulates the flow rate-distortion trade-off. Other cost functions are possible.
[0148] As long as the cost function has not reached its minimum, or a predefined number of iterations has not been reached, the performance measurement is not satisfactory, and the process is repeated from step E22. This minimization can be performed by a mechanism known as gradient descent with updating of the parameters during step E22 for the values of the feature maps and E23, E27 for the network parameters.
[0149] During step E30, if the cost function has reached its minimum, or if the desired number of iterations is reached, the training stops. If a coded version corresponding to the last simulation of the synthesis neural network parameters Wk and the feature maps FM; is available, the BO and Bl streams can be constructed from it. According to another embodiment, the actual coding of the updated synthesis neural network parameters Wk and the values of the feature maps FM; is performed at this step to produce the encoded parameters Wcket FMc; which constitute the BO and Bl streams.
[0150] At this stage, it is not possible from the BO and Bl streams to decode a specific area of the coded signal without decoding the entire coded signal.
[0151] To achieve this functionality, according to one embodiment, the coding process performs the steps as described below. These steps correspond to a transcoding, or re-encoding, of the B0 stream into a new zone-decodable stream.
[0152] During an E31 step, the FMc coded values are decoded by the FMD module to generate the FMd maps.
[0153] During an E32 step, the segmentation of the signal I(Pn) into J (greater than or equal to 2) Zo' zones is carried out.
[0154] According to a first example of segmentation, the image is divided into J regular zones, for example of identical size and shape (except possibly at the edges of the image). For example, the image can be divided into rectangular zones whose boundaries are horizontal and vertical. Such a division corresponds to the concept of "tiles" implemented by coding standards such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC.
[0155] According to a second example of segmentation, the image is divided into J zones corresponding to samples (or blocks) traversed in a lexicographical order. Such a division corresponds to the concept of "slices" implemented by coding standards such as H.264 / AVC, H.265 / HEVC and H.266 / VVC.
[0156] According to a third example of segmentation, the image is divided into semantic zones, for example a background and a foreground. To do this, a Manual definition can be performed by an operator. Segmentation can also be automatic or semi-automatic, depending on the segmentation algorithm used. It should be noted that there are no restrictions related to the type of segmentation algorithm used.
[0157] Advantageously, such a division into zones allows, during the subsequent decoding of the signal I(Pdn), the decoding of the zones to be decoded Zodj to be parallelized by distributing the decoding load equally on each of the decoders and / or processors.
[0158] At the end of step E32, this segmentation is represented as a set of J masks MZo' each allowing to identify a ZoJ zone.
[0159] For a given zone Zo', the coding process then performs steps E33 to E38.
[0160] During an E33 step, a region Rj comprising all the points of the FMd maps; colocalized with samples belonging to the MZo' mask of the ZoJ zone is obtained.
[0161] In step E34, a dilation rule defined, for example, by a structuring element ES is obtained, and in step E35, from this structuring element ES, a dilated characteristic area ZCD' of the FMd cards is determined. This dilated characteristic area ZCD' comprises at least all the points of the FMd cards necessary to decode a given ZoJ area.
[0162] According to one embodiment, the structuring element ES is defined for each map FMdi, FMd2, FMd3 as a square whose size is a function of the size of the convolutional kernels used by the convolutional neural networks of the NNSYN synthesis neural network. The ZCD* area is obtained, for example, in each map FMd; by a morphological expansion of the trace of the Rj region in each map FMd; by the structuring element ES.
[0163] According to another embodiment, the structuring element ES is defined by a number of numerical values equal to twice the dimension of the signal. Thus, there are four numerical values for a 2-dimensional signal such as an I(Pn) image corresponding to a dilation in each of the two vertical and horizontal directions. The ZCD* area is obtained in each FMd map by adding a number of points to the edge of the Rj region in the two directions associated with each dimension of the signal. In other words, for a 2-dimensional signal such as an I(Pn) image, pixels are added in both horizontal and both vertical directions.
[0164] According to another embodiment, the structuring element ES is defined by a scale factor. The ZCD* zone is obtained in each FMd map by adding a number of points to the edge of the Rj region proportionally to the size of the ZoJ zone.
[0165] During step E36, the CE module performs the coding for each zone Zo' of the values of the points in the dilated characteristic zones ZCD' in the form of data coded The coded values constitute the B2j flow associated with the Zo' zone. Note that this coding is entropic in nature, and preferably lossless.
[0166] During an E37 step, the MZo' mask of the Zo* zone is encoded as an MZocj stream.
[0167] According to one embodiment, the MZocj stream is obtained by losslessly encoding the segmentation map corresponding to the J MZoj masks (for example using the JPEG-LS algorithm defined by the international standard ISO / IEC 14495-1).
[0168] According to another embodiment, the MZoc' stream is obtained by encoding the contours of the MZo' mask.
[0169] During an E38 step, the structural element ES is optionally coded in a form denoted ESc.
[0170] The coded values MZoc' and ESc constitute the B3j flow associated with the Zo' zone.
[0171] Thus, the coded values ESc, £CJ, MZoCj and Wck constitute the compressed representation of the Zo* zone of the input signal I(Pn).
[0172] It should be noted that steps E31 to E38 can be analyzed as steps in a transcoding process of a stream B0 representing a coded image (obtained during a preliminary step EP) into two streams B2j and B3j representing the encoding of a Zo' region of the coded image. The preliminary step EP can correspond to the encoding steps (E20-E30) described above, or to any other encoding process resulting in a coded representation of a set of feature maps (and a synthetic neural network).
[0173] Figure 6 illustrates a method for encoding an I(Pn) signal used in an embodiment of the invention.
[0174] In this embodiment, there are 4 FM cards generated. In a preferred mode, there are 7.
[0175] The first FMi map has the same resolution as the I(Pn) image, and therefore contains WxH variables, where W represents the image width in pixels, and H its height. The second FM2 map has half the resolution (in each dimension) of the FMi map. Each additional map has half the resolution of the previous map. This structure reduces the number of variables in the feature maps, which facilitates coding and learning while minimizing coding costs.
[0176] The FM2 map is oversampled by the SE module by a factor of 2 in each dimension, according to a process illustrated in support of [Fig. 6]. The FM3 map is oversampled by a factor of 4 in each dimension, and the FM4 map by a factor 8 in each dimension. The FMi map is not affected by oversampling. (FMSi=FMi).
[0177] The resulting FMS maps are of the same resolution as the image I(Pn), and therefore each have WxH values, where W represents the width of the image in pixels and H its height (N=WxH).
[0178] Other types of structure are possible, for example one can use a different reduction rate of one half between the cards (one quarter, or one third, etc.).
[0179] In this embodiment, the vector Zn is a 4-tuple (zi...z4) consisting of the values extracted from the FMS maps located at the coordinates (xn, yn) of the current pixel Pn. The vector Zn, consisting of the extracted (quantized) values from the FMS maps, is processed by the NNSYN synthesis neural network to generate a second output vector. In this example, the output vector is the triplet (R, G, B) of the encoded and then decoded pixel P'n. The triplet is inserted into the decoded image I(P'n) at the positions (xn, yn) of the color components (R', G', B').
[0180] In another embodiment, not shown, the Znest vector is extracted directly from the FMi layers, with positions recalculated according to the size of the maps, and then the extracted values are optionally processed and quantified after extraction.
[0181] The [Fig.7] is a logic diagram representing an example of a decoding process for a decoded area (Zodj) of a signal (I(Pdn)) which can be implemented by the DEC decoding device of the [Fig.2] when the synthesis neural network NNSYN' is for example that presented in the [Fig.3].
[0182] During a step F20, the Bl, B2j and B3j streams are extracted from the encoded stream. They contain respectively the parameters Wck, the coded representations p^i of the points of the dilated characteristic zones ZCD' and the coded values MZoc' and optionally ESc.
[0183] During step F21, the structuring element ES and the mask MZo' of the Zodj zone are generated by decoding the coded values ESc and MZoc'. In the case where the ESc parameters are not encoded in the B3J stream, during this step F21, the structuring element is determined from the topology of the synthesis neural network NNSYN'.
[0184] During a step F22, pj^j^i maps are initialized (for example with null values) and a region Rj comprising all the points of the p^ji maps co-located (possibly taking into account a scale factor) with samples belonging to the mask MZo' of the Zo' zone is obtained.
[0185] During a step F23, a dilated characteristic area ZCD' of the py / pp1 maps is determined from the Rj region and the structuring element ES.
[0186] According to embodiments as described for the encoder: - the ZCD* zone is obtained in each map by a morphological dilation of the trace of the region Rj in each map pp^i by the structuring element ES. - The ZCD* area is obtained in each map p]\4d- by adding a number of points to the edge of the Rj region in the two directions associated with each dimension of the signal. The ES element in this case corresponds to a set (vector) of four integer values, for example (1,2,1,2) to indicate a dilation of 1 point (pixel) to the left, 2 to the right, 1 upwards, 2 downwards. This vector can be predefined, or defined for a set / type of images, in which case it does not need to be transmitted. - The ZCD* zone is obtained in each map pp^j^ji by adding a number of points to the border of the Rj region proportional to the size of the ZoJ zone. The ES element corresponds in this case, for example, to a percentage.
[0187] During a step F24, the M maps pp^^j are generated by decoding the pcj values of the points in the dilated characteristic zones ZCD'. In one embodiment, the maps p\qdJ are decoded in order (FMdi, FMd2, ... FMd4), and the values of each map in a predefined order, for example lexicographic.
[0188] According to embodiments as described for the encoder: - The p^ji cards have the same resolution as the coded I(Pn) signal, that is to say they have N=WxH values. - The ppjji cards have a resolution less than or equal to that of the coded I(Pn) signal. - Several pj^ppi cards have the same resolution, lower than the resolution of the coded I(Pn) signal.
[0189] During a step F25, according to one embodiment, the M maps of the first group pp^j^j are transformed by the module SE' to generate maps of the second group pTçJS'j at the resolution of the input images.
[0190] This step is similar to step E24, which was described for the encoder in support of [Fig. 5], and the embodiments apply. In particular: - According to one embodiment, M pM$'j cards are generated. - According to one embodiment, each card is transformed into an FMSF card - In one embodiment, at least one map pMd' has a lower resolution than the images of the image to be encoded, and the transformation operation includes oversampling so that the resulting map p^pgûtran s has the same number of samples as the input image. Oversampling consists of adding values to the maps FMS'- to achieve the resolution of the input image. It can be simple (by nearest neighbor replication) or involve interpolation (linear, polynomial, filtered, etc.).
[0191] The transformation may optionally include inverse quantization of the extracted values, if necessary. However, inverse quantization is not mandatory.
[0192] During step F26, values are extracted by the XTR' module from the FMd- or possibly pMS'j transformed maps. This extraction is performed based on the coordinates (xn, yn) of a sample to be decoded Pdn. It can also be performed based on the resolution of the map considered. This step is similar to step E25, which was described for the encoder in support of [Fig. 5], and the embodiments apply. In particular:
[0193] According to one embodiment, the characteristic vector Zdn results directly from this extraction.
[0194] In one embodiment, Zdn is an L-tuple (zb z2,..., Zj), consisting of the values of the maps or pMg'j located at the coordinates (xn, yn) of a current pixel Pdn, as will be illustrated in support of [Fig.8].
[0195] The samples to be decoded Pdn are for example processed in sequential order according to the mask Mzodj.
[0196] According to one embodiment, in step F27, a vector Zdn is constructed by the module TT' from the extracted values of the pjyjj'j maps of the first group or the pjyjS'i maps of the second group, for each sample Pdn of coordinates (xn, yn) to be decoded, as a function of the coordinates (xn, yn). This step is similar to step E26, which was described for the encoder in support of [Fig. 5], and the described embodiments apply. The extraction may include inverse quantization of the extracted values or of the constructed vector Zdn, if necessary.
[0197] During step F28, the Wdk parameters of the NNSYN' synthesis neural network are generated by decoding the Wck values of the Bl stream. This can be used at This is achieved using any known decoding technique corresponding to the encoding technique used by the encoder. The NNSYN' synthesis neural network is similar to the NNSYN synthesis network; that is, it has the same structure and parameters, except for the encoding, which can be performed with or without loss.
[0198] According to one embodiment, the B2 stream is decoded before the Bl stream, in order to have the NNSYN' synthesis neural network available before starting to decode the samples.
[0199] In step F29, the vector Zdn is processed by the NNSYN' synthesis neural network to generate as output a representative vector of the sample Pdn to be decoded, according to one embodiment a triplet which is injected into the decoded image ZodJ(Pdn) at the positions (xn, yn) of the color components (Rd, Gd, Bd). This step is similar to step E27 which was described for the encoder in support of [Fig. 5].
[0200] Fig. 8 illustrates a method for decoding a Zodj (Pdn) zone of an I(Pdn) signal used in an embodiment of the invention.
[0201] In this embodiment, there are 4 decoded cards. In a preferred mode, there are 7.
[0202] In this embodiment, the first map has the same resolution as image I, and therefore comprises WxH variables, where W represents the width of the image in pixels, and H its height. The second map pjypjj has half the resolution (in each dimension) of the map pM^f. Each additional map has half the resolution of the preceding map. This structure reduces the number of variables in the feature maps, thus facilitating decoding while minimizing the coding cost.
[0203] The FMd^ map is oversampled by a factor of 2 in each dimension, using any oversampling method available to a person skilled in the art. The pty^j map is oversampled by a factor of 4 in each dimension, and the pMd^^ map by a factor of 8 in each dimension.
[0204] The ppyJS'j maps are of the same resolution as the image to be decoded, and therefore have WxH values, where W represents the width of the image in pixels, and H its height.
[0205] In this embodiment, the vector Zdn is a 4-tuple (zi...z4) consisting of the values of the maps located at the coordinates (xn, yn) of the current pixel Pdn. The vector Zdn is optionally dequantized and then processed by the NNSYN' synthesis neural network to generate as output the triplet (R, G, B) representative of the sample Pdn to be decoded. The triplet (R, G, B) is inserted into the decoded image I (Pdn) at coordinates (xn, yn) in the color components (Rd, Gd, Bd).
[0206] The [Fig.9] is a logic diagram representing an entropic coding method for feature maps which can be implemented by the coding device of the [Fig.1] and by the coding process of the [Fig.5].
[0207] These steps constitute sub-steps of step E36 described previously in support of Figure 5. They aim to encode a current value Vd' of a point of a dilated characteristic zone ZCD' of a feature map of the first group being processed using values from the neighborhood.
[0208] In a substep E361, a neighborhood vector (Cn) is established, comprising values close to the value Vn. These neighboring values may be located in the same map and / or in a different map from the plurality M of maps. This neighborhood vector consists of a number C of values, or data, corresponding to neighborhood values (for example, C=10). These values must be known to the encoder and the decoder; therefore, they must be located in a causal neighborhood of the value Vn.
[0209] According to a first embodiment, these values are used to determine the context of an entropy encoder to encode the current value during an E363 step. This encoder can be a CAB AC (Context-adaptive binary arithmetic coding) type encoder. This type of encoder is well known to those skilled in the art. It is notably used in the H.265 / HEVC video compression standard. It is an arithmetic encoder with lossless compression. It decomposes all non-binary symbols into binary symbols. Then, for each bit, the encoder selects the most suitable probability model and uses a context to optimize the probability estimation. This context can be defined by information from neighboring elements. Adaptive or non-adaptive arithmetic coding is then applied to compress the resulting data. As is known to those skilled in the art, there are several ways to use the neighborhood vector to produce context information.For example, one can count the number of neighboring values other than zero and associate a context with each number. Alternatively, one can perform comparisons between several neighboring values and associate a given context with an ordering configuration between the neighboring values, for example by sorting the neighboring values in ascending order and associating a context with each possible order.
[0210] In a second embodiment, the neighborhood is used to predict, during a step E362, the current value from an autoregressive model. It is recalled that an autoregressive model predicts a sample from a series based on its past values. In this embodiment, the past values are constituted by the context, and the difference between the predicted variable and the actual value is quantified and then entropically coded during step E363.
[0211] At the end of the process, the current coded value Vcnde of the card being processed is coded.
[0212] The [Fig. 10] is a logic diagram representing a method for decoding feature cards which can be implemented by the decoding device of [Fig. 2] and by the decoding process of [Fig. 7]
[0213] These steps constitute sub-steps of step F24 described previously in support of Figure 7. They aim to decode a current value of a point of a dilated characteristic zone ZCD' of a feature map of the first group being processed using values from the neighborhood.
[0214] In a substep F241, a neighborhood vector (Cdn) is established, comprising values close to the value Vdn. This step is similar to the previously described step E361, and the same embodiments apply. This neighborhood vector consists of a number C of values, or data, corresponding to neighborhood values (for example, C=10) located in the same map and / or in a different map from the plurality M of maps of a point in an expanded characteristic zone ZCD* of a feature map p^ji. These values, being in a causal neighborhood of the value Vdn, are known to the decoder.
[0215] According to a first embodiment, these values are used to determine the context of an entropy decoder for decoding the current value during a step F243. This decoding is similar to that used in the encoder, for example, CAB AC. The use of the neighborhood to produce context information is similar to that chosen in the encoder. For example, one can count the number of non-zero neighboring values and associate a context with each number. Alternatively, one can perform comparisons between several neighboring values and associate a given context with an ordering configuration among the neighboring values, for example, by ranking the neighboring values in ascending order and associating a context with each possible order.
[0216] In a second embodiment, the neighborhood is used to predict the current value from an autoregressive model during a step F242. In this embodiment, the past values are constituted by the context, and the difference between the predicted variable and the actual value is decoded and then entropically dequantized during step F243.
[0217] At the end of the process, the current value Vdnde of the FMd card being processed is decoded.
[0218] Figure 11 schematically represents a TRANSC transcoding device for a coded representation of a signal in the form of independent streams B2j and B3j. into a single BO stream. The advantage of such transcoding is to obtain the most compact possible representation of the stream, in cases where the image is not to be decoded by zones, but in its entirety. Indeed, the zone-coded representation, as described previously, contains redundant information, particularly coded samples that belong to several zones because they are located within zone boundaries.
[0219] The TRANSC transcoding device of [Fig.1 1] receives as input the J data streams B2j and B3j corresponding to the J zones Zo' segmenting the signal I(Pn).
[0220] This TRANSC decoding device includes a DSEG module capable of decoding an MZo' mask of a zone, a DES module capable of decoding a structuring element ES, a DZCD module capable of determining an expanded characteristic zone ZCD' in the M characteristic maps, a DE module for decoding the characteristic maps from the characteristic zone ZCD* and the encoded data group, a COMP module for composing the characteristic maps and an FMC module for encoding the FMd characteristic maps.
[0221] With the exception of the COMP module, the other modules of the TRANSC transcoding device correspond to the modules of the decoding device of [Fig.2] referenced in a similar way and are not described again here.
[0222] The COMP module receives as input the J sets of feature maps associated with the Zo* zones of the signal I, as well as the associated Rj regions (these regions are, for example, provided by the DZCD module). From this data, the COMP module will construct a set of FMdj feature maps by assigning, for each point in a region of an FMdj feature map co-located with an Rj region, the value of the corresponding point in the map
[0223] The cards are then coded by the FMC module to produce the FMci data constituting the B0 flow.
[0224] We will now describe a second embodiment based not on a neural approach but on a classical approach using a discrete cosine transform.
[0225] Fig. 12 schematically represents, according to this second embodiment, an ENC coding device.
[0226] This ENC coding device includes a TRANS module for generating a characteristic map associated with a discrete cosine transform, a Q module for quantizing the characteristic map, and an EC module for encoding the characteristic map after its quantization.
[0227] The ENC coding device also includes a DC decoding module for the encoded feature map, a SEG module for obtaining a segmentation into zones, a CSEG module for encoding these zones, an OES module for obtaining a structuring element, an optional CES module for encoding this structuring element, a DZCD module for determining dilated feature zones in the feature maps and a CE module for encoding the value of the points of the dilated feature zones.
[0228] The ENC coding device can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be realized by the cooperation of the processor and computer program instructions stored in the aforementioned memory and designed to perform the functionalities of the module concerned, in particular as described below, when these instructions are executed by the processor.
[0229] The ENC coding device of [Fig. 12] receives as input a succession of samples to be coded, denoted Pn, for example a set of data from a grey level image denoted I(Pn).
[0230] The TRANS module for generating a feature map associated with a discrete cosine transform is configured to generate a feature map denoted FMi. To do this, the TRANS module partitions, for example, the image I(Pn) into 8x8 blocks and transforms each block by applying a discrete cosine transform, known as a two-dimensional discrete cosine transform, by applying a one-dimensional discrete cosine transform first to the rows and then to the columns of the block, or vice versa.
[0231] The Q module performs a quantification of the data from the FMi cards, for example by using a uniform quantifier.
[0232] The EC module performs an entropic coding of the quantized values of each of the transformed blocks of the FMiselon card in a predetermined traversal order to obtain the coded values FMci which constitute the compressed representation of the input signal I(Pn).
[0233] To obtain a compressed representation of the input signal I(Pn) in the form of a set of J independently coded zones, the DC module performs the decoding of the coded values FMci. The map decoded by the DC module is denoted FMdi.
[0234] The SEG module performs a segmentation into J (greater than or equal to 2) ZoJ zones of the image to be coded I(Pn)) while the CSEG module performs the actual coding of the segmentation.
[0235] The OES module obtains a structural element ES. This structural element ES is determined in this embodiment by the definition of the discrete cosine transform and by a filtering operation (presented later with reference to [Fig. 13]) carried out during the decoding of a Zo' zone.
[0236] The CES module optionally performs lossless encoding of the ES structuring element. In this case, the encoded structuring element is denoted ESc.
[0237] From this structuring element ES, for each zone Zo*, a dilated characteristic zone ZCD' of the FMdi map is determined. This dilated characteristic zone ZCD* includes at least all the points of the FMdi map necessary to decode a given zone Zo*.
[0238] The determination of the dilated characteristic zones ZCDj in the FMdi characteristic map for the Zo' zones is carried out by the DZCD module.
[0239] The CE module performs the coding for each zone Zo* of the value of the points of the dilated characteristic zone ZCD* in the form of different coded data for each zone Zo'.
[0240] In a first embodiment, the CE module encodes only the values of the points in the expanded characteristic zone ZCD*, excluding any other points in the FMdi characteristic map. In this case, the effective encoding of the segmentation is performed without loss by the CSEG module.
[0241] Alternatively, the CE module creates a feature map FMI>]For each dilated characteristic zone ZCD'. The value of a point on the feature map FMDj is equal to the value of the point on the FMDi map if that point belongs to the dilated characteristic zone ZCD* and to a predetermined value, for example zero, otherwise. The CE module then encodes the entire feature map FMQJ To obtain the different encoded data g^J for each zone Zo*.
[0242] Thus, whatever the example of embodiment chosen, the coded values ESc, and MZoCj constitute the compressed representation of the Zo* area of the image signal I(Pn).
[0243] Thus, each zone Zo* is coded independently of the other zones, which allows their subsequent decoding to be carried out independently and possibly in parallel.
[0244] Fig. 13 schematically represents a DEC decoding device for a Zodj area of a greyscale image I(Pn), called the decoding area, said decoding area Zodj comprising a plurality of Pdn samples to be decoded.
[0245] This DEC decoding device comprises a DESG module capable of decoding an MZo' mask of said zone to be decoded Zodj, a DES module capable of decoding a structuring element ES, a DZCD module capable of determining a dilated characteristic zone ZCD' in the feature map ' a DE module for decoding the feature map from the feature area ZCD' and the encoded data group, a Q1 module for inverse quantization of the values of the feature map jJ, a TRANS1 module for generating an image by applying an inverse discrete cosine transform to the feature map obtained at the output of the Q* module, a FILT low-pass filtering module, and an EXT module for extracting samples belonging to the Zodj(Pn) area of the I(Pdn) image.
[0246] The decoding device DEC receives as input a group of encoded data corresponding to the characteristic map pM^ representing the zone to be decoded Zodj, the encoded parameters MZoc' of a mask of said zone to be decoded Zodj and the encoded parameters ESc of a structuring element.
[0247] The DES module obtains a structuring element ES, for example by decoding the encoded parameters ESc if these have been encoded by the ENC encoder of the [Fig.12]
[0248] The DESG module performs the decoding of the Mzodj mask from the MZoc' encoded parameters.
[0249] The DZCD module performs the determination of the dilated characteristic zone ZCD* in the FMdi characteristic map.
[0250] The DE module performs the decoding of the point values of the dilated characteristic zone ZCD* from the coded data pcJ and the dilated characteristic zone ZCD* in order to obtain the characteristic map
[0251] The Q 1 module performs an inverse quantization corresponding to the quantization performed at the encoder and the TRANS 1 module generates an intermediate image I;(Pn) by applying to each block of the FMd feature map] a two-dimensional inverse discrete cosine transform.
[0252] The FILT module applies a low-pass filter along the blocks of the decoded area Zodj (Pdn) of the intermediate image I;(Pn) in order to reduce the discontinuities induced along these blocks by the quantization of the coefficients from the discrete cosine transform of these blocks.
[0253] In the example shown, the filter is of length 5 and the associated weights are represented as a vector (1 / 16, 3 / 16, 8 / 26, 3 / 16, 1 / 16) associated with a filter mask of dimension 5. Obviously, filters with different filter masks, for example two-dimensional or with different lengths, can be applied.
[0254] At the end of this filtering phase, the EXT module retains the samples belonging to the Mzodj mask while the others are deleted, the resulting signal then corresponding to the Zodj(Pn) area of the I(Pdn) image.
[0255] The DEC decoding device can be implemented by means of an electronic device comprising a processor and a memory, not shown; each of the modules mentioned above can then be realized by the cooperation of the processor and computer program instructions stored in the aforementioned memory and designed to perform the functionalities of the module concerned, in particular as described below, when these instructions are executed by the processor.
[0256] With reference to [Fig. 14], we will now describe an example of determining the dilated characteristic area ZCD1 associated with a zone Zo1 in the case of an image signal I(Pn) segmented into three zones Zo1, Zo2, Zo3 identified by their mask MZo1, MZo2 and MZo3 according to the second embodiment described with reference to Figures 12 and 13. [Fig. 14] thus presents two steps (1)-(2) of determining the dilated characteristic area corresponding to the zone Zo1.
[0257] This determination includes, in a step (1), obtaining a region R1 comprising all the points of the FMdi map co-located with samples belonging to the mask MZo1 of the Zo1 zone. In the example presented, the boundary of this region R1 is in the FMdi feature map identical to the boundary of the mask MZo1 in the image I(Pn).
[0258] This region Rj is then dilated, in a step (2), by applying a dilation rule to this region R.
[0259] This dilation rule includes the selection of all transformed blocks of the FMdi feature map comprising a point co-located with a sample belonging to the mask MZo1 of the Zo1 zone, the definition of a structuring element ES defining four values indicating the number of FMdi map blocks to be added to the edge of the region defined by the selected blocks along the two horizontal and vertical directions of the image I(Pn)) and the addition of the indicated blocks.
[0260] Alternatively, the dilation rule includes obtaining a structuring element ES defining four values indicating the number of points from the FMdi map to be added to the edge of the Rj region along the two horizontal and vertical directions to obtain a dilated characteristic zone ZCD' from the Rj region and selecting all transformed blocks from the FMdi characteristic map including a point belonging to the dilated characteristic zone ZCD'.
[0261] Regardless of the variant chosen, the structuring element is defined from the characteristics of the low-pass filter that will be applied by the FILT module of the DEC decoder during the decoding of the ZoJ area. For this reason, the structuring element may not be encoded by the ENC encoder and may be obtained independently by the DEC decoder.
[0262] The [Fig. 15] is a flowchart representing an example of a coding process which can be implemented by the ENC coding device of the [Fig. 12].
[0263] During a G20 step, a succession of samples to be coded, denoted Pn, for example a set of data of a grey level image denoted I(Pn) is provided as input to the process.
[0264] During a G21 step, an FMdias feature map associated with a discrete cosine transform is created by the TRANS module.
[0265] During a G22 step, the data from the FMdi feature chart are quantized by the Q module, for example using a uniform quantizer.
[0266] During a G23 step, the EC module performs an entropic coding of the quantized values of each of the transformed blocks of the FMiselon card in a predetermined traversal order to obtain the coded FMci values which constitute a compressed representation of the input signal I(Pn).
[0267] During a G24 step, the signal I (Pn) is segmented into J (greater than or equal to 2) Zo' zones. At the end of the G34 step, this segmentation is represented as a set of J masks MZo', each allowing the identification of a Zo' zone.
[0268] For a given Zo* zone, the coding process then performs steps G25 to G30.
[0269] During a G25 stage, an Rj region comprising all points of the FMdi map colocalized with samples belonging to the MZo' mask of the Zo* zone is obtained.
[0270] During a G26 step, a structuring element ES is obtained.
[0271] During a G27 step, for each zone Zo*, a dilated characteristic zone ZCD' The FMdi map is determined from the MZoj mask and the ES structuring element.
[0272] During a G28 step, the CE module performs the coding for each zone Zo* of the value of the points of the dilated characteristic zone ZCD* in the form of Eci coded data different for each zone Zo'.
[0273] During a G29 step, the MZoj mask of the Zo* zone is encoded as an MZoc' stream.
[0274] During a G30 step, the structural element ES is coded in a form denoted ESc.
[0275]
[0276]
[0277]
[0278]
[0279]
[0280]
[0281]
[0282]
[0283]
[0284]
[0285] [Fig. 16] is a flowchart representing an example of a decoding process that can be implemented by the DEC decoding device of [Fig. 13]. During an H20 step, the BF and B2j streams are extracted from the encoded stream. They contain the coded representations i and the coded values MZoc' and ESc. During an H21 step, the structuring element ES and the mask MZo' of the Zodj zone are generated by decoding the coded values ESc and MZoc'. During step H22, the DE module decodes the values of the points in the dilated characteristic zone ZCD' from the coded data and possibly from the dilated characteristic zone ZCD' in order to obtain the characteristic map Based on implementation examples as described for the coder: - The coded data represents a complete PMcpj feature map which is then decoded by the DE module. - The ppi-coded data represents only the values of the points in the dilated characteristic zone ZCD', excluding any other points in the FM^p characteristic map. In this case, the map is decoded using the p^i-coded data and the dilated characteristic zone ZCD'. During an H23 step, an inverse quantification of the values of the points of the FMdj map] is performed by the Q* module. During an H24 step, the TRANS 1 module generates an intermediate image I;(Pn) by applying a two-dimensional inverse discrete cosine transform to each block of the feature map. During an H25 step, the FILT module applies a low-pass filter along the intermediate image I;(Pn). During an H26 step, the EXT module extracts only the samples belonging to the Mzodj mask in order to obtain the decoded Zodj(Pdn) signal corresponding to the Zo' area of the I(Pdn) image. It should also be noted that the invention is not limited to the embodiments described above. Indeed, it will be apparent to those skilled in the art that various modifications can be made to the embodiments described above, in light of the information just disclosed to them. For example, the encoding and decoding methods described above may use a wavelet transform instead of a discrete cosine transform or of a neural coding method.
[0286] Furthermore, the invention can be implemented with NNSYN / NNSYN' synthetic neural networks different from those previously described. For example, the NNSYN / NNSYN' synthetic neural networks can be recurrent neural networks. In another example, the synthetic neural networks can consist of one or more convolutional neural networks, followed by an MLP, and then followed by one or more convolutional neural networks. In these examples, obtaining the Zn / Zdn vectors and / or the intermediate output vectors Vsn is adapted to the topology of the NNSYN / NNSYN' synthetic neural networks.
[0287] In the detailed presentation of the invention given above, the terms used shall not be interpreted as limiting the invention to the embodiments set forth in this description, but shall be interpreted as including all equivalents which can be foreseen by a person skilled in the art by applying their general knowledge to the implementation of the teaching which has just been disclosed to them.
Claims
1.
2. Demands A method for encoding a region (Zo^) of a signal (I(Pn)), called the region to be encoded, said region to be encoded comprising a plurality of samples (Pn) to be encoded, said encoding method comprising the following steps: - a step (E31) of obtaining a group of at least one characteristic map (FM;) representative of said signal (I(Pn)), - a step (E32, E34) of obtaining a mask (MZo'j in said signal (I(Pn)) of said zone to be coded (Zo*) and a dilation rule of a region of the group of at least one feature map, - a step (E33) of obtaining a characteristic zone (Rj) of the group of at least one characteristic map (FM;) according to said mask (MZo'j, - a step (E35) of determining a dilated characteristic zone (ZCD1) in said group of at least one feature map (FM;) by applying said dilation rule to said characteristic zone (Rj), - a step (E36) of encoding the value of the points of said dilated characteristic zone (ZCD'), and - a step (E37) of encoding said mask. A coding method according to claim 1, wherein the step of obtaining a group of at least one feature card (FM) comprises: - for at least one sample, called the current sample (Pn), of the signal to be encoded, associated with a position (xn, yn) in the signal to be encoded, a construction step (E25) of a characteristic vector (Zn) from said group of at least one feature map (FM), as a function of said position (xn, yn) of said current sample (Pn), and a processing step of said characteristic vector (Zn) by an artificial neural network, called the synthesis neural network (NNSYN) defined by a set of parameters (Wk), to provide a vector (P'n) representative of a decoded value of the current sample, and - an update step (E22, E27) of at least one value of the group of said at least one feature card and / or of at least one parameter of said network, based on a coding performance measure.
3. A method for decoding an area (Zodj) of a signal (I(Pdn)), referred to as the area to be decoded, said area to be decoded comprising a plurality of samples (Pdn) to be decoded, said decoding method comprising the following steps: - a step (F21) of decoding a mask (MZo1) of said area to be decoded (Zodj), - a step of obtaining (F21) a dilation rule for a region of the group of at least one feature map <FMDp’ - une étape (F23) de détermination d’une zone caractéristique dilatée (ZCD1) dans ledit groupe d’au moins une carte de caractéristiques (pMpjî) en fonction de ladite règle de dilatation et dudit masque (MZoj), - une étape (F24) de décodage de la valeur des points de ladite zone caractéristique dilatée (ZCD’), et - une étape de synthèse (F29) de ladite zone à décoder (Zodj) à partir desdites valeurs décodées et dudit masque (MZ©*).
4. Decoding method according to the preceding claim wherein said dilation rule of a region of the group of at least one feature map (pMD^ includes the dilation of said region by a structuring element (ES).
5. Decoding method according to the preceding claim wherein said dilation by a structuring element (SE) is a morphological dilation.
6. A decoding method according to claim 4 wherein the structuring element (SE) comprises at least one maximum distance, the expanded region comprising said region and points of said region at least one feature map of said group located at a distance from an edge of said region in at least one direction outwards from said region less than and / or equal to said maximum distance.
7. A decoding method according to claim 4 in which said dilation rule is defined by a number of parameters equal to twice the dimension of said signal (I(Pdn)), each direction of one of the dimensions of said signal being associated with a parameter, the dilated region being obtained by adding to said region, in at least one direction of a dimension of said signal, a number of points equal to said parameter associated with said at least one direction of a dimension of said signal.
8. A decoding method according to any one of claims 3 to 7 wherein said characteristic values are representative of said area to be decoded in a transformed domain, for example associated with a direct discrete cosine transform or a wavelet transform and wherein said synthesis step includes applying said characteristic values of an inverse transform associated with said direct transform.
9. A method according to any one of claims 3 to 7 wherein said characteristic values are representative of said area to be decoded in a latent domain, wherein said decoding method further comprises a step of decoding the parameters (Wdk) of a neural network (NNSYN'), said synthesis neural network, and wherein said synthesis step comprises, for at least one sample, said current sample (Pdn), of the signal to be decoded, associated with a position (xn, yn) in the signal to be decoded: - a step of constructing a characteristic vector (Zdn) from said characteristic values as a function of said position (xn, yn) of said current sample, and - a step of processing said characteristic vector (Zdn) by the synthesis neural network (NNSYN') defined by the decoded parameters (Wdk) to provide a decoded value of the current sample (Pdn).
10. A coding device for an area (Zo1) of a signal (I(Pn)), called the area to be coded, said area to be coded comprising a plurality of samples (Pn) to be coded, characterized in that said coding device is configured to implement: - a step of obtaining a group of at least one characteristic map (FM;) representative of said signal (I(Pn)), - a step of obtaining a mask (MZo1) in said signal (I(Pn)) of said area to be coded (Zo^ and a rule for dilating a region of the group of at least one feature map, - a step of obtaining a characteristic zone (Rj) of the group of at least one characteristic map (FM;) according to said mask (MZoJ), - a step of determining a dilated characteristic zone (ZCD1) in said group of at least one feature map (FM;) by applying said dilation rule to said characteristic zone (Rj), - a step of encoding the value of the points of said dilated characteristic zone (DCD*), and - a step of encoding said mask.
11. A decoding device for an area (Zodj) of a signal (I(Pdn)), referred to as the area to be decoded, said area to be decoded comprising a plurality of samples (Pdn) to be decoded, characterized in that said decoding device is configured to implement: - a step (F21) of decoding a mask (MZo1) of said zone to be decoded (Zodj), - a step of obtaining (F21) a dilation rule for a region of the group of at least one feature map <FMDp’ - a step (F23) of determining a dilated characteristic zone (ZCD*) in said group of at least one feature map (pMDÎ) according to said dilation rule and said mask (MZ©*), - a step (F24) of decoding the value of the points of said dilated characteristic zone (DCD*), and - a synthesis step (F29) of said zone to be decoded (Zodj) from said decoded values and said mask (MZ©*).
12. A computer program comprising instructions for performing the steps of an encoding process according to any one of claims 1 to 2 or a decoding process according to one of any of claims 3 to 9 when said program is executed by a computer.
Citation Information
Patent Citations
Method and device for encoding and decoding images.
FR3143245A1
Residual coding method and device, video coding method and device, and storage medium
US20240064309A1
Pre-analysis based image compression methods
US20240121445A1