Method and apparatus for decoding bitstream

By obtaining suitable neural networks from a set of compressed neural networks at the decoding device for decoding, the problem of large data transmission in the prior art is solved, and efficient media content decoding is achieved, especially suitable for medical images and video sequences.

CN120380482APending Publication Date: 2025-07-25ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380087224.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-12-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When using neural networks to encode and decode media data, it is necessary to transmit a large amount of data between the encoding device and the decoding device to configure the decoding neural network, resulting in excessive data volume and inefficiency, especially in video compression.

Method used

By obtaining suitable neural networks from a set of multiple compressed neural networks at the decoding device, decoding bitstreams are used and the neural networks are grouped to speed up the adaptation and generation process, while designing an efficient encoding and decoding network suitable for a variety of contents.

Benefits of technology

Decoding media content with good performance in compression and quality is achieved, reducing the amount of data required by decoding devices, and improving decoding efficiency, especially suitable for efficient decoding of medical images and video sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380482A_ABST
    Figure CN120380482A_ABST
Patent Text Reader

Abstract

The invention relates to a method for decoding a bitstream, the method comprising, at a decoding device: obtaining (S33) a decoded neural network from at least one group of a plurality of compressed neural networks, the compressed neural networks in the at least one group being available for the decoding device; and decoding the bit stream using the obtained neural network (S34).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of encoding and decoding media data using neural networks. More specifically, the present invention relates to a method for decoding a bitstream. The present invention also relates to a decoding device and a bitstream.

[0002] Generally, the present invention can be applied in any industrial and technical field where encoding and / or decoding of media data must be performed. Background Art

[0003] In recent years, the growing demand for media streaming has led to significant progress in compression, particularly video compression. The aim of compression algorithms is to reduce the amount of data required for media streaming or storage while maintaining acceptable visual / audio quality.

[0004] Recently, neural network-based encoders and decoders for encoding and decoding media data have been introduced to improve performance. To this end, data representing the content of media data is encoded using an artificial neural network (also referred to as an encoding neural network), and the decoding of the encoded data can be implemented by another neural network (also referred to as a decoding neural network).

[0005] However, current methods for encoding and decoding media data using neural networks typically require a large amount of data to be pre-transferred between the encoding device and the decoding device to allow the decoding device to correctly configure the neural network to be used for decoding the media data.

[0006] In video compression, some methods even require different decoding neural networks to be transmitted to decode each video sequence of the bitstream, which also represents a large amount of data.

[0007] Therefore, there is room for improvement in the field of encoding and decoding using neural networks. Summary of the Invention

[0008] To this end, the present invention first provides a method for decoding a bitstream, the method comprising, at a decoding device:

[0009] - obtaining a decoded neural network from at least one set of a plurality of compressed neural networks, the compressed neural networks in the at least one set being available for the decoding device; and

[0010] - decoding the bitstream using the obtained neural network.

[0011] As will be further detailed below, the compressed neural networks in the set are "available for" the decoding device because the compressed neural networks are stored in the memory of the decoding device or are available for the decoding device by download.

[0012] The method provides the following advantages: it provides a decoding device that is suitable for generating encoded media content with good performance in terms of compression or quality, and provides this advantage within a wide range of types of media content, while limiting the amount of data available for the decoding device.

[0013] Grouping neural networks into "groups" speeds up the search for an adapted neural network, and also speeds up the generation of an adapted decoding neural network and the decoding of the bitstream. In addition, the neural network groups can be designed such that efficient encoding and decoding networks are available for various content, thus ensuring that excellent compression performance can be achieved in all cases.

[0014] In some embodiments, the compression neural network in the at least one group is suitable for decoding a media data bitstream of content of the same type.

[0015] The media data corresponds to, for example, audio data, video data, and / or still images.

[0016] The present invention is particularly suitable for decoding medical images (such as magnetic resonance images) or medical video sequences (such as cardiac image sequences, for example for detecting irregular pulsations of a patient). In fact, medical images have specific image statistics that can be efficiently processed by a neural network encoder / decoder (for example, by only adjusting weight values).

[0017] In addition, medical images are typically 3D or even 4D (3D + time) content. Using a neural network encoder / decoder to process such images is particularly advantageous because its architecture is relatively similar to that of a neural network encoder / decoder suitable for processing images with lower dimensions (for example, 2D images).

[0018] In some embodiments, obtaining the decoded neural network includes decoding the compression neural network in the at least one group, and the compression neural network can be decoded with reference to a reference neural network.

[0019] In some embodiments, the compression neural network can be decoded with reference to a reference neural network using decoded differential data.

[0020] The feature that the compression neural network can be decoded with reference to a reference neural network provides the advantage of limiting the amount of data to be processed because only the differential data with respect to the reference neural network is encoded and / or decoded.

[0021] In some embodiments, the reference neural network belongs to the at least one group.

[0022] In some embodiments, decoding the compression neural network in the at least one group includes:

[0023] - Apply a neural network decoder to the reference neural network to obtain a decoded reference neural network; and

[0024] - Apply an incremental neural network decoder that takes as input the decoded reference neural network and decoded differential data associated with the compressed neural network to obtain the decoded neural network.

[0025] In some embodiments, the reference neural network does not belong to the at least one group.

[0026] In this case, the reference decoded neural network can be considered the default reference decoder.

[0027] In some embodiments, obtaining the decoded neural network includes independently decoding the compressed network by applying a neural network decoder that complies with neural network compression and representation standards to the compressed neural network.

[0028] The neural network compression and representation standards can correspond, for example, to Part 17 of ISO / IEC 15938.

[0029] Hereinafter, the terms "independently encode" or "independently decode" refer to the fact that a neural network is encoded or decoded without accessing another neural network (or more generally without accessing data related to another network).

[0030] In some embodiments, the bitstream includes data for signaling whether the bitstream should be decoded by a decoding device adapted to obtain a decoded neural network from at least one group of compressed neural networks.

[0031] In some embodiments, the bitstream includes data for signaling predetermined characteristics of at least one group of compressed neural networks adapted to decode the bitstream.

[0032] In some embodiments, the data for signaling the predetermined characteristics of a group of compressed neural networks indicates:

[0033] - A group including decoded neural networks that are adapted to a specific type of content, such as sports, synthetic images, medical images,...

[0034] - A group characterized by the maximum incremental decoding level of its neural decoder. For example, the data BT indicates that up to two decoded neural networks need to be decompressed before a given requested decoded neural network identified by the flag I (described below) can be decompressed;

[0035] - A group represented by the fact that a certain decoding method is required to decompress these compressed decoding neural networks (e.g., NNR decoding or ONNX decoding according to these corresponding formats);

[0036] - A group represented by the fact that all decoding neural networks in this group can be decoded independently of each other; and / or

[0037] - A group represented by the fact that decoding at least one of the neural decoders may require access to neural networks not in this group.

[0038] In some embodiments, the bitstream includes data for identifying the at least one group of compressed neural networks.

[0039] In a particular embodiment, the data for identifying the at least one group takes the form of a flag or indicator.

[0040] In some embodiments, the bitstream includes an encoded video sequence, and the bitstream further includes data for identifying the compressed neural network in the at least one group of compressed neural networks to be used for decoding the current video sequence.

[0041] In some embodiments, the bitstream further includes data for identifying the compressed neural network in the at least one group of compressed neural networks to be used for decoding the next video sequence after the current video sequence.

[0042] This feature is particularly advantageous because it allows the decoding device to be notified in advance, so that the decoding device can initiate the generation of a neural network decoder, which is used to decode the next video sequence when receiving the group of parameters, without waiting to receive the next video sequence.

[0043] In some embodiments, the bitstream includes data for signaling that the bitstream further includes data that can be used to obtain refined data of the decoded neural network.

[0044] In some embodiments, the decoding method further includes determining whether the bitstream should be decoded by a decoding device suitable for obtaining a decoded neural network from at least one group of compressed neural networks based on data for signaling whether the bitstream should be decoded by a decoding device suitable for obtaining a decoded neural network from at least one group of compressed neural networks.

[0045] In some embodiments, the decoding method further includes determining whether the decoding device is suitable for accessing at least one group of compressed neural networks with predetermined requirements based on data for signaling predetermined characteristics of a group of compressed neural networks suitable for decoding the bitstream.

[0046] In some embodiments, the decoding method further includes determining, based on data for identifying a compressed neural network to be used for decoding the current video sequence, whether the decoding device is capable of obtaining a neural network suitable for decoding the current sequence of the bitstream.

[0047] According to a second aspect, the present invention relates to a decoding device, which includes at least one processor and a memory, and a program for implementing the method for decoding a bitstream according to the present invention is stored on the memory.

[0048] Embodiments of the present invention also extend to a program that causes the computer or processor to execute the above method when running on a computer or processor, or extend to a program that causes the device to become the above device when loaded into a programmable device. The program can be provided by itself or carried by a carrier medium. The carrier medium can be a storage medium or a recording medium, or it can be a transmission medium, such as a signal. The program embodying the present invention can be transient or non-transient.

[0049] According to a third aspect, the present invention relates to a bitstream, which includes at least one of the following:

[0050] - data for signaling whether the bitstream should be decoded by a decoding device, the decoding device being adapted to obtain a decoded neural network from at least a set of compressed neural networks;

[0051] - data for signaling predetermined characteristics of a set of compressed neural networks suitable for decoding the bitstream;

[0052] - data for identifying at least a set of compressed neural networks; and

[0053] - data for identifying a compressed neural network to be used for decoding a current video sequence in the at least a set of compressed neural networks.

[0054] According to a fourth aspect, the present invention relates to a method for encoding media data, the method including, at an encoding device:

[0055] - selecting at least a set of encoding neural networks suitable for encoding the media data by applying a neural network trained to determine an optimal set of encoding neural networks according to a predetermined criterion; and

[0056] - encoding the media data by using the encoding neural networks in the selected set.

[0057] The method provides the following advantages: It provides an encoding device that accelerates the search for a suitable neural network among multiple sets of encoding neural networks configured to process a wide range of media content. The set of encoding neural networks is selected according to a predetermined criterion, which can correspond to a standard performance in terms of compression or quality.

[0058] The media data can correspond to audio data, video data, and / or still images.

[0059] In some embodiments, the encoding method further includes selecting an encoding neural network in at least one selected set by applying a neural network trained to identify the encoding neural network in the set that meets the predetermined criterion.

[0060] In some embodiments, the encoding method further includes selecting an encoding neural network in at least one selected set by:

[0061] - encoding the media data using a number of encoding neural networks in the at least one set; and

[0062] - selecting the encoding neural network in the at least one set that meets the predetermined criterion.

[0063] According to a fifth aspect, the present invention relates to a method for decoding a bitstream, the method including, at a decoding device:

[0064] - obtaining data related to at least one set of multiple neural networks;

[0065] - obtaining a decoded neural network from the obtained data, the neural network in the at least one set being available for the decoding device; and

[0066] - decoding the bitstream using the obtained neural network.

[0067] In some embodiments, the bitstream includes data for identifying the at least one set of neural networks.

[0068] In some embodiments, the neural network in the at least one set available for the decoding device is a compression neural network.

[0069] In some embodiments, the neural network in the at least one set is adapted to decode a video bitstream having the same type of content.

[0070] In some embodiments, obtaining the decoded neural network includes decoding the compression neural network in the at least one set, and the compression neural network can be decoded using decoded differential data with reference to a reference neural network.

[0071] In some embodiments, the bitstream includes data for signaling whether the bitstream should be decoded by a decoding device adapted to obtain neural networks from at least one set of neural networks.

[0072] In some embodiments, the bitstream includes data for signaling predetermined characteristics of a set of neural networks adapted to decode the bitstream.

[0073] In some embodiments, the bitstream includes an encoded video sequence, and the bitstream further includes data for identifying a neural network within the at least one set of neural networks to be used for decoding the current video sequence.

[0074] In some embodiments, the bitstream further includes data for identifying a neural network within the at least one set of neural networks to be used for decoding the next video sequence after the current video sequence.

[0075] In some embodiments, the bitstream includes data for signaling that the bitstream further includes refinement data that can be used to obtain a refined neural network.

[0076] In some embodiments, the decoding method of the fifth aspect further includes determining whether the bitstream should be decoded by a decoding device adapted to obtain neural networks from at least one set of neural networks, based on data for signaling whether the bitstream should be decoded by a decoding device adapted to obtain neural networks from at least one set of neural networks.

[0077] In some embodiments, the decoding method of the fifth aspect further includes determining whether the decoding device is adapted to access at least one set of neural networks with predetermined requirements, based on data for signaling predetermined characteristics of a set of compressed neural networks adapted to decode the bitstream.

[0078] In some embodiments, the decoding method of the fifth aspect further includes determining whether the decoding device is capable of obtaining a neural network adapted to decode the current sequence of the bitstream, based on data for identifying a neural network to be used for decoding the current video sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 Shows a particular embodiment of an encoding device according to the present invention;

[0080] Figure 2 Shows Figure 1 an example of the hardware architecture of an encoding device;

[0081] Figure 3 Shows a pair of encoding neural networks and decoding neural networks;

[0082] Figure 4 Shows the generation of a compression decoding neural network for decoding a bitstream;

[0083] Figure 5 Includes Figure 5 A, Figure 5 B, and Figure 5 C, and shows three examples of groups of compression neural networks that can be stored in the memory of a decoding device;

[0084] Figure 6 Is a flowchart of the steps of a method for encoding media data according to the present invention;

[0085] Figure 7 Depicted for illustrative purposes is the bitstream generated by the Figure 6 method for encoding media data;

[0086] Figure 8 Shows a specific embodiment of a decoding device according to the present invention;

[0087] Figure 9 Shows Figure 8 an example of the hardware architecture of an encoding device; and

[0088] Figure 10 Is a flowchart of the steps of a method for decoding a bitstream according to the present invention. Detailed Description

[0089] Figure 1 Shows a specific embodiment of an encoding device 10 according to the present invention.

[0090] The encoding device 10 includes a sequence processing module MOD_SEQ, a module MOD_DET for determining a neural network, an encoding module MOD_COD, and a module MOD_BT for generating a bitstream.

[0091] In an example, the encoding device 10 receives in input a time series S of digital images IMG for encoding in order to store it in a non-volatile memory (for future reading by a decoding device) or to send it (to a decoding device, such as the Figure 8 decoding device 20 described).

[0092] Each image received in input is represented, for example, by a 2D representation (such as a pixel matrix). In fact, each image received in input can be represented by multiple 2D representations (or pixel matrices), which respectively correspond to multiple signals, such as RGB color signals, or in a variant to a luminance signal and at least one chrominance signal.

[0093] As described with respect to Figure 6As described, the sequence processing module MOD_SEQ is configured to determine whether an image IMG received in the input corresponds to the first image of a video sequence and is configured to determine whether a new current encoding neural network should be selected to encode the image (and thus the corresponding sequence). The sequence processing module MOD_SEQ is also configured to generate a set of parameters associated with the current sequence.

[0094] As described below, the encoding device produces in the output a plurality of data units DU respectively associated with a plurality of images in a time sequence. A "video sequence" can be defined as a set of consecutive images respectively associated with a plurality of data units DU representing different images among the plurality of images. In other words, each image of the video sequence can be decoded based on the data unit DU associated with the image of the video sequence without referring to other data units associated with images that do not belong to the video sequence.

[0095] The module MOD_DET for determining the neural network is configured to determine an encoding neural network for encoding the current video sequence in the time sequence S. Different ways for implementing this step are described below.

[0096] Figure 3 Shows a pair of an encoding neural network and a decoding neural network. As Figure 3 shown, the encoding neural network EN is associated with the decoding neural network DN, which can be used to decode the data unit DU generated by the encoding neural network EN. The encoding neural network EN and the decoding neural network DN form a pair of artificial neural networks.

[0097] Each of the encoding neural network EN and the decoding neural network DN is defined by a structure that can include multiple layers and / or a plurality of weights respectively associated with the neurons of the neural network involved.

[0098] Apply a representation (e.g., 2D representation) of the current image IMG (or in a variant, a signal or data block of the current image IMG) to the input of the encoding neural network EN (e.g., on the input layer). The encoding neural network EN then outputs the encoded data represented by the Figure 3 data unit DU in. In a variant, a video sequence is applied to the input of the encoding neural network, which produces a data unit DU representing the entire video sequence.

[0099] Next, as regarding Figure 8 and Figure 10More specifically, the encoded data (here, the data unit DU) is applied to the input of the decoding neural network DN. The decoding neural network DN then outputs a representation IMG' (which can be a 2D representation) corresponding to the decoded version of the current image IMG (or in a variant, a signal or data block of the current image IMG). In a variant, the decoding neural network produces a video sequence corresponding to the decoded version of the current video sequence.

[0100] The encoding neural network EN is configured such that the amount of data included in the data unit DU is less than the amount of data of the representation of the current image IMG (or in a variant, at least on average for at least one subset of the images of the video sequence). In other words, the data unit DU is generally not only encoded but also compressed.

[0101] The encoding neural network and the decoding neural network were previously trained (e.g., by applying a large number of images to the input of the encoding neural network) to minimize the difference between the representation of the image IMG in the input and its representation IMG' in the output, and also to minimize the amount of data transmitted between the encoding neural network EN and the decoding neural network DN.

[0102] As further described below, the module MOD_DET for determining the neural network can select the encoding neural network among multiple different encoding and decoding neural networks (which respectively belong to several pairs of encoding and decoding neural networks).

[0103] The encoding module MOD_COD is configured to encode the representation of the current image IMG (or in a variant, a part of the current image or the entire video sequence) into encoded data (here encoded as the data unit DU) using the encoding neural network EN determined by the module MOD_DET.

[0104] The encoding module receives, in the input, the representation of the current image IMG, or the representation of a block of the current image IMG, or the representations of multiple images IMG (which can be a 2D representation, such as a pixel matrix).

[0105] The module MOD_BT for generating the bitstream B is configured to:

[0106] - receive, in the input, a set of parameters associated with the current video sequence and the data unit DU; and

[0107] - generate the bitstream B based on the received data.

[0108] In fact, the module MOD_BT for generating the bitstream B can include a concatenation unit for aggregating the previously listed elements in a predefined manner and an entropy encoder configured to encode the aggregated data.

[0109] The bitstream B can then be transmitted to another electronic device (e.g., the decoding device 20, such as the decoding device 20 described with respect to Figure 8 ), or stored in a non-volatile memory (such as the hard disk drive or non-volatile memory 4 of the encoding device 10) for future reading and decoding.

[0110] Figure 2 FIG. shows an example of the hardware architecture of an encoding device according to the present invention.

[0111] To this end, the encoding device 10 has the hardware architecture of a computer. As Figure 2 shown, the encoding device 10 includes a processor 1. Although shown as a single processor 1, two or more processors may be used depending on the specific needs, expectations, or specific embodiments of the encoding device 10. Generally, the processor 1 executes instructions and manipulates data to perform the operations of the encoding device 10 and any algorithms, methods, functions, procedures, processes, and programs described in this disclosure.

[0112] The encoding device 10 further includes a communication device 5. Although shown as a single communication device 5 in Figure 2 , two or more communication devices may be used depending on the specific needs, expectations, or specific embodiments of the encoding device 10. This communication device is used by the encoding device to communicate with another electronic device. Generally, the communication device 5 is operable to communicate with a radio network and includes logic encoded in software, hardware, or a combination of software and hardware. More specifically, the communication device 5 may include software that supports one or more communication protocols associated with the communication, such that the hardware of the radio network or interface is operable to transmit physical signals within and outside the shown encoding device.

[0113] The encoding device 10 further includes a random access memory 2, a read-only memory 3, and a non-volatile memory 4. The read-only memory 3 of the device constitutes a recording medium in accordance with the present invention, which can be read by the processor 1 and on which a computer program PROG_EN is recorded that is in accordance with the present invention and contains instructions for performing the steps of a method for encoding media data according to the present invention.

[0114] The program PROG_ENC defines functional modules of the device that are based on or control the above-described elements 1 to 5 of the encoding device, and these functional modules particularly include:

[0115] - a sequence processing module MOD_SEQ;

[0116] - a module MOD_DET for determining a neural network;

[0117] - an encoding module MOD_COD; and

[0118] - Module BT for generating bitstream B.

[0119] Figure 4 Shows the generation of compression - decoding neural networks, which can be used by a decoding device (such as the decoding device 20 described with respect to Figure 8 to decode a bitstream. In the example, the generation is implemented by the encoding device 10 shown in Figure 1 and the resulting compression - decoding neural networks are stored in the non - volatile memory 4 of the decoding device 20 or are available to the decoding device 20.

[0120] As described above, the encoding neural network 41 is associated with the decoding neural network 42, and both form an encoding - decoding neural network pair 40. In this example, four pairs of encoding - decoding neural networks { ; }, { ; }, { ; }, { ; } are considered. Each of these encoding - decoding neural networks is an "original" neural network, that is, they are not compressed and are not produced by decoding a compressed version of the neural network.

[0121] The decoding device of the present invention can access multiple compression - decoding neural networks, which are organized into "groups" or "families" of decoding neural networks.

[0122] In a particular embodiment, the decoding neural network is stored in the non - volatile memory 4 of the decoding device 20 in a compressed state. Storing a compressed version of the decoding neural network allows reducing the amount of data stored in the decoding device.

[0123] Furthermore, and as described further below, the decoding neural network can be encoded independently (i.e., without reference to another decoding neural network) or can be encoded with reference to at least one reference neural network.

[0124] In a particular embodiment, to determine whether a decoding neural network should be encoded with reference to another decoding neural network, the encoding device 10 calculates a distance representing the similarity between the decoding neural network to be encoded and another reference decoding neural network. If this distance is greater than a predetermined threshold (i.e., if the two decoding neural networks are very different), then the decoding neural network to be encoded is not encoded with reference to the involved reference decoding neural network. In a variant, all combinations between the decoding neural network to be encoded and candidates for the reference decoding neural network are tested (i.e., the distances representing the similarity between the decoding neural network to be encoded and each candidate are calculated), and the configuration that allows obtaining the best performance in terms of compression or quality is selected.

[0125] As further described below, a given decoding neural network can be encoded with reference to a number of reference decoding neural networks. For example, this is the case when a decoding neural network is encoded with reference to a first reference decoding neural network and when the first reference decoding neural network itself is encoded with reference to a second reference decoding neural network. In this example, the number of levels of incremental encoding is equal to two. However, there is no limit on the number of levels of incremental encoding that can be used to encode a given decoding neural network.

[0126] In a particular implementation, when a decoding neural network is to be independently encoded, a neural network encoder NNR_ENC that complies with the neural network compression and representation standard is applied to the decoding neural network involved. This neural network compression and representation standard can correspond to Part 17 of ISO / IEC 15938.

[0127] As Figure 4 shown, the decoding neural network is encoded using this neural network encoder NNR_ENC to generate a compressed version of the decoding neural network . Similarly, the decoding neural network is encoded using the neural network decoder NNR_ENC to generate a compressed version of the decoding neural network . Since they are both independently encoded, they are represented by double contour lines.

[0128] In a particular implementation, when a decoding neural network is to be encoded with reference to a reference decoding neural network, the encoding device is configured to:

[0129] - Apply a neural network decoder to the encoded version of the reference decoding neural network to obtain a decoded version of the reference decoding neural network; and

[0130] - Apply an incremental neural network encoder ICNN_ENC that takes as input the decoding neural network to be encoded and the decoded version of the reference decoding neural network to obtain a compressed version of the decoding neural network to be encoded.

[0131] If several levels of incremental encoding are needed to obtain a compressed version of the decoding neural network, these steps are iterated for each level.

[0132] As Figure 4 shown, to encode the decoding neural network with reference to the reference decoding neural network , since the decoding neural network has already been independently encoded, it is necessary to perform the operation on the decoding neural network A compressed version Apply the neural network decoder NNR_DEC that complies with the neural network compression and representation standard to obtain a decompressed version This decompressed version and the original decoded neural network are input into the incremental neural network encoder ICNN_ENC to obtain a compressed version of the decoded neural network A compressed version .

[0133] To encode with reference to the decoded neural network the decoded neural network Since the decoded neural network has been encoded with reference to Therefore, the incremental neural network decoder ICNN_DEC is to be applied. This incremental neural network decoder takes the compressed version of the decoded neural network A compressed version and the decompressed version of the decoded neural network in the input to obtain a decompressed version . This decompressed version and the original decoded neural network are input into the incremental neural network encoder ICNN_ENC to obtain a compressed version of the decoded neural network A compressed version .

[0134] The features (or "metadata") of these compressed decoded neural networks (including the reference REF1 between the compressed version of the decoded neural network A compressed version and the compressed version of the decoded neural network A compressed version ; and the reference REF2 between the compressed version of the decoded neural network A compressed version and the compressed version of the decoded neural network A compressed version ) are stored, for example, in the non-volatile memory 4 of the decoding device 20 or are available for the decoding device 20 through download.

[0135] Figure 5 Include Figure 5 A, Figure 5 B and Figure 5 C, and show three examples of groups of compressed neural networks that can be stored in the memory of the decoding device.

[0136] Groups of encoding and decoding neural network pairs

[0137] As described above, an encoding neural network is associated with a decoding neural network , and the two form an encoding-decoding neural network pair ( , ). Each pair of encoding-decoding neural networks is associated with a group (or "family") of encoding-decoding neural networks. These groups can include multiple pairs of encoding and decoding neural networks.

[0138] The encoding-decoding neural network pairs in a given group are adapted to process images of a given type of content. In other words, when these encoding-decoding neural networks process (i.e., encode and / or decode) video data of a given type of content, they achieve a performance in compression that is higher than a predetermined threshold (e.g., according to a rate-distortion optimization criterion or according to a value representing the quality of the representation decoded by these decoding neural networks).

[0139] In a particular embodiment, the predetermined threshold represents or is equal to the best performance of other encoding networks trained for various types of content.

[0140] Further details regarding the rate-distortion optimization criterion can be found in Y. Shoham, A. Gersho, "Efficient bit Allocation for an Arbitrary set of Quantizers [Efficient bit Allocation for an Arbitrary set of Quantizers]", IEEE Acoustics, Speech and Signal Processing [IEEE Transactions on Acoustics, Speech, and Signal Processing], Vol. 36, No. 9, September 1988.

[0141] In other words, a group of encoding-decoding neural network pairs is designed such that all the encoding-decoding neural networks in the group are adapted to a particular type of content, e.g., sports, movies, video conferencing, synthetic content, gaming, and / or medical imaging. In fact, the construction of a given group of encoding neural networks can be applied by training a group of encoding-decoding neural network pairs EN, DN using input training images corresponding to a particular type of content (e.g., sports), and each pair of encoding-decoding neural networks can be trained to adapt to a particular sport (e.g., football, horseback riding, underwater activities, etc.).

[0142] In a particular embodiment, a number of groups of encoding-decoding neural networks are constructed. Some groups can be designed to cover a particular type of content, and some groups can be designed to cover all expected types of content.

[0143] Group of compression decoding neural networks

[0144] Similar to how pairs of encoding and decoding neural networks are associated with a "group" that characterizes the media content types they are particularly suited for, a compressed version of the decoding neural networks available to the decoding device is also associated with a "group".

[0145] More precisely, consider for example, a pair of encoding and decoding neural networks particularly suitable for processing media content of type , then the group of pairs is associated with a corresponding group that includes the compressed versions of the decoding neural networks in the group of pairs.

[0146] In other words, consider the case of an encoding and decoding neural network pair (EN, DN) belonging to a group , then the compressed version of the decoding neural network belongs to a group of compressed decoding neural networks , and both groups and are particularly suitable for the same type of media content.

[0147] Figure 5 A shows a first example of a set of compressed neural networks that can be stored in the memory of a decoding device.

[0148] The group 51 of compressed decoding neural networks includes the compressed versions of the (independently encoded) decoding neural networks , the compressed versions of the decoding neural networks encoded with reference to the decoding neural network - or more precisely, encoded with reference to the decompressed version of the decoding neural network , the compressed versions of the decoding neural networks encoded with reference to the decoding neural network - or more precisely, encoded with reference to the decompressed version of the decoding neural network , the compressed versions of the decoding neural networks encoded with reference to the decoding neural network , and the compressed versions of the (independently encoded) decoding neural networks . .

[0149] Figure 5 B shows a second example of a set of compressed neural networks that can be stored in the memory of a decoding device.

[0150] The group 52 of compressed decoding neural networks includes the compressed versions of the (independently encoded) decoding neural networks , the compressed versions of the (independently encoded) decoding neural networks ​, (independently encoded) decoding neural network compressed version , and (independently encoded) decoding neural network compressed version .

[0151] Figure 5 C shows a third example of a set of compressed neural networks that can be stored in the memory of a decoding device.

[0152] The set 53 of compressed decoding neural networks includes the compressed version of the decoding neural network , the compressed version of the decoding neural network , and the reference decoding neural network compressed version , and the decoding neural network encoded with reference to the reference decoding neural network and the reference decoding neural network compressed version . These two compressed decoding neural networks . These two compressed decoding neural networks and are encoded with reference to a decoding neural network that does not belong to set 53 . In fact, this reference decoding neural network can be considered as the default reference decoding neural network, which is stored in the non-volatile memory 4 of the decoding device or is available for the decoding device through downloading. In this example, the number of levels of incremental encoding is equal to two.

[0153] As detailed below, the features of the decoding neural network that do not belong to a specific group and require a reference decoding neural network can be indicated in the bitstream through specific signaling.

[0154] Figure 6 is a flowchart of the steps of a method for encoding media data implemented by the encoding device 10.

[0155] This method for encoding media data is applied to the time series S of the image IMG. This time series S forms a video that should be encoded before transmission or storage.

[0156] As Figure 6 shown, the method for encoding media data includes a first step S10, during which the encoding device (here, the sequence processing module MOD_SEQ of the encoding device 10) selects the first image of the time series as the "current image".

[0157] A method for encoding media data then includes step S11, in which an encoding device 10 (here, the sequence processing module MOD_SEQ of the encoding device 10) determines whether the current image IMG is selected as the first image of a video sequence. The criteria considered for this determination can depend on the application under consideration.

[0158] In an example (which can ultimately be used when encoding media data intended to be broadcast on a television), the encoding device 10 (here, the sequence processing module MOD_SEQ) starts a new video sequence every second. Thus, the maximum time period required to wait for a portion of the bitstream related to the new video sequence is at most 1 second (e.g., which is the time period required when switching between two channels).

[0159] In a variant, a new video sequence is started when a new shot in the video is detected. This new shot can be automatically determined by analyzing the image content of a time series.

[0160] In another variant, sequences with a longer time period (e.g., between 10 seconds and 40 seconds) can be considered in order to reduce the bit rate, because encoding the first image of a video sequence requires a larger amount of data to reconstruct the image based only on this data (i.e., without accessing data associated with another video sequence).

[0161] If it is determined at step S11 that the current image IMG is selected as the first image of a video sequence (e.g., if the encoding device starts encoding a new video sequence), then the method continues at step S12.

[0162] Conversely, if it is determined at step S11 that the current image IMG is not selected as the first image of a video sequence (e.g., if the encoding device continues encoding the current sequence), then the method continues at step S14.

[0163] At step S12, the encoding device (here, the module MOD_DET for determining the neural network) determines the encoding neural network EN to be associated with the current video sequence. The encoding neural network EN thus determined becomes the "current encoding neural network".

[0164] In a particular implementation, the module MOD_DET for determining the neural network selects the current encoding neural network from a predetermined set of encoding neural networks (e.g., from a set of encoding neural networks respectively associated with decoding neural networks available for a given decoding device), and the previous data transmission between the encoding device and the decoding device can identify this set of neural networks.

[0165] In a particular embodiment, the module MOD_DET for determining the neural network selects the encoded neural network EN that achieves a performance in terms of compression higher than a predetermined threshold (e.g., according to a rate-distortion optimization criterion or according to a value representing the quality of the decoded video sequence). In a particular embodiment, the predetermined threshold represents or is equal to the best performance of other encoded networks trained for various contents.

[0166] In a particular embodiment, the current image IMG represents the content of the entire video sequence, and the module MOD_DET for determining the neural network inputs the current image IMG into a number of encoded neural networks (each encoded neural network is ultimately associated with a decoded neural network), and selects the encoded neural network EN that allows achieving the best performance in terms of compression or quality, as described above.

[0167] In a variant, a set of images of the video sequence can be successively applied to the inputs of different encoded neural networks EN in order to select the encoded neural network EN that allows achieving the best performance in terms of compression or quality.

[0168] In a particular embodiment, the encoded neural networks that can be selected by the encoding device 10 are grouped into "groups" or "families" (as explained in reference Figure 5 ), and the encoded neural networks in a given group are suitable for encoding video sequences with the same type of content (e.g., for TV programs characterized by motion, the encoded neural networks in the group are suitable for processing motion content). In some embodiments, the encoded neural networks in a given group are suitable for encoding a particular "variety of contents". For example, for a group of encoded neural networks suitable for processing motion content, some encoded neural networks are suitable for processing fast motion actions, while other encoded neural networks are suitable for processing content with slow movement, such as when a TV host appears.

[0169] Then, the module MOD_DET for determining the neural network is configured to:

[0170] - Analyze the content of the current image IMG (or multiple images of the current sequence, including the current image IMG), for example to identify the type of relevant content (e.g., motion, computer-generated images, movies, video conferences, etc.),

[0171] - According to the result of this analysis, for example, select a group of neural networks according to the type of the identified content;

[0172] - And then select the encoded neural network in the selected group. In a particular embodiment, the selected encoded neural network EN is the encoded neural network EN in the group that allows achieving the best performance in terms of compression or quality.

[0173] In a variant, the module MOD_DET for determining the neural network can determine the encoded neural network EN as a result of a machine - learning - based method. Such a machine - learning - based method includes, for example, the following steps:

[0174] - Input the current image (or multiple images of the current sequence, including the current image IMG) into a neural network trained to analyze the content of the current image IMG (or multiple images of the current sequence, including the current image IMG) in order to determine, among multiple sets of encoded neural networks, a set of encoded neural networks suitable for encoding the content type of the current image (or multiple images of the current sequence).

[0175] - Select the encoded neural network from the selected set. In a particular embodiment, the selected encoded neural network EN is the one in the set that allows for the best performance in terms of compression. In a particular embodiment, this selection is implemented by a neural network trained to select the encoded neural network from the set according to a predetermined criterion.

[0176] The encoding method then includes the step S13 of generating a set of parameters related to the current sequence.

[0177] In a particular embodiment, the set of parameters includes data AF for signaling whether the bitstream should be decoded by a decoding device suitable for obtaining a decoded neural network from at least one set of compressed neural networks. The data AF can take the form of a flag AF that signals the possibility (or impossibility) of using a decoding device suitable for obtaining a decoded neural network from at least one set of compressed neural networks. In this case, when the decoding device according to the invention should be used, the encoding device (here the processing sequence module MOD_SEQ) can set the flag AF to "1", otherwise to "0".

[0178] In a particular embodiment, the set of parameters further includes data BT for signaling predetermined characteristics of a set of compressed neural networks suitable for decoding the desired characteristics of the bitstream or the decoder. The data BT can indicate:

[0179] - A group including decoding neural networks suitable for a particular type of content, such as sports, synthetic images, medical images,...

[0180] - A group characterized by the maximum incremental decoding level of its neural decoder. For example, the data BT indicates that up to two decoding neural networks need to be decompressed before a given requested decoding neural network identified by the flag I (described below) can be decompressed;

[0181] - a group characterized by the fact that a certain decoding method is required to decompress these compressed decoding neural networks (e.g., NNR decoding or ONNX decoding according to these corresponding formats);

[0182] - a group characterized by the fact that all decoding neural networks in this group can be decoded independently of each other; and / or

[0183] - a group characterized by the fact that decoding at least one of the neural decoders may require a neural network not in this group.

[0184] In a particular embodiment, the required decoding neural network belongs to several groups, and the data BT indicates multiple criteria. For example, the required decoding neural network should be decoded independently and belongs to a group of decoding neural networks suitable for processing medical images.

[0185] In a particular embodiment, the group of parameters further includes data BI representing a group of decoding neural networks used to decode the current sequence. This data may correspond to the identifiers of a group of decoding neural networks identified by module MOD_DET at step S12.

[0186] In a particular embodiment, the group of parameters includes data I for identifying the decoding neural network DN in a predetermined group of decoding neural networks, such as the identifier of the decoding neural network DN. In a preferred embodiment, data I identifies the decoding neural network in the group of decoding neural networks identified by data BI. In this case, both data BI and data I representing the group of decoding neural networks to be used are required to identify a single decoding neural network.

[0187] In a particular embodiment, the group of parameters further includes data ERF for signaling whether there is additional refinement data in the bitstream that can be used to obtain the decoding neural network. In an example, the signaling data ERF takes the form of a flag, and the encoding device (here, the processing sequence module MOD_SEQ) can set the flag ERF to "1" when the bitstream includes additional refinement data, and set the flag to "0" when the bitstream does not include additional refinement data associated with the decoding neural network that should be used to decode the current sequence. At step S12, it is determined by the module MOD_DET for determining the neural network whether such refinement data exists.

[0188] This additional refinement data may represent additional layers to be inserted into the structure of the identified decoding neural network (identified in the bitstream by the data I for identifying the decoding neural network DN). In an example, for each additional layer, the additional refinement data describes the structure of the involved additional layer, the weights of the neurons forming the involved additional layer, and / or data for locating the positions within the structure of the decoding neural network where these additional layers should be inserted.

[0189] This set of parameters may further include data representing the image format of the current sequence, and / or data representing the maximum size of the images of the current sequence, and / or data representing the minimum size of the image blocks of the current sequence.

[0190] This set of parameters is inserted into the bitstream B to be transmitted (here by the module MOD_BT for generating the bitstream), for example, into the header of the part of the bitstream associated with the current video sequence.

[0191] The encoding method further includes a step S14 of encoding the current image IMG using the current encoding neural network, which generates a data unit DU associated with the current image IMG in the output. This data unit DU can be inserted into the bitstream B (here by the module MOD_BT for generating the bitstream).

[0192] Then, the encoding device 10 selects the next image in the time sequence as the new current image (step S15), and loops back to step S11.

[0193] Figure 7 For illustrative purposes, a bitstream generated by Figure 6 the method for encoding media data is depicted.

[0194] This bitstream B includes:

[0195] - For each video sequence, a header associated with the involved video sequence, the header including a set of parameters related to the involved video sequence;

[0196] - For each image of the involved video sequence, a data unit DU (including encoded data) relative to the involved image.

[0197] In an example and as Figure 7As shown, the set of parameters of the bitstream B includes data AF for signaling whether the bitstream should be decoded by a decoding device, which is adapted to obtain a decoded neural network from at least one set of compressed neural networks. The data AF can take the form of a flag AF, which signals the possibility (or impossibility) of using a decoding device adapted to obtain a decoded neural network from at least one set of compressed neural networks. In this case, when a decoding device according to the present invention can be used, the flag AF can be "1", otherwise it can be "0".

[0198] The set of parameters further includes data BT for signaling predetermined characteristics of a set of compressed neural networks suitable for decoding the bitstream. The data BT can indicate:

[0199] - A group including decoded neural networks that are suitable for a specific type of content, such as sports, synthetic images, medical images,...

[0200] - A group characterized by the maximum incremental decoding level of its neural decoders. For example, the data BT indicates that before a given requested decoded neural network identified by flag I (described below) can be decompressed, at most two decoded neural networks need to be decompressed;

[0201] - A group characterized by the fact that a certain decoding method is required to decompress these compressed decoded neural networks (e.g., NNR decoding or ONNX decoding according to these corresponding formats);

[0202] - A group characterized by the fact that all decoded neural networks in the group can be decoded independently of each other; and / or

[0203] - A group characterized by the fact that decoding at least one of the neural decoders may require neural networks not in the group.

[0204] In this example, the set of parameters of the bitstream B further includes data BI representing a set of decoded neural networks for decoding the current sequence. The data can correspond to identifiers of a set of decoded neural networks.

[0205] In this example, the set of parameters further includes data ERF for signaling whether there is additional refinement data ERD in the bitstream that can be used to obtain a decoded neural network. The signaling data ERF can take the form of a flag, and in this case, when the bitstream includes additional refinement data ERD, the flag ERF can be set to "1", and when the bitstream does not include additional refinement data ERD associated with the decoded neural network that should be used to decode the current sequence, the flag can be set to "0".

[0206] As already indicated above, this additional refinement data ERD can represent additional layers of a decoding neural network that should be inserted into the structure of the identified decoding neural network (identified in the bitstream by the data I for identifying the decoding neural network DN). In an example, for each additional layer, this additional refinement ERD data describes the structure of the additional layer involved, the weights of the neurons forming the additional layer involved, and / or data for locating the positions within the structure of the decoding neural network where these additional layers should be inserted.

[0207] Figure 8 Shows a specific embodiment of a decoding device according to the present invention.

[0208] The decoding device 20 includes a module MOD_AN for analyzing the bitstream, a module MOD_GEN for generating a decoded version of the decoding neural network, a module MOD_DEC for decoding the bitstream, and a module MOD_SYN for analyzing syntax elements that can be a sub-module of the module MOD_AN for analyzing the bitstream.

[0209] In an example, the decoding device receives a bitstream at the input, here the bitstream B generated by the encoding device 10. The bitstream B includes several parts, each part corresponding to a video sequence. Thus, each part includes a plurality of data units DU respectively associated with different images of the video sequence.

[0210] The module MOD_AN for analyzing the bitstream is configured to analyze the bitstream in order to identify different data in the bitstream B, such as Figure 7 the data shown. In fact, the module MOD_AN for analyzing the bitstream can be configured to implement an entropy decoding method of the bitstream B to obtain this data.

[0211] As Figure 8 shown, the module MOD_AN for analyzing the bitstream includes the module MOD_SYN for analyzing syntax elements. More precisely, the module MOD_SYN is configured to detect the data I for identifying a decoding neural network DN in a predetermined set of decoding neural networks. As already described above, this data can take the form of an identifier of the decoding neural network DN.

[0212] As indicated above, the decoding device can access a plurality of decoding neural networks, which are organized into "groups" of neural networks. The decoding neural networks in a given group are particularly suitable for processing video sequences with a given type of content. In other words, when they process video data with a given type of content, the decoding neural networks in a given group achieve performance in compression higher than a predetermined threshold (e.g., according to a rate-distortion optimization criterion or according to a value representing the quality of the representation decoded by these decoding neural networks).

[0213] Thus, in a particular embodiment, the module MOD_SYN is further configured to detect data BI representing a set of decoding neural networks for decoding the current video sequence. This data may correspond to identifiers of a set of decoding neural networks.

[0214] In a particular embodiment, the module MOD_SYN is further configured to detect data AF for signaling whether the bitstream should be decoded by a decoding device according to the present invention. As indicated above, this data AF may take the form of a flag AF that signals the possibility (or impossibility) of using such a decoding device. In an example, if the value of the flag AF is equal to "1", the module MOD_AN for analyzing the bitstream determines that the decoding device 20 may be suitable for decoding data associated with the video sequence.

[0215] In a particular embodiment, the module MOD_SYN is further configured to detect data BT for signaling predetermined characteristics of a set of compression neural networks suitable for decoding the bitstream. As already indicated above, this data BT may indicate:

[0216] - A group including decoding neural networks that are suitable for a particular type of content, such as sports, synthetic images, medical images,...

[0217] - A group characterized by the maximum incremental decoding level of its neural decoders. For example, the data BT indicates that up to two decoding neural networks need to be decompressed before a given requested decoding neural network identified by a flag I (described below) can be decompressed;

[0218] - A group characterized by the fact that a certain decoding method is required to decompress these compressed decoding neural networks (e.g., NNR decoding or ONNX decoding according to these corresponding formats);

[0219] - A group characterized by the fact that all decoding neural networks in the group can be decoded independently of each other;

[0220] - A group characterized by the fact that decoding at least one of the neural decoders may require neural networks not in the group.

[0221] Then, the module MOD_AN analyzes this data BT, and this module may determine whether the decoding device 20 is suitable for decoding the video sequence.

[0222] In a particular embodiment, the module MOD_SYN is also configured to extract from the bitstream data ERF that signals whether additional refinement data for obtaining a decoded neural network is present in the bitstream. The signaling data ERF may take the form of a flag. In an example, if the flag ERF is set to "1", the decoding device 20 (here the module MOD_AN for analyzing the bitstream) determines that the bitstream includes refinement data for obtaining a decoded neural network and extracts the refinement data. Conversely, if the flag ERF is set to "0", the decoding device 20 (here the module MOD_AN for analyzing the bitstream) determines that the bitstream does not include refinement data and no further refinement data extraction is required.

[0223] The module MOD_AN for analyzing the bitstream is further configured to identify data units within the bitstream B and transmit each data unit DU to the decoding module MOD_DN, as described below.

[0224] The decoding device 20 also includes a module MOD_GEN for generating a decoded version of the decoded neural network.

[0225] Since the decoding device 20 can access multiple compressed decoded neural networks organized in "groups", the module MOD_GEN is configured to identify a particular decoded neural network from the multiple decoded neural networks, for example, according to data I for identifying the decoded neural network DN. As already described above, this data may take the form of an identifier of the decoded neural network DN. Since these decoded neural networks are compressed, the module MOD_GEN for generating a decoded version of the decoded neural network is configured to decode the compressed decoded neural network.

[0226] More precisely, according to an embodiment, the compressed decoded neural network can be decoded independently, i.e., without data related to a decoded neural network different from the decoded neural network identified by the data I for identifying the decoded neural network. In a variant, the compressed decoded neural network can be decoded with reference to at least one reference neural network. As indicated above, this information ("metadata") related to the way the compressed decoded neural network can be decoded is stored, for example, in the non-volatile memory 4 of the decoding device 20 or is available for the decoding device by download.

[0227] Therefore, the module MOD_GEN for generating a decoded version of the decoded neural network is configured to access this information to determine whether the compressed decoded neural network should be decoded independently or, conversely, should be decoded with reference to at least one reference neural network.

[0228] In a particular embodiment, when the compressed decoding neural network is to be decoded independently, the module MOD_GEN for generating a decoded version of the decoding neural network is configured to apply a neural network decoder compliant with the neural network compression and representation standard to the compressed decoding neural network.

[0229] In a particular embodiment, when the compressed decoding neural network is to be decoded with reference to a reference neural network, MOD_GEN for generating a decoded version of the decoding neural network is configured to:

[0230] - obtain data for identifying a reference decoded neural network from an analysis of information associated with the decoded neural network identified by data I for identifying the decoded neural network;

[0231] - obtain a compressed version of the reference decoded neural network;

[0232] - apply a neural network decoder to the reference decoded neural network to obtain a decoded version of the reference decoded neural network;

[0233] - apply an incremental neural network decoder that takes as input the decoded version of the reference neural network and decoded differential data associated with the compressed neural network to obtain a decoded version of the decoded neural network.

[0234] As already indicated, the decoded neural network DN and the reference decoded neural network may have the same structure, and the decoded differential data may include data representing the difference between each weight of the decoded neural network DN and the corresponding weight of the reference decoded neural network.

[0235] If several levels of incremental decoding are required to obtain a decoded version of the decoded neural network, these steps are re-iterated for each level.

[0236] As indicated above, in a particular embodiment, the additional refinement data may represent additional layers to be inserted into the structure of the decoded neural network identified by the data I for identifying the decoded neural network DN (identified in the bitstream by the data I for identifying the decoded neural network DN). Thus, when the data ERF signals the presence in the bitstream of additional refinement data that can be used to obtain the decoded neural network, the module MOD_GEN for generating a decoded version of the decoded neural network is configured to:

[0237] - obtain a decoded version of the compressed decoding neural network identified by the data I for identifying the neural network; and

[0238] - insert the additional layers into the structure of the decoded version of the decoded neural network at positions determined according to the additional refinement data.

[0239] The module MOD_GEN for generating a decoded version of the decoding neural network is further configured to configure the decoding module MOD_DN such that the decoding module MOD_DN can decode the current data unit DU and ultimately decode at least one other data unit DU associated with the current video sequence.

[0240] Figure 9 An example of the hardware architecture of a decoding device according to the present invention is shown.

[0241] To this end, the decoding device 20 has the hardware architecture of a computer. As Figure 9 shown, the decoding device 20 includes a processor 1. Although shown as a single processor 1, two or more processors may be used according to the specific needs, expectations, or specific embodiments of the decoding device 20. Generally, the processor 1 executes instructions and manipulates data to perform the operations of the decoding device 20 and any algorithms, methods, functions, processes, flows, and programs described in this disclosure.

[0242] The decoding device 20 further includes a communication device 5. Although shown as a single communication device 5 in Figure 9 , two or more communication devices may be used according to the specific needs, expectations, or specific embodiments of the decoding device 20. The communication device is used by the decoding device 20 to communicate with another electronic device (e.g., an encoding device, such as Figure 1 the encoding device 10 shown). Generally, the communication device 5 is operable to communicate with a radio network and includes logic encoded in software, hardware, or a combination of software and hardware. More specifically, the communication device 5 may include software that supports one or more communication protocols associated with communication such that the hardware of the radio network or interface is operable to transmit physical signals inside and outside the decoding device shown.

[0243] The decoding device 20 further includes a random access memory 2, a read-only memory 3, and a non-volatile memory 4. The read-only memory 3 of the device constitutes a recording medium according to the present invention, which can be read by the processor 1 and on which a computer program PROG_DEC is recorded according to the present invention and containing instructions for performing the steps of the method for decoding a bitstream according to the present invention.

[0244] The program PROG_DEC defines functional modules of the device that are based on or control the above elements 1 to 5 of the decoding device 20, and these functional modules particularly include:

[0245] - A module MOD_AN for analyzing the bitstream;

[0246] - A module MOD_GEN for generating a decoded version of the decoding neural network;

[0247] - Module MOD_DEC for decoding a bitstream; and

[0248] - Module MOD_SYN for analyzing syntax elements.

[0249] Figure 10 is a flowchart of steps of a method for decoding a bitstream implemented by decoding device 20.

[0250] Decoding device 20 receives, in an input, Figure 6 a bitstream B generated by a method for encoding media data and presented by Figure 7 the.

[0251] As Figure 10 shown, the method for decoding a bitstream includes a first step S20, during which the decoding device receives the bitstream.

[0252] Then, the method for decoding media data includes step S21, in which decoding device 20 (here, module MOD_AN for analyzing the bitstream) determines whether there is a header related to a video sequence in the received bitstream B. In fact, this determination can correspond to detecting data of a header specific to the video sequence in the bitstream.

[0253] If so, the method for decoding media data then includes step S22, in which decoding device 20 (here, module MOD_SYN for analyzing syntax elements) extracts a set of parameters related to the current sequence from the header determined in step S21.

[0254] The part of the bitstream associated with the current video sequence includes a set of parameters related to the current sequence and other data obtained until a new header related to a new video sequence is detected (e.g., data units related to images of the current video sequence).

[0255] In a particular embodiment, the set of parameters includes a data AF for signaling whether a bitstream should be decoded by a decoding device according to the invention. This data AF may take the form of a flag AF that signals the possibility (or impossibility) of using a decoding device adapted to obtain a decoded neural network from at least one set of compressed neural networks. In this case, the method for decoding the bitstream includes step S21, in which the decoding device 20 (here the module MOD_SYN for analyzing syntax elements) obtains the flag AF. At step S23, if a decoding device according to the invention should be used (for example, if the flag AF is set to "1"), the decoding method proceeds to step S24. Otherwise, if another decoding device should be used, the decoding method proceeds to step S35. At step S35, the current video sequence is skipped. In a variant, a preconfigured decoder is used.

[0256] In a particular embodiment, the set of parameters includes a data BT for signaling predetermined characteristics of a set of compressed neural networks adapted to decode a bitstream. As indicated above, this data BT may indicate:

[0257] - A group including decoding neural networks that are adapted to a particular type of content, such as sports, synthetic images, medical images,...

[0258] - A group characterized by the maximum incremental decoding level of its neural decoders. For example, the data BT indicates that up to two decoding neural networks need to be decompressed before a given requested decoding neural network identified by a flag I (described below) can be decompressed;

[0259] - A group characterized by the fact that a certain decoding method is required to decompress these compressed decoding neural networks (for example, NNR decoding or ONNX decoding according to these respective formats);

[0260] - A group characterized by the fact that all decoding neural networks in the group can be decoded independently of each other; and / or

[0261] - A group characterized by the fact that decoding at least one of the neural decoders may require access to neural networks not in the group.

[0262] In this case, the method for decoding the bitstream includes step S24, in which the decoding device 20 (here the module MOD_SYN for analyzing syntax elements) obtains the data BT for signaling predetermined characteristics of a set of compressed neural networks.

[0263] At step S25, the decoding device 20 (here, the module MOD_AN for analyzing the bitstream) determines whether at least one set of decoding neural networks available to the decoding device 20 complies with the requirements defined by the data BT, where the data BT signals predetermined characteristics of a set of compression neural networks.

[0264] If so, the decoding method proceeds to step S26. Otherwise, if a set that complies with the requirements defined based on the data BT cannot be accessed, the decoding method proceeds to step S35.

[0265] In a particular embodiment, the set of parameters further includes data BI representing a set of decoding neural networks for decoding the current sequence. The data BI may correspond to an identifier of a set of decoding neural networks.

[0266] In this case, the method for decoding the bitstream includes step S26, in which the decoding device 20 (here, the module MOD_SYN for analyzing syntax elements) obtains the data BI representing the set of decoding neural networks to be used.

[0267] At step S27, the decoding device 20 (here, the module MOD_AN for analyzing the bitstream) determines whether a set of decoding neural networks identified by the set of identifiers can be accessed. If so, the decoding method proceeds to step S28. Otherwise, the decoding method proceeds to step S35.

[0268] In a particular embodiment, the set of parameters further includes data I for identifying the decoding neural network DN in a predetermined set of decoding neural networks, such as an identifier of the decoding neural network DN.

[0269] In this case, the method for decoding the bitstream includes step S28, in which the decoding device 20 (here, the module MOD_SYN for analyzing syntax elements) obtains the identifier of the decoding neural network.

[0270] At step S29, the decoding device 20 (here, the module MOD_AN for analyzing the bitstream) determines whether the decoding neural network for which the ID has been obtained can be accessed. If so, the decoding method proceeds to step S30. Otherwise, if the decoding neural network for which the ID has been obtained cannot be accessed, the decoding method proceeds to step S35.

[0271] In a particular embodiment, the set of parameters further includes data ERF for signaling whether there is additional refinement data ERD in the bitstream that can be used to obtain a decoding neural network. In an example, the signaling data ERF takes the form of a flag, and the decoding method thus includes step S30, in which the decoding device 20 (here, the module MOD_SYN for analyzing syntax elements) obtains the flag ERF.

[0272] At step S31, if the bitstream includes additional refinement data ERD associated with the decoding neural network to be used for decoding the current sequence (e.g., if the flag ERF is set to "1"), the decoding method proceeds to step S32. At step S32, this additional refinement data ERD is extracted from the bitstream. As already explained, this additional refinement data ERD may represent additional layers to be inserted into the structure of the identified decoding neural network (identified in the bitstream by the data I for identifying the decoding neural network DN). In the example, for each additional layer, this additional refinement data describes the structure of the involved additional layer, the weights of the neurons forming the involved additional layer, and / or data for locating the positions within the structure of the decoding neural network where these additional layers should be inserted.

[0273] Otherwise, if the bitstream does not include additional refinement data associated with the decoding neural network to be used for decoding the current sequence (e.g., if the flag ERF is set to "0"), the decoding method proceeds to step S33.

[0274] At step 33, the decoding device 20 (here, the module MOD_GEN for generating the decoded version of the decoding neural network) generates the decoded version of the decoding neural network based on the data extracted by the module MOD_AN for analyzing the bitstream and the module MOD_SYN for analyzing the syntax elements.

[0275] As already detailed, the decoding device can access multiple compressed decoding neural networks organized into "groups". Thus, the module MOD_GEN can in particular identify a specific decoding neural network based on the data I obtained at step S28.

[0276] Since these decoding neural networks are compressed, the module MOD_GEN for generating the decoded version of the decoding neural network is configured to decode the compressed decoding neural network identified by this data I.

[0277] More precisely, according to a particular embodiment, the compressed decoding neural network can be decoded independently, i.e., without data related to a decoding neural network different from the decoding neural network identified by the data I for identifying the decoding neural network. In a variant, the compressed decoding neural network can be decoded with reference to at least one reference neural network. As indicated above, this information related to the way the compressed decoding neural network can be decoded is stored, for example, in the non-volatile memory 4 of the decoding device 20 or is available for the decoding device by download.

[0278] Thus, at step S33, the decoding device (here, the module MOD_GEN for generating a decoded version of the decoding neural network) accesses this information to determine whether the compressed decoding neural network should be decoded independently or should be decoded with reference to at least one reference neural network.

[0279] In a particular embodiment, when the compressed decoding neural network should be decoded independently, the module MOD_GEN for generating a decoded version of the decoding neural network applies a neural network decoder compliant with the neural network compression and representation standard to the compressed decoding neural network.

[0280] In a particular embodiment, when the compressed decoding neural network should be decoded with reference to a reference neural network, step S33 includes the following sub-steps implemented by the module MOD_GEN for generating a decoded version of the decoding neural network:

[0281] - Obtaining data for identifying a reference decoded neural network from an analysis of information associated with the decoding neural network identified by the data I for identifying the decoding neural network;

[0282] - Obtaining a compressed version of the reference decoded neural network;

[0283] - Applying a neural network decoder to the reference decoded neural network to obtain a decoded version of the reference decoded neural network; and

[0284] - Applying an incremental neural network decoder that takes as input the decoded version of the reference neural network and the decoded differential data associated with the compressed neural network to obtain a decoded version of the decoding neural network.

[0285] If several levels of incremental coding are required to obtain a decoded version of the decoding neural network, these sub-steps are re-iterated for each level.

[0286] Furthermore, if additional refinement data ERD has been extracted from the bitstream at step S32, step S33 includes the following sub-steps:

[0287] - Obtaining a decoded version of the compressed decoding neural network identified by the data I for identifying the neural network; and

[0288] - Inserting additional layers into the structure of the decoded version of the decoding neural network at positions determined based on the additional refinement data.

[0289] A method for decoding a bitstream then includes step S34, in which the decoding device 20 (here the decoding module MOD_DN) extracts from the bitstream a data unit DU associated with each image of the current video sequence, and the decoding module MOD_DN decodes this data unit DU by means of the current decoding neural network generated at step S33 to obtain a 2D representation IMG' of the images of the sequence.

[0290] So far, the present invention has been described in the case where a part of the bitstream is associated with the current video sequence and includes a set of parameters associated with this current sequence.

[0291] In a variant, this part of the bitstream more precisely includes a set of parameters associated with the next video sequence after the current video sequence. In this particular case, the bitstream includes data (I) for identifying the compression neural network to be used for decoding the next video sequence. This allows the decoding device to be notified in advance so that the decoding device can initiate the generation of a neural network decoder that is used to decode the next video sequence when this set of parameters is received, without having to wait to receive the next video sequence.

[0292] In a variant, this part of the bitstream includes the set of parameters associated with the current video sequence and also includes a set of parameters associated with the next video sequence after the current video sequence. In this particular case, the bitstream includes data (I) for identifying the compression neural network to be used for decoding the current video sequence and data (I) for identifying the compression neural network to be used for decoding the next video sequence.

[0293] So far, the present invention has been described in the case of a set of neural network representations being a set of neural networks adapted to process images of a given type of content, that is, these codec neural networks process (i.e., encode and / or decode) video data of a given type of content such that the performance they obtain in terms of compression is higher than a predetermined threshold. However, the present invention also applies to a set of neural network representations being a set of neural networks adapted to encode images in a certain way (e.g., according to a predetermined format, according to the type of sampling and sub-sampling, or using a certain coding method).

Claims

1. A method for decoding a bitstream (B), the method comprising, at a decoding device (20): - obtaining (S33) a decoded neural network from at least one set of a plurality of compressed neural networks, the compressed neural networks in the at least one set being available for the decoding device; and - decoding (S34) the bitstream using the obtained neural network.

2. The decoding method according to claim 1, wherein, The compressed neural networks in the at least one set are adapted to decode media data bitstreams having the same type of content.

3. The decoding method according to claim 1 or 2, wherein, Obtaining the decoded neural network includes decoding the compressed neural networks in the at least one set, wherein the compressed neural network can be decoded with reference to a reference neural network.

4. The decoding method according to claim 3, wherein, The reference neural network belongs to the at least one set.

5. The decoding method according to claim 3, wherein, The reference neural network does not belong to the at least one set.

6. The decoding method according to claim 1 or 2, wherein, Obtaining the decoded neural network includes decoding the compressed neural network by applying a neural network decoder compliant with a neural network compression and representation standard to the compressed neural network.

7. The decoding method according to any one of claims 1 to 6, wherein, The bitstream (B) includes data (AF) for signaling whether the bitstream should be decoded by the decoding device, the decoding device being adapted to obtain a decoded neural network from at least one set of compressed neural networks.

8. The decoding method according to any one of claims 1 to 7, wherein, The bitstream (B) includes data (BT) for signaling predetermined characteristics of a set of compressed neural networks adapted to decode the bitstream.

9. The decoding method according to any one of claims 1 to 8, wherein, The bitstream (B) includes data (BI) for identifying the at least one set of compressed neural networks.

10. The decoding method according to any one of claims 1 to 9, wherein, The bitstream (B) includes an encoded video sequence, wherein the bitstream (B) includes data (I) for identifying a compressed neural network in the at least one set of compressed neural networks to be used for decoding the current video sequence.

11. The decoding method according to claim 10, wherein, The bitstream (B) further includes data (I) for identifying a compressed neural network in the at least one set of compressed neural networks to be used for decoding the next video sequence after the current video sequence.

12. The decoding method according to any one of claims 1 to 11, wherein, The bitstream (B) includes data (ERF) for signaling that the bitstream further includes refinement data (ERC) that can be used to obtain the decoded neural network.

13. The decoding method according to claim 12, further comprising determining (S29), based on the data (I) for identifying the compressed neural network to be used for decoding the current video sequence, whether the decoding device can obtain a neural network adapted to decode the current sequence of the bitstream.

14. A decoding device (20) comprising at least one processor (1) and a memory (3), the memory storing a program (PROG_DEC) for implementing the method for decoding a bitstream according to any one of claims 1 to 13.

15. A computer program (PROG_DEC) which, when executed by a decoding device (20), causes the decoding device to execute the method for decoding a bitstream according to any one of claims 1 to 13.

16. A bitstream (B) comprising at least one of the following: - data (AF) for signaling whether the bitstream should be decoded by a decoding device, the decoding device being adapted to obtain a decoded neural network from at least one set of compressed neural networks; - Data (BT) signaling predetermined characteristics of a set of compressed neural networks suitable for decoding the bitstream; - Data (BI) for identifying at least one set of compressed neural networks; and - Data (I) for identifying, within the at least one set of compressed neural networks, the compressed neural network to be used for decoding the current video sequence.

17. A method for decoding a bitstream, the method comprising, at a decoding device: - Obtaining (S26) data related to at least one set of multiple neural networks; - Obtaining (S33) the decoded neural networks from the obtained data, the neural networks in the at least one set being available for the decoding device; and - Decoding (S34) the bitstream using the obtained neural networks.

18. The decoding method according to claim 17, wherein, The bitstream (B) includes data (BI) for identifying the at least one set of neural networks.