Encoding and decoding of audio and / or video data
Patent Information
- Application Number
- EP2023738047
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-08
- Filing Date
- 2023-07-07
- Publication Date
- 2025-05-14
AI Technical Summary
Conventional video encoders and decoders are not compatible due to differences in encoding and decoding processes, as neural networks operate differently from traditional methods, leading to issues with syntax elements and decoding capabilities.
An AI-based encoder transmits decoding characteristics within the compressed audio/video signal, allowing decoders to identify and verify their compatibility, enabling decoding even when using AI-based encoding methods.
Ensures compatibility between AI-based encoders and decoders, allowing for efficient decoding of AI-encoded audio/video data by embedding decoding configuration information within the signal, facilitating the use of AI-based decoding processes.
Smart Images

Figure 1.1
Abstract
Description
[0001] DESCRIPTION
[0002] Title: Coding and decoding of audio and / or video data
[0003] Field of invention
[0004] The present invention relates generally to the field of audio and / or video data processing, and in particular to the coding and decoding of digital images and digital image sequences.
[0005] The coding / decoding of digital images applies in particular to images from at least one video sequence comprising:
[0006] - images from the same camera and following one another in time (2D coding / decoding),
[0007] - images from different cameras oriented according to different views (3D type coding / decoding),
[0008] - corresponding texture and depth components (3D type coding / decoding),
[0009] - etc...
[0010] The present invention applies in a similar manner to the coding / decoding of 2D or 3D type images.
[0011] The invention may in particular, but not exclusively, apply to the video coding implemented in current video coders AVC (“Advanced Video Coding” in English), HEVC (“High Efficiency Video Coding” in English), VVC (“Versatile Video Coding” in English), and their extensions (MVC (“Multiview Video Coding” in English), 3D-AVC, MV-HEVC, 3D-HEVC, etc.), and to the corresponding decoding.
[0012] Prior art
[0013] Currently, artificial intelligence approaches, particularly neural ones, are tending to develop for the compression of still image, video or audio data and many studies report spectacular results on their capacity to represent a compressed data signal efficiently.
[0014] For example, in the context of image processing, such neural approaches are now capable of addressing image compression, no longer only as a method aimed at replacing or improving a step of the classic compression approach (such as prediction or filtering), but by completely replacing the encoder and the decoder, notably using "auto-encoders". Such an auto-encoder is for example described in the document: Théo Ladune, Pierrick Philippe, "AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC", 17 Feb 2022. Such an auto-encoder comprises an encoding neural network, which takes the video frames as input and provides latent variables as output, which represent the signal representative of these compressed frames.These latent variables are then quantified and then coded by entropic coding, for example by Huffman coding or CABAC coding (“Context-adaptive binary arithmetic coding” in English), to produce the signal representative of these compressed images.
[0015] This signal is then transmitted to a decoder, which performs entropy decoding and then dequantization of the data in this signal, the entropy decoding and dequantization corresponding respectively to the entropy coding and quantization implemented in the auto-encoder. At the end of this decoding, decoded latent variables are produced. These decoded latent variables are then supplied to a decoding neural network corresponding to the encoding neural network, the decoding neural network providing the decoded images of the video as output.
[0016] It is also known that in conventional video coders, for example WC coders, there are different tasks for encoding an image or a sequence of images, a task being associated with different constraints at the coder level, in particular in terms of coding tools used, computing power, data storage, image resolution, etc. For this purpose, the coder transmits to the decoder information representative of these constraints in the form of syntax elements. A VVC decoder will be configured to know how to interpret such syntax elements and therefore be able to decode the signal received from the coder. A non-WC compliant decoder will not know how to interpret such syntax elements, and therefore will not be able to encode the signal received from the decoder.Given that encoding neural networks operate in a completely different way from conventional video coders, and therefore respond to different constraints, the syntax elements used conventionally, particularly in the WC standard, are not suitable for these encoding neural networks. For example, some WC syntax elements are defined on value ranges that are limited, whereas encoding neural networks require encoding indications of a broader nature. Subject and summary of the invention.
[0017] One of the aims of the invention is to remedy the drawbacks of the aforementioned state of the art by proposing:
[0018] - an audio and / or video encoder based on an artificial intelligence approach which is configured to transmit in the compressed audio and / or video signal one or more indications of decoding characteristics that an audio and / or video decoder must support to decode the compressed video signal,
[0019] - an audio and / or video decoder configured to receive this compressed audio and / or video signal and read this or these decoding characteristics indication(s), so as to identify in a very simple manner whether or not it is capable of decoding the compressed audio and / or video signal.
[0020] Thus, the invention advantageously allows an audio or video decoder, whether standardized (of the AVC, HEVC, WC, AAC, MPEG-H 3D Audio type, etc.) or whether it implements decoding based on artificial intelligence, to be compatible in reading mode with the information read in the coded audio and / or video data signal, even when the audio and / or video data have been encoded using a coder based on an artificial intelligence approach.
[0021] To this end, an object of the present invention relates to a method for coding audio and / or video data, implemented by a coding device configured to implement at least one step of coding the audio and / or video data using an artificial coding neural network, said coding method comprising the following:
[0022] - encode audio and / or video data,
[0023] - generate a data signal which contains the encoded audio and / or video data,
[0024] - encoding information representative of a decoding configuration to be possessed by a decoding device to decode said encoded data,
[0025] - insert the coded information into the data signal.
[0026] The invention advantageously allows an audio and / or video coder of which at least one coding step is implemented using an artificial coding neural network, and therefore requiring one or more configuration elements dedicated to the implementation of this particular coding step, to code information representative of this or these corresponding coding and therefore decoding configuration elements, with a view to transmitting this information to an audio and / or video data decoder to inform it of the decoding capabilities that this decoder must possess in order to be able to decode the audio and / or video data.
[0027] An object of the present invention also relates to a method for decoding coded audio and / or video data, implemented by a decoding device, comprising the following:
[0028] - receive a coded audio and / or video data signal,
[0029] - decoding, from the signal, information representative of a decoding configuration according to which at least one step of decoding the audio and / or video data is implemented using a decoding artificial neural network,
[0030] - check whether the decoding device has the decoding configuration corresponding to the decoded information,
[0031] - decode or not the signal, depending on the result of the verification.
[0032] The invention advantageously allows an audio and / or video decoder to identify, in the coded audio and / or video data signal that it receives, the information relating to the decoding configuration that it must have to decode the signal.
[0033] According to a particular embodiment of the aforementioned coding or decoding method, the decoding configuration belongs to:
[0034] - to a first category corresponding to at least one particular physical characteristic of hardware or software to be supported by the decoding device in order to be able to decode the signal, and / or
[0035] - to a second category corresponding to a particular characteristic of said signal, and / or
[0036] - to a third category corresponding to at least one particular processing functionality to be applied by the decoding device in order to be able to decode the signal.
[0037] Such an embodiment advantageously makes it possible, in the case where a large variety of coding / decoding configurations is used, to group these configurations by category so as to reduce the number of information items to be coded. According to a particular embodiment of the aforementioned coding or decoding method, the information representative of a decoding configuration, which is respectively coded or decoded, is associated with at least one category from among the first, second or third category.
[0038] Such an embodiment advantageously makes it possible to encode / decode in a structured manner information representative of decoding configurations of different types. Furthermore, when there are several pieces of information representative of a decoding configuration, associated with different decoding parameters or characteristics of the same category, such an embodiment makes it possible to generate more compact signaling of this information, since a single syntax element or indicator is signaled for an entire category of decoding parameters or characteristics, instead of each decoding parameter or characteristic being indicated individually in the signal.
[0039] According to a particular embodiment of the aforementioned coding or decoding method, the decoding configuration belongs to a set comprising:
[0040] - a maximum size of a data storage memory;
[0041] - a minimum number of operations per second;
[0042] - a minimal flow of latent variables;
[0043] - a particular type of electronic circuit;
[0044] - a level of precision of mathematical representation of at least one operating parameter of the decoding artificial neural network;
[0045] - the activation or not of at least one reference decoding step;
[0046] - at least one particular mathematical operator or a list of particular mathematical operators;
[0047] - a particular mathematical function;
[0048] - a number of statistical sources of entropy decoding.
[0049] According to a particular embodiment of the aforementioned coding or decoding method, the information representative of a decoding configuration is contained in a set of predefined video parameters of the coding or decoding method respectively, or, when the video data are representative of a sequence of images, in a set of parameters associated with said sequence. Such an embodiment advantageously makes it possible to use the coding syntax of existing or standardized coders to code the information representative of a configuration element. In the case, for example, of an AVC, HEVC or WC coder, the set of predefined video parameters of the coding method is, for example, the VPS (“Video Parameter Set”) and the set of parameters associated with said sequence is the SPS (“Sequence Parameter Set”).In another example, the information representative of a decoding configuration is associated with a sub-image, in particular a tile or a slice as defined for example in the HEVC standard.
[0050] According to a particular embodiment of the aforementioned coding or decoding method, the information representative of a decoding configuration comprises:
[0051] - a first value which is associated with a first configuration element of said decoding configuration, and
[0052] - a second value which is associated with the first configuration element and a second configuration element of said decoding configuration.
[0053] Such an embodiment advantageously makes it possible, when the configuration comprises several configuration elements corresponding to different decoding capacities to be supported by the decoding device, to indicate them by nesting in the signal transmitted to the decoder.
[0054] The various embodiments or features mentioned above may be added independently or in combination with each other to the coding or decoding method defined above.
[0055] An object of the present invention also relates to an audio and / or video data signal, said signal comprising:
[0056] - audio and / or video data that has been encoded by a coding device configured to implement at least one step of encoding the audio and / or video data using an artificial coding neural network,
[0057] - coded information which is representative of a decoding configuration to be possessed by a decoding device for decoding said coded audio and / or video data.
[0058] An object of the present invention also relates to a device for coding audio and / or video data, configured to implement at least one step of coding the audio and / or video data using an artificial coding neural network, said coding device being configured to implement the following:
[0059] - encode audio and / or video data,
[0060] - generating a data signal which contains the coded audio and / or video data, - encoding information representative of a decoding configuration to be possessed by a decoding device to decode said coded data,
[0061] - insert said coded information into said data signal.
[0062] Such a coding device is in particular capable of implementing the aforementioned coding method.
[0063] An object of the present invention also relates to a device for decoding coded audio and / or video data, configured to implement the following:
[0064] - receive a coded audio and / or video data signal,
[0065] - decoding, from said signal, information representative of a decoding configuration according to which at least one step of decoding the audio and / or video data is implemented using a decoding artificial neural network,
[0066] - check whether the decoding device has the decoding configuration corresponding to the decoded information,
[0067] - decode or not the said signal, depending on the result of the verification.
[0068] Such a decoding device is in particular capable of implementing the aforementioned decoding method.
[0069] The invention also relates to a computer program comprising instructions for implementing the coding or decoding method according to the invention, according to any one of the particular embodiments described previously, when said program is executed by a processor.
[0070] Such instructions may be stored permanently in a non-transitory memory medium of the coding device implementing the aforementioned coding method or of the decoding device implementing the aforementioned decoding method.
[0071] This program may use any programming language, and may be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0072] The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions of a computer program as mentioned above.
[0073] The recording medium may be any entity or device capable of storing the program. For example, the medium may include a storage medium, such as a ROM, for example a CD-ROM, a DVD-ROM, a synthetic DNA (deoxyribonucleic acid), etc. or a microelectronic circuit ROM, or a magnetic recording medium, for example a USB key or a hard disk.
[0074] Furthermore, the recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means. The program according to the invention may in particular be downloaded from a network such as the Internet.
[0075] Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the encoding or decoding method according to the invention.
[0076] Brief description of the drawings
[0077] Other characteristics and advantages will appear on reading particular embodiments of the invention, given as illustrative and non-limiting examples, and the appended drawings, among which:
[0078] - [Fig. 1] represents the main steps of a method for coding audio and / or video data, in a particular embodiment of the invention,
[0079] - [Fig. 2A] represents a mode of transport of the data coded in accordance with the coding method of [Fig. 1], in a particular embodiment of the invention,
[0080] - [Fig. 2B] represents a mode of transport of the data coded in accordance with the coding method of [Fig. 1], in another particular embodiment of the invention,
[0081] - [Fig. 3] represents a coding device implementing the coding method of Figure 1, in a particular embodiment of the invention,
[0082] - [Fig. 4] represents the main steps of a method for decoding audio and / or video data, in a particular embodiment of the invention,
[0083] - [Fig. 5] represents a decoding device implementing the decoding method of Figure 4, in a particular embodiment of the invention.
[0084] Detailed description of different embodiments of the invention
[0085] Coding of audio and / or video data
[0086] A method for encoding audio and / or video data representative of an image or a sequence of images of 2D or 3D type is described below. Such a coding method is capable of being implemented in any type of video encoder or decoder, for example compliant with the JPEG, AVC, HEVC, WC standard and their extensions (MVC, 3D-AVC, MV-HEVC, 3D-HEVC, etc.), or other, for example video encoders based on neural networks. Such a coding method is also capable of being implemented in any type of audio encoder or decoder, for example compliant with the MP3, AAC (“Advanced Audio Coding”), MPEG-H 3D Audio standard, or other, for example audio encoders based on neural networks.
[0087] With reference to Figure 1, the coding method according to the invention comprises the following:
[0088] In C1, common audio and / or video data is selected.
[0089] Such audio data is in the form of a set of current samples B c which can be:
[0090] - a one-dimensional temporal audio signal;
[0091] - part of such a signal;
[0092] - a multidimensional temporal audio signal (stereo or greater than two dimensions).
[0093] Such video data is in the form of a current set of pixels B c which can be:
[0094] - an original current image;
[0095] - a part or area of the original current image;
[0096] - a block of the current image resulting from a partitioning of this image in accordance with what is practiced in standardized coders of the AVC, HEVC or WC type.
[0097] In C2, a current prediction data set BPc, the data being for example pixels, is calculated according to an Intra, Inter, IBC (“Intra Block Copy” in English), SKIP, etc. prediction well known to those skilled in the art.
[0098] In audio coding, the data in the current prediction dataset BPc are samples.
[0099] In C3, a signal BEc is calculated, representing the difference between the current set of pixels Bc and the current set of prediction pixels BPc obtained in C2.
[0100] In C4, in the case where this BEc signal is the one which optimizes the coding with respect to a classic coding performance criterion, such as for example the minimization of the bit rate / distortion cost or the choice of the best efficiency / complexity compromise, which are criteria well known to those skilled in the art, the BEc signal is quantified and coded.
[0101] At the end of this operation, a quantified and entropically coded deviation signal BE c cod is obtained. Such entropic coding is for example by Huffman coding or CABAC coding. In the preferred embodiment the entropic coding is of the CABAC type.
[0102] In C5, an F signal or stream is generated to contain DAT data of the quantized and BE-encoded deviation signal c cod . In a manner known per se, the signal F is capable of being transmitted to a decoding device or decoder which will be described later in the description.
[0103] According to the invention, at least one of the operations C1 to C5 is implemented using a computing device based on artificial intelligence, referenced DCIA_C, which is configured to automate said at least one coding operation to make it more efficient and more adaptive. Such a computing device comprises for example a neural network or several neural networks, a support vector machine, a reasoning engine, an expert system, a fuzzy logic system, etc.
[0104] In the preferred embodiment, the computing device is an artificial coding neural network, such as for example a convolutional neural network or CNN (“Convolutional Neural Network” in English), a multi-layer perceptron, an LSTM (“Long Short Term Memory” in English), etc. Such a neural network is defined by a structure comprising for example a plurality of layers of artificial neurons and / or by a set of weights associated respectively with the artificial neurons of this network.
[0105] More particularly, in the preferred embodiment, the neural network used is a convolutional neural network. In a particular embodiment, the latter calculates in C3 the deviation signal BEc or codes the current set of pixels B ctogether with the set of prediction pixels BPc generated in C2, thus performing operations C3 and C4. Such a neural network is for example of the type described in the document: Ladune “Optical Flow and Mode Selection for Learning-based Video Coding”, IEEE MMSP 2020. In another particular embodiment, the prediction operation C2 is also implemented using a convolutional neural network and not using a conventional prediction device, for example of the VVC or CELP (“Code-Excited Linear Prediction” type in the case of audio samples. Such a neural network is described in particular in the document Théo Ladune, Pierrick Philippe, “AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC”, 17 Feb 2022.
[0106] The use of one or more computing devices based on artificial intelligence to implement a method for encoding audio and / or video data requires the encoding device that implements the encoding method to have a specific hardware or software configuration. Such a configuration is correspondingly required in a decoding device, so that the latter is capable of decoding in real time the encoded audio and / or video data signal, received from the encoding device comprising this or these computing devices. This decoding configuration belongs to one or more categories including for example:
[0107] - a first category corresponding to at least one particular physical characteristic of hardware or software to be supported by the decoding device in order to be able to decode the signal F, and / or
[0108] - a second category corresponding to one or more particular characteristics of the data signal F, when at least one coding step is implemented by the DCIA_C calculation device,
[0109] - a third category corresponding to at least one particular processing functionality to be applied by the decoding device in order to be able to decode the signal F.
[0110] As non-exhaustive examples, the first category of decoding characteristics includes:
[0111] - an electronic circuit of a specific type used by the computing device, for example a GPU (“Graphical Processing Unit” in English), a TPU (“Tensor Processing Unit” in English), a DSP (“Digital Signal Processor” in English), or any other type of suitable electronic circuit;
[0112] - a data memory size, for example a buffer memory, which the decoding device must have in order to store all the data necessary for the neural decoding of the signal F; - a minimum number of operations per second, for example the TOPS (“Tera Operations Per Second” in English) which the decoding device must be capable of implementing;
[0113] - a level of precision of representation of the parameters of the DCIA_C calculation device that the decoding device must respect in order to be able to decode the signal F, such parameters including for example the weights of the decoding artificial neural network(s) used by the decoding device, the parameters of the activation functions applied at the output of the artificial neurons, etc.;
[0114] - a number of entropic decoding sources of latent variables that the decoding device must support in order to be able to decode the signal F,
[0115] - etc... .
[0116] As non-exhaustive examples, the second category of decoding characteristics includes:
[0117] - a minimum number of latent variables to be processed per unit of time by the decoding device, when one or more neural networks are used, so that the decoding device maintains its real-time decoding capabilities;
[0118] - a maximum number of latent variables per unit of time, allowing the decoding device to check whether it has the capacity to decode this signal;
[0119] - a minimum number of bits representative of syntax elements aimed at reconstructing latent variables per unit of time (entropic rate), so that the decoding device maintains its real-time decoding capabilities;
[0120] - a maximum number of bits representative of syntax elements aimed at reconstructing latent variables per unit of time (entropic rate), so that the decoding device can check whether it has the capacity to decode this signal;
[0121] - etc.
[0122] The decoding characteristics are transmitted to the decoding device through indicators which explicitly give the values of these characteristics, or through more global indicators which indicate at once the values of several characteristics, thanks to predetermined association tables, such as for example the table T10 or the table T 11 described later.
[0123] The advantage of transmitting a flow rate of latent variables (expressed as the number of latent variables per unit of time or as the coded flow rate of these latent variables per unit of time) is to allow the complexity of the data requiring processing by a neural network to be adjusted with the capabilities of the decoding device to perform processing by a neural network. Indeed, a coded signal representative of, for example, an image or a video, can contain both coded data whose decoding implements conventional processing, typically carried out using a CPU (Central Processing Unit) type processor, and data whose decoding implements neural processing, typically carried out using a specific processor allowing a very large quantity of small calculations of the same nature to be carried out in parallel (GPU or TPU processor).By specifying the throughput characteristics of the latent variables, it is easier to adjust the sub-part of the signal that concerns the neural processing of the data and the GPU or TPU-type capabilities of the decoding device.
[0124] As non-exhaustive examples, the third category of decoding characteristics includes:
[0125] - a set of predefined and standardized logical and / or mathematical operators and / or a list of logical and / or mathematical operators to be supported by the decoding device in order to be able to decode the signal;
[0126] - a list of activation functions to be applied by the decoding device at the output of the neural network(s) to be able to decode the signal;
[0127] - a capacity to reproduce the decoding results faithful to a reference (inter-platform reproducibility) that the decoding device must have to be able to decode the signal;
[0128] - etc... .
[0129] All of these decoding characteristics are considered both in the context of coding / decoding audio data and in the context of coding / decoding video data.
[0130] According to the invention, the coding method comprises a step C6 of coding one or more ICD information items representative of a decoding configuration to be possessed by the decoding device to decode said DAT coded data, such a decoding configuration belonging to the first and / or second and / or third category mentioned above.
[0131] At the end of C6 coding, one or more ICD information cod are obtained. The ICD information(s) codare then written in C7 either in the data signal F, or in a signal F' associated with the signal F.
[0132] Referring to Figure 2A, the ICD information(s) cod are written in C7 in the signal F, in a data packet independently identifiable and decodable with respect to the DAT coded audio and / or video data, packet comprising other information which is necessary for the decoding of the DAT coded data, such as predefined audio and / or video parameters of the coding method or, when the DAT coded data is representative of a sequence of images, parameters associated with this sequence of images, in the manner of the VPS or SPS syntax respectively, as implemented for example in the WC standard. According to another example, the ICD information(s) cod are parameters associated with a sub-image, in particular a tile or a slice as defined for example in the HEVC standard.
[0133] Referring to Figure 2B, the ICD information(s) cod are written in C7 in an optional packet F', the decoding of which is not necessary for decoding DAT data, in the same way as writing information in an SEI message ("Supplemental Enhancement Information" in English) conforming for example to the WC standard.
[0134] In C8, the F signal containing the DAT coded data and the ICD coded information(s) cod , alternately the signal F containing the coded data DAT and the message F' containing the coded information ICD cod , are stored or transmitted to a decoding device which will be described later in the description.
[0135] In a preferred embodiment, each decoding configuration or characteristic represented by ICD information cod is reported individually. Different examples of ICD information coding codare represented below in corresponding syntax tables.
[0136] In the case of a particular hardware, such as for example an electronic circuit or processor of a specific type, to be supported by the decoding device, a proc dc indicator as represented in the syntax table T1 below, takes for example eight values ranging from 0 to 7 to indicate to the decoding device the type of processor to support to decode the signal F.
[0137] Regarding the size of the data memory that the decoding device must have, a buffer_size_idc flag specifies the size of this memory in bytes, or alternatively can take a predetermined number of values that are associated with predefined size limits, as in the following syntax table T2. The buffer_size_idc flag here, for example, takes eight values ranging from 0 to 7.
[0138] Regarding the minimum number of operations per second to be supported by the decoding device, an ops_idc flag specifies the number of operations per second, or alternatively can take a predetermined number of values which are associated with predefined limits, as in the following T3 syntax table, where the ops_idc flag takes for example four values 0 to 3:
[0139] Regarding the minimum number of latent variables to be processed per unit of time by the decoding device, a latent_rate_idc flag specifies this number in number of variables per second, i.e. in bitrate, or alternatively can take a predetermined number of values which are associated with predefined bitrate limits, as in the following T4 syntax table, where the latent_rate_idc flag takes for example seven values 0 to 6:
[0140] Regarding the level of precision of representation of the parameters of the DCIA_C computing device, and in the case where this device is a neural network, a precision dc indicator specifies the required precision of the weights of this network using the following predetermined association table T5, where the precisionjdc indicator takes for example eight values 0 to 7:
[0141] Regarding the ability to reproduce the results of decoding faithful to a reference (inter-platform reproducibility) that the decoding device must have, in table T6 below, a repro lag indicator is set:
[0142] - either to a first value, for example 0, to indicate that it is not necessary for the decoding device to reproduce an identical reference decoding,
[0143] - either to a second value, for example 1, to indicate that it is necessary for the decoding device to reproduce an identical reference decoding.
[0144] Regarding the number of latent variable entropy decoding sources that the decoding device must support, a sourcesjdc indicator specifies the number of sources, or alternatively can take a predetermined number of values that are associated with predefined limits of numbers of sources, as in the following table T7, where the sourcesjdc indicator takes for example four values 0 to 3:
[0145] As the set or list of predefined and standardized logical and / or mathematical operators to be supported by the decoding device, an operators dc indicator specifies the logical and / or mathematical operators to be supported on the decoder side. According to a preferred embodiment shown in the following syntax table T8, the operators dc indicator comprises five values 0 to 4, the values 0 to 3 being constructed in a nested manner such that:
[0146] - the value 0 is associated with a set of basic mathematical operators “+, -, x, / ”,
[0147] - the value 1 is associated with the set of basic mathematical operators “+, -, x, / ” and with the set of mathematical operators “x y , exp(), sqrt() »,
[0148] - the value 2 is associated with the set of basic mathematical operators “+, -, x, / ”, with the set of mathematical operators “x y , exp(), sqrt() », and to the operator
[0149] “N!”, where N is a natural number,
[0150] - the value 3 is associated with the set of basic mathematical operators “+, -, x, / ”, with the set of mathematical operators “x y , exp(), sqrt() », to the operator
[0151] “N!”, and to the set of mathematical operators “sin(), cos(), tan()”.
[0152] Of course, this way of signaling operators is not exhaustive. In other embodiments, each operator or a list of operators may be signaled individually, which generates a higher signaling cost. Regarding the list of activation functions to be applied by the decoding device, an activationsjdc flag specifies the activation functions to be supported by the decoding device. According to a preferred embodiment shown in the following T9 syntax table, the activationsjdc flag comprises three values 0 to 2, the values 0 and 1 being constructed in a nested manner such that:
[0153] - the value 0 is associated with the following list of activation functions:
[0154] F(x) = x,
[0155] G(x) = 0 if x<0, 1 otherwise
[0156] H(x) = 0 if x<0, x otherwise
[0157] - the value 1 is associated with this list of activation functions, as well as with the following list of activation functions: l(x) = 1 / (1 +exp(-x)) J(x) = tan -1 (x)
[0158] In a particular embodiment, a single leveljdc indicator specifies several decoding characteristics to be supported by the decoding device at the same time. For this purpose, a correspondence table, referenced T10 below, which is predefined on both the encoding device and decoding device sides, is generated. Table T10 matches at least one particular value of the level dc indicator with a particular latent variable rate, a particular data memory size, etc. In the example shown, the leveljdc indicator has six values 0 to 5.
[0159] T10
[0160] In a particular embodiment, a series of indicators, cat1_level_idc, cat2_level_idc, cat3_level_idc, instead of a single indicator level dc, specifies one or more decoding characteristics according to the first, second or third category to which this or these decoding characteristics belong. For this purpose, for each indicator cat1_level_idc, cat2_level_idc, cat3_level_idc, a correspondence table, predefined both on the coding device side and on the decoding device side, is generated. Such a table is shown below and bears the reference T11.
[0161] Regarding the first category, table T11 below maps at least one particular value of the cat1_level_idc flag to a particular processor type, a particular data memory size, etc. In the example shown, the cat1_level_idc flag has six values 0 to 5.
[0162] Regarding the second category, table T12 below maps at least one particular value of the cat2_level_idc indicator to a particular latent variable rate. In the example shown, the cat2_level_idc indicator has five values, 0 to 4. Regarding the third category, table T13 below maps at least one particular value of the cat3_level_idc indicator to a capability or not of reproducing decoding results, a list of particular mathematical operators, etc. In the example shown, the cat3_level_idc indicator has four values 0 to 3. A COD coder shown in schematic form is now described with reference to FIG. 3, the COD coder being adapted to implement the coding method illustrated in FIG. 1, in a particular embodiment of the invention. According to this particular embodiment, the actions executed by the coding method are implemented by computer program instructions. For this, the COD coding device has the conventional architecture of a computer and notably comprises a memory MEM_C, a processing unit UT_C, equipped for example with a processor PROC_C, and controlled by the computer program PG_C stored in memory MEM_C. The computer program PG_C comprises instructions for implementing the actions of the coding method as described above, when the program is executed by the processor PROC_C.
[0163] At initialization, the code instructions of the computer program PG_C are for example loaded into a RAM memory (not shown) before being executed by the processor PROC_C. The processor PROC_C of the processing unit UT_C implements in particular the actions of the coding method described above, according to the instructions of the computer program PG_C.
[0164] The COD encoder receives as input E_C a set of pixels or current samples B c and delivers at output S_C the transport stream F which is transmitted to a decoder using a suitable communication interface (not shown).
[0165] The COD encoder comprises a prediction device PRED configured to implement the aforementioned prediction step C2. As already explained above in the description, this prediction device:
[0166] - can be conventional and configured according to e.g. HEVC, VVC, CELP, etc. standard;
[0167] - can be a neural network, for example of the type described in the aforementioned document Théo Ladune, Pierrick Philippe, “AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC”, 17 Feb 2022,
[0168] - etc...
[0169] The COD encoder also includes the DCIA_C calculation device based on artificial intelligence which is for example of the type described in the aforementioned document: Ladune “Optical Flow and Mode Selection for Learning-based Video Coding”, IEEE MMSP 2020.
[0170] The COD coder also comprises a CICD information coding device configured to implement the aforementioned step C6 of coding one or more ICD information representative of a decoding configuration to be possessed by the decoding device to decode the signal F.
[0171] The COD encoder also comprises an IICD device configured to implement the aforementioned step C7 of writing ICD information. codobtained by the CICD device either in the data signal F or in the message F' associated with the signal F.
[0172] The COD coder also comprises a storage memory MS_C configured to store the syntax tables T1 to T12. Alternatively, this storage memory MS_C is not contained in the COD coder but is accessible by the latter by any suitable means, via a communication network for example. In one embodiment, the data signal F and the optional message F' can also be stored in the storage memory MS_C or in an additional storage memory (not shown).
[0173] Decoding of encoded audio and / or video data
[0174] A method for decoding a coded audio and / or video data signal relating to a 2D or 3D image or image sequence is described below. Such a decoding method is capable of being implemented in any type of video decoder, for example compliant with the JPEG, AVC, HEVC, WC standard and their extensions (MVC, 3D-AVC, MV-HEVC, 3D-HEVC, etc.), or other, for example video decoders based on neural networks. Such a decoding method is also capable of being implemented in any type of audio decoder, for example compliant with the MP3, AAC, MPEG-H 3D Audio standard, or other, for example audio decoders based on neural networks.
[0175] Referring to Figure 4, the decoding method according to the invention comprises the following.
[0176] In D1, the aforementioned data signal F, which contains the coded audio and / or video data DAT and the coded information ICD, is received by a decoding device DEC shown in Figure 5.cod representative of a particular decoding configuration to be possessed by the decoding device DEC to decode the signal F. Alternatively, in D1, are received by the decoding device DEC:
[0177] - the data signal F containing the DAT encoded audio and / or video data,
[0178] - the F' message containing the ICD coded information cod representative of a particular decoding configuration to be possessed by the decoding device DEC to decode the signal F.
[0179] In D2, the received F data signal is extracted from the DAT-encoded audio and / or video data and the ICD-encoded information. cod Alternatively, in D2, an extraction is carried out from the received data signal F of the coded audio and / or video data DAT and an extraction is carried out from the received message F' of the coded information ICD. cod .
[0180] According to the invention, the ICD coded information(s) codare decoded into D3. At the end of this operation, the ICD information is reconstructed, which allows the decoder to identify the decoding characteristic(s) required for it to be able to decode the DAT-encoded audio and / or video data. To this end, in a particular embodiment of the invention, the value of one or more indicators, such as for example the indicators proc dc, buffer_size_idc, ops_idc, latent_rate_idc, precision dc, repro_flag, sources dc, operators dc, activations dc, level dc, cat1_level_idc, cat2_level_idc, cat3_level_idc, is read, then matched with its associated decoding characteristic or its associated decoding characteristics in the case in particular of the indicators cat1_level_idc and cat3_level_idc. Such matching is implemented using the aforementioned correspondence tables T1 to T13 which are made accessible by the decoding device DEC.
[0181] In D4, the decoding device DEC compares the decoding characteristic(s) identified in D3 with the decoding characteristic(s) specific to it. Such a comparison is made possible by the fact that the decoding device DEC is able to access the technical characteristics of the platform on which it operates (whether software or hardware or hybrid) and the performance it is capable of achieving.
[0182] If, at the end of the comparison D4, the decoding device DEC does not have the decoding characteristic(s) identified in D3, the decoding of the DAT-encoded audio and / or video data is not implemented. The decoding is therefore abandoned (ABD).
[0183] If at the end of the comparison D4, the decoding device DEC has the decoding characteristic(s) identified in D3, the decoding of the DAT-coded audio and / or video data is implemented using a DCIA_D calculation device based on artificial intelligence, the DCIA_D calculation device implementing a decoding corresponding to the coding implemented by the DCIA_C calculation device. For this purpose, a dequantization and an entropy decoding of the DAT-coded audio and / or video data are carried out in D5. Such an entropy decoding is for example a Huffman decoding or a CABAC decoding. In the preferred embodiment, the entropy decoding is of the CABAC type. At the end of this operation, a decoded deviation signal BE c dec is obtained.
[0184] In D6, a prediction is implemented, generating the current prediction data set BPc, this data being for example pixels here but can also be samples of an audio signal.
[0185] Steps D5 and D6 can be implemented in any order or simultaneously.
[0186] In D7, a reconstructed current pixel set BDc is calculated by combining the decoded deviation signal BE c dec obtained in D5 to the set of pixels or BPc prediction samples obtained in D6.
[0187] In a manner known per se, the reconstructed current set of pixels BDc may possibly undergo filtering by a loop filter of the reconstructed signal which is well known to those skilled in the art.
[0188] Of course, in the case where the deviation signal BE cwhich was calculated during the aforementioned coding process is zero, which may be the case for the SKIP coding mode, step D2 of extracting the deviation signal BEc and step D5 of dequantization and entropy decoding are not implemented.
[0189] A DEC decoder represented in schematic form will now be described with reference to FIG. 5, the DEC decoder being adapted to implement the decoding method illustrated in FIG. 4, in a particular embodiment of the invention.
[0190] According to this particular embodiment, the actions executed by the decoding method are implemented by computer program instructions. For this, the decoding device DEC has the conventional architecture of a computer and notably comprises a memory MEM_D, a processing unit UT_D, equipped for example with a processor PROC_D, and controlled by the computer program PG_D stored in memory MEM_D. The computer program PG_D comprises instructions for implementing the actions of the decoding method as described above, when the program is executed by the processor PROC_D. The decoder DEC receives at input E_D the data signal F, possibly the message F', transmitted by the coder COD of FIG. 3, and delivers at output S_D the current decoded set of pixels or samples BDc.
[0191] The DEC decoder also comprises a DICD information decoding device, configured to implement the aforementioned step D3 of decoding one or more pieces of coded information ICD cod representative of a decoding configuration to be possessed by the DEC decoder to decode the signal F.
[0192] The DEC decoder also comprises a COMP device configured to compare in D4 the reconstructed ICD information(s) representative of a decoding configuration to be possessed by the DEC decoder to decode the signal F with the DEC decoder's own decoding characteristics.
[0193] Correspondingly to the COD coder of Figure 3, the DEC decoder also comprises a storage memory MS_D configured to store the aforementioned syntax tables T1 to T13. Alternatively, this storage memory MS_D is not contained in the DEC decoder but is accessible by the latter by any suitable means, via a communication network for example. In one embodiment, the data signal F and the optional message F' can also be stored in the storage memory MS_D or in an additional storage memory (not shown).
[0194] The decoder DEC comprises a prediction device PRED_D configured to implement the aforementioned prediction step D6. As already explained above in the description, this prediction device:
[0195] - can be conventional and configured according to e.g. HEVC, VVC, CELP, etc. standard;
[0196] - can be a neural network, for example of the type described in the aforementioned document Théo Ladune, Pierrick Philippe, “AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC”, 17 Feb 2022,
[0197] - etc...
[0198] The DEC decoder may also include an artificial intelligence-based computing device, referenced DCIA_D, which is configured to automate at least one decoding operation to make it more efficient and more adaptive. Such a computing device includes, for example, one or more artificial neural decoding networks, a support vector machine, a reasoning engine, an expert system, a fuzzy logic system, etc.
[0199] In the preferred embodiment, the DCIA_D computing device is a neural network, such as for example a convolutional neural network or CNN, a multi-layer perceptron, an LSTM, etc. More particularly, in the preferred embodiment, the neural network used is a convolutional neural network. Such a neural network is defined by a structure comprising for example a plurality of layers of artificial neurons and / or by a set of weights associated respectively with the artificial neurons of this network.
[0200] In a particular embodiment, the DCIA_D neural network combines in D7 the decoded deviation signal BE c decobtained in D5 to the set of pixels or prediction samples BPc generated in D6. Such a neural network is for example of the type described in the document: Ladune “Optical Flow and Mode Selection for Learningbased Video Coding”, IEEE MMSP 2020. In another particular embodiment, the prediction operation D6 is also implemented using a convolutional neural network and not using a conventional prediction device of the WC type for example. Such a neural network is described in particular in the document Théo Ladune, Pierrick Philippe, “AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC”, 17 Feb 2022.
[0201] It goes without saying that the embodiments described above have been given for purely indicative purposes and are in no way limiting, and that numerous modifications can easily be made by those skilled in the art without departing from the scope of the invention.
Claims
CLAIMS
1. A method of coding audio and / or video data, implemented by a coding device configured to implement at least one step of coding the audio and / or video data using a coding artificial neural network (DCIA_C), said coding method comprising the following: - encode (C2-C4) audio and / or video data, - generate (C5) a data signal which contains the encoded audio and / or video data, - encoding (C6) information representative of a decoding configuration relating to the use of said network, to be possessed by a decoding device for decoding said coded data, - insert (C7) the coded information into the data signal.
2. A method of decoding encoded audio and / or video data, implemented by a decoding device, comprising the following: - receive (D1) a coded audio and / or video data signal, - decoding (D3), from said signal, information representative of a decoding configuration according to which at least one step of decoding the audio and / or video data is implemented using a decoding artificial neural network (DCIA_D), - check (D4) whether the decoding device has the decoding configuration corresponding to the decoded information, - decode (D5-D7) or not said signal, depending on the result of the verification.
3. A method of coding according to claim 1 or decoding according to claim 2, wherein the decoding configuration belongs to: - to a first category corresponding to at least one particular physical characteristic of hardware or software to be supported by the decoding device in order to be able to decode the signal, and / or - to a second category corresponding to a particular characteristic of said signal, and / or - to a third category corresponding to at least one particular processing functionality to be applied by the decoding device in order to be able to decode the signal.
4. A method of coding or decoding according to claim 3, wherein the information representative of a decoding configuration, which is respectively coded or decoded, is associated with at least one category among the first, second or third category.
5. A method of coding according to any one of claims 1, 3 to 4 or of decoding according to any one of claims 2 to 4, wherein the decoding configuration belongs to a set comprising: - a maximum size of a data storage memory; - a minimum number of operations per second; - a minimal flow of latent variables; - a predetermined number of values associated respectively with predefined latent variable flow limits; - a particular type of electronic circuit; - a minimum number of bits representative of syntax elements aimed at reconstructing latent variables per unit of time; - a maximum number of bits representative of syntax elements aimed at reconstructing latent variables per unit of time; - a level of precision of mathematical representation of at least one operating parameter of the decoding artificial neural network; - the activation or not of at least one decoding identical to a reference decoding; - at least one particular mathematical operator or a list of particular mathematical operators; - a particular mathematical function; - a number of statistical sources of entropy decoding.
6. A method of coding according to any one of claims 1, 3 to 5 or of decoding according to any one of claims 2 to 5, in which the information representative of a decoding configuration is contained in a set of predefined video parameters of the coding or decoding method respectively, or, when the video data is representative of a sequence of images, in a set of parameters associated with said sequence.
7. A method of coding according to any one of claims 1, 3 to 6 or of decoding according to any one of claims 2 to 6, in which the information representative of a decoding configuration comprises: - a first value which is associated with a first configuration element of said decoding configuration, and - a second value which is associated with the first configuration element and with a second configuration element of said decoding configuration.
8. Device for encoding audio and / or video data, configured to implement at least one step of encoding the audio and / or video data using an artificial coding neural network, said encoding device being configured to implement the following: - encode audio and / or video data, - generate a data signal which contains the encoded audio and / or video data, - encode information representative of a decoding configuration relating to the use of said network, to be possessed by a decoding device for decoding said coded data, - inserting said coded information into said data signal.
9. Device for decoding coded audio and / or video data, configured to implement the following: - receive a coded audio and / or video data signal, - decoding, from said signal, information representative of a decoding configuration according to which at least one step of decoding the audio and / or video data is implemented using a decoding artificial neural network, - check whether the decoding device has the decoding configuration corresponding to the decoded information, - decoding or not decoding said signal, depending on the result of the verification.
10. Computer program comprising program code instructions for implementing the coding method according to any one of claims 1, 3 to 7, or the decoding method according to any one of claims 2 to 7, when executed on a computer.
11. Information medium readable by a computer, and comprising instructions of a computer program according to claim 10.