Encoding and decoding of audio and / or video data
The encoding method addresses compatibility issues by generating and inserting decoding configuration information, allowing AI-based encoders to work with standard decoders across various video encoding standards.
Patent Information
- Application Number
- JP2025500154
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-08
- Filing Date
- 2023-07-07
- Publication Date
- 2025-07-03
AI Technical Summary
Existing video encoders based on artificial intelligence approaches, such as autoencoders, do not conform to conventional video encoding standards like AVC, HEVC, and VVC, leading to compatibility issues with standard decoders.
An encoding method that includes generating and inserting information representing the decoding configuration required by a decoder, using an artificial neural network, to ensure compatibility with both AI-based and standard decoders.
Enables decoders to identify and decode AI-encoded audio and video data by providing necessary decoding configurations, ensuring compatibility across different encoding standards.
Smart Images

Figure 2025520955000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of audio and / or video data processing, and more particularly to the encoding and decoding of digital images and digital image sequences.
[0002] The encoding / decoding of digital images is particularly applied to images from at least one video sequence comprising the following. - Images that are temporally consecutive from the same camera (2D encoding / decoding), - Images from various cameras oriented in different fields of view (3D encoding / decoding), - Corresponding texture and depth components (3D encoding / decoding), - Etc.
[0003] The present invention is similarly applicable to the encoding / decoding of 2D or 3D images.
[0004] The present invention may be particularly applied to, but not limited to, current AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding) video encoders, and their extensions (MVC (Multiview Video Coding), 3D-AVC, MV-HEVC, 3D-HEVC, etc.), as well as the video encoding implemented therein and the corresponding decoding.
Background Art
[0005] Currently, artificial intelligence approaches, particularly neural approaches, tend to become more common for the compression of still images, videos, or audio data, and many studies have reported remarkable results regarding their ability to efficiently represent compressed data signals.
[0006] For example, in relation to image processing, such a neural approach is not aimed at replacing or improving the steps of classical compression approaches (such as prediction or filtering), but rather at handling image compression by completely replacing the encoder and decoder, particularly using an "autoencoder", as described, for example, in (Non-Patent Document 1). Such an autoencoder comprises an encoding neural network that takes in the images of the video at the input and supplies latent variables that represent signals representing these compressed images at the output. These latent variables are then quantized and then encoded by entropy encoding, such as Huffman encoding or CABAC (Context-Adaptive Binary Arithmetic Coding) encoding, to generate signals representing these compressed images.
[0007] This signal is then transmitted to the decoder, which performs entropy decoding of the data of this signal and then inverse quantization, and the entropy decoding and inverse quantization respectively correspond to the entropy encoding and quantization implemented in the autoencoder. At the end of this decoding, decoded latent variables are generated. These decoded latent variables are then supplied to a decoding neural network corresponding to the encoding neural network, and the decoding neural network supplies the decoded images of the video at the output.
[0008] In a conventional video encoder, for example, a VVC encoder, there are various tasks for encoding an image or an image sequence. One task is, in particular, known to be associated with various constraints in the encoder in terms of, for example, the encoding tools used, computational power, data storage, image resolution, etc. To achieve this purpose, the encoder transmits information representing these constraints in the form of syntax elements to the decoder. The VVC decoder is configured to know how to interpret such syntax elements and thus can decode the signal received from the encoder. A decoder not compliant with VVC does not know how to interpret such syntax elements and thus cannot decode the signal received from the decoder.
[0009] Considering that the encoding neural network operates in a completely different way from a conventional video encoder and thus satisfies different constraints, in particular, the syntax elements conventionally used in the VVC standard are not suitable for these encoding neural networks. For example, some VCC syntax elements are defined over a limited range of values, while the encoding neural network requires encoding instructions of a broader nature.
Prior Art Documents
Non-Patent Documents
[0010]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0011] One of the objects of the present invention is - An audio and / or video encoder based on an artificial intelligence approach, configured to transmit one or more indications of decoding features that an audio and / or video decoder should support for decoding a compressed video signal in a compressed audio and / or video signal. - An audio and / or video decoder configured to receive the compressed audio and / or video signal and read one or more of these indications of decoding features to identify whether the compressed audio and / or video signal can be decoded in a very simple way. It is to address the drawbacks of the prior art by proposing the above.
[0012] The present invention thus advantageously enables an audio or video decoder to be compatible with the information read from the encoded audio and / or video data signal in read mode, regardless of whether the audio and / or video data is encoded using an encoder based on an artificial intelligence approach or is standardized (such as types like AVC, HEVC, VVC, AAC, MPEG-H 3D audio, etc.).
Means for Solving the Problem
[0013] To achieve this object, one subject of the present invention relates to a method for encoding audio and / or video data implemented by an encoding device configured to implement at least one step of encoding audio and / or video data using an encoded artificial neural network, the encoding method comprising: - Encoding the audio and / or video data; - Generating a data signal including the encoded audio and / or video data; - Encoding information representing the decoding configuration that a decoding device should have for decoding the encoded data; - Inserting the encoded information into the data signal. includes.
[0014] Advantageously, according to the present invention, at least one encoding step is implemented using an encoded artificial neural network, and thus, an audio and / or video encoder that requires one or more dedicated components to implement this specific encoding step notifies the audio and / or video data decoder of the decoding capabilities that this decoder should have in order to be able to decode the audio and / or video data. For the purpose of transmitting this information to the audio and / or video data decoder, it is possible to encode the information representing these one or more corresponding encoding, and thus decoding components.
[0015] Another main subject of the present invention is a method for decoding encoded audio and / or video data implemented by a decoding device, comprising: - receiving an encoded audio and / or video data signal; - decoding from the signal information representing a decoding configuration in which at least one step of decoding the audio and / or video data is implemented using a decoded artificial neural network; - checking whether the decoding device has a decoding configuration corresponding to the decoded information; - depending on the result of the check, decoding or not decoding the signal. includes.
[0016] Advantageously, according to the present invention, the audio and / or video decoder is able to identify one or more information items regarding the decoding configuration that it should have in order to decode the received encoded audio and / or video data signal.
[0017] According to a specific embodiment of the above encoding or decoding method, the decoding configuration is - A first category corresponding to at least one specific physical characteristic of hardware or software supported by a decoding device so that a signal can be decoded, and / or - A second category corresponding to a specific characteristic of the signal, and / or - A third category corresponding to at least one specific processing function applied by a decoding device so that a signal can be decoded, belongs to.
[0018] Such embodiments advantageously enable these configurations to be grouped together by category to reduce the amount of information to be encoded when various encoding / decoding configurations are used. According to a specific embodiment of the encoding or decoding method described above, the information representing the decoding configuration to be encoded or decoded respectively is associated with at least one of the first, second, or third categories.
[0019] Such embodiments advantageously enable the information representing various types of decoding configurations to be encoded / decoded in a structured manner. Moreover, if there are multiple information items representing decoding configurations associated with various decoding parameters or characteristics of the same category, such embodiments enable a more compact signaling of this information because instead of each decoding parameter or characteristic being individually indicated within the signal, a single syntax element or indicator is signaled for the entire category of decoding parameters or characteristics.
[0020] According to a specific embodiment of the encoding or decoding method described above, the decoding configuration is - The maximum size of the data storage memory, and - The minimum number of operations per second, and - The minimum latent variable rate, and - A specific type of electronic circuit, and - The accuracy level of the mathematical representation of at least one operating parameter of the decoding artificial neural network, - activation or deactivation of at least one reference decoding step, and - at least one specific mathematical operator or a list of specific mathematical operators, and - a specific mathematical function, and - several entropy decoding statistical sources, and belong to a set comprising.
[0021] According to a specific embodiment of the encoding or decoding method described above, the information representing the decoding configuration is respectively included in a set of predetermined video parameters of the encoding or decoding method, or, when the video data represents an image sequence, included in a set of parameters associated with the sequence.
[0022] Such an embodiment advantageously enables the use of the encoding syntax of an existing or standardized encoder to encode the information representing the components. For example, in the case of an AVC, HEVC, or VVC encoder, the set of predetermined video parameters of the encoding method is, for example, a VPS (Video Parameter Set), and the set of parameters associated with the sequence is an SPS (Sequence Parameter Set). In another example, the information representing the decoding configuration is associated with a sub-image, in particular, a tile or a slice as defined, for example, in the HEVC standard.
[0023] According to a specific embodiment of the encoding or decoding method described above, the information representing the decoding configuration is - a first value associated with a first component of the decoding configuration, and - a second value associated with the first and second components of the decoding configuration, and comprises.
[0024] Such an embodiment advantageously enables these to be indicated by nesting them within the signal transmitted to the decoder when the configuration comprises a plurality of components corresponding to various decoding capabilities supported by the decoding device.
[0025] Any of the above-described embodiments or implementation features may be added, independently or in combination with each other, to the encoding or decoding method defined above.
[0026] Another subject of the present invention is an audio and / or video data signal, said signal being - audio and / or video data encoded by an encoding device configured to implement at least one step of encoding audio and / or video data using an encoded artificial neural network, and - encoding information representing a decoding configuration that a decoding device should have for decoding said encoded audio and / or video data, comprising.
[0027] Another subject of the present invention is a device for encoding audio and / or video data, configured to implement at least one step of encoding audio and / or video data using an encoded artificial neural network, said encoding device being - encoding audio and / or video data, - generating a data signal comprising the encoded audio and / or video data, - encoding information representing a decoding configuration that a decoding device should have for decoding said encoded data, - inserting said encoded information into said data signal, configured to implement.
[0028] Such an encoding device can, in particular, implement the encoding method described above.
[0029] Another subject of the present invention is a device for decoding encoded audio and / or video data, - receiving an encoded audio and / or video data signal, - Decoding information representing a decoding configuration in which at least one step of decoding audio and / or video data from the signal is implemented using a decoding artificial neural network; - Checking whether the decoding device has a decoding configuration corresponding to the decoded information; - Decoding or not decoding the signal according to the result of the check; is configured to implement.
[0030] Such a decoding device can, in particular, implement the above-described decoding method.
[0031] The present invention also relates to a computer program comprising instructions for implementing an encoding or decoding method according to the present invention, according to any one of the specific embodiments described above, when the program is executed by a processor.
[0032] Such instructions may be permanently stored in a non-transitory memory medium of an encoding device implementing the above-described encoding method or a decoding device implementing the above-described decoding method.
[0033] This program may use any programming language and may be in the form of source code, object code, or intermediate code between source code and object code, such as a partially compiled form, or any other desired form.
[0034] The present invention also relates to a computer-readable recording medium or information medium comprising instructions of a computer program as described above.
[0035] The recording medium may be any entity or device capable of storing a program. For example, the medium may comprise a ROM, such as a CD-ROM, a DVD-ROM, synthetic DNA (deoxyribonucleic acid), etc., or a super-small electronic circuit ROM, or magnetic recording means, such as a USB key or a hard disk, etc.
[0036] Furthermore, the recording medium may be a transmission medium such as an electrical or optical signal that can be carried via an electrical cable, an optical cable, wirelessly, or by other means. The program according to the present invention may in particular be downloaded via the Internet.
[0037] Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, and the circuit is adapted to execute or be used in the execution of the encoding or decoding method according to the present invention.
[0038] Other features and advantages will become apparent by reading the specific embodiments of the present invention given as illustrative and non-limiting examples, and the accompanying drawings.
Brief Description of the Drawings
[0039]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Embodiments for Carrying Out the Invention
[0040] Encoding of Audio and / or Video Data A method for encoding audio and / or video data representing 2D or 3D images or image sequences will be described below. Such an encoding method can be implemented, for example, in any type of video encoder or decoder, such as a neural network-based video encoder, according to, for example, JPEG, AVC, HEVC, VVC standards and their extensions (MVC, 3D-AVC, MV-HEVC, 3D-HEVC, etc.). Such an encoding method can also be implemented, for example, in any type of audio encoder or decoder, such as a neural network-based audio encoder, according to, for example, MP3, AAC (Advanced Audio Coding), MPEG-H 3D audio standard, etc.
[0041] Referring to FIG. 1, the encoding method according to the present invention includes the following.
[0042] In C1, current audio and / or video data is selected.
[0043] Such audio data is in the form of a current set B of samples c which is - a one-dimensional temporal audio signal, - a part of such a signal, - a multi-dimensional temporal audio signal (stereo or higher than 2 dimensions), and may be.
[0044] Such video data is in the form of a current set B of pixels c which is - the original current image, - a part or region of the original current image, - a block of the current image resulting from dividing this image in a manner consistent with that performed in a standardized AVC, HEVC, or VVC encoder, and may be.
[0045] In C2, a current prediction data set BP in which the data is, for example, pixelsc is calculated by predictions such as INTRA, INTER, IBC (intra-block copy), SKIP, etc., which are well-known to those skilled in the art.
[0046] In relation to audio coding, the data in the current prediction dataset BP c is a sample.
[0047] In C3, the current set B of pixels c and the current predicted set BP of pixels obtained in C2 c A signal BE representing the difference between them c is calculated.
[0048] In C4, this signal BE c is, for example, in accordance with criteria well-known to those skilled in the art, such as minimizing distortion / rate cost, or otherwise selecting the best efficiency / complexity trade-off point, etc. If it is for optimizing coding with respect to conventional coding performance criteria, the signal BE c is quantized and encoded.
[0049] At the end of this operation, the quantized and entropy-encoded difference signal BE c cod is obtained. Such entropy encoding is performed, for example, by Huffman coding or CABAC coding. In a preferred embodiment, the entropy encoding is CABAC coding.
[0050] In C5, the signal, i.e., the flow F, is generated to include the data DAT of the quantized and encoded difference signal BE c cod In a method known per se, the signal F can be transmitted to a decoding device or decoder described later in the description.
[0051] According to the present invention, at least one of operations C1 to C5 is implemented using a computing device based on artificial intelligence described as DCIA_C configured to automate it so as to make the at least one encoding operation more efficient and more adaptive. Such a computing device includes, for example, a neural network or a plurality of neural networks, a support vector machine, an inference engine, an expert system, a fuzzy logic system, etc.
[0052] In a preferred embodiment, the computing device is an encoding artificial neural network such as, for example, a convolutional neural network (CNN), a multi-layer perceptron, an LSTM (Long Short Term Memory), etc. Such a neural network is defined, for example, by a structure comprising a plurality of layers of artificial neurons and / or by a set of weights respectively associated with the artificial neurons of this network.
[0053] In particular, in a preferred embodiment, the neural network used is a convolutional neural network. In one particular embodiment, the latter calculates, at C3, the difference signal BE c or, together with the predicted set of pixels BP c generated at C2, the current set of pixels B cEncode and thus perform operations C3 and C4. Such a neural network is of the kind described, for example, in the following document, Ladune “Optical Flow and Mode Selection for Learning-based Video Coding”, IEEE MMSP 2020. In another specific embodiment, the prediction operation C2 is also implemented using a convolutional neural network and, in the case of audio samples, does not use a classical prediction device of the VVC or CELP (Code-Excited Linear Prediction) type. Such a neural network is described in particular in the document, Theo Ladune, Pierrick Philippe, “AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC”, February 17, 2022.
[0054] The use of one or more computing devices based on artificial intelligence to implement a method for encoding audio and / or video data requires that the encoding device implementing the encoding method has a specific hardware or software configuration. Such a configuration is correspondingly required in the decoding device, such that the latter can decode in real time the encoded audio and / or video data signals received from an encoding device comprising one or more of these computing devices. This decoding configuration belongs to one or more categories comprising, for example, the following. - A first category corresponding to at least one specific physical characteristic of the hardware or software supported by the decoding device so that the signal F can be decoded, and / or - A second category corresponding to one or more specific characteristics of the data signal F when at least one encoding step is performed by the computing device DCIA_C, - A third category corresponding to at least one specific processing function applied by the decoding device so that the signal F can be decoded.
[0055] As a non-exhaustive example, the first category of decoding features includes the following. - A specific type of electronic circuit used by a computing device, such as a GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), DSP (Digital Signal Processor), or any other suitable type of electronic circuit. - The data memory size that the decoding device should have to store all the data necessary for neural decoding of signal F, such as buffer memory. - The minimum number of operations per second, such as TOPS (Tera Operations Per Second) that the decoding device must be able to perform. - The level of accuracy of the representation of the parameters of computing device DCIA_C that the decoding device must conform to in order to be able to decode signal F, where such parameters include, for example, the weights of one and / or more decoding artificial neural networks used by the decoding device, the parameters of the activation function applied at the output of the artificial neurons, etc. - The number of latent variable entropy decoding sources that the decoding device must support in order to be able to decode signal F. - Etc.
[0056] As a non-exhaustive example, the second category of decoding features includes the following. - When one or more neural networks are used such that the decoding device maintains its real-time decoding ability, the minimum number of latent variables processed per unit time by the decoding device. - The maximum number of latent variables per unit time that enables the decoding device to check whether it has the ability to decode this signal. - The minimum number of bits representing the syntactic elements aimed at reconstructing the latent variables (entropy rate) per unit time such that the decoding device maintains its real-time decoding ability. - The maximum number of bits representing a syntactic element aimed at reconstructing the latent variable (entropy rate) per unit time so that it is possible to check whether the decoding device has the ability to decode this signal. - etc.
[0057] The decoding characteristics are transmitted to the decoding device by an indicator explicitly giving the values of these characteristics, or by a more global indicator showing the values of a plurality of characteristics at once using a predetermined association table such as table T10 or table T11 described below, for example.
[0058] The advantage of transmitting the latent variable rate (expressed as the number of latent variables per unit time or the coding rate of these latent variables per unit time) is to enable adjusting the complexity of the data that requires neural network processing according to the ability of the decoding device to perform neural network processing. In fact, for example, an encoded signal representing an image or video may include both encoded data whose decoding performs a conventional processing operation usually executed using a CPU (Central Processing Unit) processor, and data whose decoding performs a neural processing operation usually executed using a specific processor (GPU or TPU processor) that enables executing a very large number of small calculations of the same nature in parallel. Specifying the characteristics of the latent variable rate facilitates the adjustment of the lower part of the signal related to neural data processing and the GPU or TPU capabilities of the decoding device.
[0059] As a non-exhaustive example, the third category of decoding characteristics comprises the following. - A set of predefined and standardized logical and / or mathematical operators supported by the decoding device to enable decoding the signal, and / or a list of logical and / or mathematical operators. - A list of activation functions applied by the decoding device at the output of one and / or more neural networks to enable decoding the signal. - The ability to reproduce the results of decoding in a manner faithful to the criteria (platform - to - platform reproducibility) that a decoding device should have in order to be able to decode the signal - etc.
[0060] All of these decoding characteristics are considered in both the context of encoding / decoding audio data and in the context of encoding / decoding video data.
[0061] According to the present invention, the encoding method includes a step C6 of encoding one or more information items ICD representing a decoding configuration that a decoding device should have to decode the encoded data DAT, and such a decoding configuration belongs to the above - mentioned first and / or second and / or third category.
[0062] At the end of encoding C6, one or more information items ICD cod are obtained.
[0063] One or more information items ICD cod are then written in step C7 either to the data signal F or to a signal F' associated with the signal F.
[0064] Referring to Figure 2A, one or more information items ICD cod are written in step C7 into data packets that can be independently identified and decoded with respect to the signal F and / or with respect to the encoded audio and / or video data DAT. The packets contain other information necessary to decode the encoded data DAT, such as predetermined audio and / or video parameters of the encoding method, or, when the encoded data DAT represents an image sequence, parameters associated with this image sequence in a manner similar to the VPS or SPS syntax, for example as implemented in the VVC standard. According to another embodiment, one or more information items ICD cod are parameters associated with sub - images, in particular tiles or slices as defined, for example, in the HEVC standard.
[0065] Referring to FIG. 2B, one or more information items ICD cod are written in packet F' that does not need to be decoded to decode data DAT in C7 in the same way as writing information in an SEI (Supplemental Enhancement Information) message according to, for example, the VVC standard.
[0066] In C8, a signal F including the encoded data DAT and one or more encoded information items ICD cod or, alternatively, a message F' including a signal F including the encoded data DAT and one or more encoded information items ICD cod is stored or transmitted to a decoding device described later in this description.
[0067] In a preferred embodiment, each decoding configuration or feature represented by the information item ICD cod is signaled individually.
[0068] Various examples of the encoded information ICD cod are shown below in the corresponding syntax table.
[0069] For example, for a particular type of hardware such as a particular electronic circuit or processor supported by a decoding device, an indicator proc_idc as shown in the following syntax table T1 takes, for example, eight values in the range from 0 to 7 and indicates to the decoding device the type of processor supported for decoding signal F.
[0070]
Table 1
[0071] Regarding the size of the data memory that the decoding device should have, the indicator buffer_size_idc may specify the size of this memory in bytes, or alternatively, may take a predetermined number of values associated with predefined size limits, as in the following syntax table T2. For example, the indicator buffer_size_idc takes 8 values in the range from 0 to 7.
[0072]
Table 2
[0073] Regarding the minimum number of operations per second supported by the decoding device, the indicator ops_idc may specify the number of operations per second, or alternatively, may take a predetermined number of values associated with predefined limits, as in the following syntax table T3. In this case, the indicator ops_idc takes, for example, 4 values from 0 to 3.
[0074]
Table 3
[0075] Regarding the minimum number of latent variables processed per unit time by the decoding device, the indicator latent_rate_idc may specify this number as the number of variables per second, i.e., in terms of rate, or alternatively, may take a predetermined number of values associated with predefined rate limits, as in the following syntax table T4. In this case, the indicator latent_rate_idc takes, for example, 7 values from 0 to 6.
[0076]
Table 4
[0077] Regarding the level of precision of the representation of the parameters of computing device DCIA_C, and if this device is a neural network, the indicator precision_idc specifies the required precision of the weights of this network using the following predetermined association table T5. In this case, the indicator precision_idc takes, for example, eight values from 0 to 7.
[0078]
Table 5
[0079] Regarding the ability to reproduce the decryption result in a way that is faithful to the criteria (inter-platform reproducibility) that the decryption device should have, in the following table T6, the indicator repro_flag - A first value, for example 0, indicating that the decryption device does not need to reproduce the reference decryption identically - Or a second value, for example 1, indicating that the decryption device needs to reproduce the reference decryption identically is set to either one of them.
[0080]
Table 6
[0081] Regarding the number of latent variable entropy decryption sources that the decryption device should support, the indicator sources_idc either specifies the number of sources or, alternatively, can take a predetermined number of values associated with a predefined limit on the number of sources, as in the following table T7. Here, the indicator sources_idc takes, for example, four values from 0 to 3.
[0082]
Table 7
[0083] Regarding a set or list of predetermined standardized logical and / or mathematical operators supported by a decoding device, the indicator operators_idc specifies the logical and / or mathematical operators supported on the decoder side. According to one preferred embodiment shown in the following syntax table T8, the indicator operators_idc has five values from 0 to 4, and the values from 0 to 3 are structured in a nested manner as follows. - Value 0 is associated with the set of basic mathematical operators "+", "-", "×", "÷", - Value 1 is associated with the set of basic mathematical operators "+", "-", "×", "÷" and the set of mathematical operators "xy", "exp()", "sqrt()", - Value 2 is associated with the set of basic mathematical operators "+", "-", "×", "÷", the set of mathematical operators "xy", "exp()", "sqrt()", and the operator "N!", where N is a natural number, - Value 3 is associated with the set of basic mathematical operators "+", "-", "×", "÷", the set of mathematical operators "xy", "exp()", "sqrt()", the operator "N!", and the set of mathematical operators "sin()", "cos()", "tan()".
[0084]
Table 8
[0085] Of course, this method of signaling operators is not exhaustive. In other embodiments, each operator or list of operators is signaled individually, which may result in a higher signaling cost.
[0086] Regarding the list of activation functions applied by a decoding device, the indicator activations_idc specifies the activation functions supported by the decoding device. According to one preferred embodiment shown in the following syntax table T9, the indicator activations_idc has three values from 0 to 2, and the values 0 and 1 are structured in a nested manner as follows. - Value 0 is associated with the following list of activation functions: F(x) = x, if x < 0, then G(x) = 0; otherwise, G(x) = 1, if x < 0, then H(x) = 0; otherwise, H(x) = x. - The value 1 is associated with this list of activation functions and also with the following list of activation functions: I(x) = 1 / (1 + exp(-x)), J(x) = tan -1 (x).
[0087]
Table 9
[0088] In a particular embodiment, a single indicator level_idc simultaneously specifies a plurality of decoding features supported by the decoding device. To achieve this purpose, a corresponding table, hereinafter referred to as T10, which is predefined on both the encoding device and the decoding device sides, is generated. Table T10 maps at least one specific value of the indicator level_idc to a specific latent variable rate, a specific data memory size, etc. In the illustrated embodiment, the indicator level_idc has six values from 0 to 5.
[0089]
Table 10
[0090] In a particular embodiment, instead of a single indicator level_idc, a series of indicators cat1_level_idc, cat2_level_idc, cat3_level_idc specify one or more decoding features according to the first, second, or third category to which these one or more decoding features belong.
[0091] To achieve this objective, a correspondence table predefined on both the encoding device and the decoding device sides is generated for each of the indicators cat1_level_idc, cat2_level_idc, cat3_level_idc. Such a table is shown below and denoted by reference numeral T11.
[0092] Regarding the first category, the following table T11 maps at least one specific value of the indicator cat1_level_idc to a specific type of processor, a specific data memory size, etc. In the illustrated embodiment, the indicator cat1_level_idc has six values from 0 to 5.
[0093] [Table 11]
[0094] Regarding the second category, the following table T12 maps at least one specific value of the indicator cat2_level_idc to a specific latent variable rate. In the illustrated embodiment, the indicator cat2_level_idc has five values from 0 to 4.
[0095] [Table 12]
[0096] Regarding the third category, the following table T13 maps at least one specific value of the indicator cat3_level_idc to the ability or inability to reproduce the decoding result, a list of specific mathematical operators, etc. In the illustrated embodiment, the indicator cat3_level_idc has four values from 0 to 3.
[0097] [Table 13]
[0098] Next, referring to FIG. 3, the encoder COD shown in the form of a schematic diagram will be described. The encoder COD is designed to implement the encoding method shown in FIG. 1 in a specific embodiment of the present invention.
[0099] According to this specific embodiment, the operations performed by the encoding method are implemented by computer program instructions. For this purpose, the encoding device COD has a conventional architecture of a computer, and in particular, includes a memory MEM_C and, for example, a processor PROC_C, and a processing unit UT_C driven by a computer program PG_C stored in the memory MEM_C. The computer program PG_C includes instructions for implementing the operations of the encoding method as described above when the program is executed by the processor PROC_C.
[0100] At the time of initialization, the code instructions of the computer program PG_C are loaded into a RAM memory (not shown), for example, before being executed by the processor PROC_C. The processor PROC_C of the processing unit UT_C particularly implements the operations of the encoding method according to the instructions of the computer program PG_C.
[0101] The encoder COD receives, at the input E_C, the current set B of pixels or samples c and sends, at the output S_C, the transport flow F, which is transmitted to the decoder using an appropriate communication interface (not shown).
[0102] The encoder COD includes a prediction device PRED configured to perform the prediction step C2 described above. As already described above in this description, this prediction device - may be conventional and may be configured, for example, according to standards such as HEVC, VVC, CELP, etc. - For example, it may be a type of neural network described in the above-mentioned document, Theo Ladune, Pierrick Philippe, "AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC", on February 17, 2022. - And so on.
[0103] The encoder COD also includes an artificial intelligence-based computing device DCIA_C, which is of the type described, for example, in the above-mentioned document, Ladune "Optical Flow and Mode Selection for Learning-based Video Coding", IEEE MMSP 2020.
[0104] The encoder COD also includes an information encoding device CICD configured to perform the above-described step C6 of encoding one or more information items ICD representing the decoding configuration that a decoding device should have to decode the signal F.
[0105] The encoder COD also includes the information ICD obtained by the device CICD cod and also includes a device IICD configured to perform the above-described step C7 of writing it into either the data signal F or the message F' associated with the signal F.
[0106] The encoder COD also includes a storage memory MS_C configured to store the syntax tables T1 to T12. Alternatively, this storage memory MS_C is not included in the encoder COD but is accessible thereto, for example, via a communication network using any suitable means. In one embodiment, the data signal F and any message F' may also be stored in the storage memory MS_C or an additional storage memory (not shown).
[0107] Decoding of Encoded Audio and / or Video Data A method for decoding an encoded audio and / or video data signal related to a 2D or 3D image or image sequence will be described below. Such a decoding method can be implemented, for example, in any type of video decoder, such as a video decoder based on a neural network, according to, for example, the JPEG, AVC, HEVC, VVC standards and their extensions (MVC, 3D-AVC, MV-HEVC, 3D-HEVC, etc.). Such a decoding method can also be implemented, for example, in any type of audio decoder, such as an audio decoder based on a neural network, according to, for example, the MP3, AAC, MPEG-H 3D audio standards, etc.
[0108] Referring to FIG. 4, the decoding method according to the present invention includes the following.
[0109] In D1, the above-mentioned data signal F is received by the decoding device DEC shown in FIG. 5, and the data signal includes the encoded audio and / or video data DAT and the encoded information ICD representing a specific decoding configuration that the decoding device DEC should have for decoding the signal F. cod and.
[0110] Alternatively, in D1, the decoding device DEC receives the following. - A data signal F including encoded audio and / or video data DAT, - Encoded information ICD representing a specific decoding configuration that the decoding device DEC should have for decoding the signal F. cod A message F' including.
[0111] In D2, the encoded audio and / or video data DAT and the encoded information ICD cod are extracted from the received data signal F.
[0112] Alternatively, in D2, the encoded audio and / or video data DAT is extracted from the received data signal F, and the encoded information ICD cod is extracted from the received message F'.
[0113] According to the present invention, one or more encoded information items ICD cod are decoded at D3. At the end of this operation, the information ICD is reconstructed, which enables the decoder to identify one or more decoding functions necessary to decode the encoded audio and / or video data DAT. For this purpose, in a particular embodiment of the present invention, the values of one or more indicators such as, for example, indicator proc_idc, buffer_size_idc, ops_idc, latent_rate_idc, precision_idc, repro_flag, sources_idc, operators_idc, activations_idc, level_idc, cat1_level_idc, cat2_level_idc, cat3_level_idc, etc. are read, and then, particularly in the case of indicators cat1_level_idc and cat3_level_idc, they are mapped to their associated decoding features or other their associated decoding features. Such mapping is performed using the above-described correspondence tables T1 to T13 accessible to the decoding device DEC.
[0114] At D4, the decoding device DEC compares each of the one or more decoding features identified at D3 with one or more decoding features specific to it. Such comparison is made possible by the fact that the decoding device DEC can access the technical features of the platform (this software or hardware or hybrid) on which it operates and the performance it can achieve.
[0115] At the end of comparison D4, if the decoding device DEC does not have one or more decoding features identified at D3, the decoding of the encoded audio and / or video data DAT is not performed. Thus, the decoding is discarded (ABD).
[0116] At the end of comparison D4, if the decoding device DEC has one or more decoding features identified in D3, the decoding of the encoded audio and / or video data DAT is performed using an artificial intelligence-based computing device DCIA_D, and the computing device DCIA_D performs decoding corresponding to the encoding performed by the computing device DCIA_C.
[0117] To achieve this purpose, in D5, the inverse quantization and entropy decoding of the encoded audio and / or video data DAT are performed. Such entropy decoding is, for example, Huffman decoding or CABAC decoding. In a preferred embodiment, the entropy decoding is CABAC decoding. At the end of this operation, the decoded difference signal BE c dec is obtained.
[0118] In D6, prediction is performed to generate the current prediction data set BP c and these data are, for example, pixels here, but may also be samples of an audio signal in some cases.
[0119] Steps D5 and D6 may be performed in any order or simultaneously.
[0120] In D7, the set of current pixels BD c is calculated by combining the decoded difference signal BE c dec obtained in D5 with the prediction set BP c of pixels or samples obtained in D6.
[0121] In a method known per se, the set of current pixels BD c may optionally be filtered by a loop filter performed on the reconstructed signal, which is well known to those skilled in the art.
[0122] Regardless of the difference signal BE calculated between the above encoding methods c is zero, this may be the case of the SKIP encoding mode, but the difference signal BE c The step D2 of extracting and the steps D5 of inverse quantization and entropy decoding are not performed.
[0123] Next, referring to FIG. 5, a decoder DEC shown in the form of a schematic diagram will be described. The decoder DEC is designed to implement the decoding method shown in FIG. 4 in a specific embodiment of the present invention.
[0124] According to this specific embodiment, the operations performed by the decoding method are implemented by computer program instructions. For this purpose, the decoding device DEC has a conventional architecture of a computer, and in particular, includes a memory MEM_D and, for example, a processor PROC_D, and a processing unit UT_D driven by a computer program PG_D stored in the memory MEM_D. The computer program PG_D includes instructions for implementing the operations of the decoding method as described above when the program is executed by the processor PROC_D.
[0125] The decoder DEC receives, at the input E_D, the data signal F transmitted by the encoder COD of FIG. 3, and optionally the message F', and at the output S_D, the current set BD of decoded pixels or samples c is sent out.
[0126] The decoder DEC also includes an information decoding device DICD configured to perform the above-described step D3 of decoding one or more encoding information items ICD representing the decoding configuration that the decoder DEC should have to decode the signal F. cod
[0127] The decoder DEC also includes a device COMP configured to compare, at D4, one or more reconfigured information items ICD representing the decoding configuration that the decoder DEC should have to decode the signal F with specific decoding features of the decoder DEC.
[0128] In a method corresponding to the encoder COD of FIG. 3, the decoder DEC also includes a storage memory MS_D configured to store the above-described syntax tables T1 to T13. Alternatively, this storage memory MS_D is not included in the decoder DEC but can be accessed thereto using any suitable means, for example, via a communication network. In one embodiment, the data signal F and any message F' may also be stored in the storage memory MS_D or an additional storage memory (not shown).
[0129] The decoder DEC includes a prediction device PRED_D configured to perform the above-described prediction step D6. As already explained above in this description, this prediction device may be - conventional and may be configured, for example, according to standards such as HEVC, VVC, CELP, etc., - for example, a neural network of the kind described in the above-mentioned document, Theo Ladune, Pierrick Philippe, “AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC”, on February 17, 2022, - etc.
[0130] The decoder DEC may also include an artificial intelligence-based computing device referred to as DCIA_D, which is configured to automate at least one decoding operation and make it more efficient and more adaptable. Such a computing device includes, for example, a decoding artificial neural network or a plurality of decoding artificial neural networks, a support vector machine, an inference engine, an expert system, a fuzzy logic system, etc.
[0131] In a preferred embodiment, the computing device DCIA_D is, for example, a neural network such as a convolutional neural network (CNN), a multi-layer perceptron, an LSTM, etc. In particular, in a preferred embodiment, the neural network used is a convolutional neural network. Such a neural network is defined, for example, by a structure comprising a plurality of layers of artificial neurons and / or by a set of weights respectively associated with the artificial neurons of this network.
[0132] In a specific embodiment, the neural network DCIA_D combines, at D7, the decoded differential signal BE obtained at D5 c dec with a predicted set BP of pixels or samples generated at D6 c Such a neural network is, for example, of the type described in the following document: Ladune, "Optical Flow and Mode Selection for Learning-based Video Coding", IEEE MMSP 2020. In another specific embodiment, the prediction operation D6 is also implemented using a convolutional neural network and does not use, for example, a classical prediction device of the VVC type. Such a neural network is described in particular in the document: Theo Ladune, Pierrick Philippe, "AIC ARTIFICIAL INTELLIGENCE BASED VIDEO CODEC", February 17, 2022.
[0133] Needless to say, the embodiments described above are given purely as a completely non-limiting indication, and numerous modifications may be easily made by those skilled in the art without departing from the scope of the present invention.
Explanation of Reference Numerals
[0134] Operations C1 to C7 CICD device CNN Convolutional Neural Network COD Encoding Device COMP Device D1 - D7 Steps DAT Data DCIA_C Encoded Artificial Neural Network DCIA_D Decoded Artificial Neural Network DEC Decoding Device DICD Information Decoding Device E_C Input E_D Input F Signal F’ Signal ICDcod Information Item IICD Device level_idc Indicator MEM_C Memory MEM_D Memory MS_C Storage Memory MS_D Storage Memory ops_idc Indicator PG_C Computer Program PG_D Computer Program PRED Prediction Device PRED_D Prediction Device PROC_C Processor PROC_D Processor proc_idc Indicator repro_flag Indicator S_C Output S_D Output
Claims
1. A method for encoding said audio and / or video data, implemented by an encoding device configured to implement at least one step of encoding the audio and / or video data using a quantized convolutional neural network (DCIAC), comprising: - encoding said audio and / or video data (C2-C4); - generating a data signal comprising said encoded audio and / or video data (C5); - encoding information representing a decoding configuration for use of said network, which a decoding device should have in order to decode said encoded data (C6); - inserting said encoded information into said data signal (C7). An encoding method comprising the above steps.
2. A method for decoding encoded audio and / or video data, implemented by a decoding device, comprising: - receiving an encoded audio and / or video data signal (D1); - decoding from said signal information representing a decoding configuration in which at least one step of decoding said audio and / or video data is implemented using a decoding convolutional neural network (DCIAD) (D3); - checking whether said decoding device has said decoding configuration corresponding to said decoded information (D4); - decoding said signal (D5-D7) or not according to the result of said check. A decoding method comprising the above steps.
3. Said decoding configuration belongs to: - a first category corresponding to at least one specific physical characteristic of hardware and / or software supported by said decoding device so as to be able to decode said signal; and / or - a second category corresponding to a specific characteristic of said signal; and / or - a third category corresponding to at least one specific processing function applied by said decoding device so as to be able to decode said signal. The encoding method according to Claim 1 or the decoding method according to Claim 2.
4. The information representing the decoding configuration to be encoded or decoded respectively is associated with at least one of said first, second or third categories. The encoding or decoding method according to Claim 3.
5. Said decoding configuration includes: - the maximum size of the data storage memory; - The minimum number of operations per second, and - The minimum latent variable rate, and - A predetermined number of values respectively associated with a predetermined latent variable rate limit, and - A specific type of electronic circuit, and - The minimum number of bits representing a syntax element aimed at reconstructing the latent variable per unit time, and - The maximum number of bits representing a syntax element aimed at reconstructing the latent variable per unit time, and - The accuracy level of the mathematical expression of at least one operation parameter of the decoding artificial neural network, and - The activation or deactivation of at least one decoding identical to the reference decoding, and - At least one specific mathematical operator or a list of specific mathematical operators, and - A specific mathematical function, and - Some entropy decoding statistical sources, and belonging to a set comprising, the encoding method according to any one of claims 1 or 3 to 4 or the decoding method according to any one of claims 2 to 4.
6. The information representing the decoding configuration is respectively included in a set of predetermined video parameters of the encoding or decoding method, or, when the video data represents an image sequence, included in a set of parameters associated with the sequence, the encoding method according to any one of claims 1 or 3 to 5 or the decoding method according to any one of claims 2 to 5.
7. The information representing the decoding configuration is - A first value associated with a first component of the decoding configuration, and - A second value associated with the first and second components of the decoding configuration, and comprising, the encoding method according to any one of claims 1 or 3 to 6 or the decoding method according to any one of claims 2 to 6.
8. A device for encoding audio and / or video data, configured to implement at least one step of encoding the audio and / or video data using an encoding artificial neural network, wherein - Encoding the audio and / or video data, and - Generating a data signal including the encoded audio and / or video data, and - Encoding information representing a decoding configuration regarding the use of the network that a decoding device should have for decoding the encoded data, and - Inserting the encoded information into the data signal, and configured to implement, an encoding device. Claim 9 A device for decrypting encoded audio and / or video data, comprising: - receiving an encoded audio and / or video data signal; - decrypting, from the signal, information representing a decryption configuration in which at least one step of decrypting the audio and / or video data is implemented using a decryption artificial neural network; - checking whether the decryption device has the decryption configuration corresponding to the decrypted information; - decrypting or not decrypting the signal according to the result of the check; A decryption device configured to implement the above. Claim 10 A computer program comprising program code instructions for implementing the encoding method according to any one of Claims 1 or 3 to 7 or the decoding method according to any one of Claims 2 to 7 when executed on a computer. Claim 11 A computer-readable information medium comprising the instructions of the computer program according to Claim 10.