Decryption method and apparatus, related computer program and data stream

The method efficiently identifies and utilizes neural networks within data packets to decode audio or video content by processing first data packets to obtain and reuse neural networks, enhancing decoding efficiency.

JP2025522777APending Publication Date: 2025-07-17オランジュ
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024576526
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-29
Filing Date
2023-06-28
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing decoding methods for audio or video content using artificial neural networks lack efficient mechanisms for identifying and utilizing neural networks within data packets, leading to inefficiencies in decoding processes.

Method used

A method and apparatus that identifies a first data packet containing information indicating a predetermined type, processes second data within that packet to obtain an artificial neural network, and decodes subsequent data packets using this network to recreate audio or video content, with optional reuse of previously obtained networks.

Benefits of technology

Facilitates efficient identification and utilization of neural networks within data packets, enabling effective decoding of audio or video content by reusing obtained networks, thus optimizing decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522777000001_ABST
    Figure 2025522777000001_ABST
Patent Text Reader

Abstract

The data stream includes data packets each containing at least first data and second data. A method for decoding this data stream includes: - identifying a first data packet (12) among the aforementioned data packets, wherein the first data thereof includes information (T1) indicating a data packet of a predetermined type; - processing the second data (NNC) of the first data packet to obtain an artificial neural network; - decoding the second data (C) included in a second data packet (14) among the aforementioned data packets using at least the obtained artificial neural network, thereby creating data representing audio or video content. Also proposed are a related decoding apparatus and a computer program.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of encoding audio or video content.

Background Art

[0002] The present invention particularly relates to a decoding method and apparatus, and related computer programs and data streams.

[0003] It has been proposed to use an artificial neural network to perform all or part of the decoding of data representing audio or video content.

[0004] (Patent Document 1) discloses a decoding method in which an indicator is decoded to determine whether an artificial neural network is encoded in a received data stream or forms part of a set of predetermined artificial neural networks, and then the artificial neural network is used to decode data representing audio or video content.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0006] In relation to the present specification, the present invention is a method for decoding a data stream including data packets each including at least first data and second data, - identifying, among the aforementioned data packets, a first data packet in which the first data includes information indicating a data packet of a predetermined type; - Processing the second data of the first data packet to obtain an artificial neural network; - Decoding the second data included in the second data packet among the aforementioned data packets by using at least the obtained artificial neural network, thereby creating data representing audio or video content; A method is proposed, which is characterized by including the above steps.

[0007] In this way, data enabling the obtaining of an artificial neural network and data decodable using this artificial neural network to reproduce audio or video content are carried within individual data packets, which facilitates the identification and use of those data during decoding. The first data packet containing data enabling the obtaining of an artificial neural network is specifically identified by information indicating a predetermined type of data packet in this regard.

[0008] The second data of the first data packet may include, for example, description data of an artificial neural network, and the processing step may be a step of decoding the description data to obtain parameters of the artificial neural network.

[0009] The first data packet may also include an identifier of the artificial neural network. This identifier may be an element of a list of distinct identifiers individually associated with distinct artificial neural networks.

[0010] This method may further include receiving a third data packet whose first data includes information indicating a packet of the aforementioned predetermined type and includes the aforementioned identifier, and / or reusing the obtained artificial neural network to decode second data included in a fourth data packet among the aforementioned data packets. Therefore, the presence of the identifier in the third data packet indicates that the third data packet also includes data that can be used to obtain the artificial neural network defined in the first data packet, and thus this artificial neural network can be reused without the need to process the data of the third data packet again.

[0011] The second data packet may also include the aforementioned identifier. In other words, the identifier included in the first data packet and the identifier included in the second data packet are the same. In this case, the identifier can be used to indicate that the artificial neural network defined in the first data packet (which also includes this identifier) must be used to decode the data included in the second data packet (which includes the identifier in this case).

[0012] However, other possibilities can also be considered to indicate the neural network used for decoding.

[0013] This method can include receiving another data packet including parameters related to at least one image of the aforementioned content, and in this case, these parameters may include the aforementioned identifier. Then, the artificial neural network defined in the first data packet is used to decode the data for obtaining the aforementioned at least one image of the content.

[0014] According to one possible embodiment, the first data packet is the last packet in the data stream that precedes a second data packet in the data stream among the data packets whose first data includes information indicating a packet of the aforementioned predetermined type. Thus, in this embodiment, the artificial neural network used to decode the data of the second packet is defined within the last received packet having the predetermined type.

[0015] According to a possible embodiment, the second data packet includes a pointer to the first data packet.

[0016] Following the same concept, the method - includes a step of reading a flag within the second data packet, and - when the flag has a default value, includes a step of reading a pointer to the first data packet within the second data packet and may include.

[0017] The method - includes a step of receiving another data packet including parameters related to at least one image of the aforementioned content, - includes a step of reading a flag from among the aforementioned parameters, and - when the flag has a default value, includes a step of reading a pointer to the first data packet from among the aforementioned parameters and may further include.

[0018] The pointer may specify a position in a part of the data stream related to an image sequence different from, for example, the image sequence at least partially encoded by the second data of the second data packet.

[0019] Furthermore, the first data packet may include information indicating the encoded form of the second data of the first data packet.

[0020] The first data packet can start with a default marker, and in this case, the second data packet can also start with the aforementioned default marker. In this case, such a marker identifies the start of the data packet.

[0021] The first data of the data packet is included, for example, in the header of this data packet, while the second data can be included in the payload data of this data packet.

[0022] The present invention is an apparatus for decoding a data stream including data packets each including at least first data and second data, - identifying, from among the aforementioned data packets, a first data packet in which the first data thereof includes information indicating a data packet of a predetermined type, - processing the second data of the first data packet to obtain an artificial neural network, - decoding the second data included in the second data packet among the aforementioned data packets by using at least the obtained artificial neural network, thereby creating data representing audio or video content characterized by including a processor configured or programmed to perform the above.

[0023] The present invention also proposes a computer program including instructions executable by a processor and designed to implement the method proposed above when those instructions are executed by the processor.

[0024] Finally, the present invention is a data stream including data packets each including at least first data and second data, wherein the data packets are - a first data packet in which the first data thereof includes information indicating a data packet of a predetermined type and the second data thereof defines an artificial neural network, - Second data packets such that the second data can be decoded using at least the artificial neural network described above to create data representing audio or video content, and A data stream is proposed, characterized in that it comprises.

[0025] As described above, the first data packet can include an identifier, the data stream can include another data packet including the identifier described above, the first data includes the information described above indicating a data packet of a predetermined type, and the second data is identical to the second data of the first data packet.

[0026] Of course, the various features, modifications, and embodiments of the present invention can be related to each other according to various combinations as long as they are not incompatible or mutually exclusive.

[0027] Furthermore, various other features of the present invention will become apparent from the accompanying description made with respect to the drawings showing non-limiting embodiments of the present invention.

Brief Description of the Drawings

[0028]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0029] Note that in these drawings, structural elements and / or functional elements common to various modified forms may have the same reference numerals.

[0030] FIG. 1 shows an encoding device used within the scope of the present invention.

[0031] This encoding device includes a management module 2, an encoding module 4, a stream formation module 6, and a stream emission module 8.

[0032] Each of these modules can be implemented by a programmed processor (e.g., by instructions stored in a memory associated with the processor) that actually implements the functions described below for the relevant modules (in this example, for the processor to execute a part of the aforementioned instructions). Further, in practice, several modules can be implemented by the same processor by executing, for example, several instruction sets respectively corresponding to the various modules (by this processor). As a modification, any of the modules can be created by an application-specific integrated circuit.

[0033] The management module 2 is configured to control the operation of the encoding module 4, in particular to determine which encoding process must be used to encode the data B representing the audio or video content, as will be described below.

[0034] The encoding module 4 is configured to receive this data B representing the audio or video content as an input and, based on at least a part of the data B, generate an encoded representation C of this content as an output. The size of the encoded representation C (in units of the number of bits) is usually smaller than the size of the corresponding data B (in units of the number of bits).

[0035] In the case of video content, the data B includes values individually associated with the pixels of the images (or components of the images) of the video sequence. Thus, the data B can be the luminance values or chrominance values individually associated with the pixels of the components of the images of the associated video sequence.

[0036] In the case of audio content, the data B is data representing a sound signal in the WAV format (used, for example, to store an audio compact disc).

[0037] To generate the encoded representation C based on the data B, the encoding module 4 uses at least one artificial neural network N, N'.

[0038] According to the first possible embodiment shown in FIG. 2, data B representing audio or video content is applied as an input to artificial neural network N, and artificial neural network N generates a corresponding part of the encoded representation C as an output.

[0039] The data B applied as an input to artificial neural network N (i.e., applied to the input layer of artificial neural network N) can represent a block of an image, or a block of components of an image (e.g., a block of the luminance component or chrominance component of this image or a block of the color components of this image), or an image of a video sequence, or components of an image of a video sequence (e.g., a luminance component or chrominance component or color component), or even a series of images of a video sequence.

[0040] In this case, it can be stipulated that at least some of the neurons in the input layer respectively receive pixel values of components of an image, the value of which is represented by one of the data items B.

[0041] According to the second possible embodiment, encoding module 4 processes data B representing audio or video content in several steps, at least one of which is executed by artificial neural network N'.

[0042] Therefore, for example, as shown in FIG. 3, a previously obtained part C of the encoded representation j-1 is applied as an input to artificial neural network N', which enables prediction data P j to be generated as an output from artificial neural network N', and this prediction data is subtracted from the current data B j to obtain the part C j of the encoded representation corresponding to the current data B j as (the output from encoding module 4).

[0043] In FIG. 3, reference numeral 10 is the part C j of the corresponding encoded representation to obtain the current data Bj When processing, the previously obtained part C of the encoded representation is applied as the input to the artificial neural network N'. j-1 It represents a delay module for showing the fact that it is so.

[0044] Actually, the part C j-1 is the data B related to the image preceding the image represented by, for example, the current data B j and is previously obtained by processing (by the encoding module 4) the data B related to the image preceding the image represented by the current data B. j-1 In a modified form, the previously obtained part used as the input to the artificial neural network N' can be a part of the encoded representation corresponding to at least one block of the image adjacent to the block represented by the current data B

[0045] whose pixel values are. j During encoding, the management module 2 determines which encoding process (i.e., which process executed by the encoding module 4) must be used to encode the data set B representing the audio or video content.

[0046] Therefore, the management module 2 determines, among other things, which artificial neural networks N, N' must be used within the encoding module 4.

[0047] The data set B for which the management module 2 determines the process to use (especially the artificial neural networks N, N') depends on the relevant application. This data set B is, for example, a set of data B related to a given image or a data set B related to a given image sequence.

[0048]

[0049] The management module 2 selects, for example, from among a plurality of predetermined artificial neural networks, the artificial neural networks N, N' to be used when encoding the dataset B in order to minimize, for example, a throughput distortion criterion (considering the distortion between the size of the encoded representation C and the content represented by the data B and the content reconstructed based on the encoded representation C).

[0050] As a modified form, the management module 2 executes the step of training the artificial neural networks N, N' so as to optimize a given criterion (such as the aforementioned throughput distortion criterion) when processing the relevant dataset B, and instructs the encoding module 4 to use the artificial neural network thus trained to generate the encoded representation C based on the relevant dataset B.

[0051] Therefore, the management module 2 can create information i indicating the artificial neural network used when decoding the encoded representation C (especially intended for the stream formation module 6). In particular, the management module 2 can provide such information i for each part C of the encoded representation related to the dataset B defined above.

[0052] In some cases, the artificial neural network used to decode the encoded representation C is different from the artificial neural network N.

[0053] For example, in the case of FIG. 2 (where the artificial neural network N receives the data B as input and creates the encoded representation C as output), the artificial neural network used to decode the encoded representation C is designed (i.e., actually trained) to minimize the distortion of the data B while continuously passing through the artificial neural network N (to create the encoded representation C) and while continuously passing through the artificial neural network used for decoding, and / or to minimize the size of the encoded representation C (in the sense of the throughput distortion criterion).

[0054] When the management module 2 selects the artificial neural network N from a plurality of default artificial neural networks, the artificial neural network used for decoding is related to the selected artificial neural network N in a default manner. The information i can specify this network related to the selected artificial neural network N (e.g., within a list of artificial neural networks).

[0055] When the management module 2 obtains the artificial neural network N through a training step, this training step can enable the simultaneous training of the artificial neural network used for decoding. The information i can include the description data of the artificial neural network used for decoding (this description data is related to each neuron of this artificial neural network, for example, and can include the weights determined during the training step).

[0056] The stream formation module 6 receives the encoded representation C created by the encoding module 4 and the information i supplied by the management module 2, and constructs a data stream F based on these elements. The stream formation module 6 can, of course, actually receive other data from the encoding module 4 and / or the management module 6.

[0057] The stream formation module 6 constructs the data stream F in the form of various data packets intended to be continuously transmitted (e.g., transmitted) to the decoding device. These data packets are, for example, each a unit of the network abstraction layer (NAL unit).

[0058] The stream formation module 6 constructs various data packets according to the following description.

[0059] In the example described here, each data packet starts with a default marker M (i.e., formed by a default sequence of bits or a default pattern). Thus, in this case, each data packet starts with the same marker M, which makes it possible to identify the start of a data packet when receiving the stream. For example, it is proposed to prohibit the value corresponding to marker M (i.e., the sequence of bits forming marker M) within the data stream F except at the start of the data packet.

[0060] As a variant, other means for identifying data packets within the data stream F can be considered, such as a list enumerating the addresses of various data packets within the data stream F.

[0061] In this case, each data packet further includes a type identifier that specifies the type of the related data packet from a set of predetermined possible types.

[0062] In the example described below, at least some of the following types of data packets are used: - A data packet carrying the description data of an artificial neural network (used for decoding), specified by a first type identifier denoted as T1 below - A data packet carrying the encoded data representing audio or video content (i.e., the encoded representation C obtained by the encoding module 4 in this case), specified by a second type identifier denoted as T2 below - A data packet carrying parameters related to audio or video content or the decoding process used, specified by a third identifier denoted as T3 below - A data packet carrying encoded data representing audio or video content, but this encoded data is obtained by a different encoding process from the aforementioned encoded data included in the data packet of type T2 (thus requiring a different decoding process), and this last type of data packet is specified by a fourth identifier denoted as T4 below.

[0063] Accordingly, there are data packets containing the coded representation of the content (in the examples described herein, data packets of type T2 and T4), as well as data packets containing the description data of the artificial neural network (for example, data coded in a given format and representing this artificial neural network), here data packets of type T1, and data packets containing parameters (here data packets of type T3).

[0064] In the embodiments described herein, a data packet (data packet of type T1) containing the description data of the artificial neural network includes the following: - Marker M - A type identifier that assumes a value of T1 in this case (forming the first data of this data packet) - Optionally, an identifier NNI related to the relevant artificial neural network (this identifier NNI forms part of a set of predetermined identifiers individually related to various artificial neural networks) - Optionally, a format identifier NNF indicating the format of the description data NNC contained within the data packet - The description data NNC of the relevant artificial neural network (this description data NNC forms the second data of this data packet)

[0065] The T1 type identifier and thus the identifier NNI and / or the format identifier NNF can be included, for example, within the header of the data packet, and the description data NNC can form the payload data of the data packet for its part.

[0066] The description (or coding) format of the description data NNC (identified by the format identifier NNF if necessary) can be, for example, the NNR format (MPEG-7 part 17), the NNEF format, or the ONNX format. As a modified form, the format identifier NNF can also specify the format that can be accepted by tools for handling artificial neural networks, and further, the format of the artificial neural network identifier within a set of predetermined artificial neural networks (in that case, the data NNC includes such an identifier).

[0067] When the format used is agreed upon (predetermined) by the encoding device and the decoding device (or in other words, when a single format is used by the decoding device), the use of the format identifier NNF within the data packet is not necessary.

[0068] When the coded representation part C is created by the coding module 4 based on the data set B, the stream formation module 6 receives the information i indicating the decoding artificial neural network used to decode the part C (from the management module 2) as already shown. As will be clear from the examples shown below, in this way, the stream formation module 6 can determine based on this information i which description data NNC should be placed within a given type T1 data packet.

[0069] In an embodiment using the identifier NNI, it is assumed that all data packets (i.e., all type T1 data packets in this case) containing the description data of the artificial neural network including a given identifier NNI contain the same description data NNC.

[0070] In the examples described in this specification, the data packets (type T2 and T4 data packets) containing the coded representation of the content include the following: - Marker M - A type identifier (the type is one of the types available for the coded representation of the content) assuming a value of T2 or T4 in this case (forming the first data of this data packet) - Optionally, an identifier NNI related to an artificial neural network (in accordance with the same association rules as described above for data packets of type T1), where this artificial neural network is used to decode the encoded representation of the content included in this data packet - Optionally, a position identifier NNL such as a file pointer indicating the position within the data stream of the data packet that contains the description data of the artificial neural network used to decode the encoded representation of the content included in this data packet - Optionally, a remote description indicator DNN indicating whether the data packet containing the description data of the artificial neural network used is part of the data stream related to the current image sequence or part of the data stream related to an image sequence different from the current image sequence - The encoded representation of the relevant content C (which is part of the encoded representation generated by the encoding module 4 and forms the second data of this data packet).

[0071] The image sequence in this case is a set of images that can be obtained by decoding a part of the encoded representation of the content (here, video) without the need to access another part of the encoded representation of the content (here, video).

[0072] The type T2 and T4 identifiers, and optionally the identifier NNI and / or the position identifier NNL and / or the remote description indicator are included, for example, in the header of the data packet, and the encoded representation C can form the payload data of the data packet.

[0073] Among the data packets containing the coded representation of the content, some packets may have a specific type (in this case, the type corresponding to the identifier T4 in the example described later) for identifying the entry point within the data stream. In this case, for example, it can be stipulated that only the packets of this type (corresponding to the entry point) contain the identifier NNI or the location identifier NNL.

[0074] According to one possible modification, the identifier NNI and / or the location identifier NNL and / or the remote description indicator DNN may be included within the data packets (in the example described here, the data packets of type T3) carrying the parameters related to the image or the image sequence.

[0075] It can also be stipulated that the stream formation module 6 constructs the data stream F such that the data packets (data packets of type T2 or T4) containing the coded representation of the content and any data packet containing a given identifier NNI are preceded (within the data stream F) by a data packet of type T1 that also contains this given identifier NNI (and thus the description data of the artificial neural network specified by this given identifier NNI).

[0076] Next, various examples of possible data streams will be described with reference to FIGS. 4 to 9. The decoding of these possible data streams will be described below.

[0077] The first example of the data stream is shown in FIG. 4.

[0078] In this example, the data stream includes a data packet 12 of type T1 and a data packet 14 of type T2 (later in the data stream with respect to packet 12).

[0079] Data packet 12 includes marker M, an identifier of the type assuming the value of T1, an identifier NNI related to a given artificial neural network, a format identifier NNF indicating the format of description data NNC (described later), and this description data NNC of the given artificial neural network.

[0080] Data packet 14 includes marker M, an identifier of the type assuming the value of T2, an identifier NNI related to a given artificial neural network (the same as the identifier included in data packet 12), and a portion C of a coded representation generated by coding module 4 and decodable in particular using the given artificial neural network.

[0081] A second example of the data stream is shown in FIG. 5.

[0082] In this example, the data stream includes data packet 16, data packet 18, and data packet 20 in this order. (Other data packets may exist in the data stream between data packet 16 and data packet 18 and / or between data packet 18 and data packet 20).

[0083] Data packet 16 includes marker M, an identifier of the type assuming the value of T1, an identifier NNI related to a given artificial neural network, and description data NNC of the given artificial neural network.

[0084] Data packet 18 includes marker M, an identifier of the type assuming the value of T3 (corresponding to a data packet including parameters related to a given image or a given image sequence as described above), and an identifier NNI related to a given artificial neural network (among these parameters).

[0085] Data packet 20 includes marker M, an identifier of the type assuming the value of T2, and a portion C of a coded representation generated by coding module 4 and decodable in particular using the given artificial neural network.

[0086] A third example of a data stream is shown in FIG. 6.

[0087] In this example, the data stream includes data packet 22 and data packet 24 in the second half of the data stream.

[0088] Data packet 22 includes marker M, an identifier of the type assuming the value of T1, an identifier NNI related to a given artificial neural network, and description data NNC of the given artificial neural network.

[0089] Data packet 24 includes marker M, an identifier of the type assuming the value of T2, and a portion C of a coded representation generated by coding module 4 and decodable using a given artificial neural network in particular.

[0090] A fourth example of a data stream is shown in FIG. 7.

[0091] In this example, the data stream includes data packet 26, data packet 28, and data packet 30 in this order. (Other data packets may be present in the data stream between data packet 26 and data packet 28 and / or between data packet 28 and data packet 30).

[0092] Data packet 26 includes marker M, an identifier of the type assuming the value of T1, an identifier NNI related to a given artificial neural network, and description data NNC of the given artificial neural network.

[0093] Data packet 28 includes marker M, an identifier of the type assuming the value of T4 (corresponding to the entry point in the stream as described above), an identifier NNI related to a given artificial neural network (the same as the identifier included in data packet 26), and a portion C of a coded representation generated by coding module 4 and decodable using a given artificial neural network in particular.

[0094] The data packet 30 includes a marker M, an identifier of the type that assumes the value of T2, and another part C' of the coded representation generated by the coding module 4, which can be decoded using a given artificial neural network in particular.

[0095] A fifth example of a data stream is shown in FIG. 8.

[0096] In this example, the data stream includes data packets 32, 34, 36, and 38 in this order. (Other data packets may be present in the data stream between these various packets).

[0097] The data packet 32 includes a marker M, an identifier of the type that assumes the value of T4 (corresponding to the entry point in the stream as described above), a remote description indicator DNN, a location identifier NNL, and a part C of the coded representation generated by the coding module 4.

[0098] The data packet 34 includes a marker M, an identifier of the type that assumes the value of T2, and another part C' of the coded representation generated by the coding module 4.

[0099] The data packets 32 and 34 are related to the same image sequence S, that is, the parts C and C' of the coded representation form part of a coded data set, enabling a set of images to be decoded without relying on coded data located outside this coded data set.

[0100] The data packet 36 includes a marker M, an identifier of the type that assumes the value of T1, and the description data NNC of the artificial neural network.

[0101] The data packet 38 includes a marker M, an identifier of the type that assumes the value of T2, and a part C'' of the coded representation generated by the coding module 4.

[0102] Data packets 36 and 38 relate to the same image sequence S', which is different from the image sequence S.

[0103] The remote description indicator DNN included in data packet 32 indicates that the artificial neural network used to decode the coded representation part C (and part C') is not described by the description data included in the image sequence S, but is described outside this image sequence S (here, within the image sequence S').

[0104] Therefore, data packet 32 includes the aforementioned position identifier NNL, which here is a pointer to data packet 36 (located within the image sequence S').

[0105] Such a pointer can be, for example: - The difference in the number of bytes relative to the position of data packet 32 (this difference can be signed, i.e., it can have a positive value indicating the number of bytes in a certain direction within the data stream or a negative value indicating the number of bytes in the opposite direction) - The difference in the number of bytes relative to the start of the file (or, in other words, the data stream) - A physical memory storage address (rewritten by the decoding device when processing the data stream in the case of storing an artificial neural network, and then rewritten in all references NNL to this artificial neural network during the preprocessing step) can be.

[0106] A sixth example of a data stream is shown in FIG. 9.

[0107] In this example, the data stream includes data packet 40, data packet 42, data packet 44, and data packet 46 in this order. (Other data packets may exist within the data stream between these various packets).

[0108] Data packet 40 includes marker M, an identifier of the type assuming the value of T1, an identifier NNI related to a given artificial neural network, and description data NNC of the given artificial neural network.

[0109] Data packet 42 includes marker M, an identifier of the type assuming the value of T2, identifier NNI (the same as that included in data packet 40), and a portion C of the encoded representation generated by encoding module 4 and decodable in particular using a given artificial neural network.

[0110] Data packet 44 includes marker M, an identifier of the type assuming the value of T1, an identifier NNI related to a given artificial neural network, and description data NNC of the given artificial neural network (this data NNC is the same as the data NNC included in data packet 40).

[0111] Data packet 46 includes marker M, an identifier of the type assuming the value of T2, identifier NNI (the same as that included in data packets 40, 42, and 44), and another portion C' of the encoded representation generated by encoding module 4 and decodable in particular using a given artificial neural network.

[0112] Using another data packet 44 of type T1 that includes description data NNC of the artificial neural network identified by identifier NNI enables, for example, the decoder to read the data stream in an order other than that shown in FIG. 9 (random access) in order to start reading the data stream at a position other than that of data packet 40. Accordingly, other data packets identical to data packets 40, 44 may be present, for example, at regular intervals within the data stream.

[0113] In the example described in this specification, the data stream F constructed by the data formation module 6 is transmitted on a communication channel by the stream emission module 8 (optionally after other processing steps, such as an entropy encoding step).

[0114] As a modified form, the data stream F can be stored (e.g., on a recording device such as a hard disk of an encoding device) for subsequent reading and decoding (in this case, for example, the encoding device and the decoding device described later are the same electronic device).

[0115] FIG. 10 shows a decoding device according to the present invention.

[0116] This decoding device includes a stream reception module 50, a stream analysis module 52, a decoding module 54, and a configuration module 56.

[0117] (In this example, for the processor to execute a part of the aforementioned instructions) each of these modules can actually be implemented by a programmed processor for implementing the functions described later for the relevant modules (e.g., by instructions stored in a memory related to the processor). Further, in practice, for example, by executing several sets of instructions corresponding to various modules respectively (by this processor), several modules can be implemented by the same processor. As a modified form, any of the modules can be implemented by an application-specific integrated circuit.

[0118] The stream reception module 50 receives a data stream such as the data stream F constructed by the stream formation module 6 and emitted by the emission module 8 (e.g., via a communication channel).

[0119] According to the above-described modified form, this data stream is read onto a recording medium such as a hard disk.

[0120] The data stream F (received by the stream receiving module 50 or read onto a recording medium) is analyzed by the stream analysis module 52 as described below, whereby on the one hand data C, C', C'' that form part of the coded representation of the content, and on the other hand the artificial neural network used to decode this data C, C', C'' can be identified.

[0121] Next, as described below in the context of various examples, the configuration module 56 is designed to configure the decoding module 54 so that the decoding module decodes data C, C', C'' using the identified artificial neural network to create data B' representing audio or video content.

[0122] According to a first possible embodiment of the decoding module 54 shown in FIG. 11, data C, C', C'' (the coded representation of the content) is applied as input to the identified artificial neural network N'', and this artificial neural network generates data B' representing audio or video content as output.

[0123] The data B' created as output from the artificial neural network N'' corresponds to the data B applied as input to the artificial neural network N, and thus can represent a block of an image, or a block of components of an image (e.g., a block of the luminance component or chrominance component of this image or a block of the color components of this image), or an image of a video sequence, or components of an image of a video sequence (e.g., a luminance component or chrominance component or color component), or even a series of images of a video sequence.

[0124] In this case, at least some of the neurons in the output layer of the artificial neural network N'' each create pixel values of components of the image, and those values form one of the data items B'.

[0125] According to a second possible embodiment, the decoding module 54 processes data C, C', C'' (C in FIG. 12) in several stepsj is processed (represented as such), and at least one of the steps is performed by an artificial neural network N'.

[0126] Thus, for example, as shown in FIG. 12, a portion C of the coded representation previously received or read within the data stream j-1 is applied as an input to the artificial neural network N', which enables the generation of prediction data P as an output from the artificial neural network N', j and the prediction data is combined with the current portion C of the coded representation (e.g., by addition) to obtain a portion B' of the data representing the audio or video content j (as the output from the decoding module 4). j It should be noted that the artificial neural network N' used for decoding (shown in FIG. 12) is the same as the artificial neural network N' used for coding (see FIG. 3 above) in this case.

[0127] In FIG. 12, reference numeral 60 represents a delay module for indicating the fact that when processing the current portion C of the coded representation to obtain the corresponding portion B' of the representative data, it is the previously received (or read) portion C of the coded representation that is applied as an input to the artificial neural network N'.

[0128] In fact, as already shown for coding, the portion C j is related to a portion B' representing an image preceding the image represented by, for example, the portion B' j represented by the portion B'. j-1 As a variant, the previously received or read portion used as an input to the artificial neural network N' is the data B'

[0129] j-1 j j-1

[0130] j ​​​​​It can be a coded representation part corresponding to at least one block of an adjacent image of a block represented by the pixel value thereby.

[0131] FIG. 13 shows steps of an example of a method for decoding a data stream F.

[0132] This decoding method can be used, among other things, for the examples of data streams described above and shown in FIGS. 4 and 9. For this reason, the description of this decoding method is shown using the reference numbers of the numbers mentioned in FIGS. 4 and 9.

[0133] This method starts from step E2 where the analysis module 52 identifies the start of data packets 12, 14, 40, 42, 44, 46 (within the data stream F) by means of a marker M from which no data packet starts in this case.

[0134] Once the start of a data packet is identified, the analysis module 54 can identify the type of the data packet (step E4) by reading (and optionally decoding) the T type identifier (or first data) of this data packet in the header of this data packet, for example.

[0135] Next, the analysis module 54 determines in step E6 whether the type identified by the T type identifier is a predetermined type (corresponding to the T1 type in this case). As already shown, this predetermined type (T1 in this case) is related to data packets containing data indicating an artificial neural network.

[0136] In the case of an affirmative determination (arrow P) in step E6, the method continues to step E8. (This is especially the case when processing data packets 12, 40, 44).

[0137] In the case of a negative determination (arrow N) in step E6, the method continues to step E16 (this is especially the case when processing data packets 14, 42, 46).

[0138] In step E8, the analysis module 52 reads (and optionally decodes) the identifier NNI in the data stream F, which specifies a particular artificial neural network (the identifier NNI forms part of a set of predetermined identifiers individually associated with various artificial neural networks).

[0139] Then, in step E10 (optionally in cooperation with the configuration module 56), the analysis module 52 determines whether the artificial neural network specified by the identifier NNI is stored in the decoder, for example, after prior reception of a data packet (such as data packet 40) that already contains data indicating the artificial neural network.

[0140] In the case of an affirmative determination (arrow P) in step E10 (such as when processing data packet 44 if data packet 40 has been previously processed), the previously received and stored artificial neural network can be reused (when later transitioning to step E22 described below), so there is no need to continue processing the current data packet, and the method loops back to step E2.

[0141] However, in the case of a negative determination (arrow N) in step E10 (such as when processing data packet 12 or 40), the method reads the data NNC in the data stream F and, optionally taking into account the encoded form of this indicated data NNC as required by the format identifier NNF (in the case of data packet 12), to obtain (e.g., construct) an artificial neural network, and continues to step E12 to decode the aforementioned data (the second data of the current data packet) indicating the artificial neural network associated with the identifier NNI.

[0142] As already shown, particularly in the examples of FIGS. 4 and 9, the instruction data NNC is description data of an artificial neural network that can be decoded (by the stream analysis module 52 or the configuration module 56) to obtain the parameters of the artificial neural network, and those parameters enable the configuration module 56 to configure the decoding module 54, thereby enabling this decoding module 54 to implement an artificial neural network (specified by the identifier NNI) in particular.

[0143] The parameters of the artificial neural network obtained by decoding the description data NNC are then stored in the memory of the decoding device (for example, the memory associated with the configuration module 56) in step E14, and the method loops back to step E2 to process a new data packet of the data stream F.

[0144] If it is determined in step E6 that the type of the current data packet does not correspond to the predetermined type T1, the method continues to step E16 as described next as already shown.

[0145] In step E16, the stream analysis module 52 determines whether the type T specified by the type identifier (or the first data) of the current data packets 12, 14, 40, 42, 44, 46 forms part of the type related to data packets (corresponding to types T2 and T4 in this case) that contain a coded representation of the content.

[0146] In the case of a negative determination (arrow N) in step E16, the method continues to step E18 to decode the data packet. This is the case, for example, when the data packet contains parameters related to an image or an image sequence (data packet of type T3 in the example described in this case), and decoding the current data packet enables obtaining parameters related to the image or the current image sequence (during decoding) in this case.

[0147] The method then loops back to step E2 to process another data packet.

[0148] In the case of a positive determination (arrow P) at step E16, the method continues to step E20 where it reads (and optionally decodes) the identifier NNI within data stream F. This identifier NNI specifies an artificial neural network that is used to decode the encoded representations C, C' contained within the current packets 14, 42, 46.

[0149] The parameters that define this artificial neural network were previously obtained from data indicating this artificial neural network contained within a previously received data packet of type T1. In the example described in this case, these parameters were previously decoded based on the descriptive data NNC contained within a previously received data packet of type T1 (data packet 12 in the case of FIG. 4, data packet 40 in the case of FIG. 9, or data packet 44 if data packet 40 has not been read by the decoding device).

[0150] The thus obtained (e.g., decoded) parameters are also stored in the memory of the decoding device (the memory associated with the configuration device 56 in this case) as already explained.

[0151] The configuration module 56 can then configure the decoding module 54 with the parameters of the artificial neural network specified by the identifier NNI read within the data stream at step E20, whereby the decoding module 54 can use this artificial neural network to decode the encoded representation.

[0152] Next, the stream analysis module 52 extracts the encoded representations C, C' (second data) from the current packet and transmits them to a decoding module 54 for decoding the encoded representations C, C' using the artificial neural network (designated by the identifier NNI read in step E20) to obtain data B' representing the audio or video content (this data is, for example, the pixel values of at least a part of an image or the components of an image), which is the output from the decoding module 54.

[0153] Figure 14 shows the steps of a method that can be considered for decoding the data stream of Figure 5.

[0154] The method includes step E30 of identifying and analyzing data packet 16 using the stream analysis module 52.

[0155] This step includes, in this case, identifying the start of data packet 16 by marker M, detecting a type identifier (first data) corresponding to a predetermined type T1, and reading an identifier NNI related to the artificial neural network and description data NNC (second data) corresponding to the identifier NNI within the data stream.

[0156] Next, the method may include step E32 of decoding the data NNC to obtain the parameters of the artificial neural network and storing the obtained parameters in the memory of the decoding device (for example, the memory related to the configuration module 56).

[0157] Thereafter, during step E34, the stream analysis module 52 identifies and analyzes data packet 18.

[0158] Step E34 includes, in this case, identifying the start of data packet 18 by marker M, detecting a type identifier indicating a T3 type corresponding to a data packet containing parameters related to at least one image of the current image sequence, and reading the parameters included in data packet 18 (within the data stream), and those parameters include identifier NNI in this case.

[0159] According to one possible embodiment, the parameters included in data packet 18 of type T3 may relate only to the decoded image (i.e., the image from which the representative data is obtained by the next decoding operation by decoding module 54).

[0160] The presence of identifier NNI in data packet 18 indicates that, in this case, an artificial neural network related to identifier NNI is used to decode representative data C related to the decoded image (to obtain data B' related to at least a part of the decoded image).

[0161] According to another possible embodiment, the parameters included in data packet 18 of type T3 may relate to all images of the decoded image sequence.

[0162] The presence of identifier NNI in data packet 18 indicates that, in this case, an artificial neural network related to identifier NNI is used to decode representative data C related to various images of the current image sequence (to obtain data B' related to at least a part of one of the images of the current image sequence).

[0163] During step E34, according to one possible embodiment, the configuration module 54 can configure the decoding module 56 so that the decoding module 56 can decode data representing the data received in the data stream by using the artificial neural network specified by the identifier NNI. This configuration can actually be performed by reading the parameters of the artificial neural network in the aforementioned memory of the encoding device (see step E32 below).

[0164] Then during step E36, the stream analysis module 52 identifies and analyzes the data packet 20. In this case, this data packet 20 is considered to be related to an image to which the parameters included in the aforementioned data packet 18 are applied.

[0165] Step E36 includes, in this case, identifying the start of the data packet 20 by the marker M, detecting a type identifier indicating type T2 corresponding to the data packet including the encoded representation of the video in this case, and reading a portion C of the encoded representation (the second data of the data packet 20).

[0166] Then the method includes step E38 of using the artificial neural network specified by the identifier NNI and decoding the portion C of the encoded representation by the decoding module 54.

[0167] Figure 15 shows the steps of a method that can be considered for decoding the data stream of Figure 6.

[0168] The method includes step E40 of using the stream analysis module 52 to identify and analyze the data packet 22.

[0169] This step includes, in this case, identifying the start of the data packet 22 by the marker M, detecting a type identifier (first data) corresponding to a predetermined type T1, and reading the description data NNC of the artificial neural network (second data).

[0170] Step E40 may also include reading and / or decrypting an identifier NNI associated with this artificial neural network. As described in other embodiments, this eliminates the need to decrypt the description data NNC of the artificial neural network again when a data packet of type T1 containing the same identifier NNI is detected later.

[0171] The method then includes step E42 of decrypting the data NNC to obtain the parameters of the artificial neural network and storing the obtained parameters in the memory of the decrypting device (e.g., the memory associated with configuration module 56). As shown above, in embodiments where the identifier NNI is used within data packet 22 and the data packet of type T1 carrying this identifier NNI has already been processed, this step can be omitted.

[0172] During step E42, the configuration module 54 can configure the decrypting module 56 (with the parameters obtained as described above) so that the decrypting module 54 can decrypt subsequent representative data received in the data stream using the artificial neural network.

[0173] The method then includes step E44 of identifying and analyzing data packet 24 (by stream analysis module 52). In this case, it is considered that no data packet of type T1 is included between data packet 22 and data packet 24.

[0174] Step E44 includes, in this case, identifying the start of data packet 24 by marker M, detecting a type identifier indicating type T2 corresponding to a data packet containing the coded representation of the video in this case, and reading part C of the coded representation (the second data of data packet 24) (within the data stream).

[0175] The method then includes step E46 of using an artificial neural network represented by data NNC included in data packet 22 to decode portion C of the encoded representation by decoder module 54.

[0176] Thus, in this embodiment, the artificial neural network used to decode the encoded representation included in data packet 24 is defined within the last data packet of type T1 preceding this data packet 24 (in the indication data included therein, in this case by the description data NNC).

[0177] FIG. 16 shows the steps of a possible method for decoding the data stream of FIG. 7.

[0178] The method includes step E50 of using stream analysis module 52 to identify and analyze data packet 26.

[0179] This step includes, in this case, identifying the start of data packet 26 by marker M, detecting a type identifier (first data) corresponding to a predetermined type T1, and reading identifier NNI within the stream related to the artificial neural network and description data NNC (second data) of the artificial neural network corresponding to identifier NNI.

[0180] The method then includes step E52 of decoding data NNC to obtain the parameters of the artificial neural network and storing the obtained parameters in the memory of the decoding device (e.g., the memory related to configuration module 56).

[0181] Thereafter, during step E54, stream analysis module 52 identifies and analyzes data packet 28.

[0182] Step E54 includes, in this case, identifying the start of data packet 28 by marker M, detecting a type identifier indicating type T4 corresponding to a data packet that includes the coded representation of the content and identifies an entry point within the data stream, and reading (within the data stream) identifier NNI and the first part C of the coded representation of the content.

[0183] The method may then continue to step E56 of decoding this first part C of the coded representation of the content by decoding module 54 and by an artificial neural network associated with this identifier NNI (included within data packet 28). To do this, step E56 may optionally include the step of configuring decoding module 54 using configuration module 56 in fact and by the parameters obtained (and stored) in step E52.

[0184] Thereafter, during step E58, stream analysis module 52 identifies and analyzes data packet 30.

[0185] Step E58 includes, in this case, identifying the start of data packet 30 by marker M, detecting a type identifier indicating type T2 corresponding to a data packet that includes the coded representation of the content, and reading (within the data stream) the second part C' of the coded representation of the content.

[0186] In fact, as already shown, in this embodiment, only data packets corresponding to possible entry points within the data stream are those that include the identifier (NNI in this case) of the neural network used to decode the coded representation of the content.

[0187] The method may then continue to step E60 of decoding this second portion C' of the coded representation of the content using the decoding module 54 and by means of the artificial neural network associated with the identifier NNI included in this data packet 28 and designated as a possible entry point by the T4 type identifier included in the data packet 28.

[0188] Figure 17 shows the steps of a method that may be considered for decoding the data stream of Figure 8.

[0189] The method includes step E70 of identifying and analyzing the data packet 32 using the stream analysis module 52.

[0190] Step E70 includes, in this case, identifying the start of the data packet 32 by the marker M, detecting in this case a type identifier indicating the T4 type corresponding to the data packet containing the coded representation of the content and identifying an entry point in the data stream, and reading (in the data stream) the remote description indicator DNN, the position identifier NNL, and the first portion C of the coded representation of the content.

[0191] In fact, assume that in this case the remote description indicator DNN assumes a value (for example, a value of 1) indicating that the artificial neural network used to decode the first portion C is described outside the current sequence S. As a result, the stream analysis module 52 reads the position identifier NNL located behind (immediately in this case) the remote description indicator DNN in the data stream.

[0192] Next, the stream analysis module 52 browses the data stream according to the indication given by the position identifier NNL (for example, by browsing the byte difference indicated by the position identifier NNL or by jumping to the physical memory address indicated by the position identifier NNL) until it reads and analyzes the data packet 36 (step E72).

[0193] This step E72 includes identifying the start of data packet 36 by marker M in this case, detecting a type identifier (first data) corresponding to a predetermined T1 type, and reading the description data NNC of the artificial neural network (second data) in the data stream.

[0194] Next, the method includes step E74 of decrypting the data NNC to obtain the parameters of the artificial neural network and storing the obtained parameters in the memory of the decrypting device (for example, the memory associated with configuration module 56).

[0195] Next, the method may continue to step E76 of decrypting the first part C of the coded representation of the content using the decrypting module 54 and by the above-mentioned artificial neural network (decrypted from the description data NNC included in data packet 36). To do this, step E76 may optionally include the step of configuring the decrypting module 54 using the actually used configuration module 56 and the parameters obtained (and stored) in step E74.

[0196] Then, during step E78, the stream analysis module 52 identifies and analyzes data packet 34.

[0197] Step E78 includes identifying the start of data packet 34 by marker M in this case, detecting a type identifier indicating the T2 type corresponding to the data packet including the coded representation of the content in this case, and reading the second part C' of the coded representation of the content (in the data stream).

[0198] In fact, as already shown, in this embodiment, only the data packets corresponding to the possible entry points in the data stream are assumed to include the identifier of the neural network (NNI in this case) used to decrypt the coded representation of the content.

[0199] The method can then continue to step E80 of decoding this second part C' of the coded representation of the content using the decoding module 54 and by means of the artificial neural network obtained as described above in steps E74 and E76 by the position identifier NNL included in the data packet 32 of type T4 preceding the current data packet 34.

[0200] (When decoding sequence S') Any subsequent processing of data packet 38 is not described here.

Claims

1. A method for decoding a data stream including data packets each including at least first data and second data, comprising: - identifying (E4, E6, E30, E40, E50, E72) a first data packet (12, 16, 22, 26, 36, 40) among the data packets, wherein the first data of the first data packet includes information (T1) indicating a data packet of a predetermined type; - processing (E12, E32, E42, E52, E74) the second data (NNC) of the first data packet to obtain an artificial neural network; - decoding (E22, E38, E46, E56, E60, E76, E80) the second data (C, C') included in a second data packet (14, 20, 24, 28, 30, 32, 34, 42, 46) among the data packets by using at least the obtained artificial neural network, thereby creating data representing audio or video content. A decoding method, characterized by including the above steps.

2. The decoding method according to claim 1, wherein the second data of the first data packet includes description data (NNC) of the artificial neural network, and the processing step is a step of decoding the description data (NNC) to obtain parameters of the artificial neural network.

3. The decoding method according to claim 1 or 2, wherein the first data packet includes an identifier (NNI) of the artificial neural network.

4. The decoding method according to claim 3, wherein the identifier (NNI) is an element of a list of distinct identifiers each individually associated with a distinct artificial neural network.

5. Receiving (E70) a third data packet (44) whose first data includes information (T1) indicating the packet of the predetermined type and includes the identifier (NNI); and reusing the obtained artificial neural network to decode second data (C') included in a fourth data packet (46) among the data packets. The decoding method according to claim 4, characterized by including the above steps.

6. The decoding method according to any one of claims 3 to 5, wherein the second data packet includes the identifier (NNI).

7. Receiving another data packet (18) including parameters related to at least one image of the content, the parameters including the identifier (NNI), the decoding method according to any one of claims 4 to 6.

8. The first data packet (22) is the last packet preceding the second data packet (24) in the data stream among the data packets including information (T1) indicating that the first data thereof is a packet of the predetermined type, the decoding method according to any one of claims 1 to 5.

9. The decoding method according to any one of claims 1 to 3, wherein the second data packet (32) includes a pointer (NNL) to the first data packet (36).

10. - Reading a flag (DNN) in the second data packet (32); - When the flag has a predetermined value, reading a pointer (NNL) to the first data packet (36) in the second data packet (32); The decoding method according to any one of claims 1 to 3, comprising:

11. - Receiving another data packet including parameters related to at least one image of the content; - Reading a flag from the parameters; - When the flag has a predetermined value, reading a pointer to the first data packet from the parameters; The decoding method according to any one of claims 1 to 3, comprising:

12. The decoding method according to any one of claims 9 to 11, wherein the pointer (NNL) designates a position in a part of the data stream related to an image sequence (S') different from the image sequence (S) at least partially encoded by the second data (C) of the second data packet (32).

13. The decoding method according to any one of claims 1 to 11, wherein the first data packet includes information (NNF) indicating the coding format of the second data (NNC) of the first data packet.

14. An apparatus for decoding a data stream including data packets each including at least first data and second data, - Identifying, from among the data packets, a first data packet (12, 16, 22, 26, 36, 40) whose first data includes information (T1) indicating a data packet of a predetermined type; - Processing the second data (NNC) of the first data packet to obtain an artificial neural network; - Decoding the second data (C, C') included in a second data packet (14, 20, 24, 28, 30, 32, 34, 42, 46) among the data packets by using at least the obtained artificial neural network, thereby creating data representing audio or video content The apparatus characterized by including a processor configured or programmed to perform the above.

15. A computer program including instructions executable by a processor and designed to implement the method according to any one of Claims 1 to 13 when executed by the processor.

16. A data stream including data packets each including at least first data and second data, wherein the data packets - A first data packet (12, 16, 22, 26, 36, 40) whose first data includes information (T1) indicating a data packet of a predetermined type and whose second data (NNC) defines an artificial neural network; and - A second data packet (14, 20, 24, 28, 30, 32, 34, 42, 46) whose second data (C, C') can be decoded using at least the artificial neural network so as to create data representing audio or video content The data stream characterized by including the above.

17. The data stream according to Claim 16, wherein the first data packet (40) includes an identifier (NNI), the data stream includes another data packet (44) including the identifier (NNI), the first data of which includes the information (T1) indicating the data packet of the predetermined type, and the second data of which is the same as the second data (NNC) of the first data packet (40).

Citation Information

Patent Citations

  • Method for decoding a data stream, associated device and associated data stream

    WO2022013249A2