Method and decoding device for decoding at least part of data stream
By introducing indicators and identifiers into the data stream, the problem of decoders struggling to determine artificial neural network sets in existing technologies is solved, enabling efficient and flexible video content decoding, supporting random access and low-latency decoding, and adapting to devices with different computing capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-13
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, video content decoding methods struggle to effectively utilize artificial neural networks for efficient decoding, particularly in determining whether the decoder belongs to a predetermined set and obtaining the corresponding neural network parameters.
By introducing indicators and identifiers into the data stream, the decoder can determine whether an artificial neural network belongs to a predetermined set and retrieve the corresponding neural network parameters from local storage or a remote server based on the type of indicator, thereby decoding the video content.
It improves the efficiency and flexibility of video content decoding, supports random access and low latency decoding, adapts to devices with different computing capabilities, and simplifies the decoding process.
Smart Images

Figure CN121814947A_ABST
Abstract
Description
[0001] This patent application is a divisional application of the following invention patent application:
[0002] Application Number: 202180050577.1
[0003] Application date: July 13, 2021
[0004] Invention Title: Video Decoding Using Neural Networks Technical Field
[0005] This invention relates to the field of audiovisual content decoding technology.
[0006] Specifically, it relates to a method for decoding a data stream, as well as the associated device and data stream. Background Technology
[0007] It has been proposed to use artificial neural networks to compress data representing video content. Then, another artificial neural network can be used to perform the decoding of the compressed data, for example, in Guo Lu et al.'s paper "DVC: An End-to-end DeepVideo Compression Framework", IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 10998-11007. Summary of the Invention
[0008] In this context, the present invention proposes a method for decoding a data stream including indicators and data representing audio or video content, the method comprising the following steps:
[0009] - Decode the indicator to determine whether the artificial neural network used to decode the representative data is encoded in the data stream or belongs to a predetermined set of artificial neural networks;
[0010] - The representative data is decoded using an artificial neural network.
[0011] Such an indicator thus allows the decoder to know that it will be able to access the artificial neural network methods used to decode the data representing the content.
[0012] If the decoding method determines that the artificial neural network belongs to the predetermined set by decoding an indicator, then the decoding method may include an identifier for decoding the neural network.
[0013] This invention also proposes, in its own original form, a method for decoding a data stream comprising data representing audio or video content and an identifier of an artificial neural network from a predetermined set of artificial neural networks. The method includes the following steps:
[0014] -Decoding identifier;
[0015] - The representative data is decoded using an artificial neural network specified by a decoded identifier.
[0016] The decoding of this identifier thus indicates which artificial neural network must be used from a predetermined set of artificial neural networks, such as a predetermined set of artificial neural networks that can be accessed by an electronic decoding device, as explained below.
[0017] According to the first possibility, the decoding method may include reading parameters of the artificial neural network identified by the decoded identifier from the storage unit.
[0018] It can be provided that the storage unit stores a first set of parameters representing a first artificial neural network forming a random access decoder and / or a second set of parameters representing a second artificial neural network forming a low latency decoder.
[0019] Furthermore, the decoding method may include the step of generating an error message in the absence of data associated with the artificial neural network identified by the decoded identifier (e.g., within a storage unit such as the aforementioned storage unit).
[0020] According to the second possibility, the decoding method may include receiving parameters of an artificial neural network identified by the decoded identifier from a remote server.
[0021] In addition, the decoding method may include: if it is determined by a decoding indicator that the artificial neural network is encoded in the data stream, then decoding the data of the artificial neural network encoded in the data stream to obtain the parameters of the artificial neural network.
[0022] According to possible embodiments, the decoding method may include a (preliminary) step of transmitting a list of artificial neural networks to a device for controlling the transmission of a data stream. In some embodiments, the list of artificial neural networks may correspond to the predetermined set described above. In other words, in this case, the predetermined set of artificial neural networks may include (or be formed from) the artificial neural networks in the list.
[0023] This content can actually be the first part of a video sequence, which can include both the first and second parts.
[0024] In this case, the step of decoding the representative data that generates the first part may further include the step of decoding other data by means of another artificial neural network to generate the second part.
[0025] According to a possible embodiment, another artificial neural network has the same structure as the said artificial neural network, which simplifies the updating of the artificial neural network within the electronic decoding device.
[0026] The first part and the second part mentioned above each form, for example, two sets of images used to represent the format of the content used.
[0027] The present invention also proposes a decoding device, comprising:
[0028] - A unit for receiving a data stream, the data stream including an indicator and data representing audio or video content;
[0029] - A decoding component designed to determine, by decoding the indicator, whether the artificial neural network used to decode the representative data belongs to a predetermined set of artificial neural networks or is encoded in the data stream, and to decode the representative data by means of the artificial neural network.
[0030] The present invention also proposes a decoding device, comprising:
[0031] - A unit for receiving a data stream, the data stream including data representing audio or video content, and an identifier of an artificial neural network in a predetermined set of artificial neural networks;
[0032] - A decoding component designed to decode identifiers and, with the aid of an artificial neural network specified by the decoded identifiers, decode the representative data.
[0033] In the embodiments described below, such a decoding component includes a processor designed or programmed to decode indicators and / or identifiers, and / or a parallel processing unit designed to perform multiple operations of the same type in parallel at a given time and implement the artificial neural network described above to decode the representative data described above.
[0034] The present invention also proposes a data stream comprising data representing audio or video content, and an indicator indicating whether an artificial neural network used to decode the representative data is encoded in the data stream or belongs to a predetermined set of artificial neural networks.
[0035] Finally, the present invention proposes a data stream comprising data representing audio or video content, and an identifier of an artificial neural network within a predetermined set of artificial neural networks for decoding the representative data.
[0036] The present invention also proposes a method implemented by a decoding device for decoding at least a portion of a data stream, the at least portion comprising an indicator (IND; IND') and representative data (Fnn) representing at least one block of an image or image component, the method comprising the steps of: - decoding (E60) the indicator (IND; IND'); - if the indicator (IND; IND') indicates that an artificial neural network (18) for decoding the representative data (Fnn) is encoded in the at least portion of the data stream: then decoding the parameters of the artificial neural network (18) from the at least portion of the data stream, or, - if the indicator (IND; IND') indicates that the artificial neural network (18) for decoding the representative data (Fnn) is accessible by the decoding device: then reading the parameters of the artificial neural network (18) stored locally, or requesting the parameters of the artificial neural network (18) from a remote server; - decoding (E70) the representative data (Fnn) block by block by means of the artificial neural network (18).
[0037] The present invention also proposes a decoding device, comprising: - a unit (11) for receiving at least a portion of a data stream, the at least a portion including an indicator (IND; IND') and representative data (Fnn) representing at least one block of an image or image component; - a decoding component (14, 16) configured to: decode the indicator (IND; IND'), if the indicator (IND; IND') indicates that an artificial neural network (18) for decoding the representative data (Fnn) is encoded in the at least a portion of the data stream: then decode the parameters of the artificial neural network (18) from the at least a portion of the data stream, and if the indicator (IND; IND') indicates that the artificial neural network (18) for decoding the representative data (Fnn) is accessible to the decoding device: then read the parameters of the artificial neural network (18) stored locally, or request the parameters of the artificial neural network (18) from a remote server, and decode the representative data block by block by means of the artificial neural network (18).
[0038] Of course, different features, alternatives and embodiments of the present invention can be associated with each other in various combinations, as long as they are not incompatible or mutually exclusive. Attached Figure Description
[0039] Furthermore, various other features of the invention will be apparent from the accompanying description, which shows non-limiting embodiments of the invention, and wherein:
[0040] - Figure 1 An electronic coding device used within the framework of this invention is shown;
[0041] - Figure 2 Shown in Figure 1 A flowchart of the steps of the encoding method implemented within the electronic encoding device;
[0042] - Figure 3 Through Figure 2 The first example of a data stream obtained by this method;
[0043] - Figure 4 Through Figure 2 A second example of a data stream obtained by this method;
[0044] - Figure 5 Through Figure 2 The third example of a data stream obtained by this method;
[0045] - Figure 6 Through Figure 2 The fourth example of a data stream obtained by this method;
[0046] - Figure 7 An electronic coding device according to an embodiment of the present invention is shown; and
[0047] - Figure 8 It is shown in Figure 7 A flowchart of the steps of the decoding method implemented within the electronic decoding device. Detailed Implementation
[0048] Figure 1 An electronic coding device 2 using at least one artificial neural network 8 is shown.
[0049] The electronic coding device 2 includes a processor 4 (e.g., a microprocessor) and a parallelization processing unit 6, such as a graphics processing unit or GPU, or a tensor processing unit or TPU.
[0050] like Figure 1 The diagram illustrates that processor 4 receives data P and B representing audio or video content to be compressed, where format data P is formatted data and content data B is content data.
[0051] The format data P indicates the characteristics of the representation format of the audio or video content, such as, for video content, image size (in pixels), frame rate, binary depth of luminance information, and binary depth of chrominance information.
[0052] Content data B forms a representation of the audio or video content (uncompressed). For example, in the case of video content, for each pixel of each image in an image sequence, the content data includes data representing the pixel's luminance value and / or data representing the pixel's chrominance value.
[0053] The parallel processing unit 6 is designed to implement the artificial neural network 8 after it has been configured by the processor 4. To this end, the parallel processing unit is designed to execute multiple operations of the same type in parallel at a given time.
[0054] As described below, an artificial neural network 8 is used within the framework of processing content data B to obtain compressed data C.
[0055] In the embodiments described herein, when content data B is applied to the input of artificial neural network 8, artificial neural network 8 produces compressed data C at the output.
[0056] The content data B applied to the input of the artificial neural network 8 (that is, the input layer of the artificial neural network 8) can represent a block of an image, or a block of image components (e.g., a block of the luminance or chrominance components of the image, or a block of color components), or an image of a video sequence, or a component of an image of a video sequence (e.g., a luminance or chrominance component, or a color component), or also a series of images of a video sequence.
[0057] For example, in this case, the pixel value of the image component received by each neuron in at least some neurons of the input layer can be provided, the value being represented by a data in the content data B.
[0058] Alternatively, the processing of content data B may include the use of several artificial neural networks, as described in the aforementioned paper "DVC: An End-to-end Deep Video Compression Framework" by Guo Lu et al., published at the 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) in June 2019.
[0059] Now refer to Figure 2 An example of an encoding method implemented by electronic encoding device 2 is described.
[0060] The memory connected to processor 4 stores, for example, computer program instructions that are designed to be executed by processor 4. Figure 2 At least some steps of the method. In other words, processor 4 is programmed to implement... Figure 2 At least some of the steps.
[0061] Here Figure 2 The method begins with an optional step E2, receiving a list of artificial neural networks accessible by an electronic decoding device.
[0062] This list is, for example, generated directly from the electronic decoding device (e.g., by the processor 4 of the electronic encoding device 2) by the electronic decoding device (as described below). Figure 7 (The electronic decoding device 10 is consistent with the receiver).
[0063] As described below, an artificial neural network accessible by an electronic decoding device is an artificial neural network for which the electronic decoding device stores parameters defining the artificial neural network of interest, or which can be accessed by a remote electronic device connected to a server.
[0064] Alternatively, the list can be received by the processor 4 of the electronic device 2 from a remote server such as the server mentioned above.
[0065] Figure 2 The method continues with step E4, which involves selecting the encoding-decoding pair. As already indicated for the encoding process, both the encoding and decoding processes utilize at least one artificial neural network.
[0066] In the example described here, the encoding process is implemented by an encoding artificial neural network, and the decoding process is implemented by a decoding artificial neural network.
[0067] A collection of encoding and decoding artificial neural networks (where the output of the encoding artificial neural network is applied to the input of the decoding artificial neural network) forms, for example, an autoencoder.
[0068] The encoding-decoding pair is selected from among a number of predefined encoding-decoding pairs, that is, in this case, from among a number of encoding artificial neural network-decoding artificial neural network pairs.
[0069] When a list of artificial neural networks accessible by an electronic decoding device is received in advance (as explained above at step E2), the encoding-decoding process pair is selected, for example, in the case where the decoding process uses an encoding-decoding process pair of artificial neural networks present in the received list.
[0070] The encoding-decoding process pair can also be selected based on the intended application (e.g., indicated by the user using the user interface (not shown) of the electronic encoding device 2). For example, if the intended application is video conferencing, the selected encoding-decoding process pair includes a low-latency decoding process. In other applications, the selected encoding-decoding process pair will include a random access decoding process.
[0071] In the low latency process used to decode the video sequence, the images of the video sequence are represented, for example, by encoded data that can be sent and decoded immediately; the data can then be sent in the order in which the video images are displayed, which in this case ensures a one-frame delay between encoding and decoding.
[0072] In the random access processing used to decode the video sequence, encoded data associated with multiple images are sent in a different order than the display order of these images, which allows for increased compression. Then, encoded images that do not reference other images (so-called intraframes) can be encoded regularly, which allows decoding of the video sequence to begin from several positions in the encoded stream.
[0073] For this purpose, you can refer to the article “Overview of the High Efficiency Video Coding (HEVC) Standard” by GJ Sullivan, J.-R. Ohm, W.-J. Han and T. Wiegand, in IIEEE Transactions on Circuits and Systems for Video Technology, December 2012, Vol. 22, No. 12, pp. 1649-1668.
[0074] You can also select the encoding-decoding process pair to obtain the best possible compression-distortion trade-off.
[0075] To this end, multiple encoding-decoding pairs can be applied to the content data B, and a set that achieves a better compression-distortion trade-off can be selected.
[0076] Alternatively, the type of content can be determined (e.g., by analyzing content data B) and an encoding-decoding pair can be selected based on the determined type.
[0077] The encoding-decoding process pair can also be selected based on the computing power available at the electronic decoding device. Information representing this computing power may have been previously transmitted from the electronic decoding device to the electronic encoding device (and received by the electronic encoding device, for example, in step E2 above).
[0078] Different criteria used to select encoding-decoding pairs may be combined.
[0079] Once the encoding-decoding process pair is selected, the processor 4 configures the parallel processing unit 6 in step E6, so that the parallel processing unit 6 can implement the selected encoding method.
[0080] Step E6 specifically includes the instantiation of the encoded artificial neural network 8 used by the selected encoding process within the parallelization processing unit 6.
[0081] This instantiation can specifically include the following steps:
[0082] - Reserve the storage space required for implementing the encoded artificial neural network within the parallel processing unit 6; and / or
[0083] - Parallelize processing unit 6 by programming the weights W and activation functions of the limited encoded artificial neural network 8; and / or
[0084] - Load at least a portion of the content data B into the local memory of the parallelization processing unit 6.
[0085] Then, Figure 2 The method includes step E8 of implementing the encoding process, that is, applying content data B at the input of the encoding artificial neural network 8 (or in other words, activating the encoding artificial neural network 8 by taking content data B as input).
[0086] Therefore, step E8 allows the generation of compressed data C (here, at the output of the encoded artificial neural network 8).
[0087] The following steps involve encoding (i.e., preparing) a data stream that specifically contains compressed data C and is intended to be sent to an electronic decoding device (e.g., refer to...). Figure 7 The described electronic decoding device 10).
[0088] Therefore, the method specifically includes step E10 of encoding a first header portion Fc, which includes data features of the representation format of audio or video content (here, for example, data in a format linked to the video sequence being encoded).
[0089] These data indicators that form the first header portion Fc include, for example, image size (in pixels), frame rate, binary depth of luminance information, and binary depth of chrominance information. This data is constructed, for example, based on the aforementioned format data P (after potential reformatting).
[0090] Then, Figure 2 The method continues to step E12, determining the decoding artificial neural network (used by the decoding process selected in step E4) for an electronic decoding device capable of decoding the data stream (e.g., see below). Figure 7 The availability of the described electronic decoding device 10).
[0091] This determination can be based on the list received in step E2: in this case, processor 4 determines whether the decoding artificial neural network used by the decoding process selected in step E4 belongs to the list received in step E2. (Of course, in embodiments where the encoding-decoding process pair is systematically selected to correspond to a decoding artificial neural network that can be used in an electronic decoding device, step E12 can be omitted, and the method then proceeds to step E14.)
[0092] According to a possible embodiment, in the absence of information about the availability of the electronic decoding device for decoding the artificial neural network, the method proceeds to step E16 (in such a manner that data describing the decoding of the artificial neural network is transmitted to the electronic decoding device, as explained above).
[0093] If processor 4 determines in step E12 that the decoding artificial neural network is available for the electronic decoding device (arrow P), the method continues to step E14 below.
[0094] If processor 4 determines in step E12 that decoding the artificial neural network is not available for the electronic decoding device (arrow N), the method continues to step E16 below.
[0095] Alternatively, as a step after step E12, a choice can be made between steps E14 and E16 based on another criterion, such as a dedicated indicator stored in the electronic coding device 2 (and possibly adjusted by the user through the user interface of the electronic coding device 2), or a choice made by the user (e.g., obtained through the user interface of the electronic coding device 2).
[0096] In step E14, processor 4 encodes a second header portion including the indicator IND and a third header portion including the identifier Inn of the decoded artificial neural network.
[0097] In step E14, the indicator IND encoded in the data stream indicates that the decoded artificial neural network belongs to a predetermined set of artificial neural networks, which in this case is the set of artificial neural networks available (or accessible) to the electronic decoding device (that is, for example, the set of artificial neural networks in the list received in step E2).
[0098] The identifier Inn for decoding the artificial neural network is the identifier for the decoded artificial neural network that is conventionally defined (specifically, shared by electronic encoding and electronic decoding devices), for example, within the aforementioned predetermined set.
[0099] In step E16, processor 4 encodes a second header portion including the indicator IND' and a third header portion including the data Rc describing the decoded artificial neural network.
[0100] In step E16, the indicator IND' encoded in the data stream indicates that the decoded artificial neural network is encoded in the data stream, that is, represented by the aforementioned descriptive data Rc.
[0101] Decoding artificial neural networks is done, for example, by Rc encoding (that is, represented by Rc) of descriptive data (or data that encodes and decodes artificial neural networks) according to standards such as MPEG-7 Part 17 or formats such as JSON.
[0102] For further information, please refer to the article "DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression", presented by S. Wiedemann et al. at the 36th International Conference on Machine Learning in Long Beach, California, PMLR 97, 2019, or the article "Compact and Computationally Efficient Representation of Deep Neural Networks", presented by S. Wiedemann et al. in IEEE Transactions on Neural Networks and Learning Systems (Vol. 31, No. 3), March 2020.
[0103] After step E14, as after step E16, Figure 2 The method continues to step E18, determining the possibility that the electronic decoding device can use a decoding artificial neural network to implement the decoding process.
[0104] For example, processor 4 determines the possibility, for instance, by determining (potentially by means of prior exchange between electronic encoding and electronic decoding devices) whether the electronic decoding device includes modules or software designed to implement the decoding process, said software being adapted for the electronic decoding device to implement the decoding process when executed by the processor of the electronic decoding device.
[0105] If processor 4 determines that the electronic decoding device can perform the decoding process, the method continues to step E22 as described below.
[0106] If processor 4 determines that the electronic decoding device cannot perform the decoding process, the method executes step E20 as described below (before proceeding to step E22).
[0107] Alternatively, the choice to perform or not perform step E20 (before performing step E22) may depend on another criterion, such as a dedicated indicator stored in the electronic coding device 2 (and possibly adjusted by the user through the user interface of the electronic coding device 2), or on a choice made by the user (e.g., obtained through the user interface of the electronic coding device 2).
[0108] In step E20, processor 4 encodes a fourth header portion in the data stream containing a computer program (or code) executable by the processor of an electronic decoding device, namely Exe. (See below) Figure 8 Describe the use of the computer program EXE in electronic decoding equipment.
[0109] To accommodate execution within the electronic decoding device, the computer program selects from a library, for example, based on information related to the hardware configuration of the electronic decoding device (e.g., information received during a previous exchange between the electronic encoding device 2 and the electronic decoding device).
[0110] Then, processor 4 proceeds to step E22, where it encodes the compressed stream Fnn based on the compressed data C obtained in step E8.
[0111] In this regard, it is observed that in the description above, step E8 is described before the steps of encoding the header Fet (steps E10 to E20). However, step E8 can actually be performed before step E22.
[0112] Specifically, when step E8 allows processing only a portion of the audio or video content to be compressed (e.g., when step E8 performs processing of blocks, components, or images of the video sequence to be compressed), the implementation of steps E8 (to obtain compressed data associated with consecutive portions of the content) and E22 (to encode the obtained compressed data in the data stream) can be repeated.
[0113] Therefore, processor 4 can construct a complete data stream including header Fet and compressed stream Fnn in step E24.
[0114] The complete data stream is constructed in such a way that the header Fet and the compressed stream Fnn can be identified separately.
[0115] According to a possible embodiment, the header Fet contains an indicator of the start of the compressed stream Fnn within the complete data stream. This indicator is, for example, the position (in bits) at which the start of the compressed stream Fnn originates from the start of the complete data stream. (In other words, in this case, the header has a predetermined fixed length.)
[0116] Other means of identifying the header Fet and the compressed stream Fnn could be considered as alternatives, such as a marker (that is, a binary combination used to indicate the start of the compressed stream Fnn, and which is prohibited from being used in the rest of the data stream or at least in the header Fet).
[0117] exist Figures 3 to 6 It shows that it can be accessed through Figure 2 An example of a complete data stream obtained using this method.
[0118] As mentioned above, these data streams include the header Fet and the compressed stream Fnn.
[0119] exist Figure 3 In the case of (corresponding to the case where step E14 has been implemented but step E20 has not yet been implemented), the header includes:
[0120] - The first part, Fc, includes the data characteristics of the representation format of the audio or video content;
[0121] - The second part includes an indicator IND that indicates whether the decoded artificial neural network belongs to a predetermined set of artificial neural networks; and
[0122] - The third part includes the identifier Inn for decoding the artificial neural network.
[0123] exist Figure 4 In the case of (corresponding to the case where step E16 has been implemented but step E20 has not yet been implemented), the header includes:
[0124] - The first part, Fc, includes the data characteristics of the representation format of the audio or video content;
[0125] - The second part includes the indicator IND', which indicates that the decoded artificial neural network is encoded in the data stream; and
[0126] - The third part includes the data Rc that describes the decoding of the artificial neural network (here, the data encoding).
[0127] exist Figure 5 In the case where steps E14 and E20 have already been performed, the header includes:
[0128] - The first part, Fc, includes the data characteristics of the representation format of the audio or video content;
[0129] The second part includes an indicator IND that indicates whether the decoded artificial neural network belongs to a predetermined set of artificial neural networks;
[0130] - Part Three includes the identifier Inn for decoding artificial neural networks; and
[0131] - Part Four includes the computer program Exe.
[0132] exist Figure 6 In the case where steps E16 and E20 have already been performed, the header includes:
[0133] - The first part, Fc, includes the data characteristics of the representation format of the audio or video content;
[0134] - The second part includes the indicator IND', which indicates that the decoded artificial neural network is encoded in the data stream; and
[0135] - The third part includes the data Rc describing the decoding of the artificial neural network (here, the data encoding); and
[0136] - Part Four, which includes the computer program Exe.
[0137] The data stream constructed in step E24 can be encapsulated in a known transmission format, such as the "Packet Transport System" format or the "Byte Stream" format.
[0138] In the case of a "packet transport system" format (such as that proposed by the RTP protocol), data is encoded into identifiable packets and transmitted over the communication network. Using packet identification information provided by the network layer, the network can easily identify the boundaries of the data (images, image groups, and here the header Fet and compressed stream Fnn).
[0139] In the format “byte stream”, there are no specific groups, and the construction of step E24 must allow (e.g., using a Network Abstraction Layer (NAL) unit) additional means to identify the boundaries of related data (e.g., the boundaries between stream portions corresponding to each image, and the boundaries between the header Fet and the compressed stream Fnn), where a unique binary combination (e.g., 0x00000001) makes it possible to identify the boundaries between data.
[0140] Then, the complete data stream constructed in step E24 can be transmitted to electronic decoding device 26 in step E26 (via a communication device not shown and / or via at least one communication network), or stored within electronic encoding device 2 (for later transmission or as an alternative for later decoding, for example, within the electronic encoding device itself, in which case the electronic encoding device is designed to further implement the following reference). Figure 8 (Description of the decoding method).
[0141] When the audio or video content comprises multiple parts (e.g., multiple sets of images when the content is a video sequence), the methods in steps E4 to E24 can be implemented for each part of the content (e.g., for each set of images) in such a way that, for each part of the content (e.g., for each set of images), the following is obtained: Figures 3 to 6 The data flow is illustrated in one of the diagrams. Therefore, the compressed stream Fnn associated with each group of images can be decoded using an artificial neural network specific to the group of images of interest, and potentially different from the artificial neural networks used for other groups of images, as described below. The artificial neural networks may have the same structure (and differ only in their weights and / or the activation functions that define the specific artificial neural network).
[0142] Figure 7 An electronic decoding device 10 using at least one artificial neural network 18 is shown.
[0143] The electronic decoding device 10 includes a receiving unit 11, a processor 14 (e.g., a microprocessor), and a parallel processing unit 16, such as a graphics processing unit or GPU, or a tensor processing unit or TPU.
[0144] The receiving unit 11 is, for example, a communication circuit (such as a radio frequency communication circuit) and can receive data (and specifically, an encoded data stream) from an external electronic device such as the electronic encoding device 2 and transmit this data to the processor 14 (the receiving unit 11 is connected to the processor 14, for example, via a bus).
[0145] The electronic decoding device 10 also includes a storage unit 12, such as a memory (possibly rewritable non-volatile memory) or a hard disk drive. Although the storage unit 12 is... Figure 7 The components shown are different from those of processor 14, but storage unit 12 can be integrated into (i.e. included in) processor 14 as an alternative.
[0146] In this case, processor 14 is adapted to execute multiple instructions sequentially, for example, a computer program stored in memory unit 12.
[0147] The parallel processing unit 16 is designed to implement the artificial neural network 18 after it has been configured by the processor 14. To this end, the parallel processing unit 16 is designed to execute multiple operations of the same type in parallel at a given time.
[0148] like Figure 7 As schematically shown, processor 14 receives a data stream (e.g., via a communication device not shown in electronic decoding device 10), which includes a first dataset, here the header Fet, and a second dataset representing audio or video content, here the compressed stream Fnn.
[0149] As described below, the artificial neural network 18 is used within the framework of processing the second dataset (that is, the compressed data Fnn described herein) to obtain audio or video content corresponding to the initial audio or video content B.
[0150] Storage unit 12 can store multiple parameter sets, each parameter set defining the decoding of the artificial neural network. As described below, in this case, processor 14 can configure parallel processing unit 16 with the aid of specific parameter sets in these parameter sets, so that parallel processing unit 16 can then implement the artificial neural network defined by that specific parameter set.
[0151] Storage unit 12 may specifically store a first parameter set defining a first artificial neural network forming a random access decoder and / or a second parameter set defining a second artificial neural network forming a low latency decoder.
[0152] In this case, the electronic decoding device 10 has decoding options in advance for both situations where random access to content is desired and situations where content needs to be displayed without delay.
[0153] Now refer to Figure 8 Describes a decoding method implemented within an electronic decoding device 10 and using an artificial neural network 18 implemented by a parallel processing unit 16.
[0154] This method can be initiated by electronic decoding device 10 transmitting optional steps of a list L of artificial neural networks available to electronic decoding device 10 to a device for controlling the transmission of the data stream to be decoded. The data stream transmission control device can be, for example, electronic encoding device 2. (In this case, electronic encoding device 2 is referred to above...) Figure 2 The described step E2 receives the list L) Alternatively, the data stream transmission control device can be a dedicated server that works in conjunction with the electronic coding device 2.
[0155] The artificial neural network accessible by the electronic decoding device 10 is an artificial neural network for which the electronic decoding device 10 stores a parameter set of the artificial neural network of interest (as indicated above), or can access the parameter set by a remote electronic device connected to a server (as explained below).
[0156] Figure 8 The method includes step E52 of receiving (via electronic decoding device 10, and specifically via receiving unit 11) a data stream comprising a first parameter set (i.e., header Fet) and a second parameter set, namely the compressed stream Fnn. Receiving unit 11 transmits the received data stream to processor 14.
[0157] Then processor 14 proceeds to step E54, for example, by means of an indicator of the start of the compressed stream (already mentioned in the description of step E24) to identify the first dataset (header Fet) and the second dataset (compressed stream Fnn) within the received data stream.
[0158] Processor 14 can also identify different parts of the first dataset (header) in step E54, that is, within the header Fet: a first part Fc (including data representing the format characteristics of the content encoded by the data stream), a second part (indicator IND or IND'), a third part (identifier Inn or encoded data Rc), and possibly a fourth part (computer program Exe), as described above. Figures 3 to 6 As shown.
[0159] If executable instructions (e.g., instructions of a computer program Exe) are identified (i.e. detected) within the first data in step E54, processor 14 may initiate the execution of these executable instructions in step E56 to implement at least some steps of processing the data from the first dataset (described below). These instructions may be executed by processor 14, or alternatively by a virtual machine instantiated within electronic decoding device 10.
[0160] Figure 7 The method continues to step E58 to decode the data Fc, which is a representation format of the audio or video content, in a manner that reveals the characteristics of the format. For example, in the case of video content, decoding the data portion Fc allows the binary depth of the image size (in pixels) and / or frame rate and / or luminance information and / or chroma information to be obtained.
[0161] Then, processor 14 proceeds (in some embodiments, due to the execution of instructions identified in the first dataset in step E54 as already indicated) to step E60, which decodes the indicators IND, IND' contained in the second part of the header Fet.
[0162] If the decoding of the indicator IND, IND' present in the received data stream indicates that the artificial neural network 18 to be decoded belongs to a predetermined set of artificial neural networks (that is, if the indicator present in the first dataset is the indicator IND indicating that the decoded artificial neural network 18 belongs to a predetermined set of artificial neural networks), then the method continues with step E62 as described below.
[0163] If the decoding indication of the indicator IND, IND' present in the received data stream is to be used to decode the artificial neural network 18 encoded in the data stream (that is, if the indicator present in the first dataset is the indicator IND' indicating that the decoding artificial neural network 18 is encoded in the data stream), then the method continues to step E66 as described below.
[0164] In step E62, processor 14 performs decoding of identifier Inn (in some embodiments, due to the execution of instructions identified in the first dataset in step E54, as already indicated) (which is contained in the third part of the header Fet). As already noted, this identifier Inn is an identifier that specifies the decoding of artificial neural network 18, for example, in a predetermined set of artificial neural networks described above.
[0165] Then, the processor 14 may continue in step E64 (in some embodiments, due to the execution of instructions identified in the first dataset in step E54 as already indicated), for example, reading the parameter set associated with the decoding identifier Inn (which defines the artificial neural network identified by the decoding identifier Inn) from the storage unit 12.
[0166] According to possible embodiments, the processor 14 may generate an error message in the absence of data (specifically parameters) related to the artificial neural network identified by the decoding identifier Inn (here within the storage unit 12).
[0167] As an alternative (or in the case that the parameter set of the artificial neural network identified by the decoding identifier Inn is not stored in storage unit 12), electronic decoding device 10 may send a request for a parameter set (which includes, for example, the decoding identifier Inn) to a remote server (in some embodiments, due to the execution of an instruction identified in the first dataset in step E54, as already indicated), and receive, in step E64, a response that defines the parameter set of the artificial neural network identified by the decoding identifier Inn.
[0168] The method then proceeds to step E68, as described below.
[0169] In step E66, processor 14 continues (in some embodiments, due to the execution of instructions identified in the first dataset in step E54 as already indicated) to decode the data Rc describing the artificial neural network 18 (included here in the third part of the header Fet).
[0170] As already noted, this descriptive data (or encoded data) Rc is encoded, for example, according to standards such as MPEG-7 part 17 or formats such as JSON.
[0171] Decoding the descriptive data Rc allows the parameters of the artificial neural network used to decode the data to be obtained from a second dataset (that is, data from the compressed stream Fnn).
[0172] In this case, the method also proceeds to step E68, which will now be described.
[0173] Then, in step E68, the processor 14 configures the parallel processing unit 16 by means of the parameters defining the decoding artificial neural network 18 (the parameters obtained in step E64 or step E66) (in some embodiments, due to the execution of instructions identified in the first dataset in step E54, as already indicated), so that the parallel processing unit 16 can implement the decoding of the artificial neural network 18.
[0174] The configuration step E68 specifically includes the instantiation of the decoded artificial neural network 18 within the parallelization processing unit 16, using the parameters obtained in step E64 or step E66.
[0175] This instantiation can specifically include the following steps:
[0176] - Reserve the storage space required to implement the decoding of the artificial neural network 18 within the parallel processing unit 16; and / or
[0177] -Program the parallelization processing unit 16 with parameters (e.g., including weights W' and activation functions) that define the decoding artificial neural network 18 (parameters obtained in step E64 or step E66); and / or
[0178] - Load at least a portion of the data from the second dataset (that is, at least a portion of the data from the compressed stream Fnn) into the local memory of the parallelization processing unit 16.
[0179] As can be seen from the description of steps E58 to E68 above, the data from the first dataset Fet is therefore processed by processor 14.
[0180] Then, in step E70, processor 14 may apply (i.e. render) the data from the second dataset (here, the data from the compressed stream Fnn) to the artificial neural network 18 implemented by parallelization processing unit 16, such that the data is processed by a decoding process that at least partially utilizes the artificial neural network 18.
[0181] In the example described here, the artificial neural network 18 receives data from a second dataset Fnn as input and produces a representation I of the encoded content as output, which is adapted for reproduction on an audio or video reproduction device. In other words, at least some data from the second dataset Fnn is applied to the input layer of the artificial neural network 18, and the output layer of the artificial neural network 18 produces the aforementioned representation I of the encoded content. In the case of video content (including images or image sequences), the artificial neural network 18 thus produces at least one matrix representation I of the image as output (that is, in the output layer of the artificial neural network 18).
[0182] In some embodiments, in order to process certain data from the compressed stream Fc (e.g., corresponding to blocks or images), the artificial neural network 18 may receive at least some data generated at the output of the artificial neural network 18 during the processing of previous data in the compressed stream Fc (e.g., corresponding to previous blocks or previous images) as input. In this case, proceeding to step E72, the data generated at the output of the artificial neural network 18 is re-injected into the input of the artificial neural network 18.
[0183] Furthermore, according to other possible embodiments, the decoding process may use multiple artificial neural networks, as already mentioned above regarding the processing of content data B.
[0184] Therefore, the data from the second group (at least some of the data from the compressed stream Fnn) has been processed by relying on the processing of a portion of the data from the first group (relying on the processing of the identifier Inn of the encoded data Rc) and using an artificial neural network 18 implemented by the parallel processing unit 16.
[0185] Then, in step E74, the processor 14 determines whether the processing of the compressed stream Fnn by means of the artificial neural network 18 has ended.
[0186] In the case of negative determination (N), the method loops to step E70, which applies additional data from the compressed stream Fnn to the artificial neural network 18.
[0187] In the case of a positive determination (P), the method continues to step E76, where the processor 14 determines whether there is any more data to be processed in the received data stream.
[0188] In the case of a negative determination (N) in step E76, the method ends in step E78.
[0189] In the case of a positive determination (P) in step E76, the method loops back to step E52 to process the new portion of the data stream, such as... Figures 3 to 6 As shown in one of the diagrams.
[0190] As indicated above regarding the repeated encoding steps E4 to E24, this other portion of the data stream then includes a first dataset and a second dataset representing another set of audio or video content (e.g., in the case of video content, another set of images representing the representation format of the content used). In this case, another artificial neural network can be determined based on some of these first data (identifier Inn or encoded data Rc), as described above in steps E54 to E66, and then the parallelization processing unit 16 can be configured to implement this other artificial neural network (according to step E68 above). Then, the data from the second dataset (e.g., related to the other set of images mentioned above) from this other portion of the data stream can be decoded by means of this other artificial neural network (as described above in step E70).
[0191] The other artificial neural network mentioned earlier can have the same structure as the artificial neural network 18 described above, which simplifies the steps of configuring the parallel processing unit 16 (e.g., only updating the weights and / or activation functions that define the current artificial neural network).
Claims
1. A method implemented by a decoding device for decoding at least a portion of a data stream, said at least a portion comprising an indicator (IND; IND') and representative data (Fnn) representing at least one block of an image or image component, said method comprising the steps of: -Decode (E60) the indicator (IND; IND'); -If the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) is encoded in at least a portion of the data stream: Then, the parameters of the artificial neural network (18) are decoded from at least a portion of the data stream, or... -If the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) is accessible to the decoding device: Then read the parameters of the artificial neural network (18) stored locally, or Request the parameters of the artificial neural network (18) from the remote server; -The representative data (Fnn) is decoded block by block (E70) by means of the artificial neural network (18).
2. The decoding method according to claim 1, wherein, The at least portion is a first part of a video sequence, the video sequence including the first part and a second part, wherein the step of decoding the representative data produces the first part, and wherein the method further includes the step of decoding other data by means of another artificial neural network to produce the second part.
3. The decoding method according to claim 2, wherein, The other artificial neural network has the same structure as the artificial neural network.
4. The decoding method according to claim 2, wherein, The first part and the second part respectively form two sets of images in a format used to represent the content used.
5. The decoding method according to claim 1, comprising the step of transmitting the list (L) of the artificial neural network to a device for controlling the transmission of the data stream (E50).
6. The decoding method according to any one of claims 1 to 5, comprising: If it is determined by decoding the indicator (IND) that the artificial neural network (18) is accessible by the decoding device, then the identifier (Inn) of the neural network (18) is decoded (E62).
7. The decoding method according to claim 6, comprising reading the parameters of the artificial neural network (18) identified by the decoded identifier (Inn) in the storage unit (12).
8. The decoding method according to claim 7, wherein, The storage unit (12) stores a first set of parameters representing a first artificial neural network forming a random access decoder and a second set of parameters representing a second artificial neural network forming a low latency decoder.
9. The decoding method of claim 6, further comprising the step of generating an error message in the absence of data relating to the artificial neural network identified by the decoded identifier.
10. The decoding method according to claim 6, comprising receiving from a remote server parameters of the artificial neural network (18) identified by the decoded identifier (Inn).
11. A decoding device, comprising: - A unit (11) for receiving at least a portion of a data stream, said at least a portion including an indicator (IND; IND') and representative data (Fnn) representing at least one block of an image or image component. - Decoding components (14, 16), the decoding components being designed to: Decode the indicator (IND; IND'). If the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) is encoded in the at least portion of the data stream: then the parameters of the artificial neural network (18) are decoded from the at least portion of the data stream. If the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) is accessible to the decoding device: then the parameters of the artificial neural network (18) are read from the local storage, or the parameters of the artificial neural network (18) are requested from a remote server. The representative data is decoded block by block using the artificial neural network (18).