Parallel video decoding using neural networks

The method and device dynamically configure a parallel processing unit using a processor to decode audio or video content with an artificial neural network, addressing inefficiencies in existing methods by enhancing flexibility and efficiency in video decoding.

JP7850128B2Active Publication Date: 2026-04-22オランジュ
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
オランジュ
Filing Date
2021-07-13
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing video decoding methods are inefficient and inflexible, particularly when utilizing parallel processing units and artificial neural networks, as they lack the ability to dynamically adjust processing based on the data stream content.

Method used

A method and electronic device that utilize a processor and a parallel processing unit to decode audio or video content by processing a first dataset to configure the parallel processing unit based on a portion of the data stream, allowing for flexible and efficient decoding using an artificial neural network, which can be adjusted by memory allocation, instantiation, and assignment of weights and activation functions.

Benefits of technology

Enables flexible and efficient decoding of audio or video content by dynamically configuring the parallel processing unit, improving processing efficiency and adaptability to different content types and computing power constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850128000001
    Figure 0007850128000001
  • Figure 0007850128000002
    Figure 0007850128000002
  • Figure 0007850128000003
    Figure 0007850128000003
Patent Text Reader

Abstract

A method for decoding a data flow by an electronic device (10) including a processor (14) and parallel processing units (16) designed to perform multiple operations of the same type in parallel at a given time, the data flow including a first set of data (Fet) and a second set of data (Fnn) representing audio or video content, the decoding method comprising the steps of: processing (E70) data from the first set of data (Fet) by the processor (14); and processing (E70) data from the second set of data (Fnn) using a process that depends at least in part on the data from the first set (Fet) and using an artificial neural network (18) implemented by the parallel processing units (16) to obtain the audio or video content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the decoding of content, particularly in the technical field of audio or video content.

[0002] Particularly, the present invention relates to a method and an electronic device for decoding a data stream and a related data stream.

Background Art

[0003] Particularly in the field of video decoding, it is known to use an electronic device that includes both a processor (generally the central unit of an electronic device, i.e., the "central processing unit", abbreviated as CPU) and a parallel processing unit designed to perform a plurality of operations of the same type in parallel at a certain point in time. Such a parallel processing unit is, for example, a graphics processing unit, i.e., GPU or a tensor processing unit, i.e., TPU, which is described, for example, in the article "Google’s Tensor Processing Unit explained:this is what the future of computing looks like" by Joe Osborne of Techradar dated August 22, 2016.

[0004] Furthermore, it has been proposed to compress data representing video content by an artificial neural network. Therefore, the decoding of the compressed data can be performed by another artificial neural network, which is described, for example, in the article "DVC:An End-to-end Deep Video Compression Framework" by Guo Lu et al. in 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp.10998-11007.

Summary of the Invention

Means for Solving the Problems

[0005] In this regard, the present invention relates to a method for decoding a data stream by an electronic device including a processor and a parallel processing unit designed to perform multiple operations of the same kind in parallel at a given time, wherein the data stream includes a first dataset and a second dataset representing audio or video content, and the method is - The step of processing data from the first dataset with a processor, - A step of obtaining audio or video content by processing data from a second dataset using an artificial neural network implemented by a parallel processing unit, using a process that relies on at least a portion of the data from the first dataset, We propose a method characterized by including the following.

[0006] The step of processing data from a second set by an artificial neural network implemented by a parallel processing unit can therefore be adjusted according to the data from the first set contained in the data stream. This makes the processing of the second data for its decoding by the artificial neural network flexible and efficient.

[0007] The method, therefore, may include the step of configuring, for example, a parallel processing unit based on at least a portion of the data from a first dataset.

[0008] The configuration of a parallel processing unit may include memory allocation for the parallel processing unit, and / or instantiation of memory for parallel processing, and / or assignment of values ​​to processing operations performed on the parallel processing unit (in accordance with the aforementioned portion of the data from the first dataset) (here, in practice, the assignment of weights and / or activation functions that define an artificial neural network, these weights and / or activation functions are identified in accordance with the aforementioned portion of the data from the first dataset).

[0009] According to the first possibility, the first dataset may include data describing an artificial neural network (for example, data encoding an artificial neural network). In this case, during the configuration step, the processor may configure parallel processing units based on this descriptive data.

[0010] According to the second possibility, the electronic device may include memory units for storing multiple parameter sets, each defining multiple artificial neural networks. In this case, the first dataset may include identifiers. Therefore, in the configuration step, the processor can configure parallel processing units based on the parameter set associated with this identifier from among the multiple parameter sets.

[0011] A processor is, for example, a microprocessor, and a processor is therefore capable of executing multiple instructions from a computer program in sequence.

[0012] Furthermore, the data stream may be configured to include instructions that can be executed within an electronic device (for example, by a processor or, alternatively, by a virtual machine). Processing of the data from the first dataset can, in this case, be performed at least partially by executing at least some of these instructions.

[0013] In particular, the configuration step of the parallel processing unit can therefore be carried out by executing at least some of these instructions.

[0014] The decoding method may further include the step of identifying the first and second datasets within the data stream (for example, by utilizing predetermined binary lengths for the first and / or second datasets, or by utilizing indicators of the boundary between the first and second datasets).

[0015] The first dataset may further include data representing the format characteristics of the content encoded by the data stream.

[0016] If the content is video content (i.e., the content includes at least one image, possibly a sequence of images), then processing the data from the second set may, for example, generate at least one matrix representation of at least a portion of the image (e.g., a block or component of the image or the entire image).

[0017] The artificial neural network may receive data from a second dataset as input (i.e., at the input layer of the artificial neural network). The artificial neural network may further generate the aforementioned matrix representation as output (i.e., at the output layer of the artificial neural network).

[0018] In a conceivable embodiment, the artificial neural network may receive previously generated data as input (i.e., in the input layer) as output (i.e., in the output layer) of the artificial neural network.

[0019] The present invention also proposes an electronic device for decoding a data stream comprising a first dataset and a second dataset representing audio or video content, the electronic device, - A processor adapted to process data from the first dataset, - A parallel processing unit designed to perform multiple operations of the same type in parallel at a given time, wherein the parallel processing unit is adapted to acquire audio or video content by processing data from a second dataset using a process that depends on at least a portion of the data from a first set, and using an artificial neural network implemented by the parallel processing unit. Includes.

[0020] As described above, the processor may be further adapted to configure the parallel processing unit according to at least a part of the data from the first data set.

[0021] The parallel processing unit may be adapted to generate at least one matrix representation of at least a part of the image.

[0022] Finally, the present invention proposes a data stream including a first data set and a second data set representing audio or video content, wherein the first data set includes data that at least partially defines a process for processing data from the second data set using an artificial neural network.

[0023] As described above and as will be described below, these data that at least partially define the processing process may be an identifier of an artificial neural network (from a predetermined set of artificial neural networks) or data describing the artificial neural network (e.g., data encoding it).

[0024] Of course, the various features, alternative forms, and embodiments of the present invention may be associated with each other in various combinations as long as they do not conflict with each other or are not exclusive.

[0025] Furthermore, various other features of the present invention will become apparent from the following description with respect to the drawings below, which illustrate non-limiting embodiments of the present invention.

Brief Description of the Drawings

[0026] [Figure 1] Shows an electronic encoding device used within the framework of the present invention. [Figure 2] It is a flowchart showing the steps of the encoding method executed within the electronic encoding device of FIG. 1. [Figure 3] It is a first example of the data stream obtained by the method of FIG. 2. [Figure 4]This is a second example of a data stream obtained by the method shown in Figure 2. [Figure 5] This is a third example of a data stream obtained by the method shown in Figure 2. [Figure 6] This is a fourth example of a data stream obtained by the method shown in Figure 2. [Figure 7] This shows an electronic coding device according to an embodiment of the present invention. [Figure 8] Figure 7 is a flowchart showing the steps of the decoding method performed within the electronic decryption device. [Modes for carrying out the invention]

[0027] Figure 1 shows an electronic coding device 2 that uses at least one artificial neural network 8.

[0028] The electronic coding device 2 includes a processor 4 (e.g., a microprocessor) and a parallel processing unit 6, such as a graphics processing unit, i.e., a GPU, or a tensor processing unit, i.e., a TPU.

[0029] As schematically shown in Figure 1, the processor 4 receives data P and B, in this case format data P and content data B, representing the audio or video content to be compressed.

[0030] Format data P indicates the characteristics of the representation format of the audio or video content, such as the image size (in pixels), frame rate, binary depth of luminance information, and binary depth of chrominance information in the case of video content.

[0031] Content data B forms an (uncompressed) representation of the audio or video content. For example, in the case of video content, the content data includes, for each pixel in each image of the image sequence, data representing the luminance value of that pixel and / or data representing the color difference value of that pixel.

[0032] The parallel processing unit 6 is designed to implement the artificial neural network 8 after being configured by the processor 4. For this purpose, the parallel processing unit is designed to perform multiple operations of the same type in parallel at a given time.

[0033] As described below, the artificial neural network 8 is used within the framework of processing content data B with the aim of obtaining compressed data C.

[0034] In the embodiments described herein, when content data B is applied as input to artificial neural network 8, the artificial neural network 8 generates compressed data C as output.

[0035] The content data B applied at the input of the artificial neural network 8 (i.e., applied at the input layer of the artificial neural network 8) may also represent blocks of images, or blocks of image components (e.g., blocks of luminance or chrominance components or blocks of the color components of the image), or images in a video sequence, or components of images in a video sequence (e.g., luminance or chrominance components or color components), or a series of images in a video sequence.

[0036] For example, in this case, at least some of the neurons in the input layer may be adapted to receive pixel values ​​of image components, each represented by one of the content data B.

[0037] Alternatively, processing content data B may involve the use of several artificial neural networks, as described, for example, in the aforementioned article, "DVC: An End-to-end Deep Video Compression Framework" by Guo Lu et al. at the 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.

[0038] An example of an encoding method performed by the electronic encoding device 2 will now be explained with reference to Figure 2.

[0039] The memory linked to processor 4 stores, for example, computer program instructions, which are designed to execute at least some of the steps of the method shown in Figure 2 when these instructions are executed by processor 4. In other words, processor 4 is programmed to execute at least some of the steps of Figure 2.

[0040] The method shown in Figure 2 here begins with an optional step E2, in which a list of artificial neural networks accessible by the electronic decoding device is received.

[0041] The list is received, for example, directly from the electronic decoding device (for example, the electronic decoding device 10 in Figure 7, which will be described later) by the processor 4 of the electronic encoding device 2.

[0042] As described below, an artificial neural network accessible by an electronic decoding device is an artificial neural network in which the electronic decoding device either stores parameters that define the artificial neural network it is related to, or can access these parameters by connecting to remote electronic equipment such as a server.

[0043] Alternatively, the list can be received directly from a remote server, such as the aforementioned server, by the processor 4 of the electronic device 2.

[0044] The method in Figure 2 follows step E4, which involves selecting a coding-decoding process pair. As previously mentioned regarding the coding process, both the coding and decoding processes utilize at least one artificial neural network.

[0045] In the examples described herein, the encoding process is performed by an encoding artificial neural network, and the decoding process is performed by a decoding artificial neural network.

[0046] A set formed by an encoding artificial neural network and a decoding artificial neural network (where the output of the encoding artificial neural network is applied to the input of the decoding artificial neural network) forms, for example, an autoencoder.

[0047] The encoding process-decoding process pair is selected from, for example, a set of pre-configured encoding process-decoding process pairs, i.e., a set of encoding artificial neural network-decoding artificial neural network pairs.

[0048] If a list of artificial neural networks accessible by the electronic decoding device is received in advance (as described above in step E2), the encoding-decoding process pair is selected from among the encoding-decoding process pairs that use, for example, an artificial neural network in the list received by the decoding process.

[0049] The encoding-decoding process pair may also be selected depending on the intended application (e.g., as indicated by the user using the not-illustrated user interface of the electronic encoding device 2). For example, if the intended application is video conferencing, the selected encoding-decoding process pair includes a low-latency decoding process. For other applications, the selected encoding-decoding process pair includes a random-access decoding process.

[0050] In a low-latency process for decoding video sequences, the images in the video sequence may be represented, for example, by encoded data, which can be transmitted immediately and decoded. The data may then be transmitted in the order in which the video images are displayed, thereby ensuring a latency of one frame between encoding and decoding.

[0051] In a random access process for decoding a video sequence, encoded data for multiple images is transmitted in an order other than the display order of those images, thereby increasing compression. Images encoded without referencing other images (so-called intraframes) can then be encoded systematically, thereby allowing the decoding of the video sequence to begin from multiple points in the encoded stream.

[0052] For that purpose, please refer to the article "Overview of the High Efficiency Video Coding (HEVC) Standard" by G.S. Ullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand in IEEE Transactions on Circuits and Systems for Video Technology, vol.22, no.12, pp.1649-1668, Dec.2012.

[0053] The encoding-decoding process pair may also be selected to obtain the best compression-distortion balance.

[0054] To that end, it is possible to apply multiple encoding-decoding process pairs to content data B and select the set that achieves a better compression-distortion balance.

[0055] Alternatively, it is possible to identify the type of content (for example, by analyzing content data B) and select an encoding-decoding process pair appropriate to the identified type.

[0056] The encoding-decoding process pair may also be selected depending on the computing power available to the electronic decoding device. Information representing this computing power may have been transmitted in advance from the electronic decoding device to the electronic encoding device (for example, received by the electronic encoding device in step E2 above).

[0057] It may also be possible to combine various criteria for selecting encoding-decoding process pairs.

[0058] Once an encoding-decoding process pair is selected, the processor 4 proceeds to step E6, which configures the parallel processing unit 6 in such a way that the parallel processing unit 6 can perform the selected encoding method.

[0059] Step E6 includes the instantiation of the coding artificial neural network 8, which is used by a particularly selected coding process, within a parallel processing unit 6.

[0060] This instantiation involves the following steps, - A step of reserving memory space within the parallel processing unit 6 necessary for implementing an encoded artificial neural network, and / or - A step of programming the parallel processing unit 6 with weights W and activation functions that define the encoded artificial neural network 8, and / or - A step of loading at least a portion of content data B into the local memory of parallel processing unit 6. It may include.

[0061] The method in Figure 2 therefore includes step E8, which performs the encoding process, i.e., applying content data B to the input of the encoding artificial neural network 8 (or, in other words, activating the encoding artificial neural network 8 by taking content data B as input).

[0062] Therefore, step E8 allows us to generate compressed data C (in this case, the output of the encoded artificial neural network 8).

[0063] The following steps relate in particular to the encoding (i.e., preparation) of a data stream that includes compressed data C and is directed to an electronic decoding device (e.g., electronic decoding device 10 as described in relation to Figure 7).

[0064] The method therefore includes step E10 of encoding a first header portion Fc which includes, in particular, data characteristics of the representation format of audio or video content (in this case, for example, data associated with the format of the video sequence being encoded).

[0065] These data forming the first header section Fc indicate, for example, the image size (in pixels), frame rate, binary depth of luminance information, and binary depth of chrominance information. These data are constructed, for example, based on the format data P described above (after any possible reformatting).

[0066] The method in Figure 2 then proceeds to step E12, which identifies the availability of a decoding artificial neural network (used by the decoding process selected in step E4) for an electronic decoding device (e.g., electronic decoding device 10, described later with respect to Figure 7) capable of decoding the data stream.

[0067] This identification may be based on the list received in step E2, in which case processor 4 determines whether the decoding artificial neural network to be used by the decoding process selected in step E4 belongs to the list received in step E2. (Naturally, in embodiments in which the encoding-decoding process pair is systematically selected to correspond to the decoding artificial neural networks available to the electronic decoding device, step E12 may be omitted, and the method proceeds to step E14.)

[0068] According to a possible embodiment, if there is no information regarding the availability of the decoding artificial neural network for the electronic decoding device, the method proceeds to step E16 (so that data describing the decoding artificial neural network is sent to the electronic decoding device, as already described).

[0069] If, in step E12, processor 4 determines that the decoding artificial neural network is available for the electronic decoding device (arrow P), the method proceeds to step E14 below.

[0070] If processor 4 determines in step E12 that the decoding artificial neural network is not available to the electronic decoding device (arrow N), the method proceeds to step E16.

[0071] Alternatively, the choice between step E14 or step E16 as the step following step E12 may depend on other criteria, for example, a dedicated indicator stored in the electronic encoding device 2 (which may be adjustable by the user via the user interface of the electronic encoding device 2) or a selection made by the user (for example, obtained via the user interface of the electronic encoding device 2).

[0072] Processor 4 proceeds to step E14, which encodes a second header portion containing the indicator IND and a third header portion containing the identifier Inn of the decoded artificial neural network.

[0073] In step E14, the indicator IND encoded in the data stream indicates that the decoded artificial neural network belongs to a predetermined set of artificial neural networks, in this case, a set of artificial neural networks available (or accessible) to the electronic decoding device (i.e., the set of artificial neural networks in the list received in step E2).

[0074] The identifier Inn for a decoded artificial neural network is, by transformation, an identifier that defines this decoded artificial neural network (particularly shared by the electronic encoding device and the electronic decoding device) within, for example, the predetermined set described above.

[0075] Processor 4 proceeds to step E16, which encodes a second header portion containing the indicator IND' and a third header portion containing data Rc describing the decoded artificial neural network.

[0076] In step E16, the encoded indicator IND' in the data stream indicates that the decoded artificial neural network is encoded in the data stream, i.e., represented by the descriptive data Rc described above.

[0077] The decoded artificial neural network is encoded (i.e., represented) by descriptive data (or data encoding the decoded artificial neural network) Rc, for example, according to a standard such as MPEG-7 part 17 or a format such as JSON.

[0078] For that purpose, please refer to the article "DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression" by S. Wiedemann et al. in the Proceedings of the 36th International Conference on Machine Learning, Long Beach, California, PMLR 97, 2019, or the article "Compact and Computationally Efficient Representation of Deep Neural Networks" by S. Wiedemann et al. in IEEE Transaction on Neural Networks and Learning Systems (Vol. 31, Iss. 3), March 2020.

[0079] In step E14 following step E16, the method shown in Figure 2 proceeds to step E18, which identifies the possibility that the electronic decoding device will perform the decoding process using a decoding artificial neural network.

[0080] Processor 4 identifies this possibility, for example, by determining whether the electronic decoder is designed to perform this decoding process (perhaps through a prior exchange between the electronic coding device and the electronic decoding device) or whether it includes software suitable for the electronic decoder to perform this decoding process when the software is executed by the processor of the electronic decoder.

[0081] If processor 4 determines that the electronic decryption device can perform the decryption process, the method proceeds to step E22 described below.

[0082] If processor 4 determines that the electronic decryption device is unable to perform the decryption process, the method performs step E20 (described below) (before proceeding to step E22).

[0083] Alternatively, the choice of whether or not to perform step E20 (before performing step E22) may depend on other criteria, for example, a dedicated indicator stored in the electronic coding device 2 (and possibly adjustable by the user via the user interface of the electronic coding device 2) or a selection made by the user (for example, obtained via the user interface of the electronic coding device 2).

[0084] In step E20, processor 4 encodes a fourth header portion in the data stream that contains a computer program (exe) (or code) executable by the processor of the electronic decryption device. (The use of the computer program (exe) within the electronic decryption device will be described later with reference to Figure 8.)

[0085] To ensure that the computer program is adapted for execution within the electronic decoding device, it is selected from a library based on information about the hardware configuration of the electronic decoding device (e.g., information received through previous exchanges between the electronic encoding device 2 and the electronic decoding device).

[0086] The processor 4 then proceeds to step E22, which encodes the compressed stream Fnn based on the compressed data C obtained in step E8.

[0087] In this regard, the above explanation shows that step E8 is described before the steps for encoding the header Fet (steps E10-E20). However, in practice, step E8 can be performed immediately before step E22.

[0088] In particular, if step E8 can process only a portion of the audio or video content to be compressed (for example, if step E8 processes a block or component of a video sequence or an image to be compressed), it is possible to repeat the execution of steps E8 (to obtain compressed data relating to a contiguous portion of the content) and E22 (to encode the compressed data obtained in the data stream).

[0089] The processor 4 can therefore construct a complete data stream, including the header Fet and the compressed stream Fnn, in step E24.

[0090] The complete data stream is constructed in such a way that the header Fet and the compressed stream Fnn are individually identifiable.

[0091] In a conceivable embodiment, the header Fet includes an indicator of the start of the compressed stream Fnn within the complete data stream. This indicator, for example, represents the position of the beginning of the compressed stream Fnn from the start of the compressed data stream using bits. (In other words, the header in this case has a predetermined fixed length.)

[0092] Other methods for identifying the header Fet and compressed stream Fnn can also be considered as alternatives, such as a marker (i.e., a combination of binaries used to indicate the start of the compressed stream Fnn, the use of which is hidden in the rest of the data stream or at least in the header Fet).

[0093] Figures 3-6 show examples of complete data streams that can be obtained using the method in Figure 2.

[0094] As already explained, these data streams include a header Fet and a compressed stream Fnn.

[0095] In the case shown in Figure 3 (corresponding to the situation where step E14 is executed but step E20 is not), the header is: - The first part Fc includes the data characteristics of the representation format of audio or video content. - A second part including an indicator IND that shows the decoding artificial neural network belongs to a given set of artificial neural networks. - The third part containing the identifier Inn of the decoded artificial neural network Includes.

[0096] In the case shown in Figure 4 (corresponding to a situation where step E16 has been performed but step E20 has not), the header is: - The first part Fc includes the data characteristics of the representation format of audio or video content. - The second part includes an indicator IND' showing that the decoding artificial neural network is encoded in the data stream. - The third part contains the data Rc (the data that encodes it) that describes the decoding artificial neural network. Includes.

[0097] In the case shown in Figure 5 (corresponding to the situation where steps E14 and E20 have been performed), the header is: - The first part Fc includes the data characteristics of the representation format of audio or video content. - A second part including an indicator IND that shows the decoding artificial neural network belongs to a given set of artificial neural networks. - The third part containing the identifier Inn of the decoded artificial neural network, - The fourth part, which includes the computer program EXE. Includes.

[0098] In the case shown in Figure 6 (corresponding to the situation where steps E16 and E20 have been performed), the header is: - The first part Fc includes the data characteristics of the representation format of audio or video content. - The second part includes an indicator IND showing that the decoding artificial neural network is encoded in the data stream. - The third part contains the data Rc (the data that encodes it) that describes the decoding artificial neural network. - The fourth part, which includes the computer program EXE. Includes.

[0099] The data stream constructed in step E24 can be encapsulated in a known transmission format, such as the "Packet-Transport System" format or the "Byte-Stream" format.

[0100] In the "Packet-Transport System" format (for example, as proposed by the RTP protocol), data is encoded in identifiable packets and transmitted over the communication network. The network can easily identify the boundaries of the data (images, sets of images, and in this case, the header Fet and compressed stream Fnn) using packet identification information provided by the network layer.

[0101] In the case of the "Byte-Stream" format, there are no packets in particular, and the construction of step E24 must be able to identify the boundaries of the relevant data (e.g., between the parts of the stream corresponding to each image and, in this case, between the header Fet and the compressed stream Fnn) using additional means such as the use of a network extraction layer (NAL) unit, in which case it is possible to identify the boundaries between data by a unique binary combination (e.g., 0x00000001)).

[0102] The complete data stream constructed in step E24 may then be transmitted in step E26 to the electronic decoding device 26 (by means of communication not shown and / or through at least one communication network) or stored in the electronic encoding device 2 (for later transmission or, alternatively, for later decoding within the electronic encoding device itself (in this case, designed to further perform the decoding method described later with respect to Figure 8)).

[0103] When audio or video content consists of multiple parts (for example, multiple image sequences if the content is a video sequence), the method in steps E4-E24 can probably be performed in such a way that for each part of the content (for example, for each image sequence), a data stream is obtained as shown in one of Figures 3-6. Thus, the compressed stream Fnn for each image sequence is appropriate for the image sequence in question and can probably be decoded using a different artificial neural network than the one used for other image sequences, as will be discussed later. These artificial neural networks may have the same structure (and only differ in their weights and / or the activation function that defines a particular artificial neural network).

[0104] Figure 7 shows an electronic decoding device 10 that uses at least one artificial neural network 18.

[0105] This electronic decoding device 10 includes a receiving unit 11, a processor 14 (e.g., a microprocessor), and a parallel processing unit 16, such as a graphics processing unit, i.e., a GPU, or a tensor processing unit, i.e., a TPU.

[0106] The receiving unit 11 is, for example, a communication circuit (such as a radio frequency communication circuit), which receives data (and in particular, encoded data streams) from an external electronic device such as an electronic coding device 2, and enables communication of this data to the processor 14 (to which the receiving unit 11 is connected, for example, by a bus).

[0107] The electronic decoding device 10 also includes a storage unit 12, such as memory (which may be rewritable non-volatile memory) or a hard drive. Although the storage unit 12 is shown as a separate element from the processor 14 in Figure 7, the storage unit 12 may, as an alternative, be integrated into (i.e., included in) the processor 14.

[0108] In this case, the processor 14 is adapted to execute, for example, multiple instructions of a computer program stored in the memory unit 12 in sequence.

[0109] The parallel processing unit 16, after being configured by the processor 14, is designed to implement an artificial neural network 18. For this purpose, the parallel processing unit 16 is designed to perform multiple operations of the same type in parallel at any given time.

[0110] As schematically shown in Figure 7, the processor 14 receives a data stream (for example, via communication means not shown in the electronic decoding device 10) that includes a first dataset, in this case a header Fet, and a second dataset representing audio or video content, in this case a compressed stream Fnn.

[0111] As will be described later, the artificial neural network 18 is used to retrieve audio or video content corresponding to the initial audio or video content B within the framework of processing the second dataset (i.e., compressed data Fnn in this case).

[0112] The memory unit 12 can store multiple parameter sets, each parameter set defining a decoding artificial neural network. As will be described later, in this case, the processor 14 can configure the parallel processing unit 16 using a specific parameter set from among these parameter sets, in a manner that enables the parallel processing unit 16 to implement the artificial neural network defined by this specific parameter set.

[0113] The memory unit 12 may store, in particular, a first parameter set that defines a first artificial neural network forming a random access decoder and / or a second parameter set that defines a second artificial neural network forming a low-latency decoder.

[0114] In this case, the electronic decoding device 10 has pre-decoding options for both situations where random access to the content is desirable and situations where the content will be displayed without delay.

[0115] Here, with reference to Figure 8, we will explain a decoding method performed within the electronic decoding device 10 using an artificial neural network 18 implemented by a parallel processing unit 16.

[0116] This method may begin with a selective step of the electronic decoding device 10 sending a list L of artificial neural networks available to the electronic decoding device 10 to a device that controls the transmission of the data stream to be decoded. The data stream transmission control device may be, for example, the electronic coding device 2. (In this case, the electronic coding device 2 receives this list L in step E2 described above with respect to Figure 2.) Alternatively, the data stream transmission control device may also be a dedicated server cooperating with the electronic coding device 2.

[0117] An artificial neural network accessible by the electronic decoding device 10 is an artificial neural network that the electronic decoding device 10 can access either by storing a parameter set that defines the artificial neural network it is related to (as described above), or by connecting to a remote electronic device such as a server (as described later).

[0118] The method in Figure 8 includes step E52 of receiving a data stream (by the electronic decoding device 10, more precisely the receiving unit 11 in this case) which includes a first parameter set, namely the header Fet, and a second parameter set, namely the compressed stream Fnn. The receiving unit 11 transmits the received data stream to the processor 14.

[0119] The processor 14 then proceeds to step E54, which identifies the first dataset (header Fet) and the second dataset (compressed stream Fnn) in the received data stream, for example, by an indicator of the start of the compressed stream (already described in the description of step E24).

[0120] In step E54, the processor 14 may also identify the first dataset (header), i.e., different parts within the header Fet, which, as shown in Figures 3-6 above, are the first part Fc (containing data representing the format characteristics of the content encoded by the data stream), the second part (indicator IND or IND'), the third part (Inn or encoded data Rc), and possibly the fourth part (computer program Exe).

[0121] In step E54, if executable instructions (e.g., instructions for a computer program Exe) are identified (i.e., detected) within the first data, the processor 14 may, in step E56, execute these executable instructions and begin executing at least certain steps of processing the data from the first dataset. These instructions may be executed by the processor 14 or, alternatively, by a virtual machine instantiated in the electronic decoding device 10.

[0122] The method in Figure 7 follows step E58, in which data Fc, which is a characteristic of the representation format of the audio or video content, is decoded in a manner that obtains the characteristics of this format. In the case of video content, for example, decoding of data portion Fc makes it possible to obtain the image size (in pixels), and / or frame rate, and / or the binary depth of the luminance information, and / or the binary depth of the chrominance information.

[0123] The processor 14 then proceeds to step E60, which decodes the indicators IND, IND' contained in the second part of the header Fet (by executing the instructions identified in the first dataset in step E54 as described above in a particular embodiment).

[0124] If decoding of indicators IND and IND' in the received data stream indicates that the artificial neural network 18 used for decoding belongs to a predetermined set of artificial neural networks (i.e., if an indicator in the first dataset is an indicator IND indicating that the decoding artificial neural network 18 belongs to a predetermined set of artificial neural networks), the method proceeds to step E62 described below.

[0125] If decoding of indicators IND and IND' in the received data stream indicates that the artificial neural network 18 used for decoding is encoded within the data stream (i.e., if an indicator in the first dataset is indicator IND' indicating that the decoding artificial neural network 18 is encoded within the data stream), the method proceeds to step E66 described below.

[0126] In step E62, the processor 14 proceeds to decode the identifier Inn (here contained in the third part of the header Fet) (by executing the instruction identified in the first dataset in step E54 as described above in a particular embodiment). As described above, this identifier Inn is, for example, an identifier that specifies the decoded artificial neural network 18 in the given set of artificial neural networks described above.

[0127] The processor 14 can then proceed to step E64, which reads from within the memory unit 12 the parameter set associated with the decoded identifier Inn (by executing the instructions identified in the first dataset in step E54 as described above in a particular embodiment) (this parameter set defines the artificial neural network identified by the decoded identifier Inn).

[0128] In a conceivable embodiment, the processor 14 may be configured to generate an error message if there is no data (in particular parameters) relating to this artificial neural network identified by the decoded identifier Inn (in this case, within the memory unit 12).

[0129] Alternatively (or if no parameter set is stored in the memory unit 12 for the artificial neural network identified by the decoded identifier Inn), the electronic decoder 10 may send a request for a parameter set to a remote server (by executing an instruction identified in the first dataset in step E54 as described above in a particular embodiment) (this request may include, for example, the encoded identifier Inn), and in step E64, receive as a response a parameter set that defines the artificial neural network identified by the decoded identifier Inn.

[0130] Next, the method continues in step E68, which will be described later.

[0131] In step E66, the processor 14 proceeds to decode the data Rc (which is contained in this case in the third part of the header Fet) that describes the artificial neural network 18 (by executing instructions identified in the first dataset in step E54 as described above in a particular embodiment).

[0132] As mentioned above, these descriptive data (or encoded data) Rc are encoded according to a standard such as MPEG-7 part 17 or a format such as JSON.

[0133] By encoding the descriptive data Rc, it becomes possible to obtain the parameters that define the artificial neural network used to decode the data from the second dataset (i.e., the compressed stream Fnn data in this case).

[0134] The method, in this case, follows step E68 and is described below.

[0135] The processor 14 then proceeds to step E68, which configures the parallel processing unit 16 with parameters (parameters obtained in step E64 or step E66) that define the decoding artificial neural network 18 (by executing instructions identified in the first dataset in step E54 as described above in a particular embodiment) in such a way that the parallel processing unit 16 can implement the decoding artificial neural network 18.

[0136] This configuration step E68 specifically involves instantiating the decoding artificial neural network 18 within the parallel processing unit 16, using the parameters obtained in step E64 or step E66.

[0137] This instantiation is particularly important. - A step of reserving memory space within the parallel processing unit 16 necessary for implementing the decoding artificial neural network 18, and / or - A step of programming the parallel processing unit 16 with parameters that define the decoding artificial neural network 18 (e.g., including weights W' and activation function) (parameters obtained in step E64 or step E66), and / or - A step of loading at least a portion of the data from the second dataset (i.e., at least a portion of the data from the compressed stream Fnn) into the local memory of the parallel processing unit 16. It may include.

[0138] As can be seen from the above explanation of steps E58-E68, the data of the first dataset Fet is therefore processed by processor 14 in this manner.

[0139] The processor 14 can then, in step E70, apply (i.e., present) the data from the second dataset (in this case, the data from the compressed stream Fnn) to the artificial neural network 18 implemented by the parallel processing unit 16 in such a manner that the data is processed at least partially by a decoding process that uses the artificial neural network 18.

[0140] In the examples described herein, the artificial neural network 18 receives data from a second dataset Fnn as input and generates, as output, a representation I of encoded content adapted for playback on an audio or video playback device. In other words, at least certain data from the second dataset Fnn is applied to the input layer of the artificial neural network 18, and the output layer of the artificial neural network 18 generates the above-mentioned representation I of encoded content. In the case of video content (including a single image or a sequence of images), the artificial neural network 18 therefore generates, as output (i.e., in its output layer), at least one matrix representation I of the image.

[0141] In a particular embodiment, in order to process specific data (e.g., corresponding to a block or an image) in a compressed stream Fc, the artificial neural network 18 may receive as input at least certain data generated by the artificial neural network 18 during processing of preceding data (e.g., corresponding to a preceding block or an image) in the compressed stream Fc. In this case, it proceeds to step E72, in which the data generated at the output of the artificial neural network 18 is reinjected into the input of the artificial neural network 18.

[0142] Furthermore, according to other possible embodiments, the decoding process can utilize multiple artificial neural networks, as already mentioned with respect to the processing of content data B.

[0143] The data from the second set (in this case, at least specific data from the compressed stream Fnn) is then processed using an artificial neural network 18 implemented by the parallel processing unit 16, by a process that depends on the portion of the data from the first set (in this case, a process that depends on the identifier Inn or encoded data Rc).

[0144] The processor 14 then determines in step 74 whether the processing of the compressed stream Fnn by the artificial neural network 18 has finished.

[0145] If the specific result is negative (N), the method returns to step E70 and applies the remaining data from the compressed stream Fnn to the artificial neural network 18.

[0146] If the result is positive (P), the method proceeds to step E76, in which the processor 14 determines whether there is any data remaining in the received data stream that needs to be processed.

[0147] If the identification result in step E76 is negative (N), the method terminates in step E78.

[0148] If the result of step E76 is positive (P), the method returns to step E52 and processes the new portion of the data stream shown in one of Figures 3-6.

[0149] As shown above as an iteration of encoding steps E4-E24, this other part of the data stream also includes the first dataset and a second dataset representing other audio or video content (e.g., in the case of video content, other image groups relating to the content representation format used). In this case, as described above in steps E54-E66, another artificial neural network can be identified based on a particular one of these first data, and the parallel processing unit 16 may then be configured to implement this other artificial neural network (according to step E68 above). The data from the second dataset of this other part of the data stream (e.g., relating to the other image groups mentioned above) may then be decoded by this other artificial neural network (as described above in step E70).

[0150] The aforementioned other artificial neural network may have the same structure as that of the aforementioned artificial neural network 18, thereby simplifying the steps that constitute the parallel processing unit 16 (for example, only the weights and / or activation functions that define the current artificial neural network are updated).

Claims

1. A method for decoding a data stream by an electronic device (10) including a first processor (14) and a parallel processor (16) configured to perform multiple operations of the same kind in parallel at a given time, wherein the data stream includes a first dataset (Fet) and a second dataset (Fnn) representing audio or video content, and the decoding method is - A step of identifying the first dataset and the second dataset in the data stream by utilizing a predetermined binary length for the first dataset and / or the second dataset, - A step of obtaining an identifier or data describing an artificial neural network by processing the data from the first dataset (Fet) with the first processor (14) (E56, E58, E60, E62, E66), - The steps of obtaining the audio or video content by processing the data from the second dataset (Fnn) using an artificial neural network (18) implemented by the parallel processor (16) (E70), A decoding method characterized by including the following.

2. The decoding method according to claim 1, comprising the step (E68) of configuring the parallel processor (16) according to at least a portion of the data from the first dataset (Fet).

3. The decoding method according to claim 2, wherein the first dataset (Fet) includes data (Rc) describing the artificial neural network (18), and in the configuring step (E68), the first processor (14) configures the parallel processor (16) based on the describing data (Rc).

4. The decoding method according to claim 2, wherein the electronic device (10) includes a storage unit (12) for storing a plurality of parameter sets that each define a plurality of artificial neural networks, the first dataset (Fet) includes an identifier (Inn), and in the configuring step (E68), the first processor (14) configures the parallel processor (16) based on the parameter set associated with the identifier (Inn) from the plurality of parameter sets.

5. The decoding method according to any one of claims 1 to 4, wherein the data stream further includes an executable instruction (Exe) within the electronic device (10), and the processing of the data from the first dataset (Fet) is performed at least partially by the execution of at least a portion of the instruction (Exe).

6. The decoding method according to claim 5, which is directly or indirectly dependent on claim 2, wherein the step (E68) comprising the parallel processor (16) is performed by executing at least a portion of the instruction (Exe).

7. The decoding method according to any one of claims 1 to 6, wherein the first dataset (Fet) includes data (Fc) representing the format characteristics of the video content encoded by the data stream.

8. The decoding method according to any one of claims 1 to 7, wherein the processing of the data from the second dataset (Fnn) generates at least one matrix representation (I) of at least a portion of the image.

9. The decoding method according to any one of claims 1 to 8, wherein the artificial neural network (18) receives data from the second dataset (Fnn) as input.

10. The decoding method according to any one of claims 1 to 9, wherein the artificial neural network (18) receives previously generated data as input as the output of the artificial neural network (18).

11. An electronic device (10) for decoding a data stream comprising a first dataset (Fet) and a second dataset (Fnn) representing audio or video content, - A first processor (14) adapted to identify the first dataset and the second dataset in the data stream by utilizing a predetermined binary length for the first dataset and / or the second dataset, and to obtain an identifier or data describing an artificial neural network by processing data from the first dataset (Fet), - A parallel processor (16) configured to perform multiple operations of the same type in parallel at a specific point in time, wherein the parallel processor (16) is adapted to acquire the audio or video content by processing data from the second dataset (Fnn) using an artificial neural network (18) implemented by the parallel processor (16), An electronic device (10) including the electronic device.

12. The electronic device according to claim 11, wherein the first processor (14) is adapted to configure the parallel processor (16) according to at least a portion of the data from the first dataset (Fet).

13. The electronic device according to claim 12, comprising a storage unit (12) for storing a plurality of parameter sets that define a plurality of artificial neural networks, wherein the first dataset (Fet) includes an identifier (Inn), and the first processor (14) is adapted to configure the parallel processor (16) based on the parameter set associated with the identifier (Inn) from the plurality of parameter sets.

14. The electronic device according to any one of claims 11 to 13, wherein the parallel processor (16) is adapted to generate at least one matrix representation (I) of at least a portion of an image.

Citation Information

Patent Citations

  • Method and apparatus of neural network based processing in video coding

    US20180249158A1

  • Low rank matrix compression

    US20180293758A1

  • Coding method, decoding method, information processing method, coding device, decoding device, and information processing system

    WO2019131880A1

  • Method, apparatus and stream for volumetric video format

    WO2019191205A1