Video decoding using neural networks
By introducing indicators and identifiers into the data stream, the problem of identifying and accessing neural networks in artificial neural network decoding methods is solved, enabling more efficient video or audio content decoding and adapting to the needs of different devices and computing capabilities.
Patent Information
- Application Number
- CN202180050577.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-17
- Filing Date
- 2021-07-13
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-07-13
AI Technical Summary
In existing technologies, artificial neural network decoding methods struggle to effectively identify and access appropriate neural networks for decoding video or audio content, resulting in low decoding efficiency.
By introducing indicators and identifiers into the data stream, it is determined whether the artificial neural network belongs to a predetermined set or is encoded in the data stream, and this information is used for decoding, including reading parameters from storage units or obtaining parameters from remote servers, generating error messages, or transmitting computer programs to achieve decoding.
It improves decoding efficiency and flexibility, ensuring that electronic decoding devices can select the appropriate neural network for decoding and adapt to different computing capabilities and decoding needs.
Smart Images

Figure CN116325740B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of audiovisual content decoding.
[0002] In particular, it relates to a method for decoding a data stream, as well as an associated device and data stream. BACKGROUND
[0003] It has been proposed to compress data representative of video content with the aid of an artificial neural network. The decoding of the compressed data can then be performed with the aid of another artificial neural network, for example in the article "DVC: An End-to-end Deep Video Compression Framework" by Guo Lu et al., in IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pages 10998-11007. SUMMARY
[0004] In this context, the application proposes a method for decoding a data stream comprising an indicator and data representative of audio or video content, the method comprising the steps of:
[0005] - decoding the indicator to determine whether the artificial neural network used to decode the representative data is encoded in the data stream or belongs to a predetermined set of artificial neural networks;
[0006] - decoding the representative data by means of the artificial neural network.
[0007] Such an indicator thus allows the decoder to know the method it will be able to access to the artificial neural network used to decode the data representative of the content.
[0008] If it is determined by decoding the indicator that the artificial neural network belongs to said predetermined set, the decoding method can comprise decoding an identifier of the neural network.
[0009] The application also proposes in a manner that is original per se a method for decoding a data stream comprising data representative of audio or video content, and an identifier of an artificial neural network of a predetermined set of artificial neural networks, the method comprising the steps of:
[0010] - decoding the identifier;
[0011] - decoding the representative data with the aid of the artificial neural network specified by the decoded identifier.
[0012] The decoding of this identifier thus indicates which artificial neural network must be used in the predetermined set of artificial neural networks, for example a predetermined set of artificial neural networks accessible to the electronic decoding device, as explained below.
[0013] According to a first possibility, the decoding method can comprise reading in a storage unit the parameters of the artificial neural network identified by the decoded identifier.
[0014] It can be provided that this storage unit stores a first set of parameters representative of a first artificial neural network forming a random access decoder and / or a second set of parameters representative of a second artificial neural network forming a low latency decoder.
[0015] In addition, the decoding method can comprise a step of generating an error message in the absence of data related to the artificial neural network identified by the decoded identifier, for example within a storage unit such as the one mentioned above.
[0016] According to a second possibility, the decoding method can comprise receiving from a remote server the parameters of the artificial neural network identified by the decoded identifier.
[0017] In addition, the decoding method can comprise decoding the data encoding the artificial neural network included in the data stream in order to obtain the parameters of the artificial neural network if it is determined by decoding the indicator that the artificial neural network is encoded in the data stream.
[0018] According to a possible embodiment, the decoding method can comprise a (preliminary) step of transmitting a list of artificial neural networks to a device for controlling the transmission of the data stream. In certain embodiments, this list of artificial neural networks can correspond to the predetermined set mentioned above. In other words, in this case, the predetermined set of artificial neural networks can be formed (or by) the artificial neural networks of this list.
[0019] This content can in fact be a first part of a video sequence, wherein this video sequence can comprise said first part and a second part.
[0020] In this case, the decoding producing said representative data of said first part, the method can further comprise a step of decoding other data by means of another artificial neural network to produce a second part.
[0021] According to a possible embodiment, the other artificial neural network has the same structure as said artificial neural network, which simplifies the updating of the artificial neural networks within the electronic decoding device.
[0022] The first part and the second part mentioned above form, for example, two groups of images respectively, for a format of content used.
[0023] The application also proposes a decoding device comprising:
[0024] - a unit for receiving a data stream comprising an indicator and data representative of audio or video content;
[0025] - a decoding component designed to determine, by decoding said indicator, whether the artificial neural network used to decode said representative data belongs to a predetermined set of artificial neural networks or is encoded in the data stream, and to decode said representative data by means of the artificial neural network.
[0026] The application also proposes a decoding device comprising:
[0027] - a unit for receiving a data stream comprising data representative of audio or video content, and an identifier specifying an artificial neural network in a predetermined set of artificial neural networks;
[0028] - a decoding component designed to decode the identifier and to decode said representative data by means of the artificial neural network specified by the decoded identifier.
[0029] In the embodiments described below, such a decoding component comprises a processor designed or programmed to decode the indicator and / or the identifier, and / or a parallelization processing unit designed to perform, at a given time, a plurality of operations of the same type in parallel and to implement the artificial neural network described above to decode the representative data.
[0030] The application also proposes a data stream comprising data representative of audio or video content, and an indicator indicating whether the artificial neural network used to decode said representative data is encoded in the data stream or belongs to a predetermined set of artificial neural networks.
[0031] Finally, the application proposes a data stream comprising data representative of audio or video content, and an identifier specifying, in a predetermined set of artificial neural networks, an artificial neural network used to decode said representative data.
[0032] Of course, the different features, alternatives and embodiments of the application can be associated with each other according to various combinations, as long as they are not mutually incompatible or exclusive. BRIEF DESCRIPTION OF DRAWINGS
[0033] Moreover, various other characteristics will become apparent from the appended description made with reference to the attached drawings showing non-limiting embodiments of the application, and among which:
[0034] - Figure 1 An electronic encoding device used within the framework of the application is shown;
[0035] -Figure 2 a flowchart showing the steps of the encoding method implemented within the electronic encoding device of Figure 1
[0036] - Figure 3 is a first example of a data stream obtained by the method of Figure 2
[0037] - Figure 4 is a second example of a data stream obtained by the method of Figure 2
[0038] - Figure 5 is a third example of a data stream obtained by the method of Figure 2
[0039] - Figure 6 is a fourth example of a data stream obtained by the method of Figure 2
[0040] - Figure 7 shows an electronic encoding device according to an embodiment of the application; and
[0041] - Figure 8 is a flowchart showing the steps of the decoding method implemented within the electronic decoding device of Figure 7 DETAILED DESCRIPTION
[0042] Figure 1 An electronic encoding device 2 is shown using at least one artificial neural network 8.
[0043] The electronic encoding device 2 comprises a processor 4, for example a microprocessor, and a parallelization processing unit 6, for example a graphic processing unit or GPU, or a tensor processing unit or TPU.
[0044] As schematically shown in Figure 1 , the processor 4 receives data P, B representing an audio or video content to be compressed, here the format data P and the content data are B.
[0045] The format data P is indicative of features of the representation format of the audio or video content, for example, in the case of a video content, the image size in pixels, the frame rate, the binary depth of the luminance information and the binary depth of the chrominance information.
[0046] The content data B forms a representation of the audio or video content, here uncompressed. For example, in the case of a video content, for each pixel of each image of the sequence of images, the content data comprises data representative of the luminance value of the pixel and / or data representative of the chrominance value of the pixel.
[0047] The parallelization processing unit 6 is designed to implement the artificial neural network 8 after it has been configured by the processor 4. To this end, the parallelization processing unit is designed to perform a plurality of operations of the same type in parallel at a given time.
[0048] As described below, the artificial neural network 8 is used within the framework of the processing of the content data B intended to obtain the compressed data C.
[0049] In the embodiments described here, the artificial neural network 8 produces the compressed data C at the output when the content data B is applied to the input of the artificial neural network 8.
[0050] The content data B applied to the input of the artificial neural network 8, that is to say to the input layer of the artificial neural network 8, can represent a block of an image, or a block of an image component (for example, a block of a luminance or chrominance component of the image, or a block of a color component), or an image of a video sequence, or a component (for example, a luminance or chrominance component, or a color component) of an image of a video sequence, or also a series of images of a video sequence.
[0051] For example, in this case, each of at least some of the neurons of the input layer can be provided to receive a pixel value of an image component, said value being represented by one of the data in the content data B.
[0052] As an alternative, the processing of the content data B can comprise the use of several artificial neural networks, as described in the above-mentioned article "DVC: An End-to-end Deep Video Compression Framework" by Guo Lu et al., published in the 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
[0053] An example of an encoding method implemented by the electronic encoding device 2 will now be described with reference to Figure 2 The method implemented by the electronic encoding device 2 comprises the following steps:
[0054] The memory connected to the processor 4 stores, for example, computer program instructions designed to implement, when these instructions are executed by the processor 4, at least some of the steps of the method of Figure 2 In other words, the processor 4 is programmed to implement at least some of the steps of the method of Figure 2 In other words, the processor 4 is programmed to implement at least some of the steps of the method of
[0055] The method of Figure 2 begins with an optional step E2 of receiving a list of artificial neural networks accessible by the electronic decoding device.
[0056] This list is for example received by the processor 4 of the electronic encoding device 2 directly from an electronic decoding device, for example in accordance with the electronic decoding device 10 described hereafter. Figure 7
[0057] As described hereafter, the artificial neural network accessible by the electronic decoding device is such that the electronic decoding device stores parameters defining the artificial neural network of interest, or can access these parameters by connection to a remote electronic device such as a server.
[0058] As an alternative, this list can be received by the processor 4 of the electronic device 2 from a remote server such as the server described above.
[0059] Figure 2 The method of the electronic decoding device 10 continues with a step E4 of selecting an encoding process-decoding process pair. As already pointed out for the encoding process, the encoding process and the decoding process each use at least one artificial neural network.
[0060] In the example described here, the encoding process is implemented by an encoding artificial neural network and the decoding process is implemented by a decoding artificial neural network.
[0061] The set formed by the encoding artificial neural network and the decoding artificial neural network, applying the output of the encoding artificial neural network to the input of the decoding artificial neural network, forms for example an autoencoder.
[0062] The encoding process-decoding process pair is for example selected among a plurality of predefined encoding process-decoding process pairs, that is to say here among a plurality of encoding artificial neural network-decoding artificial neural network pairs.
[0063] When the list of artificial neural networks accessible by the electronic decoding device is received in advance, as explained above at step E2, the encoding process-decoding process pair is for example selected among the encoding process-decoding process pairs for which the decoding process uses an artificial neural network present in the received list.
[0064] The encoding process-decoding process pair can also be selected according to an intended application, for example indicated by a user using a user interface (not shown) of the electronic encoding device 2. For example, if the intended application is a videoconference, the selected encoding process-decoding process pair comprises a low latency decoding process. In other applications, the selected encoding process-decoding process pair will comprise a random access decoding process.
[0065] In a low latency process for decoding a video sequence, the images of the video sequence are represented for example by encoded data that can be immediately transmitted and decoded; the data can then be transmitted in the display order of the video images, which in this case ensures a one-frame delay between encoding and decoding.
[0066] In a random access process for decoding a video sequence, the encoded data relating to a plurality of images are transmitted in an order different from the display order of the images, which allows an increase in compression. An encoded image that is not referenced to other images (so-called intra frame) can then be regularly encoded, which allows the decoding of the video sequence to be started from several positions in the encoded stream.
[0067] To this end, reference can be made to the article by G. J. Sullivan, J.-R. Ohm, W.-J. Han and T. Wiegand, "Overview of the High Efficiency Video Coding (HEVC) Standard", in IEEE Transactions on Circuits and Systems for Video Technology, December 2012, vol. 22, no. 12, pages 1649 to 1668.
[0068] The encoding process-decoding process pair can also be chosen so as to obtain the best possible compression-distortion compromise.
[0069] To this end, a plurality of encoding process-decoding process pairs can be applied to the content data B and the set that achieves a better compression-distortion compromise is chosen.
[0070] As an alternative, the type of content can be determined (for example by analyzing the content data B) and the encoding process-decoding process pair is chosen according to the determined type.
[0071] The encoding process-decoding process pair can also be chosen according to the computing power available at the electronic decoding device. Information representative of this computing power can have been previously transmitted from the electronic decoding device to the electronic encoding device (and received by the electronic encoding device for example at step E2 described above).
[0072] The different criteria for choosing the encoding process-decoding process pair can be combined together.
[0073] Once the encoding process-decoding process pair has been chosen, the processor 4 proceeds, at step E6, with the configuration of the parallelization processing unit 6 so that the parallelization processing unit 6 can implement the chosen encoding method.
[0074] This step E6 specifically comprises instantiating, within the parallelization processing unit 6, an instance of the encoding artificial neural network 8 used by the selected encoding process.
[0075] This instantiation can specifically comprise the following steps:
[0076] - reserving, within the parallelization processing unit 6, storage space necessary to implement the encoding artificial neural network; and / or
[0077] - programming the parallelization processing unit 6 with the weights W and the activation functions defining the encoding artificial neural network 8; and / or
[0078] - loading at least part of the content data B into the local memory of the parallelization processing unit 6.
[0079] Then, Figure 2 The method of the application comprises a step E8 of implementing the encoding process, that is to say, here, applying the content data B at the input of the encoding artificial neural network 8 (or in other words, activating the encoding artificial neural network 8 with the content data B as input).
[0080] Thus, the step E8 allows to produce compressed data C (here at the output of the encoding artificial neural network 8).
[0081] The following steps concern the encoding (that is to say, the preparation) of a data stream specifically containing the compressed data C and intended to be sent to an electronic decoding device (for example, the electronic decoding device 10 described with reference to Figure 7 ).
[0082] Thus, the method specifically comprises a step E10 of encoding a first header portion Fc comprising data characteristic of the format of representation of the audio or video content (here, for example, data linked to the format of the video sequence being encoded).
[0083] These data forming the first header portion Fc indicate, for example, the image size in pixels, the frame rate, the binary depth of the luminance information and the binary depth of the chrominance information. These data are built, for example, on the basis of the format data P described above (after potential reformatting).
[0084] Then, Figure 2 The method of the application continues with a step E12 of determining the availability of a decoding artificial neural network (used by the decoding process selected at the step E4) for an electronic decoding device (for example, the electronic decoding device 10 described with reference to Figure 7 ).
[0085] This determination can be made on the basis of the list received at step E2: in this case, the processor 4 determines whether the decoding artificial neural network used by the decoding process selected at step E4 belongs to the list received at step E2. (Of course, in the embodiment in which the encoding process-decoding process pairs are systematically selected so as to correspond to a decoding artificial neural network available to the electronic decoding device, step E12 can be omitted, and the method then continues with step E14.)
[0086] According to a possible embodiment, in the absence of information on the availability of the decoding artificial neural network to the electronic decoding device, the method continues with step E16 (in such a way that the data describing the decoding artificial neural network are transmitted to the electronic decoding device, as explained above).
[0087] If the processor 4 determines at step E12 that the decoding artificial neural network is available to the electronic decoding device (arrow P), the method then continues with step E14 hereafter.
[0088] If the processor 4 determines at step E12 that the decoding artificial neural network is not available to the electronic decoding device (arrow N), the method then continues with step E16 hereafter.
[0089] As an alternative, as a step following step E12, the choice between step E14 and step E16 can be made according to another criterion, for example according to a dedicated indicator stored within the electronic encoding device 2 (and possibly adjusted by the user through the user interface of the electronic encoding device 2), or according to a choice made by the user (for example obtained through the user interface of the electronic encoding device 2).
[0090] The processor 4 proceeds at step E14 to the encoding of a second header part comprising the indicator IND and of a third header part comprising here the identifier Inn of the decoding artificial neural network.
[0091] The indicator IND encoded in the data stream at step E14 indicates that the decoding artificial neural network belongs to a predetermined set of artificial neural networks, here to the set of artificial neural networks available (or accessible) to the electronic decoding device (that is to say, for example, to the set of artificial neural networks of the list received at step E2).
[0092] The identifier Inn of the decoding artificial neural network is an identifier that defines this decoding artificial neural network by convention (in particular shared by the electronic encoding device and the electronic decoding device), for example within the predetermined set mentioned above.
[0093] The processor 4 proceeds at step E16 to the encoding of a second header part comprising the indicator IND' and of a third header part comprising here the data Rc describing the decoding artificial neural network.
[0094] The indicator IND' encoded in the data stream in step E16 indicates that the decoded artificial neural network is encoded in the data stream, that is to say, is represented by means of the descriptive data Rc described above.
[0095] The decoded artificial neural network is encoded (that is to say, is represented by Rc) by the descriptive data (or data encoding the decoded artificial neural network) Rc, for example according to a standard such as MPEG-7 part 17 or a format such as JSON.
[0096] To this end, reference can be made to the article "DeepCABAC: Context-adaptive binary arithmetic coding for deep neural network compression" by S. Wiedemann et al. in Proceedings of the 36th International Conference on Machine Learning, PMLR 97, 2019, held in Long Beach, California, or to the article "Compact and Computationally Efficient Representation of Deep Neural Networks" by S. Wiedemann et al. in IEEE Transactions on Neural Networks and Learning Systems (Volume 31, Issue 3), March 2020.
[0097] After step E14, as after step E16, Figure 2 The method continues with step E18, determining the possibility for the electronic decoding device to implement the decoding process using the decoded artificial neural network.
[0098] For example, the processor 4 determines this possibility, for example by determining (potentially with the aid of a previous exchange between the electronic encoding device and the electronic decoding device) whether the electronic decoding device comprises a module or software designed to implement the decoding process, said software being suitable for implementing the decoding process by the electronic decoding device when executed by the processor of the electronic decoding device.
[0099] If the processor 4 determines that the electronic decoding device can implement the decoding process, the method continues with step E22 described below.
[0100] If the processor 4 determines that the electronic decoding device cannot implement the decoding process, the method performs step E20 described below (before going to step E22).
[0101] As an alternative, the choice of performing or not performing step E20 (before performing step E22) can depend on another criterion, for example according to a dedicated indicator stored in the electronic encoding device 2 (and possibly adjusted by the user through the user interface of the electronic encoding device 2), or according to a choice made by the user (for example obtained through the user interface of the electronic encoding device 2).
[0102] At step E20, the processor 4 encodes in the data stream a fourth header portion containing a computer program Exe (or code) executable by the processor of the electronic decoding device. (Hereinafter, reference is made to the computer program Exe as a computer program.) Figure 8 The use of the computer program Exe in the electronic decoding device is described.
[0103] To adapt to the execution in the electronic decoding device, the computer program is selected in the library, for example according to information related to the hardware configuration of the electronic decoding device (for example information received during a previous exchange between the electronic encoding device 2 and the electronic decoding device).
[0104] Then, the processor 4 goes to step E22, encoding in the data stream Fnn the compressed data C obtained at step E8.
[0105] It is observed in this respect that in the above description, step E8 is described before the steps of encoding the header Fet (steps E10 to E20). However, step E8 can actually be performed before step E22.
[0106] In particular, when step E8 allows to process only a part of the audio or video content to be compressed (for example when step E8 performs the processing of a block, or component or image of the video sequence to be compressed), the implementation of steps E8 (to obtain the compressed data related to a successive part of the content) and E22 (to encode in the data stream the compressed data obtained) can be repeated.
[0107] Thus, the processor 4 can construct at step E24 a complete data stream comprising the header Fet and the compressed stream Fnn.
[0108] The complete data stream is constructed in such a way that the header Fet and the compressed stream Fnn can be identified separately.
[0109] According to a possible embodiment, the header Fet contains an indicator of the beginning of the compressed stream Fnn in the complete data stream. This indicator is for example the position of the beginning of the compressed stream Fnn from the position of the beginning of the complete data stream in bits. (In other words, in this case, the header has a predetermined fixed length.)
[0110] Other means for identifying the header Fet and the compressed stream Fnn can be considered as alternatives, such as a marker (that is to say, a binary combination used to indicate the beginning of the compressed stream Fnn and forbidden to be used in the rest of the data stream or at least in the header Fet).
[0111] In Figures 3 to 6 the example shown in Figure 2 illustrates an example of a complete data stream that can be obtained by the method of
[0112] As mentioned above, these data streams comprise a header Fet and a compressed stream Fnn.
[0113] In Figure 3 the case corresponding to the case where step E14 has been implemented and step E20 has not yet been implemented, the header comprises:
[0114] - a first part Fc comprising data characteristic of the format of representation of the audio or video content;
[0115] - a second part comprising an indicator IND indicating that the decoding artificial neural network belongs to a predetermined set of artificial neural networks; and
[0116] - a third part comprising an identifier Inn of the decoding artificial neural network.
[0117] In Figure 4 the case corresponding to the case where step E16 has been implemented and step E20 has not yet been implemented, the header comprises:
[0118] - a first part Fc comprising data characteristic of the format of representation of the audio or video content;
[0119] - a second part comprising an indicator IND' indicating that the decoding artificial neural network is encoded in the data stream; and
[0120] - a third part comprising data Rc describing the decoding artificial neural network (here the data encoding).
[0121] In Figure 5 the case corresponding to the case where steps E14 and E20 have been implemented, the header comprises:
[0122] - a first part Fc comprising data characteristic of the format of representation of the audio or video content;
[0123] - a second part comprising an indicator IND indicating that the decoding artificial neural network belongs to a predetermined set of artificial neural networks; and
[0124] - a third part comprising an identifier Inn of the decoding artificial neural network.
[0125] - a fourth part comprising a computer program Exe.
[0126] In the case where steps E16 and E20 have been implemented (corresponding to the case where the header comprises: Figure 6
[0127] - a first part Fc comprising data characteristic of the representation format of the audio or video content;
[0128] - a second part comprising an indicator IND' indicating that a decoding artificial neural network is encoded in the data stream; and
[0129] - a third part comprising data Rc describing a decoding artificial neural network (here the data encoding); and
[0130] - a fourth part comprising a computer program Exe.
[0131] The data stream constructed at step E24 can be encapsulated in a transmission format known per se, such as the format "Packet Transmission System" or the format "Byte Stream".
[0132] In the case of the format "Packet Transmission System" (for example proposed by the RTP protocol), the data are coded by identifiable packets and transmitted on a communication network. Using the packet identification information provided by the network layer, the network can easily identify the boundaries of the data (images, groups of images and here the header Fet and the compressed stream Fnn).
[0133] In the format "Byte Stream", there is no specific packet and the construction of step E24 must allow the use of additional means (such as the use of Network Abstraction Layer (NAL) units) to identify the boundaries of the relevant data (such as the boundaries between the stream parts corresponding to each image and here the boundaries between the header Fet and the compressed stream Fnn), where the only binary combination (for example 0x00000001) makes it possible to identify the boundaries between the data.
[0134] The complete data stream constructed at step E24 can then be transmitted at step E26 to the electronic decoding device 26 (by means of communication means not shown and / or by at least one communication network) or stored within the electronic encoding device 2 (for later transmission or as an alternative for later decoding, for example within the electronic encoding device itself, in which case the electronic encoding device is designed to further implement the decoding method described below with reference to Figure 8 ).
[0135] When the audio or video content comprises several parts (for example, groups of images when the content is a video sequence), the method from step E4 to step E24 can be implemented for each part of the content (for example, for each group of images) in such a way as to obtain, for each part of the content (for example, for each group of images), a data stream as illustrated in one of the figures Figures 3 to 6 . Thus, the compressed stream Fnn related to each group of images can be decoded using an artificial neural network specific to the group of images in question, and potentially different from the artificial neural networks used for the other groups of images, as described below. The artificial neural networks can have the same structure (and only differ in their weights and / or in the activation functions that define the specific artificial neural network).
[0136] Figure 7 An electronic decoding device 10 is shown that uses at least one artificial neural network 18.
[0137] This electronic decoding device 10 comprises a reception unit 11, a processor 14 (for example, a microprocessor) and a parallelization processing unit 16, for example a graphic processing unit or GPU, or a tensor processing unit or TPU.
[0138] The reception unit 11 is for example a communication circuit (like a radiofrequency communication circuit) and can receive data (and here in particular a coded data stream) from an external electronic device such as the electronic encoding device 2, and transmit these data to the processor 14 (the reception unit 11 is for example connected to the processor 14 by a bus).
[0139] The electronic decoding device 10 also comprises a storage unit 12, for example a memory (possibly rewritable non-volatile) or a hard disk drive. Although the storage unit 12 is shown in Figure 7 as an element distinct from the processor 14, the storage unit 12 can be integrated into (i.e. included in) the processor 14 as an alternative.
[0140] In this case, the processor 14 is adapted to successively execute several instructions of a computer program, for example stored in the storage unit 12.
[0141] The parallelization processing unit 16 is designed to implement the artificial neural network 18 after having been configured by the processor 14. To this end, the parallelization processing unit 16 is designed to perform, at a given time, several operations of the same type in parallel.
[0142] As schematically shown in Figure 7 , the processor 14 receives a data stream (for example, by means of a communication device not shown of the electronic decoding device 10) comprising a first set of data, here a header Fet, and a second set of data representative of the audio or video content, here a compressed stream Fnn.
[0143] As described below, the artificial neural network 18 is used in the framework of processing a second data set, that is to say, here, the compressed data Fnn, to obtain an audio or video content corresponding to the initial audio or video content B.
[0144] The storage unit 12 can store a plurality of parameter sets, each parameter set defining a decoding artificial neural network. As described below, in this case, the processor 14 can configure the parallelization processing unit 16 by means of a particular parameter set among these parameter sets, so that the parallelization processing unit 16 can then implement the artificial neural network defined by this particular parameter set.
[0145] The storage unit 12 can in particular store a first parameter set defining a first artificial neural network forming a random access decoder and / or a second parameter set defining a second artificial neural network forming a low latency decoder.
[0146] In this case, the electronic decoding device 10 has decoding options in advance for both the case where random access to the content is desired and the case where the content is to be displayed without delay.
[0147] Reference will now be made to Figure 8 A decoding method is described, which is implemented within the electronic decoding device 10 and uses the artificial neural network 18 implemented by the parallelization processing unit 16.
[0148] This method can begin with a selectable step of transmission by the electronic decoding device 10 of a list L of artificial neural networks available to the electronic decoding device 10 to a device for controlling the transmission of the data stream to be decoded. The data stream transmission control means can for example be the electronic encoding device 2. (In this case, the electronic encoding device 2 receives this list L as described above with reference to step E2 of the method described above) Figure 2 Alternatively, the data stream transmission control device can be a dedicated server, working in collaboration with the electronic encoding device 2.
[0149] The artificial neural networks accessible by the electronic decoding device 10 are artificial neural networks for which the electronic decoding device 10 stores a parameter set defining (as indicated above) the artificial neural network in question, or can access this parameter set by connecting to a remote electronic device such as a server (as explained below).
[0150] Figure 8 The method of Fig. 1 comprises a step E52 of reception (by the electronic decoding device 10, and here precisely by the reception unit 11) of a data stream comprising a first parameter set, that is to say, a header Fet, and a second parameter set, that is to say, a compressed stream Fnn. The reception unit 11 transmits the received data stream to the processor 14.
[0151] The processor 14 then proceeds to a step E54, identifying the first data set (header Fet) and the second data set (compressed flow Fnn) within the received data stream, for example by means of the start-of-compressed-flow indicator (already mentioned in the description of step E24).
[0152] The processor 14 can also identify, at step E54, the different parts of the first data set (header), that is, here, within the header Fet: a first part Fc (comprising data indicative of the features of the format of the content encoded by the data stream), a second part (indicator IND or IND'), a third part (identifier Inn or encoded data Rc) and possibly a fourth part (computer program Exe), as described above Figures 3 to 6
[0153] In the case where executable instructions (for example, instructions of a computer program Exe) are identified (i.e. detected) within the first data at step E54, the processor 14 can launch, at step E56, the execution of these executable instructions in order to implement at least certain steps of the processing of the data from the first data set (described below). These instructions can be executed by the processor 14 or, alternatively, by a virtual machine instantiated within the electronic decoding device 10.
[0154] Figure 7 The method continues with a step E58, in which the data Fc indicative of the format of the representation of the audio or video content are decoded in a manner that makes it possible to obtain the features of the format. For example, in the case of video content, the decoding of the data part Fc makes it possible to obtain the image size (in pixels) and / or the frame rate and / or the luminance information binary depth and / or the chroma slice information binary depth.
[0155] The processor 14 then proceeds (in certain embodiments, as a result of the execution of the instructions identified in the first data at step E54 (as already indicated)) to a step E60 of decoding the indicator IND, IND' contained in the second part of the header Fet.
[0156] If the decoding of the indicator IND, IND' present in the received data stream indicates that the artificial neural network 18 to be used for decoding belongs to a predetermined set of artificial neural networks (that is, if the indicator present in the first data is the indicator IND indicative of the decoding artificial neural network 18 belonging to the predetermined set of artificial neural networks), the method continues with a step E62 described below.
[0157] If the decoding of the indicator IND, IND' present in the received data stream indicates that the artificial neural network 18 to be used for the decoding is encoded in the data stream (that is, if the indicator present in the first data set is the indicator IND' indicating that the decoding artificial neural network 18 is encoded in the data stream), the method continues with the step E66 described hereinafter.
[0158] At step E62, the processor 14 proceeds with the decoding of the identifier Inn (here contained in the third portion of the header Fet) (in some embodiments, as a consequence of the execution of the instructions identified at step E54, as already indicated).
[0159] Then, the processor 14 can proceed at step E64 (in some embodiments, as a consequence of the execution of the instructions identified at step E54, as already indicated) for example with the reading in the storage unit 12 of the set of parameters associated with the decoding identifier Inn (this set of parameters defining the artificial neural network identified by the decoding identifier Inn).
[0160] According to a possible embodiment, the processor 14 can be provided to generate an error message in the absence (here within the storage unit 12) of the data (in particular the parameters) related to this artificial neural network identified by the decoding identifier Inn.
[0161] As an alternative (or in the absence of the set of parameters of the artificial neural network identified by the decoding identifier Inn stored in the storage unit 12), the electronic decoding device 10 can transmit (in some embodiments, as a consequence of the execution of the instructions identified at step E54, as already indicated) to a remote server a request for the set of parameters (this request including for example the decoded identifier Inn) and receive at step E64 as a reply the set of parameters defining the artificial neural network identified by the decoded identifier Inn.
[0162] Then, the method continues with the step E68 described hereinafter.
[0163] At step E66, the processor 14 proceeds with the decoding of the data Rc describing the artificial neural network 18 (here contained in the third portion of the header Fet) (in some embodiments, as a consequence of the execution of the instructions identified at step E54, as already indicated).
[0164] As already indicated, these descriptive data (or encoded data) Rc are encoded for example according to a standard such as MPEG-7 part 17 or a format such as JSON.
[0165] The decoding of the descriptive data Rc makes it possible to obtain, from the second data set (that is to say, here from the data of the compressed stream Fnn), the parameters that define the artificial neural network used to decode the data.
[0166] In this case, the method also continues with a step E68 that will now be described.
[0167] The processor 14 then configures, at step E68, the parallelization processing unit 16 by means of the parameters that define the decoding artificial neural network 18 (the parameters obtained at step E64 or at step E66), making it possible for the parallelization processing unit 16 to implement the decoding artificial neural network 18 (in certain embodiments, as a result of the execution of the instructions identified in the first data set of step E54, as already indicated).
[0168] This configuration step E68 specifically comprises the instantiation of the decoding artificial neural network 18 within the parallelization processing unit 16, using the parameters obtained at step E64 or at step E66.
[0169] This instantiation can specifically comprise the following steps:
[0170] - reserving the storage space necessary to implement the decoding artificial neural network 18 within the parallelization processing unit 16; and / or
[0171] - programming the parallelization processing unit 16 with the parameters (for example comprising the weights W' and the activation functions) that define the decoding artificial neural network 18 (the parameters obtained at step E64 or at step E66); and / or
[0172] - loading at least part of the data from the second data set (that is to say, at least part of the data from the compressed stream Fnn) into the local memory of the parallelization processing unit 16.
[0173] As can be seen from the description of steps E58 to E68 above, the data from the first data set Fet are thus processed by the processor 14.
[0174] The processor 14 can then apply, at step E70, the data from the second data set (here the data from the compressed stream Fnn) to the artificial neural network 18 implemented by the parallelization processing unit 16, making it possible for these data to be processed by at least partial use of the decoding process of the artificial neural network 18.
[0175] In the example described here, the artificial neural network 18 receives as input data from the second data set Fnn and produces as output a representation I of the encoded content, which is adapted to be reproduced on an audio or video reproduction device. In other words, at least certain data from the second data set Fnn is applied to the input layer of the artificial neural network 18 and the output layer of the artificial neural network 18 produces the above-mentioned representation I of the encoded content. In the case of video content, comprising an image or a sequence of images, the artificial neural network 18 thus produces as output (that is to say at the output layer of the artificial neural network 18) at least one matrix representation I of the image.
[0176] In certain embodiments, in order to process certain data from the compressed stream Fc (for example corresponding to a block or to an image), the artificial neural network 18 can receive as input at least certain data produced at the output of the artificial neural network 18 during the processing of previous data (for example corresponding to a preceding block or to a preceding image) in the compressed stream Fc. In this case, a step E72 is performed, in which the data produced at the output of the artificial neural network 18 are re-injected into the input of the artificial neural network 18.
[0177] Moreover, according to other possible embodiments, the decoding process can use a plurality of artificial neural networks, as already mentioned above with regard to the processing of the content data B.
[0178] Thus, the data from the second group (here at least certain data of the compressed stream Fnn) have been processed by relying on the processing of a part of the data from the first group (here on the processing of the identifier Inn of the encoded data Rc) and using the artificial neural network 18 implemented by the parallelized processing unit 16.
[0179] Then, the processor 14 determines, at a step E74, whether the processing of the compressed stream Fnn by means of the artificial neural network 18 is finished.
[0180] In the case of a negative determination (N), the method loops to a step E70 for applying other data from the compressed stream Fnn to the artificial neural network 18.
[0181] In the case of a positive determination (P), the method continues to a step E76 in which the processor 14 determines whether there is still data to be processed in the received data stream.
[0182] In the case of a negative determination (N) at the step E76, the method ends at a step E78.
[0183] In the case of a positive determination (P) at the step E76, the method loops to a step E52 for processing a new portion of the data stream, as illustrated in one of the figures Figures 3 to 6
[0184] As indicated above with respect to the repeated encoding steps E4 to E24, this other portion of the data stream then also comprises a first data set and a second data set representative of another audio or video content (e.g. in the case of video content, another set of images representative of the representation format of the used content). In this case, another artificial neural network can be determined based on certain of these first data (identifier Inn or encoded data Rc), as described above in steps E54 to E66, and then the parallelization processing unit 16 can be configured to implement this other artificial neural network (according to step E68 described above). Then, data from the second data set of this other portion of the data stream (e.g. related to the other set of images described above) can be decoded by means of this other artificial neural network (as described above in step E70).
[0185] The other artificial neural network just mentioned can have the same structure as the artificial neural network 18 described above, which simplifies the step of configuring the parallelization processing unit 16 (e.g. only updating the weights and / or activation functions defining the current artificial neural network).
Claims
1. A method for decoding a data stream, the data stream including an indicator (IND; IND') and representative data (Fnn) representing audio or video content, the method comprising the following steps: - Decode (E60) the indicator (IND; IND') to determine whether the artificial neural network (18) used to decode the representative data (Fnn) is encoded in the data stream or belongs to a predetermined set of artificial neural networks; - If the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) is encoded in the data stream, then the parameters of the artificial neural network (18) are decoded from the data stream; or if the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) belongs to a predetermined set of artificial neural networks, then the parameters of the artificial neural network (18) are read from the locally stored predetermined set or requested from a remote server. -The representative data (Fnn) is decoded (E70) by means of the artificial neural network (18).
2. The decoding method according to claim 1, wherein, The content is a first part of a video sequence, the video sequence including the first part and a second part, wherein the step of decoding the representative data produces the first part, and wherein the method further includes the step of decoding other data by means of another artificial neural network to produce the second part.
3. The decoding method according to claim 2, wherein, The other artificial neural network has the same structure as the artificial neural network.
4. The decoding method according to claim 2, wherein, The first part and the second part respectively form two sets of images in a format used to represent the content used.
5. The decoding method according to claim 1, comprising the step of transmitting the list (L) of the artificial neural network to a device for controlling the transmission of the data stream (E50).
6. The decoding method according to any one of claims 1 to 5, comprising: If it is determined by decoding the indicator (IND) that the artificial neural network (18) belongs to the predetermined set, then decode (E62) the identifier (Inn) of the neural network (18).
7. The decoding method according to claim 6, comprising reading the parameters of the artificial neural network (18) identified by the decoded identifier (Inn) in the storage unit (12).
8. The decoding method according to claim 7, wherein, The storage unit (12) stores a first parameter set representing a first artificial neural network forming a random access decoder and a second parameter set representing a second artificial neural network forming a low latency decoder.
9. The decoding method of claim 6, further comprising the step of generating an error message in the absence of data relating to the artificial neural network identified by the decoded identifier.
10. The decoding method according to claim 6, comprising receiving from a remote server parameters of the artificial neural network (18) identified by the decoded identifier (Inn).
11. The decoding method according to any one of claims 1 to 5, comprising: If it is determined by decoding the indicator (IND') that the artificial neural network (18) is encoded in the data stream, then the data (Rc) containing the encoded artificial neural network (18) in the data stream is decoded (E66) to obtain the parameters of the artificial neural network (18).
12. A decoding device, comprising: - A unit (11) for receiving a data stream, the data stream including an indicator (IND; IND') and representative data (Fnn) representing audio or video content. - Decoding components (14, 16), which are designed to determine whether the artificial neural network (18) used to decode the representative data (Fnn) belongs to a predetermined set of artificial neural networks or is encoded in the data stream by decoding the indicator (IND; IND'). If the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) belongs to the predetermined set of artificial neural networks, the parameters of the artificial neural network (18) are read from the predetermined set stored locally or requested from a remote server. Alternatively, if the indicator (IND; IND') indicates that the artificial neural network (18) used to decode the representative data (Fnn) is encoded in the data stream, the parameters of the artificial neural network (18) are decoded from the data stream, and the representative data is decoded by means of the artificial neural network (18).
Citation Information
Patent Citations
Method and apparatus of neural network based processing in video coding
CN107925762A
Image decoding method and device, image encoding method and device, and equipment
CN110401836A