Hybrid analog and digital neural network implementation in video coding process
By embedding metadata representing the reliability of the neural network inference process of simulated devices in the video data, the problem that the output results of the simulation device in the neural network inference process is not repetitive, and the reliability and accuracy of the video encoding and decoding process are achieved.
Patent Information
- Application Number
- CN202380068719.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-26
- Filing Date
- 2023-09-14
- Publication Date
- 2025-05-06
AI Technical Summary
When the simulation device executes the neural network inference process, it cannot ensure the repeatability of the output results, resulting in errors in the video encoding and decoding process.
By obtaining video data containing metadata, the metadata represents the reliability level of the neural network inference process implemented by the simulation device, and the video data is decoded using metadata to ensure the reliability of the decoding process.
It effectively solves the problem that the output results of simulation devices cannot be repeated during neural network inference, and ensures the reliability and accuracy of the video encoding and decoding process.
Smart Images

Figure CN119948876A_ABST
Abstract
Description
Technical Field
[0001] At least one of the present embodiments generally relates to methods and devices for encoding and decoding video data using a neural network inference process, and more particularly to a method for preventing errors that occur when the neural network inference process is performed by an analog device compared to when the neural network inference process is performed by a digital device. Background Art
[0002] In order to achieve high compression efficiency, video coding schemes usually use prediction and transformation to make full use of spatial and temporal redundancy in video content. During the encoding process, the picture of the video content is divided into sample blocks (i.e., pixels), and then these blocks are divided into one or more sub-blocks, hereinafter referred to as original sub-blocks. Intra-frame or inter-frame prediction is then applied to each sub-block to exploit the correlation within or between image frames. Regardless of which prediction method (intra-frame or inter-frame) is used, a predicted sub-block is determined for each original sub-block. Then, the sub-block representing the difference between the original sub-block and the predicted sub-block (usually represented as a prediction error sub-block, a prediction residual sub-block, or simply a residual sub-block) is transformed, quantized, and entropy encoded to generate an encoded video stream. In order to reconstruct the video, the compressed data is decoded by the inverse process corresponding to the transformation, quantization, and entropy encoding.
[0003] Among the recently explored video codec solutions, neural network based processing has been proposed (e.g. in the post-filtering stage or for block prediction). A major problem with neural network (NN) based codec tools is that they are computationally intensive and incur a lot of energy consumption in software or hardware implementations. Nevertheless, recent chip developments using analog designs (e.g. analog matrix processors) allow for significant reductions in the energy consumption of NN based codec tools. However, since these chips use analog designs, they may produce non-bit accurate output results. This paper defines output as a set of numerical samples, e.g. a 2D image consisting of luma samples, a 2D image consisting of chroma samples, a 2D image consisting of depth samples, a 1D list of motion vector candidates, a 1D list of intra mode candidates, a 2D motion vector field where each motion vector consists of two samples. As a reminder, the process of applying a trained NN to input data to obtain output data is called the inference process. This characteristic of analog design is problematic in video codecs, as the output results of codec tools need to be accurate and reproducible. In fact, in video codecs, it is not allowed that two applications of the same NN inference process on the same input data provide different output data (which is impossible in the case of purely digital computing).
[0004] It is desirable to propose a solution that can overcome the above problems. In particular, it is desirable to propose a solution that allows ensuring or at least checking the reliability of NN inference processes performed on simulated devices. Summary of the invention
[0005] In a first aspect, one or more of the present embodiments provide a method comprising: obtaining video data including metadata, the metadata representing a reliability level of a neural network inference process implemented by an analog device, the neural network inference process being applied to decode the video data, the analog device being an electronic circuit that cannot ensure the repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; and decoding the video data, wherein the analog device implements the neural network inference process relying on the metadata.
[0006] In an embodiment, the metadata includes a flag indicating that there is a first checksum representing an expected output of the neural network reasoning process or there is a first statistical parameter representing an expected output of the neural network reasoning process in the metadata.
[0007] In an embodiment, the metadata includes a first checksum or a first statistical parameter representing an expected output of the neural network inference process.
[0008] In an embodiment, the simulation device implementing the neural network reasoning process includes: using the simulation device to execute the neural network reasoning process at least once, and determining that a second checksum or a second statistical parameter calculated based on an output of one of the at least one execution of the neural network reasoning process by the simulation device is equal to a first checksum or a first statistical parameter respectively included in the metadata, and
[0009] Decoding the video data includes: in response to the second checksum being equal to the first checksum or in response to the second statistical parameter being equal to the first statistical parameter, decoding the video data using an output obtained by performing a neural network inference process by the simulation device.
[0010] In an embodiment, decoding the video data includes: in response to each second checksum or each second statistical parameter calculated based on the output obtained by the analog device each time performing the neural network inference process is different from the first checksum or the first statistical parameter, respectively, using the output obtained by the digital device performing the neural network inference process to decode the video data, the digital device being an electronic circuit that ensures the repeatability of the output result of the electronic circuit when the electronic circuit receives the same input data.
[0011] In an embodiment, the neural network inference process is performed multiple times, and in response to each second checksum or second statistical parameter calculated based on the output obtained by the simulation device each time the neural network inference process is performed is different from the first checksum or first statistical parameter, respectively, decoding the video data includes: decoding the video data using an output corresponding to one of the following:
[0012] The median value of the output obtained by multiple executions of the neural network inference process by the simulated device;
[0013] The arithmetic or geometric mean of the outputs of multiple executions of the neural network inference process by the simulated device; and
[0014] An output selected from among outputs obtained by a simulation device performing a neural network inference process multiple times, wherein the selection makes a second statistical parameter calculated from the selected output closest to the first statistical parameter.
[0015] In an embodiment, the video data represents scalable video including a base layer and at least one enhancement layer, and the neural network inference process is included in a portion of a decoding process related to decoding of the enhancement layer, and in response to each second checksum or second statistical parameter calculated based on the output obtained by the simulation device each time the neural network inference process is performed is different from the first checksum or first statistical parameter, respectively, decoding the video data includes: decoding the video data by replacing each output obtained by the simulation device performing the neural network inference process with an output generated from the data calculated for the base layer.
[0016] In an embodiment, in response to the flag indicating that a checksum is not present in the metadata or that a statistical parameter is not present in the metadata, a neural network inference process is performed by the simulation device.
[0017] In an embodiment, the metadata includes a flag indicating which device is used to perform the neural network inference process among an analog device and a digital device, wherein the digital device is an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data.
[0018] In a second aspect, one or more of the embodiments provide a method, comprising:
[0019] encoding the video data using an encoding process including a neural network inference process performed by a digital device, the digital device being an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data;
[0020] Calculating metadata representing a reliability level of a neural network inference process implemented by an analog device using an output of the neural network inference process performed by a digital device, the analog device being an electronic circuit that cannot ensure repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; and
[0021] The metadata is signaled in the coded video data.
[0022] In an embodiment, the metadata is signaled in one of the following:
[0023] Picture header;
[0024] Slice header;
[0025] Adaptive parameter set;
[0026] Before the first coding tree unit of the region to which the metadata applies;
[0027] After the last coding tree unit of the region to which the metadata applies;
[0028] in a SEI message attached to the coded video data; and
[0029] In a manifest file that conforms to the streaming protocol.
[0030] In an embodiment, the metadata is calculated per picture group or per picture or per slice or per tile or per sub-picture.
[0031] In an embodiment, the metadata includes checksums or statistical parameters representing expected outputs of the neural network inference process.
[0032] In an embodiment, the method further comprises:
[0033] Analyzing the consistency of the neural network reasoning process performed by the analog device by performing the neural network reasoning process multiple times using the analog device and comparing the outputs obtained by the analog device performing the neural network reasoning process multiple times with the outputs obtained by performing the neural network reasoning process by the digital device; and
[0034] In response to the number of outputs obtained by multiple executions of the neural network inference process by the analog device being equal to the output obtained by the digital device executing the neural network inference process being less than a value, inserting a flag in the metadata, the flag indicating that the metadata includes a checksum or statistical parameters, otherwise the flag indicates that the checksum or statistical parameters do not exist in the metadata.
[0035] In an embodiment, if all outputs obtained by multiple executions of the neural network inference process by the analog device are equal to the outputs obtained by the digital device executing the neural network inference process, then the flag indicates that there is no checksum or statistical parameter in the metadata.
[0036] In a third aspect, one or more of the embodiments provide a device comprising a processor configured to:
[0037] Acquiring video data including metadata indicating a reliability level of a neural network inference process implemented by an analog device, the neural network inference process being applied to decode the video data, the analog device being an electronic circuit that is unable to ensure repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; and
[0038] Decoding video data, where the analog device implements the neural network inference process relying on metadata.
[0039] In an embodiment, the metadata includes a flag indicating the presence of a first checksum in the metadata representing an expected output of the neural network reasoning process or the presence of a first statistical parameter representing an expected output of the neural network reasoning process.
[0040] In an embodiment, the metadata includes a first checksum or a first statistical parameter representing an expected output of the neural network inference process.
[0041] In an embodiment, the simulation device implementing the neural network reasoning process includes: using the simulation device to execute the neural network reasoning process at least once, and determining that a second checksum or a second statistical parameter calculated based on an output of one of the at least one execution of the neural network reasoning process by the simulation device is equal to a first checksum or a first statistical parameter respectively included in the metadata, and
[0042] Decoding the video data includes: in response to the second checksum being equal to the first checksum or in response to the second statistical parameter being equal to the first statistical parameter, decoding the video data using an output obtained by performing a neural network inference process by the simulation device.
[0043] In an embodiment, decoding the video data includes: in response to each second checksum or each second statistical parameter calculated based on the output obtained by the analog device each time performing the neural network inference process is different from the first checksum or the first statistical parameter, respectively, using the output obtained by the digital device performing the neural network inference process to decode the video data, the digital device being an electronic circuit that ensures the repeatability of the output result of the electronic circuit when the electronic circuit receives the same input data.
[0044] In an embodiment, the processor is configured to perform the neural network inference process multiple times, and in response to each second checksum or second statistical parameter calculated based on the output obtained by the simulation device each time performing the neural network inference process being different from the first checksum or first statistical parameter, respectively, decoding the video data includes: decoding the video data using an output corresponding to one of the following:
[0045] The median value of the output obtained by multiple executions of the neural network inference process by the simulated device;
[0046] The arithmetic or geometric mean of the outputs of multiple executions of the neural network inference process by the simulated device;
[0047] An output selected from among outputs obtained by a simulation device performing a neural network inference process multiple times, wherein the selection makes a second statistical parameter calculated from the selected output closest to the first statistical parameter.
[0048] In an embodiment, the video data represents scalable video including a base layer and at least one enhancement layer, and the neural network inference process is included in a portion of a decoding process related to decoding of the enhancement layer, and in response to each second checksum or second statistical parameter calculated based on the output obtained by the simulation device each time the neural network inference process is performed is different from the first checksum or first statistical parameter, respectively, decoding the video data includes: decoding the video data by replacing each output obtained by the simulation device performing the neural network inference process with an output generated from the data calculated for the base layer.
[0049] In an embodiment, in response to the flag indicating that a checksum is not present in the metadata or that a statistical parameter is not present in the metadata, the processor is configured to perform a neural network inference process using the simulation device.
[0050] In an embodiment, the metadata includes a flag indicating which device is used to perform the neural network inference process among an analog device and a digital device, where the digital device is an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data.
[0051] In a fourth aspect, one or more of the embodiments provide a device comprising a processor configured to:
[0052] encoding the video data using an encoding process including a neural network inference process performed by a digital device, the digital device being an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data;
[0053] Calculating metadata representing a reliability level of a neural network inference process implemented by an analog device using an output of the neural network inference process performed by a digital device, the analog device being an electronic circuit that cannot ensure repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; and
[0054] The metadata is signaled in the coded video data.
[0055] In an embodiment, the metadata is signaled in one of the following:
[0056] Picture header;
[0057] Slice header;
[0058] Adaptive parameter set;
[0059] Before the first coding tree unit of the region to which the metadata applies;
[0060] After the last coding tree unit of the region to which the metadata applies;
[0061] in a SEI message attached to the coded video data; and
[0062] In a manifest file that conforms to the streaming protocol.
[0063] In an embodiment, the metadata is calculated per picture group or per picture or per slice or per tile or per sub-picture.
[0064] In an embodiment, the metadata includes checksums or statistical parameters representing expected outputs of the neural network inference process.
[0065] In an embodiment, the device is further configured to:
[0066] Analyzing the consistency of the neural network reasoning process performed by the analog device by performing the neural network reasoning process multiple times using the analog device and comparing the outputs obtained by the analog device performing the neural network reasoning process multiple times with the outputs obtained by performing the neural network reasoning process by the digital device; and
[0067] In response to the number of outputs obtained by multiple executions of the neural network inference process by the analog device being equal to the output obtained by the digital device executing the neural network inference process being less than a value, inserting a flag in the metadata, the flag indicating that the metadata includes a checksum or statistical parameters, otherwise the flag indicates that the checksum or statistical parameters do not exist in the metadata.
[0068] In an embodiment, if all outputs obtained by multiple executions of the neural network inference process by the analog device are equal to the outputs obtained by the digital device executing the neural network inference process, then the flag indicates that there is no checksum or statistical parameter in the metadata.
[0069] In a fifth aspect, one or more of the present embodiments provides a signal comprising metadata representing a level of reliability of a neural network inference process of a video decoding process implemented by a simulation device.
[0070] In a sixth aspect, one or more of the present embodiments provides a computer program comprising program code instructions for implementing the method according to the first or second aspect.
[0071] In a seventh aspect, one or more of the present embodiments provides a non-transitory information storage medium storing program code instructions for implementing the methods according to the first and second aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 The implementation environment of the embodiment is schematically shown;
[0073] Figure 2 An example of the divisions that a pixel picture of an original video undergoes is schematically shown;
[0074] Figure 3 A method for encoding a video stream is schematically depicted;
[0075] Figure 4 A method for decoding an encoded video stream is schematically depicted;
[0076] Figure 5A An example of a hardware architecture capable of implementing a processing module of an encoding module or a decoding module for implementing various aspects and embodiments is schematically shown;
[0077] Figure 5B A block diagram illustrating an example of a first system for implementing various aspects and embodiments is shown;
[0078] Figure 5C A block diagram showing an example of a second system for implementing various aspects and embodiments;
[0079] Fig. 6A Schematically illustrating an embodiment of a portion of a reliable NN inference process implemented by a simulation device, performed on the encoder side;
[0080] Figure 6B schematically illustrates an embodiment of a portion of a reliable NN inference process implemented by a simulation device, performed on the decoder side; and
[0081] Figure 7 An example of application of a reliable NN inference process implemented by a simulation device is schematically shown. DETAILED DESCRIPTION
[0082] The following example embodiments are described in the context of a video format similar to VVC (ITU-T H.266 Recommendation | International Standard ISO / IEC 23090-3: Generic Video Coding) developed by a joint collaborative team of ITU-T and ISO / IEC experts, known as the Joint Video Experts Team (JVET). However, these embodiments are not limited to video encoding / decoding methods corresponding to VVC. These embodiments are particularly applicable to a variety of video formats, such as HEVC (ISO / IEC 23008-2–MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265), AVC ((ISO / CEI 14496-10), EVC (Basic Video Coding / MPEG-5), SVC (Scalable Video Coding), SHVC (Scalable High Efficiency Video Coding), AV1, AV2, and VP9.
[0083] Figure 1 The implementation environment of the embodiment is schematically shown.
[0084] exist Figure 1 In the embodiment, system 11 (which may be a camera, storage device, computer, server, or any device capable of forwarding a video stream) uses communication channel 12 to send a video stream to system 13. The video stream is encoded and sent by system 11, or received and / or stored by system 11 and then sent. Communication channel 12 is a wired (e.g., Internet or Ethernet) or wireless (e.g., WiFi, 3G, 4G, or 5G) network link.
[0085] The system 13, which may be a set-top box, for example, receives the video stream and decodes it to generate a sequence of decoded pictures.
[0086] The obtained decoded sequence of pictures is then sent using a communication channel 14 (which may be a wired or wireless network) to a display system 15. The display system 15 then displays the pictures.
[0087] In an embodiment, system 13 is included in display system 15. In this case, system 13 and display system 15 are included in a television, computer, tablet computer, smartphone, head mounted display, etc.
[0088] Figure 2 , Figure 3 and Figure 4 Examples of video formats are presented.
[0089] Figure 2 An example of the partitioning that a pixel picture 21 of an original video sequence 20 undergoes is shown. Here, a pixel is considered to consist of three components: a luminance component and two chrominance components. However, other types of pixels may also include fewer or more components, such as only a luminance component or an additional depth component or a transparency component.
[0090] The image is divided into multiple encoding entities. First, Figure 2 As shown in the figure 23 in the figure, the picture is divided into a grid of blocks called coding tree units (CTUs). A CTU consists of an N×N block of luminance samples and two corresponding blocks of chrominance samples. N is usually a power of 2, and the maximum value is, for example, "128". Secondly, the picture is divided into one or more groups of CTUs. For example, it can be divided into one or more tile rows and tiles columns, and a tile is a sequence of CTUs covering a rectangular area of the picture. In some cases, a tile can be divided into one or more blocks, each of which consists of at least one row of CTUs within the tile. On top of the concepts of slices and blocks, there is another coding entity called a slice, which can contain at least one slice of a picture or at least one block of a slice.
[0091] exist Figure 2 In the example, as shown by reference numeral 22, the picture 21 is divided into three slices S1, S2 and S3 in a raster scan slice mode, each slice includes a plurality of slices (not shown), and each slice includes only one block.
[0092] like Figure 2 As shown by reference numeral 24 in FIG. 1 , a CTU can be divided into one or more sub-blocks (referred to as coding units (CUs)) in a hierarchical tree. A CTU is the root (i.e., the parent node) of the hierarchical tree and can be divided into multiple CUs (i.e., child nodes). If each CU is not further divided into smaller CUs, it will become a leaf of the hierarchical tree; if it is further divided, it will become a parent node of a smaller CU (i.e., child node).
[0093] exist Figure 2 In the example of FIG, first, a quadtree type partition is used to partition CTU 24 into "4" square CUs. The upper left CU is a leaf of the hierarchical tree because it is not further partitioned, i.e., it is not the parent node of any other CU. The upper right CU is further partitioned into "4" smaller square CUs using a quadtree type partition again. The lower right CU uses a binary tree type partition and is vertically partitioned into "2" rectangular CUs. The lower left CU uses a ternary tree type partition and is vertically partitioned into "3" rectangular CUs.
[0094] During picture encoding, the partitioning is adaptive, and each CTU is partitioned to optimize the compression efficiency of the CTU standard.
[0095] The concepts of prediction unit (PU) and transform unit (TU) appear in HEVC. In fact, in HEVC, the coding entity used for prediction (i.e. PU) and transform (i.e. TU) can be a subdivision of CU. For example, Figure 2As shown, a CU of size 2N×2N may be partitioned into PUs 2411 of size N×2N or PUs 2412 of size 2N×N. In addition, the CU may be partitioned into "4" TUs 2412 of size N×N or TUs of size 2N×N. The “16” TUs.
[0096] It can be noted that in VVC, except for some special cases, the boundaries of TU and PU are aligned with the boundaries of CU. Therefore, a CU usually includes one TU and one PU.
[0097] In this application, the term "block" or "picture block" may be used to refer to any of CTU, CU, PU and TU. In addition, the term "block" or "picture block" may be used to refer to macroblocks, partitions and sub-blocks specified in H.264 / AVC or other video coding standards, and more generally to sample arrays with various sizes.
[0098] In this application, the terms "reconstruction" and "decoding" are used interchangeably, the terms "pixel" and "sample" are used interchangeably, and the terms "image", "picture", "sub-picture", "slice" and "frame" are used interchangeably. Usually, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.
[0099] Figure 3 Schematically depicts a method for encoding a video stream performed by an encoding module. Variations of this method for encoding are contemplated, but for clarity, the following describes Figure 3 of methods used for encoding without describing all intended variants.
[0100] Before encoding, the current original picture of the original video sequence may be preprocessed. For example, in step 301, a color transformation is applied to the current original picture (e.g., from RGB 4:4:4 to YCbCr4:2:0), or a remapping is applied to the current original picture components to obtain a signal distribution that is more adaptable to compression (e.g., using a histogram equalization for one of the color components). The picture obtained by preprocessing is hereinafter referred to as a preprocessed picture.
[0101] The encoding of the pre-processed picture begins with the partitioning of the pre-processed picture during step 302, as described with respect to Figure 2 Therefore, the pre-processed picture is divided into CTU, CU, PU, TU, etc. For each block, the encoding module determines the encoding mode between intra prediction and inter prediction.
[0102] Intra prediction consists in predicting the pixels of the current block from a prediction block originating from pixels of a reconstructed block located at the causal vicinity of the current block to be encoded according to an intra prediction method during a step 303. The result of the intra prediction is a prediction direction indicating which pixels of the neighboring blocks to use and a residual block obtained by calculating the difference between the current block and the predicted block.
[0103] Inter prediction consists in predicting the pixels of the current block from a block of pixels of a picture preceding or following the current picture, called a reference block, which picture is called a reference picture. During the encoding of the current block according to the inter prediction method, the block of the reference picture closest to the current block is determined by a motion estimation step 304 according to a similarity criterion. In step 304, a motion vector indicating the position of the reference block in the reference picture is determined. The motion vector is used during a motion compensation step 305, in which a residual block is calculated in the form of the difference between the current block and the reference block. In the first generation of video compression standards, the above-mentioned unidirectional inter prediction mode was the only inter mode available. With the development of video compression standards, the family of inter modes has grown significantly and now includes many different inter modes.
[0104] During the selection step 306 , the encoding module selects a prediction mode that optimizes compression performance among the tested prediction modes (intra-frame prediction mode, inter-frame prediction mode) according to a rate / distortion optimization criterion (ie, RDO criterion).
[0105] When the prediction mode is selected, the residual block is transformed during a step 307. Then, the transformed block is quantized during a step 309.
[0106] Note that the encoding module can skip the transformation and directly apply quantization to the untransformed residual signal. When the current block is encoded according to an intra-frame prediction mode, during step 310, the entropy encoder encodes the prediction direction and the transformed and quantized residual block. When the current block is encoded according to an inter-frame prediction, where appropriate, the motion vector of the block is predicted based on a prediction vector selected from a set of motion vector predictors, which are derived from reconstructed blocks located in the spatial and temporal neighborhood of the block to be encoded. Next, during step 310, the entropy encoder encodes the motion information in the form of a motion residual and an index for identifying the prediction vector. During step 310, the entropy encoder encodes the transformed and quantized residual block.
[0107] It should be noted that the encoding module can bypass the transformation and quantization, that is, directly apply entropy coding to the residual without applying the transformation or quantization process. The result of the entropy coding is inserted into the encoded video stream 311.
[0108] Metadata such as a SEI (Supplemental Enhancement Information) message may be appended to the coded video stream 311. A SEI message, for example defined in a standard such as AVC, HEVC or VVC, is a data container associated with a video stream and includes metadata providing information related to the video stream.
[0109] After the quantization step 309, the current block is reconstructed so that the pixels corresponding to the block can be used for future predictions. This reconstruction phase is also called a prediction loop. Therefore, inverse quantization is applied to the transformed and quantized residual block during step 312, and an inverse transform is applied during step 313. According to the prediction mode for the block obtained during step 314, the prediction block of the block is reconstructed. If the current block is encoded according to the inter-frame prediction mode, during step 316, the encoding module uses the motion vector of the current block to apply motion compensation where appropriate in order to identify the reference block of the current block. If the current block is encoded according to the intra-frame prediction mode, during step 315, the prediction block of the current block is reconstructed using the prediction direction corresponding to the current block. The prediction block and the reconstructed residual block are added to obtain the reconstructed current block.
[0110] After reconstruction, during step 317, in-loop filtering aimed at reducing coding artifacts is applied to the reconstructed blocks. This filtering is called in-loop filtering because it occurs in the prediction loop to obtain at the decoder the same reference image as at the encoder, thus avoiding drift between the encoding and decoding processes. In-loop filtering tools include deblocking filtering, SAO (Sample Adaptive Offset) and ALF (Adaptive In-Loop Filtering).
[0111] When a block is reconstructed, it is inserted into a reconstructed picture stored in a reconstructed picture memory 319, usually called a decoded picture buffer (DPB), during a step 318. The reconstructed picture stored in this way can then be used as a reference picture for other pictures to be encoded.
[0112] Figure 4 Schematically depicts the method for Figure 3 Related methods A method for decoding a coded video stream 311 encoded by a method, the method being performed by a decoding module. Variations of this method for decoding are contemplated, but for the sake of clarity, the following will be described. Figure 4 of methods for decoding without describing all expected variants.
[0113] The decoding is performed block by block. For the current block, it starts with entropy decoding of the current block during a step 410. Entropy decoding allows obtaining at least the prediction mode of the block.
[0114] If the block has been coded according to inter prediction mode, entropy decoding allows obtaining, where appropriate, the prediction vector index, the motion residual and the residual block. During step 408, the prediction vector index and the motion residual are used to reconstruct the motion vector for the current block. If the block has been coded according to intra prediction mode, entropy decoding allows obtaining the prediction direction and the residual block. Steps 412, 413, 414, 415, 416 and 417 implemented by the decoding module are identical in all respects to steps 312, 313, 314, 315, 316 and 317 implemented by the encoding module, respectively.
[0115] The decoded blocks are saved in a decoded picture, and the decoded picture is stored in DPB 419 in step 418. When the decoding module decodes a given picture, the picture stored in DPB 419 is the same as the picture stored in DPB 319 by the encoding module during encoding of the given picture. The decoded picture can also be output by the decoding module, for example, for display.
[0116] The post-processing step 421 may include an inverse color transform (e.g., from YcbCr 4:2:0 to RGB 4:4:4), an inverse mapping that performs the inverse of the remapping process performed in the pre-processing of step 301, and post-filtering to improve the reconstructed picture (e.g., based on filtering parameters provided in the SEI message).
[0117] In the recently explored video coding solutions, NN-based tools have been proposed. These NN-based tools or processes can be used at different levels.
[0118] At level "1", NN-based tools (i.e., NN reasoning processes) can be used to assist the encoder in making decisions. For example, they can be used to select the coding mode for a given CU, select the CTU or CU partitioning method, decide whether to apply a codec tool (when this tool can be activated / deactivated), select GOP size or structure, select picture type, slice or local QP adaptation. Figure 3 , which may for example include adding a NN-based decision module before steps 302 (partitioning), 307 (e.g. for selecting a transform method), 303 (e.g. for intra-coding mode selection), 306 (selecting between intra- and inter-coding, applying or not applying post-prediction filtering (e.g. bilateral filtering, bidirectional optical flow filtering), 304, 305 and 308 (inter-coding mode selection, decision between unidirectional prediction or bidirectional prediction, motion vector prediction, etc.).
[0119] In level "2", NN-based tools (i.e., NN reasoning processes) can be inserted into the traditional video decoding framework. For example, they can be used in the intra-frame prediction step (415), the motion compensation step (416), the in-loop filtering step (417), and the post-processing step (421). They can also be used in steps involving classification (e.g., selecting coding modes at block level, SAO classification, ALF classification, deblocking filter strength decision) and binary decisions (e.g., activating / deactivating codec tools at picture, slice or block level).
[0120] In level "3", the complete end-to-end video codec solution is based on NN design. This can be applied to the entire process based on NN design, or the main steps of the encoding process (e.g., intra-frame coding using NN autoencoder, motion prediction using another NN autoencoder, residual coding using another NN autoencoder).
[0121] As mentioned above, the main problem of the NN inference process is its complexity and the energy consumption it causes. The recent emergence of new analog devices (such as analog matrix processors) can provide solutions to these complexity and energy consumption problems. Analog devices basically replace the digital multiplication and / or addition steps of digital-based implementations with analog processes (not involving digital values, but continuous values composed of voltages). In the various embodiments described in this document, we define an analog device as an electronic circuit that does not ensure the repeatability of the output result of the electronic circuit when the electronic circuit receives the same input data, while a digital device is an electronic circuit that ensures the repeatability of the output result of the electronic circuit when the electronic circuit receives the same input data. In other words, if executed on an analog device, applying the same process twice to the same input data does not necessarily produce the same output data, while if executed on a digital device, applying the same process twice to the same input data will produce the same output data. In fact, the probability of an error (its output does not match the reference output derived from a mathematical fixed-point calculation) in the case of a digital device is extremely low and can be completely ignored, while this probability is very high for an analog device. One advantage of analog devices is that they can perform NN inference processes with lower power consumption than GPUs (e.g., https: / / www.mythic-ai.com / wp-content / uploads / 2021 / 06 / M1076-AMP-Product-Brief-v1.0-1.pdf claims a 10x reduction in power consumption). In addition, it is well known that analog devices often provide erroneous results that are close to the correct results.
[0122] However, video codecs cannot support approximations in processing results. Therefore, a solution is needed that can address the main problems of simulation-based implementations (e.g., simulation-based implementations cannot ensure that the same process applied to the same input data will produce the same output results).
[0123] Figure 5A , Figure 5B and Figure 5C Examples of devices, apparatuses and / or systems are described that enable implementation of various embodiments.
[0124] Figure 5A Schematically illustrates an example of a hardware architecture of a processing module 500 capable of implementing a coding module or a decoding module modified according to different aspects and embodiments, wherein the coding module or the decoding module can respectively implement Figure 3 The method used to encode and Figure 4 For example, when system 11 is responsible for encoding the video stream, the encoding module is included in the system. The decoding module is included in system 13, for example.
[0125] The processing module 500 includes: a processor or CPU (central processing unit) 5000, which includes one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architecture as non-limiting examples, connected by a communication bus 5005; a random access memory (RAM) 5001; a read-only memory (ROM) 5002; a storage unit 5003, which may include non-volatile memory and / or volatile memory (including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drive and / or optical disk drive) or a storage medium reader (such as SD (Secure Digital) card reader and / or hard disk drive (HDD)) and / or a network accessible storage device; at least one communication interface 5004, which is used to exchange data with other modules, devices or systems. The communication interface 5004 may include but is not limited to a transceiver configured to send and receive data through a communication channel. The communication interface 5004 may include, but is not limited to, a modem or a network card.
[0126] If the processing module 500 implements a decoding module, the communication interface 5004 enables the processing module 500 to receive a coded video stream and provide a decoded picture sequence, for example. If the processing module 500 implements an encoding module, the communication interface 5004 enables the processing module 500 to receive a raw picture data sequence for encoding and provide a coded video stream, for example.
[0127] The processor 5000 can execute instructions loaded into the RAM 5001 from the ROM 5002, from an external memory (not shown), from a storage medium, or from a communication network. When the processing module 500 is powered on, the processor 5000 can read instructions from the RAM 5001 and execute them. These instructions form a computer program, which, for example, enables the processor 5000 to implement the instructions related to Figure 4 Decoding methods described and / or Figure 3 The encoding method described and the Fig. 6A , Figure 6B or Figure 7 The described methods include various aspects and embodiments described herein below.
[0128] Figure 3 , Figure 4 , Fig. 6A , Figure 6B and Figure 7 Some algorithms and steps in the method may be implemented in the form of software by executing instruction sets by a programmable machine (such as a DSP (digital signal processor) or a microcontroller), or in the form of hardware by a machine or a dedicated component (such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit)). Figure 3 , Figure 4 , Fig. 6A , Figure 6B and Figure 7 Other algorithms and steps in the method can be implemented by analog devices (such as analog matrix processors). In particular, the NN reasoning process can be implemented by analog devices.
[0129] It can be seen that microprocessors, general-purpose computers, special-purpose computers, processors based on or not based on multi-core architectures, DSPs, microcontrollers, FPGAs, ASICs, analog devices (such as analog matrix processors) are all suitable for at least partially implementing Figure 3 , Figure 4 , Fig. 6A , Figure 6B and Figure 7 Method of electronic circuit.
[0130] Figure 5CA block diagram of an example of a system 13 for implementing various aspects and embodiments is shown. System 13 may be embodied as a device including various components described below and configured to perform one or more aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptops, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, networked home appliances, and head-mounted displays. The elements of system 13 may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components, either individually or in combination. For example, in at least one embodiment, system 13 includes a processing module 500 that implements a decoding module. In various embodiments, system 13 is coupled in a communication manner with one or more other systems or other electronic devices, such as a communication bus or through dedicated input and / or output ports. In various embodiments, system 13 is configured to implement one or more aspects described in this document.
[0131] Inputs to processing module 500 may be provided through a variety of input modules shown in block 531. Such input modules include, but are not limited to: (i) a radio frequency (RF) module that receives RF signals transmitted over the air by, for example, a broadcaster; (ii) a component (COMP) input module (or a set of COMP input modules); (iii) a universal serial bus (USB) input module; and / or (iv) a high-definition multimedia interface (HDMI) input module. Figure 5C Other examples not shown include composite video.
[0132] In various embodiments, the input module of block 531 has associated corresponding input processing elements known in the art. For example, the RF module may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandwidth limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again bandwidth limiting to a narrower frequency band to select (for example) a signal frequency band (which may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and bandwidth-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired packet stream. The RF module of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a bandwidth limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs a variety of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or baseband. In a set-top box embodiment, RF module and its related input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to desired frequency band again by filtering, down-conversion and perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.Adding element can include inserting element between existing element, for example inserting amplifier and analog-to-digital converter.In various embodiments, RF module comprises antenna.
[0133] In addition, the USB and / or HDMI modules may include respective interface processors for connecting the system 13 to other electronic devices via the USB and / or HDMI connections. It should be appreciated that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processing module 500, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processing module 500, as desired. The demodulated, error corrected, and demultiplexed stream is provided to the processing module 500.
[0134] The various elements of the system 13 may be arranged in an integrated housing. Within the integrated housing, the various elements may be interconnected and transmit data between each other using suitable connection means (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards). For example, in the system 13, the processing module 500 is interconnected to the other elements of the system 13 via the bus 5005.
[0135] The communication interface 5004 of the processing module 500 allows the system 13 to communicate over the communication channel 12. As described above, the communication channel 12 may be implemented in a wired and / or wireless medium, for example.
[0136] In various embodiments, a wireless network such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to the system 13. The Wi-Fi signal of these embodiments is received through a communication channel 12 and a communication interface 5004 suitable for Wi-Fi communication. The communication channel 12 of these embodiments is typically connected to an access point or router that provides access to external networks (including the Internet) to allow streaming applications and other over-the-top communications. Other embodiments use the RF connection of the input box 531 to provide streaming data to the system 13. As described above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0137] The system 13 can provide output signals to a variety of output devices, including a display system 15, a speaker 535, and other peripherals 536. The display system 15 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display system 15 can be used for a television, a tablet computer, a laptop computer, a mobile phone (mobile phone), a head-mounted display, or other devices. The display system 15 can also be integrated with other components (for example, as in a smartphone) or be independent (for example, an external display for a laptop computer). In examples of various embodiments, other peripherals 536 include one or more of an independent digital video disc (or digital versatile disc) (DVR, used for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripherals 536 that provide functions based on the output of the system 13. For example, a disk player performs the function of playing the output of the system 13.
[0138] In various embodiments, control signals are communicated between system 13 and display system 15, speaker 535, or other peripheral device 536 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that allow device-to-device control with or without user intervention. Output devices may be communicatively coupled to system 13 via dedicated connections via respective interfaces 532, 533, and 534. Alternatively, output devices may be communicatively coupled to system 13 via communication interface 5004 using communication channel 12 or via communication interface 5004 using communication with Figure 5CThe display interface 532 is connected to the system 13 through a dedicated communication channel corresponding to the communication channel 12 in the display interface 532. The display system 15 and the speaker 535 can be integrated with other components of the system 13 in a single unit in an electronic device (e.g., a television). In various embodiments, the display interface 532 includes a display driver, such as a timing controller (T Con) chip.
[0139] The display system 15 and the speaker 535 may alternatively be separate from one or more other components. In various embodiments where the display system 15 and the speaker 535 are external components, the output signal may be provided through a dedicated output connection (eg, an HDMI port, a USB port, or a COMP output).
[0140] Figure 5B A block diagram of an example of a system 11 for implementing various aspects and embodiments is shown. System 11 is very similar to system 13. System 11 can be embodied as a device including various components described below and configured to perform one or more aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptops, smart phones, tablet computers, cameras, and servers. The elements of system 11 can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, system 11 includes a processing module 500 that implements a coding module. In various embodiments, system 11 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, system 11 is configured to implement one or more aspects described in this document.
[0141] The input of the processing module 500 can be obtained through various input modules (such as those already mentioned). Figure 5C 531 as described above).
[0142] The various elements of the system 11 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected and transmit data between each other using suitable connection means (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards). For example, in the system 11, the processing module 500 is interconnected to the other elements of the system 11 via the bus 5005.
[0143] The communication interface 5004 of the processing module 500 allows the system 11 to communicate over the communication channel 12 .
[0144] In various embodiments, a wireless network such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to the system 11. The Wi-Fi signal of these embodiments is received through the communication channel 12 and the communication interface 5004 suitable for Wi-Fi communication. The communication channel 12 of these embodiments is typically connected to an access point or router that provides access to external networks (including the Internet) to allow streaming applications and other over-the-top business communications. Other embodiments use the RF connection of the input box 531 to provide streaming data to the system 11.
[0145] As described above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0146] The data provided to the system 11 may be provided in different formats. In various embodiments, these data are encoded and conform to known video compression formats, such as AV1, VP9, VVC, HEVC, AVC, EVC, SVC, SHVC, etc. In various embodiments, these data are raw data, such as provided by a picture and / or audio acquisition module connected to or included in the system 11. In this case, the processing module 500 is responsible for the encoding of these data.
[0147] System 11 may provide output signals to a variety of output devices (eg, system 13) that are capable of storing and / or decoding the output signals.
[0148] Various implementations involve decoding. "Decoding" as used in this application may encompass, for example, all or part of a process performed on a received encoded video stream (i.e., encoded video data) to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and prediction. In various embodiments, such a process also includes or alternatively includes a process performed by a decoder of the various embodiments described in this application, such as for checking the reliability of the output of a simulated implementation of the NN reasoning process.
[0149] Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will become clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0150] Various implementations involve encoding. Similar to the above discussion about "decoding", "encoding" as used in this application can cover, for example, all or part of the process performed on an input video sequence to produce an encoded video stream (i.e., produce encoded video data). In various embodiments, such a process includes one or more processes typically performed by an encoder, such as partitioning, prediction, transformation, quantization, and entropy coding. In various embodiments, such a process also includes or alternatively includes a process performed by an encoder of the various embodiments described in this application, such as for signaling information that allows checking the reliability of the output of a simulated implementation of the NN reasoning process.
[0151] Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will become clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0152] It should be noted that the syntax element names used in this article are descriptive terms. Therefore, they do not exclude the use of other syntax element names.
[0153] When a diagram is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.
[0154] Various embodiments relate to rate-distortion optimization. Specifically, in the encoding process, a balance or trade-off between rate and distortion is usually considered. Rate-distortion optimization is usually expressed as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different methods to solve the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all coding options, including all considered modes or coding parameter values, and fully evaluating their coding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save coding complexity, especially based on prediction or prediction residual signals instead of reconstructed signals to calculate approximate distortion. It is also possible to mix the two methods, such as using approximate distortion only for some possible coding options and using complete distortion for other coding options. Other methods only evaluate a subset of possible coding options. Digital implementations based on digital devices can also reduce the complexity of rate-distortion optimization. More generally, many methods use any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of coding costs and associated distortion.
[0155] The implementations and aspects described herein may be implemented as, for example, methods or processes, devices, software programs, data streams, or signals. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method), the implementation of the features discussed may also be implemented in other forms (e.g., devices or programs). Devices may be implemented as, for example, appropriate hardware, software, and firmware. Methods may be implemented as, for example, processors, which generally refer to processing devices, such as computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate information communication between end users.
[0156] Reference to "an embodiment" or "an embodiment" or "an implementation" or "an implementation" and other variations thereof means that the specific features, structures, characteristics, etc. associated with the embodiment are included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations thereof appearing in multiple places in the present application do not necessarily refer to the same embodiment.
[0157] Furthermore, the present application may involve "determining" various information. Determining information may include one or more of the following, such as estimating information, calculating information, predicting information, retrieving information from a memory, or obtaining information, such as from another device, module, or from a user.
[0158] Furthermore, the present application may involve "accessing" various information. Accessing information may include one or more of the following, such as receiving information, retrieving information (e.g., retrieving from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0159] Furthermore, the present application may involve "receiving" a variety of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of the following, such as accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally involves in some way during an operation such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0160] It should be understood that the use of any of the following " / ", "and / or", and "at least one", such as in the case of "A / B", "A and / or B", and "at least one of A and B", is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many items as listed as will be apparent to one of ordinary skill in this and related arts.
[0161] In addition, as used herein, the term "signal" refers to, among other things, indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals the use of some coding tools. Thus, in an embodiment, the same parameters are used on the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicitly signal) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters as well as other parameters, signaling (implicit signaling) can be performed without transmission to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various embodiments by avoiding transmission of any actual function. It should be understood that signaling can be done in a variety of ways. For example, in various embodiments, one or more grammatical elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the word "signal" is mentioned above, the word "signal" can also be used as a noun here.
[0162] As will be appreciated by one of ordinary skill in the art, implementations may generate a variety of formatted signals to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method, or data generated by one of the implementations. For example, a signal may be formatted to carry an encoded video stream and an SEI message of the embodiment. Such a signal may, for example, be formatted as an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may, for example, include encoding a video stream and modulating a carrier with the encoded video stream. The information carried by the signal may, for example, be analog or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
[0163] In the embodiments described below, the modules of the decoding process (e.g. Figure 4 The invention provides a module for decoding a process based on a NN inference process performed by an analog device. Solutions are proposed for handling errors that occur in the NN inference process when it is performed by an analog device compared to the same NN inference process performed by a digital device. These solutions include two aspects:
[0164] • A signaling process performed at the encoder side, which includes inserting metadata derived from the output of the NN inference process performed by a digital device at the encoder side into the encoded video data.
[0165] The decoding process includes:
[0166] ○ Decoding metadata from the encoded video data;
[0167] ○ Based on the NN inference process executed on the simulated device and the decoded metadata, apply a reliable NN inference process to ensure that the output of the reliable process produces the expected output.
[0168] Fig. 6A An embodiment of a portion of a reliable NN inference process implemented by a simulation device, performed on the encoder side, is schematically presented.
[0169] Fig. 6A The process is performed by, for example, the processing module 500 of the system 11 in Figure 3 The encoding process shown is performed during the Fig. 6A The process is performed during in-loop filtering (step 317), out-of-loop filtering (e.g., performed during the pre-processing step 301), frame rate up-conversion or frame rate down-conversion (e.g., performed during the pre-processing step 301), and picture-level motion field derivation (e.g., performed during the motion estimation step 304 or the motion compensation step 305). In the following, we take in-loop filtering (e.g., SOA of ALF) based on the NN inference process as an example (step 317).
[0170] In step 601, the processing module 500 applies a NN inference process that implements in-loop filtering to the image. In step 601, the NN inference process is performed by a digital device. In this case, the output of the NN inference process is systematically correct and corresponds to an accurate expected result.
[0171] In step 602, the processing module 500 calculates metadata representing the reliability level of the NN reasoning process implemented by the analog device. In an embodiment, the metadata representing the reliability level of the NN reasoning process implemented by the analog device is a checksum or statistical parameter calculated based on the output samples obtained by the NN reasoning process implemented by the digital device in step 601. Various types of checksums or statistical parameters can be calculated in step 602. The checksum includes an error detection code. The error detection code includes, for example, a parity byte and a parity word, a cyclic redundancy check (CRC), a Reed-Solomon code, an MD5sum, etc. Statistical parameters include values representing the signal, such as minimum value, maximum value, mean value, variance, statistical moments higher than "2" order, histograms (full precision (each signal codeword corresponds to a histogram value, for example, there are "1024" histogram values for a 10-bit signal) or low precision (each signal codeword set has a histogram value, for example, there are 16 histogram values for a 10-bit signal, one histogram value splices the interval values of 64 consecutive codewords, or splices the interval values of more consecutive codewords when there is overlap between each interval)).
[0172] In step 603, the processing module 500 signals metadata in the encoded video stream (i.e., the encoded video data). For an in-loop filtering process such as SOA or ALF, since this is a picture-level process, the metadata is associated with the picture corresponding to the currently processed picture. The metadata can be signaled at different levels, for example:
[0173] In the picture header;
[0174] In the slice header;
[0175] In the adaptation parameter set (as defined in VVC);
[0176] Before the first CTU of the region to which the metadata applies;
[0177] After the last CTU of the metadata applicable area;
[0178] In a SEI message attached to the coded video data; and
[0179] In case of streaming protocols such as DASH, in the manifest file (Media Presentation Description).
[0180] Metadata can also be defined and signaled by region in a global structure (e.g., picture header, slice header, adaptation parameter set). For example, a picture can be split into rectangular regions specified in this global structure, and metadata can be signaled for each rectangular region.
[0181] It can be noted that if the SEI message containing the metadata is missing, the decoder will assume that the output obtained by the simulated device performing the NN inference process is systematically correct.
[0182] exist Fig. 6A In the above embodiment, the NN inference process is an in-loop filtering process, such as SOA or ALF. Since this is a picture-level process, it is natural to calculate the metadata (i.e., checksum) on the entire picture. However, the process can be applied at other granularity levels, which means that the checksum can be calculated and signaled at different levels: by picture, by slice, by tile, by sub-picture, by rectangular area.
[0183] When delays are acceptable for application use cases (e.g. non-live streaming applications), metadata calculation and signaling may even be performed per GOP (Group of Pictures) (i.e. metadata is calculated for all pictures of a GOP as a whole), per Intra-period (metadata is calculated for all pictures of an Intra-period as a whole), per consecutive Group of Pictures (e.g. per "4" consecutive Groups of Pictures in decoding (preferably) or display order).
[0184] In one variant, in step 601, in addition to using a digital device to apply an NN inference process that implements in-loop filtering to the picture to obtain an output output_d, the processing module 500 also uses an analog device to perform the NN inference process that implements in-loop filtering on the picture multiple times. For example, the analog device is used to iterate the NN inference process N times, where N is equal to 3, for example. Each iteration provides an output output_a(k), where k∈[0;N-1]. Based on an analysis of the consistency of output_a(k) relative to output_d, the encoder can decide whether to indicate metadata signal notification by inserting a flag in the encoded video data. If no signal notification is performed, the decoding module can safely use the analog device to apply the NN inference process.
[0185] In one variant, based on the analysis of the consistency of output_a(k), the encoder can indicate which type of device should be used to implement the NN reasoning process between analog devices and digital devices by inserting a flag in the encoded video data. On the decoder side, based on the value of the decoded flag, the decoder uses an analog device or a digital device to apply the NN reasoning process.
[0186] The consistency analysis of outputs output_a(0), ..., output_a(N-1) may include:
[0187] Check if all output_a(k) are the same as output_d.
[0188] Or whether the number of outputs output_a(k) (k=0..N-1) that is the same as output_d is higher than a threshold (eg, "9" times out of "10" iterations, output_a(k) is equal to output_d).
[0189] Or whether the maximum absolute value of the error between the output_a(k) sample and the output_d sample is higher than a threshold (eg, "1").
[0190] Or whether the average absolute value of the error between the output_a(k) samples and the output_d samples is above a threshold (e.g., "1" for a 10-bit signal and "0.25" for an 8-bit signal).
[0191] The benchmarks listed above are listed independently, but several of them can be combined and checked together to derive a flag to be inserted into the encoded video data, which indicates which type of device should be used to implement the NN reasoning process between analog devices and digital devices. For example, when the first benchmark is true, the flag can be set to "1" to indicate that the decoding module can confidently use analog devices to apply the NN reasoning process. Conversely, when the first benchmark is false, it is set to "0" to indicate that the decoding module must use digital devices to apply the NN reasoning process.
[0192] In one variant, the decision is based on the gradient of each sample at the output with respect to the input. As is well known, NN-based systems consist of multiple NN layers, such as convolutional layers. Each NN layer can be described as a function that first multiplies the input by a tensor, adds a vector called a bias, and then applies a nonlinear function to the resulting value. The shape (and other characteristics) of a tensor and a nonlinear function are called the architecture of the network. The values of the tensor and the bias are usually called weights. The parameters of the weights and (if applicable) the nonlinear function are called parameters. The architecture and parameters define a model. The parameters of the NN model are trained by applying an iterative learning process on a large amount of data. The iterative learning process includes modifying the parameters of each layer of the NN model based on the gradients of these parameters with respect to the intermediate inputs of that layer. The process also includes calculating the gradient of the output sample with respect to the input sample. The gradient value of each sample at the output with respect to the input represents the sensitivity of the output value to the intermediate calculations. When the gradient magnitude is low, this means that the output is less sensitive than when the gradient magnitude is high. The flag can be set to "1" to indicate that the decoding module can confidently use the analog device to apply the NN reasoning process when one or all of the following benchmarks are true:
[0193] The maximum absolute value of the gradient is above a threshold (e.g. "0.1").
[0194] The average absolute value of the gradient is above a threshold (e.g. "0.01").
[0195] Figure 6B An embodiment of a portion of a reliable NN inference process implemented by a simulation device, performed on the decoder side, is schematically presented.
[0196] Figure 6B The process is performed by, for example, the processing module 500 of the system 13 in Figure 4 is performed during the decoding process shown. For example, Figure 6B The process is performed during in-loop filtering (step 417), out-of-loop filtering (e.g., performed during post-processing step 421), frame rate up-conversion (e.g., performed during pre-processing step 421), and image-level motion field derivation (e.g., performed during motion compensation step 305).
[0197] In the following, we take in-loop filtering (such as SAO or ALF) based on the NN inference process as an example (step 417).
[0198] In step 611 , the processing module 500 obtains the encoded video data 311 , which includes metadata representing the reliability level of the NN inference process implemented by the simulation device, which metadata was calculated in step 602 and signaled in step 603 .
[0199] In step 612, the processing module 500 decodes the video data until step 417. In step 417, the processing module 500 performs the decoding process using the NN inference process that implements the in-loop filtering process. During step 417, the simulation device implements the NN inference process relying on the metadata. Figure 7 An example of an embodiment of step 612 is described.
[0200] Figure 7 An application example of a reliable NN inference process implemented by a simulation device is schematically shown.
[0201] Figure 7 The process of FIG. 1 shows an embodiment of step 612 performed by the processing module 500 of the system 13 .
[0202] In the embodiment, we Figure 7 It is assumed that metadata includes information about Fig. 6A The checksum described.
[0203] In step 6120, the processing module 500 decodes the metadata and obtains a checksum. The checksum is calculated based on the output of the NN inference process implemented by the digital device 317 to implement the in-loop filtering process.
[0204] In step 6121, the processing module 500 uses the analog device to apply the NN inference process implementing the in-loop filtering process and obtain an output. In the case of in-loop filtering, the output is, for example, the entire picture.
[0205] In step 6122, the processing module 500 calculates a checksum based on the output of the NN inference process implemented by the simulation device to implement the in-loop filtering process. The checksum calculated in step 6122 is of the same type as the checksum calculated in step 602 (i.e., calculated using the same checksum calculation algorithm).
[0206] In step 6123 , the processing module 500 compares the checksum calculated in step 6122 with the signaled checksum decoded in step 6120 .
[0207] In step 6124, if the two checksums are the same, the processing module 500 proceeds to step 6125. Otherwise, the processing module 500 proceeds to step 6126.
[0208] In step 6125, the processing module 500 uses the output of the NN inference process implemented by the analog device to perform in-loop filtering for decoding. In this case, the picture produced by the in-loop filtering process is inserted into the DPB 419 in step 418 so that it can be used as a reference picture for temporal prediction.
[0209] In step 6126, the processing module 500 increases the iteration counter by one unit.
[0210] In step 6127, the processing module 500 determines whether the iteration counter reaches the maximum number of iterations Nmax. The maximum number of iterations Nmax is, for example, equal to "10" iterations. In this case, applying the in-loop filtering process multiple times is not a problem because when performed by an analog device, the process consumes much less time and energy than when performed by a digital device.
[0211] If the iteration number indicated by the iteration counter is lower than the maximum iteration number Nmax, the processing module 500 loops back to step 6121 .
[0212] In a first variant, the maximum number of iterations Nmax is chosen such that, on average, out of a number of iterations equal to the maximum number of iterations Nmax, at least one iteration allows obtaining a calculated checksum equal to the signal checksum.
[0213] In a second variation, if after a number of iterations equal to the maximum number of iterations Nmax, there is no calculated checksum equal to the signal checksum, then a process is applied in step 6128 to determine the best output among the Nmax outputs of the plurality of iterations. For example, the best output is:
[0214] the median of the outputs output_a(k) obtained by applying the NN inference process Nmax times by the simulated device, where k=0...Nmax-1 (e.g., for each sample position, the sample value is calculated as the median of the samples at the same position from output_a(k) (k=0...Nmax-1);
[0215] The arithmetic or geometric mean of output output_a(k) (e.g., for each sample position, the sample value is calculated as the arithmetic or geometric mean of the samples at the same position from output_a(k) (k=0...Nmax-1));
[0216] an output selected among the Nmax outputs output_a(k) that makes the calculated statistical parameter and the signaled statistical parameter most similar. This variant applies when the metadata contains the statistical parameter. For example, if the statistical parameter is a histogram, the selected output corresponds to the output output_a(k) having a histogram that is closest to the histogram represented by the signaled checksum. For example, the well-known Kullback-Leibler or Jensen-Shannon divergence can be used to calculate the distance between two histograms.
[0217] The best output is then selected to continue the decoding process. In the case of a NN inference process implementing in-loop filtering, the best output corresponds to the picture inserted into the DPB 419 in step 418.
[0218] In a third variant, the maximum number of iterations Nmax is at least equal to "1". In this variant, if each calculated checksum is different from the signaled checksum, the output of the NN reasoning process is discarded in step 6128, and the processing module 500 applies the NN reasoning process using the digital device in step 6129.
[0219] In the variant, Figure 7 The process includes additional steps 6118 and 6119.
[0220] In step 6118 , the processing module 500 decodes a flag indicating whether a checksum signal notification is performed in the encoded video data 311 from the encoded video data 311 .
[0221] If no checksum signaling is performed in the encoded video data, then in step 6119, the processing module 500 decides to directly apply step 6125. Otherwise, when the flag indicates that a checksum signaling is performed in the encoded video data, the processing module applies the process corresponding to steps 6120 to 6129.
[0222] In a variant, step 612 includes only step 6118 and step 6119. In step 6118, the processing module 500 decodes a flag from the encoded video data 311, the flag indicating which device is used to apply the NN reasoning process between the digital device and the analog device. In step 6119, if the decoded flag indicates that the NN reasoning process is implemented using the analog device, the processing module 500 decides so, otherwise the NN reasoning process is implemented using the digital device.
[0223] Scalable coding involves encoding video at different quality levels in terms of distortion (SNR scalability), frame rate (temporal scalability), and spatial resolution (spatial scalability). The encoded scalable video typically includes multiple layers, including a base layer that displays the lowest quality, lowest frame rate, and lowest spatial resolution, and at least one enhancement layer that enhances the quality, frame rate, and spatial resolution of the base layer. The enhancement layer typically relies on the base layer, and the enhancement of spatial resolution, temporal resolution, and distortion is obtained by combining the decoding results of the base layer with the decoding results of the enhancement layer. Many scalable codecs (encoders and decoders) include multiple interconnected decoding stages, each decoding stage corresponding to a scalable layer.
[0224] In an embodiment applicable to a scalable codec, a NN inference process implemented by an analog device is introduced in at least one decoding stage corresponding to an enhancement layer. In this embodiment, when decoding each enhancement layer including the NN inference process implemented by the analog device, the NN inference process is applied. Figure 7 process.
[0225] In a first variant, when there is no calculated checksum equal to the signaled checksum, the processing module 500
[0226] No digital equipment is used to implement the NN inference process;
[0227] The best output is not determined from the Nmax outputs calculated using simulation equipment;
[0228] Instead, the output of the NN inference process performed by the analog device for the enhancement layer is replaced with the output generated from the data calculated for the base layer. For example, when the NN inference process implements an in-loop filter, the output selected in step 6128 is a picture derived from the picture of the base layer, for example, resampled to the spatial resolution of the enhancement layer in the case of spatial scalability, or interpolated in the picture of the base layer in the case of temporal scalability. The picture derived from the picture of the base layer is then inserted into the DPB of the relevant enhancement layer. Otherwise, the output of the NN inference process performed by the analog device that implements in-loop filtering is used in step 6125.
[0229] It can be noted that in the case of a NN inference process implementing in-loop filtering, the decoding result of the enhancement layer before in-loop filtering is usually combined with the base layer data to generate an intermediate enhancement layer picture, and the in-loop filtering is applied to the intermediate enhancement layer picture to obtain the final enhancement layer picture. Figure 7 The process is applied to the combined result of the decoded base layer data and the intermediate result of the enhancement layer decoding.
[0230] In the variant, Figure 7 The process is performed before combining the base layer data with the partially decoded enhancement layer data. For example, this is the case in the case of inter-layer prediction when the enhancement layer and the base layer have different spatial (or temporal) resolutions. Inter-layer prediction means that the enhancement layer data is at least partially predicted from the base layer data. The inter-layer prediction data may include samples, motion information, etc. When the two layers have different spatial (or temporal) resolutions, inter-layer prediction may require upsampling (or interpolation) of the base layer data. In this variant, the upsampling (or interpolation) of the base layer data is implemented by a NN inference process performed by an analog device. Figure 7 The process is used to check the reliability of the upsampled (or interpolated) data output by the NN inference process performed by the simulation device. If the calculated checksum is equal to the signaled checksum (steps 6123, 6124), step 6125 is applied and the output upsampled (or interpolated) data is combined with the enhancement layer data. Otherwise, the output of the enhancement layer is generated from the base layer without using the enhancement layer data (step 6129).
[0231] We have described many embodiments above. The features of these embodiments may be provided individually or in any combination. In addition, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination across multiple claim categories and types:
[0232] • A bitstream or signal comprising one or more of said syntax elements or variants thereof.
[0233] • Creating and / or sending and / or receiving and / or decoding a bitstream or signal comprising one or more of said syntax elements or variants thereof.
[0234] A television, set-top box, mobile phone, tablet computer or other electronic device that executes at least one of the described embodiments.
[0235] A television, set-top box, mobile phone, tablet computer or other electronic device that performs at least one of the described embodiments and displays (eg, using a monitor, screen or other type of display) a resulting image.
[0236] A television, set-top box, mobile phone, tablet computer, or other electronic device that tunes (eg, using a tuner) a channel to receive a signal including encoded video data and performs at least one of the described embodiments.
[0237] A television, set-top box, mobile phone, tablet computer, or other electronic device that wirelessly receives (eg, using an antenna) a signal including encoded video data and performs at least one of the described embodiments.
[0238] A server, camera, cell phone, tablet or other electronic device that wirelessly transmits (eg, using an antenna) a signal including the encoded video data and performs at least one of the described embodiments.
[0239] A server, camera, cell phone, tablet or other electronic device that transmits a signal including encoded video data by tuning (eg, using a tuner) a channel and performs at least one of the described embodiments.
Claims
1. A method comprising: Acquiring (611) video data including metadata indicating a reliability level of a neural network inference process implemented by an analog device, the neural network inference process being applied to decode the video data, the analog device being an electronic circuit that cannot ensure repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; as well as; The video data is decoded (612), wherein the simulation device implements the neural network inference process in dependence on the metadata.
2. The method of claim 1 , wherein the metadata includes a flag indicating that there is a first checksum representing an expected output of the neural network reasoning process or there is a first statistical parameter representing an expected output of the neural network reasoning process in the metadata.
3. The method of claim 1 or 2, wherein the metadata comprises a first checksum or a first statistical parameter representing an expected output of the neural network inference process.
4. The method according to claim 3, wherein the simulation device implements the neural network reasoning process comprises: executing the neural network reasoning process at least once using the simulation device, and determining that a second checksum or a second statistical parameter calculated based on an output of one of the at least one execution of the neural network reasoning process by the simulation device is equal to the first checksum or the first statistical parameter, respectively, contained in the metadata, and Decoding the video data includes: in response to the second checksum being equal to the first checksum or in response to the second statistical parameter being equal to the first statistical parameter, decoding the video data using an output obtained by the simulation device performing the neural network inference process.
5. The method of claim 4, wherein decoding the video data comprises: In response to each second checksum or each second statistical parameter calculated based on the output obtained by the analog device each time performing the neural network inference process being different from the first checksum or the first statistical parameter, respectively, the video data is decoded using the output obtained by the digital device performing the neural network inference process, wherein the digital device is an electronic circuit that ensures repeatability of the output result of the electronic circuit when the electronic circuit receives the same input data.
6. The method of claim 4, wherein the neural network inference process is performed multiple times, and in response to each second checksum or second statistical parameter calculated based on the output obtained by the simulation device each time the neural network inference process is performed being different from the first checksum or the first statistical parameter, respectively, decoding the video data comprises: The video data is decoded using an output corresponding to one of the following: a median value of outputs obtained by the simulation device executing the neural network inference process multiple times; an arithmetic or geometric mean of outputs obtained by the simulation device executing the neural network inference process multiple times; as well as An output selected from outputs obtained by the simulation device performing the neural network inference process multiple times, wherein the selection makes the second statistical parameter calculated from the selected output closest to the first statistical parameter.
7. The method of claim 4, wherein the video data represents scalable video including a base layer and at least one enhancement layer, and the neural network inference process is included in a portion of a decoding process associated with decoding the enhancement layer, and wherein in response to each second checksum or second statistical parameter calculated based on an output obtained by each execution of the neural network inference process by the simulation device being different from the first checksum or the first statistical parameter, respectively, decoding the video data comprises: The video data is decoded by replacing each output obtained by the simulation device performing the neural network inference process with an output generated from the data calculated for the base layer.
8. The method of claim 2, wherein the neural network inference process is performed by the simulation device in response to the flag indicating that the first checksum is not present in the metadata or the first statistical parameter is not present in the metadata.
9. The method according to claim 1, wherein the metadata includes a flag indicating which device is used to perform the neural network inference process among an analog device and a digital device, wherein the digital device is an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data.
10. A method comprising: encoding video data using an encoding process including a neural network inference process performed by a digital device, the digital device being an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data (601); Calculating (602) metadata representing a reliability level of the neural network reasoning process implemented by an analog device using an output obtained by the digital device executing the neural network reasoning process, the analog device being an electronic circuit that cannot ensure repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; and The metadata is signaled in the encoded video data (603).
11. The method of claim 10, wherein the metadata is signaled in one of the following: Picture header; Slice header; Adaptive parameter set; before the first coding tree unit of the region to which the metadata applies; after the last coding tree unit of the region to which the metadata applies; In a SEI message attached to the coded video data; as well as In a manifest file that conforms to the streaming protocol.
12. The method according to claim 10 or 11, wherein the metadata is calculated per picture group or per picture or per slice or per tile or per sub-picture.
13. A method according to claim 10, 11 or 12, wherein the metadata includes a checksum or statistical parameters representing the expected output of the neural network inference process.
14. A method according to any one of the preceding claims 10 to 12, wherein the method further comprises: Analyzing the consistency of the neural network reasoning process performed by the analog device by performing the neural network reasoning process multiple times using the analog device and comparing the outputs obtained by the analog device performing the neural network reasoning process multiple times with the outputs obtained by the digital device performing the neural network reasoning process; as well as In response to the number of outputs obtained by the analog device performing the neural network inference process multiple times that is equal to the output obtained by the digital device performing the neural network inference process being less than a value, inserting a flag in the metadata, the flag indicating that the metadata includes a checksum or a statistical parameter, otherwise the flag indicates that the checksum or the statistical parameter does not exist in the metadata.
15. The method of claim 14, wherein the flag indicates that the checksum or the statistical parameter is not present in the metadata if all outputs obtained by the analog device performing the neural network inference process multiple times are equal to the outputs obtained by the digital device performing the neural network inference process.
16. A device comprising a processor, the processor being configured to: Acquiring video data including metadata indicating a reliability level of a neural network inference process implemented by an analog device, the neural network inference process being applied to decode the video data, the analog device being an electronic circuit that fails to ensure repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; and The video data is decoded, wherein the simulation device implements the neural network inference process in dependence on the metadata.
17. The apparatus of claim 16, wherein the metadata includes a flag indicating that there is a first checksum in the metadata representing an expected output of the neural network reasoning process or there is a first statistical parameter representing an expected output of the neural network reasoning process.
18. The apparatus of claim 16 or 17, wherein the metadata comprises a first checksum or a first statistical parameter representing an expected output of the neural network inference process.
19. The device according to claim 18, wherein the simulation device implements the neural network reasoning process comprises: executing the neural network reasoning process at least once using the simulation device, and determining that a second checksum or a second statistical parameter calculated based on an output of one of the at least one execution of the neural network reasoning process by the simulation device is equal to the first checksum or the first statistical parameter, respectively, contained in the metadata, and Decoding the video data includes: in response to the second checksum being equal to the first checksum or in response to the second statistical parameter being equal to the first statistical parameter, decoding the video data using an output obtained by the simulation device performing the neural network inference process.
20. The apparatus of claim 19, wherein decoding the video data comprises: In response to each second checksum or each second statistical parameter calculated based on the output obtained by the analog device each time performing the neural network inference process being different from the first checksum or the first statistical parameter, respectively, the video data is decoded using the output obtained by the digital device performing the neural network inference process, the digital device being an electronic circuit that ensures repeatability of the output result of the electronic circuit when the electronic circuit receives the same input data.
21. The apparatus of claim 19, wherein the neural network inference process is performed multiple times, and in response to each second checksum or second statistical parameter calculated based on the output obtained by each execution of the neural network inference process by the simulation device being different from the first checksum or the first statistical parameter, respectively, decoding the video data comprises: The video data is decoded using an output corresponding to one of the following: a median value of outputs obtained by the simulation device executing the neural network inference process multiple times; an arithmetic or geometric mean of outputs obtained by the simulation device executing the neural network inference process multiple times; An output selected from outputs obtained by the simulation device performing the neural network inference process multiple times, wherein the selection makes the second statistical parameter calculated from the selected output closest to the first statistical parameter.
22. The apparatus of claim 19, wherein the video data represents scalable video comprising a base layer and at least one enhancement layer, and the neural network inference process is included in a portion of a decoding process associated with decoding the enhancement layer, and wherein in response to each second checksum or second statistical parameter calculated based on an output of each execution of the neural network inference process by the simulation device being different from the first checksum or the first statistical parameter, respectively, decoding the video data comprises: The video data is decoded by replacing each output obtained by the simulation device performing the neural network inference process with an output generated from the data calculated for the base layer.
23. The apparatus of claim 17, wherein the neural network inference process is performed by the simulation apparatus in response to the flag indicating that the first checksum is not present in the metadata or the first statistical parameter is not present in the metadata.
24. The device of claim 16, wherein the metadata includes a flag indicating which device is used to perform the neural network inference process between an analog device and a digital device, the digital device being an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data.
25. An apparatus comprising an electronic circuit configured to: encoding the video data using an encoding process including a neural network inference process performed by a digital device, the digital device being an electronic circuit that ensures repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; Calculating metadata representing a reliability level of the neural network reasoning process implemented by an analog device using an output obtained by executing the neural network reasoning process by the digital device, the analog device being an electronic circuit that cannot ensure repeatability of an output result of the electronic circuit when the electronic circuit receives the same input data; as well as The metadata is signaled in the encoded video data.
26. The apparatus of claim 25, wherein the metadata is signaled in one of: Picture header; Slice header; Adaptive parameter set; before the first coding tree unit of the region to which the metadata applies; after the last coding tree unit of the region to which the metadata applies; In a SEI message attached to the coded video data; as well as In a manifest file that conforms to the streaming protocol.
27. The apparatus according to claim 25 or 26, wherein the metadata is calculated per picture group or per picture or per slice or per tile or per sub-picture.
28. The apparatus of claim 25, 26 or 27, wherein the metadata comprises a checksum or statistical parameters representing an expected output of the neural network inference process.
29. The device according to any one of the preceding claims 25 to 28, wherein the device is further configured to: Analyzing the consistency of the neural network reasoning process performed by the analog device by performing the neural network reasoning process multiple times using the analog device and comparing the outputs of the neural network reasoning process performed multiple times by the analog device with the outputs of the neural network reasoning process performed by the digital device; and In response to the number of outputs obtained by the analog device performing the neural network inference process multiple times that is equal to the output obtained by the digital device performing the neural network inference process being less than a value, inserting a flag in the metadata, the flag indicating that the metadata includes a checksum or a statistical parameter, otherwise the flag indicates that the checksum or the statistical parameter does not exist in the metadata.
30. The apparatus of claim 29, wherein the flag indicates that the checksum or the statistical parameter is not present in the metadata if all outputs of multiple executions of the neural network inference process by the analog device are equal to outputs of the neural network inference process by the digital device.
31. A signal comprising metadata indicating a reliability level of a neural network inference process of a video decoding process implemented by a simulation device.
32. A computer program comprising program code instructions for implementing the method according to any preceding claim from 1 to 15.
33. A non-transitory information storage medium storing program code instructions for implementing the method according to any one of the preceding claims 1 to 15.