Method or apparatus for scaling tensor of feature data using interpolation filter
By adopting scalable neural network-based transformation and rescaling processing of feature data tensors in video decoding, the problem of insufficient optimization of computer vision algorithms in the prior art is solved, and more efficient video decoding and better visual inference performance are achieved.
Patent Information
- Application Number
- CN202380070844.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-07
- Filing Date
- 2023-10-02
- Publication Date
- 2025-05-13
AI Technical Summary
Existing video encoding and decoding methods fail to effectively optimize compression schemes for computer vision algorithms, resulting in performance bottlenecks in machine and human vision applications.
By one approach, videos are decoded using scalable neural network-based transformations and rescaling the tensors of feature data to adapt to the input requirements of NN-based visual inference tasks. The method includes obtaining a tensor of an image data sample reconstructed from a bitstream, applying a neural network-based feature synthesis process, adjusting the size of the feature tensor, and performing visual inference processing on it to generate inference results.
This method improves the efficiency and performance of video decoding, especially in hybrid machine/human vision applications, enhances encoding efficiency and improves codec consistency.
Smart Images

Figure CN119999196A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 414,053, filed on October 7, 2020, which is incorporated herein by reference in its entirety. Technical Field
[0003] At least one of the present embodiments generally relates to a method or apparatus for video encoding or decoding, and more particularly to a method or apparatus for decoding a video using a scalable NN-based transform, which decoding also includes rescaling tensors of feature data intended to be fed to an NN-based visual reasoning task. Background Art
[0004] Conventional compression standards achieve low bitrates by transforming and degrading video content using methods optimized to maintain signal fidelity or visual quality. To this end, conventional image and video coding schemes typically employ prediction (including motion vector prediction) and transforms to exploit spatial and temporal redundancy in video content. Typically, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations, and then the difference between the original image and the predicted image (usually denoted as prediction error or prediction residual) is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to entropy coding, quantization, transform, and prediction.
[0005] In recent years, new image and video compression methods based on neural networks (NNs) have been developed. In contrast to traditional methods that apply predefined prediction patterns and transformations, NN-based methods rely on many parameters that are learned on large datasets during a training phase by iteratively minimizing a loss function using some kind of gradient descent algorithm. In the case of compression, the loss function is defined by a rate-distortion cost, where the rate represents an estimate of the bitrate of the encoded bitstream and the distortion quantifies the quality of the decoded video with respect to some visual quality metric for the original input. Traditionally, the quality of the decoded input image is optimized, for example based on a measure of mean squared error or an approximation of the visual quality perceived by humans.
[0006] However, an increasing amount of visual content is now also analyzed directly by machines via deep learning based computer vision algorithms. Existing methods for encoding and decoding show some limitations, as the compression schemes are not optimized for computer vision algorithms. Therefore, there is a need to improve the state of the art by proposing compression schemes for images and videos that are targeted for both human and machine consumption. Summary of the invention
[0007] The shortcomings and deficiencies of the prior art are addressed and overcome by the general aspects described herein.
[0008] According to a first aspect, a method is provided. The method includes performing scalable video decoding by obtaining a tensor of reconstruction data representing image data samples partially reconstructed from a base layer of a bitstream, the tensor of reconstruction data including the number of channels of two-dimensional data; applying a neural network-based feature synthesis process to the tensor of reconstruction data to generate a tensor of input features representing features of the image data samples, the tensor of input features including the number of channels of two-dimensional data; adjusting the size of the tensor of input features to generate a tensor of output features, the tensor of output features including the number of channels of two-dimensional data; and applying a neural network-based visual reasoning process to the tensor of output features to generate a set of reasoning results. Advantageously, adjusting the size of the tensor of input features includes applying at least one interpolation filter to the tensor of input features to make at least one dimension of the tensor of input features suitable for the neural network-based visual reasoning process.
[0009] According to another aspect, an apparatus is provided. The apparatus comprises one or more processors, wherein the one or more processors are configured to implement a method for video decoding according to any variant thereof. According to another aspect, an apparatus for video decoding comprises a component for implementing a method for video decoding according to any variant thereof.
[0010] According to another general aspect of at least one embodiment, a 2D interpolation filter is applied to each channel of a tensor of input features to resize a spatial dimension of the tensor.
[0011] According to another general aspect of at least one embodiment, at least one convolutional filter is applied to a tensor of input features to scale the number of channels of the tensor.
[0012] According to another general aspect of at least one embodiment, information representing filters to be used in feature tensor resizing (filter type, filter coefficients, index of the filter in a predetermined filter set) is parsed from metadata of the bitstream.
[0013] According to another general aspect of at least one embodiment, there is provided an apparatus comprising an apparatus according to any one of the decoding embodiments; and at least one of: (i) an antenna configured to receive a signal comprising a video block, (ii) a frequency limiter configured to limit the received signal to a frequency band comprising the video block, or (iii) a display configured to display an output representing the video block.
[0014] According to another general aspect of at least one embodiment, there is provided a non-transitory computer-readable medium containing data content generated according to any of the described decoding embodiments or variations.
[0015] According to another general aspect of at least one embodiment, there is provided a signal comprising video data generated according to any of the described decoding embodiments or variants.
[0016] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described decoding embodiments or variations.
[0017] According to another general aspect of at least one embodiment, there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to carry out any of the described decoding / decoding embodiments or variants.
[0018] These and other aspects, features and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments, which is to be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In the drawings, several examples of embodiments are illustrated.
[0020] Figure 1 Illustrated is a block diagram of an example apparatus in which various aspects of the embodiments may be implemented.
[0021] Figure 2 A block diagram of an embodiment of a conventional video encoder is illustrated.
[0022] Figure 3 A block diagram of an embodiment of a conventional video decoder is illustrated.
[0023] Figure 4 A block diagram of an embodiment of an end-to-end neural network based video compression scheme is illustrated.
[0024] Figure 5 A block diagram of an embodiment of a basic pipeline for NN-based machine vision processing is illustrated.
[0025] Figure 6 A block diagram of an embodiment of a basic pipeline for NN-based video compression and machine vision processing is illustrated.
[0026] Figure 7 Illustrated is a block diagram of a variant embodiment of a basic pipeline of NN-based video compression and machine vision processing, in which various aspects of the embodiments can be implemented.
[0027] Figure 8 A block diagram illustrating a detailed embodiment of a basic pipeline for NN-based vision processing is shown.
[0028] Fig. 9 and Fig.10 A general tensor scaling method according to a general aspect of at least one embodiment is illustrated.
[0029] Fig.11 A general decoding method implementing tensor scaling according to a general aspect of at least one embodiment is illustrated.
[0030] Fig.12 , Fig.13 , Fig.14 , Fig.15 A variant embodiment of a general tensor scaling method is illustrated.
[0031] Fig.16 A general method for implementing parsing information related to tensor scaling filtering is illustrated in accordance with a general aspect of at least one embodiment.
[0032] Fig.17 Two remote devices communicating over a communication network are shown in accordance with an example of the present principles in which various aspects of the embodiments may be implemented.
[0033] Fig.18 The syntax of a signal according to an example of the present principles is shown. DETAILED DESCRIPTION
[0034] Various embodiments relate to video coding systems, wherein, in at least one embodiment, it is proposed to adapt video decoding tools to hybrid machine / human vision applications. Different embodiments are proposed below, introducing some tool modifications to increase coding efficiency and improve codec consistency when both applications are targeted. Among other things, a decoding method and a decoding device implementing a tensor resizing module based on the principle are proposed.
[0035] The present aspects are described in the context of the ISO / MPEG Working Group 2 known as Video Coding for Machines (VCM) and in the context of JPEG-AI. Video Coding for Machines (VCM) is an MPEG activity that aims to standardize the bitstream format generated by compressing a video stream or previously extracted features. The bitstream should enable a variety of machine vision tasks by embedding the information necessary to perform a variety of tasks at the receiver, such as segmentation, object tracking, and reconstruction of video content for human consumption. In parallel, JPEG is standardizing JPEG-AI, which is expected to involve an end-to-end NN-based image compression method that can also be optimized for some machine analysis tasks. With use cases such as video surveillance, autonomous vehicles, smart cities, etc. already ubiquitous, it is easy to envision other similar-style standards and upcoming systems of the VCM paradigm in the near future.
[0036] The present aspects are not limited to those standardization efforts, and may be applied, for example, to other standards and recommendations (whether pre-existing or developed in the future), as well as extensions of any such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.
[0037] The acronyms used in this article reflect the current state of video coding development and should therefore be considered examples of nomenclature that may be renamed at a later stage while still denoting the same technology.
[0038] Figure 1 The block diagram of the example of the system in which various aspects and embodiments can be implemented is illustrated. System 100 can be embodied as the equipment including various components described below, and is configured to perform one or more aspects described in the present application. Examples of such equipment include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances and servers. The elements of system 100 can be embodied in a single integrated circuit, multiple ICs and / or discrete components, either individually or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in the present application.
[0039] The system 100 includes at least one processor 110, which is configured to execute instructions loaded therein, for implementing various aspects described in the present application, for example. The processor 110 may include embedded memory, input and output interfaces, and various other circuits known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include a non-volatile memory and / or a volatile memory, including but not limited to an EEPROM, a ROM, a PROM, a RAM, a DRAM, an SRAM, a flash memory, a disk drive, and / or an optical drive. As a non-limiting example, the storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device.
[0040] The system 100 includes an encoder / decoder module 130, which is configured to process data to provide encoded video or decoded video, for example, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents (one or more) modules that can be included in the device to perform encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, the encoder / decoder 130 module represents (one or more) modules that can be included in the device for performing machine vision processing (or network) on the decoded data to complete the reasoning output, thereby implementing a decoding tool for hybrid machine / human vision applications using NN-based tools. Additionally, the encoder / decoder module 130 can be implemented as a separate element of the system 100, or can be incorporated into the processor 110 as a combination of hardware and software known to those skilled in the art.
[0041] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described in the present application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in the present application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations and operation logic, tensors, networks, or filter weights.
[0042] In several embodiments, memory inside the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be a memory 120 and / or a storage device 140, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations (such as for HEVC or VVC).
[0043] Input to the elements of system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0044] In various embodiments, the input device of block 105 has associated corresponding input processing elements, as known in the art. For example, the RF portion may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again limiting a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF portion of various embodiments includes one or more elements to perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the frequency band of expectation again by filtering, down-conversion and perform frequency selection.Various embodiments rearrange the order of above-mentioned (and other) elements, remove some in these elements, and / or add other elements of performing similar or different functions.Adding element can include inserting element between existing element, for example, inserting amplifier and analog-to-digital converter.In various embodiments, the RF part comprises antenna.
[0045] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 100 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or in the processor 110 as desired. Similarly, various aspects of USB or HDMI interface processing may be implemented in a separate interface IC or in the processor 110 as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and the encoder / decoder 130, operating in combination with memory and storage elements to process the data streams as desired for presentation on an output device.
[0046] The various elements of system 100 may be provided within an integrated housing within which the various elements may be interconnected and transmit data therebetween using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit boards.
[0047] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.
[0048] In various embodiments, data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received through a communication channel 190 and a communication interface 150 suitable for Wi-Fi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Other embodiments provide streaming data to the system 100 using a set-top box that passes data through an HDMI connection of the input box 105. Still other embodiments provide streaming data to the system 100 using an RF connection of the input box 105.
[0049] The system 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripherals 185. In various examples of embodiments, the other peripherals 185 include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functions based on the output of the system 100. In various embodiments, control signals are transmitted between the system 100 and the display 165, the speaker 175, or other peripherals 185 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 100 via dedicated connections through the corresponding interfaces 160, 170, and 180. Alternatively, the output devices can be connected to the system 100 using a communication channel 190 via the communication interface 150. The display 165 and the speaker 175 can be integrated into a single unit with other components of the system 100 in an electronic device (e.g., a television). In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.
[0050] For example, if the RF portion of input 105 is part of a stand-alone set-top box, the display 165 and speaker 175 may alternatively be separate from one or more other components. In various embodiments where the display 165 and speaker 175 are external components, the output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.
[0051] Figure 2 An example video encoder 200 is illustrated, such as a VVC (Versatile Video Coding) encoder. Figure 2 It is also possible to illustrate an encoder that has improved the VVC standard or an encoder that uses a technique similar to VVC.
[0052] In this application, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "encoded" or "coded" may be used interchangeably, and the terms "image", "picture" and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.
[0053] Before being encoded, the video sequence may undergo a pre-encoding process (201), for example, applying a color transform to an input color picture (e.g., converting from RGB 4:4:4 to YCbCr4:2:0), or performing a remapping of input picture components in order to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.
[0054] In encoder 200, a picture is encoded by encoder elements as described below. The picture to be encoded is partitioned (202) and processed in units such as CUs. Each unit is encoded using, for example, intra mode or inter mode. When a unit is encoded in intra mode, it performs intra prediction (260). In inter mode, motion estimation (275) and compensation (270) are performed. The encoder decides (205) which of intra mode or inter mode to use to encode the unit and indicates the intra / inter decision by, for example, a prediction mode flag. For example, a prediction residual is calculated by subtracting (210) the predicted block from the original image block.
[0055] The prediction residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal. The encoder may bypass both the transform and quantization, i.e., the residual is directly encoded without applying the transform or quantization process.
[0056] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored at a reference picture buffer (280).
[0057] Figure 3 A block diagram of an example video decoder 300, such as a VVC decoder, is illustrated. In the decoder 300, a bitstream is decoded by decoder elements, as described below. The video decoder 300 generally performs the same Figure 2 The encoder 200 also typically performs video decoding as part of encoding the video data.
[0058] In particular, the input to the decoder includes a video bitstream that may be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other encoding information. Picture partition information indicates how the picture is partitioned. Thus, the decoder can divide (335) the picture according to the decoded picture partition information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (355) to reconstruct the image block. The prediction block can be obtained (370) from intra-frame prediction (360) or motion compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored at a reference picture buffer (380).
[0059] The decoded picture may further undergo post-decoding processing (385), such as an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which performs the inverse of the remapping process performed in the pre-encoding process (201). The post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0060] At least some embodiments relate to a method for decoding video using a scalable NN-based transform, the decoding also including rescaling / resizing a tensor of feature data intended to be fed to a NN-based visual reasoning task. By enabling the resizing of an estimated deep feature tensor to another tensor of a different size, specific input size constraints imposed by a visual network can be achieved.
[0061] Figure 4 FIG. 1 is a block diagram of an embodiment of an end-to-end neural network based video compression scheme.a (Also called analytical transformation) transforms the input image X into the latent space: Y = g a (X). In most neural network-based compression frameworks, the latent representation Y is formed in a three-dimensional tensor (called the latent tensor). Then, Y is quantized (Q) and entropy encoded (EC) into a binary stream (bitstream) for storage or transmission. At the decoder, the bitstream is entropy decoded (ED) to obtain the quantized version of Y Decoder network g s (also called synthetic transform) generates the reconstructed input: ——Implicit Representation from Quantization An approximation of the original X of . For the sake of completeness, those skilled in the art will appreciate that in this figure, other modules (such as super priors and context prediction for further improving rate-distortion performance) are omitted to keep the processing pipeline in the provided figure simple. For clarity, the same form of omission will apply to the remaining figures provided in this document.
[0062] Figure 5 A block diagram of an embodiment of a basic pipeline of neural network-based machine vision processing is shown. NN-based visual reasoning tasks (visual networks) will be such as reconstructing images Most object detection and segmentation networks require the input to be resized to a specific resolution or to satisfy constraints before performing inference in order to maximize task accuracy. This is because these networks either need to run on images of a predefined size or be trained on images that make it easier for the algorithm to output bounding boxes associated with object categories to the identified objects. As such, First it is resized to and is thus fed into the vision network to output a set of inference results T.
[0063] Further optimizing existing video encoders for direct machine consumption (such as computer vision networks) is not easy because Figure 2 and Figure 3 On or Figure 4 The hand-coded tools of the compression schemes illustrated above are optimized only for rate-distortion (RD) cost in standard codecs. The performance of NN-based computer vision algorithms can be deteriorated by artifacts such as ringing, blocking artifacts and loss of high spatial frequencies produced by classic standard codecs for human consumption.
[0064] Figure 6 A block diagram of an embodiment of a basic pipeline for NN-based video compression and machine vision processing is illustrated. According to an embodiment for machine video coding (VCM), Figure 4 and Figure 5The schemes are concatenated, where the compressed input X is reconstructed and used as input to the vision network for reasoning in a sequential “chain” fashion. In this pipeline, the encoder g a and decoder g s And possibly an end-to-end compression network for vision networks can be jointly optimized for the two tasks considered (which are machine reasoning and input reconstruction) by maximizing the overall task accuracy.
[0065] Figure 7 A block diagram of a variant embodiment of the basic pipeline of NN-based video compression and machine vision processing is illustrated. Recently, H. Choi and IVBajic introduced a scalable architecture of NN-based compression for VCM in "Scalable Image Coding for Humans and Machines," (IEEE Transactions on Image Processing, vol. 31, pp. 2739-2754, 2022). Figure 7 As shown in Figure 6 Different, analytical transformation g a Two latent tensors are generated, which are then quantized to obtain and where C1 and C2 denote the number of channels of the first (basic) and second (enhanced) latent representations, respectively. s is the scaling factor between the spatial resolution of the latent tensor and the input image. This scaling factor is usually represented by g a OK, g a It usually consists of several convolutional layers with a stride of 2. Successively, the independently encoded latent tensors (i.e., the first and second bit streams) are used as input to g on the decoder side. a and f s Advantageously, the architecture is designed to support functional scalability from some simpler task (e.g., object detection) to more complex tasks (e.g., input reconstruction). For example, reconstructing each pixel with signal fidelity is desirable for input reconstruction, but not necessary for object detection. Therefore, the first bitstream carries an implicit representation for object detection. information, and the second bitstream conveys the remaining implicit representation It will include Together with the enhanced information used for input reconstruction. Since Choi described both two and three tasks supporting the VCM architecture, Figure 7 The VCM framework proposed in is not limited to two task variants, and those skilled in the art will easily adapt the described framework to more than two tasks. For the basic task (e.g., object detection), the feature synthesis module f s (also called “latent space transformation” in Choi) only Take as input to estimate the deep feature tensor Then input this feature tensor into v e In, v e Corresponds to the part of the visual network starting from the lth layer to the last layer. The visual network outputs the set T. In Choi, since the front part of the visual network from the input layer to the l-1th layer has been omitted by the proposed architecture, the decoder advantageously requires less computation to perform the inference task. In order to reconstruct It is necessary s Will Both are taken as input.
[0066] However, Choi's VCM architecture proposes that Figure 7 The estimated feature tensor generated when the input X is compressed to its original scale as shown in The problem of incompatible spatial resolutions between and the expected feature tensor resolution of F is obtained by feeding X' (the resized input) into Figure 6 The calculation is done in the front-end visual network implemented in the sequential scheme of Due to the inconsistent dimensions between F and F, it is unlikely that the compression and vision task pipelines can achieve the best performance for all tasks at the same time.
[0067] Figure 8 A block diagram illustrating a detailed embodiment of a basic pipeline for NN-based machine vision processing is shown. Figure 8 Detailed example dimensions of the associated intermediate tensors are presented. Input Resized to to comply with the input resolution N×M constraint imposed by the pre-trained vision network. Then, at an intermediate layer l of the vision network, the feature tensor F turns out to have The dimensions of C l denotes the number of channels at layer l, and k corresponds to the scaling factor between the spatial resolution of the tensor and the input image, which is usually determined by the front-end vision network involving several convolutional layers with a stride of 2 and some pooling operations. Figure 7 The best performance on vision tasks is expected to be achieved when the compression framework shown in takes the resized X' instead of X as input, since Figure 8 When the input of the front-end visual network shown in , the feature tensor size satisfies v e The dimensions expected as input to layer l, so f s Generate of However, in this scenario, g s This will produce a reconstructed input with a resolution of 3×N×M Unless combined with an auxiliary post-processing module The size is resized to the original resolution of 3×H×W, otherwise the encoded bitstream can only reconstruct the input at the resized resolution 3×N×M.
[0068] A similar problem of resolution inconsistency occurs when encoding X instead of X'. In this case, the reconstructed Needs to be resized.
[0069] This is addressed and solved by the general aspects described herein, which involve a tensor for adjusting input features The size of the output feature tensor is The method outputs a tensor of features suitable for visual reasoning processing based on neural networks (v e ) dimensions. Advantageously, Fig.11 As shown above, for the case where (one or more) deep feature tensors are generated for various vision task algorithms, a resizing module is implemented on the decoder side.
[0070] Fig. 9 A general tensor scaling method according to a general aspect of at least one embodiment is illustrated. Fig. 9 The block diagram partially shows that for example Fig.11 Modules of a decoder method implemented in an exemplary decoder.
[0071] At the decoder, the decoded ED of the first (base) layer bitstream and the neural network-based feature synthesis process f s The feature tensor is intended to be fed into a neural network-based visual reasoning process v e , to produce a result T such as segmentation, object detection, object tracking. Advantageously, the decoder also includes the proposed resizing operation so that the size of the feature tensor is adapted to the expected size of the tensor at the input of the NN-based visual reasoning processing. Therefore, when more than two tasks are supported (which means that the decoder can support more than one visual task and input reconstruction task), the proposed resizing operation is applied to any task pipeline so that the size of the tensor of a given task pipeline (one visual task) is independent of the size of the tensor of another given task pipeline (input reconstruction task). Fig. 9 As shown above, ED first reconstructs the tensor of reconstruction data from the layer 1 bitstream The tensor of the reconstructed data partially represents the image data samples to be reconstructed. Subsequently, is fed to the feature synthesis module f s To generate the feature tensor The generated tensor Can be with rescaled spatial resolution A tensor of r, where r is a tensor of architecture f s and relative to v e The number of input channels C l The resizing module then uses an interpolation filter to adjust to obtain the size The information of these filters can be conveyed by signaling the indices associated with the standard filters shared in the decoder or by encoding the filter coefficients in some form of bitstream. is used as v e to complete the reasoning task and finally obtain the output T.
[0072] Fig.10 Another general tensor scaling method according to a general aspect of at least one embodiment is illustrated. Fig.10 The block diagram partially shows that for example Fig.11 In this variant embodiment, f s Generates a channel number equal to C1 Rather than Fig.10 The one shown in will be C l In this variant, further information about the set of convolutional filters with or without bias parameters is signaled in the bitstream so that the resizing module not only performs spatial resolution resizing but also performs convolution operations to generate a spatial resolution of Tensor of .
[0073] According to yet another variant, depending on the visual tasks to be supported at each task layer, v e There may be different constraints on the input size and the number of input channels. Therefore, appropriate information about interpolation filters and / or convolution filters can be carried in each bitstream for different task layers and applied to the feature tensor.
[0074] Fig.11 A general decoding method (1100) for implementing tensor scaling according to a general aspect of at least one embodiment is illustrated. Fig.11 In a preliminary step not shown above, a bitstream is received. As explained in Choi's scheme, the bitstream may include scalable data representing a video image, including a base layer of image data intended for computer vision tasks, an enhancement layer representing additional image data intended for human vision. The bitstream may also include additional metadata for processing the received bitstream. In a first step 1110, a tensor of reconstruction data is obtained. The tensor of reconstruction data partially represents the image data samples to be reconstructed. The tensor of reconstruction data includes a size of C of two-dimensional data 1 In a second step 1120, a NN-based feature synthesis process is applied to the tensor of reconstructed data to generate a tensor of input features representing features of the image data samples. The input feature tensor consists of size C of two-dimensional data l number of channels, where r is the rescaling ratio. In step 1130, the input feature tensor The tensor that is resized / rescaled to generate the output features Advantageously, resizing allows generating a tensor of output features It includes a size that matches the expected tensor size in the neural network-based visual reasoning processing 1140 (actually at the defined layer to generate a set of reasoning results T) C of two-dimensional data l Number of channels.
[0075] According to another embodiment, the decoding method further comprises obtaining 1150 information about an interpolation filter to apply the resizing according to any of the signaling variants described below.
[0076] According to another embodiment, the decoding method further includes, in step 1160, obtaining a tensor of enhanced data The tensor of augmented data relative to the tensor of reconstructed data Complementarily representing the image data samples to be reconstructed, the tensor of the reconstructed data includes a size of C of two-dimensional data 2 number of channels (with the same notation as above). According to the NN-based image synthesis process 1170, the decoding method generates a reconstructed image of size H x W, for example to be reproduced on a display for human vision.
[0077] According to yet another embodiment, the decoding may further output additional tensors intended for different visual reasoning tasks The NN-based feature synthesis process 1120 and resizing process 1130 are therefore instantiated additional times (parallel steps) to output additional tensors intended for different visual reasoning tasks with their tensor size requirements. where j here is a scaling factor between the spatial resolution of the tensor and the input image for the relevant visual reasoning task.
[0078] Various embodiments of a general decoding method are described in the following.
[0079] Fig.12A general tensor rescaling method (1130) according to general aspects of at least one embodiment is illustrated.
[0080] According to the first variant, the number of channels (C l ) and the number of channels of the output feature tensor (C l ). In one variant, for each channel, a 2D interpolation filter is applied with an input tensor of size The 2D data is rescaled to a size of Instead, for each channel, a 1D separable filter is applied to the 2D data of size The two dimensions of the 2D data are generated in any order with the sizes In another variant, parameters representing the 2D interpolation filtering, such as coefficients of the interpolation filter, are parsed from metadata embedded in the bitstream. Fig.12 shows the use of a 2D interpolation kernel parsed from the corresponding bitstream The sequence is used to adjust the input feature tensor The size of the output feature tensor is An example where K i Used to adjust tensors It is also possible to use separable filters for each channel instead of a 2D kernel, and this can be indicated in the bitstream.
[0081] Fig.13 Another general tensor rescaling method (1130) according to general aspects of at least one embodiment is illustrated.
[0082] In yet another variant, a set of 2D interpolation filters, such as bilinear filters, bicubic filters, trilinear filters, is predefined in the decoder. The index indicating the 2D interpolation filter among the set of interpolation filters can be parsed from metadata embedded in the bitstream.
[0083] like Fig.13 As shown above, the input feature tensor is reshaped using pre-existing filters in the decoder to obtain the output feature tensor In this variant, only index j is parsed from the bitstream, and the corresponding interpolation filter with j is then applied to All channels in the output According to yet another variant, it is possible to select different interpolation filters for different channels by parsing multiple filter indices from the bitstream.
[0084] In combination Fig.12 and Fig.13In a variation of the embodiment, each channel can have a different filter type among predefined filters (bilinear, bicubic, etc.) and custom adaptive filters with coefficients. In this case, a separate filter index associated with each channel is signaled, which covers any of the filter embodiments. For the case where the filter index indicates the use of an adaptive filter, the filter coefficients subsequently transmitted should be correctly parsed for that channel.
[0085] Fig.14 Another general tensor rescaling method (1130) according to general aspects of at least one embodiment is illustrated.
[0086] According to the second variant, the number of channels of the input feature tensor C1 and the number of channels of the output feature tensor C l is different. Therefore, resizing the tensor of input features also includes applying at least one convolution filter to the tensor of input features to scale the number of channels. In fact, for and Use cases with a different number of channels (i.e., C1 instead of C l ), the tensor must be resized along the channel axis in addition to the spatial axis. One way to resize the number of channels is to perform convolution operations with the associated filter information that should be transmitted in the bitstream. Fig.14 shows a method for converting the input feature tensor using parsing filter information obtained from the bitstream. Resize to The convolution operation is shown before the spatial interpolation operation. However, it is also possible to exchange the order of these operations. If necessary, the parsed convolution filter may include a filter set with a kernel size and a bias term. The output of the convolution block then generates a filter with channel C l The intermediate tensor is then spatially resized by the parsed interpolation filters obtained from the bitstream to produce According to yet another variant, the resized network can be compared to Figure 12-14 The resizing network may include, but is not limited to, more than one convolutional layer or any type of trainable layer and activation layer before interpolation or even after the interpolation filter. In this case, not only the filter coefficients but also all corresponding weights of the layers constituting the resizing network should be signaled to the decoder. According to a non-limiting example, the layers of the resizing network may include fully connected layers, convolutional layers, deconvolutional layers, pooling layers (such as max pooling, average pooling), activation layers.
[0087] Fig.15 Another general tensor rescaling method (1130) is illustrated in accordance with a general aspect of at least one embodiment. Fig.15 In the variant shown above, the input feature tensor is resized to produce For this spatial interpolation operation, there are pre-existing filters in the decoder. For the convolution operation, apply the above for Fig.14 The same process described above is used to generate a channel number equal to C l The intermediate tensor is then spatially resized by the interpolation filter corresponding to index j parsed from the bitstream to produce If Fig.13 described.
[0088] Fig.16 A general method for implementing parsing information related to tensor scaling filtering is illustrated in accordance with a general aspect of at least one embodiment.
[0089] According to a variant embodiment, a process of parsing filter information from a bitstream to adjust the size of the associated tensor is described. The process can be implemented when parsing a separate bitstream associated with each task pipeline. In a first step 1610, a flag indicating the need to adjust the size of the input tensor for the visual reasoning task is received and parsed. In response to the need to adjust the size of the input tensor being true (the flag is equal to 1), further information about the resizing filter is parsed in 1620. In response to the need to adjust the size of the input tensor at the output of the synthesis module (the flag is inferred to be 0), the method ends. When the flag is equal to 1, it may be necessary to parse the target size or resolution to be generated by the resizing module to comply with the associated visual task network. In addition, there is a flag indicating whether the target size needs to be parsed. If the flag is equal to 0, the relevant information can be inferred 1630 by referring to the configuration associated with the visual network. If the flag is equal to 1, the target dimension to be implemented by the resizing module can be parsed 1640. For example, the process parses 1640 the number of channels, width, and height of the output feature tensor by the resizing module. For example, the process parses 1640 the scaling factor for each of the tensor dimensions (channels, width, and height) between the input and output channels. Subsequently, num_filter_minus_1 is parsed 1650 to specify how many filters will be applied to resize the input feature tensor. Therefore, the actual number of filters to be applied corresponds to num_filter_minus_1+1. In one variant, the filter types can then be parsed 1660 continuously. To give a few examples, the filter types can be convolution filters, interpolation filters, etc. Depending on the filter type, a related parsing process may be involved to parse the filter coefficients and related information about the filter. After parsing the information associated with the filter (i.e., filter type and coefficients, etc.), the same parsing process is repeated to parse the remaining filter information until the number of filter sets parsed 1670 meets num_filter_minus_1+1. The parsed filter information is the same order in which the filters are applied to the input tensor in the resizing module.
[0090] Standard methods for describing, compressing and transmitting parameters of neural networks have been standardized. For example, the so-called MPEG Neural Network Representation standard (NNR) provides tools and syntax for compressing and transmitting neural networks. One can envision using NNR as a means to transmit the parameters of the proposed filters, since they usually consist of convolution operations supported by NNR. If compression is not required, for example in cases where the size of the kernel parameters is negligible, then exchange formats such as ONNX or NNEF can also be used as a syntax for specifying filter structures and their parameter values.
[0091] The syntax for machine vision processing described in this way may include, for example, additional information related to the image, a portion of the image, or the bitstream itself, which can be shared by both human and machine vision tasks. For example, the additional information may include, but is not limited to, the padding size of the input image, the input image resolution. This information is usually present in the bitstream used for input reconstruction (traditional image / video codecs), and those skilled in the art appreciate that such information is also needed for visual task bitstreams, and since the present principle is applicable to scalable bitstreams, such information should be particularly useful for visual task bitstreams. According to another variant, in the case where the encoder only encodes the region of interest, it is necessary for the decoder to know the upper left corner coordinates, width and height of the encoded region. For example, the additional information may include, but is not limited to, the upper left corner coordinates of the encoded region, the width and height of the encoded region. According to another variant, the details of the visual network configuration and network architecture should be useful, especially the layers that interface with the encoder and decoder (resize module) and the layers that are signaled as additional data.
[0092] Fig.17 Two remote devices communicating over a communication network are shown according to an example of the present principles, in which various aspects of the embodiments can be implemented. Fig.17 In the example of the present principles shown in FIG. 1 , in the context of transmission between two remote devices A and B via a communication network NET, device A comprises a processor associated with memories RAM and ROM, which are configured to implement as combined Figure 7 Any embodiment of the method for NN scalable coding described above, and device B includes a processor associated with a memory RAM and a ROM, which are configured to implement as combined Figure 7 , Figure 9-11 or Fig.16 Any embodiment of the method for NN scalable decoding for hybrid human / machine vision applications. According to one example, the network is a broadcast network suitable for broadcasting / transmitting encoded images from device A to decoding devices including device B. The signal intended for transmission by device A carries at least one scalable bitstream, which at least one scalable bitstream includes encoded data representing at least one image and metadata allowing application of resizing information. According to another embodiment, a coding method and coding device are proposed, which embed signaling information at a decoder and on a tensor resizing module implemented based on the present principles.
[0093] Fig.18An example of the syntax of such a signal is shown when at least one encoded image is transmitted via a packet-based transport protocol. Each transmitted packet P includes a header H and a payload PAYLOAD. The payload PAYLOAD can carry the above-mentioned scalable bitstream including metadata related to machine vision applications. In one variant, the payload includes scalable neural network-based encoded data representing image data samples for neural network-based visual reasoning processing and associated metadata, wherein the associated metadata includes at least one of the following: an indication of resizing a tensor of input features; an indication of whether the resizing is inferred from a configuration associated with an expected dimension for the neural network-based visual reasoning processing or the resizing is embedded from associated metadata of the bitstream; one or more parameters representing the resized dimensions of the tensor of output features; an indication of several interpolation filters used in the resizing; and one or more parameters representing one of the several interpolation filters used in the resizing.
[0094] Additional Examples and Information
[0095] Various methods are described herein, and each method includes one or more steps or actions for realizing the method.Unless the correct operation of the method requires the specific order of steps or actions, otherwise the order and / or use of specific steps and / or actions can be modified or combined.Additionally, terms such as "first", "second" etc. can be used to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding" in various embodiments.Unless specifically required, the use of such terms does not mean the sequencing of the operation to modification.Therefore, in this example, the first decoding does not need to be performed before the second decoding, and can occur in a time period overlapping with the second decoding, such as before the second decoding, during the second decoding, or with the second decoding.
[0096] Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.
[0097] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0098] Various embodiments relate to decoding. "Decoding" as used in this application may encompass, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Based on the context of the specific description, whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to a broader decoding process will be clear and is considered to be well understood by those skilled in the art.
[0099] Various embodiments relate to encoding.In a manner similar to the discussion above regarding "decoding", "encoding" as used in this application may encompass all or part of the processes performed on an input video sequence, for example, to produce an encoded bitstream.
[0100] Note that the grammatical elements as used herein are descriptive terms. As such, they do not exclude the use of other grammatical element names.
[0101] The embodiments and aspects described herein may be implemented as, for example, pieces of information that may be transmitted or stored, such as, for example, syntax. The information may be packaged or arranged in various ways, including, for example, ways common in video standards, such as placing the information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or a slice header), or SEI message. Other ways are also available, including, for example, ways common to system-level or application-level standards, such as placing the information in one or more of the following:
[0102] - SDP (Session Description Protocol), a format for describing multimedia communication sessions, used for the purpose of session announcement and session invitation, such as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transport;
[0103] - DASH MPD (Media Presentation Description) descriptors (e.g. as used in DASH and transmitted over HTTP), which are associated with a representation or a set of representations to provide additional characteristics to the content representation;
[0104] -RTP header extensions, e.g. as used during RTP streaming;
[0105] - ISO Base Media File Format, e.g. as used in OMAF and using boxes, which are object-oriented building blocks defined by a unique type identifier and a length, also called "atoms" in some specifications;
[0106] - HLS (HTTP Live Streaming) manifests transmitted over HTTP. A manifest may be associated with, for example, a version or set of versions of the content to provide characteristics of the version or set of versions.
[0107] The various embodiments and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if only discussed in the context of a single embodiment form (e.g., discussed only as a method), the embodiments of the features discussed can also be implemented in other forms (e.g., a device or program). The device can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a device, such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.
[0108] Reference to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearance of the phrase "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations appearing in various places throughout this application are not necessarily all referring to the same embodiment.
[0109] Additionally, the present application may involve "determining" various pieces of information. Determining information may include one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory, for example.
[0110] Furthermore, the present application may refer to "accessing" various pieces of information. Accessing information may include one or more of, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0111] Additionally, the present application may involve "receiving" various pieces of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of, for example, accessing information or retrieving information (e.g., from a memory). Furthermore, during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is typically involved in one way or another.
[0112] It should be appreciated that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," any of the following uses of " / ," "and / or," and "at least one of" are intended to encompass selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As further examples, in the case of "A, B, and / or C," and "at least one of A, B, and C," such wording is intended to encompass selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many items as listed, as will be apparent to one of ordinary skill in this and related arts.
[0113] In addition, as used herein, the word "signal" refers to indicating something to the corresponding decoder in addition to other matters. For example, in some embodiments, the encoder signals the quantization matrix used for dequantization. In this way, in an embodiment, the same parameters are used at both the encoder side and the decoder side. Therefore, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters and other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding the transmission of any actual function, bit saving is achieved in various embodiments. It should be appreciated that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the word "signaling" is mentioned above, the word "signal" can also be used as a noun in this article.
[0114] It will be apparent to one of ordinary skill in the art that embodiments may generate a variety of signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method, or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.
[0115] We have described several embodiments. Features of these embodiments may be provided individually or in any combination across various claim categories and types. In addition, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination across various claim categories and types:
[0116] - In a scalable NN-based decoder and / or encoder, adapt the size of feature tensors intended for machine vision tasks.
[0117] -Select filters to apply to resize feature tensors in the decoder and / or encoder.
[0118] -Signaling information related to the resizing of the feature tensor to be applied in the decoder.
[0119] - Deriving information related to the filtering process for use in resizing the feature tensor, which is applied in the decoder and / or encoder.
[0120] - Insertion of syntax elements in the signaling that enable the decoder to identify the filtering process to use, such as a filter index.
[0121] - Based on these syntax elements, at least one filtering process is selected to be applied at the decoder.
[0122] - a bitstream or signal comprising one or more of said syntax elements or their variants.
[0123] - A bitstream or signal comprising syntax conveying information generated according to any of the described embodiments.
[0124] - Insertion of syntax elements in the signaling that enable the decoder to apply the feature tensor resizing process in a manner corresponding to that used by the encoder.
[0125] - creating and / or transmitting and / or receiving and / or decoding a bitstream or signal comprising one or more of said syntax elements or variants thereof.
[0126] - created and / or transmitted and / or received and / or decoded according to any of the described embodiments.
[0127] - A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any one of the described embodiments.
[0128] -A TV, set-top box, cellular phone, tablet or other electronic device that performs a scalable NN-based decoding process suitable for adjusting the size of feature tensors intended for machine vision tasks according to any of the described embodiments.
[0129] -A TV, set-top box, cellular phone, tablet computer, or other electronic device that executes a scalable NN-based decoding process suitable for resizing feature tensors intended for machine vision tasks in accordance with any of the described embodiments and displays images intended for human vision (e.g., using a monitor, screen, or other type of display).
[0130] -A TV, set-top box, cellular phone, tablet computer, or other electronic device that selects (e.g., using a tuner) a channel to receive a signal comprising an encoded image and performs a NN-based scalable decoding process suitable for adjusting the size of a feature tensor intended for a machine vision task in accordance with any of the described embodiments.
[0131] -A TV, set-top box, cellular phone, tablet, or other electronic device that receives a signal comprising an encoded image over the air (e.g., using an antenna) and performs a NN-based scalable decoding process suitable for adjusting the size of a feature tensor intended for a machine vision task in accordance with any of the described embodiments.
Claims
1. A method comprising: Get a tensor of reconstruction data representing image data samples partially reconstructed from the base layer of the bitstream The tensor of the reconstructed data includes the number of channels (C 1 ); Tensor for reconstruction data Applying neural network-based feature synthesis processing (f s ) to generate a tensor of input features representing the characteristics of the image data samples The input feature tensor includes the number of channels of two-dimensional data (C l ); Resize the input feature tensor The size of the output feature tensor is The output feature tensor includes the number of channels of two-dimensional data; and Tensor for output features Applying neural network-based visual reasoning processing (v e ) to generate a set of inference results (T); Wherein adjusting the size of the tensor of input features includes applying at least one interpolation filter to the tensor of input features so that at least one dimension of the tensor of input features is suitable for neural network-based visual reasoning processing.
2. A device comprising a memory and one or more processors, wherein the one or more processors are configured to Get a tensor of reconstruction data representing image data samples partially reconstructed from the base layer of the bitstream The tensor of the reconstructed data includes the number of channels (C 1 ); Tensor for reconstruction data Applying neural network-based feature synthesis processing (f s ) to generate a tensor of input features representing the characteristics of the image data samples The input feature tensor includes the number of channels of two-dimensional data (C l ); Resize the input feature tensor The size of the output feature tensor is The output feature tensor includes the number of channels of two-dimensional data; and Tensor for output features Applying neural network-based visual reasoning processing (v e ) to generate a set of inference results (T); In order to adjust the size of the tensor of input features, at least one interpolation filter is applied to the tensor of input features so that at least one dimension of the tensor of input features is suitable for visual reasoning processing based on a neural network.
3. The method according to claim 1 or the apparatus according to claim 2, wherein the number of channels of the tensor of input features is equal to the number of channels of the tensor of output features, and wherein adjusting the size of the tensor of input features further comprises applying a 2D interpolation filter to each channel of the tensor of input features.
4. The method of claim 1 or the apparatus of claim 2, wherein the number of channels of the tensor of input features and the number of channels of the tensor of output features are equal, and wherein adjusting the size of the tensor of input features further comprises applying the same 2D interpolation filter to each channel of the tensor of input features.
5. The method according to any one of claims 3 and 4 or the apparatus according to any one of claims 3 and 4, further comprising obtaining parameters representing the 2D interpolation filter from metadata of the bitstream. 6 . The method according to claim 4 or the apparatus according to claim 4 , further comprising obtaining an index from metadata of a bitstream, the index indicating an interpolation filter among the interpolation filter set.
7. The method of claim 1 or the apparatus of claim 2, wherein the number of channels of the tensor of input features and the number of channels of the tensor of output features are different, and wherein adjusting the size of the tensor of input features further comprises applying at least one convolution filter to the tensor of input features to scale the number of channels.
8. The method according to claim 7 or the apparatus according to claim 7, further comprising obtaining parameters representing at least one convolution filter from metadata of the bitstream.
9. The method according to any one of claims 5, 6, and 8 or the apparatus according to any one of claims 5, 6, and 8, wherein obtaining parameters representing at least one interpolation filter from metadata of the bitstream further comprises: Parse instructions to adjust the input features of the tensor The size of the sign, A tensor that adjusts input features in response to instructions a flag for resizing, parsing a flag indicating whether the resizing is inferred from a configuration associated with the expected dimensions for the neural network-based visual reasoning processing or whether the resizing is obtained from metadata in the bitstream; and Responsive to a flag indicating that resizing is obtained from the metadata of the bitstream, parse the tensor representing the output features The method further comprises: resolving parameters of a resized dimension of the image, parsing an indication of a number of interpolation filters used in the resizing, and parsing parameters representing an interpolation filter among the number of interpolation filters used in the resizing.
10. The method according to any one of claims 1, 3-9 or the device according to any one of claims 2-9, further comprising Get a tensor of enhancement data representing image data samples partially reconstructed from the enhancement layer of the bitstream Tensor for reconstruction data and tensors of augmented data Applying neural network-based image synthesis processing (g s ) to generate the reconstructed image.
11. A method comprising: Get the input features tensor representing the features of the image data samples The input feature tensor includes the number of channels of two-dimensional data (C l );and Resize the input feature tensor The size of the output feature tensor is The output feature tensor includes a tensor suitable for neural network-based visual reasoning processing (v e )’s number of channels for two-dimensional data; The resizing of the input feature tensor includes applying at least one interpolation filter to the input feature tensor to make the dimension of the input feature tensor suitable for the neural network-based visual reasoning processing (v e ) dimension.
12. An apparatus comprising a memory and one or more processors, wherein the one or more processors are configured to Get the input features tensor representing the features of the image data samples The input feature tensor includes the number of channels of two-dimensional data (C l );and Resize the input feature tensor The size of the output feature tensor is The output feature tensor includes a tensor suitable for neural network-based visual reasoning processing (v e )’s number of channels for two-dimensional data; Wherein, in order to adjust the size of the tensor of input features, at least one interpolation filter is applied to the tensor of input features so that the dimension of the tensor of input features is suitable for the visual reasoning processing based on the neural network (v e ) dimension.
13. The method according to claim 11 or the apparatus according to claim 12, wherein the number of channels of the tensor of input features and the number of channels of the tensor of output features are equal, and wherein in order to adjust the size of the tensor of input features, a 2D interpolation filter is applied to each channel of the tensor of input features.
14. The method of claim 11 or the apparatus of claim 12, wherein the number of channels of the tensor of input features and the number of channels of the tensor of output features are equal, and wherein in order to adjust the size of the tensor of input features, the same 2D interpolation filter is applied to each channel of the tensor of input features.
15. The method of claim 11 or the apparatus of claim 12, wherein the number of channels of the tensor of input features and the number of channels of the tensor of output features are different, and wherein in order to adjust the size of the tensor of input features, at least one convolution filter is applied to the tensor of input features to scale the number of channels.
16. A non-transitory program storage device readable by a computer, tangibly embodying a program of instructions executable by the computer for performing the method according to any one of claims 1, 3-10.
17. A bitstream comprising scalable neural network-based encoded data representing image data samples for neural network-based visual reasoning processing and associated metadata, wherein the associated metadata comprises at least one of: Instructions for resizing the input feature tensor; an indication of whether the resizing is inferred from a configuration associated with an expected dimension for the neural network-based visual reasoning processing or whether the resizing is embedded from associated metadata of the bitstream; One or more parameters representing the resized dimensions of the output feature tensor; an indication of a number of interpolation filters to use in the resizing; and Represents one or more parameters of an interpolation filter among several interpolation filters used in resizing.
18. A computer readable medium comprising a bitstream according to claim 17.