A high-level syntax for priority signaling in neural network compression

By introducing an advanced syntax signaling mechanism, users can specify priorities for neural network aspects, solving the problem of suboptimal resource utilization in existing technologies and achieving more efficient neural network compression, which is suitable for devices with limited computing power, memory and power.

CN114746870BActive Publication Date: 2025-10-03NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080083224.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-02
Filing Date
2020-10-01
Publication Date
2025-10-03
Estimated Expiration
2040-10-01

AI Technical Summary

Technical Problem

When compressing neural networks, existing technologies find it difficult to dynamically adjust the compression bit rate and accuracy based on user priority requirements, resulting in suboptimal resource utilization.

Method used

By introducing an advanced syntax signaling mechanism, users can specify the priority information of different aspects of the neural network, thereby guiding the compression process to more effectively save or reduce the number of bits for unimportant aspects while retaining the accuracy of important aspects.

Benefits of technology

This enables more efficient compression of neural networks without sacrificing important aspects of accuracy, optimizing resource utilization, especially for devices with limited computing power, memory, and power resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114746870B_ABST
    Figure CN114746870B_ABST
Patent Text Reader

Abstract

Apparatus, methods, and computer programs for compressing a neural network are disclosed. An apparatus includes means for receiving information from a second device, wherein the information includes at least one parameter configured to compress a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and means for compressing the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device. The apparatus may also include means for receiving the neural network from the second device, wherein the received neural network is a compressed neural network, and wherein the received compressed neural network has been compressed by the second device prior to compressing the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary and non-limiting embodiments relate generally to computing and more particularly to neural networks. Background Art

[0002] Performing video encoding and decoding is known. Summary of the Invention

[0003] According to one aspect, an apparatus comprises: means for receiving information from a second device, wherein the information comprises at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and means for compressing the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device.

[0004] According to one aspect, an apparatus comprises: means for sending information from the apparatus to a second device, wherein the information comprises at least one parameter configured to compress a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and means for receiving a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter.

[0005] According to one aspect, an apparatus includes: circuitry configured to receive information from a second device, wherein the information includes at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and circuitry configured to compress the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device.

[0006] According to one aspect, an apparatus comprises: circuitry configured to send information from the apparatus to a second device, wherein the information comprises at least one parameter configured to be used to compress a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and circuitry configured to receive a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter.

[0007] According to one aspect, a method includes: receiving information from a second device, wherein the information includes at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and compressing the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The foregoing aspects and other features are explained in the following description taken in conjunction with the accompanying drawings, in which:

[0009] Figure 1 schematically illustrates an electronic device employing an embodiment of the examples described herein;

[0010] Figure 2 schematically illustrates a user equipment suitable for employing embodiments of the examples described herein;

[0011] Figure 3 Also schematically shown are electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections;

[0012] Figure 4 A schematic diagram showing an example of a neural network;

[0013] Figure 5 is a signaling diagram for compressing a neural network based on the example methods described herein;

[0014] Figure 6 is an example method for compressing neural networks;

[0015] Figure 7 is another example method for compressing neural networks;

[0016] Figure 8 is another signaling diagram for compressing a neural network based on the example methods described herein;

[0017] Figure 9 Two example methods for compressing neural networks are shown;

[0018] Figure 10 An image having portions identifiable by a neural network is shown;

[0019] Figure 11 An image is shown, wherein a center of the image or a bounding box is identified, and the bounding box is identified by, for example, a neural network;

[0020] Figure 12 is another signaling diagram for compressing a neural network based on the example methods described herein;

[0021] Figure 13 is another example method for compressing neural networks; and

[0022] Figure 14 is another example method for compressing neural networks. DETAILED DESCRIPTION

[0023] The following acronyms and abbreviations that may be found in the specification and / or drawings are defined as follows:

[0024] 3GPP Third Generation Partnership Project

[0025] 4G fourth generation broadband cellular network technology

[0026] 5G fifth-generation cellular network technology

[0027] 802.x series of IEEE standards dealing with local area networks and metropolitan area networks

[0028] aka also known as

[0029] ASIC

[0030] CDMA Code Division Multiple Access

[0031] DCT Discrete Cosine Transform

[0032] DSP Digital Signal Processor

[0033] FDMA Frequency Division Multiple Access

[0034] FPGA Field Programmable Gate Array

[0035] GSM Global System for Mobile Communications

[0036] H.222.0 MPEG-2 Systems, a generic coding standard for moving pictures and associated audio information

[0037] H.26x ITU-T video coding standard series

[0038] IBC Internal Block Replication

[0039] ID or id identifier

[0040] IEC International Electrotechnical Commission

[0041] IEEE Institute of Electrical and Electronics Engineers

[0042] IMD Integrated Messaging Device

[0043] IMS Instant Messaging Service

[0044] IoT

[0045] IP Internet Protocol

[0046] ISO International Organization for Standardization

[0047] ISOBMFF ISO Base Media File Format

[0048] ITU International Telecommunication Union

[0049] ITU-T ITU Telecommunication Standardization Sector

[0050] MMS Multimedia Messaging Service

[0051] MPEG Moving Picture Experts Group

[0052] MPEG-2 ITU-defined H.222 / H.262

[0053] MSE mean square error

[0054] NAL Network Abstraction Layer

[0055] net network

[0056] NN Neural Network

[0057] NNR Neural Network Representation

[0058] PC

[0059] PDA Personal Digital Assistant

[0060] PID Packet Identifier

[0061] PLC power cable connection

[0062] PSNR Peak signal-to-noise ratio

[0063] RFID radio frequency identification

[0064] SMS text messaging service

[0065] SSIM structural similarity index metric

[0066] TCP-IP Transmission Control Protocol-Internet Protocol

[0067] TDMA Time Division Multiple Access

[0068] TS transport stream

[0069] TV

[0070] UICC Universal Integrated Circuit Card

[0071] UMTS Universal Mobile Telecommunications System

[0072] USB Universal Serial Bus

[0073] WLAN Wireless Local Area Network

[0074] A neural network (NN) is a set of algorithms or computational graphs composed of several computational layers. Each layer consists of one or more units, each of which performs a computation, such as a basic calculation. A unit is connected to one or more other units, and these connections may have associated weights. Weights can be used to scale the signals passing through the associated connections. Weights can be learnable parameters, meaning values ​​that can be learned from training data. There may also be other learnable parameters, such as those of a batch normalization layer.

[0075] Figure 4 A schematic diagram of an example of a neural network 100 is shown in FIG. In this schematic diagram, the neural network 100 includes a plurality of elements 102-114. The elements may include the units mentioned above and may be assigned various different features or components or segments of the neural network 100, such as aspects or tasks of the neural network. Each element may have one or more layers, as shown by 106 in 104 and 111, 112, and 114 in 110.

[0076] The two most widely used neural network architectures are feedforward and recurrent. A feedforward neural network has no feedback loop: each layer takes input from one or more previous layers and provides its output as input to one or more subsequent layers. Furthermore, units within a layer can take input from units in one or more previous layers and provide their output to one or more subsequent layers.

[0077] The initial layers (those close to the input data) extract semantically low-level features, such as edges and textures in an image, while the intermediate and final layers extract more high-level features. Following the feature extraction layer, there may be one or more layers that perform a task such as classification, semantic segmentation, object detection, denoising, style transfer, super-resolution, etc. In recurrent neural networks, there may be feedback loops, making the network stateful, meaning it can remember information or state.

[0078] Neural networks are being used in a growing number of applications across many different types of devices, such as mobile phones. Examples include image and video analysis and processing, social media data analysis, and device usage data analysis.

[0079] One property of neural networks (and other machine learning tools) is their ability to learn properties from input data; in a supervised or unsupervised manner. This learning can be the result of a training algorithm, or of a meta-level neural network providing the training signal.

[0080] Typically, a training algorithm involves changing some property of a neural network so that its output is as close as possible to the desired output. For example, in the case of classifying objects in an image, the output of the neural network can be used to derive a class or category index that indicates the class or category to which the object in the input image belongs. Training can be performed by minimizing or reducing the error (also called loss) of the output. Examples of losses are mean squared error, cross entropy, etc. In recent deep learning techniques, training is an iterative process where, at each iteration, the algorithm modifies the weights of the neural network to gradually improve the network's output, i.e., to gradually reduce the loss.

[0081] As used herein, the terms "model," "neural network," "neural net," and "network" are used interchangeably, and the weights of a neural network are sometimes also referred to as learnable parameters or simply parameters.

[0082] Training a neural network is an optimization process, but the ultimate goal is different from the typical goal of optimization. In optimization, the only goal is to minimize a functional or function. In machine learning, the goal of the optimization or training process is to make the model learn the properties of the data distribution from a limited training data set. In other words, the goal is to learn to use a limited training data set so that the learning generalizes to previously unseen data, that is, data that was not used to train the model. This is often called generalization. In practice, the data can be divided into at least two sets, a training set and a validation set. The training set can be used to train the network, that is, to modify its learnable parameters to minimize the loss. The validation set can be used to check the performance of the network on data that was not used to minimize the loss, as an indication of the final performance of the model. Specifically, the error on the training set and the validation set can be monitored during the training process to understand the following:

[0083] If the network is learning — in this case, the training set error may decrease, otherwise the model is underfitting.

[0084] If the network is learning to generalize—in this case, the validation set error may also decrease to not be much higher than the training set error. If the training set error is low, but the validation set error is much higher than the training set error, or does not decrease, or even increases, then the model may be overfitting. This means that the model has just memorized the properties of the training set and performs well only on that set, but not on sets that were not used to tune its parameters.

[0085] Neural networks can be used to compress and decompress data such as images. An architecture used for such tasks is an autoencoder, which is a neural network consisting of two parts: a neural encoder and a neural decoder (we will refer to them simply as encoder and decoder in this article, although we are referring to algorithms that are learned from data rather than hand-tuned). The encoder can take an image as input and can produce a code that requires fewer bits than the input image. This code may have been obtained through a binarization or quantization process after the encoder. The decoder can take this code and reconstruct the image that was input to the encoder. The encoder and decoder can be trained to minimize a combination of bit rate and distortion, where the distortion can be a mean squared error (MSE), PSNR, SSIM, or similar metric. These distortion metrics can be inversely proportional to human visual perception of quality.

[0086] Neural network compression can refer to the compression of the weights of a neural network; in terms of bits, it may be the largest part required to represent a neural network. Relative to the weights, the other part, namely the architecture definition, can be considered negligible (or require very few bits to represent), especially for large neural networks (i.e., NNs with many layers and weights). It can be assumed that the input to the compression system is the original trained network; it is trained using at least the task loss. As used herein, the task loss refers to the main loss function that the network needs to minimize in order to be trained to achieve the desired output.

[0087] There may be a need to compress a neural network for various reasons, for example, to reduce the bit rate required to transmit the network over a communication channel, or to reduce storage requirements, or to reduce memory consumption at runtime, or to reduce computational complexity at runtime, etc. The performance of an algorithm for compressing a neural network may be based on a reduction in the number of bits required to represent the network and a reduction in task performance. The compression algorithm may reduce the number of bits (referred to herein as the bit rate) as much as possible while minimizing the reduction in task performance, where task performance may be the performance of the task for which the network was trained, for example, the classification accuracy for a classifier or the MSE for a network performing regression.

[0088] There are several approaches to NN compression. Some of them are based on quantization of the weights, others on pruning (removing) small values, others on low-rank decomposition of the weight matrix, and still others (generally the most successful) involve a training or retraining step. For the latter, each neural network that needs to be compressed can be retrained. This can include retraining the neural network to be compressed with a different loss relative to the task loss used to originally train the network, for example, with a combination of at least one task loss and a compression loss. The compression loss can be calculated for the weights, for example to force pruning or sparsification (i.e., to force many weights to have low values) or to force easier quantization (i.e., to force the values ​​of the weights to be close to the quantized values). With the most powerful current hardware acceleration, such retraining can take up to a week.

[0089] Suitable devices and possible mechanisms for video / image encoding processes according to example embodiments are described in more detail below. Figure 1 and 2 ,in Figure 1 An example block diagram of an apparatus 50 is shown. The apparatus may be an Internet of Things (IoT) apparatus configured to perform various functions, such as collecting information through one or more sensors, receiving or transmitting information, analyzing information collected or received by the apparatus, etc. The apparatus 50 may include a video encoding system, which may incorporate a codec. Figure 2 The layout of an apparatus 50 according to an example embodiment is shown.

[0090] The electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag or other low-power device. However, it should be understood that the embodiments may be implemented in any electronic device or apparatus that can process data through a neural network.

[0091] Device 50 may include a housing 30 for enclosing and protecting the device. Device 50 may also include a display 32, for example, in the form of a liquid crystal display. In other embodiments, the display may be any suitable display technology suitable for displaying images or video. Device 50 may also include a keypad 34. In other embodiments, any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.

[0092] The device 50 may include a microphone 36 or any suitable audio input that may be a digital or analog signal input. The device 50 may also include an audio output device, for example, which may be any of the following: an earpiece 38, a speaker, or an analog audio or digital audio output connection. The device 50 may also include a battery (or in other embodiments, the device may be powered by any suitable mobile energy device, such as a solar cell, a fuel cell, or a clockwork generator). The device 50 may also include a camera 42 capable of recording or capturing images and / or video. The device 50 may also include an infrared port for short-range line-of-sight communication with other devices. In other embodiments, the device 50 may also include any suitable short-range communication solution, such as a Bluetooth wireless connection or a USB / Firewire wired connection.

[0093] The apparatus 50 may include a controller 56, a processor, or a processor circuit (e.g., the controller 56 may be a processor) for controlling the apparatus 50. The controller 56 may be connected to a memory 58, which may store data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may also be connected to a codec circuit 54, which may be adapted to perform encoding and / or decoding of audio and / or video data or to assist in encoding and / or decoding performed by the controller 56. The apparatus 50 may also include a card reader 48 and a smart card 46, such as a UICC and a UICC reader, for providing user information and adapted to provide authentication information for authenticating and authorizing a user on a network.

[0094] The device 50 may include a radio interface circuit 52 connected to a controller 56 and adapted to generate wireless communication signals, for example, for communicating with a cellular communication network, a wireless communication system, or a wireless local area network. The device 50 may also include an antenna 44 connected to the radio interface circuit 52 for transmitting radio frequency signals generated at the radio interface circuit 52 to other devices and / or for receiving radio frequency signals from other devices.

[0095] The apparatus 50 may include a camera capable of recording or detecting individual frames, which are then passed to a codec 54 or a controller 56 for processing. The apparatus may receive video image data from another device for processing before transmitting and / or storing it. The apparatus 50 may also receive images wirelessly or via a wired connection for encoding / decoding. The structural elements of the apparatus 50 described above represent examples of modules for performing corresponding functions.

[0096] about Figure 3, shows an example of a system in which embodiments of the present invention may be utilized. System 10 includes a plurality of communication devices that may communicate over one or more networks. System 10 may include any combination of wired and / or wireless networks, including but not limited to wireless cellular telephone networks (such as GSM, UMTS, CDMA, 4G, 5G networks, etc.), wireless local area networks (WLANs) (e.g., as defined by any of the IEEE 802.x standards), Bluetooth personal area networks, Ethernet local area networks, token ring local area networks, wide area networks, and the Internet. System 10 may include wired and wireless communication devices and / or apparatuses 50 suitable for implementing example embodiments. For example, Figure 3 The illustrated system shows a representation of the mobile telephone network 11 and the Internet 28. Connectivity to the Internet 28 may include, but is not limited to, long-range wireless connections, short-range wireless connections, various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication paths.

[0097] The example communication devices shown in system 10 may include, but are not limited to, electronic device or apparatus 50, a combination personal digital assistant (PDA) and mobile phone 14, PDA 16, integrated messaging device (IMD) 18, desktop computer 20, notebook computer 22. Figure 3 As shown, PDA 16, IMD 18, desktop computer 20, and notebook computer 22 can access Internet 28 via wireless or wired link / interface 2. When device 50 is carried by a mobile individual, device 50 can be stationary or mobile. Device 50 can also be located in a mode of transportation, including but not limited to a car, truck, taxi, bus, train, ship, airplane, bicycle, motorcycle, or any similar suitable mode of transportation.

[0098] Embodiments may also be implemented in set-top boxes (i.e., digital TV receivers (which may or may not have display or wireless capabilities)), in tablet computers or (laptop) personal computers (PCs) (which have hardware and / or software to process neural network data), in various operating systems, and in chipsets, processors, DSPs and / or embedded systems that provide hardware / software based encoding.

[0099] Some or further devices may send and receive calls and messages and communicate with a service provider via a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile phone network 11 and the Internet 28. The system may include additional communication devices and various types of communication devices.

[0100] Communication devices may communicate using various transmission technologies, including but not limited to Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol-Internet Protocol (TCP-IP), Short Message Service (SMS), Multimedia Messaging Service (MMS), email, Instant Messaging Service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband Internet of Things, and any similar wireless communication technologies. Communication devices involved in implementing various embodiments of the present invention may communicate using various media, including but not limited to radio, infrared, laser, cable connections, and any suitable connection.

[0101] In telecommunications and data networks, a channel can refer to a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a wire, while a logical channel can refer to a logical connection on a multiplexed medium that is capable of carrying multiple logical channels. A channel can be used to transmit information signals (e.g., a bit stream) from one or more senders (or transmitters) to one or more receivers.

[0102] Example embodiments may also be implemented in so-called IoT devices. For example, the Internet of Things (IoT) can be defined as the interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and will enable many areas of embedded systems (such as wireless sensor networks, control systems, home / building automation, etc.) to be included in the Internet of Things (IoT). In order to utilize the Internet of Things (IoT), devices are provided with an IP address as a unique identifier. IoT devices can be equipped with a radio transmitter, such as a WLAN or Bluetooth transmitter, or an RFID tag. Alternatively, IoT devices can access an IP-based network via a wired network (such as an Ethernet-based network) or a power line connection (PLC).

[0103] The MPEG-2 transport stream (TS), specified in ISO / IEC 13818-1 or, equivalently, in ITU-T Recommendation H.222.0, is a format for carrying audio, video, and other media, as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify elementary streams within a TS (also known as a packetized elementary stream). Therefore, logical channels within an MPEG-2 TS can be considered to correspond to specific PID values. Available media file format standards include the ISO Base Media File Format (ISO / IEC 14496-12, which may be abbreviated as ISOBMFF) and the file format for NAL unit structured video (ISO / IEC 14496-15), which is derived from ISOBMFF.

[0104] A video codec consists of an encoder that transforms the input video into a compressed representation suitable for storage / transmission, and a decoder that decompresses the compressed video representation back into a viewable form. The video encoder and / or video decoder can also be separate from each other, i.e., they do not need to form a codec. Typically, the encoder discards some information from the original video sequence in order to represent the video in a more compact form (i.e., at a lower bit rate).

[0105] Some hybrid video encoders, such as many encoder implementations of ITU-T H.263 and H.264, encode video information in two stages. First, the pixel values ​​in a certain picture region (or "block") are predicted, for example, by a motion compensation module (which finds and indicates an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or a spatial module (which uses the pixel values ​​surrounding the block to be encoded in a specified manner). Second, the prediction error, that is, the difference between the predicted pixel block and the original pixel block, is encoded. This is accomplished by transforming the difference in pixel values ​​using a specified transform (such as the discrete cosine transform (DCT) or a variant thereof), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the trade-off between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bit rate).

[0106] In temporal prediction, the prediction source may be a previously decoded picture (aka reference picture). In intra block copy (IBC; aka intra block copy prediction and current picture reference), prediction may be applied similarly to temporal prediction, but the reference picture may be the current picture and only previously decoded samples may be referenced in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal prediction, but the reference picture may be a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer only to temporal prediction, while in other cases, inter prediction may be collectively referred to as temporal prediction and any one of intra block copy, inter-layer prediction, and inter-view prediction, provided that they are performed with the same or similar temporal prediction process. Inter prediction or temporal prediction may sometimes be referred to as motion compensation or motion compensated prediction.

[0107] Inter-frame prediction, also known as temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter-frame prediction, the prediction source is a previously decoded picture. Intra-frame prediction exploits the fact that adjacent pixels within the same picture may be correlated. Intra-frame prediction can be performed in the spatial domain or the transform domain, meaning that sample values ​​or transform coefficients can be predicted. Intra-frame prediction can be used in intra-frame coding without inter-frame prediction.

[0108] One result of the encoding process is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy coded more efficiently if they are first predicted from spatially or temporally adjacent parameters. For example, a motion vector can be predicted from spatially adjacent motion vectors, and only the difference relative to the motion vector predictor can be encoded. Prediction of coding parameters and intra-frame prediction can be collectively referred to as in-picture prediction.

[0109] Using features as described herein, a high-level syntax for signaling information about Neural Network Representation (NNR) standards can be included with respect to user preferences. A user can be considered an entity (human or machine) requesting compression of a certain neural network. Traditionally, neural network compression is performed using average precision along with bitrate minimization as the primary guiding metrics. Average precision as used herein generally refers to the accuracy averaged over different aspects of a neural network or the task it solves. For example, for a classifier, one aspect is the number of classes and the accuracy is averaged over all classes. For object localization (or detection), one aspect is the center of the bounding box and another is the size of the box.

[0110] Using the features described herein, a user can request neural network compression by including priority information for different aspects of the neural network in the request. In an additional or alternative example embodiment, the priority information can be sent to the user from the device performing the neural network compression.

[0111] In one example embodiment, assume that a first device (device A) is configured to perform compression of at least one neural network. Also assume that a second device (device B), the "user," is a device that needs to compress the neural network for any reason (e.g., due to limited resources in terms of computing power, memory, or power). Device A can be a physical entity, such as a server, or simply an abstract entity, such as a part of a larger device in which device B also resides. The neural network may already be on device A, or the neural network may have been sent by device B, or device A may have obtained the neural network through a third-party entity. It can also be assumed that it is impossible to achieve very high compression rates without sacrificing some accuracy in the network. This is a very general assumption that should apply to most neural networks. Exceptions may be, for example, when the expected output of the neural network can be determined even without analyzing the input (e.g., data with a very imbalanced class distribution). The syntax described below is also understandable to both parties.

[0112] Using features described herein, a user can send signaling information to device A. This signaling information can be configured to indicate which aspects of the neural network can be preserved in terms of accuracy, and optionally to what degree. As used herein, "accuracy" refers to any suitable metric measuring the quality of an aspect of a neural network. Furthermore, in some cases, there may be multiple degrees of accuracy for determining the quality of a neural network, and signaling can take into account one or more of these multiple degrees of accuracy. Compression inevitably results in a decrease in accuracy for one or more aspects. Through this signaling, during compression by device A, bits can be saved / reduced more for aspects identified as unimportant to the user in the user's signaling information. Thus, device A is configured to compress the neural network more for aspects identified as unimportant to the user by device B; compression means saving or reducing bits through compression. Device A is configured to compress the neural network less for aspects identified as important to the user by device B. The signaling information received by device A from the user (device B) can be used to determine aspects of the neural network that are important to the user, and device A can then use this information during compression of the neural network to reduce the number of bits deleted for those important aspects. This improves the accuracy of the compressed neural network for those important aspects identified in the signaling information. Using features as described herein, the distinction between content to be retained and content that can be "corrupted" (not retained) is not only binary (e.g., not necessarily binary), but can be of different classes and even further precision. Alternatively or additionally, the signaling information can be used to determine one or more aspects of the neural network that are unimportant to the user, whereupon device A can use this information during compression of the neural network to increase the number of bits deleted for those unimportant aspects (reducing the precision of the neural network for those aspects).

[0113] In the case of a classifier, one aspect could be which classes need to be retained. In the case of object detection / localization, one aspect could be the center of the bounding box, another the size of the bounding box. Other examples could be features of semantic segmentation maps, features of natural language generated from images (e.g., in image captioning), etc. Priority information could, for example, take one of the following forms (or a combination thereof):

[0114] For each aspect, multiple subsets are associated with different priorities for each subset. The priorities may be ranks.

[0115] • For each aspect, multiple subsets are associated with specific allowed degradation ranges.

[0116] For some use cases, the following are some non-limiting examples of this signaling.

[0117] Example 1: Classifying images into N classes

[0118] The user specifies a subset of classes S1 with priority 1, S2 with priority 2, and S3 with priority 3. The signaling may consist of the following dictionary: {'c1':1,'c2':3,'c3':3,'c4':1,'c5':2} – where subset S1 includes classes ‘c1’ and ‘c4’, subset S2 includes class ‘c5’, and subset S3 includes classes ‘c2’ and ‘c3’.

[0119] Alternatively, the user specifies that for subset S1, it (eg, the user) can accept no degradation in accuracy, while for subsets S2 and S3, it can accept a maximum degradation of 10% in accuracy. Example: {'c1': 0, 'c2': 10, 'c3': 10, 'c4': 0, 'c5': 10}.

[0120] When the server receives priority information from the user, it (e.g., the server) can compress the neural network that satisfies the priority information. The following is an example of an image classifier:

[0121] Regarding the first item mentioned above, the server (device A) may compress the neural network in such a way that the priority 1 class will be penalized much less than priority 2. Similarly, the server may compress the neural network in such a way that the priority 2 class will be penalized much less than priority 3. This is just an example and should not be considered limiting.

[0122] For the second item mentioned above as an alternative, the server (device A) can compress the neural network in such a way that the classes in subset S1 have 0 degradation in accuracy, while the classes in subsets S2 and S3 have a maximum degradation in accuracy of 10%. Again, this is just an example and should not be considered limiting.

[0123] Example 2: Object Detector

[0124] The center of the user-specified bounding box has a priority of 1, while the size of the box has a priority of 2.

[0125] Alternatively, the user may specify that the center may have no degradation margin, and the size may have a degradation margin of 10 pixels. Example: {'center':0,'size':10}.

[0126] Using the features as described herein, the features need not be limited to any particular algorithm used by the server to compress the neural network.

[0127] Features as described herein can be used to signal priority information to a user. In this additional example embodiment, a server can send information to a user to allow the user to prioritize one or more aspects of a neural network. The prioritization of one or more aspects can be an option made by the user. The server can send information to the user about a mapping between different compression hyperparameters and resulting priorities. This can be sent in-band or out-of-band (relative to a compressible model). For example, the server can first process the neural network to make the neural network more compressible (e.g., more robust to sparsification), and then send the neural network to the user along with the associated mapping. The mapping can associate different sparsification thresholds with different prioritizations. Here is an example:

[0128] {0.05:{'c1':1,'c2':3,'c3':3,'c4':1,'c5':2},0.1:{'c1':2,'c2':3,'c3':3,'c4':1,'c5':2}}.

[0129] The user can then choose a sparsification threshold based on which classes the user considers to be more important. For example, if class 'c1' is very important to the user, 0.05 can be used for the threshold, thereby sparsifying the weights.

[0130] This example embodiment may be useful when the size of the neural network input to the user's device is not an issue (e.g., when channel bandwidth or memory are not an issue), but rather the main issue is inference-stage resources (e.g., memory, computational power, and power at inference time, which may even vary and thus be dynamic (e.g., resource availability may vary over time due to many processes running on the user's device)). In these cases, the user can decide how much it wants to compress (or further compress) the neural network. The input neural network may be a more compressible version of the received neural network, e.g., trained or fine-tuned using a compression loss on the weights (which makes them more robust to compression), and / or already compressed to a certain extent (so that the user can compress it further).

[0131] This embodiment (signaling priority information to users) consists of Figure 12 The signaling diagram is shown in Figure 12, at 704, device A (e.g., encoder) sends / signals the neural network and priority information to device B (e.g., decoder). In some examples, device A sends / signals the priority information without sending the neural network. As mentioned, the priority information signaled at 704 can be a mapping between different compression hyperparameters and result priorities, and / or the priority information signaled at 704 can be a mapping associating different sparsification thresholds with different priority rankings, or information related to unification or decomposition. The information signaled at 704 can also be information similar to the signaling information provided in other embodiments, such as similar to Figure 5 The signaling information provided at 200 or at Figure 8 200' and as described throughout this document. At 706, device B acts as an encoder and compresses or further compresses the neural network based on the received signaling information provided at 704. It is optional for device B to further compress the neural network in the sense that device A compressed the neural network at 702, where the compression at 702 is performed, for example, before the neural network and / or signaling information is sent to device B at 704. The compression at 702 is optional, as shown by the dashed line. It is also optional for device B to request the neural network and / or signaling information from device A at 700 (the optionality is indicated by the dashed line).

[0132] It should be noted that Figure 12 An embodiment is shown in which priority information is signaled from an encoder or network (eg, device A) to a user (eg, device B), however this embodiment has already been reflected in, for example, Figure 5 For example, in Figure 5 In the example, device B may be a server or encoder rather than a user device, and at 200, device B sends signaling information and / or a neural network (initially compressed or uncompressed) to device A, where device A is a user device or decoder rather than a server device or encoder. Then at 206, device A acts as an encoder and compresses or further compresses the neural network. Figure 5 At 208, device A may send an acknowledgment of receiving the information at 200 or even the compressed or further compressed neural network to device B.

[0133] Similarly in Figure 8 In , device B may be a serving device or encoder rather than a user device, wherein device B sends a neural network (initial compressed or uncompressed) and / or priority / signaling information to device A at 200', wherein device A is a user device or decoder. Figure 8 In , at 206, the user device or decoder also acts as an encoder, compressing or further compressing the neural network. Figure 8At 300 in FIG. 1 , for example, before device B sends signaling information to device A at 200 ′, user equipment or decoder device A sends a request for signaling information to device B. Figure 8 At 208, device A may send to device B an acknowledgment of receipt of the information at 200' or even the compressed or further compressed neural network.

[0134] Figure 13 It is used based on Figure 12 Another example method for compressing a neural network is provided by a signaling diagram shown in FIG. The method optionally includes, at 802, compressing the neural network by a first device. The first device can be, for example, an encoder. At 804, the method includes sending, by the first device, a neural network (e.g., uncompressed) or a compressed neural network and information to a second device, wherein the information includes at least one parameter configured to compress or further compress the neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network. The second device can be, for example, a decoder. At 806, the method includes compressing or further compressing the neural network by the second device, wherein the neural network is compressed or further compressed based at least in part on the at least one parameter received from the first device.

[0135] In the Moving Picture Experts Group (MPEG) Neural Network Representation (NNR), a high-level syntax is needed. One of the aspects that a high-level syntax may support is the preference of a user (who requests compression) for certain aspects of the neural network (NN) or certain aspects of the task that the neural network (NN) solves.

[0136] Also refer to Figure 5-7 , an example method will be further described. In this example, device A is configured to compress a neural network. Figure 5 As shown in 200 in FIG, device B may send a request for a neural network to device A. Figure 6 As shown in 202, the request may include signaling information from device B to device A, wherein the information includes at least one parameter configured to compress the neural network, wherein the at least one parameter identifies a first aspect or task of the neural network, such as Figure 4 108 of . As used herein, a "task" of a neural network may sometimes be referred to as just an "aspect" of the neural network. Figure 7 204 and Figure 5 As shown in 200 in FIG. 1 , device A may receive signaling information from device B, where the signaling information includes at least one parameter configured to be used for compressing a neural network. Figure 7 206 and Figure 5As shown in 206, device A may then compress the neural network, wherein during the compression, a first aspect of the neural network has less loss (relative to at least one other aspect of the neural network) based at least in part on at least one parameter received from device B. Figure 5 and 6 As shown in 208, device A may then send and device B may receive a compressed neural network, wherein the compressed neural network includes a first aspect or task (e.g., Figure 4 104), and at least one second aspect or task having more losses than the first aspect 108 (e.g. Figure 4 110 in the ).

[0137] Also refer to Figure 8 , shows another example, where before device B sends signaling information, device A sends a request 300 to device B to send signaling information. Then, device B can send a reply request 200' with the signaling information to device A. Also refer to Figure 9 , two example methods are shown, wherein a first request for a neural network is sent by device B and received by device A, as shown in block 400. In one example method, as shown in block 402, device A sends a reply request to device B, wherein the reply request is configured to request device B to send information to a first device, e.g., device A. In another example method, as shown in block 404, device A sends a compressed neural network, the compressed neural network, the reply request, and a mapping associating the compressed neural network with at least one parameter to device B. As shown in block 406, device B sends, and device A receives, a response to the reply request, wherein the response includes a value for the at least one parameter. As shown in block 408, device A may then compress 206 the neural network based at least in part on at least the at least one parameter received from device B (e.g., a second device), wherein a first aspect of the neural network has a lower loss than at least one other aspect of the neural network.

[0138] As mentioned above, features as described herein may be applied to images. For example, Figure 10An image 500 is shown having parts 502, 504, 506. The neural network may be configured to identify different types of parts, such as a person 502, a house 504, and a dog 506. Some parts, such as living subjects 502 and 506, may be grouped together as subsets in some aspects of the neural network. 502 and 506 may be given a first classification, 504 may be given a second, different classification, and each classification may be given a different priority with respect to loss when the neural network is compressed. Device B may be able to specify to device A that aspects or tasks of the neural network associated with items in the image (e.g., associated with person 502) should have no loss or should have a loss or degradation of no less than a predetermined value. See also Figure 11 , an image 600 is shown in which the center 602 of the image is identified and a bounding box 604 or 606 is identified. Signaling information sent by device B to device A may include values ​​for parameters related to center 602 and / or bounding box 604 and / or bounding box 606. For example, the signaling information may specify that the area around center 602 can only be degraded to a limit of 20%, while the area around bounding box 604 or 606 can be degraded to a limit of 50%. These are merely examples to help understand the features described herein and should not be considered limiting.

[0139] Compression is applied to the neural network rather than to a specific aspect of the neural network. One aspect or task of the neural network, for example, is the size of the bounding box 604 or 606 (e.g., in the case of an object detection neural network). Using the features described herein, compression of the neural network can be accomplished such that the size of the bounding box 604 or 606 has a greater decrease in accuracy relative to the center 602 of the image or the bounding box, where the decrease in accuracy at 604 or 606 and 602 is caused by compression of the neural network.

[0140] Figure 14

[0026] Another example method for compressing a neural network based on the examples described herein is provided. At 202, the method includes sending information from a first device to a second device, wherein the information includes at least one parameter configured to compress a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network. At 207, the method includes receiving, by the first device, a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter.

[0141] References to "computers," "processors," and in some examples, "controllers" should be understood to encompass not only computers with different architectures (e.g., single / multi-processor architectures and sequential (von Neumann) / parallel architectures), but also specialized circuits (e.g., field programmable gate arrays (FPGAs), application specific circuits (ASICs), signal processing devices, and other processing circuits). References to computer programs, instructions, code, etc. should be understood to encompass software or firmware for a programmable processor (e.g., the programmable content of a hardware device (e.g., instructions for a processor)), or configuration settings for a fixed-function device, gate array, or programmable logic device, etc.

[0142] Memory 58 may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. Memory 58 may include a database for storing data.

[0143] As used in this application, the term "circuitry" may refer to one or more or all of the following:

[0144] (a) pure hardware circuit implementation (e.g., implementation using analog and / or digital circuits) and

[0145] (b) a combination of hardware circuitry and software (and / or firmware), such as (as applicable):

[0146] (i) a combination of processors, and

[0147] (ii) the processor / software portion (including a digital signal processor), software and memory that work together to enable the device to perform its various functions, and

[0148] (c) Circuitry (such as a microprocessor or portion of a microprocessor) that requires software (such as firmware) for operation, even if the software or firmware is not physically present.

[0149] As another example, as used herein, the term "circuitry" would also cover an implementation of merely a processor (or multiple processors) or portion of a processor and its (or their) accompanying software and / or firmware. For example, the term "circuitry" would also cover a baseband integrated circuit or an applications processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, cellular network device, or other network equipment, if applicable to the particular element.

[0150] An example method may be provided that includes: receiving, by a first device, information from a second device, wherein the information includes at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and compressing, by the first device, the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device.

[0151] Other aspects of the method may include the following. The at least one first aspect or task may include a separate aspect or task of the neural network. Compression of the neural network may result in the accuracy of at least one second aspect or task of the neural network being lower than the accuracy of the at least one first aspect or task. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information for preventing any reduction in accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, and wherein the at least one first aspect or task may include one of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting associated with the at least one first aspect or task. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., the center of the image or the center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The method may also include transmitting the compressed neural network compressed by the first device to the second device. The method may also include: sending, by the first device to the second device, a request configured to request the second device to send information to the first device. The request may identify a plurality of priorities for at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The first device may send the compressed neural network to the second device, along with the request and the mapping associating the compressed neural network with the at least one parameter.

[0152] Other aspects of the method may include the following. The method may also include receiving, by the first device, a neural network from the second device. The neural network received by the first device may be a compressed neural network. The received compressed neural network may have been compressed by the second device before the neural network was compressed by the first device. The information may include a sparsification performance map specifying a mapping between at least one sparsification threshold and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network (including at least one first aspect or task of the neural network). Each of the at least one sparsification threshold may be mapped to a separate accuracy of the at least one accuracy for each of at least one class. Each of the at least one sparsification threshold may be mapped to an overall accuracy considering each of the at least one class. Each of the at least one class predicted by the neural network may be ordered based on an order of output of the neural network or an order specified during training of the neural network. The information may include a unified performance map specifying a mapping between at least one unified threshold and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network (including at least one first aspect or task of the neural network). Each of the at least one unified threshold may be mapped to a separate accuracy in the at least one accuracy for each of the at least one class. Each of the at least one unified threshold may be mapped to an overall accuracy that considers each of the at least one class. Each of the at least one class predicted using the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The information may include a decomposition performance map that specifies a mapping between at least one MSE threshold between at least one decompressed tensor and at least one original tensor and at least one accuracy of the neural network. At least one accuracy may be provided separately for different aspects of the output of the neural network, including at least one first aspect or task of the neural network. Each of the at least one MSE threshold may be mapped to a separate accuracy in the at least one accuracy for each of the at least one class. Each of the at least one MSE threshold may be mapped to an overall accuracy that considers each of the at least one class. Each of the at least one class predicted using the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The first device may be an encoder, and the second device may be a decoder. The first device may be a decoder, and the second device may be an encoder.

[0153] An example embodiment may be provided in an apparatus comprising: at least one processor; and at least one non-transitory memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to: cause receipt of information from a second device, wherein the information comprises at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and compress the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device.

[0154] Other aspects of the apparatus may include the following. At least one first aspect or task may include a separate aspect or task of the neural network. Compression of the neural network may result in the accuracy of at least one second aspect or task of the neural network being lower than the accuracy of the at least one first aspect or task. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information for preventing any reduction in accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, wherein the at least one first aspect or task may include a subset of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting associated with the at least one first aspect or task. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., the center of the image or the center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The at least one memory and the computer program code may also be configured, with the at least one processor, to cause the apparatus to send the compressed neural network to the second device. The at least one memory and the computer program code may also be configured, with the at least one processor, to cause the apparatus to send a request to the second device, wherein the request is configured to request the second device to send the information to the apparatus. The request may identify a plurality of priorities for the at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The apparatus may send the compressed neural network to the second device, along with the request and the mapping associating the compressed neural network with the at least one parameter.

[0155] Other aspects of the apparatus may include the following. The at least one memory and the computer program code may be further configured, together with the at least one processor, to cause the apparatus to receive the neural network from the second device. The received neural network may be a compressed neural network. The received compressed neural network may have been compressed by the second device before compressing the neural network. The information may include a sparsification performance map specifying a mapping between at least one sparsification threshold and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network (including the at least one first aspect or task of the neural network). Each of the at least one sparsification threshold may be mapped to a separate accuracy of the at least one accuracy for each of the at least one class. Each of the at least one sparsification threshold may be mapped to an overall accuracy that considers each of the at least one class. Each of the at least one class predicted using the neural network may be ranked based on an order of outputs of the neural network or an order specified during training of the neural network. The information may include a unified performance map specifying a mapping between at least one unified threshold and at least one accuracy of the neural network. At least one accuracy measure may be provided separately for different aspects of the output of the neural network (including the at least one first aspect or task of the neural network). Each of the at least one uniform threshold may be mapped to a separate accuracy measure from the at least one accuracy measure for each of the at least one class. Each of the at least one uniform threshold may be mapped to an overall accuracy measure that considers each of the at least one class. Each of the at least one class predicted using the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The information may include a decomposition performance map that specifies a mapping between at least one mean square error (MSE) threshold value between at least one decompressed tensor and at least one original tensor and at least one accuracy measure of the neural network. The at least one accuracy measure may be provided separately for different aspects of the output of the neural network (including the at least one first aspect or task of the neural network). Each of the at least one MSE threshold value may be mapped to a separate accuracy measure from the at least one accuracy measure for each of the at least one class. Each of the at least one MSE threshold value may be mapped to an overall accuracy measure that considers each of the at least one class. Each of the at least one class predicted using the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The apparatus may be an encoder and the second device may be a decoder. The apparatus may be a decoder and the second device may be an encoder.

[0156] Example embodiments may provide a non-transitory program storage device readable by a machine, tangibly embodying a program of instructions executable by the machine for performing operations comprising: receiving, by a first device, information from a second device, wherein the information comprises at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and compressing, by the first device, the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device.

[0157] Other aspects of the non-transitory program storage device may include the following. The at least one first aspect or task may include a separate aspect or task of a neural network. Compression of the neural network may result in the accuracy of at least one second aspect or task of the neural network being lower than the accuracy of the at least one first aspect or task. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information to prevent any reduction in accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, and wherein the at least one first aspect or task may include one of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting associated with the at least one first aspect or task. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., the center of the image or the center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The operation may also include transmitting the compressed neural network compressed by the first device to the second device. The operations may also include sending, by the first device, a request to the second device, wherein the request is configured to request the second device to send the information to the first device. The request may identify a plurality of priorities for at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The first device may send the compressed neural network to the second device, along with the request and the mapping associating the compressed neural network with the at least one parameter.

[0158] Other aspects of the non-transitory program storage device may include the following. The operations may also include receiving, by the first device, a neural network from the second device. The neural network received by the first device may be a compressed neural network. The received compressed neural network may have been compressed by the second device before the neural network was compressed by the first device. The information may include a sparsification performance map specifying a mapping between at least one sparsification threshold and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network (including at least one first aspect or task of the neural network). Each of the at least one sparsification threshold may be mapped to a separate accuracy of the at least one accuracy for each of the at least one class. Each of the at least one sparsification threshold may be mapped to an overall accuracy that considers each of the at least one class. Each of the at least one class predicted by the neural network may be ranked based on an order of outputs of the neural network or an order specified during training of the neural network. The information may include a unified performance map specifying a mapping between at least one unified threshold and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network (including at least one first aspect or task of the neural network). Each of the at least one uniform threshold may be mapped to a separate accuracy in the at least one accuracy for each of the at least one class. Each of the at least one uniform threshold may be mapped to an overall accuracy that considers each of the at least one class. Each of the at least one class predicted by the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The information may include a decomposition performance map that specifies a mapping between at least one MSE threshold between at least one decompressed tensor and at least one original tensor and at least one accuracy of the neural network. The at least one accuracy may be provided for different aspects of the output of the neural network, including at least one first aspect or task of the neural network. Each of the at least one MSE threshold may be mapped to a separate accuracy in the at least one accuracy for each of the at least one class. Each of the at least one MSE threshold may be mapped to an overall accuracy that considers each of the at least one class. Each of the at least one class predicted by the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The first device may be an encoder, and the second device may be a decoder. The first device may be a decoder, and the second device may be an encoder.

[0159] Example embodiments may provide an apparatus comprising: means for receiving information from a second device, wherein the information comprises at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and means for compressing the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device.

[0160] Other aspects of the apparatus may include the following. The at least one first aspect or task may include a separate aspect or task of the neural network. Compression of the neural network may result in the accuracy of at least one second aspect or task of the neural network being lower than the accuracy of the at least one first aspect or task. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information for preventing any reduction in accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, wherein the at least one first aspect or task may include one of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting associated with the at least one first aspect or task. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., the center of the image or the center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The apparatus may also include means for transmitting the compressed neural network to the second device. The apparatus may further include means for sending a request to the second device, wherein the request is configured to request the second device to send information to the apparatus. The request may identify a plurality of priorities for at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The apparatus may send the compressed neural network to the second device, along with the request and the mapping associating the compressed neural network with the at least one parameter.

[0161] Other aspects of the apparatus may include the following. The apparatus may further include means for receiving a neural network from the second device. The received neural network may be a compressed neural network. The received compressed neural network may have been compressed by the second device prior to compression of the neural network. The information may include a sparsification performance map specifying a mapping between at least one sparsification threshold and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network (including the at least one first aspect or task of the neural network). Each of the at least one sparsification threshold may be mapped to a separate accuracy of the at least one accuracy for each of the at least one class. Each of the at least one sparsification threshold may be mapped to an overall accuracy that considers each of the at least one class. Each of the at least one class predicted by the neural network may be ranked based on an order of the output of the neural network or an order specified during training of the neural network. The information may include a unified performance map specifying a mapping between at least one unified threshold and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network (including the at least one first aspect or task of the neural network). Each of the at least one uniform threshold may be mapped to a separate accuracy in the at least one accuracy for each of the at least one class. Each of the at least one uniform threshold may be mapped to an overall accuracy that takes into account each of the at least one class. Each of the at least one class predicted by the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The information may include a decomposition performance map that specifies a mapping between at least one MSE threshold between at least one decompressed tensor and at least one original tensor and at least one accuracy of the neural network. The at least one accuracy may be provided separately for different aspects of the output of the neural network, including at least one first aspect or task of the neural network. Each of the at least one MSE threshold may be mapped to a separate accuracy in the at least one accuracy for each of the at least one class. Each of the at least one MSE threshold may be mapped to an overall accuracy that takes into account each of the at least one class. Each of the at least one class predicted by the neural network may be ranked based on an output order of the neural network or an order specified during training of the neural network. The apparatus may be an encoder, and the second device may be a decoder. The apparatus may be a decoder, and the second device may be an encoder.

[0162] An example method may be provided, comprising: sending information from a first device to a second device, wherein the information includes at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and receiving, by the first device, a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter.

[0163] The method may also include applying the compressed neural network to the image by the first device. The at least one first aspect or task may include a separate aspect or task of the neural network. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information for preventing a reduction in accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, wherein the at least one first aspect or task may include one of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., the center of the image or the center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The method may also include receiving, by the first device, a request from the second device, wherein the request is configured to request the first device to send information to the second device. The request may identify a plurality of priorities for the at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The second device may send a first different compressed neural network to the first device, along with the request and the mapping associating the first different compressed neural network with the at least one parameter. The method may also include wherein at least one second aspect or task in the compressed neural network has a greater reduction in accuracy than the at least one first aspect or task.

[0164] Example embodiments may provide an apparatus comprising: at least one processor; and at least one non-transitory memory comprising computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to: cause information to be sent from the apparatus to a second device, wherein the information comprises at least one parameter configured for compressing a neural network, wherein the at least one parameter relates to at least one first aspect or task of the neural network; cause receipt of a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter.

[0165] The at least one memory and the computer program code may also be configured, with the at least one processor, to cause the apparatus to use the compressed neural network for an image. The at least one first aspect or task may include a separate aspect or task of the neural network. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information for preventing a reduction in accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, wherein the at least one first aspect or task may include one of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., the center of the image or the center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The at least one memory and the computer program code may also be configured to, with the at least one processor, cause the apparatus to receive a request from the second device, wherein the request is configured to request the apparatus to send the information to the second device. The request may identify a plurality of priorities for the at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The second device may send a first different compressed neural network to the apparatus along with the request and the mapping associating the first different compressed neural network with the at least one parameter. The apparatus may also include wherein at least one second aspect or task in the compressed neural network has a greater reduction in accuracy than the at least one first aspect or task.

[0166] Example embodiments may be provided with a non-transitory program storage device readable by a machine, tangibly embodying a program of instructions executable by the machine for performing operations comprising: sending information from a first device to a second device, wherein the information comprises at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and receiving, by the first device, from the second device, a compressed neural network, wherein the compressed neural network has been compressed based on the at least one parameter.

[0167] The operations may also include applying, by the first device, the compressed neural network to the image. The at least one first aspect or task may include a separate aspect or task of the neural network. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information for preventing a reduction in accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, wherein the at least one first aspect or task may include one of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., a center of the image or a center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The operations may also include receiving, by the first device, a request from the second device, wherein the request is configured to request the first device to send information to the second device. The request may identify a plurality of priorities for the at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The second device may send a first different compressed neural network to the first device, along with the request and the mapping associating the first different compressed neural network with at least one parameter. The non-transitory program storage device may also include, wherein at least one second aspect or task of the compressed neural network has a greater reduction in accuracy than the at least one first aspect or task.

[0168] Example embodiments may provide an apparatus comprising: means for sending information from the apparatus to the second device, wherein the information comprises at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and means for receiving a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter.

[0169] The apparatus may further include means for applying the compressed neural network to an image. The at least one first aspect or task may include a separate aspect or task of the neural network. The information may include an identification of the at least one first aspect or task. The information may include an identification of an image classification. The at least one parameter may include a priority value. The information may include an identification of at least one portion of the image. The at least one parameter may include information for preventing degradation of accuracy of the at least one first aspect or task. The information may include an identification of a plurality of class subsets, wherein the at least one first aspect or task may include one of the plurality of subsets. The at least one parameter may include a percentage or value less than one. The at least one parameter may include a compression value or setting. The information may include an image location on the image, and the at least one parameter may include a compression setting for the image location. The image location may include at least one of a center (e.g., the center of the image or the center of a bounding box) or a bounding box. The at least one parameter may include a pixel value. The at least one parameter may include a degradation value or degradation range. The apparatus may further include means for receiving a request from a second device, wherein the request is configured to request the apparatus to send information to the second device. The request may identify a plurality of priorities for the at least one parameter. The request may identify different aspects or tasks of the neural network, including the at least one first aspect or task. The request may include a mapping. The second device may send a first different compressed neural network to the apparatus along with the request and the mapping associating the first different compressed neural network with at least one parameter. The apparatus may also include wherein at least one second aspect or task in the compressed neural network has a greater reduction in accuracy than the at least one first aspect or task.

[0170] An example apparatus may include circuitry configured to receive information from a second device, wherein the information includes at least one parameter configured for compressing a neural network, wherein the at least one parameter relates to a first aspect or task of at least one neural network; and circuitry configured to compress the neural network, wherein the neural network is compressed based at least in part on the at least one parameter received from the second device.

[0171] An example apparatus may include circuitry configured to send information from the apparatus to the second device, wherein the information includes at least one parameter configured to compress a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network; and circuitry configured to receive a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter. The apparatus may also include circuitry configured to receive a compressed neural network from the second device, wherein the compressed neural network has been compressed based on the at least one parameter. The apparatus may also include circuitry configured to receive a compressed neural network from the second device, wherein the compressed neural network has a greater reduction in accuracy than the at least one first aspect or task in the compressed neural network.

[0172] It should be understood that the foregoing description is illustrative only. Those skilled in the art may devise various alternatives and modifications. For example, the features recited in the various dependent claims may be combined with one another in any suitable combination. Furthermore, features from the different embodiments described above may be selectively combined to form new embodiments. Therefore, this description is intended to encompass all such alternatives, modifications, and variations that fall within the scope of the appended claims.

Claims

1. A device for neural network compression, comprising: at least one processor; as well as at least one non-transitory memory comprising computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to at least: receiving information from a second device, wherein the information includes at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network, wherein the information further comprises at least one of: a sparsification performance map, a unified performance map, or a decomposition performance map, wherein the sparsification performance map specifies a mapping between at least one sparsification threshold and at least one accuracy of the neural network, the unified performance map specifies a mapping between at least one unified threshold and at least one accuracy of the neural network, and the decomposition performance map specifies a mapping between at least one MSE threshold between at least one decompressed tensor and at least one original tensor and at least one accuracy of the neural network, wherein the at least one accuracy measure is provided separately for different aspects of the output of the neural network, the different aspects of the output of the neural network comprising the at least one first aspect or task of the neural network; and The neural network is compressed based at least in part on the information received from the second device and the at least one parameter.

2. The device according to claim 1, wherein Each of the at least one sparsification threshold is mapped to a separate one of the at least one precision for each of the at least one class.

3. The device according to claim 2, wherein Each of the at least one sparsification threshold is mapped to an overall accuracy considering each of the at least one class.

4. The device according to claim 1, wherein Each of the at least one unified threshold is mapped to a separate one of the at least one precision for each of the at least one class.

5. The device according to claim 4, wherein Each of the at least one uniform threshold is mapped to an overall accuracy considering each of the at least one class.

6. The device according to claim 1, wherein Each of the at least one MSE threshold is mapped to a separate one of the at least one precision for each of the at least one class.

7. The device according to claim 6, wherein Each of the at least one MSE threshold is mapped to an overall accuracy considering each of the at least one class.

8. The device according to any one of claims 2 to 7, wherein Each of the at least one class predicted using the neural network is ordered based on an output order of the neural network or an order specified during training of the neural network.

9. A method for neural network compression, comprising: receiving information from a second device, wherein the information includes at least one parameter configured for compressing a neural network, wherein the at least one parameter is related to at least one first aspect or task of the neural network, wherein the information further comprises at least one of: a sparsification performance map, a unified performance map, or a decomposition performance map, wherein the sparsification performance map specifies a mapping between at least one sparsification threshold and at least one accuracy of the neural network, the unified performance map specifies a mapping between at least one unified threshold and at least one accuracy of the neural network, and the decomposition performance map specifies a mapping between at least one MSE threshold between at least one decompressed tensor and at least one original tensor and at least one accuracy of the neural network, wherein the at least one accuracy measure is provided separately for different aspects of the output of the neural network, the different aspects of the output of the neural network comprising the at least one first aspect or task of the neural network; and The neural network is compressed based at least in part on the information received from the second device and the at least one parameter.

Citation Information

Patent Citations

  • Data processing method, end device, cloud device, and end-cloud collaboration system

    WO2018121282A1