High-level syntax for compressed representations of neural networks

By encoding and decoding neural networks using advanced bitstream syntax, the shortcomings of compression and updates in neural network exchange formats are resolved, enabling efficient transmission and detailed characteristic description of neural network data transmission.

CN115210715BActive Publication Date: 2026-06-02NOKIA TECHNOLOGIES OY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2020-12-31
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing neural network exchange formats lack high-level syntax for compression, cannot effectively transmit and update entire or partial neural networks, and lack information describing the characteristics of neural networks.

Method used

The neural network is encoded or decoded using a high-level bitstream syntax, and the neural network data is transmitted through a serialized bitstream. Metadata is included in the information unit to enable compression and partial updates.

Benefits of technology

It enables efficient transmission and partial updates of neural network data, provides detailed information on neural network characteristics, and supports progressive compression and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115210715B_ABST
    Figure CN115210715B_ABST
Patent Text Reader

Abstract

An apparatus comprising: means for encoding or decoding a high-level bitstream syntax for at least one neural network; wherein the high-level bitstream syntax comprises at least one information unit having metadata or compressed neural network data of a portion of the at least one neural network; and wherein a serialized bitstream comprises one or more of the at least one information unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary and non-limiting embodiments generally relate to multimedia transmission and neural networks, and more specifically, to a high-level syntax for compressed representations of neural networks. Background Technology

[0002] A standardized format is known to be provided for the exchange of neural networks. Summary of the Invention

[0003] According to one aspect, an apparatus includes means for encoding or decoding an advanced bitstream syntax for at least one neural network; wherein the advanced bitstream syntax includes at least one information unit having metadata or compressed neural network data as a portion of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit.

[0004] According to one aspect, an apparatus includes at least one processor; and at least one non-transitory memory including computer program code; wherein the at least one memory and the computer program code are configured, together with the at least one processor, to cause the apparatus to at least: encode or decode a high-level bitstream syntax for at least one neural network; wherein the high-level bitstream syntax includes at least one information unit having metadata or compressed neural network data as a portion of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit.

[0005] According to one aspect, a method includes encoding or decoding a high-level bitstream syntax for at least one neural network; wherein the high-level bitstream syntax includes at least one information unit having metadata or compressed neural network data as a portion of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit.

[0006] According to one aspect, a machine-readable non-transitory program storage device is provided, which tangibly embodies a machine-executable program of instructions for performing operations, the operations including: encoding or decoding a high-level bitstream syntax for at least one neural network; wherein the high-level bitstream syntax includes at least one information unit having metadata or compressed neural network data of a portion of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit. Attached Figure Description

[0007] The foregoing aspects and other features are explained in the following description in conjunction with the accompanying drawings, wherein:

[0008] Figure 1 An electronic device employing an embodiment of the examples described herein is illustrated schematically;

[0009] Figure 2 A user equipment suitable for employing the examples described herein is illustrated schematically;

[0010] Figure 3 Electronic devices employing embodiments of the examples described herein are also schematically illustrated, these electronic devices being connected using wireless and wired network connections;

[0011] Figure 4 A block diagram of a general-level encoder is shown schematically.

[0012] Figure 5 This is a block diagram illustrating the interface between the encoder and decoder according to the example described herein;

[0013] Figure 6 An example structure of a compressed neural network (NNR) bitstream is shown;

[0014] Figure 7 This is an example diagram illustrating how an NNR bitstream can be composed of several NNR units of different types;

[0015] Figure 8 An example topology description of AlexNet is shown, which uses the Neural Network Exchange Format (NNEF) topology graph format;

[0016] Figure 9 It is an example apparatus configured to implement a high-level syntax for compressed representations of neural networks;

[0017] Figure 10 It is an example method for implementing a high-level syntax for compressed representations of neural networks; and

[0018] Figure 11 This is a block diagram of one possible, non-limiting system in which exemplary embodiments can be practiced. Detailed Implementation

[0019] The following acronyms and abbreviations, which can be found in the instruction manual and / or accompanying drawings, are defined as follows:

[0020] 3GP 3GPP file format

[0021] 3GPP Third Generation Partnership Project

[0022] 3GPP TS 3GPP Technical Specification

[0023] 4CC Four-letter code

[0024] 4G fourth-generation broadband cellular network technology

[0025] 5G fifth-generation cellular network technology

[0026] 5GC 5G Core Network

[0027] ACC accuracy

[0028] AI (Artificial Intelligence)

[0029] AIoT (Artificial Intelligence of Things)

[0030] aka is also known as

[0031] AMF Access and Mobility Management Functions

[0032] AVC (Advanced Video Coding)

[0033] CDMA Code Division Multiple Access

[0034] CE Core Experiment

[0035] CU Central Unit

[0036] DASH is an HTTP-based dynamic adaptive streaming service.

[0037] Discrete Cosine Transform (DCT)

[0038] DSP Digital Signal Processor

[0039] DU Distributed Unit

[0040] The evolved Node B (e.g., LTE base station) from eNB (or eNodeB)

[0041] EN-DC E-UTRA-NR Dual Connection

[0042] The en-gNB or En-gNB provides the UE with NR user plane and control plane protocol termination and acts as an auxiliary node in the EN-DC.

[0043] E-UTRA evolved Universal Terrestrial Radio Access, namely LTE radio access technology

[0044] FDMA (Frequency Division Multiple Access)

[0045] f(n) is a fixed-pattern bit string of n bits written in a left-to-right manner.

[0046] Interface between F1 or F1-C CU and DU control interfaces

[0047] gNB (or gNodeB) is a base station used for 5G / NR, that is, a node that provides NR user plane and control plane protocol termination to the UE and connects to the 5GC through the NG interface.

[0048] GSM Global Mobile Communication System

[0049] The official names of the H.222.0 MPEG-2 system are ISO / IEC 13818-1 and ITU-T Rec.H.222.0.

[0050] The H.26x ITU-T domain video coding standards family

[0051] HLS Advanced Syntax

[0052] IBC Internal Block Copy

[0053] ID identifier

[0054] IEC International Electrotechnical Commission

[0055] IEEE Institute of Electrical and Electronics Engineers

[0056] I / F interface

[0057] IMD Integrated Messaging Device

[0058] IMS Instant Messaging Service

[0059] I / O Input / Output

[0060] IoT (Internet of Things)

[0061] IP Internet Protocol

[0062] ISO International Organization for Standardization

[0063] ISOBMFF ISO Basic Media File Format

[0064] ITU (International Telecommunication Union)

[0065] ITU-T Telecommunication Standardization Sector

[0066] LTE Long Term Evolution

[0067] LZMA Lempel–Ziv–Markov chain compression

[0068] LZMA2 is a simple container format that can include uncompressed data and LZMA data.

[0069] LZO Lempel–Ziv–Oberhumer compression

[0070] LZW Lempel–Ziv–Welch compression

[0071] MAC Media Access Control

[0072] Mdat Media Data Box

[0073] MME (Mobility Management Entity)

[0074] MMS Multimedia Messaging Service

[0075] Moov Movie Box

[0076] MP4 is a file format used for MPEG-4 Part 14 files.

[0077] MPEG Moving Picture Experts Group

[0078] MPEG-2 is H.222 / H.262 as defined by the ITU.

[0079] MPEG-4 is the audio and video coding standard used in ISO / IEC 14496.

[0080] MSB (Most Significant Bit)

[0081] NAL Network Abstraction Layer

[0082] NDU NN Compressed Data Unit

[0083] ng or NG, next generation

[0084] ng-eNB or NG-eNB, the next generation of eNB

[0085] NN Neural Network

[0086] NNEF Neural Network Exchange Format

[0087] NNR Neural Network Representation

[0088] NR New Radio (5G Radio)

[0089] Num

[0090] N / W or NW network

[0091] ONNX Open Neural Network Exchange

[0092] PB protocol buffer

[0093] PC (Personal Computer)

[0094] PDA (Personal Digital Assistant)

[0095] PDCP (Packet Data Convergence Protocol)

[0096] PHY physical layer

[0097] PID group identifier

[0098] PLC power line communication

[0099] PSNR (Peak Signal-to-Noise Ratio)

[0100] RAM (Random Access Memory)

[0101] RAN (Radio Access Network)

[0102] RFC Request for Comments

[0103] RFID (Radio Frequency Identification)

[0104] RFM Reference Frame Memory

[0105] RLC Radio Link Control

[0106] RRC Radio Resource Control

[0107] RRH Remote Radio Head

[0108] RU radio unit

[0109] Rx receiver / receiver

[0110] SDAP Service Data Adaptation Protocol

[0111] SGW Service Gateway

[0112] SMF Session Management Function

[0113] SMS Short Message Service

[0114] st(v) is a null-terminated string encoded as UTF-8 characters as specified in ISO / IEC 10646.

[0115] SVC (Scalable Video Coding)

[0116] Interface between S1 eNodeB and EPC

[0117] TCP / IP Transmission Control Protocol - Internet Protocol

[0118] TDMA (Time Division Multiple Access)

[0119] trak track box

[0120] TS transport stream

[0121] TV

[0122] Tx transmitter / send

[0123] UE User Equipment

[0124] ue(v) is the syntax element of the unsigned integer Exp-Golomb encoding, with the left bit first.

[0125] UICC General Integrated Circuit Card

[0126] UMTS (Universal Mobile Telecommunications System)

[0127] u(n) is an unsigned integer using n bits.

[0128] UPF User Plane Functions

[0129] URI (Uniform Resource Identifier)

[0130] URL Uniform Resource Locator

[0131] UTF-8 8-bit Unicode Transformation Format

[0132] WLAN (Wireless Local Area Network)

[0133] Interconnection interface between two eNodeBs in an X2 LTE network

[0134] Xn is the interface between two NG-RAN nodes.

[0135] The following describes in detail suitable apparatus and possible mechanisms for a video / image encoding process according to embodiments. In this regard, reference is made first to... Figure 1 and Figure 2 ,in Figure 1 An example block diagram of device 50 is shown. This device can be an Internet of Things (IoT) device configured to perform various functions, such as collecting information through one or more sensors, receiving or transmitting information, analyzing information received or collected by the device, and so on. The device may include a video encoding system that can incorporate a codec. Figure 2 The layout of the equipment according to an example embodiment is shown. Figure 1 and Figure 2 The components will be described below.

[0136] Electronic device 50 may be, for example, a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other low-power device. However, it should be understood that the exemplary embodiments described herein can be implemented within any electronic device or apparatus that can process data via a neural network.

[0137] The device 50 may include a housing 30 for attaching and protecting the device. The device 50 may also include a display 32 in the form of a liquid crystal display. In other embodiments of the examples described herein, the display may be any suitable display technology suitable for displaying images or video. The device 50 may also include a keypad 34. In other embodiments of the examples described herein, any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data input system as part of a touch-sensitive display.

[0138] The equipment may include a microphone 36 or any suitable audio input that may be a digital or analog signal input. Equipment 50 may also include an audio output device, which in the example embodiments described herein may be any of the following: a handset 38, a speaker, or an analog or digital audio output connection. Equipment 50 may also include a battery (or, in other examples described herein, the device may be powered by any suitable mobile energy device such as a solar cell, fuel cell, or clockwork generator). The equipment may also include a camera capable of recording or capturing images and / or video. Equipment 50 may also include an infrared port for short-range line-of-sight communication with other devices. In other embodiments, equipment 50 may also include any suitable short-range communication solution, such as Bluetooth wireless connectivity or USB / FireWire wired connectivity.

[0139] Equipment 50 may include a controller 56, a processor, or processor circuitry for controlling equipment 50. Controller 56 may be connected to memory 58, which, in the example embodiments described herein, may store data in the form of image and audio data and / or may also store instructions for implementation on controller 56. Controller 56 may also be connected to codec circuitry 54 adapted to perform encoding and / or decoding of audio and / or video data, or to assist in encoding and / or decoding performed by the controller.

[0140] The equipment 50 may also include a card reader 48 and a smart card 46, such as a UICC and a UICC card reader, for providing user information and suitable for providing authentication information for user authentication and authorization on the network.

[0141] Equipment 50 may include a radio interface circuit 52 connected to a controller and adapted to generate wireless communication signals, such as for communicating with a cellular communication network, a wireless communication system, or a wireless local area network. Equipment 50 may also include an antenna 44 connected to the radio interface circuit 52 for transmitting radio frequency signals generated at the radio interface circuit 52 to other equipment and / or for receiving radio frequency signals from other equipment.

[0142] Equipment 50 may include a camera capable of recording or detecting individual frames, which are then passed to codec 54 or a controller for processing. The equipment may receive video image data from another device for processing before transmission and / or storage. Equipment 50 may also receive images wirelessly or via a wired connection for encoding / decoding. The structural elements of the above-described equipment 50 represent examples of apparatuses for performing the corresponding functions.

[0143] about Figure 3 This document illustrates an example of a system in which embodiments of the examples described herein can be utilized. System 10 includes multiple communication devices that can communicate via one or more networks. System 10 may include any combination of wired or wireless networks, including but not limited to wireless cellular telephone networks (e.g., GSM, UMTS, CDMA, LTE, 4G, 5G networks, etc.), such as wireless local area networks (WLANs) defined by any IEEE 802.x standard, Bluetooth personal area networks, Ethernet LANs, token ring LANs, wide area networks, and the Internet.

[0144] System 10 may include wired and wireless communication devices and / or equipment 50, which are adapted to implement embodiments of the examples described herein.

[0145] For example, Figure 3 The system shown illustrates a representation of mobile phone network 11 and Internet 28. Connection to Internet 28 can include, but is not limited to, long-range wireless connections, short-range wireless connections, and various wired connections (including, but not limited to, telephone lines, cables, power lines, and similar communication paths).

[0146] The example communication devices shown in System 10 may include, but are not limited to, electronic devices or equipment 50, a combination of a personal digital assistant (PDA) and a mobile phone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, and a laptop computer 22. Equipment 50 may be stationary or mobile when carried by a mobile individual. Equipment 50 may also be positioned in a transportation mode, including but not limited to cars, trucks, taxis, buses, trains, ships, airplanes, bicycles, motorcycles, or any similar suitable mode of transportation.

[0147] The embodiments may also be implemented in the following: set-top boxes (i.e., digital TV receivers) that may or may not have display or wireless capabilities, tablet computers or (laptop) personal computers (PCs) with hardware and / or software to process neural network data, various operating systems, and chipsets, processors, DSPs and / or embedded systems that provide hardware / software-based encoding.

[0148] Some or more equipment can send and receive calls and messages, and communicate with service providers via a wireless connection 25 to base station 24. Base station 24 can connect to network server 26, which allows communication between mobile phone network 11 and Internet 28. The system may include additional communication equipment and various types of communication devices. Interface 2 is configured to provide access to Internet 28, for example for integrating a messaging device (IMD) 18.

[0149] Communication devices can communicate using various transmission technologies, including but not limited to Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol-Internet Protocol (TCP-IP), Short Message Service (SMS), Multimedia Messaging Service (MMS), email, Instant Messaging Service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband Internet of Things, and any similar wireless communication technologies. Communication devices relating to various embodiments of the present invention can communicate using various media, including but not limited to radio, infrared, laser, cable connections, and any suitable connection.

[0150] In telecommunications and data networks, a channel can refer to a physical channel or a logical channel. A physical channel can refer to a physical transmission medium such as a line, while a logical channel can refer to a logical connection on a multiplexed medium, capable of transmitting multiple logical channels. A channel can be used to transmit information signals (such as bit streams) from one or more transmitters to one or more receivers.

[0151] Implementations can also be carried out in so-called IoT devices. For example, the Internet of Things (IoT) can be defined as the interconnection of uniquely identifiable embedded computing devices within existing internet infrastructure. The convergence of various technologies has enabled and will enable many areas of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc., to be incorporated into the Internet of Things (IoT). To utilize the internet, IoT devices are assigned an IP address as a unique identifier. IoT devices can be equipped with radio transmitters, such as WLAN or Bluetooth transmitters or RFID tags. Alternatively, IoT devices can access IP-based networks via wired networks, such as Ethernet-based networks or power line connections (PLCs).

[0152] The MPEG-2 transport stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0, is a format used to carry audio, video, and other media, as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify the elementary stream (also known as a packetized elementary stream) within the TS. Therefore, logical channels within an MPEG-2 TS can be considered to correspond to specific PID values.

[0153] Available media file format standards include the ISO Basic Media File Format (ISO / IEC 14496-12, which can be abbreviated as ISOBMFF) and the file format derived from ISOBMFF for NAL unit structured video (ISO / IEC 14496-15).

[0154] A video codec consists of an encoder that transforms the input video into a compressed representation suitable for storage / transmission, and a decoder that decompresses the compressed video representation back into a visual form. The video encoder and / or video decoder can also be separate entities, i.e., they do not need to form a codec. Typically, the encoder discards some information from the original video sequence to represent the video in a more compact form (i.e., at a lower bit rate).

[0155] Typical hybrid video encoders, such as many encoder implementations of ITU-T H.263 and H.264, encode video information in two stages. First, pixel values ​​in a picture region (or “block”) are predicted, for example, by a motion compensation module (which locates and indicates regions closely corresponding to the block being encoded in one of the previously encoded video frames) or by a spatial module (which uses pixel values ​​around the block to be encoded in a specified manner). Second, the prediction error (i.e., the difference between the predicted pixel block and the original pixel block) is encoded. This can be done by transforming the differences in pixel values ​​using a specified transform (e.g., a Discrete Cosine Transform (DCT) or a variant thereof), quantization coefficients, and entropy encoding of the quantization coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size or transmission bit rate).

[0156] In temporal prediction, the prediction source can be a previously decoded image (aka the reference image). In intra-block copy (IBC; also known as intra-block copy prediction and current image reference), prediction can be applied similarly to temporal prediction, but the reference image can be the current image, and only previously decoded samples can be referenced during the prediction process. Inter-layer or inter-view prediction can be applied similarly to temporal prediction, but the reference image can be a decoded image from another scalable layer or from another view. In some cases, inter-frame prediction can refer only to temporal prediction, while in others, inter-frame prediction can be collectively referred to as temporal prediction as well as any of intra-block copy, inter-layer prediction, and inter-view prediction, provided they are performed using the same or similar process as temporal prediction. Inter-frame prediction or temporal prediction is sometimes referred to as motion compensation or motion compensation prediction.

[0157] Inter-frame prediction (also known as temporal prediction, motion compensation, or motion-compensated prediction) reduces temporal redundancy. In inter-frame prediction, the prediction source is the previously decoded image. Intra-frame prediction takes advantage of the fact that adjacent pixels within the same image may be related. Intra-frame prediction can be performed in the spatial or transform domain, i.e., it can predict sample values ​​or transform coefficients. Intra-frame prediction can be used in intra-frame coding when inter-frame prediction is not applied.

[0158] One outcome of the encoding process is a set of encoded parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are first predicted from spatially or temporally adjacent parameters. For example, motion vectors can be predicted from spatially adjacent motion vectors, and only the differences relative to the motion vector predictor can be encoded. The prediction of encoded parameters and intra-frame prediction can be collectively referred to as in-picture prediction.

[0159] Figure 4 A block diagram showing the general structure of a video encoder is provided. Figure 4 An encoder for two layers is presented, but it should be understood that the encoder presented can be similarly extended to encode more than two layers. Figure 4 A video encoder is shown, comprising a first encoder portion 500 for a base layer and a second encoder portion 502 for an enhancement layer. Each of the first encoder portion 500 and the second encoder portion 502 may include similar elements for encoding an input image. Encoder portions 500 and 502 may include pixel predictors 302 and 402, prediction error encoders 303 and 403, and prediction error decoders 304 and 404. Figure 4Embodiments of pixel predictors 302 and 402 are also shown, which include inter-frame predictors 306 and 406, intra-frame predictors 308 and 408, mode selectors 310 and 410, filters 316 and 416, and reference frame memories 318 and 418. The pixel predictor 302 of the first encoder section 500 receives a base layer image of the video stream to be encoded at both the inter-frame predictor 306 (which determines the difference between the image and the motion-compensated reference frame 318) and the intra-frame predictor 308 (which determines the prediction for the image patch based only on the processed portion of the current frame or picture). The outputs of both the inter-frame predictor and the intra-frame predictor are passed to the mode selector 310. The intra-frame predictor 308 may have more than one intra-frame prediction mode. Thus, each mode can perform intra-frame prediction and provide the prediction signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer image 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives the enhancement layer image of the video stream to be encoded at both the inter-frame predictor 406 (which determines the difference between the image and the motion-compensated reference frame 418) and the intra-frame predictor 408 (which determines the prediction for the image patch based only on the processed portion of the current frame or image). The outputs of both the inter-frame predictor and the intra-frame predictor are passed to the mode selector 410. The intra-frame predictor 408 may have more than one intra-frame prediction mode. Thus, each mode can perform intra-frame prediction and provide the prediction signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer image 400.

[0160] Depending on which encoding mode is selected to encode the current block, the output of inter-frame predictors 306, 406, or the output of one of the optional intra-frame predictor modes, or the output of the surface encoder within the mode selector, is passed to the output of mode selectors 310, 410. The output of the mode selector is passed to the first summing devices 321, 421. The first summing devices can subtract the output of pixel predictors 302, 402 from the base layer image 300 / enhancement layer image 400 to generate a first prediction error signal 320, 420, which is input to the prediction error encoders 303, 403.

[0161] Pixel predictors 302 and 402 further receive a combination of the predicted representations of image blocks 312 and 412 and the outputs 338 and 438 of the prediction error decoders 304 and 404 from the preliminary reconstructors 339 and 439. The preliminary reconstructed images 314 and 414 can be passed to intra-frame predictors 308 and 408 and filters 316 and 416. Filters 316 and 416 receiving the preliminary representations can filter the preliminary representations and output final reconstructed images 340 and 440, which can be stored in reference frame memories 318 and 418. Reference frame memory 318 can be connected to inter-frame predictor 306 to be used as a reference image for comparison with future base layer images 300 during inter-frame prediction operations. According to some embodiments, when the base layer is selected and indicated as a source for inter-layer sample prediction and / or inter-layer motion information prediction of the enhancement layer, reference frame memory 318 can also be connected to inter-frame predictor 406 to be used as a reference image for comparison with future enhancement layer images 400 during inter-frame prediction operations. In addition, the reference frame memory 418 can be connected to the inter-frame predictor 406 to be used as a reference image for comparison with the future enhancement layer image 400 during inter-frame prediction operations.

[0162] According to some embodiments, when the base layer is selected and indicated as a source for predicting the filtering parameters of the enhancement layer, the filtering parameters of the filter 316 from the first encoder section 500 can be provided to the second encoder section 502.

[0163] The prediction error encoders 303 and 403 include transform units 342 and 442 and quantizers 344 and 444. Transform units 342 and 442 transform the first prediction error signal 320 and 420 into the transform domain. This transform is, for example, a DCT transform. Quantizers 344 and 444 quantize the transform domain signal, for example, DCT coefficients, to form quantized coefficients.

[0164] Prediction error decoders 304 and 404 receive the outputs from prediction error encoders 303 and 403, and perform the reverse processing of prediction error encoders 303 and 403 to generate decoded prediction error signals 338 and 438. The decoded prediction error signals 338 and 438, when combined with the prediction representations of image blocks 312 and 412 at second summing devices 339 and 439, produce preliminary reconstructed images 314 and 414. The prediction error decoder can be considered to include dequantizers 346 and 446 and inverse transform units 348 and 448. Dequantizers 346 and 446 dequantize quantized coefficient values ​​(e.g., DCT coefficients) to reconstruct the transform signal, and inverse transform units 348 and 448 perform an inverse transform on the reconstructed transform signal, wherein the output of inverse transform units 348 and 448 contains the reconstructed blocks. The prediction error decoder may also include block filters that can filter the reconstructed blocks based on further decoding information and filter parameters.

[0165] Entropy encoders 330 and 430 receive the outputs of prediction error encoders 303 and 403, and can perform appropriate entropy coding / variable-length coding on the signal to provide error detection and correction capabilities. The outputs of entropy encoders 330 and 430 can be inserted into the bitstream, for example, via multiplexer 508.

[0166] Figure 5 This is a block diagram 500 illustrating the interface between an encoder 502 implementing neural network encoding 503 and a decoder 504 implementing neural network decoding 505, according to the example described herein. The encoder 502 may embody a device, software method, or hardware circuitry. The encoder 502 has the objective of compressing input data 511 (e.g., input video) into compressed data 512 (e.g., a bitstream), thereby minimizing the bit rate and maximizing the accuracy of the analysis or processing algorithm. To this end, the encoder 502 uses an encoder or compression algorithm, such as to perform neural network encoding 503.

[0167] A general analysis or processing algorithm may be part of the decoder 504. The decoder 504 uses a decoder or decompression algorithm, such as performing neural network decoding 505, to decode compressed data 512 (e.g., compressed video) encoded by the encoder 502. The decoder 504 produces decompressed data 513 (e.g., reconstructed data).

[0168] Encoder 502 and decoder 504 can be entities that implement abstractions, can be separate entities or the same entity, or can be part of the same physical device.

[0169] The analysis / processing algorithm can be any algorithm, traditional or learned from data. In the case of algorithms learned from data, it is assumed that the algorithm can be modified or updated, for example, using optimization via gradient descent. An example of a learning algorithm is a neural network.

[0170] ISO Basic Media File Format Available media file format standards include the ISO Basic Media File Format (ISO / IEC 14496-12, abbreviated as ISOBMFF), the MPEG-4 file format (ISO / IEC 14496-14, also known as MP4), the file format for structured video in NAL (Network Abstraction Layer) units (ISO / IEC 14496-15), and the 3GPP file format (3GPP TS 26.244, also known as 3GP). ISOBMFF is the basis from which all the above file formats are derived (excluding ISOBMFF itself).

[0171] Some concepts, structures, and specifications of ISOBMFF are described below as examples of container file formats based on which embodiments can be implemented. Aspects of the examples described herein are not limited to ISOBMFF, but are given for a possible basis upon which the examples described herein can be partially or fully implemented.

[0172] The basic building blocks in the ISO Basic Media File Format are called boxes. Each box has a header and a payload. The box header indicates the type of box and its size in bytes. Boxes can enclose other boxes, and the ISO file format specifies which box types are allowed within a particular type of box. Furthermore, the presence of some boxes in a file may be mandatory, while the presence of others may be optional. Additionally, for some box types, multiple boxes may be allowed in a file. Therefore, the ISO Basic Media File Format can be considered to specify a hierarchical structure of boxes.

[0173] According to the ISO Basic Media File Format, the file includes media data and metadata encapsulated in boxes. Each box is identified by a four-character code (4CC) and begins with a header indicating the type and size of the box.

[0174] In files conforming to the ISO Basic Media File Format, media data can be provided in one or more instances of MediaDataBox('mdat'), and MovieBox('moov') can be used to wrap metadata for timed media. In some cases, both "mdat" and "moov" boxes may be required for a file to be operational. A "moov" box can contain one or more tracks, and each track can reside in a corresponding TrackBox("trak"). Each track is associated with a handler, identified by a four-character code that specifies the track type. Video, audio, and image sequence tracks can be collectively referred to as media tracks, and they contain the basic media stream. Other track types include cue tracks and timed metadata tracks.

[0175] Tracks include samples, such as audio or video frames. For video tracks, media samples may correspond to coded images or access units. A media track refers to a sample formatted according to a media compression format (and its encapsulation of the ISO Basic Media File Format). A cue track refers to a cue sample containing instructions for constructing data packets for transmission via the indicated communication protocol. A timing metadata track may refer to a sample describing the referenced media and / or cue sample.

[0176] Movie clips can be used, for example, when recording content to an ISO file to avoid data loss in case of recording application crashes, insufficient storage space, or other accidents. Without movie clips, data loss may occur because the file format may require all metadata (such as movie boxes) to be written to a contiguous area of ​​the file. Furthermore, while recording the file, there may not be enough storage space (e.g., random access memory (RAM)) to buffer movie boxes against the available storage size, and recalculating the contents of the movie boxes when closing the movie may be too slow. Additionally, movie clips allow for simultaneous recording and playback of the file using a regular ISO file parser. Moreover, for progressive downloads, such as receiving and playing the file simultaneously while using movie clips, a shorter initial buffering duration may be required, and the initial movie box size is smaller compared to a structured file with the same media content but without movie clips.

[0177] Movie clip features enable the segmentation of metadata into multiple pieces, whereas metadata might otherwise reside within a movie box. Each piece can correspond to a specific time segment of a track. In other words, movie clip features allow for the interweaving of file metadata and media data. Therefore, the size of a movie box may be limited to the use cases described above.

[0178] In some examples, media samples for movie clips can reside in an `mdat` box. However, for the metadata of a movie clip, a `moof` box can be provided. The `moof` box can include information about the duration of a specific playback time that was previously in a `moov` box. The `moov` box itself may still represent a valid movie, but in addition, it can include an `mvex` box indicating that the movie clip might follow in the same file. The movie clip can be extended in time with the rendering associated with the `moov` box.

[0179] Within a movie clip, there can be a set of track segments, including any position from zero to multiple for each track. Track segments can further include any position from zero to multiple track runs, with each document in its document representing a consecutive sample run for that track (and thus similar to a chunk). Within these structures, many fields are optional and can be defaulted. Metadata that can be included in a moof box may be restricted to a subset of metadata that can be included in a moof box and may be encoded differently in some cases. Detailed information about boxes that can be included in a moof box can be found in the ISOBMFF specification.

[0180] A self-contained movie clip can be defined as consisting of consecutive moof boxes and mdat boxes in file order, where the mdat box contains a sample of the movie clip (the moof box provides metadata for it) and does not contain samples of any other movie clips (i.e., any other moof boxes).

[0181] A media segment can include one or more self-contained movie clips. Media segments can be used for delivery, such as in MPEG-DASH, for example, streaming.

[0182] The track reference mechanism can be used to associate tracks with each other. A TrackReferenceBox contains boxes, each providing a reference from its containing track to a set of other tracks. These references are identified by the box type of the containing box (i.e., the box's four-character code).

[0183] The ISO Basic Media File Format includes three mechanisms for timing metadata that can be associated with a specific sample: sample groups, timing metadata tracks, and sample ancillary information. Derivative specifications can provide similar functionality through one or more of these three mechanisms.

[0184] Sample grouping in ISO basic media file formats and their derivatives (such as AVC and SVC file formats) can be defined as assigning each sample in a track as a member of a sample group based on a grouping criterion. Sample groups are not limited to consecutive samples and can contain non-adjacent samples. Since a sample in a track may have more than one sample group, each sample group can have a type field to indicate the type of group. Sample groups can be represented by two linked data structures: (1) a SampleToGroupBox (sbgp box) representing the assignment of samples to sample groups; and (2) a SampleGroupDescriptionBox (sgpd box) containing sample group entries for each sample group, describing the attributes of that group. Multiple instances of SampleToGroupBox and SampleGroupDescriptionBox may exist based on different grouping criteria. These instances can be distinguished by a type field used to indicate the type of group. SampleToGroupBox may include a grouping_type_parameter field, which can be used, for example, to indicate the subtype of the group.

[0185] The MPEG compression representation of the neural network standard (MPEG NNR-ISO / IEC 15938-17) aims to provide a standardized way to compress and distribute “neural networks” (hereafter referred to as NN). This is an important aspect of the era of AI-enabled Internet of Things (AIoT) devices and ecosystems, where billions of connected IoT devices may be intelligent and have AI components (e.g., connected cars, home automation systems, smartphones, cameras, etc.).

[0186] Several "exchange formats" have been defined in the industry. ONNX (https: / / onnx.ai / ) or NNEF (https: / / www.khronos.org / nnef) can be listed as the two most well-known. However, they lack any aspect of neural networks in terms of compression, and do not define a flexible and well-structured high-level syntax for such compressed NN data. They provide topological information and links from topological elements to neural network weights and / or coefficients.

[0187] MPEG NNR has the following use cases, which typically require resolution through potentially advanced bitstream syntax:

[0188] • Transmission of the entire or part of a neural network (e.g., layer by layer)

[0189] • Update the neural network in a timely manner

[0190] • Provides information about certain characteristics of the compressed neural network (e.g., accuracy, compression method, compression ratio, etc.).

[0191] • Access compressed neural network representations in a progressive manner

[0192] • Reusable weights or portions of a neural network (in many cases, only the initial layers of the neural network are used, for example when the neural network is used only as a feature extractor and the extracted features are used by other neural networks or portions of other neural networks).

[0193] • Retraining and federated learning of neural networks

[0194] The examples described in this paper introduce a high-level syntax for neural networks used in MPEG NNR compression. It provides granular access to relevant NN information at the layer or sub-layer level (i.e., filters, kernels, biases, etc.). It also enables the transmission of compressed NN representations over communication channels, as well as partial updates to the compressed neural network. It can also be used with existing exchange formats to provide compressed representations of neural network weights and coefficients.

[0195] The examples described in this paper improve the interchangeability and storage of compressed neural networks. The examples also improve decoder-side operations by providing a well-defined high-level syntax for exchanging compressed neural network information.

[0196] The examples described in this article define a high-level bitstream syntax for compressed neural networks, which includes the following logical concepts:

[0197] - Compressed NN global metadata signaling

[0198] -NN topology-level metadata signaling

[0199] -NN quantization parameter signaling

[0200] -Metadata signaling related to compressed NN data

[0201] - Compressed NN data signaling

[0202] A basic bitstream syntax is defined. This bitstream is initiated by global metadata spanning the entire compressed neural network. The bitstream consists of information units containing metadata and / or compressed neural network data for specific parts of the entire neural network. This partitioning is related to the topological information of the neural network, with a granularity down to the uniquely identifiable parts of the neural network (layers, filters, kernels, etc.).

[0203] Some novel aspects of the examples described in this paper can be listed as a method, including:

[0204] • Signal transmission of computational structural information of neural networks (i.e., topology and quantization)

[0205] • Global information about the neural network is transmitted as metadata. This metadata can include NN type information, compressed representation information, compressed representation accuracy information, etc.

[0206] • Divide the neural network into independent or subordinate decodeable units.

[0207] • Signal transmission of segmentation information and metadata associated with segmentation units

[0208] • Transmit relevant information signals about the segmented neural network units to the structural information of the neural network.

[0209] • Store the encoded representation of the segmented neural network units and the associated metadata used to decode such units.

[0210] A (quantized) neural network (NN) is typically represented by the following information:

[0211] -General level parameters of neural networks

[0212] - The network topology, which has unique identifiers for different topology elements.

[0213] - (If the weights of the quantized neural network) Information about the quantization

[0214] - Variables representing topological elements and their corresponding data types.

[0215] Exchange formats such as NNEF define the above information and store it in files within a predefined directory tree structure.

[0216] In the example described in this article, this data structure is defined as independent of the underlying file storage mechanism, but in an interoperable manner, so that this exchange mechanism can utilize compressed NN bit streams and convert them into their own data and storage structures.

[0217] Compressed NNN bitstreams (hereinafter referred to as NNR bitstreams) have the following characteristics: Figure 6 The structure shown.

[0218] NNR bitstream 602 is composed of NNR units. NNR unit 604 may include the following information:

[0219] -NNR Unit Size: This information can be the total size of the signaled NNR unit in bytes. The size of this field can be 16 bits or 32 bits, and it can be indicated by the most significant bit (MSB) of the first byte.

[0220] - NNR Unit Header: This field may contain information about the type of data carried in the payload, general metadata about the NNR unit it carries, etc. A more complete list of such metadata elements is given in the following sections.

[0221] - NNR Unit Payload: It carries compressed or uncompressed NN-related data. Such data can be one of the following: NNR parameter set (i.e., NN global information), topology data, quantized data, or compressed / uncompressed NN data.

[0222] NNR units can be cascaded to form a serialized bitstream that can represent N / N. An NNR encoder can generate such a serialized bitstream. This bitstream is then carried through a transmission channel to a decoder for decoding.

[0223] In another embodiment, such an NNR bitstream can be stored as a file in a virtual or non-virtual directory tree structure. For example, NNEF can be used to carry compressed NN data. In this case, the NN variables can be compressed using NNR encoding and stored as an NNR bitstream in a file within a predetermined directory structure.

[0224] In some embodiments, the topology data, quantized data, and compressed NN data can be all or part of the data; this means the data is either in a single NNR unit or split across multiple NNR units. In the latter case, the NNR unit header may contain information indicating such partial storage. This information may be represented by a counter that counts backward to indicate the number of splits used. For example, if the data is split into two NNR units, the first NNR unit would contain counter number 2, while the second NNR unit could contain counter number 1. In another embodiment, the counter may represent one less value than the number of splits.

[0225] In another embodiment, a flag in the NNR cell header may indicate a segmentation and another flag may indicate the last NNR cell belonging to the segmented data.

[0226] In another embodiment, the segmented NNR units may have the same identifier in their NNR unit headers, which indicates the NN level information to which the segment belongs. This could be a unique ID, a unique string, a relative or absolute URI or URL, etc.

[0227] If the topology information is segmented, multiple topology data NNR units can be aggregated to form the final NN topology. The NN compressed data unit (NDU) payload can belong to the entire neural network, an entire layer, or a portion thereof.

[0228] The variables / weights in the payload of an NN compressed data unit can be mapped to topology elements via unique references or labels. Such data can be carried in the NNR NDU unit header (i.e., in the NNR unit header of an NNR unit whose payload contains compressed data). This data can be a unique ID, a unique string, a relative or absolute URI or URL, etc.

[0229] Each NNR unit header may contain information about the type of data in the NNR unit data payload 606. The following table provides an example of such data units and their enumerated data values:

[0230]

[0231] In the table above, enumerations, types, and identifiers are given as examples only, and other values ​​can be used.

[0232] Figure 7 Figure 700 is an example illustrating how an NNR bitstream can be composed of several NNR units of different types (e.g., NNR units 1 to 6).

[0233] The order in which the units exist can vary and can be chosen by the NNR encoder and content creator. The NNR start unit 702 may be absent if the start of NN data can be inferred through other means (e.g., file storage, storage format indication, etc.). In another embodiment, the NNR start unit 702 may have a payload with a fixed or variable number of bytes that indicates a start code. (See reference...) Figure 6 Reference mark 606 (“NNR NN Data Start Cell Payload”).

[0234] NNR NN-level parameter set The NNR parameter set payload 704 can consist of all or a subset of the following information:

[0235] 1. Indicator, indicating whether the bitstream internally carries the topology as a data unit or the bitstream provides a reference / URL / Id for that topology.

[0236] 2. Model accuracy and other performance metrics [test-accuracy, test-dataset-ID / URL, bitrate, etc.]. If the server has a set of pre-compressed models, this information can be used by both the server and client sides, and based on the client's request, the server selects the closest matching compressed NN model.

[0237] 3. In addition, based on the information in the preceding field (5), the server can generate a list to be sent to the client, and then the client can select its preferred compression model.

[0238] 4. Update indicator (whether it is an update to the NN)

[0239] 5. Update Reference: The baseline NN version on which weights are applied for updates. This may be a unique NN version ID.

[0240] 6. A sparsification indicator (possibly a single bit) indicating whether sparsification has been performed.

[0241] 7. Sparse indicator tensor (indicating which cells or weights are zero)

[0242] 8. A quantization indicator (possibly a single bit) indicating whether quantization has been performed.

[0243] 9. Quantization step size (if using scalable unified quantization)

[0244] 10. Quantitative Chart

[0245] 11. Maximum memory size required on the decoder side (this is usually related to the maximum number of layers and output activations): some examples might be the maximum number of fully connected parameters in a layer, the maximum number of convolutional parameters in a layer, etc.

[0246] 12. Network-level priority information (maps decoder-side compression parameters to finer-grained performance metrics. For example, in the case of a classifier, how each class is affected by certain sparsity thresholds).

[0247] 13. Network-level entropy information, such as entropy models and context models. This might be the ID of one of the context models available on the decoder side.

[0248] 14. Information about the input type (e.g., media type; required input size; range; normalization; numerical precision; etc.)

[0249] As an example, the following NNEF model can be encoded by NNR to have the following characteristics: Figure 8 The head shown. In particular, Figure 8 Example topology description 800 of AlexNet is shown. AlexNet uses the Neural Network Exchange Format (NNEF) topology graph format.

[0250] The corresponding NNR network parameter set, NNR cell 704, can be defined as follows:

[0251] • NNR unit header: "NNR_NPS"

[0252] • Payload:

[0253] 1. [Does the file explicitly carry the topology or use it as an ID?]

[0254] 2. [Model accuracy and other performance metrics: test-accuracy, test-dataset-ID / URL, bitrate] ACC = 37%, dataset = www.imagenet.com / 1234, file_size = 23MB

[0255] 3. [Update flag] 0

[0256] 4. [Update Reference: The baseline NN version on which weights are updated. This can be a unique NN version ID]0

[0257] 5. [Sparseness flag] 0

[0258] 6. [Sparseness indicator tensor (indicates which cells or weights are zero)] 0 (This can be included in the NNR cells of each layer)

[0259] 7. [Quantitative Indicator] 0

[0260] 8. [Quantization step size (if it is scalable unified quantization)]: None

[0261] 9. [Quantization graph (if based on non-uniform quantization based on codebook)]: dict(0:0.2,1:0.77,2:-0.8,....)

[0262] 10. [Maximum memory size required on the decoder side: maximum number of fully connected parameters, maximum number of convolutional parameters] 12345, 54321

[0263] 11. [Further sparsification at the decoder's side network level (mapping the sparsification threshold to the resulting performance metric)] dict(0.01:37,0.05:40,0.1:49)

[0264] 12. [Network-level priority information (maps decoder-side compression parameters to finer-grained performance metrics)] dict(0.01:dict(class1:35,class2:20,class3:54,class4:43),0.05:dict(class1:50,class2:25,class3:43,class4:35),0.1:dict(class1:40,class2:60,class3:54,class4:36))

[0265] 13. [Network-level entropy information, such as entropy models and context models] context_model_3

[0266] 14. [Information about the input type (e.g., media type; required input size and shape; range; normalization; numerical precision; etc.)] "image", [3,224,224], 0-255, dict("mean": [0.47,0.47,0.47,], "variance": [0.5,0.5 0.5]), 24 bits.

[0267] The NNR_NPSNNR unit payload can be formatted to include both fixed and variable length information for each of the aforementioned information elements.

[0268] NNR Topology Data Unit The NNR topology data unit 706 can include the following information in the NNR unit header:

[0269] - The NNR unit type is NNR_TPL

[0270] - The topology format enumeration is NNEF: This field indicates the actual format of the stored topology information. Possible values ​​are shown in the table below:

[0271]

[0272] - Whether the topology has been further compressed. This information may include an enumeration of the following compression indicators:

[0273]

[0274]

[0275] - Partial Information Flag: Indicates that the information in this data unit is partial.

[0276] - Last data segment marker: Indicates that this is the last data unit of the partial information. Alternatively,

[0277] - Counter: Indicates the index of the partial information. A value of 0 indicates no partial information, while a value greater than 0 indicates the index of the partial information. This counter can count backwards to initially indicate the total number of segments.

[0278] The NNR data unit payload may contain topology data units in compressed or uncompressed format, or in segmented or unsegmented format, as indicated in the NNR unit header data.

[0279] NNR Quantization Data Unit The data unit 708 may contain a header of the same type as that in the NNR topology data unit, but the data unit type is marked as "NNR_QNT". It may contain the same fields defined above. The NNR data unit payload may contain quantized data in compressed or uncompressed format, or in split or unsplit format, as indicated in the NNR unit header data. An example of quantized data is a dictionary or lookup table mapping quantized values ​​to dequantized values.

[0280] NNR compressed NN data unit NNR Compressed Data Units (NDUs, also known as CDUs) 710 and 712 can contain all or part of the information of data elements belonging to the NN topology or graph.

[0281] An NDU can be identified by an NNR cell header containing the type "NNR_NDU". The NNR cell header of an NDU may contain a subset or all of the following information:

[0282] 1. NDU number or identifier

[0283] 2. NDU element activation / disabling

[0284] 3. NDU-specific quantization plot. If empty, an NN-level quantization plot can be used.

[0285] 4. Priority information for NDUs (mapping decoder-side compression parameters to result performance metrics). Priority information can be relative priority information between different NDUs.

[0286] 5.1 bits: A flag indicating whether the matrix has been decomposed.

[0287] 6. Additional information about the decomposed matrix

[0288] 7. NDU-level entropy information, such as entropy models and context models.

[0289] 8. A flag indicating that the NDU can be decoded independently.

[0290] 9. Map elements in the NDU to an array of unique identifiers for topological elements (e.g., in the case of NNEF, this is a list of labels that exist in the NDU).

[0291] 10. NDU Counter: Backward count of the number of relevant NDUs (e.g., partially carried NN encoded variables): Default value may be 0.

[0292] All or part of the NNR-compressed data can be carried as a payload.

[0293] Using the same AlexNet example as above, several variables can be stored in the NNR NDU, as shown below. In this implementation example,

[0294] • NNR NDU unit:

[0295] • Header: “NNR_NDU”, base_id = "alexnet_v2 / conv1 / ", id_list = {['kernel', enum(type)], ['bias', enum(type)]} / / …. enum(type) corresponds to one of the supported data types, such as float32, float64, int16, int32, uint16, int16, int8, uint8, etc.

[0296] Other information in the header includes:

[0297] 1. [NDU Number] 1

[0298] 2. [NDU Activation or Disabling] 0

[0299] 3. [NDU-specific quantization plot. If empty, use an NN-level quantization plot.] Empty

[0300] 4. [Priority information for NDU (mapping decoder-side compression parameters to result performance metrics)] 1

[0301] 5. [Has the matrix been decomposed?] 0

[0302] 6. [Additional information about the decomposed matrix] None

[0303] 7. [NDU-level entropy information, such as entropy models and context models.] context_model_3

[0304] 8. [Independently Decodeable NDU Flag] 1

[0305] 9. NDU counter: 0 (Alternatively, NDU_Partial_Flag = False, NDU_Last_Flag = False)

[0306] - Payload: A compressed representation of the NNR variables, such as those listed in the ID / tag array in the header.

[0307] `base_id` indicates the basic identifier for different NN variables in the payload. `id_list` contains variable IDs and type pairs.

[0308] NNR bitstream advanced syntax The data structures and information in the following sections are given as examples, and their names and values ​​are variable. Their order or cardinality in the high-level syntax may also differ in different implementations.

[0309] Bitstream type descriptors: The following descriptors specify the parsing process for each syntax element:

[0310] -b(8): Bytes with any bit string pattern (8 bits). The parsing process for this descriptor is specified by the return value of the function read_bits(8).

[0311] -f(n): A fixed-pattern bit string (from left to right) written using n bits, with the left bit first. The parsing process for this descriptor is specified by the return value of the function read_bits(n).

[0312] -i(n): Uses an n-bit signed integer. When n is "v" in the syntax table, the number of bits varies depending on the values ​​of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits(n), which is interpreted as a two's complement integer representation with the most significant bit written first. Specifically, the parsing process for this descriptor is as follows:

[0313]

[0314] -st(v): Null-terminated strings are encoded as UTF-8 characters according to ISO / IEC 10646. The parsing process is as follows: st(v) starts from a byte-aligned position in the bitstream, reads and returns a series of bytes from the bitstream, starting from the current position and continuing to the next byte-aligned byte (excluding 0x00), and advances the bitstream pointer by (stringLength + 1) * 8 bits, where stringLength equals the number of bytes returned. The st(v) syntax descriptor is only used when the current position in the bitstream is a byte-aligned position.

[0315] -u(n): Uses an n-bit unsigned integer. When n is "v" in the syntax table, the number of bits varies in a way that depends on the values ​​of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits(n), which is interpreted as a binary representation of an unsigned integer, with the most significant bit written first.

[0316] -ue(v): Unsigned integer zero-order Exp-Golomb encoded syntax element, left bit first.

[0317] Byte alignment: The following data structures are assumed to be byte-aligned. To enable this alignment, the byte_alignment() data structure is appended to other data structures.

[0318]

[0319] NNR bitstream The following data structures are newly defined in the context of compressed neural network high-level bitstream syntax.

[0320] NNR Unit Syntax The NNR unit consists of size information, header information, and payload information.

[0321]

[0322] The more_data_in_nnr_unit() function is defined as follows:

[0323] - If there is more data following the current nnr_unit, that is, the decoded data in the current nnr_unit up to now is less than numBytesInNNRUnit, then the return value of more_data_in_nnr_unit() is TRUE.

[0324] Otherwise, the return value of more_data_in_nnr_unit() is FALSE.

[0325] In some embodiments, each piece of information within an NNR unit may have its size information. In another embodiment, only a subset of such information may exist in the NNR unit. In yet another embodiment, each NNR unit may have a start code and an end code for marking the beginning and end of these NNR units. The start code may be a predefined bit pattern. The bitstream syntax allows the start code to be identifiable, i.e., the beginning of the NNR unit can be found by searching for the bit pattern of the start code in the bitstream.

[0326] NNR Unit Size Syntax The NNR unit size can indicate the total size of the NNR units in bytes. It can provide overall NNR unit size information, including nnr_unit_size(). In some embodiments, it can indicate only the size of the header and payload.

[0327]

[0328] The `nnr_unit_size_flag` indicates the number of bits used as the data type for `nnr_unit_size`. If this value is 0, then `nnr_unit_size` is a 15-bit unsigned integer value; otherwise, it is a 31-bit unsigned integer value. In another embodiment, for some NNR unit types, such as NR units of type "NNR_STR" (NN start indicator), `nnr_unit_size_flag` can be forced to be set to 0.

[0329] NNR Unit Header Syntax: The NNR unit header can provide information about the NNR unit type and additional related metadata.

[0330]

[0331]

[0332] The `nnr_unit_type` parameter indicates the type of the NNR unit. The following NNR units can be defined.

[0333] NNR unit type Type identifier value NN-level parameter set data unit NNR_NPS 0x01 NN topology or graphical data unit NNR_TPL 0x02 NN Quantized Data Unit NNR_QNT 0x03 NN Compressed Data Unit NNR_NDU 0x04 NN Compressed Network Data Start Unit NNR_STR 0x00 reserve 0x05-0xFF

[0334] It should be noted that the table above is an example, and many more NNR data units can be defined. Furthermore, the type identifier and value are given as examples, and other identifiers and values ​​can be defined. In another embodiment, nnr_unit_type can be defined with less than 8 bits. In all the examples below, "(count-1)" can be replaced with a variable that can be called "countMinusOne" or a similar name, which can indicate a value one less than the count variable.

[0335] NNR parameter set unit header Below is an example syntax for nnr_parameter_set_unit_header().

[0336]

[0337] The `topology_flag` flag indicates the presence of topological information in the NN high-level syntax bitstream. When set to 1, it indicates that the topology is in the bitstream and carried along with the NNR unit type "NNR_TPL". If it is 0, it means that the topology is referenced externally via an ID, URI, URL, etc.

[0338] When nn_update_flag is set to 1, it indicates that the NNR unit can be used for partial updates of a previous NN with idupdate_nn_id.

[0339] `nn_id` (and `update_nn_id`) is a unique identifier for the NNR-encoded NN. In another embodiment, this field can be a null-terminated string that may contain absolute or relative URIs or URLs, or a unique string. In the NNEF context, this string reference can correspond to one or more NNEF variables that are stored as a ".dat" file.

[0340] When set to 1, sparsification_flag indicates that sparsification is applied to the NN.

[0341] The `sprsification_tensor()` function contains information about which units or weights are zero. It can have the following syntax:

[0342]

[0343] In `sparsification_tensor()`: `compressed_flag` indicates whether compression is applied to `sparsification_data()`. The `compression_format` enumeration is used to specify the compression algorithm for `sparsification_data()`. In another embodiment, there may be an explicitly defined order in the sparse data representation that directly maps to the order of the weights. In yet another embodiment, `sparsification_tensor()` may reside within the NNR unit payload of an NNR unit.

[0344] When set to 1, `decomposition_flag` indicates that at least one matrix in the NN compressed data unit contains decomposition. `quantization_flag` can indicate the presence of quantization information. `quantization_step_size` can indicate the step size interval for uniform quantization of the scalar.

[0345] The `quantization_map()` function can handle non-uniform quantization schemes based on codebooks for signal propagation. It may have the following syntax:

[0346]

[0347] `quantization_map_data()` can be an array (e.g., a dictionary) of the form `{[index<integer>:value<floating-point>]}`, where the index can be a quantization value indicator and the second value can be a signed floating-point value corresponding to that quantization value index. In another embodiment, each index can indicate a range of quantization steps. In another embodiment, these values ​​can be of type 8-bit, 16-bit, 32-bit, or 64-bit floating-point values. In another embodiment, the quantization map can be carried in the NNR cell payload of the NNR cell.

[0348] `max_memory_requirement` can indicate the maximum memory required to run a neural network for an NNR decoder or inference device. In one embodiment, this value can be indicated as a concatenation of two values: the maximum number of fully connected parameters and the maximum number of convolutional parameters in a layer or a portion of a layer (e.g., in a convolutional kernel). In another embodiment, this value can be indicated as a 64-bit value and can be an unsigned integer or a floating-point number coerced to an unsigned integer.

[0349] The `spareification_performance_map()` function can signal a mapping between different sparsification thresholds and the resulting NN inference accuracy based on a chosen accuracy reporting scheme. In the following example, accuracy is a value between 0 and 100, where the threshold is a floating-point value.

[0350]

[0351] The count can signal the number of information tuples present in the data structure. In another embodiment, `sparsification_performance_map()` can be carried in the NNR unit payload of the NNR unit. `sparsification_threshold` can signal a sparsification threshold; when applied to weights (i.e., weights or parameters below `sparsification_threshold` are zeroed); the inference accuracy is `nn_accuracy` (scaled from 0 to 100). In another embodiment, `nn_accuracy` is a relative value between the entries listed in the data structure.

[0352] The `dataset_id()` function can provide information about which datasets and which versions of the datasets were used to compute performance measurements. In another embodiment, this data structure can also exist in the header or payload of the `NNR_NDU` unit to indicate the accuracy level of different thresholds when applied to data structures within the `NNR` compressed data unit.

[0353] In another embodiment, sparsification_performance_map() may include the model’s task-related performance (e.g., classification accuracy or PSNR in image compression), the performance degradation or gain relative to the original non-sparse model, and the compression ratio in terms of non-zero ratio obtained by applying sparsification to the model.

[0354] accuracy_information() can provide information about the accuracy of compressed NNs on different datasets.

[0355]

[0356] `dataset_information` is the absolute or relative URI or URL of the dataset, relative to its calculated accuracy. `nn_dataset_accuracy` is an accuracy value between 0 and 100. In another embodiment, `nn_accuracy` is the ratio between the uncompressed and compressed NN accuracy when testing against a dataset. In this case, the value can be signaled as a floating-point value.

[0357] `priority_map()` can signal information about how different aspects of a neural network's performance are affected by some compression parameters (e.g., sparsification with different thresholds). For example, in the case of a classifier neural network, this information could include a dictionary or lookup table mapping a set of sparsification thresholds to a set of corresponding accuracy values ​​for each class.

[0358]

[0359]

[0360] The `compression_parameter` can signal one or more compression parameter values, such as a sparsity threshold or the number of quantization points. In one embodiment, different `compression_parameters` can be applied to different components (e.g., variables) of the NN signaled in the NDU.

[0361] The accuracies_per_aspect can take the form of aspect-based granularity, signaling information about the accuracy of the neural network when the compression_parameter is applied. The aspect of the neural network can be, for example, the class in the case of a classifier neural network (thus providing accuracy for each class), or the relationship between the bounding box size and the bounding box center in the case of a detector neural network (thus providing accuracy for the bounding box size and the bounding box center separately), etc.

[0362] dataset_id() can provide information about which dataset(s) and which version of the dataset were used when calculating performance measurements.

[0363] nn_entropy_information() can signal information about which entropy model or context model is used on the decoder side among all available entropy models or context models.

[0364]

[0365] The nn_input_type_information() function can signal information about the type of input data being received, such as media type (image, audio frame, etc.), size and shape required for the input data structure, range, normalization method and parameters, numerical precision, and so on.

[0366]

[0367] NNR Topology Unit Header This data structure transmits topology-related header data.

[0368]

[0369] `topology_storage_format` indicates the actual format of the stored topology information. Possible values ​​are shown in the table below:

[0370]

[0371] The `topology_compressed_flag` can indicate whether the topology has been further compressed.

[0372] compression_format can contain an enumeration of the following compression directives:

[0373]

[0374] `partial_flag` indicates that the information in this data unit is partial. `last_flag` indicates that this is the last data unit of partial information. `counter` indicates the index of the partial information. A value of 0 indicates no partial information, and a value greater than 0 indicates the index of a partial information. This counter can count backwards to initially indicate the total number of splits. In another embodiment, `partial_flag` and `last_flag` may not exist when the counter exists in the data structure. Conversely, they may also not exist.

[0375] NNR quantization unit header This header information may be very similar to the NNR topology unit header.

[0376]

[0377] `quantization_storage_format` can have the same syntax and semantics as `topology_storage_format`. `quantization_compressed_flag` can have the same syntax and semantics as `topology_compressed_flag`.

[0378] NNR Compressed Data Unit Header The NNR compressed data unit header provides information about the NNR compressed data units that follow it. Its data structure and semantics may be as follows:

[0379]

[0380]

[0381] `base_id` is a unique string that can be used to indicate the base URI, URL, root, or similar information of a directory tree structure. When concatenated with the `id_name` value of an `id_list` element, it provides a unique identifier for the compressed NN data unit element. In the NNEF context, this unique identifier can correspond to an NNEF variable saved as a ".dat" file.

[0382] id_list can provide a list of uniquely identifiable neural network topology elements that exist within compressed NN data units.

[0383]

[0384]

[0385] `count` indicates the number of entities listed in the data structure. `id_name` provides a unique identifier for compressed NN data unit elements that may span a portion of the compressed data. In one embodiment, such an identifier may correspond to a variable identifier in the NNEF topology graph. The interpretation of this field may depend on the compressed data format (i.e., NNEF, ONNX, MPEG, etc.).

[0386] `data_type` can be an enumerated data type value. Possible values ​​can be (but are not limited to): binary, uint, int, floating-point numbers with 1 bit, 4 bits, 8 bits, 16 bits, 32 bits, and 64 bits of precision. `data_size` can indicate the number of parameters or weights belonging to this id when the compressed N data units are uncompressed. In another embodiment, this value can indicate the byte size corresponding to these parameters or weights.

[0387] NDU_quantization_map() can handle non-uniform quantization schemes based on codebooks for signal transmission. It may have the following syntax:

[0388]

[0389] `quantization_map_data()` can be an array (e.g., a dictionary) of the form `{[index<integer>:value<floating-point>]}`, where the index can be a quantization value indicator, and the second value can be a signed floating-point value corresponding to that quantization value index. In another embodiment, each index can indicate a range of quantization steps. In yet another embodiment, these values ​​can be of type 8-bit, 16-bit, 32-bit, or 64-bit floating-point values.

[0390] NDU_priority_map() can signal information about how different aspects of a neural network's performance are affected by some compression parameters (e.g., sparsification with different thresholds or different levels of precision used for quantization). For example, in the case of a classifier neural network, this information could include a dictionary or lookup table mapping a set of sparsification thresholds to corresponding groups of accuracy by class.

[0391]

[0392] `compression_parameter` can signal one or more compression parameter values, such as the sparsity threshold or the number of quantization points. `accuracies_per_aspect` can signal information about the accuracy of a neural network at the granularity of each aspect when `compression_parameter` is applied. An aspect of the neural network could be, for example, the class in the case of a classifier neural network (therefore, accuracy is provided per class), or the bounding box size relative to the bounding box center in the case of a detector neural network (therefore, accuracy is provided separately for the bounding box size and the bounding box center), etc. `dataset_id()` can provide information about which datasets and which versions of the datasets were used when computing performance measurements.

[0393] NDU_entropy_information() can signal information about which entropy model or context model is used on the decoder side among all available entropy models or context models. Context models can be used in lossless coding steps (such as arithmetic coding) to estimate the probability of the next symbol to be encoded or decoded.

[0394]

[0395]

[0396] NDU_decomposition_information() can signal information about the decomposition method used to compress the variables considered in this NNR unit and its parameters.

[0397]

[0398] The `decomposition_method` signal transmits information about the decomposition method. Its default value may be 0, indicating that this field is not used.

[0399]

[0400] The `decomposition_parameters()` signal transmits information about specific parameters of the decomposition method specified by `decomposition_method`. These parameters are useful on the decoder side for reconstructing the original data structure or for the inference process. Such information may include the dimensions of the matrices produced by the decomposition.

[0401] NNR Start Unit HeaderThe NNR start unit header can indicate the start of the compressed NN bitstream. It has a unique signature so that it can be identified when parsing the bitstream from any bit index. In another embodiment, some start code emulation prevention scheme can be applied to the NNR compressed bitstream such that the value is not emulated and does not exist in any other position in the bitstream. The emulation prevention scheme can, for example, add emulation prevention bytes to the bitstream at byte positions where start code emulation would otherwise occur. In another embodiment, such emulation prevention can be applied to the entire bitstream value of NNR unit size + NNR unit header + NNR unit payload. Such a value with the following definition of nnr_start_code might be as follows: 0x000C00F0F0F0F0F0F0F0F0 (2-byte size + 1-byte unit type + 8-byte header data). In one embodiment, the decoder or another entity identifies the NNR unit, for example, from the start code, and then removes the start code emulation prevention from the NN bitstream or individual NNR units, for example, by identifying bytes added to the bitstream to avoid start code emulation and removing those bytes from the bitstream. Then the decoding of the NNR unit can be completed without considering start code emulation bytes or similar interference syntax.

[0402] nnr_start_unit_header(){ descriptor nnr_start_code u(64) }

[0403] `nnr_start_code` can indicate the start of the NNR bitstream. This value can be a 64-bit value, such as 0xF0F0F0F0F0F0F0F0. This value is given as an example, and other values ​​can be defined. The NN start indicator NNR unit may not have an NNR data payload. In another embodiment, `nnr_start_code` can be stored as the payload, and `nnr_start_header` can be empty.

[0404] NNR unit payload. The following provides an example syntax for nnr_unit_payload.

[0405]

[0406]

[0407] The nnr_parameter_set_payload() function can be empty or filled with some data structures that have already been defined in the NNR parameter set unit header.

[0408] `nnr_topology_unit_payload()` can be a partial or complete representation of the topology. It can be compressed or uncompressed. Its format is defined in the NNR topology unit header. The NNR decoder is expected to provide this information to higher-level components, either before or after decompression. The NNR decoder may not understand the data format of this structure unless it is defined by the same entity that defines the NNR decoder.

[0409] `nnr_quantization_unit_payload()` can be a partial or complete representation of the quantization parameters. It can be compressed or uncompressed. Its format is defined in the NNR quantization unit header. The NNR decoder is expected to provide this information to higher-level components, either before or after decompression. The NNR decoder may not understand the data format of this structure unless it is defined by the same entity that defines the NNR decoder.

[0410] nnr_start_unit_payload() can be empty, or it can contain the NNR start code as defined above.

[0411] `nnr_data_unit_payload()` is an NNR compressed data unit. Its compression scheme is defined by the same entity that defines the NNR decoder. An NNR compressed data unit can be decoded independently or depend on other compressed data units. The NNR compressed data unit header information provides the necessary metadata about the various characteristics of the compressed data unit.

[0412] NNR decoding process When the NNR decoder receives an NNR-encoded bitstream, it is expected to perform the following steps (the order of parsing elements may change):

[0413] 1. Parse and check if an NNR cell of type NNR_STR exists.

[0414] 2. Once an NNR_STR cell is found, start parsing the next NNR cell by reading the cell size, header information, and payload.

[0415] 3. Identify and parse the topology NNR unit, and provide the topology information to the decoding entity.

[0416] 4. Identify and parse the quantized NNR units and provide topology information to the decoding entity.

[0417] 5. Identify and parse NNR compressed data units and provide them to the NNR decoding process. Decoding can be performed unit by unit or by combining multiple NNR units together.

[0418] In another embodiment, the NNR decoder can simply parse the bitstream from the beginning without paying attention to the NNR start data unit. The presence of such a data unit can be signaled to the NNR decoder in other ways.

[0419] In some embodiments of the examples described herein, the decoder may determine to perform (further) compression on the neural network data based on information provided via the disclosed high-level grammar. Specifically, priority information in the HLS is used to determine which compression parameter values ​​to apply based on given requirements. For example, the requirements may be acceptable overall accuracy or acceptable accuracy for a subset of classes (in the case of a classifier neural network).

[0420] Figure 9 Example device 900, which can be implemented in hardware, is configured to implement a high-level grammar for compressed representation of neural networks based on the examples described herein. Device 900 includes a processor 902 and at least one non-transitory memory 904 including computer program code 905, wherein at least one memory 904 and computer program code 905 are configured to utilize at least one processor 902 to enable the device to implement a high-level grammar 906 based on the examples described herein. Device 900 optionally includes a display or I / O 908 for displaying content during rendering. Device 900 optionally includes one or more network (NW) interface (I / F) 910s. The one or more NW I / F 910s can be wired and / or wireless and communicate via the Internet / other networks using any communication technology. The one or more NW I / F 910s can include one or more transmitters and one or more receivers. The one or more N / W I / F 910s can include standard well-known components such as amplifiers, filters, frequency converters, (de)modulators and encoder / decoder circuitry, and one or more antennas.

[0421] Equipment 900 can be a remote, virtual, or cloud-based equipment. Equipment 900 can be an encoder or decoder, or both. Memory 904 can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. Memory 904 can include a database for storing data. Equipment 900 does not need to include every feature mentioned, or may include other features. Equipment 900 can correspond to or Figure 1 and Figure 2 Another embodiment of the equipment 50 shown, or Figure 3 Any of the equipment shown. Equipment 900 can correspond to... Figure 11 The equipment shown is or Figure 11Another embodiment of the equipment shown includes a UE 110, a RAN node 170, or a network element 190.

[0422] Figure 10 Example method 1000 is for implementing a high-level grammar for compressed representation of neural networks. In 1002, the method includes encoding or decoding a high-level bitstream grammar for at least one neural network. In 1004, the method includes wherein the high-level bitstream grammar includes at least one information unit of compressed neural network data having metadata or being a portion of said at least one neural network. In 1006, the method includes wherein a serialized bitstream includes one or more of said at least one information unit.

[0423] Turning Figure 11 The figure illustrates a block diagram of one possible, non-limiting example in which these examples can be practiced. It shows a user equipment (UE) 110, a radio access network (RAN) node 170, and one or more network elements 190. Figure 1 In this example, User Equipment (UE) 110 wirelessly communicates with Wireless Network 100. The UE is a wireless device that can access Wireless Network 100. UE 110 includes one or more processors 120, one or more memories 125, and one or more transceivers 130 interconnected via one or more buses 127. Each of the one or more transceivers 130 includes a receiver Rx 132 and a transmitter Tx 133. The one or more buses 127 may be address, data, or control buses and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optic cables, or other optical communication devices. The one or more transceivers 130 are connected to one or more antennas 128. The one or more memories 125 include computer program code 123. UE 110 includes a module 140, comprising one or both of portions 140-1 and / or 140-2, which may be implemented in various ways. Module 140 may be implemented in hardware as module 140-1, for example, as part of one or more processors 120. Module 140-1 may also be implemented as an integrated circuit or via other hardware such as a programmable gate array. In another example, module 140 may be implemented as module 140-2, which is implemented as computer program code 123 and executed by one or more processors 120. For example, one or more memories 125 and computer program code 123 may be configured, together with one or more processors 120, to cause user equipment 110 to perform one or more of the operations described herein. UE 110 communicates with RAN node 170 via radio link 111.

[0424] In this example, RAN node 170 is a base station that provides access to wireless network 100 for wireless devices (e.g., UE 110). RAN node 170 can be, for example, a base station for 5G, also known as New Radio (NR). In 5G, RAN node 170 can be an NG-RAN node, which is defined as a gNB or ng-eNB. A gNB is a node that provides NR user plane and control plane protocol termination to the UE and is connected to the 5GC (e.g., network element 190) via an NG interface. An ng-eNB is a node that provides E-UTRA user plane and control plane protocol termination to the UE 110 and is connected to the 5GC via an NG interface. An NG-RAN node can include multiple gNBs, and a gNB can also include a central unit (CU) (gNB-CU) 196 and a distributed unit (DU) (gNB-DU), where DU 195 is shown. Note that DU 195 can include a radio unit (RU) or be coupled to and control the RU. A gNB-CU is a logical node that hosts the Radio Resource Control (RRC), SDAP, and PDCP protocols of a gNB, or controls the RRC and PDCP protocols of an en-gNB that controls the operation of one or more gNB-DUs. The gNB-CU terminates the F1 interface connected to the gNB-DU. The F1 interface is shown as reference 198, although reference 198 also shows links between remote elements and centralized elements of RAN node 170, such as the link between gNB-CU 196 and gNB-DU 195. A gNB-DU is a logical node that hosts the RLC, MAC, and PHY layers of a gNB or en-gNB, and its operation is partially controlled by the gNB-CU. One gNB-CU supports one or more cells. A cell is supported by only one gNB-DU. The gNB-DU terminates the F1 interface 198 connected to the gNB-CU. Note that DU 195 is considered to include transceiver 160, for example, as part of an RU; however, some examples in this regard could allow transceiver 160 to be part of a separate RU, for example, under the control of and connected to DU 195. RAN node 170 could also be an eNB (evolved NodeB) base station for LTE (Long Term Evolution), or any other suitable base station or node.

[0425] RAN node 170 includes one or more processors 152, one or more memories 155, one or more network interfaces (N / WI / F) 161, and one or more transceivers 160 interconnected via one or more buses 157. Each of the one or more transceivers 160 includes a receiver Rx 162 and a transmitter Tx 163. The one or more transceivers 160 are connected to one or more antennas 158. The one or more memories 155 include computer program code 153. CU 196 may include processor 152, memory 155, and network interface 161. Note that DU 195 may also contain its own one or more memories and processors, and / or other hardware, but these are not shown.

[0426] RAN node 170 includes module 150, which includes one or both of portions 150-1 and / or 150-2, which can be implemented in various ways. Module 150 can be implemented in hardware as module 150-1, for example, as part of one or more processors 152. Module 150-1 can also be implemented as an integrated circuit or through other hardware such as a programmable gate array. In another example, module 150 can be implemented as module 150-2, which is implemented as computer program code 153 and executed by one or more processors 152. For example, one or more memories 155 and computer program code 153 are configured, together with one or more processors 152, to cause RAN node 170 to perform one or more of the operations described herein. Note that the functionality of module 150 can be distributed, for example, distributed between DU 195 and CU 196, or implemented solely in DU 195.

[0427] One or more network interfaces 161 communicate on the network, for example, via links 176 and 131. Two or more gNBs 170 can communicate using, for example, link 176. Link 176 can be wired, wireless, or a combination of both, and can implement, for example, an Xn interface for 5G, an X2 interface for LTE, or other suitable interfaces for other standards.

[0428] One or more buses 157 may be address, data, or control buses and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optic or other optical communication equipment, wireless channels, etc. For example, one or more transceivers 160 may be implemented as a Remote Radio Header (RRH) 195 for LTE or a Distributed Unit (DU) 195 for a gNB implementation of 5G, wherein other elements of the RAN node 170 may be physically located in a different location from the RRH / DU 195, and one or more buses 157 may be partially implemented as, for example, fiber optic cables or other suitable network connections to connect other elements of the RAN node 170 (e.g., Central Unit (CU) 196, gNB-CU) to the RRH / DU 195. Reference 198 also indicates those suitable network links.

[0429] Note that the description in this article indicates that a "cell" performs functions, but it should be clear that the equipment forming the cell can perform these functions. A cell constitutes part of a base station. That is, each base station can have multiple cells. For example, a single carrier frequency and associated bandwidth may have three cells, each covering one-third of a 360-degree area, thus the coverage area of ​​a single base station is approximately elliptical or circular. Furthermore, each cell can correspond to a single carrier, and a base station can use multiple carriers. So if each carrier has 3 120-degree cells and 2 carriers, then the base station has a total of 6 cells.

[0430] Wireless network 100 may include one or more network elements 190, which may include core network functions and provide connectivity to further networks such as telephone networks and / or data communication networks (e.g., the Internet) via one or more links 181. Such core network functions for 5G may include Access and Mobility Management Functions (AMF) and / or User Plane Functions (UPF) and / or Session Management Functions (SMF). Such core network functions for LTE may include MME (Mobility Management Entity) / SGW (Serving Gateway) functions. These are merely example functions that can be supported by network element 190, and it should be noted that both 5G and LTE functions can be supported. RAN node 170 is coupled to network element 190 via link 131. Link 131 may be implemented as, for example, an NG interface for 5G, or an S1 interface for LTE, or other suitable interfaces for other standards. Network element 190 includes one or more processors 175, one or more memories 171, and one or more network interfaces (N / WI / F) 180 interconnected via one or more buses 185. The one or more memories 171 include computer program code 173. One or more memories 171 and computer program code 173 are configured, together with one or more processors 175, to cause network element 190 to perform one or more operations.

[0431] Wireless network 100 can implement network virtualization, which is the process of combining hardware and software network resources and network functions into a single software-based management entity, a virtual network. Network virtualization involves platform virtualization and is often used in conjunction with resource virtualization. Network virtualization is categorized externally as combining many networks or parts of networks into a virtual unit, or internally as providing network-like functionality to a software container on a single system. Note that the virtualized entity created by network virtualization is still implemented to some extent using hardware such as processors 152 or 175 and memories 155 and 171, and this virtualized entity also produces technical effects.

[0432] Computer-readable storage devices 125, 155, and 171 can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic storage devices and systems, optical storage devices and systems, fixed storage, and removable storage. Computer-readable storage devices 125, 155, and 171 can be means for performing storage functions. Processors 120, 152, and 175 can be of any type suitable for the local technical environment and can include one or more of general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), and processors based on multi-core processor architectures, as non-limiting examples. Processors 120, 152, and 175 can be means for performing functions such as controlling UE 110, RAN node 170, network element 190, and other functions described herein.

[0433] Typically, various embodiments of user equipment 110 may include, but are not limited to, cellular phones (e.g., smartphones), tablet computers, personal digital assistants (PDAs) with wireless communication capabilities, portable computers with wireless communication capabilities, image capture devices (e.g., digital cameras) with wireless communication capabilities, gaming devices with wireless communication capabilities, music storage and playback devices with wireless communication capabilities, internet devices that allow wireless internet access and browsing, tablet computers with wireless communication capabilities, and portable units or terminals that combine these functions.

[0434] One or more of modules 140-1, 140-2, 150-1, and 150-2 can be configured to implement a high-level syntax for compressed representations of neural networks based on the examples described herein. Computer program code 173 can also be configured to implement a high-level syntax for compressed representations of neural networks based on the examples described herein.

[0435] References to “computer,” “processor,” etc., should be understood to encompass not only computers with different architectures (e.g., single / multiprocessor architectures and sequential (von Neumann) / parallel architectures) but also special-purpose circuits (e.g., field-programmable gate arrays (FPGAs), special-purpose circuits (ASICs), signal processing devices, and other processing circuits). References to computer programs, instructions, code, etc., should be understood to encompass software or firmware for programmable processors (e.g., programmable content of hardware devices (e.g., instructions for processors)), or configuration settings for fixed-function devices, gate arrays, or programmable logic devices, etc.

[0436] As used herein, the term "circuit" may refer to any of the following: (a) a hardware circuit implementation, such as an implementation in analog and / or digital circuitry; and (b) a combination of circuitry and software (and / or firmware), such as (if applicable): (i) a combination of processors, or (ii) a portion of processor / software, including a digital signal processor, software, and memory, which work together to enable an apparatus to perform various functions; and (c) a circuit that requires software or firmware to operate, such as a microprocessor or a portion of a microprocessor, even if the software or firmware is not physically present. This description of "circuit" applies to the use of the term in this application. As a further example, as used herein, the term "circuit" will also cover only an implementation of a processor (or multiple processors) or a portion of a processor and its accompanying software and / or firmware. For example, if applicable to a particular element, the term "circuit" will also cover a baseband integrated circuit or application processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, cellular network equipment, or other network equipment.

[0437] An example apparatus includes: at least one processor; and at least one non-transitory memory including computer program code; wherein the at least one memory and the computer program code are configured, together with the at least one processor, to cause the apparatus to at least: encode or decode a high-level bitstream syntax for at least one neural network; wherein the high-level bitstream syntax includes at least one information unit having metadata or compressed neural network data as a portion of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit.

[0438] The equipment may also include the ability to store the serialized bitstream as a file in a virtual or non-virtual directory tree structure, or to send it as a bitstream through a data pipeline.

[0439] The apparatus may further include, wherein the portion of the at least one neural network is at least one of the following: a layer, a filter, a kernel, a bias, a quantization weight, a tensor, or any other data structure that is a recognizable portion of the at least one neural network.

[0440] The equipment may also include, wherein the information unit comprises: a unit size for signal transmission of the information unit; a unit payload carrying compressed or uncompressed data related to the neural network; and a unit header having information about the type of data carried by the unit payload.

[0441] The equipment may also include, wherein the unit payload includes at least one of the following: a set of parameters including global information about the neural network; topology data; compressed or uncompressed neural network data unit payload; complete or partial neural network data; quantized data; or payload data associated with the start code.

[0442] The equipment may also include, wherein at least one of topology data, compressed or uncompressed neural network data unit payload, or quantized data is divided into multiple information units.

[0443] The equipment may also include, wherein the unit head includes information for indicating the segmentation.

[0444] The equipment may also include a counter that indicates the number of segments used, by counting backwards.

[0445] The equipment may also include a marker in the cell header indicating the segmentation, and another marker in the cell header indicating the last information cell belonging to the segmented data.

[0446] The device may also include multiple information units having the same identifier in their respective unit heads to indicate the neural network level information to which the segment belongs.

[0447] The equipment may also include a neural network exchange format for carrying compressed neural network data.

[0448] The equipment may also include a compressed neural network data unit payload that is mappable to topology data via a reference or tag within a unit header associated with the compressed network data unit payload, wherein the reference or tag includes at least one of the following: a unique identifier, a unique string, or a relative or absolute Uniform Resource Identifier or Locator.

[0449] The device may also include a cell header indicating the start of the serialized bit stream.

[0450] The equipment may also include an encoder that provides a serialized bit stream to a decoder via a transmission channel.

[0451] The equipment may also include, wherein the high-level syntax decoding includes: parsing the unit by reading the size of at least one information unit, the unit header associated with the information unit, and the payload associated with the information unit; and identifying and parsing at least one of topological data, quantized data, or compressed or uncompressed data associated with the information unit.

[0452] The apparatus may further include, wherein the at least one memory and the computer program code are further configured, together with the at least one processor, to cause the apparatus to at least perform: checking the presence of a start unit, the start unit indicating the start of the bit stream and the start of decoding at that start of the bit stream.

[0453] The equipment may further include, wherein the at least one memory and the computer program code are further configured, together with the at least one processor, to cause the equipment to at least: further compress compressed data associated with the information unit.

[0454] Other aspects of the equipment may include the following. The information unit may include: a unit size of bytes for signal transmission; a unit payload carrying compressed or uncompressed data and related metadata associated with the at least one neural network; and a unit header containing information about the data type carried by the unit payload and related metadata. The unit payload may include at least one of the following: a parameter set including global metadata and information about at least one neural network; neural network topology information and related data; compressed or uncompressed neural network data, whether complete or partial; quantized data; or a compressed neural network bitstream start indicator or payload data associated with a start code. The neural network unit header may include information indicating the segmentation. A counter value of 0 may indicate no partial information, and a counter value greater than 0 may indicate an index of partial information. Multiple information units may have the same unique identifier in their respective unit headers to indicate the neural network topology element-level information to which the segment belongs. The unique identifier may be a Khronos Neural Network Exchange Format (NNEF) variable identifier or label in the NNEF topology diagram. Topology information may include Khronos Neural Network Exchange Format (NNEF) topology information. Multiple information units may have flags in their unit headers indicating whether these information units are independently decodable. The parameter set may contain flags indicating the presence and carrying of topological units in the compressed neural network bitstream. The parameter set may contain flags indicating whether sparsity is applied to the at least one neural network. The parameter set may contain a sparsity performance graph data structure whose signal transmission maps between different sparsity thresholds and the resulting neural network inference accuracy. The resulting neural network inference accuracy may correspond to the performance of the at least one neural network in terms of output accuracy. The unit payload or header may contain a quantization mapping data structure whose signal transmission codebook includes mappings between quantized values ​​and corresponding dequantized values. The unit header may indicate the neural network unit type, indicating the start of the serialized bitstream, wherein the serialized bitstream is a compressed or uncompressed neural network bitstream. Decoding of the high-level syntax may include: parsing the unit by reading the size of at least one information unit, the unit header associated with the information unit, and the payload associated with the information unit; identifying and parsing at least one of the topological data, quantization data, start code indicator data, parameter set data, or compressed or uncompressed data associated with the information unit.

[0455] An example apparatus includes: means for encoding or decoding an advanced bitstream syntax for at least one neural network; wherein the advanced bitstream syntax includes at least one information unit having metadata or compressed neural network data as a part of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit.

[0456] Other aspects of the equipment may include the following. The serialized bitstream may be stored as a file in a virtual or non-virtual directory tree structure, or transmitted as a bitstream through a data pipeline. The portion of the at least one neural network is at least one of the following: a layer, filter, kernel, bias, quantization weight, tensor, or any other data structure that is an identifiable part of the at least one neural network. The information unit may include: a unit size that signals the size of the information unit; a unit payload carrying compressed or uncompressed data and related metadata associated with the at least one neural network; and a unit header having information about the data type carried by the unit payload and related metadata. The unit payload may include at least one of the following: a parameter set including global metadata and information about the at least one neural network; neural network topology information and related data; compressed or uncompressed neural network data, which may be complete or partial; quantization data; or payload data associated with a compressed neural network bitstream start indicator or start code. At least one of the topology data, compressed or uncompressed neural network data unit payload, or quantization data may be divided into multiple information units. The neural network unit header may include information indicating the division. The information indicating the division may be represented by a counter that counts backwards to indicate the number of divisions used. A counter value of 0 indicates the absence of partial information, while a counter value greater than 0 indicates an index of partial information. Multiple information units may have the same unique identifier in their respective unit headers to indicate the level of information at the neural network topology element level to which they belong. This unique identifier may be a Khronos Neural Network Exchange Format (NNEF) variable identifier or label in the NNEF topology graph. Topology information may include Khronos Neural Network Exchange Format (NNEF) topology information. Multiple information units may have flags in their unit headers indicating whether these information units are independently decodable. The parameter set may contain flags indicating the presence and carrying of topology units in the compressed neural network bitstream. The parameter set may contain flags indicating whether sparsity is applied to the at least one neural network. The parameter set may contain a sparsity performance graph data structure whose signal transmission maps between different sparsity thresholds and the resulting neural network inference accuracy. The resulting neural network inference accuracy may correspond to the performance of the at least one neural network in terms of output accuracy. The unit payload or header may contain a quantization mapping data structure whose signal transmission codebook includes mappings between quantized values ​​and corresponding dequantized values. Neural network exchange formats can be used to carry compressed neural network data.The compressed neural network data unit payload can be mapped to topology data via a reference or tag within a unit header associated with the compressed network data unit payload, wherein the reference or tag includes at least one of the following: a unique identifier, a unique string, or a relative or absolute Uniform Resource Identifier or locator. The unit header can indicate the type of neural network unit, indicating the start of a serialized bitstream, wherein the serialized bitstream is a compressed or uncompressed neural network bitstream. The encoder can provide the serialized bitstream to the decoder via a transmission channel. Decoding of the high-level syntax can include: parsing the unit by reading the size of at least one information unit, the unit header associated with the information unit, and the payload associated with the information unit; identifying and parsing at least one of the topology data, quantization data, start code indicator data, parameter set data, or compressed or uncompressed data associated with the information unit. The apparatus may also include means for checking the presence of a start unit, which indicates the start of the bitstream and the start of decoding at the beginning of the bitstream. The apparatus may also include means for further compressing the compressed data associated with the information unit.

[0457] An example method includes: encoding or decoding a high-level bitstream syntax for at least one neural network; wherein the high-level bitstream syntax includes at least one information unit having metadata or compressed neural network data as a portion of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit.

[0458] Other aspects of the method may include the following. The serialized bitstream may be stored as a file in a virtual or non-virtual directory tree structure, or transmitted as a bitstream through a data pipeline. The portion of the at least one neural network is at least one of the following: a layer, filter, kernel, bias, quantization weight, tensor, or any other data structure of at least one identifiable portion of the neural network. The information unit may include: a unit size that signals the size of the information unit; a unit payload carrying compressed or uncompressed data and related metadata associated with the neural network; and a unit header having information about the data type carried by the unit payload and related metadata. The unit payload may include at least one of the following: a parameter set including global metadata and information about at least one neural network; neural network topology information and related data; compressed or uncompressed neural network data, which may be complete or partial; quantization data; or a compressed neural network bitstream start indicator or payload data associated with a start code. At least one of the topology data, compressed or uncompressed neural network data unit payload, or quantization data may be segmented into multiple information units. The neural network unit header may include information indicating the segmentation. The information indicating the segmentation may be represented by a counter that counts backwards to indicate the number of segments used. A counter value of 0 indicates the absence of partial information, while a counter value greater than 0 indicates an index of partial information. Multiple information units can have the same unique identifier in their respective unit headers to indicate the level of information at the neural network topology element level to which the segment belongs. This unique identifier can be a Khronos Neural Network Exchange Format (NNEF) variable identifier or label in the NNEF topology graph. Topology information can include Khronos Neural Network Exchange Format (NNEF) topology information. Multiple information units can have flags in their unit headers indicating whether these information units are independently decodable. The parameter set can contain flags indicating the presence and carrying of topology units in the compressed neural network bitstream. The parameter set can contain flags indicating whether sparsity is applied to at least one neural network. The parameter set can contain a sparsity performance graph data structure that signals a mapping between different sparsity thresholds and the resulting neural network inference accuracy. The resulting neural network inference accuracy can correspond to the performance of the at least one neural network in terms of output accuracy. The unit payload or header can contain a quantization mapping data structure that signals a codebook including a mapping between quantized values ​​and corresponding dequantized values. Neural network exchange formats can be used to carry compressed neural network data.The compressed neural network data unit payload can be mapped to topology data via a reference or tag within the unit header associated with the compressed network data unit payload, wherein the reference or tag includes at least one of the following: a unique identifier, a unique string, or a relative or absolute Uniform Resource Identifier or locator. The unit header can indicate the type of neural network unit, indicating the start of a serialized bitstream, wherein the serialized bitstream is a compressed or uncompressed neural network bitstream. The encoder can provide the serialized bitstream to the decoder via a transport channel. Decoding of the high-level syntax can include: parsing the unit by reading the size of at least one information unit, the unit header associated with the information unit, and the payload associated with the information unit; identifying and parsing at least one of the topology data, quantization data, start code indicator data, parameter set data, or compressed or uncompressed data associated with the information unit. The method can also include checking for the presence of a start unit indicating the start of the bitstream and the start of decoding at the beginning of the bitstream. The method can also include further compressing the compressed data associated with the information unit.

[0459] A machine-readable example non-transitory program storage device is provided, which tangibly embodies a machine-executable program of instructions for performing operations, said operations including: encoding or decoding a high-level bitstream syntax for at least one neural network; wherein the high-level bitstream syntax includes at least one information unit having metadata or compressed neural network data as a portion of the at least one neural network; and wherein a serialized bitstream includes one or more of the at least one information unit.

[0460] Other aspects of the non-transitory program storage device may include the following. The serialized bitstream may be stored as a file in a virtual or non-virtual directory tree structure, or transmitted as a bitstream via a data pipeline. The portion of the at least one neural network is at least one of the following: a layer, filter, kernel, bias, quantization weight, tensor, or any other data structure of at least one identifiable portion of the neural network. The information unit may include: a unit size that signals the size of the information unit; a unit payload carrying compressed or uncompressed data and related metadata associated with the neural network; and a unit header having information about the data type carried by the unit payload and related metadata. The unit payload may include at least one of the following: a parameter set including global metadata and information about at least one neural network; neural network topology information and related data; compressed or uncompressed neural network data, which may be complete or partial; quantization data; or payload data associated with a compressed neural network bitstream start indicator or start code. At least one of the topology data, compressed or uncompressed neural network data unit payload, or quantization data may be divided into multiple information units. The neural network unit header may include information indicating the division. Information indicating the segmentation can be represented by a counter that counts backwards to indicate the number of segments used. A counter value of 0 indicates no partial information, while a counter value greater than 0 indicates an index of partial information. Multiple information units can have the same unique identifier in their respective unit headers to indicate the neural network topology element-level information to which the segment belongs. This unique identifier can be a Khronos Neural Network Exchange Format (NNEF) variable identifier or label in the NNEF topology graph. Topology information can include Khronos Neural Network Exchange Format (NNEF) topology information. Multiple information units can have flags in their unit headers indicating whether these information units are independently decodable. The parameter set can contain flags indicating the presence and carrying of topology units in the compressed neural network bitstream. The parameter set can contain flags indicating whether sparsity is applied to the at least one neural network. The parameter set can contain a sparsity performance graph data structure that maps the signal transmission between different sparsity thresholds and the resulting neural network inference accuracy. The resulting neural network inference accuracy can correspond to the performance of the at least one neural network in terms of output accuracy. The cell payload or header may contain a quantization-mapped data structure with a signal transmission codebook that includes a mapping between quantized values ​​and their corresponding dequantized values. Compressed neural network data can be carried using a neural network exchange format.The compressed neural network data unit payload can be mapped to topology data via a reference or tag within the unit header associated with the compressed network data unit payload, wherein the reference or tag includes at least one of the following: a unique identifier, a unique string, or a relative or absolute Uniform Resource Identifier or locator. The unit header can indicate the type of neural network unit, indicating the start of a serialized bitstream, wherein the serialized bitstream is a compressed or uncompressed neural network bitstream. The encoder can provide the serialized bitstream to the decoder via a transmission channel. Decoding of the high-level syntax can include: parsing the unit by reading the size of at least one information unit, the unit header associated with the information unit, and the payload associated with the information unit; identifying and parsing at least one of the topology data, quantization data, start code indicator data, parameter set data, or compressed or uncompressed data associated with the information unit. Operation of the non-transitory program storage device can also include checking for the presence of a start unit indicating the start of the bitstream and the start of decoding at the start of the bitstream. Operation of the non-transitory program storage device can also include further compressing the compressed data associated with the information unit.

[0461] It should be understood that the above description is illustrative only. Those skilled in the art can devise various alternatives and modifications. For example, the features recited in the various dependent claims can be combined with each other in any suitable combination. Furthermore, features from the different embodiments described above can be selectively combined to form new embodiments. Therefore, this description is intended to cover all such alternatives, modifications, and variations falling within the scope of the appended claims.

Claims

1. An apparatus comprising: At least one processor; as well as At least one non-temporary memory, including computer program code; The at least one non-transitory memory and the computer program code are configured, together with the at least one processor, to cause the apparatus to perform at least the following operations: Encode or decode high-level bitstream syntax for at least one neural network; The high-level bitstream syntax includes at least one information unit, which has metadata or compressed neural network data as a part of the at least one neural network; and The serialized bit stream includes one or more information units from the at least one information unit; The information unit includes: a unit size that transmits signals of the size of the information unit; a unit payload carrying compressed or uncompressed data and related metadata related to the at least one neural network; and a unit header having information about the data type carried by the unit payload and related metadata. The unit payload includes at least one of the following: a set of parameters including global metadata and information about the at least one neural network; neural network topology information and related topology data; compressed or uncompressed neural network data, which may be complete or partial; quantized data; or a compressed neural network bitstream start indicator or payload data associated with a start code.

2. The equipment according to claim 1, wherein, The serialized bitstream is stored as a file in a virtual or non-virtual directory tree structure, or sent as a bitstream through a data pipeline.

3. The equipment according to claim 1, wherein, The portion of the at least one neural network is at least one of the following: a layer, a filter, a kernel, a bias, a quantized weight, a tensor, or any other data structure of at least one identifiable portion of the neural network.

4. The equipment according to claim 1, wherein, At least one of the topology data, the compressed or uncompressed neural network data unit payload, or the quantized data is divided into multiple information units.

5. The equipment according to claim 4, wherein, The neural network unit head includes information for indicating the segmentation.

6. The equipment according to claim 5, wherein, The information indicating the segmentation is represented by a counter that counts backwards to indicate the number of segments used.

7. The equipment according to claim 5, wherein, A counter value of 0 indicates that there is no partial information, and a counter value greater than 0 indicates the index of the partial information.

8. The equipment according to claim 4, wherein, The multiple information units have the same unique identifier in their respective unit headers to indicate the neural network topology element-level information to which the segment belongs.

9. The equipment according to claim 8, wherein, The unique identifier is the Khronos neural network exchange format NNEF variable identifier or label in the NNEF topology graph.

10. The equipment according to claim 4, wherein, The topology information includes Khronos neural network exchange format NNEF topology information.

11. The equipment according to claim 4, wherein, The plurality of information units have a flag in their unit header indicating whether such information unit is independently decodeable.

12. The equipment according to claim 1, wherein, The parameter set contains flags indicating the presence and carrying of topological units in the compressed neural network bitstream.

13. The equipment according to claim 1, wherein, The parameter set includes flags indicating whether sparsification will be applied to the at least one neural network.

14. The equipment according to claim 1, wherein, The parameter set includes a sparsification performance graph data structure, which transmits a mapping between different sparsification thresholds and the resulting neural network inference accuracy.

15. The equipment according to claim 14, wherein, The obtained neural network inference accuracy corresponds to the performance of the at least one neural network in terms of output accuracy.

16. The equipment according to claim 1, wherein, The unit payload or header contains a quantization mapping data structure, the quantization mapping data structure signal transmission codebook, the codebook including a mapping between quantized values ​​and corresponding dequantized values.

17. The equipment according to claim 1, wherein, The compressed neural network data is carried using a neural network exchange format.

18. The equipment according to claim 1, wherein, The compressed neural network data unit payload is mappable to the topology data via a reference or tag in the unit header associated with the compressed network data unit payload, wherein the reference or tag includes at least one of the following: a unique identifier, a unique string, or a relative or absolute Uniform Resource Identifier or Locator.

19. The equipment according to claim 1, wherein, The unit header indicates the type of neural network unit, which indicates the start of the serialized bitstream, wherein the serialized bitstream is a compressed or uncompressed neural network bitstream.

20. The equipment according to claim 1, wherein, The encoder provides the serialized bit stream to the decoder via a transmission channel.

21. The equipment according to claim 1, wherein, The decoding of the high-level bitstream syntax includes: The unit is parsed by reading the size of at least one information unit, the unit header associated with the information unit, and the payload associated with the information unit; and Identify and parse at least one of the following associated with the information unit: topology data, quantization data, start code indicator data, parameter set data, or compressed or uncompressed data.

22. The equipment according to claim 21, wherein, The at least one non-transitory memory and the computer program code are further configured, together with the at least one processor, to cause the device to perform at least the following operations: The presence of a start unit is checked, the start unit indicating the start of the bitstream and the start of decoding at the start of the bitstream.

23. The equipment according to claim 21, wherein, The at least one non-transitory memory and the computer program code are further configured, together with the at least one processor, to cause the device to perform at least the following operations: Further compress the compressed data associated with the information unit.

24. A method comprising: Encode or decode high-level bitstream syntax for at least one neural network; The high-level bitstream syntax includes at least one information unit, which has metadata or compressed neural network data as a part of the at least one neural network; and The serialized bit stream includes one or more information units from the at least one information unit; The information unit includes: a unit size that transmits signals of the size of the information unit; a unit payload carrying compressed or uncompressed data and related metadata related to the at least one neural network; and a unit header having information about the data type carried by the unit payload and related metadata. The unit payload includes at least one of the following: a set of parameters including global metadata and information about the at least one neural network; neural network topology information and related topology data; compressed or uncompressed neural network data, which may be complete or partial; quantized data; or a compressed neural network bitstream start indicator or payload data associated with a start code.

25. The method according to claim 24, wherein, The serialized bitstream is stored as a file in a virtual or non-virtual directory tree structure, or sent as a bitstream through a data pipeline.

26. The method according to claim 24, wherein, The portion of the at least one neural network is at least one of the following: a layer, a filter, a kernel, a bias, a quantization weight, a tensor, or any other data structure of at least one identifiable portion of the neural network.

27. The method according to claim 24, wherein, At least one of the topology data, the compressed or uncompressed neural network data unit payload, or the quantized data is divided into multiple information units.

28. The method according to claim 27, wherein, The neural network unit head includes information for indicating the segmentation.

29. The method according to claim 28, wherein, The information indicating the segmentation is represented by a counter that counts backwards to indicate the number of segments used.

30. The method according to claim 28, wherein, A counter value of 0 indicates that there is no partial information, and a counter value greater than 0 indicates the index of the partial information.

31. The method according to claim 27, wherein, The multiple information units have the same unique identifier in their respective unit headers to indicate the neural network topology element-level information to which the segment belongs.

32. The method according to claim 31, wherein, The unique identifier is the Khronos neural network exchange format NNEF variable identifier or label in the NNEF topology graph.

33. The method according to claim 27, wherein, The topology information includes Khronos neural network exchange format NNEF topology information.

34. The method according to claim 27, wherein, The plurality of information units have a flag in their unit header indicating whether such information unit is independently decodeable.

35. The method according to claim 24, wherein, The parameter set contains flags indicating the presence and carrying of topological units in the compressed neural network bitstream.

36. The method according to claim 24, wherein, The parameter set includes flags indicating whether sparsification is applied to the at least one neural network.

37. The method according to claim 24, wherein, The parameter set includes a sparsification performance graph data structure, which transmits a mapping between different sparsification thresholds and the resulting neural network inference accuracy.

38. The method according to claim 37, wherein, The obtained neural network inference accuracy corresponds to the performance of the at least one neural network in terms of output accuracy.

39. The method according to claim 24, wherein, The unit payload or header contains a quantization mapping data structure, the quantization mapping data structure signal transmission codebook, the codebook including a mapping between quantized values ​​and corresponding dequantized values.

40. The method of claim 24, wherein, The compressed neural network data is carried using a neural network exchange format.

41. The method according to claim 24, wherein, The compressed neural network data unit payload is mappable to the topology data via a reference or tag in the unit header associated with the compressed network data unit payload, wherein the reference or tag includes at least one of the following: a unique identifier, a unique string, or a relative or absolute Uniform Resource Identifier or Locator.

42. The method according to claim 24, wherein, The unit header indicates the type of neural network unit, which indicates the start of the serialized bitstream, wherein the serialized bitstream is a compressed or uncompressed neural network bitstream.

43. The method according to claim 24, wherein, The encoder provides the serialized bit stream to the decoder via a transmission channel.

44. The method of claim 24, wherein, The decoding of the high-level bitstream syntax includes: The unit is parsed by reading the size of at least one information unit, the unit header associated with the information unit, and the payload associated with the information unit; and Identify and parse at least one of the following associated with the information unit: topology data, quantization data, start code indicator data, parameter set data, or compressed or uncompressed data.

45. The method of claim 44, further comprising: The presence of a start unit is checked, the start unit indicating the start of the bitstream and the start of decoding at the start of the bitstream.

46. ​​The method of claim 44, further comprising: Further compress the compressed data associated with the information unit.

47. A machine-readable, non-transitory program storage device, tangibly embodying a program executable by the machine to perform operations, said operations including: Encode or decode high-level bitstream syntax for at least one neural network; The high-level bitstream syntax includes at least one information unit, which has metadata or compressed neural network data as a part of the at least one neural network; and The serialized bit stream includes one or more information units from the at least one information unit; The information unit includes: a unit size that transmits signals of the size of the information unit; a unit payload carrying compressed or uncompressed data and related metadata related to the at least one neural network; and a unit header having information about the data type carried by the unit payload and related metadata. The unit payload includes at least one of the following: a set of parameters including global metadata and information about the at least one neural network; neural network topology information and related data; compressed or uncompressed neural network data, which may be complete or partial; quantized data; or a compressed neural network bitstream start indicator or payload data associated with a start code.