Dynamic feature size adaptation in divisible deep neural networks
Patent Information
- Application Number
- JP2023544040
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-05
- Filing Date
- 2022-02-03
- Publication Date
- 2026-09-03
- Estimated Expiration
- 2042-02-03
Smart Images

Figure 0007914995000021 
Figure 0007914995000022 
Figure 0007914995000023
Abstract
Description
[Technical Field]
[0001] This embodiment generally relates to dynamic feature size adaptation in a divisible deep neural network (DNN). [Background technology]
[0002] Artificial intelligence is a crucial functional block in many technological fields today. This is due to the resurgence of neural networks, particularly in the form of deep neural networks (DNNs). Modern DNNs are often computationally intensive, making it difficult to run them on mobile phones or other edge devices with low processing power. This is often addressed by transferring data from the mobile device to a cloud server, where all the computations are performed. [Overview of the project]
[0003] According to one embodiment, a device is presented which comprises a Wireless Transmit / Receive Unit (WTRU), the WTRU comprising: a receiver configured to receive a portion of a deep neural network (DNN) model, which is located before a split point of the DNN model and includes a neural network for compressing features at that split point of the DNN model; one or more processors configured to obtain a compression factor of the neural network, determine in response to the compression factor which nodes in the neural network should be connected, and perform inference using the portion of the DNN model to construct the neural network and generate compressed features in response to the determination; and a transmitter configured to transmit the compressed features to another WTRU.
[0004] In another embodiment, a device is presented which comprises a wireless transmit / receive unit (WTRU), the wireless transmit / receive unit (WTRU) being a receiver configured to receive a portion of a deep neural network (DNN) model, which lies after a split point of the DNN model and includes a neural network for extending the features at that split point of the DNN model, and is also configured to receive one or more features output from another WTRU; and one or more processors configured to obtain a compression factor of the neural network, determine in response to the compression factor which nodes in the neural network should be connected, and in response to the determination, configure the neural network and perform inference using the portion of the DNN model, with the one or more features output from another WTRU as input to the neural network.
[0005] In another embodiment, a method is presented which includes a method carried out by a wireless transmit / receive unit (WTRU), the method comprising: receiving a portion of a deep neural network (DNN) model which lies before a split point of the DNN model and includes a neural network for compressing features at that split point of the DNN model; obtaining a compression factor of the neural network; determining which nodes in the neural network should be connected in response to the compression factor; configuring the neural network in response to the decision; performing inference using the portion of the DNN model to generate compressed features; and transmitting the compressed features to another WTRU.
[0006] In another embodiment, a method is presented which includes receiving a portion of a deep neural network (DNN) model, which is located after a split point of the DNN model and includes a neural network for extending the features at that split point of the DNN model; receiving one or more features output from another WTRU; obtaining a compression factor of the neural network; determining which nodes in the neural network should be connected in response to the compression factor; constructing the neural network in response to the decision; and performing inference using the portion of the DNN model, with the one or more features output from the other WTRU as input to the neural network.
[0007] Further embodiments include systems configured to carry out the methods described herein. Such systems may include a processor and a non-temporary computer storage medium that stores instructions which, when executed on the processor, operate to carry out the methods described herein. [Brief explanation of the drawing]
[0008] [Figure 1A] This is a system diagram showing an exemplary communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] This is a system diagram showing an exemplary wireless transmit / receive unit (WTRU) that may be used in the communication system shown in Figure 1A, according to one embodiment. [Figure 2] This paper presents a mechanism for distributed AI between two devices without feature size compression. [Figure 3A] These examples show DNNs with one, two, and three candidate partitions for feature compression, respectively. [Figure 3B] These examples show DNNs with one, two, and three candidate partitions for feature compression, respectively. [Figure 3C]DNNs having one, two, and three candidate splits for feature compression, respectively, are shown. [Figure 4] A DNN having a single split for feature compression is shown. [Figure 5A] A feature size compression mechanism for distributed AI between two devices, device 1 and device 2, using a bandwidth-reducer (BWR) and a bandwidth-expander (BWE) that support a single compression coefficient, is shown. [Figure 5B] A feature size compression mechanism that supports a plurality of compression coefficients is shown. [Figure 6A] Total inference latency without BWR and BWE is shown. [Figure 6B] Total inference latency with BWR and BWE is shown, where the size of intermediate data can be reduced. [Figure 7] A process of dynamically switching between splits and compression factor (CF) configurations according to an embodiment is shown. [Figure 8A] It is shown that device 1 and device 2 estimate their computing power and transmission channels. [Figure 8B] Reception of an AI / ML model from each of the devices is shown. [Figure 8C] Inference time operation of the device is shown. [Figure 9] A method with a single split in a DNN for adaptive feature compression according to an embodiment is shown. [Figure 10] An exemplary DySw capable of reducing and expanding an input of size 4 is shown. [Figure 11] Connections of the DySw configuration shown in Fig. 9 are shown. DETAILED DESCRIPTION OF EMBODIMENTS
[0009] Figure 1A shows an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, message transmission, and broadcast to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may use one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique-word OFDM (UW-OFDM), resource block filtering OFDM, and filter bank multicarrier (FBMC).
[0010] As shown in Figure 1A, the communication system 100 may include radio transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104, CN 106, public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it will be understood that the disclosed embodiments assume any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a radio environment. For example, WTRU102a, 102b, 102c, and 102d, any of which may be referred to as “station” and / or “STA”, may be configured to transmit and / or receive radio signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscriber-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, radio sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other radio devices operating in an industrial and / or automated processing chain context), consumer electronics devices, and devices operating in commercial and / or industrial radio networks. Any of WTRU102a, 102b, 102c, and 102d may interchangeably be referred to as UE.
[0011] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks such as CN 106, the Internet 110, and / or other networks 112. As an example, base stations 114a and 114b may be base transceiver stations (BTS), node B, eNodeB, home node B, home eNodeB, gNB, NR nodeB, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are shown as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.
[0012] Base station 114a may be part of RAN 104, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), and relay nodes. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be licensed spectra, unlicensed spectra, or a combination of licensed and unlicensed spectra. Cells may provide coverage of radio services to a particular geographic area that may be relatively fixed or change over time. Cells may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver per sector of the cell. In one embodiment, the base station 114a may use multiple-input multiple output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0013] Base stations 114a and 114b may communicate with one or more WTRUs 102a, 102b, 102c, and 102d via an air interface 116, which may be any suitable radio communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).
[0014] More specifically, as described above, the communication system 100 may be a multiple access system and may use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 114a of RAN 104 and WTRU 102a, 102b, and 102c may implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) and may establish an air interface 116 using wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Advanced HSPA (HSPA+). HSPA may include High-Speed Downlink Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0015] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish an air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0016] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as NR radio access, which can establish an air interface 116 using New Radio (NR).
[0017] In one embodiment, base station 114a and WTRU 102a, 102b, 102c may implement multiple radio access technologies. For example, base station 114a and WTRU 102a, 102b, 102c may implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Thus, the air interface utilized by WTRU 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNB and gNB).
[0018] In other embodiments, base stations 114a and WTRUs 102a, 102b, and 102c may implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity, WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access, WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0019] The base station 114b in Figure 1A may be, for example, a wireless router, home node B, home e-node B, or access point, and may utilize any suitable RAT to facilitate wireless connectivity in local areas such as offices, homes, vehicles, campuses, industrial facilities, aerial corridors (for use by drones), roads, etc. In one embodiment, the base station 114b and WTRU 102c, 102d may implement wireless technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and WTRU 102c, 102d may implement wireless technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base stations 114b and WTRUs 102c, 102d may establish picocells or femtocells using cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in Figure 1A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not need to access the internet 110 via CN 106.
[0020] RAN104 may communicate with CN106, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have various quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 may provide call control, billing services, mobile location-based services, prepaid calls, internet connectivity, video distribution, etc., and / or high-level security functions such as user authentication. Although not shown in Figure 1A, it will be understood that RAN104 and / or CN106 may communicate directly or indirectly with other RANs using the same RAT or different RAT as RAN104. For example, in addition to being connected to RAN104 which may utilize NR radio technology, CN106 may also communicate with another RAN (not shown) using GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0021] CN106 may also function as a gateway to WTRU102a, 102b, 102c, and 102d for access to PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a public switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices, which use common communication protocols such as the transmission control protocol (TCP), the user datagram protocol (UDP), and / or the Internet protocol (IP) of the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs that may use the same RAT as RAN104 or a different RAT.
[0022] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multimode capability (for example, WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different radio networks via different radio links). For example, WTRU 102c shown in Figure 1A may be configured to communicate with base station 114a, which may use cellular-based radio technology, and base station 114b, which may use IEEE 802 radio technology.
[0023] Figure 1B is a system diagram showing an exemplary WTRU102. As shown in Figure 1B, the WTRU102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU102 may include any partial combination of the aforementioned elements while maintaining consistency with one embodiment.
[0024] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to a transceiver 120 which may be coupled to a transmit / receive element 122. Figure 1B shows the processor 118 and transceiver 120 as separate components, but it will be understood that the processor 118 and transceiver 120 may be integrated together in an electronic package or chip.
[0025] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of radio signals.
[0026] Although the transmit / receive element 122 is shown as a single element in Figure 1B, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may utilize MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals via the air interface 116.
[0027] The transceiver 120 may be configured to modulate the signal transmitted by the transmit / receive element 122 and demodulate the signal received by the transmit / receive element 122. As described above, the WTRU 102 may have multimode capability. Therefore, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0028] The processor 118 of the WTRU102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and may receive user input from these. The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and may store data in memory. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 may access information from memory not physically located on the WTRU 102, such as on a server or home computer (not shown), and store data in memory.
[0029] The processor 118 may receive power from the power supply 134, but may also be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 may be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, a fuel cell, etc.
[0030] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 may acquire location information by any preferred location determination method while maintaining consistency with one embodiment.
[0031] The processor 118 may be further coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functions, and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, and the like. The peripheral device 138 may include one or more sensors, which may be one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, compass sensor, proximity sensor, temperature sensor, time sensor, geolocation sensor, altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.
[0032] WTRU102 may include a full-duplex radio in which the transmission and reception of some or all of the signals (e.g., associated with specific subframes for both UL (e.g., transmission) and downlink (e.g., reception) may be in parallel and / or simultaneous. The full-duplex radio may include an interference management unit for reducing and / or substantially eliminating self-interference via hardware (e.g., chokes) or signal processing via a processor (e.g., via a separate processor (not shown) or processor 118). In one embodiment, WRTU102 may include a half-duplex radio for the transmission and reception of any of the signals (e.g., associated with specific subframes for either UL (e.g., transmission) or downlink (e.g., reception)).
[0033] Although the WTRU is described as a wireless terminal device in Figures 1A and 1B, in certain representative embodiments, such a terminal device is expected to be able to use a wired communication interface with a communication network (for example, temporarily or permanently).
[0034] With regard to Figures 1A-1B and the corresponding descriptions therein, one or more of the functions described herein with respect to one or more of the WTRU102a-d, base stations 114a-b, eNode-B160a-c, MME162, SGW164, PGW166, gNB180a-c, AMF182a-b, UPF184a-b, SMF183a-b, DN185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0035] Emulation devices may be designed to implement testing of one or more other devices in a laboratory and / or operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless network to test other devices in a communications network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless network. Emulation devices may be directly coupled to another device for testing purposes and / or may perform testing using terrestrial radio communication.
[0036] One or more emulation devices may perform one or more functions, including all of the above, while not implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test laboratory test scenario, and / or in a wired and / or wireless communication network that is not deployed (e.g., for testing purposes), to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) may be used by the emulation device to transmit and / or receive data.
[0037] As mentioned above, DNN operation is often handled by transferring data from mobile devices to cloud servers, where all computations are performed. However, this requires bandwidth, is time-intensive (due to transmission latency), and raises data privacy concerns. One way to solve this is to perform all computations on the user device (e.g., a mobile phone) via a lightweight, low-precision DNN. Another method is to use a high-precision DNN but share the computations across one or more mobile devices and / or the cloud.
[0038] Flexible AI methods Model compression techniques are widely used to run DNN models only on user devices. They reduce the model memory footprint and runtime, making them adaptable to specific devices. However, there are cases where the device on which the model will run is unknown in advance, and even when the device is known, its available resources may change over time due to other processes, for example. To overcome these problems, a family of so-called flexible AI models has recently been proposed. These models can instantly adapt to available resources, for example, by enabling early classification termination, model width adaptation (slimming), or switchable model weight quantization.
[0039] Distributed AI method Some so-called distributed AI methods partition the model between two or more devices (i.e., WTRUs), or between a device and the cloud / edge. For example, Figure 2 shows a mechanism for distributed AI between two devices, device 1 and device 2, without feature size compression. In distributed AI, intermediate data (features) that can be of very high dimension need to be transmitted. This adds latency to processing and is not always possible due to bandwidth limitations of the corresponding transmission network. To overcome this problem, methods have been proposed to reduce feature size through a bottleneck. Figure 3A shows a DNN with one candidate partition for feature compression, where a1, a2, or a3 can be used as the partition point. Figure 3B shows a DNN with two candidate partitions (e.g., a1 and a2) for feature compression. Figure 3C shows a DNN with three candidate partitions (e.g., c1, c2, and c3) for feature compression.
[0040] Without introducing any limitations, features can be considered as individual measurable properties or characteristics of data that can be used to represent a phenomenon. One or more features may be associated with a machine learning algorithm, a neural network, and / or one of its layers, as inputs and / or outputs. For example, features may be organized as vectors. For instance, features associated with a wireless use case may include time, transmitter identity, and measurements on a reference signal (RS).
[0041] For example, features associated with the algorithm used to process positioning information may include values associated with positioning RS (PRS) measurement, quantities such as Reference Signal Receive Power (RSRP), quantities such as Reference Signal Receive Quality (RSRQ), quantities related to Received Signal Strength Indication (RSSI), quantities related to time difference measurement based on signals from separate sources (e.g., in the case of time-based positioning methods), quantities related to angle of arrival measurement, quantities related to beam quality, and / or values associated with output from sensors (WTRU rotation, imaging from a camera, etc.).
[0042] For example, features associated with algorithms used to process Channel State Information (CSI) may include measurements of quantities associated with reception, such as Channel State Reference Signal (CSI-RS), Synchronization Signal Block (SSB), Precoding Matrix Indication (PMI), Rank Indicator (RI), Channel Quality Indicator (CQI), RSRP, RSRQ, and RSSI.
[0043] For example, features associated with algorithms used to process beam management and selection may include quantities associated with measurements similar to those used to process one or more parameters related to CSI, Transmit / Receive Point (TRP) identification information (identity, ID), beam ID, and / or Beam Failure Detection (BFD) (e.g., determining a threshold for sufficient beam quality).
[0044] Similarly, any method described herein may be applied to specific stages of AI / ML processing, for example, hyperparameters used for machine learning algorithms for training or inference, or may further include specific parameter settings for such purposes.
[0045] Figure 4 shows a DNN with a single partition (a2 and b2, respectively) for feature compression, where the feature size is reduced from (a) 4 to 2 and (b) 4 to 3. Specifically, (a3) is a subnetwork that achieves the feature size reduction from 4 to 2, and (b3) is a subnetwork that achieves the reduction from 4 to 3. In existing research, the DNN is trained from scratch with feature compressors and expanders for each compression factor. Note that the compression factor is the ratio of the feature size at the output of the compressor to the feature size at the input to the compressor. This means that whenever the compression factor needs to be changed, the device and the cloud server must tune and download a new model from the cloud server. Figure 5A shows a feature size compression mechanism for distributed AI between two devices, device 1 and device 2, using a bandwidth reducer (BWR, 510) and a bandwidth expander (BWE, 520) that supports a single compression factor. However, these methods cannot adapt bottlenecks to different transmission network bandwidths.
[0046] To provide flexibility in the distributed AI paradigm, we introduce a Flexible and Distributed AI (FD-AI) approach. The proposed approach is distributed because the DNN can be partitioned across two or more devices. Furthermore, the proposed approach is flexible because it can select a partition point from several possible candidates depending on the available resources within the devices. In addition, the feature size transmitted at each partition point can be compressed to fit the available network bandwidth for transmission.
[0047] In one embodiment, we propose a switchable bottleneck subnetwork that is part of a DNN architecture. The bottleneck subnetwork is switchable because it can adapt to different transmit network bandwidths during inference. In the proposed design, we have one bottleneck subnetwork with layers for reducing feature size and another set of layers for restoring the feature size to its original size. These bottleneck subnetworks can be incorporated into one or more partition locations of any existing DNN. For brevity, in the following description, we consider a DNN having a single partition with one set of bottleneck subnetworks for feature size reduction and expansion.
[0048] In one embodiment, the first device may be either an edge device or a cloud server, and the second device may be either an edge device or a cloud server. More generally, the methods described herein may be applicable to any device that exchanges data over a communication link. Such a device may include processing of segmented neural network or autoencoder functions. The methods described herein may be applicable to processing in a device for, for example, end-user applications (e.g., audio, video, etc.) or functions related to processing for data transmission and / or reception. More generally, such a device may be a mobile terminal, a radio access network node such as a gNB, etc. Such a communication link may be a radio link and / or interface such as a 3GPP Uu, 3GPP sidelink, or Wifi link.
[0049] The DNN layers up to the split point, which have a feature size reduction layer in the bottleneck subnetwork, are loaded into the first device. The remaining portion, i.e., the bottleneck subnetwork expander and the rest of the DNN after the split point, are loaded into the second device. We refer to the bottleneck subnetwork, including the reducer and expander, as a Dynamic Feature Size Switch (DySw). Features sent to the second device are extracted in the middle of the DySw. We call the DNN that achieves this a Dynamic Switchable Feature Size Network (DyFsNet). DyFsNet is generally applicable to any DNN architecture, such as convolutional neural networks (CNNs), and is novel in design and training. Inference in DyFsNet is simple and tunable (with respect to the split point and available network bandwidth).
[0050] Figure 5B shows an example of a feature size compression mechanism that uses a bandwidth reducer (BWR) and a bandwidth expander (BWE) to support multiple compression factors for distributed AI between two devices, device-1 and device-2, where K1, K2, ..., K N This specifies the compression factor within the trainable BWR(530) and BWE(540), which can be exclusively and dynamically switched during inference.
[0051] More specifically, devices 1 and 2, optionally together with a server, monitor channel conditions and device states and select compression factors and feature sizes at the splitting points. Device 1 receives the first portion of the DNN model up to the splitting point, and device 2 receives the remaining portion of the DNN model. Device 1 performs inference to compute features from the input, which are then compressed by the BWR. Different compression factors can be obtained by controlling which nodes are connected in the BWR(530), as will be explained in more detail in relation to Figures 10 and 11. Device 2 receives the compressed features and expands them by the BWE(540). Similar to the BWR, the BWE's compression factor can be controlled by controlling the node connections within the BWE. Device 2 then continues inference and provides the final output.
[0052] Network bandwidth limitations introduce additional latency to the entire inference process. Figure 6A shows the total inference latency without BWR and BWE. Figure 6B shows the total inference latency with BWR and BWR, where the size of intermediate data can be reduced.
[0053] As described above, we propose a method to reduce the size of intermediate data at different points within a DNN model in order to limit throughput requirements on the communication network while maintaining near-perfect prediction accuracy. Figure 7 shows a process for dynamically switching between partitioning / compression factor (CF) configurations according to one embodiment.
[0054] During the model training and partition / CF estimation phase (710), the DyFsNet model is trained for different partitions and CFs. This can now be done offline on a cloud server. The trained model is stored on the cloud server and available for download to the device. The (server-side) orchestrator manages the selection of trained models and the coordination of transmission to the end device on request. Here, we assume that bandwidth information is available. Based on this, the CF is estimated as the ratio of feature size to available bandwidth.
[0055] For example, an orchestrator or external control system determines the partitioning locations of the DNN based on the computing power of the end devices (e.g., device-1 and device-2). This information is communicated to the devices that load the DNN for processing according to the partitioning information.
[0056] In the model deployment phase (720), the trained partitioned models are received by the device. Once received, they are loaded into the device for inference.
[0057] The status of the network (e.g., bandwidth) and / or devices (e.g., available processing power) is monitored (730). The devices monitor the network channels between them and coordinate CF between them. This is done without the involvement of a server.
[0058] Once a consensus is reached between devices, CF selection (740) is performed, which therefore affects the feature size at the partition location. Note that the available CF options depend on the number of channels in the filter of the DNN layer in which the partition is implemented. Typically, the CF is selected to approximate the available bandwidth, not to precisely match it.
[0059] The partitioned model inference is performed on a first device and a second device (750). For example, the first device computes intermediate features using the DNN up to the partition, compresses the features, and transfers the compressed features to the second device. The second device receives the compressed features, does not compress them, and continues the DNN inference. In one embodiment, where the device is a wireless terminal device and / or the communication link of the device is a wireless air interface (e.g., NR Uu, sidelink, etc.), the device can perform at least one of the following: -Initiate the adaptations proposed herein. For example, a device may adapt the partitioning processing points, feature dimensions, compression coefficient, inference latency, processing requirements, accuracy of functionality, or any other aspects proposed herein. - The device may trigger such a fit for AI processing when it determines at least one of the following in relation to the L1 / physical (PHY) layer operation: ○The device may determine that a change in radio characteristics has occurred, where such characteristics may affect the transmitted data rate through the interface, such as a change in cell identity, a change in carrier frequency, a change in bandwidth part (BWP), a change in the number of BWP and / or physical resource blocks (PRB) of the cell, a change in subcarrier spacing (SCS), a change in the number of aggregate carriers available for transmission, a change in available transmit power, or a change in metering. The device may determine that a change in operating state has occurred via the wireless interface, such as a change in the control channel resource (CORESET) or identity, where a first identity may be associated with a first threshold, and a second identity may be associated with a second threshold. The device may determine that a change exceeds a certain possible configured threshold indicating a degradation of channel quality and may implement a modification that reduces the data rate associated with AI processing. Conversely, the device may determine that the radio condition is improving and may implement a modification that increases the data rate associated with AI processing. For example, this can be applied to physical layer functions of devices such as CSI autoencoding. - The device may trigger such a suitability for AI processing if it determines at least one of the following regarding L2 / Medium Access Control (MAC) layer operation: The device may determine that a change has occurred in data processing or information bearers (e.g., data radio bearers, signaling radio bearers), and such characteristics may affect the transmitted data rate via interfaces available for AI processing, such as logical channel prioritization parameters (e.g., Packet Delay Budget (PDB), Prioritized Bit Rate (PBR), changes in TTI duration / numerology, changes in associated QoS flow IDs, and mapping restrictions to sets of resources that enable different data rates). The device may determine that a change exceeds a certain possible configured threshold indicating a decrease in the data rate available for AI processing, and may implement a modification that reduces the data rate associated with the AI processing. Conversely, the device may determine that the available data rate is increasing, and may implement a modification that increases the data rate associated with the AI processing. For example, this can be applied to system-level functions such as the positioning capabilities of a device. For example, this could be applied to DRBs associated with a particular data radio bearer (DRB) and / or DRB type, e.g., DRBs associated with a particular AI-enabled application where a change in the DRB or its characteristics may trigger the adaptation of AI-based processing in the associated application layer. - The device may trigger such a suit for AI processing if it determines at least one of the following with respect to L3 / Radio Resource Control (RRC) layer operation: The device may determine that a configuration change has occurred that affects one or more of the L1 / L2 configurations, such as the embodiments described above that may change the available data rate. The device may, for example, determine that it has received a reconfiguration message for mobility and / or that it should be applied (for example, for a conditional handover command), where the message may include instructions for an applicable data rate for AI processing and / or its associated radio bearer. ○The device can determine that a radio link failure (RLF) or other radio link failure has occurred. The device may determine that a change exceeds a certain possible configured threshold indicating a decrease in the data rate available for AI processing and may implement a modification that reduces the data rate associated with AI processing. Conversely, the device may determine that the available data rate is increasing and may implement a modification that increases the data rate associated with AI processing. Alternatively, it may determine that the event itself may be associated with an increase (e.g., adding a cell to the device's configuration connectivity, e.g., dual connectivity) or a decrease (e.g., removing a cell to the RLF and / or the device's configuration connectivity) in the data rate available for AI processing. - The device may trigger such a suitability for AI processing when it determines at least one of the following in relation to the available processing resources. The device may determine that a change in available hardware processing has occurred, for example, based on a change in the number of instantiated and / or active AI processes, a change in dynamic device capabilities, or a change in processing requirements for AI processing (e.g., inference latency, accuracy). ○ A device may determine that a change in its power state has occurred. For example, a device may determine that it has transitioned from a first state to a second state, where such a state may relate to an RRC connectivity state (IDLE, INACTIVE, or CONNECTED), a DRX state (active, inactive), or a different configuration thereof. The device may determine that a change exceeds a certain possible configured threshold indicating a decrease in available processing resources. Conversely, the device may determine that there is an increase in available processing resources and implement adaptations that may increase the data rate associated with the AI processing. Similarly, a particular state may be associated with a specific AI processing level, split point configuration, and / or data rate. - A device may trigger such a fitting for AI processing when it decides to receive control signaling according to at least one of the following: The device may receive control information indicating either an increase or decrease in the data rate available for AI processing / AI processing. This may be implicitly based on signaled values and / or modifications of control channel properties of values as described above for L1, L2, L3 processing and / or power saving management, or it may be explicitly based on instructions in control messages. Such control information may be received in L1 signals, L1 messages, e.g., DCI on PDCCH, in L2 MAC control elements, or in RRC messages. ○ The control information may include specific split point configurations applied to a given AI processing, hyperparameter settings, target resolution, target accuracy, target feature vector, etc.
[0060] Figures 8A, 8B, and 8C provide further diagrams of the process. Figure 8A shows devices 1 and 2 (840, 860) estimating their computing power and transmission channels (850). These estimates are communicated to the operator / edge / cloud (820, 830) to request a suitable AI / ML model (810).
[0061] Figure 8B shows the reception of AI / ML models from each device. The operator / cloud / edge selects a model and transmits it over the network (830), and the requested model is received by devices 1 and 2.
[0062] Figure 8C illustrates the inference time operation of the device. Device-1 computes features and then, based on channel conditions, sends feature sizes of appropriate dimensions to Device-2. Device-1 performs inference on the input data (870). The input data may be one or more images from device memory, or images captured live from the device's camera, or audio data on device memory, or audio data captured live from the device's microphone, or any other data that needs to be processed by the DNN. Device-1 outputs an intermediate or initial output (880) processed by the DNN, such as in the case of an MSDNet type DNN. Information necessary for further processing of the features is also communicated to Device-2 via channel (850). Device-2 receives the features and further continues inference and CF switching as needed, providing a final output (890). Additionally, Device-1 sends the features along with control information to Device-2 for further processing. Device-2 receives the features and control information and continues inference.
[0063] Figure 9 illustrates the proposed method of using a single partition in a DNN for feature compression. Figure 9(a) shows a co-trained subnetwork DySw(a3) with no selected compression coefficients. Figure 9(b) shows a co-trained subnetwork (b3) with selected feature compression coefficients of 4:2. Figure 9(c) shows a co-trained subnetwork (c3) with selected feature compression coefficients of 4:3. Note that the DNNs in Figures 9(a), 9(b), and 9(c) are the same (single) DNN.
[0064] DySw can be trained together with the entire DNN. Alternatively, a DNN without DySw can be pre-trained, and the DySw subnetwork can be added. Note that in this alternative solution, the pre-trained DNN is extended with the DySw(a3) subnetwork, and the training is for DySw only, while keeping the pre-trained DNN (and its weights) invariant (i.e., fixed).
[0065] As shown in Figure 9, DySw can be reconfigured to fit multiple compression factors. Reconfiguration is achieved through the connection details of the DySw nodes. For example, in the case of a DySw subnetwork as shown in Figure 10, we can maintain a 4x3 matrix specifying the node connections as shown in Figure 11. Each element in the matrix (E ij The ) indicates whether input node i is connected to output node j, where "0" means disconnected and "1" means connected. The matrices shown in Figures 11(a), 11(b), and 11(c) correspond to Figures 9(a), 9(b), and 9(c), respectively. Specifically, Figure 9(a) specifies that no input nodes are connected to any output nodes, Figure 9(b) specifies that only two of the output nodes (output nodes-2 and 3) are connected to the input nodes, and Figure 9(c) specifies that all input nodes are connected to the outputs. Figure 11 shows the connections on the reducer side, and the expander can maintain matrices corresponding to different compression coefficients. In one embodiment, the shape of the expander-side matrix is transposed (relative to that of the reducer-side), but the number of rows with all zeros remains the same.
[0066] As shown in Figure 8, the device adjusts the CF. In one embodiment, an orchestrator or external control system informs device-1 about the available bandwidth. Based on the bandwidth information, device-1 determines which CF to use. Device-1 then switches the DySw to achieve feature size compression corresponding to the determined CF. Device-1 can also communicate which CF it is using, and so device-2 switches its side of the DNN to match the communicated information.
[0067] In one embodiment, after a CF is selected, device-1 determines which connections between nodes should be disabled to provide the selected CF, and device-2 also determines which connections should be disabled accordingly to properly implement the expansion. The CF determines how many output nodes will be connected to the input nodes, but the method and number of these connections are determined through learning.
[0068] As described above, Figure 10 shows an exemplary DySw that can reduce and expand an input of size 4. While Figure 10 shows a single-layer “reduction” block for simplification, it should be noted that reductions are not limited to a single layer. The illustrated DySw can perform 4:3, 4:2, and 4:1 compressions, as well as corresponding expansions (i.e., 1:4, 2:4, and 3:4). A DySw design may have additional layers as needed, such as a BatchNorm layer for better training. Here, only reductions (BWR shown to the left of the dotted line) and expansions (BWE shown to the right of the dotted line) are shown. Nonlinearity is inherent in the layers. A BatchNorm layer may be an optional layer required for efficient training and is therefore not shown here.
[0069] More generally, a typical DySw comprises four types of layers: feature dimensionality reducer and expander layers, a nonlinear layer, and a batch normalization (BatchNorm) layer. Of these layers, the BatchNorm layer is optional. A simple DySw is shown in Figure 10.
[0070] The DySw used in DNN classifiers can be trained using conventional task-specific losses, such as cross-entropy loss for classification tasks or mean squared error loss for regression tasks. The DySw can be used for any task, i.e., classification, detection, or segmentation, and in any DNN architecture, i.e., CNNs, GANs, autoencoders, etc. Training a DySw involves learning the reducer-expander layer weights and the parameters of the batch normalization layer (also referred to as "BatchNorm"). BatchNorm is used for faster convergence of training.
[0071] DySw training allows for additional constraints on the loss objective. As an example, we show the addition of reconstruction loss across DySw. Reconstruction loss penalizes the parallax between the input to DySw and the output to DySw. DySw is an auxiliary, optional entity that can be added to a trained DNN.
[0072] In DySw, the reduction factor can be switched on the fly during inference. In DyFsNet, the training iterations are modified to co-learn shared DySw weights using multiple reduction factors, as will be further detailed below.
[0073] DySw training can be offline or online, performed on the cloud / operator / edge, or it can be federated training on the device. Here, we describe the architecture and training of a partitioned DNN in the case of a single partition between two devices having DySw. The training mechanism described herein can be extended to multiple partition cases. Below, we describe in detail the architecture of the partitioned DNN, the DySw layer and the architecture of DyFsNet (DNN with a DySw layer), as well as different loss functions and their training.
[0074] Consider a split at the end of the l-th layer, where Device-1 processes up to layer l, and Device-2 processes from layer l+1 onwards. Let a part of the DNN in Device-1 be h device1 , and similarly let a part of the DNN in Device-2 be h device2 . The input to the DNN can be any type of data, but herein, let the input X be a color image such that X∈R {W×H×3} , where W and H are the width and height, respectively, and 3 represents the number of color channels (e.g., RGB). The feature tensor (or simply feature) in the split is y l ∈R {M×N×C} , where M, N, and C represent the width, height, and number of channels thereof, respectively. The feature y l is transmitted to Device-2 via a wireless network, and Device-2 takes y l as input and generates an output Y. Therefore, y l = h device1 (X), and Y = h device2 (y l ).
[0075] DySw is a subnetwork represented by h DySw . The parameter of h DySw is θ DySw . If the reducer (first part) and expander (second part) of DySw are referred to as BWR and BWE, exemplary implementations of such a reducer and expander can include a convolutional layer, a non-linear layer (ReLu), and a batch normalization layer (BatchNorm), as summarized below.
[0076]
Mathematical Formula
[0077] A DNN comprising DySw is referred to as DyFsNet. DyFsNet is represented as h. Let θ be the parameter of h. The subnetwork of DyFsNet before the split point is
[0078]
Mathematical Formula
[0079]
number
[0080] DySw switches between various compression factors (CFs) for feature size. CF switching is indexed by K. The K-indexed intermediate outputs in the DyFsNet partition are as follows:
[0081]
number
[0082]
number
[0083]
number
[0084]
number
[0085]
number
[0086] The configuration offers two types of monitoring, one of which is ground truth labeling.
[0087]
number
[0088]
number
[0089] A DyFsNet trained from the start:
[0090]
number
[0091] DyFsNet trained using pre-trained initialization:
[0092]
number
[0093] A multi-partition DyFsNet trained from the start:
[0094]
number
[0095] Multi-split DyFsNet trained from pre-trained initialization:
[0096]
Math
[0097] DyFsNet Training Algorithm (X i , Y i ) ∈ D is the dataset, wherein X i and Y i are data and the supervision thereof, respectively, i∈{0, 1, ..., N} is an index, N is the number of training samples, and Num-of-epochs is the number of training epochs. Herein, we provide a training algorithm for a classifier that uses global loss, i.e., cross-entropy and KD. KD-based loss can be of four types through distillation from: i) the output of DySw without compression (i.e., DySw with K=1), ii) the output of DySw with a current lower compression factor (i.e., distillation from DySw with K=K1 to DySw with K=K2, where K1<K2), iii) an affine combination of the uncompressed DySw output and the closest compressed DySw output(s), or iv) the output of a completely different DNN architecture that has been sufficiently trained for the same task.
[0098] The overall algorithm is as follows. a. Calculate the loss of DyFsNet for the uncompressed configuration of DySw. In our example, this is, but not limited to, cross-entropy loss. b. Perform backpropagation and accumulate gradients for the uncompressed configuration of DySw. c. Select N r numbers of CF in the range from 1 to C, where 1 represents no compression and C represents maximum compression. d. When CF=2 to N r : i. Calculate the DyFsNet loss for distillation type (i), (ii), (iii), or (iv). ii. Perform backpropagation and accumulate the gradient of DySw. e. Update the weights using the accumulated gradient.
[0099] In one example, the following pseudocode is used.
[0100] KD from uncompressed (K=1) DySw output:
[0101] [Table 1]
[0102] KD is the output of a DySw with K=K1 to a DySw with K=K2, where K1 <K2である:
[0103] [Table 2]
[0104] KD is the output of a DySw with K=K1 to a DySw with K=K2, where K1 <K2である:
[0105] [Table 3]
[0106] We tested the proposed idea for an image classification task using the well-known MSDNet model. This model has several CNN blocks that can perform classification on the output of any block. We want to split this large network at different block ends and send the corresponding features to a second device (or cloud). Table 1 shows the feature dimensions of the MSDNet at the end of each block in the ImageNet dataset.
[0107] [Table 4]
[0108] Here, we demonstrate the usefulness of feature size reduction through an example of data rate requirements in a typical DNN. The data rate required to transmit features corresponding to a single 224×224×3 image generated by a DNN used for image classification (MSDNet) ranges from 13 Mbps to 0.5 Gbps. This is a challenging data rate for transmission over wireless networks. In a preliminary implementation of our approach using the MSDNet model, we were able to reduce feature size by 50% with a maximum accuracy loss of 1%.
[0109] Below, we describe our implementation of DySw on MSDNet for CIFAR-100, where the DNN is divided into seven locations, and the feature size at each division location (each unit being 16 bits) is shown in Table 2. We achieved compression factors of 1, 2, 4, and 10.
[0110] [Table 5]
[0111] This study investigates the effects of adding bandwidth reducers and expanders to MSDNet. Table 3 shows the results for the baseline (no bandwidth reducers / expanders) and for reduction factors of 1, 2, 4, and 10 with bandwidth reducers / expanders. Reduction factors 1, 2, 4, and 10 correspond to 100%, 50%, 25%, and 10% of the original bandwidth, respectively. It can be seen that the accuracy of the bandwidth-reduced MSDNet is almost the same as the baseline MSDNet with no reduction. Note that the accuracy is for all six blocks (0-6) and for the compression implementation at the end of all scales. In other words, by adding a new bandwidth reducer / expander at each split point, features can be significantly reduced to support feature transmission, while classification accuracy remains largely unchanged.
[0112] [Table 6]
[0113] Methods for switchable precision networks that reference the precision of CNN weights exist. There has also been research on switchable multiwidth CNNs. However, unlike those, we propose a switchable feature bandwidth network that can switch between different feature bandwidths at inference time. This switchability is useful for addressing bandwidth constraints of communication channels between devices, device clouds, or other combinations thereof. For example, this mechanism can be used independently of the CNN architecture and can be used seamlessly with existing models performing different machine learning tasks such as ResNet, AlexNet, DenseNet, SoundNet, and VGG. This mechanism can also be used independently of other types of feature compression techniques such as weight quantization.
[0114] The proposed method deals with efficient bandwidth for transmission for distributed AI, with provisions for switching between multiple feature bandwidths. During distributed inference at edge devices, each device only needs to load a portion of the AI model once, but the input / output features communicated between them can be flexibly configured according to the available transmission bandwidth by enabling / disabling connections between nodes in DySw. Other parameters of the DNN remain the same when several nodes are connected or disconnected to achieve the desired compression factor. That is, the same DNN model is used for different compression factors, and there is no need to download a new DNN model to fit the compression factor or network bandwidth.
[0115] AI processing can be used, for example, on images captured by a basic mobile phone camera, or on images captured by a smart TV camera for UI interaction via gesture detection, but is not limited to these. The proposed method can be used in a variety of scenarios. For example, the AI model can be partitioned between the device and the cloud. Below, we list some possible use scenarios. 1. The AI model is split between two devices. For example, a user might want to process data captured on a smartwatch, where some of the processing is done on the watch and the rest on the user's mobile phone. 2. The AI model is divided across multiple devices and, in some cases, the cloud. For example, a user might want to process the feed from a smart CCTV camera quickly on the camera itself, and then handle the detailed processing in the cloud or on a local server. 3. Similar to Use Case 3, but with speech / audio processing using a computationally capable microphone instead of a CCTV camera. 4. Share the processing of medical data in the diagnostic room and the cloud. 5. A terminal device capable of communicating via a wireless link, wherein AI processing relates to the transmission and / or reception functions of a wireless processing chain (e.g., CSI compression, CSI autocoding, positioning determination, etc.). 6. A terminal device capable of communicating via a wireless link, wherein AI processing relates to scheduling or data processing functions, for example, to QoS processing (e.g., user plane data rate adaptation).
[0116] Various numerical values are used in this application. Specific values are provided for illustrative purposes only, and the embodiments described are not limited to these specific values.
[0117] While features and elements are described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented in computer programs, software, or firmware embedded on computer-readable media for execution by a computer or processor. Examples of non-temporary computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with software can be used to implement video encoders, video decoders, or both, or radio frequency transceivers for use in UEs, WTRUs, terminals, base stations, RNCs, or any host computer.
[0118] Furthermore, the embodiments described above include other devices, including processing platforms, computing systems, controllers, and processors. These devices may include at least one central processing unit ("Central Processing Unit, CPU") and memory. According to the convention of those skilled in the art in the field of computer programming, references to operations and symbolic representations of arithmetic or instructions may be performed by various CPUs and memories. Such operations and arithmetic or instructions may be referred to as "execution," "computer execution," or "CPU execution."
[0119] Those with ordinary art in the art will understand that operations and symbolically represented arithmetic or instructions involve the manipulation of electrical signals by the CPU. The electrical system represents data bits that can cause a resulting transformation or reduction of electrical signals, and maintains these data bits in memory locations in the memory system, thereby reconfiguring or otherwise modifying the CPU's operations and processing of other signals. The memory locations where the data bits are maintained are physical locations having specific electrical, magnetic, or optical properties that correspond to or represent the data bits. It should be understood that exemplary embodiments are not limited to the platforms or CPUs described above, and other platforms and CPUs may support the methods provided.
[0120] Data bits may also be maintained on computer-readable media, including magnetic disks, optical disks, and any other volatile (e.g., Random Access Memory ("RAM")) or CPU-readable non-volatile (e.g., Read-Only Memory ("ROM")) mass storage systems. The computer-readable media may include cooperative or interconnected computer-readable media distributed among multiple interconnected processing systems, which may reside exclusively on a processing system or be local or remote to the processing system. Typical embodiments are not limited to the memory described above, and it is understood that other platforms and memories may support the methods described.
[0121] In exemplary embodiments, any of the operations, processes, etc., described herein may be implemented as computer-readable instructions stored on a computer-readable medium. These computer-readable instructions may be executed by processors in mobile devices, network elements, and / or any other computing devices.
[0122] The use of hardware or software is generally (though not always, in certain situations the choice between hardware and software can be significant) a design choice involving a cost-effectiveness trade-off. Various vehicles (e.g., hardware, software, and / or firmware) may exist in which the processes and / or systems and / or other technologies described herein may be effective, and the preferred vehicle may vary depending on the context in which the processes and / or systems and / or other technologies are deployed. For example, if the implementer determines that speed and accuracy are paramount, the implementer may choose a primarily hardware and / or firmware vehicle. If flexibility is paramount, the implementer may choose a primarily software implementation. Alternatively, the implementer may choose any combination of hardware, software, and / or firmware.
[0123] The detailed description above illustrates various embodiments of devices and / or processes through the use of block diagrams, flowcharts, and / or examples. Those skilled in the art will understand that, insofar as such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, each function and / or operation in such block diagrams, flowcharts, or examples may be implemented individually and / or collectively by a wide range of hardware, software, firmware, or substantially any combination thereof. Suitable processors include, by example, GPUs (graphics processing units), general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.
[0124] While the features and elements are provided above in specific combinations, it will be understood by those with ordinary art in the art that each feature or element can be used individually or in any combination with other features and elements. This disclosure is not limited in terms of the specific embodiments described in this application, which are intended to be illustrative of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from the spirit and scope of the invention. Any elements, actions, or instructions used in the description of this application should not be construed as important or essential to the invention unless expressly presented as such. In addition to those enumerated herein, functionally equivalent methods and apparatus within the scope of this disclosure will be apparent to those skilled in the art from the above description. Such modifications and variations are intended to fall within the scope of the appended claims. This disclosure is limited only by the terms of the appended claims, and is limited along with the full scope of the equivalents for which such claims are entitled. It should be understood that this disclosure is not limited to any particular method or system.
[0125] It should also be understood that the terms used herein are for the purpose of describing only specific embodiments and are not intended to be limiting.
[0126] In certain representative embodiments, some parts of the subject matter described herein may be implemented via application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, it will be recognized by those skilled in the art that some aspects of the embodiments disclosed herein can be equivalently implemented in an integrated circuit, in whole or in part, as a single computer program running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or substantially any combination thereof, and that designing circuits and / or writing software and / or firmware code is within the scope of the art of those skilled in the art in light of this disclosure. In addition, it will be understood by those skilled in the art that the mechanisms of the subject matter described herein may be distributed as various forms of program products, and that the exemplary embodiments of the subject matter described herein are applicable regardless of the particular type of signal-carrying medium used to actually carry out the distribution. Examples of signal-carrying media include, but are not limited to, recordable media such as floppy disks, hard disk drives, CDs, DVDs, digital tapes, and computer memory, as well as transmitting media such as digital and / or analog communication media (e.g., optical fiber cables, waveguides, wired communication links, wireless communication links, etc.).
[0127] The subject matter described herein may, in some cases, depict different components that are contained within or connected to other different components. Such illustrated architectures are merely examples, and it should be understood that in practice, many other architectures can be implemented to achieve the same function. Conceptually, any arrangement of components to achieve the same function is effectively “associated” in such a way that the desired function can be achieved. Therefore, any two components combined herein to achieve a particular function, regardless of architecture or intermediate components, can be seen as “associated” with each other in such a way that the desired function can be achieved. Similarly, any two components thus associated can be considered “operably connected” or “operably coupled” with each other to achieve the desired function, and any two components that can be associated in such a way can be considered “operably coupled” with each other to achieve the desired function. Specific examples of operably coupled components include, but are not limited to, physically matable and / or physically interacting components, and / or wirelessly interactable and / or wirelessly interacting components, and / or logically interacting and / or logically interactable components.
[0128] With regard to the use of substantially any plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or singular to plural as appropriate to the context and / or use. For clarity purposes, various singular / plural rearrangements may be explicitly described herein.
[0129] In general, it will be understood by those skilled in the art that the terms used herein, and in particular in the appended claims (e.g., the main text of the appended claims), are generally intended to be “non-limiting” terms (for example, the term “includes” should be interpreted as “includes but not limited to,” the term “has” should be interpreted as “has at least,” and the term “includes” should be interpreted as “includes but not limited to.”). Furthermore, it will be understood by those skilled in the art that if a particular number of claims introduced are intended to be described, such intent is explicitly stated in the claims, and if such statement is not present, such intent does not exist. For example, if only one item is intended, the term “single” or similar wording may be used. To aid understanding, the following description of the appended claims and / or herein may include the use of the introductory phrases “at least one” and “one or more” to introduce the description of the claims. However, the use of such phrases should not be interpreted as meaning that the introduction of a claim description by the indefinite article "a" or "an" limits any particular claim containing such introduced description to embodiments containing only one such description, even if the same claim contains the introductory phrase "one or more" or "at least one" and an indefinite article such as "a" or "an" (for example, "a" and / or "an" should be interpreted as meaning "at least one" or "one or more"). The same applies to the use of definite articles used to introduce a claim description. In addition, it will be recognized by those skilled in the art that even if a particular number of descriptions in an introduced claim are explicitly stated, such description should be interpreted as meaning at least the number stated (for example, the simple statement "two descriptions" without other modifiers means at least two descriptions or two or more descriptions).Furthermore, when a notation similar to "at least one of A, B, and C" is used, such a structure is generally intended to mean what a person skilled in the art would understand (for example, "a system having at least one of A, B, and C" includes, but is not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together). When a notation similar to "at least one of A, B, or C" is used, such a structure is generally intended to mean what a person skilled in the art would understand (for example, "a system having at least one of A, B, or C" includes, but is not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together). It will be further understood by those skilled in the art that any substantially any disjunct word and / or phrase presenting two or more alternative terms in the specification, claims, or drawings should be understood as construed to include the possibility of including one of the terms, either of the terms, or both of the terms. For example, the phrase “A or B” should be understood to include the possibility of “A” or “B” or “A and B.” Furthermore, as used herein, the term “any of ~” followed by a list of multiple items and / or a list of categories of multiple items is intended to include “any of,” “any combination of,” “any number of,” and / or “any number of combinations of,” of the items and / or categories of items, individually or in combination with other items and / or categories of other items. Furthermore, as used herein, the term “set / group” or “cluster” is intended to include any number of items, including zero. In addition, as used herein, the term “number” is intended to include any number, including zero.
[0130] In addition, if any feature or aspect of the present disclosure is described in terms of the Markush group, a person skilled in the art will recognize that the present disclosure is also described in terms of any individual member or subgroup of a member of the Markush group.
[0131] For all purposes, including providing written explanations, as will be understood by those skilled in the art, all scopes disclosed herein also encompass any possible sub-scopes and combinations of sub-scopes. Any enumerated scope can be readily recognized as sufficiently explainable and enable that the same scope can be broken down into at least equal 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 10, etc. As a non-limiting example, each scope described herein can readily be broken down into the lower third, the middle third, the upper third, etc. Also, as will be understood by those skilled in the art, all words such as “up to,” “at least,” “greater than,” and “less than” include the number mentioned and mean a scope that can be further broken down into sub-scopes as described above. Finally, as will be understood by those skilled in the art, a scope includes each individual element. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so on.
[0132] Furthermore, unless otherwise specifically stated, the claims should not be read as being limited to the order or elements provided. In addition, in any claim, the use of the term “means for” is intended to appeal to Section 112, paragraph 6 of the U.S. Patent Act, or the means-plus-function claim format, and no claim without the term “means for” is intended to appeal in that way.
[0133] The system is intended to be implemented in software on a microprocessor / general-purpose computer (not shown). In certain embodiments, one or more functions of the various components may be implemented in software that controls the general-purpose computer.
[0134] In addition, although the present invention is illustrated and described herein with reference to specific embodiments, it is not intended to be limited to the details shown. Rather, various modifications can be made in detail within the scope of the claims and their equivalents, without departing from the present invention.
Claims
1. A wireless transceiver unit (WTRU), A receiver configured to receive a portion of a deep neural network (DNN) model, wherein the portion lies before a division point of the DNN model, and the portion of the DNN model includes a neural network for compressing the features at the division point of the DNN model; One or more processors, Obtaining the compression coefficient of the aforementioned neural network, In response to the compression coefficient, determine which nodes in the neural network should be connected, In response to the aforementioned decision, the neural network is constructed, To generate compressed features, inference is performed using the aforementioned part of the DNN model, One or more processors configured to perform, A transmitter configured to transmit the compressed features to another WTRU, WTRU equipped with.
2. The WTRU according to claim 1, wherein the transmitter is further configured to transmit an indication of the acquired compression coefficient to the other WTRU.
3. The WTRU according to claim 1, wherein one or more processors are configured to determine which nodes in the neural network should be connected when the compression coefficient is adjusted.
4. The WTRU according to claim 1, wherein at least one of the division point and the compression coefficient is configured based on one or more of (1) physical layer operation, (2) media access control layer operation, (3) wireless resource control layer operation, (4) available processing resources, and (5) control signaling.
5. The WTRU according to claim 1, wherein at least one of the division point and the compression coefficient is configured based on the transmission data rate.
6. A method performed by a wireless transceiver unit (WTRU), Receiving a portion of a deep neural network (DNN) model, wherein the portion lies before a split point of the DNN model, and the portion of the DNN model includes a neural network for compressing the features at the split point of the DNN model. Obtaining the compression coefficient of the aforementioned neural network, In response to the compression coefficient, determine which nodes in the neural network should be connected, In response to the aforementioned decision, the neural network is constructed, To generate compressed features, inference is performed using the aforementioned part of the DNN model, The compressed features are transmitted to another WTRU, A method that includes this.
7. The method of claim 6, further comprising transmitting the obtained compression factor indication to the other WTRU.
8. The method of claim 6, wherein which nodes in the neural network should be connected is determined when the compression coefficient is adjusted.
9. The method of claim 6, wherein only one DNN model is loaded into the WTRU for different compression coefficients.
10. The method of claim 6, wherein at least one of the division point and the compression coefficient is based on one or more of (1) physical layer operation, (2) media access control layer operation, (3) wireless resource control layer operation, (4) available processing resources, and (5) control signaling.
11. A wireless transceiver unit (WTRU), A receiver configured to receive a portion of a deep neural network (DNN) model, wherein the portion lies after a split point of the DNN model, and the portion of the DNN model includes a neural network for extending the features at the split point of the DNN model, and the receiver is also configured to receive one or more features output from another WTRU. One or more processors, Obtaining the compression coefficient of the aforementioned neural network, In response to the compression coefficient, determine which nodes in the neural network should be connected, In response to the aforementioned decision, the neural network is constructed, Using one or more of the features output from another WTRU as input to the neural network, inference is performed using the part of the DNN model, One or more processors configured to perform, WTRU equipped with.
12. The WTRU according to claim 11, wherein the receiver is further configured to receive a signal indicating the compression coefficient.
13. The WTRU according to claim 11, wherein one or more processors are configured to determine which nodes in the neural network should be connected when the compression coefficient is adjusted.
14. The WTRU according to claim 11, wherein only one DNN model is loaded into the WTRU for different compression coefficients.
15. The WTRU of claim 11, wherein at least one of the division point and the compression coefficient is configured based on one or more of (1) physical layer operation, (2) media access control layer operation, (3) wireless resource control layer operation, (4) available processing resources, and (5) control signaling.
16. A method performed by a first wireless transceiver unit (WTRU), Receiving a portion of a deep neural network (DNN) model, wherein the portion lies after a split point of the DNN model, and the portion of the DNN model includes a neural network for extending the features at the split point of the DNN model. Receiving one or more features output from the second WTRU, Obtaining the compression coefficient of the aforementioned neural network, In response to the compression coefficient, determine which nodes in the neural network should be connected, In response to the aforementioned decision, the neural network is constructed, Using one or more features output from the second WTRU as input to the neural network, inference is performed using the part of the DNN model, A method that includes this.
17. The method of claim 16, further comprising receiving a signal indicating the compression coefficient.
18. The method of claim 16, wherein which nodes in the neural network should be connected is determined when the compression coefficient is adjusted.
19. The method of claim 16, wherein only one DNN model is loaded into the second WTRU for different compression coefficients.
20. The method of claim 16, wherein at least one of the division point and the compression coefficient is based on one or more of (1) physical layer operation, (2) media access control layer operation, (3) wireless resource control layer operation, (4) available processing resources, and (5) control signaling.
Citation Information
Patent Citations
Data analysis system, method and program
JP2019191635A
System and methods to share machine learning functionality between cloud and an IoT network
US20200372412A1
Machine-learned model switching system, edge device, machine-learned model switching method, and program
WO2019193660A1