Split neural network computations

By collaborating to split the neural network computing load between the client device and the networked computing device, the problem of the client device being limited in resources is solved and it is difficult to perform complex DNN operations is achieved, thereby achieving efficient neural network operations and improved user experience.

CN120153377APending Publication Date: 2025-06-13GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079253.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-01
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Due to resource constraints, it is difficult for the client device to implement sufficiently complex deep neural network (DNN) operations to provide the desired results.

Method used

By collaborating between the client device and the networked computing device, the neural network computing load is split, so that the neural network architecture of the neural network is distributed on two or more devices, and processed by each device step by step to generate the final result output.

Benefits of technology

This enables client devices to perform complex neural network operations while meeting resource limitations, improves performance efficiency and user experience, while reducing the need for battery power and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153377A_ABST
    Figure CN120153377A_ABST
Patent Text Reader

Abstract

A client device (104) and one or more such networked computing devices (112) coordinate to implement a client-initiated and designed split neural network computing scheme in which a neural network architecture of a neural network (116) is distributed and "split" over the client device and the one or more networked computing devices, each device operates to perform at least one successive split neural network portion (122, 124) of the entire neural network operation using the input data (130), where intermediate result data (136) generated thereby is then transmitted to a next device in the sequence until a final result output (140) is generated by a final split neural network portion in the sequence. The device that generates the final result output may then take one or more actions based on the final result output, or forward the final result output to the device that initiates the split neural network computing operation for further processing.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] User equipment, wearable devices, and other client devices in cellular networks and other wireless networks are increasingly adopting neural networks to perform certain operations at various protocol layers to achieve improved performance efficiency or enhanced user experience. For example, a deep neural network (DNN) can be adopted at a client device to provide image analysis, transmission environment-aware radio frequency signaling, and others. However, client devices are typically resource-constrained, which generally makes it impractical to implement a neural network complex enough to provide the desired results. For example, a user device may not have enough battery power, enough computing resources, or enough network capacity to perform DNN operations within an allotted time frame or with the expected complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] The present disclosure is better understood by reference to the accompanying drawings, and numerous features and advantages of the present disclosure become apparent to those skilled in the art. The same reference numerals are used in different drawings to indicate similar or identical items.

[0003] Figure 1 is a diagram illustrating an example wireless system that employs a split neural network computing scheme to split the neural network computing load among two or more devices, according to some embodiments.

[0004] Figure 2 is a diagram illustrating an example of a split neural network configuration using a client-first implementation, according to some embodiments.

[0005] Figure 3 is a diagram illustrating another example of a split neural network configuration using an interleaved configuration, according to some embodiments.

[0006] Figure 4 is a diagram illustrating, according to some embodiments, Figure 1 an example hardware configuration of a user equipment of a wireless system.

[0007] Figure 5 is a diagram illustrating, according to some embodiments, Figure 1 an example hardware configuration of a server of a wireless system.

[0008] Figure 6 is a diagram illustrating a machine learning (ML) module that employs split neural network portions, according to some embodiments.

[0009] Figure 7 is a flowchart illustrating an example method for split neural network computing in a wireless system, according to some embodiments.

[0010] Figure 8Is a flowchart illustrating an example of a process for negotiating a split neural network configuration between at least two devices according to some embodiments.

[0011] Figure 9 Is a flowchart illustrating an example of performing machine learning operations using a neural network split across two or more devices according to some embodiments.

[0012] Figure 10 Is an illustration according to some embodiments of Figures 7 to 9 The ladder signaling diagram of the example operations of the method. Detailed Description

[0013] Due to various resource constraints, client devices typically avoid implementing complex DNNs or other neural networks. However, the network to which a client device is attached is typically able to access servers and other computing devices with significantly fewer resource constraints. To take advantage of this, in at least one embodiment, a client device and one or more such networked computing devices cooperate to implement a split neural network computing scheme in which the neural network architecture of a neural network is distributed or "split" across the client device and one or more networked computing devices such that each device operates to perform at least one successive part of the overall DNN operation (hereinafter "split neural network part" or simply "split part"), where the resulting intermediate result data is then transmitted to the next device in the sequence until a final result output is generated by the final split neural network part in the sequence. The device that generates the final result output can then take one or more actions based on the final result output or forward the final result output to the device that initiated the split neural network computing operation for further processing.

[0014] In at least one embodiment, this split neural network configuration is initiated by a client device. Thus, in response to the initiation of an operation involving the use of a neural network, one of the client device or the networked computing device determines a relevant set of current local conditions of the client device, such as its available computing resources (e.g., processor and memory), current battery condition, current network condition, current thermal condition, and others. Moreover, in some implementations, there may be execution requirements associated with the operation, such as a latency condition indicating a maximum time period allowed for generating the final result output, a quality of service (QoS) condition indicating a maximum amount of time allowed to complete the operation, or an accuracy condition indicating a minimum level of accuracy or other performance level in the final result output. In the case where the client device determines that the latency condition, QoS condition, accuracy condition, or other such execution requirements can be met given the current local conditions of the client device, the client device can choose to perform the operation entirely locally; that is, perform the operation using the full local neural network architecture of the neural network.

[0015] However, in cases where one or more specified execution requirements appear impossible to meet with only local execution, in at least some embodiments, either or both of the client device or the networked computing device determine a split neural network configuration, in which the neural network architecture of the neural network is dispatched or "split" between the client device and the networked computing device (or networked computing devices). The specific split between the devices can be based on current local conditions, execution requirements, or a combination thereof. For example, in cases with relatively low local computing resource availability, low battery reserve, and / or short latency requirements, the split neural network configuration can dispatch more layers of the neural network architecture to one or more networked computing devices and fewer layers to the client device, while in cases with relatively high local computing resource availability, high battery reserve conditions, and long latency conditions, the split neural network configuration can dispatch more layers of the neural network architecture to itself and fewer layers to one or more networked computing devices.

[0016] In some embodiments, the initiating device then initiates a negotiation with other devices intended to be included in the split neural network configuration by sending a split computing request with the proposed split neural network configuration. The other devices can then consider their own resource constraints and execution requirements when determining whether to accept the proposed split neural network configuration or counter-propose with a different split neural network configuration that is more compatible with the resource constraints or execution requirements. When the client device and one or more networked computing resources have agreed on an agreed-upon split neural network configuration, each of the client device and the one or more networked computing resources implements its corresponding portion of the neural network architecture (i.e., its corresponding "split portion") at the corresponding machine learning (ML) module. Thereafter, the device at the head of the sequence of the split neural network configuration receives input data for operation, processes the input data at its ML module, and provides the output as intermediate result data to the device in the sequence that has the next split portion of the neural network architecture, and so on.

[0017] In some implementations, the split is a two-way split, where one device (e.g., a networked computing device) has the first sequence of layers or other initial portion (e.g., "first stage") of a neural network architecture, and another device (e.g., a client device) has the last sequence of layers or other final portion (e.g., "last stage") of the neural network architecture. In other implementations, the split is an N-way split (N > 2), where each of the N stages is assigned to a corresponding device. In some embodiments, each stage is assigned to a different device, such as the first stage being assigned to one networked computing device, the intermediate stages being assigned to a second networked computing device, and the last stage being assigned to a client device. In other embodiments, the stages may be interleaved among a group of M devices (M < N). For example, the first stage and the last stage may be assigned to the client device, and the intermediate stages may be assigned to the networked computing devices.

[0018] In the case of implementing neural network operations using this split computing method, when the client device determines that its current local conditions may prevent a local-only solution from satisfying one or more conditions or goals, the client device can opportunistically utilize the resources of one or more available networked computing devices to assist in performing the neural network operations.

[0019] For ease of description, the systems and techniques are described herein in the example scenario of a cellular network, in which a user equipment (UE) operates as the client device and one or more servers operate together as the networked computing devices. However, these references are for illustrative purposes only, and it should be understood that the references to the UE, server, or cellular network are similarly applicable, respectively, to other client devices, other networked computing devices, and other networks, unless otherwise specified.

[0020] Moreover, for ease of illustration and to facilitate understanding, the techniques of the present disclosure are described in example implementations in which the UE or other client device is the initiator of the partitioning of the neural network via a split neural network configuration and is thus the device that triggers the negotiation process. However, in other embodiments, it may be the networked computing device that initiates the splitting of the neural network between the networked computing device and the client device (and in some cases, also between one or more additional networked computing devices). Thus, using the criteria provided herein, references and descriptions of the initiation or negotiation of a split neural network configuration by the client device are to be understood as applying, alternatively, to the initiation or negotiation of a split neural network configuration by the networked computing device.

[0021] Figure 1FIG. 0 illustrates an example wireless communication network 100 employing a split neural network computing scheme according to some embodiments. In the depicted example, wireless communication network 100 is a cellular network including network infrastructure 102 wirelessly connected to one or more client devices such as UE 104. Network infrastructure 102 includes a core network 106, which is coupled to one or more wide area networks (WANs) 108 or other packet data networks (PDNs) (such as the Internet). Core network 106 is further connected to at least one base station (BS) 110. BS 110 supports wireless communication with one or more wireless client devices such as UE 104 via radio frequency (RF) signaling using one or more applicable RATs as specified by one or more communication protocols or standards. Thus, BS 110 operates as a wireless interface between one or more wireless devices and various networks and services provided by network infrastructure 102 such as packet switched (PS) data services, circuit switched (CS) services, and others.

[0022] BS 110 may operate using any of a variety of RATs, such as operating as a NodeB (or base transceiver station (BTS)) of a Universal Mobile Telecommunications System (UMTS) RAT (also known as “3G”), operating as an enhanced NodeB (eNodeB) of a 3rd Generation Partnership Project (3GPP) Long Term Evolution (LTE) RAT, operating as a 5G Node B (“gNB”) of a 3GPP 5th Generation (5G) New Radio (NR) RAT, and others. UE 104, in turn, represents any of a variety of client devices operable to communicate with BS 110 via a suitable RAT, including, for example, a mobile cellular phone, a tablet computer or laptop computer, a desktop computer, a video game system, a server, a network-enabled appliance, a network-enabled automotive communication system, a network-enabled smartwatch, or other wearable device, and others. Core network 106 and / or WAN 108 includes or is capable of network access to one or more networked computing devices 112 that provide computing resources for core network 106, WAN 108, and / or UE 104, such as one or more servers (or server farms), one or more workstations, and the like. In the examples described herein, one or more networked computing devices 112 include one or more servers, and thus, for ease of reference, networked computing devices 112 are also referred to herein as “servers 112”.

[0023] In at least one embodiment, UE 104 is configured to implement a machine learning (ML) module (see Figure 6), the ML module can be configured to implement one or more neural networks (or portions thereof) to provide neural network / machine learning capabilities to one or more software applications executing at the UE 104. Examples of such neural networks include deep neural networks (DNNs) such as convolutional neural networks (CNNs), artificial neural networks (ANNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and others. For example, a software application of the UE 104 can employ an ML module that implements a CNN to provide image detection / classification or an RNN to provide speech detection. However, depending on the complexity or computational resources required to implement an entire neural network, the UE 104 may not have sufficient resources to locally execute neural network operations in a timely manner. For example, for real-time speech detection and translation, the UE 104 may not have sufficient computational resources to execute speech detection and translation operations fast enough to facilitate a satisfactory user experience. Alternatively or in addition, executing neural network operations entirely at the UE 104 may impose an unreasonable or unacceptable burden on the UE 104. For illustration, using a DNN at the UE 104 to perform a series of image classification operations may result in excessive battery consumption on the UE 104 or may cause an unmanageable heat load at the UE 104.

[0024] Thus, in at least one embodiment, the UE 104 and the network infrastructure 102 together may implement a split neural network computing scheme, in which a split neural network configuration 114 is used to implement a neural network 116 that is distributed or "split" across the UE 104 and one or more servers 112 or other networked computing devices of the network 100, such that each device in the split neural network configuration 114 operates to perform a part of the overall neural network operation (i.e., a "split neural network part" or "split part"), where the intermediate results thus generated are then transmitted to the next device in the sequence until a final result output is generated by the final split part in the sequence. The split parts assigned to the corresponding devices may include any subset or partitioning of the elements of the neural network employed by the split neural network configuration 114. For illustration, many neural networks are configured as a sequence of layers, the sequence of layers including an input layer, one or more hidden layers, and an output layer, where each layer has one or more nodes (e.g., neurons and / or perceptrons). In such cases, the split neural network parts may include all of the nodes of one or more layers of the neural network, a portion of the nodes of one or more layers of the neural network, or a combination thereof. In some neural networks, the layers / nodes are further organized, and the splitting of the neural network into corresponding split neural network parts may occur according to such organization. For example, a GAN typically consists of a generator and a discriminator, and given this composition, a split of the GAN across multiple devices may occur, such that the layers and nodes associated with the generator are assigned to one or more devices, and the layers and nodes associated with the discriminator are assigned to one or more other devices.

[0025] In an implementation, the neural network is split such that the split neural network configuration 114 is implemented as a sequence of neural network parts representing the split of the entire neural network, whereby each split neural network part receives the initial input or the output from a previous split neural network part in the sequence as its input and, depending on its position in the sequence, provides its output as the input to the next split neural network part in the sequence or as the final result output. In a two-way split, the first device may be assigned, for example, the input layer and one or more hidden layers after the input layer, and the second device may be assigned, for example, the remaining hidden layers and the output layer. Which device is assigned which split part may depend on various factors, such as comparative computational resource availability, the type of neural network operation being performed, the use of the output, etc. For example, for an image classification operation where the image classification result is to be used by an application at the UE 104, the server 112 may be assigned the initial split part with the input layer and most of the hidden layers, while the UE 104 may be assigned the final split part with the remaining hidden layers and the output layer. In an N-way split (N>2), each of the M devices (M>1) is assigned a split part, and in the case where N>M, at least one device is assigned multiple split parts.

[0026] As described below, in some embodiments, the UE 104 operates to initiate the split of the neural network for performing a neural network operation and determines how to perform the split, either unilaterally or as a result of negotiation with the network infrastructure 102. The decision of the UE 104 to split the neural network operation between the UE 104 and one or more servers 112 (or other networked computing devices) may be based on various factors, including the scope or magnitude of the operation to be performed, the local resources of the UE 104 available for performing the neural network operation, one or more execution requirements (e.g., latency, QoS, accuracy threshold), network bandwidth or other network parameters, and others. As a general guideline for some implementations, the portion of the neural network that the UE 104 seeks to offload to one or more servers 112 is inversely proportional to the local resources available to the UE 104 and / or the strictness of the execution requirements, all else being equal. However, note that in other embodiments, the server 112 may alternatively use similar techniques as described below to initiate the split of the neural network.

[0027] When the UE 104 proposes / determines this particular split, the UE 104 can use any of a variety of techniques to notify the server 112 of the split neural network portions that it has been assigned. In some embodiments, the server 112 (and the UE 104) can access a repository 118 containing a representation 120 of the neural network, and the UE 104 can notify the server 112 of the split neural network portions assigned to the server 112 and other relevant details, such as initial weights, the manner in which the split neural network portions are to be implemented at the server 112, and so on. For example, each representation 120 can be identified by a corresponding identifier, which is then referenced by the UE 104 and the server 112. The representation 120 can refer to a particular implementation of a type of neural network (e.g., an RNN with a particular number of layers, nodes, weights, etc.), and the UE 104 and the server 112 use additional information to effect the split of this particular implementation into the determined split portions, or the representation 120 can represent a predefined split configuration template, such as an RNN with a particular number of layers, nodes, weights, etc., and a predefined split of these elements among multiple split neural network portions. Although shown as a separate component (e.g., a separate server in the network infrastructure 102), in some embodiments, the repository 118 can be implemented, in whole or in part, at the server 112. In other embodiments, the UE 104 can provide the server 112 with information specifying the manner in which the server 112 is to implement the corresponding split neural network portions, such as a data structure representing the structure of the split neural network portions, their input format, their output format, weights or other parameters for different nodes in the structure, and one or more other data structures.

[0028] As an illustration of this split neural network computing scheme, Figure 1Depicts an example split neural network configuration 114 for neural network 116, in which neural network 116 is split into two parts, namely, split neural network parts 122, 124, where split neural network part 122 is assigned the input layer and the first five hidden layers of example neural network 116 (and is thus referred to as the "initial split neural network part 122"), and split neural network part 124 is assigned the last two hidden layers and the output layer of example neural network 116 (and is thus referred to as the "final split neural network part 124"). Under this configuration, neural network 116 is split into the following sequence: initial neural network part 122 → final neural network part 124. For this example, it is assumed that UE 104 negotiated this split with server 112 using the techniques described below based on the current local resources available to UE 104, the execution requirements for the operations to be performed (e.g., ML-assisted image classification, voice detection, etc.), and as a result of negotiation with server 112. Further, for this implementation, the initial split neural network part 122 is assigned to server 112, and the final split neural network part 124 is assigned to UE 104. Thus, the initial split neural network part 122 is implemented at server ML module 126 at server 112, and the final split neural network part 124 is implemented at UE ML module 128 at UE 104.

[0029] To initiate an ML-assisted operation using this example configuration, a software application of UE 104 generates input data 130-1 for input into neural network 116. For example, if the ML-assisted operation is image classification, the input data 130-1 can include, for example, an image captured by a camera of UE 104. As another example, if the ML-assisted operation is speech detection and translation, the input data 130-1 can include, for example, speech captured and digitized by a microphone of UE 104. The radio frequency (RF) interface 132 of UE 104 transmits the input data (as input data 130-2) to the corresponding RF interface 134 of BS 110, and then the BS transmits the input data (as input data 130-3) to server 112. At server 112, the input data 130-3 is provided as input to server ML module 126, which processes the input data 130-3 at the input layer and the first five hidden layers of neural network 116 according to the initially assigned split neural network portion 122. Accordingly, the output 136 of the nodes of the fifth hidden layer serves as intermediate result data 138-1 of the initially assigned split neural network portion 122 (and thus as the output of server ML module 126). Thereby, server 112 transmits the intermediate result data (as intermediate result data 138-2) to BS 110, which in turn wirelessly transmits the intermediate result data 138-2 to UE 104.

[0030] UE 104 provides the intermediate result data (as intermediate result data 138-3) as input to UE ML module 128, which implements the final split neural network portion 124. Accordingly, the output 136 represented in the intermediate result data 138-3 is provided as input to the corresponding nodes of the sixth hidden layer of neural network 116 and is processed by the sixth, seventh, and eighth hidden layers and the output layer to generate a final result output 140, which represents the final result output of neural network 116 in the case where the intermediate result data 138-3 is the input. The software application of UE 104 that initiated the ML-assisted operation and / or another software application of UE 104 can then perform one or more actions in response to the final result output 140. For example, if the ML-assisted operation is a speech detection and translation operation, the final result output 140 can represent the translated representation of the input speech, and thus the software application can manipulate UE 104 to output an audio version of the translated representation via a speaker of UE 104 and / or display a text representation via a display of UE 104.

[0031] Although Figure 1Depicts a specific two-way split of neural network 116 that assigns an initial portion to server 112 and a final portion to UE 104, but the split of any given neural network can be implemented in a variety of ways depending on needs and scenarios. For example, UE 104 can implement the initial portion of the split neural network configuration while server 112 implements the final portion. Figure 2 Shows an example of this scenario, where split neural network configuration 214 splits neural network 216 into an initial split portion 202 implemented at UE ML module 228 of UE 104 and a final split portion 204 implemented at server ML module 226 of server 112. Using this method, a software application at UE 104 provides input data 230 as input to UE ML module 228, whereby a first subset of the initial and hidden layers of neural network 216 processes input data 230. The output 236 of the last hidden layer in initial split portion 202 is provided as intermediate result output 238, which is transmitted via base station 110 to server 112. Server 112 provides intermediate result output 238 as input to the first hidden layer in final split portion 204, whereby this hidden layer, subsequent hidden layers, and the output layer process intermediate result output 238 to generate final result output 240. Final result output 240 can then be transmitted back to UE 104 for use by one or more software applications at the UE, and / or final result output 240 can be used by server 112 or some other component of network infrastructure 102 to perform one or more remote actions 242. For example, neural network 116 can be used to generate enhanced display information for an augmented reality (AR) display, and thus one or more remote actions 242 can be to display the generated enhanced display information at the AR display (one result of final result output 240). As another example, neural network 116 can be used to generate virtual reality (VR) content, such as VR display content or VR audio content (an example of final result output 240), and one or more remote actions 242 can be to display or output the generated VR content. As yet another example, neural network 116 can be used to improve or enhance the cellular connectivity between UE 104 and BS 110 (e.g., by controlling some aspect of the UE's RF antenna array based on input sensor data), and thus one or more remote actions 242 can include, for example, modifying the operation of the RF antenna array based on final result output 240.

[0032] Further, as described above, in some implementations, the split configuration can be an N-way split (N > 2), where each of the M devices obtains at least one split neural network configuration. Figure 3An example of such a situation is shown, where the split neural network configuration 314 splits the neural network 316 into three split parts: an initial split part 302 implemented at the UE ML module 328 of the UE 104, an intermediate split part 304 implemented at the server ML module 326 of the server 112, and a final split part 306 implemented at the UE ML module 328 of the UE 104 (or at a second instance of the UE ML module of the UE 104). In this example, a software application at the UE 104 provides the input data 330 as an input to the UE ML module 228 for processing by the initial split part 302. The output 336-1 of the last hidden layer in the initial split part 302 is provided as a first intermediate result output 338-1, which is transmitted via the base station 110 to the server 112. The server 112 provides the first intermediate result output 338-1 as an input to the first hidden layer in the intermediate split part 304, whereby this hidden layer and subsequent hidden layers in the intermediate split part 304 process the first intermediate result output 338-1. The output 336-2 of the last hidden layer in the intermediate split part 304 is provided as a second intermediate result output 338-2, which is transmitted via the base station 110 to the UE 104. Then, the UE 104 provides the second intermediate result output 338-2 as an input to the first hidden layer in the final split part 306, and this hidden layer, subsequent hidden layers, and the output layer in the final split part 306 process this input to generate a final result output 340. The final result output 340 can then be processed by one or more software applications at the UE 104, in response to which one or more actions are taken at the UE 104.

[0033] Figure 4 An example hardware configuration for the UE 104 (as a representative client device) according to some embodiments is shown. It should be noted that the depicted hardware configuration represents the processing components and communication components most directly related to the neural network-based processes described herein, and certain components that are well known and often implemented in such electronic devices, such as displays, peripheral devices, external power supply devices, and others, are omitted.

[0034] In the depicted configuration, UE 104 includes an RF interface 402 that has one or more antennas 404 and one or more modems to support one or more radio access technologies (RATs), such as 3rd Generation Partnership Project (3GPP) 4th Generation Long Term Evolution (4G LTE)-compatible RATs, 3GPP 5th Generation New Radio (5G NR)-compatible RATs, Institute of Electrical and Electronics Engineers (IEEE) 802.11-compatible RATs, and others. UE 104 further includes one or more processors 406, a set of sensors 408, a user interface (UI) 410, and one or more batteries 412 or other power sources. The one or more processors 406 can include, for example, one or more central processing units (CPUs), graphics processing units (GPUs), tensor processing units (TPUs), or other application-specific integrated circuits (ASICs), among others. For illustration purposes, processor 406 can include an application processor (AP) utilized by UE 104 to execute an operating system and various user-level software applications, as well as one or more processors utilized by the modem or baseband processor of RF interface 402. The set of sensors 408 can include, for example, satellite positioning sensors such as Global Positioning System (GPS) sensors, Global Navigation Satellite System (GNSS) sensors, Inertial Measurement Unit (IMU) sensors, visual odometry sensors, accelerometers, gyroscopes, barometers, altimeters, tilt sensors or other inclinometers, ultra-wideband (UWB)-based sensors, and others. Other examples of the types of sensors in the set of sensors 408 can include sensors for determining the current operating state of UE 104, such as battery level sensors, thermal sensors, screen mode sensors, and others. UI 410 includes components for interfacing with the user, such as a display, keyboard or touchpad, buttons, microphone, speaker, and others.

[0035] UE 104 further includes one or more computer-readable media 414, which include any of a variety of media used by an electronic device to store data and / or executable instructions, such as random access memory (RAM), read-only memory (ROM), cache, flash memory, solid-state drive (SSD), or other mass storage devices, etc. For ease of illustration and brevity, given the frequent use of system memory or other memory to store data and instructions for execution by processor 406, computer-readable media 414 are referred to herein as "memory 414", but it should be understood that references to "memory 414" should equally apply to other types of storage media, unless otherwise specified. One or more memories 414 of UE 104 are used to store one or more sets of executable software instructions and associated data for manipulating one or more processors 406 and other components of UE 104 to perform the various functions described herein and attributable to UE 104. The set of executable software instructions includes, for example, an operating system (OS) 416 and various drivers (not shown), as well as one or more user-level software applications 418, which manipulate the hardware of UE 104 through OS 416 or otherwise interact with the hardware of the UE.

[0036] The set of executable software instructions further includes one or more of a neural network management module 420 and a split computing management module 422. The neural network management module 420 implements one or more ML modules (e.g., ML module 128, Figure 1 ), which utilize one or more neural networks (or split parts thereof) for UE 104, as described in detail below. The split computing management module 422 operates to monitor the status of the local resources of UE 104, such as battery status, available memory status, available processor resources, network status, and others, and based on these local resource statuses (and / or other considerations), determines whether to implement the neural network as a solely local neural network (i.e., entirely at UE 104) or as a split neural network configuration for performing one or more operations using the neural network. Moreover, in the case where a split neural network configuration is to be adopted, the split computing management module 422 can operate to manage the negotiation for the split parts implemented by one or more servers 112 of the network infrastructure 102, and coordinate with the neural network management module 420 to implement the split parts allocated to UE 104 in the negotiated split neural network configuration.

[0037] To facilitate the operation of the UE 104 as described herein, one or more memories 414 of the UE 104 may further store data associated with such operations. This data may include, for example, one or more neural network architecture configurations 424, as well as user data, multimedia data, beamforming codebooks, software application configuration information, and other device data (not shown). Each neural network architecture configuration 424 includes one or more data structures that contain data and other information representing the corresponding architecture and / or parameter configuration used by the neural network management module 420 to form the corresponding neural network or its split parts for the UE 104. The information included in the neural network architecture configuration 424 includes, for example, parameters specifying the following: fully connected layer neural network architecture, convolutional layer neural network architecture, recurrent neural network layer, number of connected hidden neural network layers, input layer architecture, output layer architecture, number of nodes utilized by the neural network, coefficients (e.g., weights and biases) utilized by the neural network, kernel parameters, number of filters utilized by the neural network, stride / pooling configuration utilized by the neural network, activation function for each neural network layer, interconnections between neural network layers, neural network layers to skip, etc. Thus, the neural network architecture configuration 424 includes any combination of neural network formation configuration elements that can be used to create a neural network architecture configuration (e.g., a combination of one or more neural network formation configuration elements) defining and / or forming a DNN or other neural network.

[0038] Figure 5 An example hardware configuration for a server 112 (as a representative networked computing resource) is shown in accordance with some embodiments. It should be noted that the depicted hardware configuration represents the processing and communication components most directly related to the neural network-based processes described herein, and certain components that are well known and often implemented in such electronic devices are omitted. It should further be noted that although the illustrated implementation is shown as a single server, the functionality and thus the hardware components may alternatively be distributed across multiple servers or other networked computing devices and may be distributed in a manner that performs the functions described herein. Thus, references to the functionality of a single server 112 may also apply to equivalent functionality in multiple servers 112 or other networked computing devices, unless otherwise specified.

[0039] In the depicted configuration, server 112 includes a network interface 502 for wired and / or wireless connectivity to BS 110 and other components of network infrastructure 102. Server 112 further includes one or more processors 506, such as, for example, one or more CPUs, GPUs, TPUs, or other ASICs, among others. Server 112 further includes one or more computer-readable media 508, which include any of a variety of media used by an electronic device to store data and / or executable instructions, such as RAM, ROM, cache, flash memory, SSD, or other mass storage devices, among others. Similar to memory 414 of UE 104, for purposes of illustration and brevity, and given the frequent use of system memory or other memory to store data and instructions for execution by processor 506, computer-readable media 508 are referred to herein as "memory 508", but it should be understood that references to "memory 508" should equally apply to other types of storage media, unless otherwise specified.

[0040] One or more memories 508 of server 112 are used to store one or more sets of executable software instructions and associated data that manipulate one or more processors 506 of server 112 and other components to perform various functions described herein and attributed to server 112, either individually or as a collection of servers 112. The sets of executable software instructions include, for example, OS 510 and various drivers (not shown), as well as various software applications. The sets of executable software instructions further include one or more of neural network management module 512 and split computing management module 514. Similar to neural network management module 420 and split computing management module 422 of UE 104, these modules 512, 514 operate to implement a neural network or a split portion thereof at server 112. In particular, neural network management module 512 implements one or more ML modules (e.g., ML module 126, Figure 1 ) that utilize one or more neural networks (or split portions thereof) for server 112, while split computing management module 514 operates to negotiate with UE 104 for implementing split portions at server 112 and to coordinate with neural network management module 512 to implement the split portions assigned to server 112 in the negotiated split neural network configuration.

[0041] One or more memories 508 further store various information, such as one or more neural network architecture configurations 524 representing trained neural network architecture configurations that can be employed at an ML module of server 112 (e.g., ML module 126, Figure 1 )). Thus, similar to Figure 4The neural network architecture configuration 424, and each neural network architecture configuration 524 includes one or more data structures that contain data and other information representing the corresponding architecture and / or parameter configuration used by the neural network management module 512 of the server 112 to form the corresponding part or the whole of the neural network.

[0042] Figure 6 FIG. illustrates an example machine learning (ML) module 600 for implementing a neural network or a split part thereof according to some embodiments. As described herein, both the UE 104 and the server 112 each implement one or more ML modules to implement one or more split neural network parts. The ML module 600 thus illustrates an example module for implementing one or more of these split neural network parts (or the entire neural network).

[0043] In the depicted example, the ML module 600 implements at least a portion 602 of a deep neural network (DNN) model having a group of connected nodes (e.g., neurons and / or perceptrons) organized into one or more layers. The nodes between the layers can be configured in various ways, such as a partially connected configuration where a first subset of nodes in the first layer is connected to a second subset of nodes in the second layer; a fully connected configuration where each node in the first layer is connected to each node in the second layer, etc. Neurons process input data to produce continuous output values, such as any real number between 0 and 1. In some cases, the output value indicates the proximity of the input data to the desired class. Perceptrons perform linear classification on the input data, such as binary classification. Whether neurons or perceptrons, the nodes can use a variety of algorithms to generate output information based on adaptive learning. Using the portion 602, the ML module 600 performs a variety of different types of analysis, including simple linear regression, multiple linear regression, logistic regression, stepwise regression, binary classification, multi-class classification, multivariate adaptive regression splines, locally estimated scatterplot smoothing, etc.

[0044] When the portion 602 includes the entire DNN model, two or more layers include an input layer, one or more hidden layers, and an output layer. When the portion 602 includes a split part of the DNN model, the layers included in the portion 602 depend on the position of the split part in the sequence that together constitutes the split part of the DNN model. In the case where the portion 602 is an initial split part, the portion 602 includes an initial layer and one or more adjacent hidden layers immediately following the initial layer. In the case where the portion 602 is a final split part, the portion 602 includes an output layer and one or more adjacent hidden layers immediately preceding the output layer. In the case where the portion 602 is an intermediate split part, the portion 602 includes one or more adjacent hidden layers.

[0045] For ease of explanation, the depicted portion 602 of the DNN model includes three layers 604, 606, and 608. For the initial split portion implementation, these three layers can be implemented as, for example, an input layer and two adjacent hidden layers after the input layer; for the intermediate split portion implementation, these three layers can be implemented as three adjacent hidden layers; or for the final split portion implementation, these three layers can be implemented as an output layer and two adjacent hidden layers before the output layer. However, in an actual implementation, the split portion implemented by the ML module 600 will typically have a larger number of layers. Each layer has an arbitrary number of nodes, where the number of nodes between layers can be the same or different. That is, the input layer can have the same number and / or a different number of nodes as the output layer, the output layer can have the same number and / or a different number of nodes as one or more hidden layers, and so on. For example, node 610 corresponds to one of several nodes included in the input layer (which can be represented by layer 604), where the node performs separate and independent computations. As further described, the node receives input data and uses one or more algorithms to process the input data to produce output data. Typically, the algorithms include weights and / or coefficients that change based on adaptive learning. Thus, the weights and / or coefficients reflect the information learned by the neural network. In some cases, each node can determine whether to pass the processed input data to one or more subsequent nodes. For illustration, after processing the input data, node 610 can determine whether to pass the processed input data to one or both of nodes 612 and 614 in an adjacent hidden layer (layer 608 in this example). Alternatively or additionally, node 610 passes the processed input data to the nodes based on the layer connection architecture. This process can be repeated through multiple layers until portion 602 uses the nodes of the final layer (layer 608) in portion 602 (e.g., node 616) to generate an output. To facilitate the input of the received data and the output of the generated output data, the ML module 600 can further include an input interface 618 and an output interface 620. The input interface 618 operates to receive input data 622 and distribute each data included in the input data 622 to the corresponding nodes of the first layer 604 according to a certain predetermined format. Similarly, the output interface 620 operates to receive the respective outputs of the nodes of the final layer 608 and provide the respective outputs as output data 624 according to a certain predetermined format. The formats for the input interface 618 and the output interface 620 depend on the neural network adopted and the layers of the neural network implemented in portion 602. For example, for the initial split portion, the input data 622 represents the initial data input for processing by the neural network, and thus the input interface 618 operates to distribute the initial data to the nodes of the input layer (layer 604) accordingly.However, for the intermediate split portion or the final split portion, the input data 622 represents the output from the upstream hidden layer, and thus the input data 622 represents intermediate result data, and the input interface 618 is configured to distribute each data of the intermediate result data as inputs to the nodes of the first hidden layer (layer 604) according to the neural network architecture. Similarly, for the initial split portion or the intermediate split portion, the output data 624 is the respective output of the nodes of the last hidden layer (layer 608) in the ML module 600, and thus represents intermediate result data, while for the final split portion, the output data 624 is the final result output provided by the output layer (layer 608) of the neural network.

[0046] A neural network can also adopt a variety of architectures, which determine what nodes in the neural network are connected, how data is advanced and / or retained in the neural network, what weights and coefficients are used to process the input data, how the data is processed, and so on. These various factors together describe the neural network architecture configuration, such as the neural network architecture configuration briefly described above. For illustration, a recurrent neural network (such as a long short-term memory (LSTM) neural network) forms a loop between node connections to retain information from a previous part of an input data sequence. The recurrent neural network then uses the retained information for subsequent parts of the input data sequence. As another example, a feedforward neural network passes information to forward connections without forming a loop to retain information. Although described in the context of node connections, it should be understood that the neural network architecture configuration can include a variety of parameter configurations that affect how part 602 or other neural networks process input data.

[0047] The neural network architecture configuration of a neural network can be characterized by various architectures and / or parameter configurations. For illustration, consider an example where part 602 implements a part or the whole of a CNN. Generally speaking, a convolutional neural network corresponds to a type of DNN in which layers use convolutional operations to process data to filter the input data. Therefore, the CNN architecture configuration can be characterized by, for example, pooling parameters, kernel parameters, weights, and / or layer parameters.

[0048] Pooling parameters correspond to parameters of a pooling layer that reduces the dimensionality of input data within a specified convolutional neural network. For illustration, a pooling layer can combine the outputs of nodes at a first layer into the node inputs at a second layer. Alternatively or additionally, the pooling parameters specify how and where in the data processing layer the neural network pools data. For example, a pooling parameter indicating "max pooling" configures the neural network to pool by selecting the maximum value from a grouping of data generated by nodes of the first layer, and use that maximum value as the input to a single node of the second layer. A pooling parameter indicating "average pooling" configures the neural network to generate an average value from a grouping of data generated by nodes of the first layer, and use that average value as the input to a single node of the second layer.

[0049] Kernel parameters indicate the filter size (e.g., width and height) for processing input data. Alternatively or additionally, the kernel parameters specify the type of kernel method used to filter and process the input data. For example, a support vector machine corresponds to a kernel method that uses regression analysis to identify and / or classify data. Other types of kernel methods include Gaussian processes, canonical correlation analysis, spectral clustering methods, and so on. Thus, the kernel parameters can indicate the filter size and / or the type of kernel method to be applied in the neural network. Weight parameters specify the weights and biases used by an algorithm within a node to classify input data. In some implementations, the weights and biases are learned parameter configurations, such as those generated from training data. Layer parameters specify layer connections and / or layer types, such as a fully connected layer type indicating that each node in a first layer (e.g., output layer 608) is connected to each node in a second layer (e.g., hidden layer 606), a partially connected layer type indicating which nodes in the first layer are disconnected from the second layer, an activation layer type indicating which filters and / or layers are activated within the neural network, and so on. Alternatively or additionally, the layer parameters specify the type of node layer, such as a normalization layer type, a convolutional layer type, a pooling layer type, and so on.

[0050] Although described in the context of pooling parameters, kernel parameters, weight parameters, and layer parameters, it should be understood that other parameter configurations can be used to form a DNN consistent with the guidelines provided herein. Thus, a neural network architecture configuration can include any suitable type of configuration parameter for a DNN that can be applied to affect how the DNN processes input data to generate output data.

[0051] Now turning to Figures 7 to 10, Method 700 for implementing a split neural network configuration for a neural network is shown according to some embodiments. The neural network is used to perform ML operations on behalf of a software application or other process of a UE or other client device. For illustrative purposes, method 700 and the accompanying examples are described in the example scenario of wireless communication network 100, where UE 104 operates as a client device and server 112 represents one or more collaborative networked computing devices, but this method is not limited to this specific implementation scenario. Further, Figure 10 The ladder diagram 1000 of Figures 7 to 9 is referenced in the following description of method 700 as an example of the operation of method 700 for a 2-way split neural network configuration to facilitate understanding. Also, it should be noted that the order of operations described with reference to method 700 is for illustrative purposes only, and operations can be performed in a different order, and further, one or more operations can be omitted, or one or more additional operations can be included in the method as shown.

[0052] As explained above, in an implementation, a software application of UE 104 (e.g., user-level application 418 or OS 416 or other kernel / driver-level process) utilizes a neural network to perform a corresponding operation, and the result is utilized by this software application or another software application to perform one or more actions. In the case where UE 104 has sufficient resources to fully locally implement the neural network while meeting the corresponding goals (e.g., maximum latency goal or battery consumption goal), UE 104 can implement the entire neural network at the ML module of UE 104 and use this ML module configured in this way to perform operations. However, in the case where UE 104 cannot fully locally implement the neural network while meeting one or more specified goals, UE 104 alternatively can seek to implement the neural network as a split neural network configuration that utilizes the computing resources of one or more servers 112 of network infrastructure 102.

[0053] Therefore, method 700 is initiated at block 702, where a software application of UE 104 initiates the execution of an ML operation (block 1002, Figure 10 ), and this ML operation utilizes the identified neural network to generate an output result based on input data provided by or referenced by the software application. Using the example from above, this operation can be an image classification operation using a specific CNN, where the input to the CNN is image data representing the image to be classified. The initiation of the ML operation can include, for example, the software application making a call to an application programming interface (API) of OS 416 to provide support for such an operation. At the same time, the split computing management module 422 monitors the status of the local resources of UE 104 (block 1004, Figure 10) and other indicators such as available memory, processor utilization, remaining battery capacity, current device usage, and current or near-future local resource availability. Thus, in response to the initiation of the ML operation, at block 704, the split computing management module 422 determines the current condition of the UE 104 based on these monitoring operations. Further, the split computing management module 422 identifies any execution requirements related to the ML operation, such as execution latency limits, execution power consumption limits, computing resource consumption limits, and others.

[0054] At block 706, the split computing management module 422 then determines whether only a local implementation of the identified CNN or a split configuration is more suitable for the determined current condition and any execution requirements. For example, when there is sufficient local power and computing resources available at the UE 104, the time limit for generating the final result output is relatively long, and the wireless network is bandwidth-constrained, the split computing management module 422 may determine to implement the neural network entirely locally and perform the ML operation only at the UE 104. However, if there are insufficient computing or power resources available at the UE 104, there is a strong wireless connection to the BS 110, and there is a relatively short time limit, the split computing management module 422 may determine to implement the neural network as a split configuration in which one or more parts of the neural network are assigned to the UE 104 and one or more other parts of the neural network are assigned to the server 112 (or are assigned to multiple servers 112). This determination can be made in any of a variety of ways. For example, the determination can be implemented by an algorithm (e.g., via a weighted sum equation), via a look-up table (LUT), via another trained neural network that is less resource-intensive, and others. For illustration, the UE 104 can model the performance of the proposed split configuration to determine whether it meets the UE-side power consumption limit while also meeting the QoS (e.g., latency) requirements. If so, the UE 104 can select the proposed split configuration; if not, the UE 104 can select another proposed split configuration for similar modeling.

[0055] In the case where the split computing management module 422 determines that only a local implementation of the neural network is appropriate or sufficient, then at block 708, the split computing management module 422 instructs the neural network management module 420 to implement an ML module configured to implement the entire architecture configuration of the identified neural network, and the UE 104 uses this ML module to perform the ML operation.

[0056] Returning to block 706, when the split computing management module 422 determines that the split configuration of the neural network is appropriate (block 1006, Figure 10)In the case of , at block 710, the split computing management module 422 performs a negotiation process to determine an acceptable split between the UE 104 and the server 112 (or multiple servers 112) of the neural network given the respective resource constraints of the given UE 104 and the server 112.

[0057] For temporary reference Figure 8 , according to some embodiments, an example implementation of the negotiation process of block 710 is shown. In this example method, the negotiation process follows the determination of the implementation of the split configuration, wherein the split computing management module 422 determines or identifies at block 802 a specific neural network architecture to be implemented for performing ML operations. This may include determining the type of neural network (e.g., CNN, RNN, ANN, GAN, etc.) and the parameters of the neural network type, such as the number of layers, the number of nodes in each layer, the operations performed at each node, the weights and other parameters of each node, the connections between nodes, etc. Further recall that although the following describes the negotiation process for a client-initiated split, in other embodiments, it may be the server 112 that initiates the split, in which case a similar process may be implemented, but with the roles of the server 112 and the UE 104 reversed with respect to the initiator and the responder.

[0058] In the case where a neural network architecture is identified, at block 804, the split computing management module 422 then determines a proposed split configuration for the neural network architecture based on the current UE conditions, execution requirements, etc. Similar to the determination of whether to split the neural network, this proposed split configuration can be determined or selected algorithmically via a LUT or other selection structure, via smaller neural networks, and others. In some embodiments, multiple predetermined split options are available, and the split computing management module 422 selects one of the predetermined split options as the proposed split configuration. For example, one split option can be a two-way split that evenly divides the hidden layers between the server 112 and the UE 104, another split option can be a two-way split that assigns a larger share of the hidden layers to the server 112 compared to the share of the hidden layers assigned to the UE 104, and yet another split option can be a three-way split that assigns an initial split portion and a final split portion with a small proportion of the hidden layers to the UE 104 and assigns an intermediate split portion with the majority of the hidden layers to the server 112. In this example, the UE 104 can select the split option that best matches the current UE conditions while meeting the indicated execution requirements. In other embodiments, the split computing management module 422 can dynamically determine the proposed split, for example, by allocating multiple layers to the portion assigned to the server 112 inversely proportional to the available computing resources of the UE 104. In the case where the proposed split is identified, the UE 104 then transmits a split request 1008 ( Figure 10 ) to the server 112 via the BS 110. The split request 1008 can include a description or identifier of the proposed split (e.g., a data structure that identifies relevant parameters that fully describe the split portion that the UE 104 is proposing for the server 112 to implement, or an identifier of a predefined split option that can be obtained from the repository 118 as a representation 120 of the neural network). The split request 1008 can also include information related to the requested split proposal, such as the current conditions of the UE 104 and / or the execution requirements that led to the specific proposed split, and others. Additionally or alternatively, the server 112 can obtain some or all of the current conditions of the UE 104 from the most recent radio resource control (RRC) UE capability information message transmitted by the UE 104 in response to an RRC UE capability query message from the BS 110.

[0059] At block 806, the split calculation management module 514 of server 112 (or BS 110 or other network components acting as intermediaries for server 112) evaluates the proposed split represented by split request 1008 to determine whether to accept the proposed split or counter-propose with a modified proposed split. When determining whether to accept the proposed split, the split calculation management module 514 may consider the current resource constraints of the server itself and the indicated execution requirements. For example, in a case where server 112 is able to contribute the resources required to implement the portion of the neural network allocated to server 112 according to the proposal and will be able to utilize the proposed split to meet the indicated execution requirements, the split calculation management module 514 may accept the proposed split and indicate this acceptance at block 808 by transmitting a split calculation permission message 1010 ( Figure 10 ) to UE 104 via BS 110.

[0060] Returning to block 806, in a case where the split calculation management module 514 determines that server 112 cannot allocate the computing resources required to implement the proposed split or cannot meet the indicated execution requirements given the proposed split, then at block 810, the split calculation management module 514 may determine a split counter-proposal for a modified split configuration proposed for the neural network. For example, if server 112 cannot contribute sufficient resources for the original proposed split, the split calculation management module 514 may counter-propose with a counter-proposal split that allocates fewer layers of the neural network architecture to server 112. As another example, if server 112 can allocate sufficient resources, but the split calculation management module 514 determines that the latency requirement or other execution requirements cannot be met given the original proposed split, the split calculation management module 514 may counter-propose a modified split that meets the latency requirement or makes other execution requirements more likely to be met. Server 112 then transmits a representation of the split counter-proposal to UE 104. At block 812, the split calculation management module 422 of UE 104 evaluates the counter-proposal and, if acceptable, transmits an acceptance message to server 112 at block 814, in response to which server 112 issues a split calculation permission message 1010. If the counter-proposal is not acceptable, at block 816, UE 104 may terminate the negotiation and fall back to attempting to perform the ML operation entirely locally (i.e., by implementing the entire neural network at the ML module of UE 104) or by returning a no operation (NOP) or error message to the initiating software application to indicate that the ML operation cannot be performed under the current conditions. Alternatively, in some embodiments, one or more rounds of counter-counter-proposals for the split may be conducted between UE 104 and server 112 until a mutually acceptable split is identified or a threshold number of counter-proposals have been transmitted.

[0061] Returning toFigure 7 , in the case where the server 112 and the UE 104 agree on the proposed split configuration, the server 112 and the UE 104 continue to implement their respective split parts. However, to do so, the server 112 and the UE 104 need to implement the neural network architecture details (e.g., layers, nodes, connections, weights, and other parameters) of each split part as corresponding ML modules. In some cases, the negotiation process provides the distribution of this information. For example, when the negotiation process results in mutual agreement to adopt a predetermined split neural network configuration from, for example, the repository 118, once the agreed-upon split neural network configuration is determined, the split part implementation details can be explicitly included in the negotiation messaging or subsequently obtained by either or both of the server 112 or the UE 104 from the repository 118. Alternatively, when the split neural network configuration is self-organized by the UE 104 or the counter-proposed split from the server 112 is accepted by the UE 104, the implementation details of the split parts to be implemented by the server 112 and / or the UE 104 can be included in the negotiation messaging (e.g., included in the split request 1008 or the permission message 1010, respectively). However, in some embodiments, the negotiation can involve negotiation of the overall split but without specific details such as weights or other parameters for a particular node or for a particular layer. In such cases, as represented by block 712, the UE 104 transmits one or more configuration messages to the server 112 via the BS 110, the one or more configuration messages including the implementation details of the split part (hereinafter referred to as the "server-side split part") of the neural network to be implemented by the server 112 according to the negotiated split configuration. With the implementation details in place, at block 714, the UE 104 configures one or more of the ML modules (e.g., the UE ML module 128, Figure 1 ) to implement one or more split parts (hereinafter, the "UE-side split parts") assigned to the UE 104, and at block 716, the server 112 configures one or more of the ML modules (e.g., the server ML module 126, Figure 1 ) to implement one or more server-side split parts. In the case where the UE 104 and the server 112 are configured to implement their respective split neural network parts of the agreed-upon split neural network configuration for the neural network, at block 718, the server 112 and the UE 104 operate together to perform split calculation ML operations of the neural network distributed between the server 112 and the UE 104.

[0062] Figure 9Shows an example implementation of the split - computation ML operation execution process of block 718 according to some embodiments. As described above, a split neural network configuration can include a sequence of two or more split neural network parts, and depending on the implementation, the initial split part in the sequence can be assigned to UE 104 or server 112, and the final split part in the sequence can likewise be assigned to the UE or the server. Thus, for the purpose of describing the example implementation of the split - computation ML operation execution process below Figure 9 in reference to the order in which the devices in the sequence of the split configuration participate, one of UE 104 or server 112 that implements the initial split part in a given split configuration is referred to herein as the "first device", and the other of UE 104 or server 112 that implements the next split part after the initial split part in the sequence is referred to as the "second device". In the example showing this implementation Figure 10 server 112 is the first device and UE 104 is the second device, and the split configuration is a two - way split configuration, where server 112 implements the initial split part and UE 104 implements the final split part.

[0063] A typical neural network receives input data, processes the input data on various layers or other structures of the neural network, and then provides a final result output due to this processing of the input data. Depending on the purpose of the ML operation being performed and the source of the input data of the neural network implemented as a split configuration by UE 104 and server 112, the input data can be supplied by UE 104, by server 112, by another component, or a combination thereof. Regardless of the source, to start the execution of the ML operation, the input data is provided to the first device that implements the initial split part. Thus, at block 902, one or more sources of the input data transmit or otherwise provide access to the input data to be used for the ML operation to the first device. For example, in a split configuration where server 112 implements the initial split part (i.e., server 112 is the first device and UE 104 is the second device) but the ML operation involves input data obtained by UE 104 (e.g., an image captured by UE 104), UE 104 can transmit the image data to server 112 via BS 110. As another example, the image data can be an image sourced from a web page (e.g., the ML operation is an image - matching search), and thus, instead of directly providing the image data, UE 104 transmits a web address or other pointer to the image and the web page, and server 112 uses the pointer to obtain the image data from the web page.

[0064] It should be understood that in a cellular network or other wireless network, such as in Figure 1In the illustrated wireless communication network 100, BS 110 and UE 104 may adopt one or more resource allocation schemes to facilitate the uplink or downlink transmission of data between UE 104 and BS 110 in an efficient and timely manner. Such schemes may be appropriate in cases where UE 104 is to provide a part or all of the input data to server 112 for input at the initial layer of the neural network at server 112, and particularly so when there are maximum latency requirements imposed on the ML operations. Thus, after the split computing negotiation process and before transmitting the input data from UE 104 to server 112 via BS 110, UE 104 may notify BS 110 of its uplink transmission requirements for the timely wireless transmission of the input data from UE 104 to BS 110. For illustration purposes, with temporary reference to Figure 10 , after receiving the split computing permission message 1010, at block 1012, UE 104 identifies the input data to be provided by UE 104 to server 112, determines the uplink (UL) data requirements for transmitting this input data, and caches the input data in one or more UL buffers. At block 1012, UE 104 transmits a UL transmission request for the input data to BS 110 in the form of a UL buffer status report (BSR) message 1014 to notify BS 110 of the amount of data in its UL buffer for UL transmission. In response to the UL BSR message 1014, BS 110 transmits a UL grant to UE 104 in the form of, for example, a UL downlink control information (DCI) message 1016, which notifies UE 104 of the parameters to be used when UE 104 transmits the data (including the input data) in its UL buffer to BS 110, such as physical layer resource allocation, power control commands, and others. At block 1018, UE 104 wirelessly transmits the input data as UL data to BS 110 using the parameters indicated by the UL DCI message 1016, and BS 110 forwards the input data to server 112.

[0065] Returning to the reference Figure 9 , in the case where the input data is provided to the first device, at block 904, the first device receives the input data and processes (block 1020, Figure 10 ) the input data at the initial split portion implemented at the ML module of the first device to generate output data representing the intermediate result of the ML operation. At block 906, the intermediate result (block 1022, Figure 10 ) is transmitted to the second device. At block 908, the second device receives the intermediate result and provides the received intermediate result as input data to the next split portion implemented at the ML module of the second device, which processes (block 1024, Figure 10)Input data is used to generate output data. As shown by decision block 910, in the case where the next split portion is the last split portion in the sequence of split portions, the output data represents the final result of the neural network and thus represents the final result of the ML operation. And thus at block 912, one or both of the final results are transmitted to the first device such that the first device can take one or more actions in response to the final result, or at block 914, the second device can take one or more actions in response to the final result (block 1026, Figure 10 ). However, in the case where the next split portion is an intermediate split portion in the sequence of split portions, the output data represents an intermediate result. And thus at block 916, the intermediate result is transmitted to the first device (or to a third device in the case where three or more devices are implemented in the split neural network configuration). Thereby, these intermediate results are provided as input data to the next split configuration in the sequence to generate output data. In the case where the next split configuration is the final split configuration in the sequence, this output data can in turn be the final result of the neural network and be processed accordingly. Or in the case where one or more additional split portions remain in the sequence, the output data can be provided to the next device to be used as input data at a subsequent split configuration in the sequence, and so on, until the final split configuration in the sequence outputs the final result of the neural network.

[0066] The embodiments of the present disclosure can also be better understood by considering the following non-limiting examples: Example 1: A computer-implemented method, in a first device, the computer-implemented method includes: splitting a neural network into a split neural network configuration for at least the first device and a second device based on a set of one or more current conditions of the first device, the split neural network configuration specifying the distribution of elements of the neural network architecture of the neural network between a first neural network portion for the first device and a second neural network portion for the second device; implementing the first neural network portion at the first device; transmitting a representation of the second neural network portion to the second device; and processing first data at the first neural network portion to generate a first output. Example 2: The method according to Example 1, wherein: the first data includes input data for the neural network; the first output includes intermediate result data; and the method further includes: transmitting the intermediate result data to the second device for processing by the second neural network portion. Example 3: The method according to Example 2, further includes: receiving a final result output of the neural network from the second device; and performing at least one action at the first device in response to the final result output. Example 4: The method as described in Example 1 further includes: receiving the first data from the second device, the first data including intermediate result data generated by the second neural network portion at the second device; and wherein the first output is the final result output of the neural network. Example 5: The method as described in Example 4 further includes: performing at least one action at the first device based on the final result output of the neural network. Example 6: The method as described in Example 4 further includes: transmitting input data from the first device to the second device for processing by the second neural network portion at the second device. Example 7: The method as described in any one of Examples 1 to 6, wherein the set of one or more current conditions includes at least one of the following: battery condition; network condition; thermal condition; computing resource availability; or specified quality of service condition. Example 8: The method as described in any one of Examples 1 to 7 further includes: determining a latency condition for performing neural network operations using the neural network; and further determining, based on the latency condition, the split neural network configuration for at least the first device and the second device. Example 9: The method as described in Example 8 further includes: determining, based on at least one of the set of one or more current conditions or the latency condition, whether to perform the neural network operations entirely at the client device or by using the split configuration of the neural network; and in response to determining to use the split configuration of the neural network, determining the split neural network configuration for at least the first device and the second device. Example 10: The method as described in any one of Examples 1 to 9 further includes: transmitting a split computing request to the second device, the split computing request proposing the split neural network configuration; and wherein the first neural network portion is implemented at the first device in response to receiving permission for the split computing request from the second device. Example 11: The method as described in any one of Examples 1 to 9, wherein determining the split neural network configuration for at least the first device and the second device includes: transmitting a split computing request to the second device, the split computing request proposing a different split neural network configuration based on the set of one or more current conditions; receiving a split computing counter-proposal from the second device, the split computing request counter-proposal proposing the split neural network configuration; and wherein the first neural network portion is implemented at the first device in response to accepting the split computing counter-proposal based on the set of one or more current conditions. Example 12: The method as described in one of Examples 10 or 11, wherein the split computing request further indicates a latency condition for performing neural network operations using the split neural network configuration. Example 13. The method as described in any one of Examples 1 to 12, wherein the representation of the second neural network portion includes at least one of the following: data describing the neural network architecture of the second neural network portion; or an identifier of a predetermined split neural network configuration for the neural network represented in a repository accessible to the second device. Example 14. A computer-implemented method, in a second device, the computer-implemented method comprising: receiving, from a first device, an indication of a split neural network configuration for at least the first device and the second device, the split neural network configuration specifying a distribution of elements of a neural network architecture of a neural network between a first neural network portion for at least the first device and a second neural network portion for the second device; implementing the second neural network portion at the second device; receiving first data from the first device; and processing the first data at the second neural network portion to generate a first output. Example 15: The method as described in Example 14, wherein: the first data includes input data for the neural network; the first output includes intermediate result data; and the method further comprises: transmitting the intermediate result data to the first device for processing at the first neural network portion. Example 16: The method as described in Example 14, wherein: the first data includes intermediate result data generated by the first neural network portion at the first device; and the first output is the final result output of the neural network. Example 17: The method as described in Example 16, further comprising at least one of the following: performing at least one action at the second device in response to the final result output; or transmitting the final result output to the first device. Example 18: The method as described in any one of Examples 14 to 17, wherein: receiving the indication of the split neural network configuration includes receiving a split computing request proposing the split neural network configuration; the method further comprises determining whether to permit the split computing request; and wherein the second neural network portion is implemented at the second device in response to determining to permit the split computing request. Example 19: The method as described in Example 18, further comprising: receiving, from the first device, an indication of a latency condition for performing neural network operations using the neural network; and determining whether to permit the split computing request is based on the latency condition. Example 20: The method according to any one of Examples 1 to 19, wherein the second neural network portion is implemented at the second device in response to receiving a representation of the second neural network portion from the first device, the representation of the second neural network portion including at least one of the following: data representing a neural network architecture of the second neural network portion; or an identifier of a predetermined split neural network configuration for the neural network represented in a repository accessible to the second device. Example 21: The method according to any one of Examples 1 to 20, wherein the first device is a user equipment, and the second device is one or more servers connected to a network wirelessly accessible by the user equipment. Example 22: The method according to any one of Examples 1 to 21, wherein the neural network is a deep neural network (DNN) model. Example 23: The method according to any one of Examples 1 to 22, wherein the split neural network configuration assigns a first set of successive layers of the neural network to the first neural network portion, and assigns a second set of successive layers of the neural network to the second neural network portion, the second set being adjacent to the first set. Example 24: The method according to any one of Examples 1 to 22, wherein the split neural network configuration includes a two-way split having an initial neural network portion and a final neural network portion, wherein the initial neural network portion is assigned to one of the first device or the second device, and the final neural network portion is assigned to the other of the first device or the second device. Example 25: The method according to any one of Examples 1 to 22, wherein the split neural network configuration includes a three-way split having an initial neural network portion, an intermediate neural network portion, and a final neural network portion, wherein the initial neural network portion and the final neural network portion are assigned to one of the first device or the second device, and the intermediate neural network portion is assigned to the other of the first device or the second device. Example 26. A device, comprising: a network interface; at least one processor, the at least one processor being coupled to the network interface; and a memory storing executable instructions configured to manipulate the at least one processor to execute the method according to any one of Examples 1 to 25. Example 27: The device according to Example 26, wherein the first device is a user equipment of a cellular network, and the second device is one or more servers of a network wirelessly accessible by the user equipment. Example 28: The device according to Example 26 or 27, wherein the device is the first device. Example 29: The apparatus as described in Example 26 or 27, wherein the apparatus is the second apparatus.

[0067] In some embodiments, certain aspects of the techniques described above may be implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored or otherwise tangibly embodied on a non-transitory computer-readable storage medium. The software may include instructions and certain data that, when executed by one or more processors, cause the one or more processors to perform one or more aspects of the techniques described above. The non-transitory computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as flash memory, a cache, a random access memory (RAM), or one or more other non-volatile memory devices. The executable instructions stored on the non-transitory computer-readable storage medium may be in source code, assembly language code, object code, or another instruction format that is interpreted or otherwise executable by one or more processors.

[0068] A computer-readable storage medium may include any storage medium or combination of storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact disc (CD), digital versatile disc (DVD), Blu-ray disc), magnetic media (e.g., floppy disk, magnetic tape, or magnetic hard drive), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. The computer-readable storage medium may be embedded in the computing system (e.g., system RAM or ROM), fixedly attached to the computing system (e.g., magnetic hard drive), removably attached to the computing system (e.g., optical disc or universal serial bus (USB)-based flash memory), or coupled to the computer system via a wired or wireless network (e.g., network-attached storage device (NAS)).

[0069] It should be noted that not all activities or elements described above in the general description are required, some parts of a specific activity or apparatus may not be required, and one or more further activities may be performed or one or more further elements may be included in addition to those activities or elements described. Further still, the order in which the activities are listed is not necessarily the order in which the activities are performed. Additionally, concepts have been described with reference to specific embodiments. However, those of ordinary skill in the art will understand that various modifications and changes can be made without departing from the scope of the disclosure set forth in the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the disclosure.

[0070] Advantages, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, an advantage, a solution to a problem, and any feature that may cause any advantage, advantage, or solution to occur or become more pronounced should not be construed as a critical, required, or essential feature of any or all of the claims. Moreover, the specific embodiments disclosed above are merely illustrative, since the disclosed subject matter may be modified and practiced in different but equivalent manners that will be apparent to those skilled in the art benefiting from the teachings herein. Except as described in the appended claims, no limitation is intended as to the details of the construction or design shown herein. It is, therefore, evident that the specific embodiments disclosed above may be altered or modified and all such variations are considered to be within the scope of the disclosed subject matter. Accordingly, the protection sought herein is as set forth in the appended claims.

Claims

1. A computer-implemented method, in a first device, the computer-implemented method comprises: Based on a set of one or more current conditions of the first device, splitting a neural network into a split neural network configuration for at least the first device and a second device, the split neural network configuration specifying a distribution of elements of a neural network architecture of the neural network between a first neural network portion for at least the first device and a second neural network portion for the second device; Implementing the first neural network portion at the first device; Transmitting a representation of the second neural network portion to the second device; and Processing first data at the first neural network portion to generate a first output.

2. The method according to claim 1, wherein: The first data includes input data for the neural network; The first output includes intermediate result data; and The method further comprises: Transmitting the intermediate result data to the second device for processing by the second neural network portion.

3. The method according to claim 2, further comprises: Receiving a final result output of the neural network from the second device; and Performing at least one action at the first device in response to the final result output.

4. The method according to claim 1, further comprises: Receiving the first data from the second device, the first data including intermediate result data generated by the second neural network portion at the second device; and wherein the first output is a final result output of the neural network.

5. The method according to claim 4, further comprises: Performing at least one action at the first device based on the final result output of the neural network.

6. The method according to claim 4, further comprises: Transmitting input data from the first device to the second device for processing by the second neural network portion at the second device.

7. The method according to any one of claims 1 to 6, wherein, The set of one or more current conditions includes at least one of the following: battery condition; network condition; thermal condition; computing resource availability; or specified quality of service condition.

8. The method according to any one of claims 1 to 7, further comprises: Determining a latency condition for performing a neural network operation using the neural network; and Further determining the split neural network configuration for at least the first device and the second device based on the latency condition.

9. The method according to claim 8, further comprises: Based on at least one of the set of one or more current conditions or the latency condition, determining whether to perform the neural network operation entirely at the first device or by using the split configuration of the neural network; and In response to determining to use the split configuration of the neural network, determining the split neural network configuration for at least the first device and the second device.

10. The method according to any one of claims 1 to 9, further comprises: Transmit a split calculation request to the second device, the split calculation request proposing the split neural network configuration; And wherein, the first neural network part is implemented at the first device in response to receiving permission for the split calculation request from the second device.

11. The method according to any one of claims 1 to 9, wherein, determining the split neural network configuration for at least the first device and the second device includes: Transmit a split calculation request to the second device, the split calculation request proposing a different split neural network configuration based on the set of one or more current conditions; Receive a split calculation counter-proposal from the second device, the split calculation counter-proposal proposing the split neural network configuration; and wherein, the first neural network part is implemented at the first device in response to accepting the split calculation counter-proposal based on the set of one or more current conditions.

12. The method according to one of claims 10 or 11, wherein, the split calculation request further indicates a latency condition for performing neural network operations using the split neural network configuration.

13. The method according to any one of claims 1 to 12, wherein, the representation of the second neural network part includes at least one of the following: data describing the neural network architecture of the second neural network part; or an identifier of a predetermined split neural network configuration for the neural network represented in a repository accessible to the second device.

14. A computer-implemented method, in a second device, the computer-implemented method comprises: Receive an indication of a split neural network configuration for at least a first device and the second device from a first device, the split neural network configuration specifying a distribution of elements of a neural network architecture of a neural network between a first neural network part for at least the first device and a second neural network part for the second device; Implement the second neural network part at the second device; Receive first data from the first device; And Process the first data at the second neural network part to generate a first output.

15. The method according to claim 14, wherein: the first data includes input data for the neural network; the first output includes intermediate result data; and the method further comprises: Transmit the intermediate result data to the first device for processing at the first neural network part.

16. The method according to claim 14, wherein: the first data includes intermediate result data generated by the first neural network part at the first device; and the first output is the final result output of the neural network.

17. The method according to claim 16, further comprising at least one of the following: Perform at least one action at the second device in response to the final result output; or Transmit the final result output to the first device.

18. The method according to any one of claims 14 to 17, wherein: Receiving the indication of the split neural network configuration includes receiving a split computing request that proposes the split neural network configuration; The method further includes determining whether to permit the split computing request; and wherein, the second neural network portion is implemented at the second device in response to determining to permit the split computing request.

19. The method according to claim 18, further comprising: Receiving an indication of a latency condition for performing a neural network operation using the neural network from the first device; and Determining whether to permit the split computing request is based on the latency condition.

20. The method according to any one of claims 1 to 19, wherein, The second neural network portion is implemented at the second device in response to receiving a representation of the second neural network portion from the first device, the representation of the second neural network portion including at least one of the following: data representing the neural network architecture of the second neural network portion; or an identifier of a predetermined split neural network configuration for the neural network represented in a repository accessible to the second device.

21. The method according to any one of claims 1 to 20, wherein, The first device is a user equipment, and the second device is one or more servers connected to a network that can be wirelessly accessed by the user equipment.

22. The method according to any one of claims 1 to 21, wherein, The neural network is a deep neural network (DNN) model.

23. The method according to any one of claims 1 to 22, wherein, The split neural network configuration assigns a first set of successive layers of the neural network to the first neural network portion, and assigns a second set of successive layers of the neural network to the second neural network portion, the second set being adjacent to the first set.

24. The method according to any one of claims 1 to 22, wherein, The split neural network configuration includes a two-way split having an initial neural network portion and a final neural network portion, wherein the initial neural network portion is assigned to one of the first device or the second device, and the final neural network portion is assigned to the other of the first device or the second device.

25. The method according to any one of claims 1 to 22, wherein, The split neural network configuration includes a three-way split having an initial neural network portion, an intermediate neural network portion, and a final neural network portion, wherein the initial neural network portion and the final neural network portion are assigned to one of the first device or the second device, and the intermediate neural network portion is assigned to the other of the first device or the second device.

26. An apparatus, comprising: A network interface; At least one processor, the at least one processor being coupled to the network interface; and A memory storing executable instructions configured to manipulate the at least one processor to perform the method according to any one of claims 1 to 25.

27. The apparatus according to claim 26, wherein, The first device is a user equipment of a cellular network, and the second device is one or more servers of a network to which the user equipment can wirelessly access.

28. The device according to claim 26 or 27, wherein, the device is the first device.

29. The device according to claim 26 or 27, wherein, the device is the second device.