Method, apparatus, and system for artificial intelligence (AI) model splitting

The method for AI model splitting in wireless communication systems addresses the inefficiencies in existing techniques by allowing flexible model computation splitting between devices, thereby enhancing performance and resource management for AI in future wireless technologies.

WO2025092059A1PCT designated stage expired Publication Date: 2025-05-08HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/108012
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-07-27
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing communication techniques struggle to maximize the efficient use of signal space in wireless communications, and there is a need to address issues related to the use of Artificial Intelligence (AI) in wireless networks, particularly for future generations of wireless technologies.

Method used

A method for AI model splitting, where a first device transmits a service request message to a second device for performing a task based on a model, and based on receiving a service response message including model split information, the first device performs the task. The model split information indicates a split of model computation for performing the task between the first device and the second device, allowing for flexible model splitting and improved performance.

Benefits of technology

The proposed solution enables flexible model splitting, thereby improving the performance of AI models in wireless communication systems. It allows for efficient resource allocation and latency management, enhancing the capability of AI in future wireless technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024108012_08052025_PF_FP_ABST
    Figure CN2024108012_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Example embodiments relate to a sensing measurement and reporting. In an aspect, a first device (201) transmits a service request message to a second device (202) for performing a task based on a model. Based on receiving a service response message including model split information from the second device (202), the first device (201) performs the task based on the model and the model split information, and the model split information indicates a split of model computation for performing the task at least between the first device (201) and the second device (202). Based on receiving a deny notification message from the second device (202), the first device (201) suspends the performing of the task based on the model.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, APPARATUS, AND SYSTEM FOR ARTIFICIAL INTELLIGENCE (AI) MODEL SPLITTING

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The present application claims priority to U.S. Provisional Application No. 63 / 594,107, filed on October 30, 2023, and entitled “Service request and handshake procedures for artificial intelligence (AI) and machine learning (ML) ” , the disclosure of which is incorporated by reference herein.FIELD

[0003] Example embodiments of the present disclosure generally relate to the field of communication and in particular, to methods, devices, apparatuses and computer readable storage medium for a split of a model.BACKGROUND

[0004] Artificial intelligence (AI) , and in particular deep machine learning, is a wide-ranging branch of computer science concerned with building smart machines capable of performing tasks that typically require human intelligence. It is expected that the introduction of AI will create a paradigm shift in virtually every sector of the tech industry and AI is expected to play a role in advancement of networking and communication technologies. For example, existing communication techniques, which rely on classical analytical modeling of channels, have enabled wireless communications to take place at close to the theoretical Shannon limit. However, existing techniques may be unsatisfactory to further maximize efficient use of the signal space. AI is expected to help address this challenge. Other aspects of wireless communication may benefit from the use of AI, particularly in future generations of wireless technologies, such as technologies in advanced 5G and future systems, and beyond. For the use of AI in a wireless network, there are still some issues to be addressed.SUMMARY

[0005] In general, example embodiments of the present disclosure provide a solution for a split of a model.

[0006] In a first aspect, there is provided a method. The method comprises transmitting, at a first device to a second device, a service request message for performing a task based on a model, based on receiving, from the second device, a service response message including model split information, performing the task based on the model and the model split information, and the model split information indicates a split of model computation for performing the task at least between the first device and the second device, and based on receiving a deny notification message from the second device, suspending the performing of the task based on the model. As such, the model may be split flexibly. Therefore, the performance for the model is improved.

[0007] In some embodiments, the first device may further transmit at least one of computation capability or computation resource of the first device to the second device for performing the task based on the model. In this way, the split of model computation may be determined based on the computation capability or computation resource of the first device.

[0008] In some embodiments, the service request message may comprise information indicating a service latency for performing the task, a priority of the task, or a combination of the above-mentioned two items. In this way, resource conflicts and latency may be considered.

[0009] In some embodiments, the first device may further determine computation time to complete a first part of the task by the first device at a reference model split point, and determine a residual latency based on a service latency for  performing the task and the computation time. In this way, the computation capability of the first device may not be reported to the second device.

[0010] In some embodiments, the first device may further estimate amounts of model computation at multiple candidate model split points, and select, based on the amounts of model computation, a candidate model split point from the multiple candidate model split points as the reference model split point. In this way, the reference model split point may be provided to the second device.

[0011] In some embodiments, the residual latency may comprise first transmission time for the first device to transmit, to the second device, an intermediate result of performing a first part of the task by the first device, computation time at the second device for performing a second part of the task, and second transmission time for the second device to transmit a result of performing the task to the first device. In this way, the communication time and the computation time may be considered into the latency budget.

[0012] In some embodiments, the service request message may comprise information indicating the residual latency, the reference model split point, a priority of the task, or any combination of two or more of the above-mentioned items. In this way, the reference model split point may be provided to the second device.

[0013] In some embodiments, the service response message may further comprise resource assignment for transmission of an intermediate result of performing a first part of the task by the first device. In this way, the resource for transmission of an intermediate result may be provided to the first device.

[0014] In some embodiments, performing the task based on the model and the model split information may comprise based on the model performing a first part of the task indicated by the model split information, transmitting an intermediate result of performing the first part of the task to the second device, and receiving a result of performing the task from the second device. In this way, the task may be performed at the first device and the second device.

[0015] In some embodiments, the deny notification message comprises information indicating a backoff timer, and the first device may further transmit, to the second device, a further service request message for performing the task after the backoff timer expires. In this way, the service request message may be resent in a predefined period.

[0016] In some embodiments, the backoff timer may be defined based on a priority of the service request message and a random number. In this way, a period of transmitting the service request message may be associated with the priority of the service request message.

[0017] In some embodiments, the first device may further transmit a first confirmation message to the second device for the model split information. In this way, the first device may confirm the decision of the model split.

[0018] In some embodiments, the first device may further transmit a registration request message for registering the task to the second device, and receive a model notification message indicating the model to be used for performing the task from the second device. In this way, the model to be used may be indicated to the first device.

[0019] In some embodiments, the registration request message may comprise a service latency for performing the task, an accuracy requirement for the task, or a combination of the above-mentioned two items. In this way, the service latency and the accuracy requirement may be indicated to the first device.

[0020] In some embodiments, the first device may further transmit a second confirmation message for the model to be used to the second device. In this way, the first device may confirm the model to be used.

[0021] In some embodiments, the model to be used for performing the task is determined by a third device, and the first device may further transmit a third confirmation message for the model to be used to the third device. In this way, the first device may confirm the model to be used with the third device.

[0022] In some embodiments, the model notification message may comprise information indicating multiple devices which are to participate in performing the task, wherein the multiple devices include the second device.

[0023] In some embodiments, the first device may further transmit a fourth confirmation message for the model to be used to a fourth device among the multiple devices. In this way, the first device may confirm the model to be used with the fourth device.

[0024] In some embodiments, the first device may further receive an indication of an anchor device for the task among the multiple devices from the second device or the third device. In this way, the anchor device may be indicated to the first device.

[0025] In some embodiments, the first device may be a terminal device, the second device may be an access network device, the third device may be a core network device or an access network device, or the fourth device may be an access network device. In this way, the terminal device may perform the task with one or more network devices.

[0026] In a second aspect, there is provided a method at a second device. The method comprising: receiving a service request message for performing a task based on a model at a second device from a first device, determining whether the task is to be completed within a service latency in case of a split of model computation for performing the task at least between the first device and the second device, based on determining that the task is to be completed within the service latency, transmitting, to the first device, a service response message including model split information, wherein the model split information indicates the split of the model computation, and based on determining that the task is not to be completed within the service latency, transmitting a deny notification message to the first device. As such, the model may be split flexibly. Therefore, the performance for the model is improved.

[0027] In some embodiments, the second device may further receive, from the first device, at least one of computation capability or computation resource of the first device for performing the task based on the model. In this way, the split of model computation may be determined based on the computation capability or computation resource of the first device. sd

[0028] In some embodiments, the service request message may comprise information indicating the service latency for performing the task, a priority of the task, or any combination of two or more of the above-mentioned items. In this way, resource conflicts and latency may be considered.

[0029] In some embodiments, determining whether the task is to be completed within the service latency may comprises determining an estimated latency for performing the task based on the at least one of the computation capability or the computation resource, and comparing the estimated latency and the service latency. In this way, the split of model computation may be determined based on the computation capability or computation resource of the first device.

[0030] In some embodiments, the estimated latency may comprise an estimation of computation time for the first device to perform a first part of the task in case of the computation capability of the first device and a model split point, an estimation of computation time for the second device to perform a second part of the task in case of computation capability of the second device and the model split point, and an estimation of communication time between the first device and the second device in case of a data volume associated with the model split point. In this way, the communication time and the computation time may be considered into the estimated latency.

[0031] In some embodiments, the service request message may comprise information indicating a residual latency determined based on the service latency for performing the task and computation time to complete a first part of the task by the first device at a reference model split point, the reference model split point, a priority of the task, or any combination of two or more of the above-mentioned items. In this way, the reference model split point may be provided to the second device.

[0032] In some embodiments, the residual latency may comprise first transmission time for the first device to transmit, to the second device, an intermediate result of performing a first part of the task by the first device, computation time at the second device for performing a second part of the task, and second transmission time for the second device to transmit a result of performing the task to the first device. In this way, the communication time and the computation time may be considered into the latency budget.

[0033] In some embodiments, determining whether the task is to be completed within the latency may comprise determining an estimated latency by adding estimations of the first transmission time, the computation time, and the second transmission time, and comparing the estimated latency and the residual latency. In this way, whether the task is to be completed may be determined based on the estimated latency and the residual latency.

[0034] In some embodiments, the service response message may further comprise resource assignment for transmission of an intermediate result of performing a first part of the task by the first device. In this way, the resource for transmission of an intermediate result may be provided to the first device.

[0035] In some embodiments, the second device may further perform the task based on the model and the model split information based on transmitting the service response message to the first device. In this way, the task may be performed at the second device.

[0036] In some embodiments, performing the task based on the model and the model split information may comprise: receiving, from the first device, an intermediate result of performing a first part of the task, performing, based on the intermediate result and the model, a second part of the task indicated by the model split information; and transmitting a result of performing the task to the first device. In this way, the task may be performed at the first device and the second device.

[0037] In some embodiments, performing the task based on the model and the model split information may comprise transmitting an indication to at least one other device that the at least one other device performs a third part of the task based on the model, and receiving an intermediate result of performing the third part of the task from the at least one other device. In this way, the task may be performed at the first device and the second device.

[0038] In some embodiments, the second device may further determine an estimated latency for performing the task by including computation time for the at least one other device to perform the third part of the task and an information exchange latency between the second device and the at least one other device. In this way, the communication time with the at least one other device may be considered into the estimated latency.

[0039] In some embodiments, the deny notification message comprises information indicating a backoff timer, and the second device may further receive a further service request message from the first device for performing the task after the backoff timer expires. In this way, the service request message may be resent in a predefined period.

[0040] In some embodiments, the backoff timer may be defined based on a priority of the service request message and a random number. In this way, a period of transmitting the service request message may be associated with the priority of the service request message.

[0041] In some embodiments, the second device may further receive a first confirmation message for the model split information from the first device. In this way, the first device may confirm the decision of the model split.

[0042] In some embodiments, the second device may further receive a registration request message from the first device for registering the task, determine a target model as the model to be used for performing the task, and transmit a model notification message indicating the model to be used for performing the task to the first device. In this way, the model to be used may be indicated to the first device.

[0043] In some embodiments, determining the target model as the model may comprise deciding, by the second device, the target model as the model to be used for performing the task. In this way, the model to be used may be determined by the second device.

[0044] In some embodiments, determining the target model as the model may comprise forwarding the registration request message to a third device, and receiving an indication from the third device that the target model is the model to be used for performing the task. In this way, the model to be used may be determined by the third device.

[0045] In some embodiments, the second device may further receive an indication from the third device that a fourth device is to participate in performing the task based on the model. In this way, the second device may be notified of a fourth device participating in the task.

[0046] In some embodiments, the registration request message may comprise a service latency for performing the task, an accuracy requirement for the task, or a combination of the above-mentioned two items. In this way, the service latency and the accuracy requirement may be indicated to the first device.

[0047] In some embodiments, the second device may further receive a second confirmation message from the first device for the model to be used. In this way, the first device may confirm the model to be used.

[0048] In some embodiments, the second device may further transmit an indication that the model is to be used for performing the task to a fourth device which is to participate in performing the task. In this way, the model to be used may be indicated to the fourth device.

[0049] In some embodiments, the model notification message may comprise information indicating multiple devices which are to participate in performing the task, wherein the multiple devices include the second device and the fourth device. In this way, the first device may be notified of multiple devices participating in the task.

[0050] In some embodiments, the second device may further transmit an indication of an anchor device to the first device for the task among the multiple devices. In this way, the anchor device may be indicated to the first device.

[0051] In some embodiments, the first device may be a terminal device, the second device may be an access network device, the third device may be a core network device or an access network device, or the fourth device may be an access network device. In this way, the terminal device may perform the task with one or more network devices.

[0052] In a third aspect, there is provided a first device. The first device comprises a transceiver and a processor communicatively coupled with the transceiver. The processor is configured to transmit, at a first device to a second device, a service request message for performing a task based on a model, based on receiving, from the second device, a service response message including model split information, perform the task based on the model and the model split information, wherein the model split information indicates a split of model computation for performing the task at least between the first device and the second device, and based on receiving a deny notification message from the second device, suspend the performing of the task based on the model.

[0053] In a fourth aspect, there is provided a second device. The second device comprises a transceiver and a processor communicatively coupled with the transceiver. The processor is configured to receive, at a second device from a first device, a service request message for performing a task based on a model, determine whether the task is to be completed within a service latency in case of a split of model computation for performing the task at least between the first device and the second device, based on determining that the task is to be completed within the service latency, transmit, to the first device, a service response message including model split information, wherein the model split information indicates the split of the model computation, and based on determining that the task is not to be completed within the service latency, transmit a deny notification message to the first device.

[0054] In a fifth aspect, there is provided a non-transitory computer readable medium comprising computer program stored thereon, the computer program, when executed on at least one processor, causing the at least one processor to perform the method of any one of the first aspect or second aspect.

[0055] In a sixth aspect, there is provided a chip comprising at least one processing circuit configured to perform the method of any one of the first aspect or second aspect.

[0056] In a seventh aspect, there is provided a computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions which, when executed, cause an apparatus to perform the method of any one of the first aspect or second aspect.

[0057] It is to be understood that the summary section is not intended to identify key or essential features of embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will become easily comprehensible through the following description.BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Some example embodiments will now be described with reference to the accompanying drawings, in which:

[0059] FIG. 1A illustrates an example communication system in which example embodiments of the present disclosure may be implemented;

[0060] FIG. 1B illustrates an example communication system in which example embodiments of the present disclosure may be implemented;

[0061] FIG. 1C illustrates an example of an electronic device (ED) and base stations related to some embodiments of the present disclosure;

[0062] FIG. 1D illustrates an example of units or modules in a device related to some embodiments of the present disclosure;

[0063] FIG. 2 illustrates an example signaling chart illustrating an example process according to some embodiments of the present disclosure;

[0064] FIG. 3 illustrates a first example procedure of proposed solution according to some embodiments of the present disclosure;

[0065] FIG. 4 illustrates a second example procedure of proposed solution according to some embodiments of the present disclosure;

[0066] FIG. 5 illustrates a third example procedure of proposed solution according to some embodiments of the present disclosure;

[0067] FIG. 6 illustrates a fourth example procedure of proposed solution according to some embodiments of the present disclosure;

[0068] FIG. 7 illustrates a fifth example procedure of proposed solution according to some embodiments of the present disclosure;

[0069] FIG. 8 illustrates a sixth example procedure of proposed solution according to some embodiments of the present disclosure;

[0070] FIG. 9 illustrates a flowchart of a method implemented at a first device according to some embodiments of the present disclosure;

[0071] FIG. 10 illustrates a flowchart of a method implemented at a second device according to some embodiments of the present disclosure;

[0072] FIG. 11 is a block diagram of a device that may be used for implementing some embodiments of the present disclosure;

[0073] FIG. 12 is a schematic diagram of a structure of an apparatus in accordance with some embodiments of the present disclosure; and

[0074] FIG. 13 is a schematic diagram of a structure of an apparatus in accordance with some embodiments of the present disclosure

[0075] Throughout the drawings, the same or similar reference numerals represent the same or similar elements.DETAILED DESCRIPTION

[0076] Principle of the present disclosure will now be described with reference to some example embodiments. It is to be understood that these embodiments are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.

[0077] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0078] References in the present disclosure to “one embodiment” , “an embodiment” , “an example embodiment” , and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0079] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0080] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0081] FIG. 1A illustrates an example communication system 100A in which example embodiments of the present disclosure may be implemented. Referring to FIG. 1A, as an illustrative example without limitation, a simplified schematic illustration of a communication system is provided. The communication system 100A comprises a radio access network 120. The radio access network 120 may be a next generation radio access network, or a legacy (e.g. 5G, 4G, 3G or 2G) radio access network. One or more communication electric device (ED) 110a, 110b, 110c, 110d, 110e, 110f, 110g, 110h, 110i, 110j (generically referred to as 110) may be interconnected to one another or connected to one or more network nodes (170a, 170b, generically referred to as 170) in the radio access network 120. A core network 130 may be a part of  the communication system and may be dependent or independent of the radio access technology used in the communication system 100A. Also the communication system 100A comprises a public switched telephone network (PSTN) 140, the internet 150, and other networks 160.

[0082] FIG. 1B illustrates an example communication system in which example embodiments of the present disclosure may be implemented. In general, the communication system 100B enables multiple wireless or wired elements to communicate data and other content. The purpose of the communication system 100B may be to provide content, such as voice, data, video, signaling and / or text, via broadcast, multicast and unicast, etc. The communication system 100B may operate by sharing resources, such as carrier spectrum bandwidth, between its constituent elements. The communication system 100B may include a terrestrial communication system and / or a non-terrestrial communication system. The communication system 100B may provide a wide range of communication services and applications (such as earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery and mobility, etc. ) . The communication system 100B may provide a high degree of availability and robustness through a joint operation of a terrestrial communication system and a non-terrestrial communication system. For example, integrating a non-terrestrial communication system (or components thereof) into a terrestrial communication system can result in what may be considered a heterogeneous network comprising multiple layers. Compared to conventional communication networks, the heterogeneous network may achieve better overall performance through efficient multi-link joint operation, more flexible functionality sharing, and faster physical layer link switching between terrestrial networks and non-terrestrial networks.

[0083] The terrestrial communication system and the non-terrestrial communication system could be considered sub-systems of the communication system. In the example shown in FIG. 1B, the communication system 100B includes electronic devices (ED) 110a, 110b, 110c, 110d (generically referred to as ED 110) , radio access networks (RANs) 120a-120b, a non-terrestrial communication network 120c, a core network 130, a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160. The RANs 120a-120b include respective base stations (BSs) 170a-170b, which may be generically referred to as terrestrial transmit and receive points (T-TRPs) 170a-170b. The non-terrestrial communication network 120c includes an access node 172, which may be generically referred to as a non-terrestrial transmit and receive point (NT-TRP) 172, or a sensing agent 172.

[0084] Any ED 110 may be alternatively or additionally configured to interface, access, or communicate with any T-TRP 170a-170b and NT-TRP 172, the Internet 150, the core network 130, the PSTN 140, the other networks 160, or any combination of the preceding. In some examples, ED 110a may communicate an uplink and / or downlink transmission over a terrestrial air interface 190a with T-TRP 170a. In some examples, the EDs 110a, 110b, 110c and 110d may also communicate directly with one another via one or more sidelink air interfaces 190b. In some examples, ED 110d may communicate an uplink and / or downlink transmission over a non-terrestrial air interface 190c with NT-TRP 172.

[0085] The air interfaces 190a and 190b may use similar communication technology, such as any suitable radio access technology. For example, the communication system 100B may implement one or more channel access methods, such as code division multiple access (CDMA) , space division multiple access (SDMA) , time division multiple access (TDMA) , frequency division multiple access (FDMA) , orthogonal FDMA (OFDMA) , Direct Fourier Transform spread OFDMA (DFT-OFDMA) or single-carrier FDMA (SC-FDMA) in the air interfaces 190a and 190b. The air interfaces 190a and 190b may utilize other higher dimension signal spaces, which may involve a combination of orthogonal and / or non-orthogonal dimensions.

[0086] The non-terrestrial air interface 190c can enable communication between the ED 110d and one or multiple NT-TRPs 172 via a wireless link or simply a link. For some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of EDs 110 and one or multiple NT-TRPs 172 for multicast transmission.

[0087] The RANs 120a and 120b are in communication with the core network 130 to provide the EDs 110a 110b, and 110c with various services such as voice, data, and other services. The RANs 120a and 120b and / or the core network 130 may be in direct or indirect communication with one or more other RANs (not shown) , which may or may not be directly served by core network 130, and may or may not employ the same radio access technology as RAN 120a, RAN 120b or both. The core network 130 may also serve as a gateway access between (i) the RANs 120a and 120b or EDs 110a 110b, and 110c or both, and (ii) other networks (such as the PSTN 140, the Internet 150, the sensing agent 172, and the other networks 160) . In addition, some or all of the EDs 110a 110b, and 110c may include functionality for communicating with different wireless networks over different wireless links using different wireless technologies and / or protocols. Instead of wireless communication (or in addition thereto) , the EDs 110a 110b, and 110c may communicate via wired communication channels to a service provider or switch (not shown) , and to the Internet 150. PSTN 140 may include circuit switched telephone networks for providing plain old telephone service (POTS) . Internet 150 may include a network of computers and subnets (intranets) or both, and incorporate protocols, such as Internet Protocol (IP) , Transmission Control Protocol (TCP) , User Datagram Protocol (UDP) . EDs 110a 110b, and 110c may be multimode devices capable of operation according to multiple radio access technologies, and incorporate multiple transceivers necessary to support such.

[0088] FIG. 1C illustrates an example of an electronic device (ED) and a base station related to some embodiments of the present disclosure. As shown in FIG. 1C, another example of an ED 110 and a base station 170a, 170b and / or 170c is provided. The ED 110 is used to connect persons, objects, machines, etc. The ED 110 may be widely used in various scenarios, for example, cellular communications, device-to-device (D2D) , vehicle to everything (V2X) , peer-to-peer (P2P) , machine-to-machine (M2M) , machine-type communications (MTC) , Internet of things (IOT) , virtual reality (VR) , augmented reality (AR) , mixed reality (MR) , metaverse, digital twin, industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc.

[0089] Each ED 110 represents any suitable end user device for wireless operation and may include such devices (or may be referred to) as a user equipment / device (UE) , a wireless transmit / receive unit (WTRU) , a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a station (STA) , a machine type communication (MTC) device, a personal digital assistant (PDA) , a smartphone, a laptop, a computer, a tablet, a wireless sensor, a consumer electronics device, a smart book, a vehicle, a car, a truck, a bus, a train, or an IoT device, wearable devices such as a watch, head mounted equipment, a pair of glasses, an industrial device, or apparatus (e.g. communication module, modem, or chip) in the forgoing devices, among other possibilities. Future generation EDs 110 may be referred to using other terms. The base station 170a and 170b is a T-TRP and will hereafter be referred to as T-TRP 170. Also shown in FIG. 1C, a NT-TRP will hereafter be referred to as NT-TRP 172. Each ED 110 connected to T-TRP 170 and / or NT-TRP 172 can be dynamically or semi-statically turned-on (i.e., established, activated, or enabled) , turned-off (i.e., released, deactivated, or disabled) and / or configured in response to one of more of: connection availability and connection necessity.

[0090] The ED 110 includes a transmitter 111 and a receiver 113 coupled to one or more antennas 104. Only one antenna 104 is illustrated. One, some, or all of the antennas 104 may alternatively be panels. The transmitter 111 and the receiver 113 may be integrated, e.g. as a transceiver. The transceiver is configured to modulate data or other content for transmission by at least one antenna 104 or network interface controller (NIC) . The transceiver is also configured to demodulate data or other content received by the at least one antenna 104. Each transceiver includes any suitable structure for generating signals for wireless or wired transmission and / or processing signals received wirelessly or by wire. Each antenna 104 includes any suitable structure for transmitting and / or receiving wireless or wired signals.

[0091] The ED 110 includes at least one memory 115. The memory 115 stores instructions and data used, generated, or collected by the ED 110. For example, the memory 115 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by one or more  processing unit (s) (e.g., a processor 117) . Each memory 115 includes any suitable volatile and / or non-volatile storage and retrieval device (s) . Any suitable type of memory may be used, such as random access memory (RAM) , read only memory (ROM) , hard disk, optical disc, subscriber identity module (SIM) card, memory stick, secure digital (SD) memory card, on-processor cache, and the like.

[0092] The ED 110 may further include one or more input / output devices (not shown) or interfaces (such as a wired interface to the Internet 150 in FIG. 1) . The input / output devices permit interaction with a user or other devices in the network. Each input / output device includes any suitable structure for providing information to or receiving information from a user, such as through operation as a speaker, a microphone, a keypad, a keyboard, a display, or a touch screen, etc.

[0093] The ED 110 includes the processor 117 for performing operations including those operations related to preparing a transmission for uplink transmission to the NT-TRP 172 and / or the T-TRP 170, those operations related to processing downlink transmissions received from the NT-TRP 172 and / or the T-TRP 170, and those operations related to processing sidelink transmission to and from another ED 110. Processing operations related to preparing a transmission for uplink transmission may include operations such as encoding, modulating, transmit beamforming, and generating symbols for transmission. Processing operations related to processing downlink transmissions may include operations such as receive beamforming, demodulating and decoding received symbols. Depending upon the embodiment, a downlink transmission may be received by the receiver 113, possibly using receive beamforming, and the processor 117 may extract signaling from the downlink transmission (e.g. by detecting and / or decoding the signaling) . An example of signaling may be a reference signal transmitted by the NT-TRP 172 and / or by the T-TRP 170. In some embodiments, the processor 117 implements the transmit beamforming and / or the receive beamforming based on the indication of beam direction, e.g. beam angle information (BAI) , received from the T-TRP 170. In some embodiments, the processor 117 may perform operations relating to network access (e.g. initial access) and / or downlink synchronization, such as operations relating to detecting a synchronization sequence, decoding and obtaining the system information, etc. In some embodiments, the processor 117 may perform channel estimation, e.g. using a reference signal received from the NT-TRP 172 and / or from the T-TRP 170.

[0094] Although not illustrated, the processor 117 may form part of the transmitter 111 and / or part of the receiver 113. Although not illustrated, the memory 115 may form part of the processor 117.

[0095] The processor 117, the processing components of the transmitter 111 and the processing components of the receiver 113 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory (e.g. in the memory 115) . Alternatively, some or all of the processor 117, the processing components of the transmitter 111 and the processing components of the receiver 113 may each be implemented using dedicated circuitry, such as a programmed field-programmable gate array (FPGA) , a graphical processing unit (GPU) , a Central Processing Unit (CPU) or an application-specific integrated circuit (ASIC) .

[0096] The T-TRP 170 may be known by other names in some implementations, such as a base station, a base transceiver station (BTS) , a radio base station, a network node, a network device, a device on the network side, a transmit / receive node, a Node B, an evolved NodeB (eNodeB or eNB) , a Home eNodeB, a next Generation NodeB (gNB) , a transmission point (TP) , a site controller, an access point (AP) , a wireless router, a relay station, a remote radio head, a terrestrial node, a terrestrial network device, a terrestrial base station, a base band unit (BBU) , a remote radio unit (RRU) , an active antenna unit (AAU) , a remote radio head (RRH) , a central unit (CU) , a distributed unit (DU) , a positioning node, among other possibilities. The T-TRP 170 may be a macro BS, a pico BS, a relay node, a donor node, or the like, or combinations thereof. The T-TRP 170 may refer to the forgoing devices or refer to apparatus (e.g. a communication module, a modem, or a chip) in the forgoing devices.

[0097] In some embodiments, the parts of the T-TRP 170 may be distributed. For example, some of the modules of the T-TRP 170 may be located remote from the equipment that houses the antennas 256 for the T-TRP 170, and may be  coupled to the equipment that houses the antennas 256 over a communication link (not shown) sometimes known as front haul, such as common public radio interface (CPRI) . Therefore, in some embodiments, the term T-TRP 170 may also refer to modules on the network side that perform processing operations, such as determining the location of the ED 110, resource allocation (scheduling) , message generation, and encoding / decoding, and that are not necessarily part of the equipment that houses the antennas 256 of the T-TRP 170. The modules may also be coupled to other T-TRPs. In some embodiments, the T-TRP 170 may actually be a plurality of T-TRPs that are operating together to serve the ED 110, e.g. through the use of coordinated multipoint transmissions.

[0098] The T-TRP 170 includes at least one transmitter 181 and at least one receiver 183 coupled to one or more antennas 256. Only one antenna 256 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas 256 may alternatively be panels. The transmitter 181 and the receiver 183 may be integrated as a transceiver. The T-TRP 170 further includes a processor 182 for performing operations including those related to: preparing a transmission for downlink transmission to the ED 110, processing an uplink transmission received from the ED 110, preparing a transmission for backhaul transmission to the NT-TRP 172, and processing a transmission received over backhaul from the NT-TRP 172. Processing operations related to preparing a transmission for downlink or backhaul transmission may include operations such as encoding, modulating, precoding (e.g. multiple input multiple output (MIMO) precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the uplink or over backhaul may include operations such as receive beamforming, demodulating received symbols and decoding received symbols. The processor 182 may also perform operations relating to network access (e.g. initial access) and / or downlink synchronization, such as generating the content of synchronization signal blocks (SSBs) , generating the system information, etc. In some embodiments, the processor 182 also generates an indication of beam direction, e.g. BAI, which may be scheduled for transmission by a scheduler 184. The processor 182 performs other network-side processing operations described herein, such as determining the location of the ED 110, determining where to deploy the NT-TRP 172, etc. In some embodiments, the processor 182 may generate signaling, e.g. to configure one or more parameters of the ED 110 and / or one or more parameters of the NT-TRP 172. Any signaling generated by the processor 182 is sent by the transmitter 181. Note that “signaling” , as used herein, may alternatively be called control signaling. Signaling may be transmitted in a physical layer control channel, e.g. a physical downlink control channel (PDCCH) , in which case the signaling may be known as dynamic signaling. Signaling transmitted in a downlink physical layer control channel may be known as Downlink Control Information (DCI) . Siganling transmitted in an uplink physical layer control channel may be known as Uplink Control Information (UCI) . Signaling transmitted in a sidelink physical layer control channel may be known as Sidelink Control Information (SCI) . Signaling may be included in a higher-layer (e.g., higher than physical layer) packet transmitted in a physical layer data channel, e.g. in a physical downlink shared channel (PDSCH) , in which case the signaling may be known as higher-layer signaling, static signaling, or semi-static signaling. Higher-layer signaling may also refer to Radio Resource Control (RRC) protocol signaling or Media Access Control –Control Element (MAC-CE) signaling.

[0099] The scheduler 184 may be coupled to the processor 182. The scheduler 184 may be included within or operated separately from the T-TRP 170. The scheduler 184 may schedule uplink, downlink, and / or backhaul transmissions, including issuing scheduling grants and / or configuring scheduling-free ( “configured grant” ) resources. The T-TRP 170 further includes a memory 185 for storing information and data. The memory 185 stores instructions and data used, generated, or collected by the T-TRP 170. For example, the memory 185 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by the processor 182.

[0100] Although not illustrated, the processor 182 may form part of the transmitter 181 and / or part of the receiver 183. Also, although not illustrated, the processor 182 may implement the scheduler 184. Although not illustrated, the memory 185 may form part of the processor 182.

[0101] The processor 182, the scheduler 184, the processing components of the transmitter 181 and the processing components of the receiver 183 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory, e.g. in the memory 185. Alternatively, some or all of the processor 182, the scheduler 184, the processing components of the transmitter 181 and the processing components of the receiver 183 may be implemented using dedicated circuitry, such as a FPGA, a GPU, a CPU, or an ASIC.

[0102] Although the NT-TRP 172 is illustrated as a drone only as an example, the NT-TRP 172 may be implemented in any suitable non-terrestrial form, such as high altitude platforms, satellite, high altitude platform as international mobile telecommunication base stations and unmanned aerial vehicles, which forms will be discussed hereinafter. Also, the NT-TRP 172 may be known by other names in some implementations, such as a non-terrestrial node, a non-terrestrial network device, or a non-terrestrial base station. The NT-TRP 172 includes a transmitter 186 and a receiver 187 coupled to one or more antennas 108. Only one antenna 108 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas may alternatively be panels. The transmitter 186 and the receiver 187 may be integrated as a transceiver. The NT-TRP 172 further includes a processor 188 for performing operations including those related to: preparing a transmission for downlink transmission to the ED 110, processing an uplink transmission received from the ED 110, preparing a transmission for backhaul transmission to T-TRP 170, and processing a transmission received over backhaul from the T-TRP 170. Processing operations related to preparing a transmission for downlink or backhaul transmission may include operations such as encoding, modulating, precoding (e.g. MIMO precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the uplink or over backhaul may include operations such as receive beamforming, demodulating received symbols and decoding received symbols. In some embodiments, the processor 188 implements the transmit beamforming and / or receive beamforming based on beam direction information (e.g. BAI) received from the T-TRP 170. In some embodiments, the processor 188 may generate signaling, e.g. to configure one or more parameters of the ED 110. In some embodiments, the NT-TRP 172 implements physical layer processing, but does not implement higher layer functions such as functions at the medium access control (MAC) or radio link control (RLC) layer. As this is only an example, more generally, the NT-TRP 172 may implement higher layer functions in addition to physical layer processing.

[0103] The NT-TRP 172 further includes a memory 189 for storing information and data. Although not illustrated, the processor 188 may form part of the transmitter 186 and / or part of the receiver 187. Although not illustrated, the memory 189 may form part of the processor 188.

[0104] The processor 188, the processing components of the transmitter 186 and the processing components of the receiver 187 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory, e.g. in the memory 189. Alternatively, some or all of the processor 188, the processing components of the transmitter 186 and the processing components of the receiver 187 may be implemented using dedicated circuitry, such as a programmed FPGA, a GPU, a CPU, or an ASIC. In some embodiments, the NT-TRP 172 may actually be a plurality of NT-TRPs that are operating together to serve the ED 110, e.g. through coordinated multipoint transmissions.

[0105] The T-TRP 170, the NT-TRP 172, and / or the ED 110 may include other components, but these have been omitted for the sake of clarity.

[0106] FIG. 1D illustrates an example of units or modules in a device related to some embodiments of the present disclosure. One or more steps of the embodiment methods provided herein may be performed by corresponding units or modules, according to FIG. 1D. FIG. 1D illustrates units or modules in a device, such as in the ED 110, in the T-TRP 170, or in the NT-TRP 172. For example, a signal may be transmitted by a transmitting unit or by a transmitting module. A signal may be received or input by a receiving unit or by a receiving module. A signal may be processed by a processing unit or a processing module. Other steps may be performed by an artificial intelligence (AI) or machine learning (ML)  module. The respective units or modules may be implemented using hardware, one or more components or devices that execute software, or a combination thereof. For instance, one or more of the units or modules may be an integrated circuit, such as a programmed FPGA, a GPU, a CPU, or an ASIC. It will be appreciated that where the modules are implemented using software for execution by a processor for example, the modules may be retrieved by a processor, in whole or part as needed, individually or together for processing, in single or multiple instances, and that the modules themselves may include instructions for further deployment and instantiation.

[0107] While not shown, the transmitting module and the receiving module may be part of, or combined into, a transceiver module. A transceiver module may also be known as an interface module, or simply an interface, for inputting and outputting operations.

[0108] Additional details regarding the EDs 110, the T-TRP 170, and the NT-TRP 172 are known to those of skill in the art. As such, these details are omitted here.

[0109] To support the use of AI in a wireless network, an appropriate AI framework is needed. However, traditional communication technologies, such as 5G-related technologies, only consider AI use cases to improve network performance. Traditional communication technologies do not support a network for providing an AI service to user equipment (UE) . Considering the expected massive number of devices which have data and computing capability in future networks, there exists a desire to facilitate the provisioning of AI services in these future networks. Some promising technologies that leverage these devices with data and computing capability include distributed training and / or inference technologies, for example.

[0110] Aspects of this disclosure relate to procedures to establish an AI service between a user device and the network. The procedures include at least some requests and / or configurations and / or confirmation exchanged over the air interface between one user device and one base station.

[0111] Applications based on Artificial intelligence (AI) , especially in the form of large neural networks, are widely adopted in future devices and networks. Considering the size, power, and weight constraint, the computation and storage capabilities at the device side are usually limited compared with that in the network. To proliferate the application of AI based applications on more diversified devices, it is becoming a trend to support split AI computation between the user device and the network. In this case, the network serves as a platform to serve the AI computation services requested from user devices.

[0112] In some aspects of this invention, the AI computation services mainly focus on AI inference services, which may comprise a series of AI inference tasks requested from the user devices. An AI inference task can be seen as an AI computation process along a given AI model, from input data to output results. In the case of distributed AI inference between the user device and the network, the input and some layers of the AI model (could be zero) are computed at the user device, and the rest of the layers in the AI model up to the output are computed at the network. The decision of the model split point and the corresponding AI computation assignment between the two nodes is part of the AI model management. For instance, if an AI model has N layers and the split point is after n layers, then the first node, say Node A, will compute through input to layer n, and the second node, say Node B, will complete the rest of the model, i.e. from layer n+1 to N, until the output of the inference results. It is a must step that how to establish the service and make the user device and network both understand which part to compute so that they can jointly complete an AI inference task before any AI service is provided from the network to the user device.

[0113] According to embodiments of the present disclosure, there is provided a solution for a split of a model. In an aspect, a first device transmits a service request message to a second device for performing a task based on a model. Based on receiving a service response message including model split information from the second device, the first device performs the task based on the model and the model split information, and the model split information indicates a split of  model computation for performing the task at least between the first device and the second device. Based on receiving a deny notification message from the second device, the first device suspends the performing of the task based on the model. Therefore, the split of the model may be configured by the second device. Therefore, the performance of the model is improved. Principles and implementations of embodiments of the present disclosure will be described in detail below with reference to FIGS. 2-13.

[0114] FIG. 2 illustrates an example signaling chart illustrating an example process according to some embodiments of the present disclosure. The process 200 may involve a first device 201 and a second device 202. The process 200 may further involve at least one of a third device 203 or a fourth device 204. The first device 201 in FIG. 2 may be an example of the communication electric device 110 or the network node 170 in FIG. 1A. The second device 202 in FIG. 2 may be an example of the communication electric device 110 or the network node 170 in FIG. 1A. For example, the first device 201 may be a terminal device. In addition, the second device 202 may be an access network device. The third device 203 may be a core network device or an access network device. The fourth device 204 may be an access network device. It would be appreciated that although the process flow 200 has been described in the communication system 100A of FIG. 1A, this process may be likewise applied to other communication scenarios.

[0115] In the process flow 200, the first device 201 transmits 210 a service request message 215 to the second device 202 for performing a task based on a model. Correspondingly, the second device 202 receives 220 the service request message 215 from the first device 201. The model may be an AI model, and the task may be an AI service. A common understanding of AI model selected for a given AI service may be aligned among involved AI nodes (e.g., the first device 201 and / or the second device 202) by AI model life cycle management (LCM) . Such AI model LCM may involve further alignment for supported parameters of the selected AI model, e.g. to model splitting related parameters representing different trade-off among computation / communication resources required by each AI Node.

[0116] For the determination of the model split information, both the first device 201 and the second device 202 are assumed to have obtained common knowledge of which model to use for the AI inference task. For the case that it is not assumed that the first device 201 and the second device 202 have obtained common knowledge of which model to use for the task of AI inference, a process to obtain the common knowledge of which model to use is needed. This process could be done before the tasks, e.g. when a terminal device registers for an AI service from the network. In an embodiment, it can be assumed that the model selection is done by a model manager, which is a logical function node that may lie in radio access networks (e.g. second device 202) or core networks.

[0117] In some embodiments, the first device 201 may transmit a registration request message for registering the task to the second device 202. In addition, the registration request message may comprise a service latency for performing the task, an accuracy requirement for the task, or a combination of above two items. After receiving the registration request message from the first device 201, the second device 202 may determine a target model as the model to be used for performing the task. Then the second device 202 may transmit a model notification message to the first device 201. The model notification message indicates the model to be used for performing the task. Correspondingly, the first device 201 may receive the model notification message from the second device 202.

[0118] Additionally, the registration request message may comprise a service latency for performing the task, an accuracy requirement for the task, or a combination of above two items.

[0119] The model manager may be inside the second device 202, and may also be outside the second device 202. For instance, the model manager may be inside the third device 203. For determining the target model as the model, in an example, the second device 202 may decide the target model as the model to be used for performing the task.

[0120] In another example, the second device 202 may forward 205 the registration request message 206 to the third device 203, and receive 223 an indication of the target model 222 to be used for performing the task from the third device  203. On other side of the communication, after receiving 207 the registration request message 206 from the second device 202, the third device 203 may determine 208 the model to be used for performing the task. Then the third device 203 may transmit 221 the indication of the target model 222 to the second device 202.

[0121] For the case of multiple devices participate in performing the task, the model notification message may comprise information indicating multiple devices which are to participate in performing the task, and the multiple devices include the second device 202. The multiple devices may further include the fourth device 204. For instance, the model manager is not inside the second device 202, and there is another device (e.g., the third device 203 and / or a fourth device 204) that will be participating in the task of AI inference service.

[0122] If multiple network devices (e.g., the third device 203 and / or a fourth device 204) are involved in the AI inference, and the model manager is inside the third device 203, the second device 202 may be aware of the fourth device 204. The second device 202 may receive 226 an indication that the fourth device 204 is to participate in performing the task 225 based on the model from the third device 203. Correspondingly, the third device 203 may transmit 224 the indication that the fourth device 204 is to participate in performing the task 225 based on the model to the second device 202

[0123] In addition, the second device 202 may transmit 231 an indication of the target model 232 is to be used for performing the task to the fourth device 204 which is to participate in performing the task. If multiple network devices (e.g., the third device and / or a fourth device) are involved in the AI inference, and the model manager is inside the second device 202, the second device 202 may notify the fourth device about the selection of the model for a given AI service from the first device 201 when sending the decision of which model to use to the first device 201. Correspondingly, the fourth device 204 may receive 233 the indication of the target model 232 from the second device 202.

[0124] In some embodiments, the second device 202 may transmit 228 an indication of an anchor device 229 for the task among the multiple devices to the first device 201. On the other side of the communication, the first device 201 may receive 230 an indication of an anchor device for the task among the multiple devices from the second device 202 or the third device 203. If multiple devices may be involved in the AI inference service (e.g. the second device 202 and the third device 203) , the model manager may also specify the anchor device for the task of the AI service. In an example, the model manager is inside the second device 202, the second device 202 may indicate the anchor device to the first device 201. In another example, the device that the first device 201 register service through (e.g. the second device 202) can be the anchor device. The first device 201 may request a task of AI inference through the second device 202, and the second device 202 may decide the model split point between the first device 201 and the second device 202, as well as a model split point among the second device 202 and the third device 203.

[0125] In an example, the first device 201 may be aware of the second device 202 assigning the inference computation among the second device 202 and the third device. The first device 201 may not aware of the third device 203.

[0126] In some embodiments, the first device 201 may further transmit at least one of computation capability or computation resource of the first device 201 to the second device 202 for performing the task based on the model. Correspondingly, the second device 202 may receive the at least one of computation capability or computation resource of the first device from the first device 201 for performing the task based on the model. For instance, the first device 201 is willing to report its computational complexity to the second device 202, and the second device 202 may decide a model split point as part of the AI model management. The first device 201 may report its maximal computation capability and available computation resource before the service request message 215 for the AI service, or together with the service request message 215.

[0127] In addition, the service request message 215 may comprise information indicating a service latency for performing the task, a priority of the task, or a combination of above two items. For example, the service request message for the AI service may be associated with a service latency requirement for the AI service, and optionally a service priority for the AI service.

[0128] In some embodiments, the first device 201 may not transmit computation capability or computation resource of the first device 201 to the second device 202. The first device 201 may determine computation time to complete a first part of the task by the first device 201 at a reference model split point. Based on a service latency for performing the task and the computation time, the first device 201 may determine a residual latency. For instance, the first device 201 is not willing to report its computational complexity to second device 202, instead, it tells the second device 202 a latency budget (i.e., the service latency) exclude its own computational time, i.e., the residual latency. The second device 202 decides whether the AI service could be provided within the requested latency budget as well as determines the model split point as part of the AI model management.

[0129] Alternatively or additionally, the first device 201 may estimate amounts of model computation at multiple candidate model split points. Based on the amounts of model computation, the first device 201 may select a candidate model split point from the multiple candidate model split points as the reference model split point. For example, the first device 201 may determine the computation time it may take to complete the AI model inference at a reference model split point. The reference model split point may be obtained by estimating the amount of computation at different split points and then the first device 201 may select one or more split points it can afford.

[0130] Additionally, the residual latency may include first transmission time for the first device to transmit to the second device, an intermediate result of performing a first part of the task by the first device, computation time at the second device for performing a second part of the task, and second transmission time for the second device to transmit a result of performing the task to the first device. The first device 201 may include the information of requested latency of the service from the network side when requesting the AI service, which is the overall service latency requirement minus the computation time at the first device 201. In other words, the residual latency may include communication time from first device 201 to second device 202 for the transmission of the intermediate results, the computation time at the second device 202, as well as the transmission time of the inference results from the second device 202 to the first device 201 (which may also include resource scheduling time) .

[0131] In the case that the first device 201 does not transmit the computation capability or the computation resource of the first device 201 to the second device 202, the service request message 215 may comprise information indicating the residual latency, the reference model split point, a priority of the task, or any combination of two or more of the above-mentioned items. The reference model split point may be included in the service request message 215 so that the second device 202 could know from which layer in the AI model to continue the inference, and optionally a service priority. In addition, the granularity and candidate values for process timing of determining the residual latency or selecting the reference model split point may be pre-defined.

[0132] Continuing with reference to FIG. 2, the second device 202 determines 234 whether the task is to be completed within a service latency in case of a split of model computation for performing the task at least between the first device 201 and the second device 202.

[0133] As one embodiment, in order to determine whether the task is to be completed within the service latency, the second device 202 may determine an estimated latency for performing the task based on the at least one of the computation capability or the computation resource, and compare the estimated latency and the service latency. For the case that the computation capability or computation resource of the first device 201 are transmitted to the second device 202, the second device 202 may estimate the latency to see if the task could be completed within the service latency. If the  estimated latency is greater than the residual latency, the task cannot be completed within the service latency. If the estimated latency is less than or equal to the residual latency, the task can be completed within the service latency.

[0134] Additionally, the estimated latency may include an estimation of computation time for the first device 201 to perform a first part of the task in case of the computation capability of the first device 201 and a model split point, and an estimation of computation time for the second device to perform a second part of the task in case of computation capability of the second device 202 and the model split point. The first part of the task may refer to the model inference part computed at the first device 201 and the first part of the task may refer to the model inference part computed at the second device 202. For example, the estimated latency may include the estimation of computation time needed from both the first device 201 and second device 202 based on their computation capabilities and the potential model management (e.g. the split point of the model) . In addition, the estimated latency may further include an estimation of communication time between the first device and the second device in case of a data volume associated with the model split point. For example, the estimated latency may include the estimation of communication time given the data volume associated with the model management decision (e.g. the split point) .

[0135] As another embodiment, in order to determine whether the task is to be completed within the service latency, the second device 202 may determine an estimated latency by adding estimations of the first transmission time, the computation time, and the second transmission time. Then the second device 202 may compare the estimated latency and the residual latency. For the case that the computation capability or computation resource of the first device 201 does not transmitted to the second device 202, the second device 202 may estimate the latency for performing the task and the communication between the first device 201 and the second device 202. If the estimated latency is greater than the residual latency, the task cannot be completed within the service latency. If the estimated latency is less than or equal to the residual latency, the task can be completed within the service latency.

[0136] Continuing with reference to FIG. 2, based on determining that the task is to be completed within the service latency, the second device 202 transmits 235 a service response message 236 including model split information to the first device 201, and the model split information indicates the split of the model computation. For example, if second device 202 estimates that the task could be completed within the latency requirement, or the task could be completed within the latency requirement with a good probability, then it will respond the service respond message 236 with a model split decision.

[0137] In some embodiments, the service response message 236 may further comprise resource assignment for transmission of an intermediate result of performing a first part of the task by the first device 201. For example, the communication resource may be assigned by the second device 202 for the transmission of the output of AI model inference part computed at the first device 201, which will serve as input for the AI model inference to be computed at the second device 202.

[0138] After transmitting the service response message 236 to the first device, the second device 202 may perform the task based on the model and the model split information. In addition, the second device 202 may receive an intermediate result of performing a first part of the task from the first device. Based on the intermediate result and the model, the second device 202 may perform a second part of the task indicated by the model split information. The second device 202 then may transmit a result of performing the task to the first device 201.

[0139] Alternatively or additionally, the first device 201 may transmit 241 a first confirmation message 242 for the model split information to the second device 202. Correspondingly, the second device 202 may receive 243 the first confirmation message 242 from the first device 201. The first device 201 may confirm the model split decision by sending a dedicated message. Then the task of AI inference may be performed jointly by the first device 201 and the second device 202.

[0140] In addition, the first device 201 may transmit 244 a second confirmation message 245 for the model to be used to the second device 202. Correspondingly, the second device 202 may receive 246 the second confirmation message 245 from the first device 201. After which, the first device 201 and the second device 202 may obtain the common knowledge of which model to use.

[0141] If the model to be used for performing the task is determined by the third device 203, the first device 201 may further transmit 247 a third confirmation message 248 for the model to be used to the third device 203. In a further example, the first device 201 may transmit the third confirmation message to the second device 202, and the second device 202 may forward the third confirmation message to the third device. Correspondingly, the third device 203 may receive 249 the third confirmation message 248.

[0142] If multiple network devices including the fourth device 204 are involved in the AI inference, the first device 201 may transmit 250 a fourth confirmation message 251 for the model to be used to the fourth device 204 among the multiple devices. Correspondingly, the fourth device 204 may receive 252 the fourth confirmation message 251. For instance, when sending a confirmation to the second device 202 (with the model manager) , the first device 201 may also send a confirmation / handshake information to the fourth device 204.

[0143] It is to be understood that the positions of the first confirmation message 242, the second confirmation message 245, the third confirmation message 248 and the fourth confirmation message 251 in FIG. 2 are not limited, e.g., the second confirmation message 245, the third confirmation message 248 and the fourth confirmation message 251 may be transmitted before the first confirmation message 242.

[0144] Continuing with reference to FIG. 2, based on receiving 240 the service response message 236 including model split information from the second device 202, the first device 201 performs 253 the task based on the model and the model split information, and the model split information indicates a split of model computation for performing the task at least between the first device and the second device.

[0145] In some embodiments, in order to perform the task based on the model and the model split information, the first device 201 may perform a first part of the task indicated by the model split information based on the model. The first device 201 may transmit an intermediate result of performing the first part of the task to the second device 202. After that, the first device 201 may receive a result of performing the task from the second device 202.

[0146] The task may be performed between the first device 201 and the second device 202, and may also be performed among more than two devices. In some embodiments, in order to perform the task based on the model and the model split information, the second device 202 may transmit an indication to at least one other device. The indication indicates that the at least one other device performs a third part of the task based on the model. After that, the second device 202 may receive an intermediate result of performing the third part of the task from the at least one other device. The model may be split into three parts or more parts and the at least one other device may participate the performing of the task. In an example, the second device 202 may delegate some of the AI inference computation to a group of other devices that are not seen by first device 201.

[0147] In addition, the second device 202 may determine an estimated latency for performing the task by including computation time for the at least one other device to perform the third part of the task and an information exchange latency between the second device 202 and the at least one other device. For the second device 202, it needs to consider the computation time from all participating devices and the potential information exchange latency between these devices when the second device 202 estimates the overall latency needed to complete a task.

[0148] Continuing with reference to FIG. 2, based on determining that the task is not to be completed within the service latency, the second device 202 transmits 260 a deny notification message 265 to the first device 201. Based on receiving 270 the deny notification message 265 from the second device 202, the first device 201 suspends 275 the  performing of the task based on the model. If the second device 202 estimates that the task could not be completed before the latency budget, e.g., the estimated latency is less than or equal to the service latency, it will deny the service request message with the deny notification message 265.

[0149] Alternatively or additionally, the deny notification message may comprise information indicating a backoff timer. For example, the second device 202 may send a recommended backoff timer to the first device 201 as a response. In addition, the backoff timer may be defined based on a priority of the service request message and a random number. The backoff timer can be defined with explicit rules, e.g. based on the priority of request and random number jointly. For example, the higher the priority, the shorter the timer would be.

[0150] If the second device 202 denies the service request, the first device 201 may request the task again. In some embodiments, the first device 201 may transmit a further service request message for performing the task to the second device 202 after the backoff timer expires. Correspondingly, the second device 202 may receive the further service request message from the first device for performing the task after the backoff timer expires.

[0151] In general, some embodiments of the process 200 relate to establishing the AI service between two or more devices (e.g. the user device and the network) . The split of a model depends on flexible conditions which may include the willingness and level of the devices to report their computational capabilities, and / or where the decision maker lies for AI model management, etc. In this way, the model may be flexibly split according to the actual situation on the user equipment side and the network side.

[0152] FIG. 3 illustrates an example procedure of proposed solution according to some embodiments of the present disclosure. The procedure 300 may involve a node A 301 and a node B 302. It is understood that the process 300 can be considered as a more specific example of the process 200 in FIG. 2. Thus, the node A 301 in FIG. 3 may be an example of the first device 201 in FIG. 2, the node B 302 in FIG. 3 may be an example of the second device 202 in FIG. 2.

[0153] As shown in FIG. 3, it is assumed that a common understanding of the selection of model for a given AI task is aligned between the node A 301 and the node B 302 at 311. At 313, the node A 301 may report its computation capability to the node B 302. The computation capability may be reported before the AI service request or together with the AI service request at 315.

[0154] At 315, the node A 301 may request an AI reference service with a service latency requirement. The request may further associate with the computation capability of the node A 301 and a service priority. At 317, the node B 302 estimates if the task could be finished within the given latency budget.

[0155] If the node B 302 estimates the task could be completed within the latency requirement, or the task could be completed within the latency requirement with a good probability, then it will respond the service request with a model split decision at 319. The node B 302 may further respond the service request with the communication resource assignment for the transmission of the output of AI model inference part computed at the node A 301. At 321, the node A 301 may transmit a confirmation of model management, e.g., a model split point, to the node A 301. Then the node A 301 and the node B 302 start the task according to the confirmed model split at 323.

[0156] If the node B 302 estimate that the task could not be completed before the latency budget, the node B 302 responds the service request with a service deny notification at 325. The node B 302 may further send a recommended backoff timer to the node A 301 as response.

[0157] FIG. 4 illustrates an example procedure of proposed solution according to some embodiments of the present disclosure. The procedure 400 may involve a node A 401 and a node B 402. It is understood that the process 400 can be considered as a more specific example of the process 200 in FIG. 2. Thus, the node A 401 in FIG. 4 may be an example of the first device 201 in FIG. 2, the node B 402 in FIG. 4 may be an example of the second device 202 in FIG. 2.

[0158] As shown in FIG. 4, it is assumed that a common understanding of the selection of model for a given AI task is aligned between the node A 401 and the node B 402 at 411. At 413, the node A 401 may request an AI reference service with a residual latency budget and a reference split point. The request may further associate with a service priority. At 415, the node B 402 estimates if the task could be finished within the given latency budget. The estimation can be done using the knowledge of the computation and / or communication resource available.

[0159] If the node B 402 estimates the task could be completed within the residual latency requirement, or the task could be completed within the residual latency requirement with a good probability, then it will respond the service request with a model split decision at 417. The node B 402 may further respond the service request with the communication resource assignment for the transmission of the output of AI model inference part computed at the node A 401, which will serve as input for the AI model inference to be computed at the node B 402. At 419, the node A 401 may transmit a confirmation of model management, e.g., a model split point, to the node A 401. Then the node A 401 and the node B 402 start the task according to the confirmed model split at 421.

[0160] If the node B 402 estimate that the task could not be completed before the latency budget, the node B 402 responds the service request with a service deny notification at 423. The node B 402 may further send a recommended backoff timer to the node A 401 as response. The node A 401 may send another service request after the recommended timer expires.

[0161] FIG. 5 illustrates an example procedure of proposed solution according to some embodiments of the present disclosure. The procedure 500 may involve a node A 501 and a node B 502. It is understood that the process 500 can be considered as a more specific example of the process 200 in FIG. 2. Thus, the node A 501 in FIG. 5 may be an example of the first device 201 in FIG. 2, the node B 502 in FIG. 5 may be an example of the second device 202 in FIG. 2.

[0162] As shown in FIG. 5, the model manager is inside the node B 502. At 511, the node A 501 registers an AI inference service through the node B 502. The node A 501 sends registration request and provide optionally the service latency and / or accuracy requirement to the node B 502.

[0163] At 513, the model manager in the node B 502 decides the AI model to be used for the given AI service, and then the node B 502 notifies the node A 501. At 515, the node A 501 may send a confirmation of the model to the node B 502. At 517, the node A 501 and the node B 502 obtain the common knowledge of which model to use.

[0164] FIG. 6 illustrates an example procedure of proposed solution according to some embodiments of the present disclosure. The procedure 600 may involve a node A 601, a node B 602, and a model manager 603. It is understood that the process 600 can be considered as a more specific example of the process 200 in FIG. 2. Thus, the node A 601 in FIG. 6 may be an example of the first device 201 in FIG. 2, the node B 602 in FIG. 6 may be an example of the second device 202 in FIG. 2.

[0165] As shown in FIG. 6, the model manager is outside the node B 602. The model manage may be in the core network, or in another node which the node A 601 was not originally connected to or not the anchor node for the node A 601 to apply AI service with.

[0166] At 611, the node A 601 sends registration request and provide optionally the service latency and / or accuracy requirement to the node B 602. At 613, the node B 602 forward the AI service registration request to the model manager 603. At 615, the model manager 603 decides the AI model to be used for the given AI service, and then notifies the node B 602.

[0167] At 617, the decision from the model manager will be send to the node A 601 through the forwarding of the node B 602. At 619, the node A 601 may send a confirmation to the node B 602, and the node B 602 may forward the confirmation to the model manager at 621. In addition, the node A 601 may send the confirmation directly to the model  manager. At 623, the node A 601 and the node B 602 obtain the common knowledge of which model to use for a given AI service.

[0168] For the procedure 600, the model manager in another Node, e.g. Node C, will make decisions and align parameters of selected AI model to be used by the node A 601 and the node B 602 by considering the maximal and / or available communication / computation resources of the node A 601 and the node B 602.

[0169] In the procedures 300, 400, 500 and 600, it is assumed that node B is the only collaborative Node to perform the AI services jointly with node A. In the case that there are other Node, say Node C, to be involved in the joint AI inference task, the above procedures may change. It is easy to extend from one other node, i.e. node C, to a group of nodes, say {C1, C2, …, Cn} .

[0170] FIG. 7 illustrates an example procedure of proposed solution according to some embodiments of the present disclosure. The procedure 700 may involve a node A 701, a node B 702 with a model manager and a node C 703. It is understood that the process 700 can be considered as a more specific example of the process 200 in FIG. 2. Thus, the node A 701 in FIG. 7 may be an example of the first device 201 in FIG. 2, the node B 702 in FIG. 7 may be an example of the second device 202 in FIG. 2.

[0171] As shown in FIG. 7, the model manager is inside the node B 702, and there is another node C 703 that will be participating in the AI inference service as a collaborative node. At 711, the node A 701 registers an AI inference service through the node B 702. The node A 701 sends registration request and provide optionally the service latency and / or accuracy requirement to the node B 702.

[0172] At 713, the model manager in the node B 702 decides the AI model to be used for the given AI service, and then the node B 702 notifies the node A 701. At 714, the node B 702 also notify the node C 703 about the selection of the model for an AI inference service for the node A 701.

[0173] At 715, the node A 701 may send a confirmation of the model to the node B 502. At 717, the node A 701 may also send a confirmation / handshake information to the node C 703. At 719, the node A 701, the node B 502 and the node C 703 obtain the common knowledge of which model to use.

[0174] FIG. 8 illustrates an example procedure of proposed solution according to some embodiments of the present disclosure. The procedure 800 may involve a node A 801, a node B 802, a model manager 803 and a node C 804. It is understood that the procedure 800 can be considered as a more specific example of the process 200 in FIG. 2. Thus, the node A 801 in FIG. 8 may be an example of the first device 201 in FIG. 2, the node B 802 in FIG. 8 may be an example of the second device 202 in FIG. 2.

[0175] As shown in FIG. 8, the model manager is outside the node B 802, and there is another node C 804 that will be participating in the AI inference service as a collaborative node. At 811, the node A 801 registers an AI inference service through the node B 802. The node A 801 sends registration request and provide optionally the service latency and / or accuracy requirement to the node B 802. At 813, the node B 802 forward the AI service registration request to the model manager 803. At 815, the model manager 803 decides the AI model to be used for the given AI service, and then notifies the node B 802 and may optionally notify another collaborative node, i.e., the node C 804. At 817, the node B 802 also notify the node C 803 as a collaborative node for an AI inference service for the node A 801.

[0176] At 821, the node A 801 may send a confirmation of the model to the node B 602, and the node B 802 may forward the confirmation to the model manager at 823. At 825, the node A 801 may also send a confirmation / handshake information to the node C 803. At 827, the node A 801, the node B 802 and the node C 803 obtain the common knowledge of which model to use.

[0177] If multiple network nodes may be involved in the AI inference service (e.g. the node B 802 and the node C 803) , the model manager may also specify an anchor node for the AI service. For instance, the node B 802 may be the anchor node, and the node A 801 may request AI inference task through the node B 802. The node B 802 may decide the model split point between the node A 801 and the node B 802, as well as among the node B 802 and the node C 804. It is not necessary for the node A 801 to be aware of the node C 804. The node A 801 may be aware of the node B 802 assigning the inference computation among the node B 802 and the node C 804.

[0178] In the procedure 800, the model manager 803 will make decisions and align parameters of selected AI model to be used by all involved collaborative the node A 801, the node B 802 and the node C 804 by considering the maximal and / or available communication / computation resources of the node A 801, the node B 802 and the node C 804.

[0179] FIG. 9 shows a flowchart of an example method 900 implemented at a first device in accordance with some embodiments of the present disclosure. For the purpose of discussion, the method 900 will be described from the perspective of the communication electric device 110 with reference to FIG. 1A. It is to be understood that the method 900 may include additional acts not shown and / or may omit some shown acts, and the scope of the present disclosure is not limited in this regard.

[0180] At block 910, the first device transmits, at a first device to a second device, a service request message for performing a task based on a model. At block 920, based on receiving, from the second device, a service response message including model split information, the first device performs the task based on the model and the model split information, wherein the model split information indicates a split of model computation for performing the task at least between the first device and the second device. At block 930, based on receiving a deny notification message from the second device, the first device suspends the performing of the task based on the model. The first device may be configured to or operable to support other implementations of method 900.

[0181] FIG. 10 shows a flowchart of an example method 1000 implemented at a second device in accordance with some embodiments of the present disclosure. For the purpose of discussion, the method 1000 will be described from the perspective of the network node 170120 with reference to FIG. 1A. It is to be understood that the method 1000 may include additional acts not shown and / or may omit some shown acts, and the scope of the present disclosure is not limited in this regard.

[0182] At block 1010, the second device receive, at a second device from a first device, a service request message for performing a task based on a model. At block 1020, the second device determines whether the task is to be completed within a service latency in case of a split of model computation for performing the task at least between the first device and the second device. At block 1030, based on determining that the task is to be completed within the service latency, the second device transmits, to the first device, a service response message including model split information, wherein the model split information indicates the split of the model computation. At block 1040, based on determining that the task is not to be completed within the service latency, the second device transmits a deny notification message to the first device. The second device may be configured to or operable to support other implementations of method 1000.

[0183] FIG. 11 is a block diagram of a device 1100 that may be used for implementing some embodiments of the present disclosure. In some embodiments, the device 1100 may be an element of communications network infrastructure, such as a base station (for example, a NodeB, an evolved Node B (eNodeB, or eNB) , a next generation NodeB (sometimes referred to as a gNodeB or gNB) , a home subscriber server (HSS) , a gateway (GW) such as a packet gateway (PGW) or a serving gateway (SGW) or various other nodes or functions within a core network (CN) or a Public Land Mobility Network (PLMN) . In other embodiments, the device 1100 may be a device that connects to the network infrastructure over a radio interface, such as a mobile phone, smart phone or other such device that may be classified as a User Equipment (UE) . In some embodiments, the device 1100 may be a Machine Type Communications (MTC) device (also referred to as a machine-to-machine (M2M) device) , or another such device that may be categorized as a UE despite not  providing a direct service to a user. In some embodiments, the device 1100 may be a road side unit (RSU) , a vehicle UE (V-UE) , pedestrian UE (P-UE) or an infrastructure UE (I-UE) . In some scenarios, the device 1100 may also be referred to as a mobile device, a term intended to reflect devices that connect to mobile network, regardless of whether the device itself is designed for, or capable of, mobility. Specific devices may utilize all of the components shown or only a subset of the components, and levels of integration may vary from device to device. Furthermore, the device 1100 may contain multiple instances of a component, such as multiple processors, memories, transmitters, receivers, etc.

[0184] The device 1100 typically includes a processor 1102, such as a Central Processing Unit (CPU) , and may further include specialized processors such as a Graphics Processing Unit (GPU) or other such processor, a memory 1104, a network interface 1106 and a bus 1108 to connect the components of the device 1100. The device 1100 may optionally also include components such as a mass storage device 1110, a video adapter 1112, and an I / O interface 1116 (shown in dashed lines) .

[0185] The memory 1104 may comprise any type of non-transitory system memory, readable by the processor 1102, such as static random access memory (SRAM) , dynamic random access memory (DRAM) , synchronous DRAM (SDRAM) , read-only memory (ROM) , or a combination thereof. In an embodiment, the memory 1104 may include more than one type of memory, such as ROM for use at boot-up, and DRAM for program and data storage for use while executing programs. The bus 1108 may be one or more of any type of several bus architectures including a memory bus or memory controller, a peripheral bus, or a video bus.

[0186] The device 1100 may also include one or more network interfaces 1106, which may include at least one of a wired network interface and a wireless network interface. As illustrated in FIG. 11, network interface 1106 may include a wired network interface to connect to a network 1122, and also may include a radio access network interface 1120 for connecting to other devices over a radio link. When the device 1100 is a network infrastructure element, the radio access network interface 1120 may be omitted for nodes or functions acting as elements of the PLMN other than those at the radio edge. When the device 1100 is infrastructure at the radio edge of a network, both wired and wireless network interfaces may be included. When the device 1100 is a wirelessly connected device, such as a User Equipment, radio access network interface 1120 may be present and it may be supplemented by other wireless interfaces such as WiFi network interfaces. The network interfaces 1106 allow the device 1100 to communicate with remote entities such as those connected to network 1122.

[0187] The mass storage 1110 may comprise any type of non-transitory storage device configured to store data, programs, and other information and to make the data, programs, and other information accessible via the bus 1108. The mass storage 1110 may comprise, for example, one or more of a solid state drive, hard disk drive, a magnetic disk drive, or an optical disk drive. In some embodiments, the mass storage 1110 may be remote to the device 1100 and accessible through use of a network interface such as interface 1106. In the illustrated embodiment, the mass storage 1110 is distinct from memory 1104 where it is included, and may generally perform storage tasks compatible with higher latency, but may generally provide lesser or no volatility. In some embodiments, the mass storage 1110 may be integrated with a heterogeneous memory 1104.

[0188] The optional video adapter 1112 and the I / O interface 1116 (shown in dashed lines) provide interfaces to couple the device 1100 to external input and output devices. Examples of input and output devices include a display 1114 coupled to the video adapter 1112 and an I / O device 1118 such as a touch-screen coupled to the I / O interface 1116. Other devices may be coupled to the device 1100, and additional or fewer interfaces may be utilized. For example, a serial interface such as Universal Serial Bus (USB) (not shown) may be used to provide an interface for an external device. Those skilled in the art will appreciate that in embodiments in which the device 1100 is part of a data center, I / O interface 1116 and Video Adapter 1112 may be virtualized and provided through network interface 1106.

[0189] FIG. 12 is a schematic diagram of a structure of an apparatus 1200 in accordance with some embodiments of the present disclosure. As shown in FIG. 12, the apparatus 1200 includes a transmitting unit 1202, a performing unit 1204, and a receiving unit 1204. The apparatus 1200 may be applied to the communication system as shown in FIG. 1A, and may implement any of the methods provided in the foregoing embodiments. Optionally, a physical representation form of the apparatus 1200 may be a communication device, for example, a network device or UE. Alternatively, the apparatus 1200 may be another apparatus that can implement a function of a communication device, for example, a processor or a chip inside the communication device. Specifically, the apparatus 1200 may be some programmable chips such as a field-programmable gate array (FPGA) , a complex programmable logic device (CPLD) , an application-specific integrated circuit (ASIC) , or a system on a chip (SOC) .

[0190] In some embodiments, the transmitting unit 1202 may be configured to transmit, to a second device, a service request message for performing a task based on a model. The performing unit 1204 may be configured to perform the task based on the model and the model split information based on receiving, from the second device, a service response message including model split information, wherein the model split information indicates a split of model computation for performing the task at least between the first device and the second device. The suspending unit 1206 may be configured to suspend the performing of the task based on the model based on receiving a deny notification message from the second device.

[0191] In some other embodiments, the apparatus 1200 can include various other units or modules which may be configured to perform various operations or functions as described in connection with the foregoing method embodiments. The details can be obtained referring to the detailed description of the foregoing method embodiments and are not described herein again.

[0192] FIG. 13 is a schematic diagram of a structure of an apparatus 1300 in accordance with some embodiments of the present disclosure. As shown in FIG. 13, the apparatus 1300 includes a receiving unit 1302, a determining unit 1304, a first transmitting unit 1306 and a second transmitting unit 1308. The apparatus 1300 may be applied to the communication system as shown in FIG. 1A, and may implement any of the methods provided in the foregoing embodiments. Optionally, a physical representation form of the apparatus 1300 may be a communication device, for example, a network device or UE. Alternatively, the apparatus 1300 may be another apparatus that can implement a function of a communication device, for example, a processor or a chip inside the communication device. Specifically, the apparatus 1300 may be some programmable chips such as a field-programmable gate array (FPGA) , a complex programmable logic device (CPLD) , an application-specific integrated circuit (ASIC) , or a system on a chip (SOC) .

[0193] In some embodiments, the receiving unit 1302 may be configured to receive, from a first device, a service request message for performing a task based on a model. The determining unit 1304 may be configured to determine whether the task is to be completed within a latency in case of a split of model computation for performing the task at least between the first device and the second device. The first transmitting unit 1306 may be configured to transmit, to the first device, a service response message including model split information based on determining that the task is to be completed within the latency, wherein the model split information indicates the split of the model computation. The second transmitting unit 1308 may be configured to transmit a deny notification message to the first device based on determining that the task is not to be completed within the latency.

[0194] In some other embodiments, the apparatus 1300 can include various other units or modules which may be configured to perform various operations or functions as described in connection with the foregoing method embodiments. The details can be obtained referring to the detailed description of the foregoing method embodiments and are not described herein again.

[0195] It should be noted that division into the units or modules in the foregoing embodiments of the present disclosure is an example, and is merely logical function division. In actual implementation, there may be another  division manner. In addition, function units in embodiments of the present disclosure may be integrated into one processing unit, or may exist alone physically, or two or more units may be integrated into one unit. The integrated unit may be implemented in a form of hardware, or may be implemented in a form of a software function unit.

[0196] When the integrated unit is implemented in a form of a software function unit and sold or used as an independent product, the integrated unit may be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the present disclosure essentially, or all or some of the technical solutions may be implemented in a form of a software product. The computer software product is stored in a storage medium and includes several instructions for instructing a computer device (which may be a personal computer, a server, or a network device) or a processor to perform all or some of the steps of the methods described in embodiments of the present disclosure. The foregoing storage medium includes any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (ROM) , a random access memory (RAM) , a magnetic disk, or an optical disc.

[0197] Based on the foregoing embodiments, an embodiment of this application further provides a computer program. When the computer program is run on a computer, the computer is enabled to perform any of the methods provided in the foregoing embodiments.

[0198] Based on the foregoing embodiments, an embodiment of this application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a computer, the computer is enabled to perform the any of the methods provided in the foregoing embodiments. The storage medium may be any usable medium that can be accessed by a computer. By way of example and not limitation, the computer-readable medium may include a RAM, a ROM, an Electrically Erasable Programmable Read-Only Memory (EEPROM) , a Compact Disc Read-Only Memory (CD-ROM) or another optical disk storage, a magnetic disk storage medium or another magnetic storage device, or any other medium that can be used to carry or store expected program code in a form of an instruction or a data structure and that can be accessed by a computer.

[0199] Based on the foregoing embodiments, an embodiment of the present disclosure further provides a chip. The chip is configured to read a computer program stored in a memory, to implement any of the methods provided in the foregoing embodiments.

[0200] Based on the foregoing embodiments, an embodiment of the present disclosure provides a chip system. The chip system includes a processor, configured to support a computer apparatus in implementing functions related to communication devices in the foregoing embodiments. In a possible design, the chip system further includes a memory, and the memory is configured to store a program and data that are necessary for the computer apparatus. The chip system may include a chip, or may include a chip and another discrete component.

[0201] A person skilled in the art should understand that embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may be in a form of a hardware-only embodiment, a software-only embodiment, or an embodiment combining software and hardware aspects. In addition, the present disclosure may be in a form of a computer program product implemented on one or more computer-usable storage media (including but not limited to a magnetic disk memory, a CD-ROM, an optical memory, and the like) including computer-usable program code.

[0202] The present disclosure is described with reference to the flowcharts and / or block diagrams of the method, the device (system) , and the computer program product according to the present disclosure. It should be understood that computer program instructions may be used to implement each process and / or each block in the flowcharts and / or the block diagrams and a combination of a process and / or a block in the flowcharts and / or the block diagrams. These computer program instructions may be provided for a general-purpose computer, a dedicated computer, an embedded  processor, or a processor of another programmable data processing device to generate a machine, so that the instructions executed by a computer or a processor of another programmable data processing device generate an apparatus for implementing a specific function in one or more processes in the flowcharts and / or in one or more blocks in the block diagrams.

[0203] These computer program instructions may alternatively be stored in a computer-readable memory that can indicate a computer or another programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate an artifact that includes an instruction apparatus. The instruction apparatus implements a specific function in one or more processes in the flowcharts and / or in one or more blocks in the block diagrams.

[0204] These computer program instructions may alternatively be loaded onto a computer or another programmable data processing device, so that a series of operations and steps are performed on the computer or the another programmable device, to generate computer-implemented processing. Therefore, the instructions executed on the computer or the another programmable device provide steps for implementing a specific function in one or more processes in the flowcharts and / or in one or more blocks in the block diagrams.

[0205] It is clear that a person skilled in the art may make various modifications and variations to the present disclosure without departing from the protection scope of the present disclosure. Thus, the present disclosure is intended to cover these modifications and variations, provided that they fall within the scope of the claims of the present disclosure and their equivalent technologies.

Claims

1.A method comprising:transmitting, at a first device to a second device, a service request message for performing a task based on a model;based on receiving, from the second device, a service response message including model split information, performing the task based on the model and the model split information, wherein the model split information indicates a split of model computation for performing the task at least between the first device and the second device; andbased on receiving a deny notification message from the second device, suspending the performing of the task based on the model.2.The method of claim 1, further comprising:transmitting, to the second device, at least one of computation capability or computation resource of the first device for performing the task based on the model.3.The method of claim 2, wherein the service request message comprises information indicating at least one of the following:a service latency for performing the task, ora priority of the task.4.The method of claim 1, further comprising:determining computation time to complete a first part of the task by the first device at a reference model split point; anddetermining a residual latency based on a service latency for performing the task and the computation time.5.The method of claim 4, further comprising:estimating amounts of model computation at multiple candidate model split points; andselecting, based on the amounts of model computation, a candidate model split point from the multiple candidate model split points as the reference model split point.6.The method of claim 4 or 5, wherein the residual latency comprises:first transmission time for the first device to transmit, to the second device, an intermediate result of performing a first part of the task by the first device,computation time at the second device for performing a second part of the task, andsecond transmission time for the second device to transmit a result of performing the task to the first device.7.The method of any of claims 4-6, wherein the service request message comprises information indicating at least one of the following:the residual latency,the reference model split point, ora priority of the task.8.The method of any of claims 1-7, wherein the deny notification message comprises information indicating a backoff timer, and the method further comprises:transmitting, to the second device, a further service request message for performing the task after the backoff timer expires.9.The method of any of claims 1-8, further comprising:transmitting, to the second device, a registration request message for registering the task; andreceiving, from the second device, a model notification message indicating the model to be used for performing the task.10.The method of claim 9, wherein the registration request message comprises at least one of the following:a service latency for performing the task, oran accuracy requirement for the task.11.A method comprising:receiving, at a second device from a first device, a service request message for performing a task based on a model;determining whether the task is to be completed within a service latency in case of a split of model computation for performing the task at least between the first device and the second device;based on determining that the task is to be completed within the service latency, transmitting, to the first device, a service response message including model split information, wherein the model split information indicates the split of the model computation; andbased on determining that the task is not to be completed within the service latency, transmitting a deny notification message to the first device.12.The method of claim 11, further comprising:receiving, from the first device, at least one of computation capability or computation resource of the first device for performing the task based on the model.13.The method of claim 12, wherein the service request message comprises information indicating at least one of the following:the service latency for performing the task, ora priority of the task.14.The method of claim 12 or 13, wherein determining whether the task is to be completed within the service latency comprises:determining an estimated latency for performing the task based on the at least one of the computation capability or the computation resource; andcomparing the estimated latency and the service latency.15.The method of claim 14, wherein the estimated latency comprises:an estimation of computation time for the first device to perform a first part of the task in case of the computation capability of the first device and a model split point,an estimation of computation time for the second device to perform a second part of the task in case of computation capability of the second device and the model split point, andan estimation of communication time between the first device and the second device in case of a data volume associated with the model split point.16.The method of claim 11, wherein the service request message comprises information indicating at least one of the following:a residual latency determined based on the service latency for performing the task and computation time to complete a first part of the task by the first device at a reference model split point,the reference model split point, ora priority of the task.17.The method of claim 16, wherein the residual latency comprises:first transmission time for the first device to transmit, to the second device, an intermediate result of performing a first part of the task by the first device,computation time at the second device for performing a second part of the task, andsecond transmission time for the second device to transmit a result of performing the task to the first device.18.The method of claim 17, wherein determining whether the task is to be completed within the service latency comprises:determining an estimated latency by adding estimations of the first transmission time, the computation time, and the second transmission time; andcomparing the estimated latency and the residual latency.19.The method of any of claims 11-18, wherein the deny notification message comprises information indicating a backoff timer, and the method further comprises:receiving, from the first device, a further service request message for performing the task after the backoff timer expires.20.The method of any of claims 11-19, further comprising:receiving, from the first device, a registration request message for registering the task;determining a target model as the model to be used for performing the task; andtransmitting, to the first device, a model notification message indicating the model to be used for performing the task.21.The method of claim 20, wherein determining the target model as the model comprises:forwarding the registration request message to a third device; andreceiving, from the third device, an indication that the target model is the model to be used for performing the task.22.A first device comprising:a transceiver; anda processor communicatively coupled with the transceiver,wherein the processor is configured to:transmit, via the transceiver, at a first device to a second device, a service request message for performing a task based on a model;based on receiving, from the second device, a service response message including model split information, perform the task based on the model and the model split information, wherein the model split information indicates a split of model computation for performing the task at least between the first device and the second device; andbased on receiving a deny notification message from the second device, suspend the performing of the task based on the model.23.A second device comprising:a transceiver; anda processor communicatively coupled with the transceiver,wherein the processor is configured to:receive, via the transceiver and at a second device from a first device, a service request message for performing a task based on a model;determine whether the task is to be completed within a latency in case of a split of model computation for performing the task at least between the first device and the second device;based on determining that the task is to be completed within the latency, transmit, via the transceiver to the first device, a service response message including model split information, wherein the model split information indicates the split of the model computation; andbased on determining that the task is not to be completed within the latency, transmit, via the transceiver, a deny notification message to the first device.24.A non-transitory computer readable medium storing instructions, which when executed by at least one processor, cause the at least one processor to perform the method of any of claims 1-21.25.An apparatus comprising at least one processing circuit configured to perform the method of any of claims 1-21.26.An apparatus for wireless communication, the apparatus comprising:at least one processor; anda non-transitory computer readable medium storing instructions which, when executed by the at least one processor, cause the apparatus to perform the method of any of claims 1-21.27.A computer program comprising computer-executable instructions which, when executed, cause an apparatus to perform the method of any of claims 1-21.

Citation Information

Patent Citations

  • A neural network model training method and device by using a trusted execution environment

    CN111260053A

  • Exception handling for collaborating process models

    US20080127205A1

  • Method and system for neural network execution distribution

    US20220083386A1

  • Communication method and apparatus and electronic device

    WO2022233294A1