Method, apparatus and system for artificial intelligence (AI) model splitting

By splitting model computation tasks among devices and utilizing service request and response messages for model splitting, the problem of inflexible allocation of model computation resources in existing communication technologies is solved, thereby improving model performance and resource utilization efficiency.

CN122055705APending Publication Date: 2026-05-15HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-07-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing communication technologies struggle to maximize the use of signal space, and wireless communication suffers from inflexible allocation of model computation resources.

Method used

By splitting model computation tasks among devices, using service request and response messages to split the model, and taking into account computing power and latency, model computation resources can be flexibly allocated.

Benefits of technology

It enables flexible allocation of model computation, improving model performance and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122055705A_ABST
    Figure CN122055705A_ABST
Patent Text Reader

Abstract

Exemplary embodiments relate to perceptual measurement and reporting. In an aspect, a first device (201) sends a service request message for performing a task based on a model to a second device (202). Based on receiving a service response message from the second device (202) comprising model split information, the first device (201) performs the task based on the model and the model split information, and the model split information indicates a model for performing the task to compute a split at least between the first device (201) and the second device (202). Based on receiving a rejection notification message from the second device (202), the first device (201) suspends execution of the task based on the model.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 594,107, filed October 30, 2023, entitled “Service request and handshake procedures for artificial intelligence (AI) and machine learning (ML)”, the disclosure of which is incorporated herein by reference. Technical Field

[0002] Exemplary embodiments of the present invention relate generally to the field of communications, and more particularly to methods, apparatuses, devices, and computer-readable storage media for model decomposition. Background Technology

[0003] Artificial intelligence (AI), particularly deep machine learning, is a broad branch of computer science dedicated to building intelligent machines capable of performing tasks that typically require human intelligence. The introduction of AI is expected to bring about a paradigm shift in almost all areas of the technology industry, and it is poised to play a role in driving the development of network and communication technologies. For example, existing communication technologies rely on traditional channel analysis modeling, enabling wireless communication to approach the theoretical Shannon limit. However, current technologies are insufficient to further maximize the efficiency of signal space utilization. AI is expected to help address this challenge. Other aspects of wireless communication can also benefit from the use of AI, particularly in future generations of wireless technologies, such as advanced 5G technologies and future systems, as well as technologies further into the future. Several issues remain to be addressed regarding the use of AI in wireless networks. Summary of the Invention

[0004] Overall, exemplary embodiments of the present invention provide a solution for model splitting.

[0005] In a first aspect, a method is provided. The method includes: sending a service request message for performing a model-based task from a first device to a second device; performing the task based on the model and the model splitting information received from the second device, the model splitting information indicating a split of model computation used for performing the task, at least between the first and second devices; and suspending the model-based task execution based on receiving a rejection notification message from the second device. Therefore, the model can be flexibly split. Consequently, model performance is improved.

[0006] In some embodiments, the first device may also send at least one of its computing power or computing resources for performing tasks based on the model to the second device. In this way, the breakdown of model computation can be determined based on the computing power or computing resources of the first device.

[0007] In some embodiments, the service request message may include information indicating the service latency for performing a task, the task priority, or a combination of both. This allows resource conflicts and latency to be considered.

[0008] In some embodiments, the first device may further determine the computation time for completing a first portion of the task at the reference model split point, and determine the remaining latency based on the service latency and computation time used to perform the task. In this way, the computing power of the first device may not be reported to the second device.

[0009] In some embodiments, the first device may further estimate the computational cost of the model at multiple candidate model split points, and select one candidate model split point from the multiple candidate model split points as a reference model split point based on the computational cost. In this way, the reference model split point can be provided to the second device.

[0010] In some embodiments, the remaining latency may include a first transmission time for the first device to send an intermediate result of a first part of a task performed by the first device to the second device, a computation time required to perform a second part of the task at the second device, and a second transmission time for the second device to send the result of the task to the first device. In this way, communication time and computation time can be taken into account in the latency budget.

[0011] In some embodiments, the service request message may include information indicating remaining latency, a reference model split point, task priority, or any combination of two or more of the foregoing. In this way, the reference model split point can be provided to the second device.

[0012] In some embodiments, the service response message may further include resource allocation for transmitting intermediate results of a first portion of a task performed by the first device. In this way, resources for transmitting intermediate results can be provided to the first device.

[0013] In some embodiments, performing a task based on a model and model splitting information may include: performing a first part of a task indicated by the model splitting information based on the model; sending intermediate results of performing the first part of the task to a second device; and receiving the results of performing the task from the second device. In this way, the task can be performed at both the first and second devices.

[0014] In some embodiments, the rejection notification message includes information instructing a rollback timer, and the first device may also send another service request message to the second device to perform the task after the rollback timer expires. In this way, the service request message can be resent within a predefined time period.

[0015] In some embodiments, the rollback timer can be defined based on the priority of the service request message and a random number. In this way, the period during which the service request message is sent can be associated with the priority of the service request message.

[0016] In some embodiments, the first device may also send a first confirmation message to the second device regarding the model splitting information. In this way, the first device can confirm the decision to split the model.

[0017] In some embodiments, the first device may also send a registration request message for registering a task to the second device and receive a model notification message from the second device indicating the model to be used to perform the task. In this way, the first device can be instructed on the model to be used.

[0018] In some embodiments, the registration request message may include service latency for performing a task, accuracy requirements for the task, or a combination of both. In this way, service latency and accuracy requirements can be indicated to the first device.

[0019] In some embodiments, the first device may also send a second confirmation message to the second device regarding the model to be used. In this way, the first device can confirm the model to be used.

[0020] In some embodiments, the model used to perform the task is determined by a third device, and the first device may also send a third confirmation message to the third device regarding the model to be used. In this way, the first device can confirm the model to be used with the third device.

[0021] In some embodiments, the model notification message may include information indicating a plurality of devices that will participate in performing the task, wherein the plurality of devices includes a second device.

[0022] In some embodiments, the first device may also send a fourth confirmation message of the model to be used to a fourth device among a plurality of devices. In this way, the first device can confirm the model to be used with the fourth device.

[0023] In some embodiments, the first device may also receive instructions from a second or third device regarding anchor devices used for a task from among a plurality of devices. In this manner, anchor devices can be instructed to the first device.

[0024] In some embodiments, the first device may be a terminal device, the second device may be an access network device, the third device may be a core network device or an access network device, or the fourth device may be an access network device. In this way, the terminal device can perform tasks with one or more network devices.

[0025] In a second aspect, a method is provided at a second device. The method includes: receiving a service request message from a first device for performing a model-based task at the second device; determining whether the task will be completed within a service delay, provided that the model computation used to perform the task is split at least between the first and second devices; based on the determination that the task will be completed within the service delay, sending a service response message including model splitting information to the first device, wherein the model splitting information indicates the splitting of the model computation; and based on the determination that the task will not be completed within the service delay, sending a rejection notification message to the first device. Therefore, the model can be flexibly split. Consequently, the performance of the model is improved.

[0026] In some embodiments, the second device may also receive from the first device at least one of the computing power and / or computing resources of the first device for performing tasks based on the model. In this way, the breakdown of model computation can be determined based on the computing power or computing resources of the first device.

[0027] In some embodiments, the service request message may include information indicating service latency for performing a task, task priority, or any combination of two or more of the foregoing. This allows resource conflicts and latency to be taken into account.

[0028] In some embodiments, determining whether a task will be completed within the service latency may include: determining an estimated latency for performing the task based on at least one of computing power or computing resources; and comparing the estimated latency with the service latency. In this way, the breakdown of model computation can be determined based on the computing power or computing resources of the first device.

[0029] In some embodiments, the estimated latency may include: an estimate of the computation time for the first device to perform a first part of a task, given the computing power of the first device and the model split point; an estimate of the computation time for the second device to perform a second part of a task, given the computing power of the second device and the model split point; and an estimate of the communication time between the first and second devices, given the amount of data associated with the model split point. In this way, both communication and computation times can be taken into account in the estimated latency.

[0030] In some embodiments, the service request message may include information indicating remaining latency, a reference model split point, task priority, or any combination of two or more of the foregoing, the remaining latency being determined based on the service latency for performing the task and the computation time for the first device to complete a first portion of the task at the reference model split point. In this way, the reference model split point can be provided to the second device.

[0031] In some embodiments, the remaining latency may include a first transmission time for the first device to send an intermediate result of a first part of a task performed by the first device to the second device, a computation time for the second part of the task performed at the second device, and a second transmission time for the second device to send the result of the task to the first device. In this way, communication time and computation time can be taken into account in the latency budget.

[0032] In some embodiments, determining whether a task will be completed within the delay may include: determining an estimated delay by adding estimates of a first transmission time, a computation time, and a second transmission time; and comparing the estimated delay with the remaining delay. In this way, it can be determined whether the task will be completed based on the estimated delay and the remaining delay.

[0033] In some embodiments, the service response message may further include resource allocation for transmitting intermediate results of a first portion of the task performed by the first device. In this way, resources for transmitting intermediate results can be provided to the first device.

[0034] In some embodiments, based on sending a service response message to the first device, the second device can also execute tasks based on the model and model splitting information. In this way, the task can be executed at the second device.

[0035] In some embodiments, performing a task based on a model and model splitting information may include: receiving an intermediate result of a first part of the task execution from a first device; performing a second part of the task indicated by the model splitting information based on the intermediate result and the model; and sending the result of the task execution to the first device. In this way, the task can be performed at both the first and second devices.

[0036] In some embodiments, performing a task based on a model and model splitting information may include: sending an instruction to at least one other device indicating that the at least one other device performs a third part of the task based on the model; and receiving an intermediate result from the at least one other device of the third part of the task. In this way, the task can be performed at both a first device and a second device.

[0037] In some embodiments, the second device may also determine the estimated latency for performing the task by including: the computation time of at least one other device performing a third portion of the task; and the information exchange latency between the second device and the at least one other device. In this way, the communication time with the at least one other device can be taken into account in the estimated latency.

[0038] In some embodiments, the rejection notification message includes information instructing a rollback timer, and the second device may also receive an additional service request message from the first device to perform a task after the rollback timer expires. In this way, the service request message can be retransmitted within a predefined time period.

[0039] In some embodiments, the rollback timer can be defined based on the priority of the service request message and a random number. In this way, the period during which the service request message is sent can be associated with the priority of the service request message.

[0040] In some embodiments, the second device may also receive a first confirmation message from the first device regarding the model splitting information. In this way, the first device can confirm the decision to split the model.

[0041] In some embodiments, the second device may also receive a registration request message from the first device for registering a task, determine the target model as the model to be used to perform the task, and send a model notification message to the first device indicating the model to be used to perform the task. In this way, the first device can be instructed on the model to be used.

[0042] In some embodiments, determining the target model may include having a second device determine the target model as the model to be used to perform the task. In this way, the model to be used can be determined by the second device.

[0043] In some embodiments, determining the target model may include: forwarding a registration request message to a third device; and receiving from the third device an indication that the target model is the model to be used to perform the task. In this way, the model to be used can be determined by the third device.

[0044] In some embodiments, the second device may also receive from the third device an instruction indicating that a fourth device will participate in the model-based task execution. In this way, the second device can be notified that a fourth device is participating in the task.

[0045] In some embodiments, the registration request message may include service latency for performing a task, accuracy requirements for the task, or a combination of both. In this way, service latency and accuracy requirements can be indicated to the first device.

[0046] In some embodiments, the second device may also receive a second confirmation message from the first device regarding the model to be used. In this way, the first device can confirm the model to be used.

[0047] In some embodiments, the second device may also send an instruction to a fourth device that will participate in performing the task, indicating that the model will be used to perform the task. In this way, the fourth device can be instructed on the model to be used.

[0048] In some embodiments, the model notification message may include information indicating multiple devices to participate in the task, wherein the multiple devices include a second device and a fourth device. In this way, the first device can be notified that multiple devices are participating in the task.

[0049] In some embodiments, the second device may also send instructions to the first device regarding anchor devices among a plurality of devices used for the task. In this manner, anchor devices can be instructed to the first device.

[0050] In some embodiments, the first device may be a terminal device, the second device may be an access network device, the third device may be a core network device or an access network device, or the fourth device may be an access network device. In this way, the terminal device can perform tasks with one or more network devices.

[0051] In a third aspect, a first device is provided. The first device includes a transceiver and a processor communicatively coupled to the transceiver. The processor is configured to: send a service request message for a model-based task execution to a second device from the first device; execute the task based on the model and the model splitting information received from the second device, wherein the model splitting information indicates a split of model computation for performing the task, at least between the first device and the second device; and suspend the model-based task execution based on a rejection notification message received from the second device.

[0052] In a fourth aspect, a second device is provided. The second device includes a transceiver and a processor communicatively coupled to the transceiver. The processor is configured to: receive, at the second device, a service request message from a first device for performing a model-based task; determine, if the model computation for performing the task is split at least between the first and second devices, whether the task will be completed within a service delay; based on the determination that the task will be completed within the service delay, send a service response message to the first device including model splitting information, wherein the model splitting information indicates the splitting of the model computation; and based on the determination that the task will not be completed within the service delay, send a rejection notification message to the first device.

[0053] In a fifth aspect, a non-transient computer-readable medium is provided, including a computer program stored thereon, which, when executed on at least one processor, causes at least one processor to perform the method according to either the first or second aspect.

[0054] In a sixth aspect, a chip is provided, including at least one processing circuit for performing a method according to either the first or the second aspect.

[0055] In a seventh aspect, a computer program product is provided, which is tangibly stored on a computer-readable medium and includes computer-executable instructions that, when executed, cause a device to perform the method according to either the first or the second aspect.

[0056] It should be understood that the summary section is not intended to identify key or essential features of embodiments of the invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0057] Some exemplary embodiments will now be described with reference to the accompanying drawings, in which: Figure 1A An exemplary communication system in which exemplary embodiments of the present invention may be implemented is shown; Figure 1B An exemplary communication system in which exemplary embodiments of the present invention may be implemented is shown; Figure 1C Examples of electronic devices (EDs) and base stations related to some embodiments of the present invention are shown; Figure 1D Examples of units or modules in a device related to some embodiments of the present invention are shown; Figure 2 An exemplary signaling diagram of an exemplary process according to some embodiments of the present invention is shown; Figure 3 A first exemplary process for a proposed solution according to some embodiments of the present invention is shown; Figure 4 A second exemplary process is shown for the proposed solution according to some embodiments of the present invention; Figure 5 A third exemplary process is shown for the proposed solution according to some embodiments of the present invention; Figure 6 A fourth exemplary process is shown for the proposed solution according to some embodiments of the present invention; Figure 7A fifth exemplary process is shown for a proposed solution according to some embodiments of the present invention; Figure 8 A sixth exemplary process is shown for a proposed solution according to some embodiments of the present invention; Figure 9 A flowchart of a method implemented at a first device according to some embodiments of the present invention is shown; Figure 10 A flowchart of a method implemented at a second device according to some embodiments of the present invention is shown; Figure 11 This is a block diagram of a device that can be used to implement some embodiments of the present invention; Figure 12 This is a schematic diagram of the structure of a device according to some embodiments of the present invention; Figure 13 This is a schematic diagram of the structure of a device according to some embodiments of the present invention.

[0058] In the accompanying drawings, the same or similar reference numerals denote the same or similar elements. Detailed Implementation

[0059] The principles of the invention will now be described with reference to some exemplary embodiments. It should be understood that these embodiments are described merely to illustrate and assist those skilled in the art in understanding and implementing the invention, and do not impose any limitations on the scope of the invention. The invention described herein can be implemented in various ways other than those described below.

[0060] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0061] In this invention, references to "an embodiment," "embodiment," "exemplary embodiment," etc., indicate that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in connection with an embodiment, it is understood that, whether explicitly described or not, the influence of other embodiments on such feature, structure, or characteristic is within the knowledge of those skilled in the art.

[0062] It should be understood that although terms such as "first," "second," etc., may be used herein to describe various elements, these elements should not be limited by these terms. The terms used are merely used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. The term "and / or" as used herein includes any and all combinations of one or more of the listed terms.

[0063] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein include the plural meaning. It should also be understood that the terms “comprises,” “comprising,” “has,” “having,” “includes,” and / or “including” are used herein to indicate the presence of the stated features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0064] Figure 1A An exemplary communication system 100A, in which an exemplary embodiment of the present invention can be implemented, is shown. References Figure 1A As a non-limiting illustrative example, a simplified schematic diagram of a communication system is provided. Communication system 100A includes a radio access network 120. Radio access network 120 may be a next-generation radio access network or a traditional (e.g., 5G, 4G, 3G, or 2G) radio access network. One or more communication electronic devices (EDs) 110a, 110b, 110c, 110d, 110e, 110f, 110g, 110h, 110i, 110j (generally referred to as 110) may interconnect with each other or connect to one or more network nodes (170a, 170b, generally referred to as 170) in radio access network 120. Core network 130 may be part of the communication system and may depend on or be independent of the radio access technology used in communication system 100A. Furthermore, communication system 100A includes a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160.

[0065] Figure 1BAn exemplary communication system in which exemplary embodiments of the present invention can be implemented is illustrated. Generally, communication system 100B enables multiple wireless or wired components to transmit data and other content. The purpose of communication system 100B may be to provide content such as voice, data, video, signaling, and / or text via broadcast, multicast, and unicast. Communication system 100B can operate by sharing resources such as carrier spectrum bandwidth among its constituent units. Communication system 100B may include terrestrial communication systems and / or non-terrestrial communication systems. Communication system 100B can provide a wide range of communication services and applications (e.g., earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery, and mobility). Communication system 100B can provide high availability and robustness through the joint operation of terrestrial and non-terrestrial communication systems. For example, integrating non-terrestrial communication systems (or components thereof) into terrestrial communication systems can form a multi-layered heterogeneous network. Compared to traditional communication networks, heterogeneous networks can achieve better overall performance through efficient multi-link joint operation between terrestrial and non-terrestrial networks, more flexible function sharing, and faster physical layer link switching.

[0066] Terrestrial communication systems and non-terrestrial communication systems can be considered subsystems of a communication system. Figure 1B In the example shown, communication system 100B includes electronic devices (EDs) 110a, 110b, 110c, and 110d (generally referred to as ED 110), radio access networks (RANs) 120a and 120b, a non-terrestrial communication network 120c, a core network 130, a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160. RANs 120a and 120b include corresponding base stations (BSs) 170a and 170b, which are generally referred to as terrestrial transmit and receive points (T-TRPs) 170a and 170b. The non-terrestrial communication network 120c includes an access node 172, which may be referred to as a non-terrestrial transmit and receive point (NT-TRP) 172 or a sensing agent 172.

[0067] Alternatively or additionally, any ED 110 can be used to connect, access, or communicate with any T-TRP 170a and 170b, as well as NT-TRP 172, Internet 150, core network 130, PSTN 140, other networks 160, or any combination thereof. In some examples, ED 110a can perform uplink and / or downlink transmissions with T-TRP 170a via terrestrial air interface 190a. In some examples, ED 110a, 110b, 110c, and 110d can also communicate directly with each other via one or more sidelink air interfaces 190b. In some examples, ED 110d can perform uplink and / or downlink transmissions with NT-TRP 172 via non-terrestrial air interface 190c.

[0068] Air interfaces 190a and 190b can use similar communication technologies, such as any suitable wireless access technology. For example, communication system 100B can implement one or more channel access methods in air interfaces 190a and 190b, such as code division multiple access (CDMA), space division multiple access (SDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), direct Fourier transform spread OFDMA (DFT-OFDMA), or single-carrier FDMA (SC-FDMA). Air interfaces 190a and 190b can utilize other higher-dimensional signal spaces, which may involve combinations of orthogonal and / or non-orthogonal dimensions.

[0069] The non-terrestrial air interface 190c enables communication between the ED 110d and one or more NT-TRP 172s via a wireless link or simply via a link. In some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of ED 110s and one or more NT-TRP 172s for multicast transmission.

[0070] RANs 120a and 120b communicate with core network 130 to provide various services, such as voice, data, and other services, to EDs 110a, 110b, and 110c. RANs 120a and 120b and / or core network 130 may communicate directly or indirectly with one or more other RANs (not shown), which may or may not be directly served by core network 130, and may or may not use the same radio access technology as RANs 120a and / or RAN 120b. Core network 130 may also serve as a gateway access between (i) RANs 120a and 120b or EDs 110a, 110b, and 110c or both RANs and EDs and (ii) other networks (e.g., PSTN 140, Internet 150, Sensing Agent 172, and other networks 160). Additionally, some or all of ED 110a, 110b, and 110c may include the ability to communicate with different wireless networks via different wireless links using different wireless technologies and / or protocols. ED 110a, 110b, and 110c may communicate with a service provider or exchange (not shown) via a wired communication channel and with the Internet 150, rather than wirelessly (or also wirelessly). PSTN 140 may include a circuit-switched telephone network for providing plain old telephone service (POTS). The Internet 150 may include a network of computers and subnets (intranets) or both, incorporating protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), and User Datagram Protocol (UDP). ED 110a, 110b, and 110c may be multimode devices capable of operating according to multiple wireless access technologies, incorporating multiple transceivers required to support such operation.

[0071] Figure 1C Examples of electronic devices (EDs) and base stations related to some embodiments of the present invention are shown. Figure 1CAs shown, another example of ED 110 and base stations 170a, 170b, and / or 170c is provided. ED 110 is used to connect people, objects, machines, etc. ED 110 can be widely used in various scenarios, such as cellular communication, device-to-device (D2D), vehicle-to-everything (V2X), peer-to-peer (P2P), machine-to-machine (M2M), machine-type communications (MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), mixed reality (MR), metaverse, digital twins, industrial control, autonomous driving, telemedicine, smart grids, smart furniture, smart offices, smart wearables, smart transportation, smart cities, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery, and mobility, etc.

[0072] Each ED 110 represents any end-user equipment suitable for wireless operation and may include (or be referred to as): user equipment / device (UE), wireless transmit / receive unit (WTRU), mobile station, fixed or mobile subscriber unit, cellular phone, station (STA), machine type communication (MTC) device, personal digital assistant (PDA), smartphone, laptop, computer, tablet, wireless sensor, consumer electronics, smartbook, vehicle, automobile, truck, bus, train, or IoT device, wearable device (e.g., watch, head-mounted device, glasses), industrial equipment, or devices within the aforementioned equipment (e.g., communication module, modem, or chip), etc. Future generations of ED 110 may be referred to using other terms. Base stations 170a and 170b are T-TRPs and will be referred to hereinafter as T-TRP 170. Figure 1CThe diagram also shows NT-TRP, which will be referred to below as NT-TRP 172. Each ED 110 connected to T-TRP 170 and / or NT-TRP 172 can be dynamically or semi-statically enabled (i.e., established, activated, or enabled), disabled (i.e., released, deactivated, or disabled), and / or configured in response to one or more of connectivity availability and connectivity necessity.

[0073] ED 110 includes a transmitter 111 and a receiver 113 coupled to one or more antennas 104. Only one antenna 104 is shown. One, some, or all of the antennas 104 may also be panels. The transmitter 111 and receiver 113 may, for example, be integrated as a transceiver. The transceiver is used to modulate data or other content for transmission by at least one antenna 104 or via a network interface controller (NIC). The transceiver is also used to demodulate data or other content received by at least one antenna 104. Each transceiver includes any suitable structure to generate signals for wireless or wired transmission and / or process signals received wirelessly or wiredly. Each antenna 104 includes any suitable structure to transmit and / or receive wireless or wired signals.

[0074] ED 110 includes at least one memory 115. Memory 115 stores instructions and data used, generated, or collected by ED 110. For example, memory 115 may store software instructions or modules executed by one or more processing units (e.g., processor 117) for implementing some or all of the functions and / or embodiments described herein. Each memory 115 includes any suitable one or more volatile and / or non-volatile storage and retrieval devices. Any suitable type of memory can be used, such as random access memory (RAM), read-only memory (ROM), hard disk, optical disk, subscriber identity module (SIM) card, memory stick, secure digital (SD) memory card, and on-processor cache, etc.

[0075] ED 110 may also include one or more input / output devices (not shown) or interfaces (e.g., a wired interface connected to the Internet 150 in Figure 1). The input / output devices support interaction with the user or other devices in the network. Each input / output device includes any structure suitable for (e.g., by operation) providing or receiving information from the user, such as a speaker, microphone, numeric keypad, keyboard, display, or touchscreen.

[0076] ED 110 includes a processor 117 for performing various operations, including operations related to: preparing to transmit uplink transmissions to NT-TRP 172 and / or T-TRP 170; processing downlink transmissions received from NT-TRP 172 and / or T-TRP 170; and processing lateral link transmissions to and from another ED 110. Processing operations related to preparing to transmit uplink transmissions may include operations such as encoding, modulation, transmit beamforming, and generating symbols for transmission. Processing operations related to processing downlink transmissions may include operations such as receive beamforming, demodulation, and decoding of received symbols. According to an embodiment, the downlink transmissions may be received by receiver 113, possibly using receive beamforming, and processor 117 may extract signaling from the downlink transmissions (e.g., by detecting and / or decoding signaling). For example, an example of signaling may be a reference signal transmitted by NT-TRP 172 and / or T-TRP 170. In some embodiments, processor 117 performs transmit beamforming and / or receive beamforming based on beam direction indications received from T-TRP 170, such as beam angle information (BAI). In some embodiments, processor 117 may perform operations related to network access (e.g., initial access) and / or downlink synchronization, such as operations related to detecting synchronization sequences, decoding, and acquiring system information. In some embodiments, processor 117 may perform channel estimation, for example, using reference signals received from NT-TRP 172 and / or T-TRP 170.

[0077] Although not shown, processor 117 may be part of transmitter 111 and / or receiver 113. Although not shown, memory 115 may be part of processor 117.

[0078] The processing components of processor 117, transmitter 111, and receiver 113 can each be implemented using one or more processors, which are the same or different, to execute instructions stored in memory (e.g., memory 115). Alternatively, some or all of the processing components of processor 117, transmitter 111, and receiver 113 can each be implemented using dedicated circuitry such as a field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), or application-specific integrated circuit (ASIC).

[0079] In some implementations, the T-TRP 170 can use other names, such as base station, basetransceiver station (BTS), wireless base station, network node, network device, network-side device, transmit / receive node, NodeB, evolved NodeB (eNodeB or eNB), home eNodeB, next-generation NodeB (gNB), transmission point (TP), site controller, access point (AP), wireless router, relay station, remote radio head, ground node, ground network device, ground base station, base band unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), location node, etc. The T-TRP 170 can be a macro BS, pico BS, relay node, or host node, or a combination thereof. T-TRP 170 may refer to the aforementioned device or a component within the aforementioned device (e.g., a communication module, modem, or chip).

[0080] In some embodiments, the various parts of T-TRP 170 can be distributed. For example, some modules of T-TRP 170 may be located remotely from the device housing the antenna 256 for T-TRP 170 and may be coupled to the device housing the antenna 256 via a communication link (not shown) sometimes referred to as a fronthaul (e.g., a common public radio interface, CPRI). Therefore, in some embodiments, the term T-TRP 170 may also refer to modules on the network side that perform processing operations such as determining the location of ED 110, resource allocation (scheduling), message generation, and encoding / decoding; these modules are not necessarily part of the device housing the antenna 256 of T-TRP 170. These modules may also be coupled to other T-TRPs. In some embodiments, T-TRP 170 may actually be multiple T-TRPs that operate together (e.g., by using coordinated multicast) to service ED 110.

[0081] T-TRP 170 includes at least one transmitter 181 and at least one receiver 183 coupled to one or more antennas 256. Only one antenna 256 is shown in the figure to avoid congestion. One, some, or all of the antennas 256 may also be panels. Transmitter 181 and receiver 183 may be integrated as a transceiver. T-TRP 170 also includes a processor 182 for performing various operations, including operations related to: preparing to transmit downlink transmissions to ED 110; processing uplink transmissions received from ED 110; preparing to transmit backlink transmissions to NT-TRP 172; and processing transmissions received from NT-TRP 172 via backlink. Processing operations related to preparing to transmit downlink or backlink transmissions may include operations such as encoding, modulation, precoding (e.g., multiple-input multiple-output (MIMO) precoding), transmit beamforming, and generating symbols for transmission. Processing operations related to processing transmissions received in the uplink or via backlink may include operations such as receive beamforming, demodulating received symbols, and decoding received symbols. Processor 182 can also perform operations related to network access (e.g., initial access) and / or downlink synchronization, such as generating the contents of a synchronization signal block (SSB), generating system information, etc. In some embodiments, processor 182 also generates beam direction indications, such as BAI, that scheduler 184 can schedule for transmission. Processor 182 performs other network-side processing operations described herein, such as determining the location of ED 110, determining the location for deploying NT-TRP 172, etc. In some embodiments, processor 182 can generate signaling, such as for configuring one or more parameters of ED 110 and / or one or more parameters of NT-TRP 172. Any signaling generated by processor 182 is transmitted by transmitter 181. Note that "signaling" as used herein can also be referred to as control signaling. Signaling can be transmitted in a physical layer control channel (e.g., a physical downlink control channel (PDCCH)), in which case the signaling can be referred to as dynamic signaling. Signaling transmitted in the downlink physical layer control channel is called downlink control information (DCI). Signaling transmitted in the uplink physical layer control channel is called uplink control information (UCI). Signaling transmitted in the sidelink physical layer control channel is called sidelink control information (SCI).Signaling can be included in higher-layer (e.g., above the physical layer) data packets transmitted in physical layer data channels (e.g., physical downlink shared channel, PDSCH). In this case, the signaling can be referred to as higher-layer signaling, static signaling, or semi-static signaling. Higher-layer signaling can also refer to Radio Resource Control (RRC) protocol signaling or Media Access Control-Control Element (MAC-CE) signaling.

[0082] Scheduler 184 may be coupled to processor 182. Scheduler 184 may be included within T-TRP 170 or may operate separately from it. Scheduler 184 may schedule uplink, downlink, and / or backlink transmissions, including issuing scheduling authorizations and / or configuring unscheduled (“configured authorization”) resources. T-TRP 170 also includes memory 185 for storing information and data. Memory 185 stores instructions and data used, generated, or collected by T-TRP 170. For example, memory 185 may store software instructions or modules for implementing some or all of the functions and / or embodiments described herein and executed by processor 182.

[0083] Although not shown, processor 182 may form part of transmitter 181 and / or receiver 183. Furthermore, although not shown, processor 182 may implement scheduler 184. Although not shown, memory 185 may form part of processor 182.

[0084] The processing components of processor 182, scheduler 184, transmitter 181, and receiver 183 can each be implemented using one or more processors, which are the same or different, to execute instructions stored in memory (e.g., memory 185). Alternatively, some or all of the processing components of processor 182, scheduler 184, transmitter 181, and receiver 183 can be implemented using dedicated circuitry such as FPGA, GPU, CPU, or ASIC.

[0085] Although the NT-TRP 172 is illustrated only as an example of a drone, it can be implemented using any suitable non-terrestrial form, such as an aerial platform, satellite, an aerial platform as an international mobile telecommunications base station, and unmanned aerial vehicles, which will be discussed below. Furthermore, in some implementations, the NT-TRP 172 may be referred to by other names, such as a non-terrestrial node, a non-terrestrial network device, or a non-terrestrial base station. The NT-TRP 172 includes a transmitter 186 and a receiver 187 coupled to one or more antennas 108. Only one antenna 108 is shown in the figure to avoid congestion. One, some, or all of the antennas may also be panels. The transmitter 186 and receiver 187 may be integrated as a transceiver. The NT-TRP 172 also includes a processor 188 for performing various operations, including those related to: preparing downlink transmissions to ED 110; processing uplink transmissions received from ED 110; preparing return transmissions to T-TRP 170; and processing transmissions received from T-TRP 170 via return transmissions. Processing operations related to preparing downlink or backhaul transmissions may include operations such as encoding, modulation, precoding (e.g., MIMO precoding), transmit beamforming, and generating symbols for transmission. Processing operations related to processing transmissions received in the uplink or via backhaul may include operations such as receive beamforming, demodulating received symbols, and decoding received symbols. In some embodiments, processor 188 performs transmit beamforming and / or receive beamforming based on beam direction information (e.g., BAI) received from T-TRP 170. In some embodiments, processor 188 may generate signaling, for example, to configure one or more parameters of ED 110. In some embodiments, NT-TRP 172 implements physical layer processing but does not implement higher-layer functions, such as medium access control (MAC) layer or radio link control (RLC) layer functions. Since this is only an example, NT-TRP 172 may more generally implement higher-layer functions in addition to physical layer processing.

[0086] The NT-TRP 172 also includes a memory 189 for storing information and data. Although not shown, a processor 188 may be part of a transmitter 186 and / or a receiver 187. Although not shown, the memory 189 may be part of a processor 188.

[0087] The processing components of processor 188, transmitter 186, and receiver 187 can each be implemented using one or more processors, which may be the same or different, to execute instructions stored in memory (e.g., memory 189). Alternatively, some or all of the processing components of processor 188, transmitter 186, and receiver 187 can be implemented using dedicated circuitry such as a programmable FPGA, GPU, CPU, or ASIC. In some embodiments, NT-TRP 172 may actually be multiple NT-TRPs that operate together, for example, through coordinated multicast transmissions, to service ED 110.

[0088] T-TRP 170, NT-TRP 172 and / or ED 110 may include other components, but for clarity these components have been omitted.

[0089] Figure 1D Examples of units or modules in a device related to some embodiments of the present invention are shown. One or more steps of the methods of the embodiments provided herein can be performed by... Figure 1D The corresponding unit or module provided will be executed. Figure 1D Units or modules in devices such as ED110, T-TRP 170, or NT-TRP 172 are illustrated. For example, signals can be transmitted by a transmitting unit or transmitting module. Signals can be received or input by a receiving unit or receiving module. Signals can be processed by a processing unit or processing module. Other steps can be performed by an artificial intelligence (AI) module or a machine learning (ML) module. The corresponding units or modules can be implemented using hardware, one or more components or devices executing software, or a combination thereof. For example, one or more of these units or modules can be integrated circuits such as programmable FPGAs, GPUs, CPUs, or ASICs. It should be understood that if these modules are implemented, for example, using software executed by a processor, the processor can retrieve these modules, in whole or in part, as needed, individually or collectively for processing, in one or more instances, and these modules themselves can include instructions for further deployment and instantiation.

[0090] Although not shown, the transmitting and receiving modules can be part of or combined with the transceiver module. The transceiver module can also be called an interface module, or simply an interface, and is used for input and output operations.

[0091] Additional details regarding ED 110, T-TRP 170, and NT-TRP 172 are known to those skilled in the art. Therefore, these details are omitted here.

[0092] To support the use of AI in wireless networks, a suitable AI framework is needed. However, traditional communication technologies, such as 5G-related technologies, only consider AI use cases to improve network performance. Traditional communication technologies do not support networks used to provide AI services to user equipment (UE). Considering the large number of devices with data and computing capabilities expected in future networks, it is hoped that AI services can be provided in these future networks. Some promising technologies for leveraging the data and computing capabilities of these devices include, for example, distributed training and / or inference techniques.

[0093] Various aspects of this invention relate to the process of establishing AI services between a user equipment and a network. These processes include exchanging at least some requests and / or configurations and / or acknowledgments between a user equipment and a base station via an air interface.

[0094] Applications based on artificial intelligence (AI), particularly in the form of large neural networks, will be widely adopted in future devices and networks. Given constraints in size, power, and weight, the computing and storage capabilities on the device side are typically limited compared to those in the network. To popularize AI-based applications across a wider range of devices, supporting the splitting of AI computing between user devices and the network is becoming a trend. In this context, the network acts as a platform, providing AI computing services requested from user devices.

[0095] In some aspects of this invention, AI computing services primarily focus on AI inference services, which may include a series of AI inference tasks requested from a user device. An AI inference task can be viewed as an AI computation process along a given AI model from input data to an output result. In the case of distributed AI inference between a user device and the network, the input to the AI ​​model and certain layers (which can be zero) are computed at the user device, while the remaining layers of the AI ​​model up to the output are computed at the network. The decision of the model split point between two nodes and the corresponding AI computation allocation are part of AI model management. For example, if an AI model has N layers, and the split point is after layer n, then the first node (e.g., node A) will compute from the input to layer n, and the second node (e.g., node B) will complete the rest of the model, from layer n+1 to layer N, until the output of the inference result. Before providing any AI services from the network to the user device, it is essential to establish services and ensure that both the user device and the network understand which part of the AI ​​inference task they are jointly computing.

[0096] According to an embodiment of the present invention, a solution for model splitting is provided. In one aspect, a first device sends a service request message to a second device for performing a task based on a model. Based on a service response message received from the second device including model splitting information, the first device performs the task based on the model and the model splitting information, wherein the model splitting information indicates the splitting of model computation used for performing the task, at least between the first and second devices. Based on a rejection notification message received from the second device, the first device suspends the model-based task execution. Therefore, model splitting can be configured by the second device. Thus, model performance is improved. The following will combine... Figures 2 to 13 The principles and implementation methods of the embodiments of the present invention are described in detail.

[0097] Figure 2 An exemplary signaling diagram of an exemplary process according to some embodiments of the present invention is shown. Process 200 may involve a first device 201 and a second device 202. Process 200 may also involve at least one of a third device 203 or a fourth device 204. Figure 2 The first device 201 in the middle can be Figure 1A Examples of communication electronic devices 110 or network nodes 170. Figure 2 The second device 202 in the middle can be Figure 1A Examples of communication electronic devices 110 or network nodes 170. For example, the first device 201 may be a terminal device. Additionally, the second device 202 may be an access network device. The third device 203 may be a core network device or an access network device. The fourth device 204 may be an access network device. It should be understood that, although already... Figure 1A The processing flow 200 is described in the communication system 100A, but this process can also be applied to other communication scenarios.

[0098] In processing flow 200, the first device 201 sends a service request message 215 (210) to the second device 202 for performing a task based on a model. Correspondingly, the second device 202 receives the service request message 215 (220) from the first device 201. The model can be an AI model, and the task can be an AI service. The involved AI nodes (e.g., the first device 201 and / or the second device 202) can reach a consensus on the AI ​​model selected for a given AI service through AI model lifecycle management (LCM). This AI model LCM can involve further alignment of supporting parameters for the selected AI model, such as modeling split-related parameters representing different trade-offs between the computational / communication resources required by each AI node.

[0099] For determining model splitting information, it is assumed that both the first device 201 and the second device 202 have reached a consensus on which model to use for the AI ​​inference task. In cases where it is not assumed that the first device 201 and the second device 202 have reached a consensus on which model to use for the AI ​​inference task, a process is required to obtain consensus on which model to use. This process can be completed before the task, for example, when the terminal device registers for the AI ​​service from the network. In one embodiment, it can be assumed that model selection is performed by a model manager, which may be a logical functional node located in the radio access network (e.g., the second device 202) or the core network.

[0100] In some embodiments, the first device 201 may send a registration request message for registering a task to the second device 202. Additionally, the registration request message may include service latency for task execution, accuracy requirements for the task, or a combination of both. After receiving the registration request message from the first device 201, the second device 202 may determine the target model as the model to be used for task execution. Then, the second device 202 may send a model notification message to the first device 201. The model notification message indicates the model to be used for task execution. Accordingly, the first device 201 may receive the model notification message from the second device 202.

[0101] Additionally, the registration request message may include service latency for performing the task, accuracy requirements for the task, or a combination of both.

[0102] The model manager can be located either inside or outside the second device 202. For example, the model manager can be located inside the third device 203. In order to identify the target model, in one example, the second device 202 can identify the target model as the model that will be used to perform the task.

[0103] In another example, the second device 202 may forward the registration request message 205 206 to the third device 203 and receive from the third device 203 an indication 222 223 of the target model to be used to perform the task. At the other end of the communication, after receiving the registration request message 207 206 from the second device 202, the third device 203 may determine 208 the model to be used to perform the task. Then, the third device 203 may send the indication 221 222 of the target model to the second device 202.

[0104] In cases where multiple devices participate in the task, the model notification message may include information indicating the multiple devices that will participate in the task, and these multiple devices may include the second device 202. The multiple devices may also include a fourth device 204. For example, the model manager may not be located within the second device 202, and another device (e.g., a third device 203 and / or a fourth device 204) may participate in the AI ​​inference service task.

[0105] If AI inference involves multiple network devices (e.g., third device 203 and / or fourth device 204), and the model manager is located within third device 203, then second device 202 is aware of fourth device 204. Second device 202 can receive from third device 203 an instruction 226 indicating that fourth device 204 will participate in the model-based task execution. Accordingly, third device 203 can send to second device 202 an instruction 224 indicating that fourth device 204 will participate in the model-based task execution.

[0106] Additionally, the second device 202 can send an instruction 232, 231, to the fourth device 204, which will participate in performing the task, regarding the target model to be used for the task. If the AI ​​inference involves multiple network devices (e.g., a third and / or a fourth device), and the model manager is located within the second device 202, the second device 202 can notify the fourth device of the model selection for a given AI service from the first device 201 when sending a decision to the first device 201 regarding which model to use. Accordingly, the fourth device 204 can receive an instruction 232, 233, regarding the target model from the second device 202.

[0107] In some embodiments, the second device 202 may send an indication 229 of anchor devices for a task from among 228 or more devices to the first device 201. At the other end of the communication, the first device 201 may receive an indication of anchor devices for a task from among 230 or more devices from either the second device 202 or the third device 203. If the AI ​​inference service may involve multiple devices (e.g., the second device 202 and the third device 203), the model manager may also specify anchor devices for the AI ​​service's tasks. In one example, the model manager is located within the second device 202, which may indicate anchor devices to the first device 201. In another example, the device used by the first device 201 to register the service (e.g., the second device 202) may be an anchor device. The first device 201 may request AI inference tasks through the second device 202, and the second device 202 may determine the model split point between the first device 201 and the second device 202, as well as the model split point between the second device 202 and the third device 203.

[0108] In one example, the first device 201 may know that the second device 202 is responsible for the inference computation between the second device 202 and the third device. The first device 201 may not know the third device 203.

[0109] In some embodiments, the first device 201 may also send at least one of its computing power or computing resources for performing tasks based on the model to the second device 202. Correspondingly, the second device 202 may receive at least one of the first device's computing power or computing resources for performing tasks based on the model from the first device 201. For example, the first device 201 may be willing to report its own computational complexity to the second device 202, and the second device 202 may determine the model split point as part of AI model management. The first device 201 may report its maximum computing power and available computing resources before or along with the service request message 215 for the AI ​​service.

[0110] Additionally, service request message 215 may include information indicating the service latency required to perform the task, the task priority, or a combination of both. For example, a service request message for an AI service may be associated with a service latency requirement for the AI ​​service, and optionally with a service priority for the AI ​​service.

[0111] In some embodiments, the first device 201 may not transmit its computing power or computing resources to the second device 202. The first device 201 may determine the computation time required for completing a first portion of the task at a reference model split point. Based on the service latency and computation time used to perform the task, the first device 201 may determine the remaining latency. For example, the first device 201 may not report its own computational complexity to the second device 202; instead, it tells the second device 202 that the latency budget (i.e., service latency) does not include its own computation time (i.e., remaining latency). The second device 202 determines whether the AI ​​service can be provided within the requested latency budget and identifies the model split point as part of AI model management.

[0112] Alternatively or additionally, the first device 201 can estimate the computational cost of the model at multiple candidate model split points. Based on the computational cost, the first device 201 can select one candidate model split point as a reference model split point. For example, the first device 201 can determine the computational time that might be required to complete AI model inference at the reference model split point. The reference model split point can be obtained by estimating the computational cost at different split points, and then the first device 201 can select one or more split points that the reference model split point can provide.

[0113] Additionally, the remaining latency may include a first transmission time required for the first device to send the intermediate results of the first part of the task performed by the first device to the second device, the computation time required for the second part of the task to be performed at the second device, and a second transmission time for the second device to send the results of the task to the first device. The first device 201 may include information about the service latency requested by the network side when requesting AI services; this latency is the total service latency requirement minus the computation time at the first device 201. In other words, the remaining latency may include the communication time from the first device 201 to the second device 202 for transmitting intermediate results, the computation time at the second device 202, and the transmission time of the inference results from the second device 202 to the first device 201 (and may also include resource scheduling time).

[0114] If the first device 201 does not transmit its computing power or resources to the second device 202, the service request message 215 may include information indicating remaining latency, a reference model split point, task priority, or any combination of two or more of the above. The reference model split point may be included in the service request message 215, allowing the second device 202 to know which layer of the AI ​​model to continue inference from, and the available service priorities. Furthermore, the granularity and candidate values ​​for determining the remaining latency or selecting the reference model split point during processing may be predefined.

[0115] Continue to refer to Figure 2 The second device 202 determines whether the task will be completed within the service delay, given that model calculations are split between at least the first device 201 and the second device 202 in order to perform the task.

[0116] As one embodiment, to determine whether a task will be completed within the service latency, the second device 202 may determine an estimated latency for executing the task based on at least one of computing power or computing resources, and compare the estimated latency with the service latency. When the computing power or computing resources of the first device 201 are sent to the second device 202, the second device 202 may estimate the latency to see if the task can be completed within the service latency. If the estimated latency is greater than the remaining latency, the task cannot be completed within the service latency. If the estimated latency is less than or equal to the remaining latency, the task can be completed within the service latency.

[0117] Furthermore, the estimated latency may include: an estimate of the computation time for the first device 201 to execute a first portion of the task, given the computing power of the first device 201 and the model split point; and an estimate of the computation time for the second device 202 to execute a second portion of the task, given the computing power of the second device 202 and the model split point. The first portion of the task may refer to the model inference portion computed at the first device 201, and the first portion of the task may refer to the model inference portion computed at the second device 202. For example, the estimated latency may include estimates of the computation time for both the first device 201 and the second device 202 based on their computing power and potential model management (e.g., the model split point). Additionally, the estimated latency may also include an estimate of the communication time between the first and second devices, given the amount of data associated with the model split point. For example, the estimated latency may include an estimate of the communication time, given the amount of data associated with a model management decision (e.g., the split point).

[0118] In another embodiment, to determine whether a task will be completed within the service delay, the second device 202 can determine an estimated delay by adding the estimated values ​​of the first transmission time, the computation time, and the second transmission time. The second device 202 can then compare the estimated delay with the remaining delay. If the computing power or resources of the first device 201 are not sent to the second device 202, the second device 202 can estimate the delay required to execute the task and the communication between the first device 201 and the second device 202. If the estimated delay is greater than the remaining delay, the task cannot be completed within the service delay. If the estimated delay is less than or equal to the remaining delay, the task can be completed within the service delay.

[0119] Continue to refer to Figure 2 Based on the determination that the task will be completed within the service latency, the second device 202 sends a service response message 236, including model splitting information, to the first device 201. The model splitting information indicates the splitting of model computation. For example, if the second device 202 estimates that the task can be completed within the latency requirement, or that the task has a high probability of being completed within the latency requirement, the second device will respond with a service response message 236 carrying a model splitting decision.

[0120] In some embodiments, the service response message 236 may further include resource allocation for transmitting intermediate results of a first portion of a task performed by the first device 201. For example, communication resources may be allocated by the second device 202 for transmitting the output of an AI model inference portion computed at the first device 201, which will be used as input to AI model inference computed at the second device 202.

[0121] After sending service response message 236 to the first device, the second device 202 can execute a task based on the model and model splitting information. Additionally, the second device 202 can receive intermediate results from the first device of the first part of the task execution. Based on the intermediate results and the model, the second device 202 can execute the second part of the task indicated by the model splitting information. Then, the second device 202 can send the results of the task execution to the first device 201.

[0122] Alternatively or additionally, the first device 201 may send a first confirmation message 242 (241) regarding model splitting information to the second device 202. Correspondingly, the second device 202 may receive the first confirmation message 242 (243) from the first device 201. The first device 201 may confirm the model splitting decision by sending a dedicated message. The AI ​​inference task can then be jointly performed by the first device 201 and the second device 202.

[0123] Additionally, the first device 201 can send a second confirmation message 245 (244) to the second device 202 regarding the model to be used. Correspondingly, the second device 202 can receive the second confirmation message 245 (246) from the first device 201. Afterwards, the first device 201 and the second device 202 can obtain common knowledge about which model to use.

[0124] If the model used to perform the task is determined by the third device 203, the first device 201 can also send a third confirmation message 247, indicating the model to be used, to the third device 203. In another example, the first device 201 can send a third confirmation message to the second device 202, and the second device 202 can forward the third confirmation message to the third device. Accordingly, the third device 203 can receive the third confirmation message 248.

[0125] If the AI ​​inference involves multiple network devices, including the fourth device 204, the first device 201 can send a fourth confirmation message 251 (250) to the fourth device 204 among the multiple devices, indicating the model to be used. Accordingly, the fourth device 204 can receive the fourth confirmation message 251 (252). For example, when sending confirmation to the second device 202 (which has a model manager), the first device 201 can also send confirmation / handshake information to the fourth device 204.

[0126] It should be understood that the first confirmation message 242, the second confirmation message 245, the third confirmation message 248, and the fourth confirmation message 251 are in Figure 2 The position in the message is not restrictive; for example, the second confirmation message 245, the third confirmation message 248, and the fourth confirmation message 251 can be sent before the first confirmation message 242.

[0127] Continue to refer to Figure 2 Based on receiving a service response message 236 including model splitting information from the second device 202, the first device 201 performs task 253 based on the model and the model splitting information, and the model splitting information indicates the splitting of model computation for performing the task at least between the first device and the second device.

[0128] In some embodiments, in order to perform a task based on a model and model splitting information, the first device 201 may perform a first part of the task indicated by the model splitting information based on the model. The first device 201 may send intermediate results of performing the first part of the task to the second device 202. Afterwards, the first device 201 may receive the results of performing the task from the second device 202.

[0129] The task can be executed between the first device 201 and the second device 202, or between two or more devices. In some embodiments, to execute the task based on a model and model splitting information, the second device 202 can send an instruction to at least one other device. This instruction indicates that at least one other device performs a third part of the task based on the model. The second device 202 can then receive intermediate results from the execution of the third part of the task from the at least one other device. The model can be split into three or more parts, and at least one other device can participate in the execution of the task. In one example, the second device 202 can delegate some AI inference computation to a set of other devices that are not visible to the first device 201.

[0130] Additionally, the second device 202 can determine the estimated latency for performing the task by including: the computation time of at least one other device performing the third part of the task; and the information exchange latency between the second device 202 and the at least one other device. For the second device 202, when estimating the total latency required to complete the task, it needs to consider the computation time from all participating devices and the potential information exchange latency between these devices.

[0131] Continue to refer to Figure 2 Based on the determination that the task cannot be completed within the service latency, the second device 202 sends a rejection notification message 265 (260) to the first device 201. Upon receiving the rejection notification message 265 (270) from the second device 202, the first device 201 suspends task execution based on the model (275). If the second device 202 estimates that the task cannot be completed before the latency budget, for example, if the estimated latency is less than or equal to the service latency, then the second device will use the rejection notification message 265 to reject the service request message.

[0132] Alternatively or additionally, the rejection notification message may include information instructing a fallback timer. For example, the second device 202 may send a recommended fallback timer to the first device 201 in response. Furthermore, the fallback timer may be defined based on the priority of the service request message and a random number. The fallback timer may be defined using explicit rules, such as a combination of request priority and a random number. For example, a higher priority will result in a shorter timer.

[0133] If the second device 202 rejects the service request, the first device 201 may request the task again. In some embodiments, the first device 201 may send an additional service request message to the second device 202 for performing the task after the rollback timer expires. Accordingly, the second device 202 may receive an additional service request message from the first device for performing the task after the rollback timer expires.

[0134] Generally, some embodiments of process 200 involve establishing AI services across two or more devices (e.g., user equipment and a network). The splitting of the model depends on flexible conditions that may include the devices' willingness and level of reporting their computing power, and / or the location of the decision-maker managing the AI ​​model, etc. In this way, the model can be flexibly split according to the actual conditions on both the user equipment and network sides.

[0135] Figure 3 An exemplary process for a proposed solution according to some embodiments of the present invention is illustrated. Process 300 may involve node A 301 and node B 302. It should be understood that process 300 can be considered as Figure 2 A more specific example of process 200. Therefore, Figure 3 Node A 301 in the data can be Figure 2 Example of the first device 201 in the example, Figure 3 Node B 302 in the data can be Figure 2 Example of the second device 202 in the example.

[0136] like Figure 3 As shown in 311, it is assumed that node A 301 and node B 302 have reached a consensus on model selection for a given AI task. In 313, node A 301 can report its computing power to node B 302. The computing power can be reported before the AI ​​service request, or it can be reported together with the AI ​​service request in 315.

[0137] At 315, Node A 301 can request an AI inference service with service latency requirements. This request can also be associated with Node A 301's computing power and service priority. At 317, Node B 302 estimates whether the task can be completed within a given latency budget.

[0138] If node B 302 estimates that the task can be completed within the latency requirement, or that there is a high probability that the task can be completed within the latency requirement, then node B 302 will respond to the service request at 319 using the model split decision. Node B 302 can also respond to the service request using the allocation of communication resources for transmitting the output of the AI ​​model inference part computed at node A 301. At 321, node A 301 can send an acknowledgment of model management (e.g., model split point) to node B 302. Then, nodes A 301 and B 302 begin the task at 323 based on the acknowledged model split.

[0139] If node B 302 estimates that the task cannot be completed before the latency budget, node B 302 responds to the service request using a service denial notification at 325. Node B 302 may also send a recommended backoff timer to node A 301 as a response.

[0140] Figure 4 Exemplary processes illustrating the proposed solutions of some embodiments of the present invention are shown. Process 400 may involve node A 401 and node B 402. It should be understood that process 400 can be considered as Figure 2 A more specific example of process 200. Therefore, Figure 4 Node A 401 in the data can be Figure 2 Example of the first device 201 in the example, Figure 4 Node B 402 in the data can be Figure 2 Example of the second device 202 in the example.

[0141] like Figure 4 As shown in section 411, it is assumed that node A 401 and node B 402 have reached a consensus on model selection for a given AI task. In section 413, node A 401 can request an AI reference service with a remaining latency budget and a reference split point. This request can also be associated with a service priority. In section 415, node B 402 estimates whether the task can be completed within the given latency budget. This estimation can be made using knowledge of available computational and / or communication resources.

[0142] If node B 402 estimates that the task can be completed within the remaining latency requirement, or that there is a high probability that the task can be completed within the remaining latency requirement, then node B 402 will respond to the service request at 417 using the model split decision. Node B 402 can also respond to the service request using the allocation of communication resources for transmitting the output of the AI ​​model inference portion computed at node A 401, which will be used as the input to the AI ​​model inference computed at node B 402. At 419, node A 401 can send an acknowledgment of model management (e.g., model split point) to node B 402. Then, nodes A 401 and B 402 begin the task at 421 based on the acknowledged model split.

[0143] If node B 402 estimates that the task cannot be completed before the latency budget, node B 402 responds to the service request with a service denial notification at 423. Node B 402 may also send a recommended backoff timer to node A 401 as a response. Node A 401 may send another service request after the recommended timer expires.

[0144] Figure 5 An exemplary process for a proposed solution according to some embodiments of the present invention is illustrated. Process 500 may involve node A 501 and node B 502. It should be understood that process 500 can be considered as Figure 2 A more specific example of process 200. Therefore, Figure 5 Node A 501 in the data can be Figure 2 Example of the first device 201 in the example, Figure 5 Node B 502 in the middle can be Figure 2 Example of the second device 202 in the example.

[0145] like Figure 5 As shown, the model manager is located inside node B 502. At 511, node A 501 registers the AI ​​inference service through node B 502. Node A 501 sends a registration request and optionally provides service latency and / or accuracy requirements to node B 502.

[0146] In step 513, the model manager in node B 502 decides which AI model to use for a given AI service, and then node B 502 notifies node A 501. In step 515, node A 501 can send confirmation of the model to node B 502. In step 517, nodes A 501 and B 502 reach a consensus on which model to use.

[0147] Figure 6Exemplary processes illustrating the proposed solutions of some embodiments of the present invention are shown. Process 600 may include node A 601, node B 602, and model manager 603. It should be understood that process 600 can be considered as... Figure 2 A more specific example of process 200. Therefore, Figure 6 Node A 601 in the data can be Figure 2 Example of the first device 201 in the example, Figure 6 Node B 602 in the data can be Figure 2 Example of the second device 202 in the example.

[0148] like Figure 6 As shown, the model manager is outside of node B 602. The model manager can be in the core network, or in another node that node A 601 is not initially connected to, or it is not the anchor node that node A 601 depends on when applying AI services.

[0149] At 611, Node A 601 sends a registration request and optionally provides service latency and / or accuracy requirements to Node B 602. At 613, Node B 602 forwards the AI ​​service registration request to Model Manager 603. At 615, Model Manager 603 determines the AI ​​model to be used for the given AI service and then notifies Node B 602.

[0150] At 617, the decision from the model manager is forwarded to node A 601 via node B 602. At 619, node A 601 can send an acknowledgment to node B 602, and node B 602 can forward this acknowledgment to the model manager at 621. Alternatively, node A 601 can send an acknowledgment directly to the model manager. At 623, nodes A 601 and B 602 reach a consensus on which model to use for a given AI service.

[0151] For process 600, the model manager in another node (e.g., node C) will make a decision and align the parameters of the selected AI model to be used by nodes A 601 and B 602 by taking into account the maximum and / or available communication / computing resources of nodes A 601 and B 602.

[0152] In processes 300, 400, 500, and 600, it is assumed that node B is the only collaborating node that jointly performs AI services with node A. The above processes can be modified when other nodes (such as node C) participate in the joint AI inference task. Expanding from another node (i.e., node C) to a set of nodes (e.g., {C1, C2, ..., Cn}) is straightforward.

[0153] Figure 7Exemplary processes of the proposed solutions, representing some embodiments of the present invention, are illustrated. Process 700 may involve node A 701, node B 702 with a model manager, and node C 703. It should be understood that process 700 can be considered as... Figure 2 A more specific example of process 200. Therefore, Figure 7 Node A 701 in the data can be Figure 2 Example of the first device 201 in the example, Figure 7 Node B 702 in the middle can be Figure 2 Example of the second device 202 in the example.

[0154] like Figure 7 As shown, the model manager is located within node B 702, and another node C 703 will participate in the AI ​​inference service as a collaborating node. At 711, node A 701 registers for the AI ​​inference service through node B 702. Node A 701 sends a registration request and optionally provides service latency and / or accuracy requirements to node B 702.

[0155] At 713, the model manager in node B 702 decides on the AI ​​model to be used for a given AI service, and then node B 702 notifies node A 701. At 714, node B 702 also notifies node C 703 of the model selection for the AI ​​inference service of node A 701.

[0156] At 715, node A 701 can send a confirmation of the model to node B 502. At 717, node A 701 can also send confirmation / handshake information to node C 703. At 719, nodes A 701, B 502, and C 703 reach a consensus on which model to use.

[0157] Figure 8 Exemplary processes illustrating the proposed solutions of some embodiments of the present invention are shown. Process 800 may involve node A 801, node B 802, model manager 803, and node C 804. It should be understood that process 800 can be considered as... Figure 2 A more specific example of process 200. Therefore, Figure 8 Node A 801 in the data can be... Figure 2 Example of the first device 201 in the example, Figure 8 Node B 802 in the middle can be Figure 2 Example of the second device 202 in the example.

[0158] like Figure 8As shown, outside of node B 802, there is another node C 804 that will participate in the AI ​​inference service as a collaborating node. At 811, node A 801 registers for the AI ​​inference service through node B 802. Node A 801 sends a registration request and optionally provides service latency and / or accuracy requirements to node B 802. At 813, node B 802 forwards the AI ​​service registration request to model manager 803. At 815, model manager 803 determines the AI ​​model to be used for the given AI service and then notifies node B 802, and may optionally notify another collaborating node, node C 804. At 817, node B 802 also notifies node C 803 as a collaborating node for the AI ​​inference service used by node A 801.

[0159] At 821, node A 801 can send a confirmation of the model to node B 802, and node B 802 can forward this confirmation to the model manager at 823. At 825, node A 801 can also send confirmation / handshake information to node C 803. At 827, nodes A 801, B 802, and C 803 reach a consensus on which model to use.

[0160] If the AI ​​inference service may involve multiple network nodes (e.g., node B 802 and node C 803), the model manager can also specify anchor nodes for the AI ​​service. For example, node B 802 can be the anchor node, and node A 801 can request AI inference tasks through node B 802. Node B 802 can determine the model splitting points between node A 801 and node B 802, and between node B 802 and node C 804. Node A 801 knows that node C 804 is not required. Node A 801 knows to allocate inference computation to node B 802 between node B 802 and node C 804.

[0161] In process 800, the model manager 803 will make decisions by considering the maximum and / or available communication / computing resources of nodes A 801, B 802, and C 804 and align the parameters of the selected AI model to be used by all participating nodes A 801, B 802, and C 804.

[0162] Figure 9 A flowchart illustrating an exemplary method 900 implemented at a first device according to some embodiments of the present invention is shown. For ease of discussion, reference will be made to... Figure 1A Method 900 is described from the perspective of communication electronic device 110. It should be understood that method 900 may include additional actions not shown and / or some actions shown may be omitted, and the scope of the invention is not limited thereto.

[0163] At box 910, the first device sends a service request message for performing a model-based task to the second device. At box 920, based on a service response message received from the second device including model splitting information, the first device performs the task based on the model and the model splitting information, wherein the model splitting information indicates the splitting of model computation used to perform the task, at least between the first and second devices. At box 930, based on a rejection notification message received from the second device, the first device suspends the model-based task execution. The first device may be used or operable to support other implementations of method 900.

[0164] Figure 10 A flowchart illustrating an exemplary method 1000 implemented at a second device according to some embodiments of the present invention is shown. For ease of discussion, reference will be made to... Figure 1A Method 1000 is described from the perspective of network node 170. It should be understood that method 1000 may include additional actions not shown and / or some actions shown may be omitted, and the scope of the invention is not limited thereto.

[0165] At box 1010, the second device receives a service request message from the first device for a model-based task execution. At box 1020, the second device determines whether the task will be completed within the service delay, provided that the model computation used to execute the task is split at least between the first and second devices. At box 1030, based on the determination that the task will be completed within the service delay, the second device sends a service response message to the first device including model splitting information, wherein the model splitting information indicates the splitting of model computation. At box 1040, based on the determination that the task will not be completed within the service delay, the second device sends a rejection notification message to the first device. The second device may be used or operable to support other implementations of method 1000.

[0166] Figure 11This is a block diagram of a device 1100 that can be used to implement some embodiments of the present invention. In some embodiments, device 1100 may be an element of a communication network infrastructure, such as a base station (e.g., a NodeB, an evolved NodeB (eNodeB or eNB), a next-generation NodeB (sometimes called a gNodeB or gNB)), a home subscriber server (HSS), a packet gateway (PGW), or a serving gateway (SGW), or various other nodes or functions within a core network (CN) or a Public Land Mobility Network (PLMN). In other embodiments, device 1100 may be a device connected to the network infrastructure via a wireless interface, such as a mobile phone, a smartphone, or other such device that can be classified as User Equipment (UE). In some embodiments, device 1100 may be a Machine Type Communications (MTC) device (also known as a machine-to-machine (M2M) device), or another such device that, although not providing direct service to a user, can still be classified as a UE. In some embodiments, device 1100 may be a road side unit (RSU), a vehicle UE (V-UE), a pedestrian UE (P-UE), or an infrastructure UE (I-UE). In some scenarios, device 1100 may also be referred to as a mobile device, regardless of whether the device itself is designed to be mobile or capable of being mobile; this term is intended to refer to a device connected to a mobile network. A particular device may utilize all or only a subset of the components shown, and the level of integration may vary from device to device. Furthermore, device 1100 may contain multiple instances of components, such as multiple processors, memories, transmitters, receivers, etc.

[0167] Device 1100 typically includes a processor 1102 such as a central processing unit (CPU), and may also include a dedicated processor such as a graphics processing unit (GPU) or other such processor, memory 1104, a network interface 1106, and a bus 1108 for connecting components of device 1100. Device 1100 may also optionally include components such as a mass storage device 1110, a video adapter 1112, and an I / O interface 1116 (shown in dashed lines).

[0168] Memory 1104 may include any type of non-transitory system memory readable by processor 1102, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), read-only memory (ROM), or a combination thereof. In one embodiment, memory 1104 may include more than one type of memory, such as ROM used at power-on and DRAM used to store programs and data during program execution. Bus 1108 may be one or more of any type of bus architecture, including a memory bus or memory controller, a peripheral bus, or a video bus.

[0169] Device 1100 may also include one or more network interfaces 1106, which may include at least one of wired network interfaces and wireless network interfaces. For example... Figure 11 As shown, network interface 1106 may include a wired network interface for connecting to network 1122, and may also include a wireless access network interface 1120 for connecting to other devices via a wireless link. When device 1100 is a network infrastructure element, the wireless access network interface 1120 may be omitted for nodes or functions used as PLMN elements rather than elements at the wireless edge. When device 1100 is infrastructure at the wireless edge of the network, both wired and wireless network interfaces may be included. When device 1100 is a wirelessly connected device (e.g., a user equipment), the wireless access network interface 1120 may be present, and other wireless interfaces such as a WiFi network interface may be used as supplements. Network interface 1106 enables device 1100 to communicate with remote entities, such as remote entities connected to network 1122.

[0170] Mass storage 1110 may include any type of non-transitory storage device for storing data, programs, and other information and making such data, programs, and other information accessible via bus 1108. Mass storage 1110 may include one or more of, for example, solid-state drives, hard disk drives, disk drives, or optical disk drives. In some embodiments, mass storage 1110 may be located remotely from device 1100 and may be accessed via a network interface such as interface 1106. In the illustrated embodiment, mass storage 1110 differs from the memory 1104 that includes it and typically performs storage tasks compatible with higher latency, but typically offers lower volatility or no volatility. In some embodiments, mass storage 1110 may be integrated with heterogeneous memory 1104.

[0171] Optional video adapter 1112 and I / O interface 1116 (shown in dashed lines) provide interfaces for coupling device 1100 to external input and output devices. Examples of input and output devices include a display 1114 coupled to video adapter 1112 and an I / O device 1118, such as a touchscreen, coupled to I / O interface 1116. Other devices may be coupled to device 1100, and additional or fewer interfaces may be used. For example, a serial interface such as Universal Serial Bus (USB) (not shown) may be used to provide interfaces for external devices. Those skilled in the art will understand that in embodiments where device 1100 is part of a data center, I / O interface 1116 and video adapter 1112 may be virtualized and provided via network interface 1106.

[0172] Figure 12 This is a schematic diagram of the structure of the device 1200 according to some embodiments of the present invention. For example... Figure 12 As shown, the device 1200 includes a transmitting unit 1202, an executing unit 1204, and a receiving unit 1204. The device 1200 can be applied to applications such as... Figure 1AThe communication system shown can implement any of the methods provided in the foregoing embodiments. Optionally, the physical representation of device 1200 can be a communication device, such as a network device or a UE. Alternatively, device 1200 can be other devices that can implement the functions of a communication device, such as a processor or chip inside a communication device. Specifically, device 1200 can be some programmable chip, such as a field-programmable gate array (FPGA), a complex programmable logic device (CPLD), an application-specific integrated circuit (ASIC), or a system on a chip (SOC).

[0173] In some embodiments, the sending unit 1202 can be used to send a service request message for performing a task based on a model to a second device. The execution unit 1204 can be used to: perform a task based on the model and the model splitting information received from the second device, wherein the model splitting information is used to indicate the splitting of model computation for performing the task at least between the first device and the second device. The pausing unit 1206 can be used to: pause the performance of the task based on the model based on a rejection notification message received from the second device.

[0174] In some other embodiments, the apparatus 1200 may include various other units or modules that can be used to perform the various operations or functions described in conjunction with the foregoing method embodiments. Details can be obtained by referring to the detailed description of the above method embodiments, and will not be repeated herein.

[0175] Figure 13 This is a schematic diagram of the structure of the device 1300 according to some embodiments of the present invention. For example... Figure 13 As shown, the device 1300 includes a receiving unit 1302, a determining unit 1304, a first transmitting unit 1306, and a second transmitting unit 1308. The device 1300 can be applied to applications such as... Figure 1AThe communication system shown can implement any of the methods provided in the foregoing embodiments. Optionally, the physical representation of device 1300 can be a communication device, such as a network device or a UE. Alternatively, device 1300 can be other devices that can implement the functions of a communication device, such as a processor or chip inside a communication device. Specifically, device 1300 can be some programmable chip, such as a field-programmable gate array (FPGA), a complex programmable logic device (CPLD), an application-specific integrated circuit (ASIC), or a system on a chip (SOC).

[0176] In some embodiments, the receiving unit 1302 may be configured to receive a service request message from the first device for performing a task based on a model. The determining unit 1304 may be configured to: determine whether the task will be completed within a time delay, provided that model computation is split at least between the first device and the second device for the purpose of performing the task. The first sending unit 1306 may be configured to: based on the determination that the task will be completed within the time delay, send a service response message including model splitting information to the first device, wherein the model splitting information indicates the splitting of model computation. The second sending unit 1308 may be configured to: based on the determination that the task cannot be completed within the time delay, send a rejection notification message to the first device.

[0177] In some other embodiments, the apparatus 1300 may include various other units or modules that can be used to perform the various operations or functions described in conjunction with the foregoing method embodiments. Details can be obtained by referring to the detailed description of the above method embodiments, and will not be repeated herein.

[0178] It should be noted that the division of units or modules in the foregoing embodiments of the present invention is exemplary and merely a logical functional division. In actual implementation, another division method may also be used. Furthermore, the functional units in the embodiments of the present invention can be integrated into one processing unit, exist physically independently, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0179] When an integrated unit is implemented as a software functional unit and sold or used as an independent product, the integrated unit can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially or in whole or in part, can be implemented in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for instructing a computer device (which may be a personal computer, server, or network device) or processor to execute all or part of the steps of the methods described in the embodiments of the present invention. The aforementioned storage medium includes any medium capable of storing program code, such as a USB flash drive, removable hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0180] Based on the foregoing embodiments, this application also provides a computer program. When the computer program is run on a computer, it causes the computer to perform any of the methods provided in the foregoing embodiments.

[0181] Based on the foregoing embodiments, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a computer, it causes the computer to perform any of the methods provided in the foregoing embodiments. The storage medium can be any available medium that can be accessed by a computer. By way of example and not limitation, a computer-readable medium may include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and can be accessed by a computer.

[0182] Based on the foregoing embodiments, this invention also provides a chip. This chip is used to read a computer program stored in a memory to implement any of the methods provided in the foregoing embodiments.

[0183] Based on the foregoing embodiments, this invention provides a chip system. The chip system includes a processor for supporting a computer device in implementing functions related to the communication device in the foregoing embodiments. In one possible design, the chip system further includes a memory for storing programs and data necessary for the computer device. The chip system may include a chip, or it may include a chip and other discrete components.

[0184] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of purely hardware embodiments, purely software embodiments, or embodiments combining software and hardware aspects. Additionally, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) including computer-usable program code.

[0185] This invention is described with reference to flowchart illustrations and / or block diagrams of the methods, apparatus (systems), and computer program products provided by this invention. It should be understood that computer program instructions can be used to implement each process and / or block in the flowchart illustrations and / or block diagrams, as well as combinations of processes and / or blocks in the flowchart illustrations and / or block diagrams. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to generate a machine such that the instructions, which execute on the processor of the computer or other programmable data processing apparatus, generate means for implementing a specific function in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0186] Computer program instructions may also be stored in a computer-readable storage medium capable of instructing a computer or another programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of writing including instruction means. The instruction means implement a particular function in one or more processes in a flowchart and / or one or more blocks in a block diagram.

[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to perform a series of operations and steps on the computer or other programmable apparatus, thereby generating a computer-implemented process. Therefore, these instructions, which execute on a computer or other programmable apparatus, provide steps for implementing one or more processes in a flowchart and / or one or more boxes in a block diagram.

[0188] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its scope. Therefore, this invention is intended to cover such modifications and variations, provided that they fall within the scope of the claims of this invention and their equivalents.

Claims

1. A method comprising: Send a service request message for the model-based task to the second device from the first device; Based on a service response message received from the second device that includes model splitting information, the task is performed based on the model and the model splitting information, wherein the model splitting information indicates that the model computation for performing the task is split at least between the first device and the second device. as well as Based on the rejection notification message received from the second device, the execution of the task based on the model is suspended.

2. The method according to claim 1, further comprising: Send to the second device at least one of the computing power or computing resources of the first device for performing the task based on the model.

3. The method of claim 2, wherein the service request message includes information indicating at least one of the following: Service latency for performing the task, or The priority of the task.

4. The method according to claim 1, further comprising: Determine the computation time for the first part of the task to be completed by the first device at the split point of the reference model; The remaining latency is determined based on the service latency used to perform the task and the computation time.

5. The method according to claim 4, further comprising: Estimate the computational cost of the model at the split points of multiple candidate models; Based on the computational cost of the model, one candidate model splitting point is selected from the plurality of candidate model splitting points as the reference model splitting point.

6. The method according to claim 4 or 5, wherein the remaining delay comprises: The first device sends a first transmission time to the second device the intermediate result of the first part of the task performed by the first device. The second portion of the calculation time is used to perform the task at the second device. A second transmission time is used for the second device to send the result of performing the task to the first device.

7. The method according to any one of claims 4 to 6, wherein the service request message includes information indicating at least one of the following: The remaining delay, The reference model split point, or The priority of the task.

8. The method according to any one of claims 1 to 7, wherein the rejection notification message includes information indicating a rollback timer, and the method further comprises: After the rollback timer expires, an additional service request message for performing the task is sent to the second device.

9. The method according to any one of claims 1 to 8, further comprising: Send a registration request message to the second device for registering the task; Receive a model notification message from the second device indicating that the model will be used to perform the task.

10. The method of claim 9, wherein the registration request message comprises at least one of the following: Service latency for performing the task, or The accuracy requirements of the task.

11. A method comprising: The second device receives a service request message for the model-based task execution from the first device. Determine whether the task will be completed within the service delay if the model computation used to perform the task is split at least between the first device and the second device; Based on the determination that the task will be completed within the service delay, a service response message including model splitting information is sent to the first device, wherein the model splitting information indicates the splitting of the model computation. Based on the determination that the task will not be completed within the service delay, a rejection notification message is sent to the first device.

12. The method of claim 11, further comprising: Receive from the first device at least one of the computing power or computing resources of the first device for performing the task based on the model.

13. The method of claim 12, wherein the service request message includes information indicating at least one of the following: The service latency used to perform the task, or The priority of the task.

14. The method of claim 12 or 13, wherein determining whether the task will be completed within the service delay comprises: Based on at least one of the computing power or the computing resources, determine the estimated latency for performing the task; The estimated latency is compared with the service latency.

15. The method of claim 14, wherein the estimated delay comprises: Given the computing power of the first device and the model split point, an estimate of the computation time for the first device to perform the first part of the task. Given the computing power of the second device and the model split point, the estimated computation time for the second device to execute the second part of the task is as follows: Given the amount of data associated with the model split point, an estimate of the communication time between the first device and the second device.

16. The method of claim 11, wherein the service request message includes information indicating at least one of the following: The remaining latency is determined based on the service latency used to perform the task and the computation time of the first device at the reference model split point to complete the first part of the task. The reference model split point, or The priority of the task.

17. The method of claim 16, wherein the remaining delay comprises: The first device sends a first transmission time to the second device the intermediate result of the first part of the task performed by the first device. The computation time for performing the second part of the task at the second device. The second device sends a second transmission time to the first device the result of performing the task.

18. The method of claim 17, wherein determining whether the task will be completed within the service delay comprises: The estimated delay is determined by adding the first transmission time, the calculation time, and the estimated value of the second transmission time. as well as The estimated delay is compared with the remaining delay.

19. The method of any one of claims 11 to 18, wherein the rejection notification message includes information indicating a rollback timer, and the method further comprises: After the rollback timer expires, another service request message for performing the task is received from the first device.

20. The method according to any one of claims 11 to 19, further comprising: Receive a registration request message from the first device for registering the task; The target model is determined as the model that will be used to perform the task; A model notification message is sent to the first device, indicating that the model will be used to perform the task.

21. The method of claim 20, wherein determining the target model as the model comprises: Forward the registration request message to a third device; as well as The third device receives an instruction indicating that the target model is the model to be used to perform the task.

22. A first device, comprising: transceiver; The processor is communicatively coupled to the transceiver. The processor is configured as follows: The transceiver sends a service request message for the model-based task execution to the second device from the first device. Based on a service response message received from the second device that includes model splitting information, the task is performed based on the model and the model splitting information, wherein the model splitting information indicates that the model computation for performing the task is split at least between the first device and the second device. as well as Based on the rejection notification message received from the second device, the execution of the task based on the model is suspended.

23. A second device, comprising: transceiver; The processor is communicatively coupled to the transceiver. The processor is configured as follows: The transceiver receives a service request message for a model-based task execution from the first device at the second device. Determine whether the task will be completed within a time delay if the model computation for performing the task is split at least between the first device and the second device; Based on the determination that the task will be completed within the time delay, a service response message including model splitting information is sent to the first device via the transceiver, wherein the model splitting information indicates the splitting of the model computation. Based on the determination that the task will not be completed within the time delay, a rejection notification message is sent to the first device via the transceiver.

24. A non-transient computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 21.

25. An apparatus comprising at least one processing circuit configured to perform the method according to any one of claims 1 to 21.

26. An apparatus for wireless communication, the apparatus comprising: At least one processor; as well as A non-transient computer-readable medium storing instructions that, when executed by the at least one processor, cause the apparatus to perform the method according to any one of claims 1 to 21.

27. A computer program product comprising computer-executable instructions that, when executed, cause a device to perform the method according to any one of claims 1 to 21.