Network unit and terminal device
By implementing the task offload method in network units and terminal devices, the problem of insufficient computing resources of user equipment is solved, the task completion efficiency is improved and the delay is reduced.
Patent Information
- Application Number
- CN202510213205.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively optimize the computing resources on user equipment, resulting in an extended task completion time, especially in multi-task scenarios.
By implementing a task unloading method in the network unit and the terminal device, the terminal device sends an unload request message, and the network unit determines and sends an unloading policy, thereby optimizing task allocation and execution.
It improves the overall efficiency of task completion and reduces the delay in task completion time, and is suitable for multi-user and multi-task scenarios.
Smart Images

Figure CN120201494A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of mobile communication technology, and in particular to a network unit and a terminal device. Background Art
[0002] A variety of tasks can be generated on the user equipment (UE) in the network, but the computing resources available to the UE are limited. Other servers in the network need to provide appropriate offloading strategies for the tasks to reduce the time it takes to complete the tasks. How to optimize and improve the efficiency of task completion is a technical problem that needs to be solved urgently. Summary of the invention
[0003] The present application mainly provides a network unit and a terminal device, and the technical solution of the present application is implemented as follows:
[0004] In a first aspect, an embodiment of the present application provides a network unit, the network unit comprising:
[0005] transceiver; and
[0006] a processor coupled to the transceiver, wherein the processor is configured to:
[0007] Receiving an uninstall request message via a transceiver, where the uninstall request message is a non-access layer message or a computing layer message, and the uninstall request information message includes relevant information of the terminal device and / or task information of the first task;
[0008] Determining an uninstallation strategy for the first task based on the uninstallation request message;
[0009] A task assignment message is transmitted via the transceiver, the task assignment message including an offloading policy.
[0010] In a second aspect, an embodiment of the present application provides a terminal device, the terminal device comprising:
[0011] transceiver; and
[0012] a processor coupled to the transceiver, wherein the processor is configured to:
[0013] Transmitting an offloading request message via a transceiver, the offloading request message being a non-access layer message or a computing layer message, the offloading request information message including information of the terminal device and / or task information of the first task;
[0014] A task allocation message is received via a transceiver, where the task allocation message includes an offloading policy; the offloading policy is used to indicate a first offloading subtask in a first task that is allocated to a terminal device.
[0015] In a third aspect, an embodiment of the present application provides a task offloading method, the method comprising:
[0016] Receive an offloading request message, where the offloading request message is a non-access stratum message or a computing stratum message, and the offloading request message includes information of the terminal device and / or task information of a first task;
[0017] Based on the offloading request message, determine an offloading strategy for the first task;
[0018] Transmit a task allocation message, where the task allocation message includes the offloading strategy.
[0019] In a fourth aspect, an embodiment of the present application provides a task offloading method, and the method includes:
[0020] Transmit an offloading request message, where the offloading request message is a non-access stratum message or a computing stratum message, and the offloading request message includes information of the terminal device and / or task information of a first task;
[0021] Receive a task allocation message, where the task allocation message includes an offloading strategy; the offloading strategy is used to indicate a first offloading subtask allocated to the terminal device in the first task.
[0022] In a fifth aspect, an embodiment of the present application provides a task offloading system, including: at least one access node, at least one network unit as described in the first aspect, and at least one terminal device as described in the second aspect. Description of the Drawings
[0023] Figure 1 Is a topological diagram of a network provided by an embodiment of the present application;
[0024] Figure 2 Is a flowchart of a task offloading method provided by the present application Figure 1 ;
[0025] Figure 3 Is a flowchart of a task offloading method provided by the present application Figure 2 ;
[0026] Figure 4 Is a flowchart of a task offloading method provided by the present application Figure 3 ;
[0027] Figure 5 Is a flowchart of a task offloading method provided by the present application Figure 4 ;
[0028] Figure 6 Is a flowchart of a task offloading method provided by the present application Figure 5 ;
[0029] Figure 7 Is a schematic diagram of a prediction window provided by an embodiment of the present application;
[0030] Figure 8 A schematic diagram of a first task execution process provided by an embodiment of the present application;
[0031] Figure 9 A schematic diagram of a first task offloading process provided by an embodiment of the present application;
[0032] Figure 10 A schematic diagram of a protocol stack provided by an embodiment of the present application Figure 1 ;
[0033] Figure 11 A schematic diagram of a protocol stack provided by an embodiment of the present application Figure 2 ;
[0034] Figure 12 A schematic diagram of the format of an offloading request message provided by an embodiment of the present application;
[0035] Figure 13 A schematic diagram of the process of a task offloading method provided by the present application Figure 6 ;
[0036] Figure 14 A schematic diagram of the process of a task offloading method provided by the present application Figure 7 ;
[0037] Figure 15 A schematic diagram of the process of a task offloading method provided by the present application Figure 8 ;
[0038] Figure 16 A schematic diagram of the process of a task offloading method provided by the present application Figure 9 ;
[0039] Figure 17 A schematic diagram of the process of a task offloading method provided by the present application Figure 10 ;
[0040] Figure 18 A schematic diagram of reserved resources provided by the present application;
[0041] Figure 19 A schematic diagram of the process of a task offloading method provided by the present application Figure 10 One;
[0042] Figure 20 A schematic diagram of the process of a task offloading method provided by the present application Figure 10 Two;
[0043] Figure 21 A schematic diagram of a scheduling linked list provided by the present application;
[0044] Figure 22 A schematic diagram of the process of a task offloading method provided by the present applicationFigure 10 III;
[0045] Figure 23 A schematic diagram for comparing the delay violation rate provided for this application Figure 1 ;
[0046] Figure 24 A schematic diagram for comparing the delay violation rate provided for this application Figure 2 ;
[0047] Figure 25 A schematic diagram for comparing the delay violation rate provided for this application Figure 3 ;
[0048] Figure 26 A schematic diagram for comparing the delay violation rate provided for this application Figure 4 ;
[0049] Figure 27 An optional structural schematic diagram of the device provided in the embodiment of this application. Detailed implementation manners
[0050] In order to more comprehensively understand the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and explanation purposes and are not used to limit the embodiments of this application.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0052] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0053] It should also be noted that the terms "first / second / third" related to the embodiments of this application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.
[0054] Communication and Computing Convergence for the 6th Generation Mobile Networks (6G) refers to the deep integration of communication and computing capabilities, aiming to provide more efficient and flexible services by optimizing the communication network architecture and computing resource allocation. 6G is not just a simple communication system. It will natively support communication, sensing, and computing services, becoming the network information foundation to support the efficient and sustainable development of future society.
[0055] Artificial Intelligence (AI) applications require strict reasoning capabilities, including real-time response, accuracy, and efficiency. Considering the limited computing resources available on the UE, it is necessary to utilize the Mobile Edge Computing (MEC) function, use servers in the network for collaborative reasoning, and provide appropriate offloading strategies for the first tasks to reduce the possibility of violating latency when executing these tasks. As a typical use case of Communication and Computing Convergence, the Computation Offloading service in the 6G network refers to transferring some or all of the tasks that originally needed to be processed locally on user devices (such as smartphones, tablets, augmented reality glasses, etc.) to remote servers or edge nodes for processing through a high-speed and low-latency communication network. Then, the processing results are quickly sent back to the user device. This method can effectively reduce the computing burden on the terminal device, extend the battery life, and improve the user experience.
[0056] Remote servers or edge nodes are equipped with powerful computing units, such as Graphics Processing Units (GPUs) or hardware accelerators for custom-designed Deep Neural Networks (DNNs), such as Neural Processing Units (NPUs), which can utilize parallel capabilities to concurrently process multiple tasks. When processing the same type of first tasks, multiple first tasks can share the same model parameter cache, thus reducing the total latency when performing batch reasoning on these tasks and significantly improving the utilization rate of the computing resources of the reasoning server.
[0057] The following uses a specific example to illustrate how to use the 6G network to provide a computation offloading service for users:
[0058] Consider a user wearing a pair of lightweight Augmented Reality (AR) glasses to watch a virtual concert. This scenario requires extremely high graphic quality and also needs to interact with the live audience, which means a large amount of real-time data processing capabilities are required.
[0059] The application in the AR glasses recognizes that the currently running virtual concert experience requires complex image rendering. The application determines that the local resources are insufficient to efficiently complete these rendering tasks and decides to offload some or all of the rendering work to a remote computing node or a cloud server.
[0060] With the support of the 6G network, the AR glasses can quickly search for and select the most suitable edge computing node or cloud server for image rendering.
[0061] The user's motion capture data (such as head rotation, gestures, etc.), environmental information (such as surrounding light conditions), and the required 3D models and other resources are compressed and efficiently uploaded to the selected remote computing node via the 6G network.
[0062] After receiving the data from the AR glasses, the remote computing node starts to execute the image rendering task using its powerful GPU resources. The rendering process may involve a series of complex computational operations such as lighting simulation, shadow generation, and material mapping.
[0063] The completed rendered frames are sent back to the user's AR glasses at high speed via the 6G network in the form of an encoded video stream.
[0064] The AR glasses decode the received video stream and overlay it onto the user's actual field of view, providing an immersive visual experience.
[0065] Figure 1 A topology diagram of a network provided for the embodiments of the present application. As Figure 1 shown, the present application mainly considers the following scenarios:
[0066] There are multiple UEs in the 6G network 10, each with different computing tasks, such as image rendering tasks, object recognition tasks, etc. Exemplarily, user device 1 1031 has task 1 1032, user device 2 1041 has task 2 1042, user device 3 1051 has task 3 1043, etc.
[0067] Among them, each computing task may be completed by multiple layers of Artificial Intelligence Markup Language (AIML), such as DNN deep neural network. There are different requirements for the size of the AIML model, the number of parameters required during the calculation process, and the calculation latency.
[0068] Exemplarily, a lightweight model for image classification such as MobileNetV2 is approximately 14 MB and requires less than 100 milliseconds for real-time processing. A lightweight model for speech recognition such as DeepSpeech2 is approximately 65 MB, and the latency generally needs to be controlled within 200 milliseconds to maintain the feeling of natural conversation.
[0069] As Figure 1 shown, the server 1021 can provide computing services. For example, a partial subtask of task 1 1032 of user device 1 1031 is offloaded to the server 1021 through a wireless access point (AP) 101; a partial subtask of task 2 1042 of user device 2 1041 is offloaded to the server 1021 through AP 101; or, a partial subtask of task 3 1043 of user device 3 1051 is offloaded to the server 1021 through AP 101. The server 1021 calculates task 1022 based on the tasks offloaded by each user device.
[0070] To minimize the completion time of tasks on each user device, it is crucial to carefully design the offloading strategy for each task. Based on this, the embodiments of the present application provide a network unit, a terminal device, and a task offloading method. First, the offloading request message sent by the terminal device includes information related to the terminal device and / or task information of the first task, enabling the network unit to quickly determine the terminal device that needs offloading services and the relevant task information. Second, the network unit determines the offloading strategy for the first task according to the content of the offloading request message, and further sends a task allocation message including the offloading strategy to the terminal device. Through this message interaction, the overall efficiency of task completion is improved, and the latency of task completion time is reduced.
[0071] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0072] In a first aspect, in an embodiment of the present application, Figure 2 is a schematic flowchart of a task offloading method provided by the present application Figure 1 . The task offloading method can be applied to a network unit. As Figure 2 shown, the method may include:
[0073] S201, receiving an offloading request message. The offloading request message is a non-access stratum message or a computing stratum message, and the offloading request information message includes information related to the terminal device and / or task information of the first task.
[0074] S202, determining an offloading strategy for the first task based on the offloading request message.
[0075] S203, send a task assignment message, where the task assignment message includes an offloading policy.
[0076] In a second aspect, an embodiment of the present application provides a task offloading method. Figure 3 The following is a flowchart illustration of a task offloading method provided by the present application. Figure 2 This task offloading method can be applied to a terminal device. As Figure 3 shown, this method may include:
[0077] S301, send an offloading request message. The offloading request message is a non-access stratum message or a computing stratum message, and the offloading request message includes information of the terminal device and / or task information of a first task.
[0078] S302, receive a task assignment message. The task assignment message includes an offloading policy; the offloading policy is used to indicate a first offloading subtask assigned to the terminal device in the first task.
[0079] In a third aspect, an embodiment of the present application provides a task offloading method applied to a task offloading system including a network unit and a terminal device. As Figure 4 shown, this method may include:
[0080] S401, the terminal device sends an offloading request message to the network unit. The offloading request message is a non-access stratum message or a computing stratum message, and the offloading request message includes information of the terminal device and / or task information of a first task.
[0081] S402, the network unit determines an offloading policy for the first task based on the offloading request message.
[0082] S403, the network unit sends a task assignment message to the terminal device. The task assignment message includes an offloading policy; the offloading policy is used to indicate a first offloading subtask assigned to the terminal device in the first task.
[0083] Next, a description is given of the task offloading methods shown in Figure 2 , Figure 3 and Figure 4 .
[0084] In the embodiments of the present application, the terminal device may be a communication terminal device. Exemplarily, it may be a UE. The network unit may include a network computing node with computing network function (Computing NF), such as a multi-access / mobile edge computing (MEC) device, a network data analytics function (NWDAF), or a dedicated network function node for a computing service.
[0085] Wherein, before the terminal device sends an offloading request message to the network unit, the terminal device establishes a connection with the network unit. This connection may be a connection of the control plane, a connection of the user plane, or a connection of the computing plane.
[0086] In the embodiments of the present application, the terminal device may generate different first tasks under different application conditions. Exemplarily, the first task may be a task generated by a third-party application, such as an image rendering task, an object recognition task, a video processing task, or an image classification task, etc. Or, the first task may be a task generated by a communication-related application, such as UE traffic prediction, UE trajectory prediction, etc.
[0087] The terminal device encapsulates the task information of the first task, and / or the relevant information of the terminal device itself, in the format of a non-access stratum message or a computing layer message to generate an offloading request message, and sends it to the network unit. The computing layer is a new protocol layer for interacting with computing task-related information.
[0088] It should be noted that each first task may be split into multiple subtasks. These subtasks may be executed on the terminal device or on the network unit, which is characterized by the offloading strategy generated by the network unit. Exemplarily, if the first task is completed by a multi-layer AI model, the subtasks may be one or more layers of the model.
[0089] Exemplarily, when the first task is a task of a lightweight model for image classification, such as MobileNetV2, it may be split into the following subtasks:
[0090] (1) Front-end (shallow) processing: Usually includes preprocessing of the input image and the first few convolutional operations. These layers are mainly responsible for extracting basic features such as edges and textures.
[0091] (2) Middle-layer processing: Continues with more complex feature extraction, such as shapes and object parts.
[0092] (3) Back-end (deep) processing: The last few layers focus on learning high-level features and finally making classification decisions.
[0093] Exemplarily, in the task of a lightweight model where the first task is about speech recognition such as DeepSpeech2, it can be split into the following subtasks:
[0094] (1) Front-end processing: Includes audio preprocessing (such as sampling rate conversion, normalization) and feature extraction (such as Mel spectrogram or Mel-Frequency Cepstral Coefficients (MFCCs)). This part can be executed on the terminal device to reduce the amount of data transmitted.
[0095] (2) Middle processing: The feature vectors corresponding to each audio segment are sent to the edge server or the cloud for further processing.
[0096] (3) Back-end processing: The final decoding step, such as Connectionist Temporal Classification (CTC) decoding, can be completed in the cloud and the result is returned to the terminal device.
[0097] It should be noted that the splitting of the first task can be performed on the terminal device side or on the network unit side, and is determined according to the actual processing capabilities of the terminal device and the network unit.
[0098] Further, after the aforementioned terminal device sends an offloading request message to the network unit in the 6G network as needed, after receiving the offloading request message, the network unit evaluates the requirements of the first task and the computing resources of each device according to the first task information, the relevant information of the terminal device, etc., and generates an offloading strategy for the first task. For example, the offloading strategy may include which subtasks in the first task will be calculated on the terminal device and which subtasks will be offloaded to the network unit.
[0099] Finally, the network unit carries the generated offloading strategy for the first task in the task assignment message and sends it to the terminal device, so that the terminal device knows the first offloading subtasks in the first task that it needs to execute.
[0100] It should be noted that one current solution is as follows: after a new first task is generated by the terminal device, the terminal device decides whether to process all subtasks locally and then queue the first task into the local computing task queue for processing; or unload the entire first task to the server and then send the input data of the first task to the uplink transmission queue for waiting to be transmitted; or randomly select the number of subtasks to be processed locally, add the local computing subtasks to the local computing task queue for processing, and at the same time send the input data of the next subtask to the uplink transmission queue for waiting to be transmitted. Compared with this solution, in the embodiments of the present application, when the terminal device generates a new first task, it sends the task information of the first task, the relevant information of the terminal device, etc. to the network element in the unloading request message through the uplink channel. The network element makes a decision to determine the unloading strategy of the first task and passes this decision back to the mobile device through the downlink channel. In this way, in the case where multiple users all generate the first task, the network element uniformly coordinates different first tasks as much as possible to determine their respective corresponding unloading strategies, which can ensure that all first tasks are completed within a relatively short time range and reduce latency.
[0101] In the embodiments of the present application, first, the unloading request message sent by the terminal device includes the information related to the terminal device and / or the task information of the first task, enabling the network element to quickly determine the terminal device that needs unloading service and the relevant task information. Second, the network element determines the unloading strategy of the first task according to the content of the unloading request message, and further sends the task allocation message including the unloading strategy to the terminal device. Through this message interaction, the overall efficiency of task completion is improved and the delay of task completion time is reduced.
[0102] In another embodiment of the present application, a task unloading method is provided. This method can be applied to a task unloading system including a Radio Access Network (RAN) node, which can also be called an access network device, a terminal device, and a network element. As Figure 5 shown, this task unloading method may further include:
[0103] S404, the network element sends a radio resource request message (Nran_Event_Exposure_Subscribe Request) to the access network device. The radio resource request message is used to request a radio resource response message and indicate the content included in the radio resource response message.
[0104] S405, the network element receives the radio resource response message sent by the access network device.
[0105] Among them, the radio resource response message includes at least one of the current radio resource information of the terminal device, the predicted radio resource information of the terminal device, the current radio resource information of the serving cell of the terminal device, and the predicted radio resource information of the serving cell of the terminal device.
[0106] S406. The network unit determines the offloading strategy for the first task based on the offloading request message and the radio resource response message.
[0107] In the embodiments of the present application, based on Figure 5 the task offloading method shown, the interaction process between the access network device and the network unit is as Figure 6 shown, where:
[0108] S4051. The access network device sends a first radio resource response (Nran_Event_Exposure_Subscribe Response) message to the network unit.
[0109] S4052. The access network device sends a second radio resource response (Nran_Event_Exposure_Notify) message to the network unit.
[0110] It should be noted that the network unit will call the services provided by the access network device, such as the RAN event exposure service (RANevent exposure service). The access network device will, based on the radio resource request message sent by the network unit, send the current or predicted radio resource information of one or more terminal devices to the network unit once or periodically. The radio resource information of the terminal device may include, for example, uplink / downlink packet delay (Uplink / Downlink PacketDelay) information; or, send the current or predicted radio resource information of the cells where one or more terminal devices are located to the network unit; the radio resource information of the cell may include, for example, the radio resource availability / occupation of the cell at present or in a future period of time. Among them, the radio resource availability / occupation of the cell may be represented by an integer value, for example, any value between 0 and 100, and 100 represents that the radio resource is completely idle / occupied. Or, the radio resource availability / occupation may be represented by the occupancy (0-100) of each synchronization signal block (SynchronizationSignal Block, SSB) or each physical resource block (Physical Resource Block, PRB). Or, the radio resource information of the cell may further include, for example, the average uplink / downlink packet delay of all terminal devices in the cell.
[0111] It should be noted that both the air interface resource response message and the air interface resource request message can be messages in the format of the Hyper Text Transfer Protocol Secure (HTTPs).
[0112] In the embodiments of the present application, the above air interface resource response message may include a first air interface resource response message and a second air interface resource response message. Among them, the first air interface resource response message may indicate that the access network node accepts the request of the network unit, and further, based on the indication of the network unit, sends the second air interface resource response message to the network unit once or periodically. Among them, the second air interface resource response message may include the current or predicted air interface resource information of the foregoing terminal device, such as the current or predicted air interface resource information of the service cell of the terminal device.
[0113] Alternatively, in some embodiments, if the access network device determines that the air interface resource response message is a one-time message, the first air interface resource response message and the second air interface resource response message may be combined into one air interface resource response message, and the current or predicted air interface resource information of the foregoing terminal device and the cell may be carried in this message and sent to the network unit.
[0114] In the embodiments of the present application, the network unit also comprehensively evaluates the channel carrying capacity of each terminal device based on the task information of the first task in the comprehensive offloading request message and the air interface resource response message, that is, the number of bits that each terminal device can upload per time slot per unit frequency of the single model, or calculates the average channel carrying capacity of each terminal device by using the dynamic moving average method, so as to determine the efficiency of the first task uploaded by the terminal device, and further determine the offloading strategy of the first task.
[0115] In some embodiments, the air interface resource request message includes at least one of the following:
[0116] The identifier of at least one terminal device;
[0117] The identifier of the cell where at least one terminal device is located;
[0118] Radio availability prediction indicator;
[0119] Periodic information;
[0120] Prediction window.
[0121] Among them, the radio availability prediction indicator may include: an indicator for requesting the prediction of the availability / occupancy of radio resources.
[0122] Among them, the periodic information may indicate whether the access network device should send the air interface resource response message periodically or once.
[0123] Among them, the prediction window may refer to the time length corresponding to whether the predicted radio resources in the air interface resource response message are available.
[0124] In some embodiments, when the periodic information included in the air interface resource request message is sent periodically, the air interface resource response message includes at least one of the following:
[0125] The air interface resource information corresponding to each of at least one terminal device within the prediction window;
[0126] The air interface resource information corresponding to each of at least one terminal device within the time before the time corresponding to the next air interface resource response message;
[0127] The average air interface resource information corresponding to each of at least one cell corresponding to at least one terminal device within the prediction window;
[0128] The average air interface resource information corresponding to each of at least one cell corresponding to at least one terminal device within the time before the time corresponding to the next air interface resource response message.
[0129] Based on the content of the above air interface resource request message, the access network device sends an air interface resource response message to the network unit, and the air interface resource response message may include the above content.
[0130] Among them, corresponding to the air interface resource request message, the air interface resource response message includes the air interface resource information corresponding to each of one or more terminal devices indicated by the air interface resource request message. Exemplarily, the air interface resource information of the terminal device may include the uplink / downlink packet delay information of the terminal device.
[0131] In some embodiments, the average air interface resource information of each cell where at least one terminal device indicated in the air interface resource request message is located is included in the air interface resource response message. It should be understood that if only the identifiers of at least one terminal device are included in the air interface resource request message, the access network device may determine the air interface resource information corresponding to these terminal devices; or, in the case where the air interface resource includes the identifier of the terminal device and the identifier of the cell, the access network device may respectively determine the current or predicted air interface resource information of the terminal device, and the current or predicted air interface resource information of the cell.
[0132] It should be noted that if the prediction window is represented in the form of a duration or time interval, it means that the predicted air interface resources provided by each air interface resource response message regarding the terminal device or cell are within a time range starting from the time point of this air interface resource response message to the future, the air interface resource information corresponding to the terminal device, or the average air interface resource information of multiple terminal devices in the cell.
[0133] It should also be noted that if the prediction window is represented in a way that includes a start time value and an end time value. Then in the scenario where the radio resource response message is sent periodically, each predicted prediction window is related to the time point of the periodically triggered radio resource message. As Figure 7 shown, in the case where the radio resource request message indicates that the access network device should send the radio resource response message periodically, the time interval between the first adjacent radio resource response message and the second radio resource response message is the prediction window.
[0134] It should also be noted that in the case where the radio resource request message indicates that the access network device should send the radio resource response message periodically, if the content of the prediction window field in the radio resource request message is empty, it means that the prediction information carried in each radio resource response message is about the radio resource information of the terminal device or the average radio resource information of the cell during the time period from the current time point to the next radio resource response message.
[0135] In this way, the access network device sends a radio resource response message to the network unit based on the radio resource request message of the network unit, so that the network unit can master the current and predicted radio resource conditions of the terminal device, and / or, the current and predicted radio resource conditions of the cell, which is convenient for accurately determining the offloading strategy and improving the efficiency of task execution.
[0136] In some embodiments, based on the offloading strategy, the first offloading subtask assigned to the terminal device in the first task, and / or, the second offloading subtask assigned to the network unit in the first task are determined.
[0137] In the embodiments of the present application, the first task may be composed of a series of multiple subtasks. Exemplarily, for the input data, subtask 1, subtask 2, subtask 3, subtask 4... subtask N are sequentially executed until subtask N outputs the final calculation result. The network unit determines the second offloading subtask to be executed by itself and the first offloading subtask that needs to be executed by the terminal device based on the offloading strategy, and informs the corresponding terminal device of the first offloading subtask through the downlink.
[0138] Figure 8 It is a schematic diagram of a first task execution process provided by the embodiments of the present application. As Figure 8As shown, taking the mobilenet-v2 model as an example, the complete model can be split into 8 parts, corresponding to 8 subtasks respectively. Among them, in one offloading strategy, the first offloading subtask includes subtasks 1 to 3, and the second offloading subtask includes subtasks 4 to 8. That is to say, in step S501, input data; in step S502, execute subtask 1; in step S503, execute subtask 2; in step S503, execute subtask 3, all of which are calculated on the terminal device side, and the generated intermediate variable matrix [28, 28, 32] will be uplinked from the terminal device to the network unit through the air interface as the input of subtask 4. In step S505, execute subtask 4; in step S506, execute subtask 5; in step S507, execute subtask 6; in step S508, execute subtask 7; in step S509, execute subtask 8, all of which are calculated on the network unit side, and the finally generated calculation result will be downlinked to the terminal device. Among them, for step S506, it may include sub-steps: S5061, convolution (convolution layer or Conv) 1×1; S5062, convolution 3×3; S5063, convolution 1×1; S5064, addition. It should be noted that subtask 1 may be Conv + bottleneck module (bottleneckmodule or B) 1, subtasks 2 to 7 respectively correspond to B2 - B7; subtask 8 may be a classification layer (classification layer or CLS).
[0139] In the embodiments of the present application, there may be multiple terminal devices. The first tasks generated by different terminal devices may be the same, but the corresponding offloading strategies are different. After the network unit determines the offloading strategies corresponding to different terminal devices, it sends them to the corresponding terminal devices. The terminal device has the ability to independently execute the initial subtasks of the first task, and offload the subsequent subtasks to the network unit. The network unit is equipped with a GPU and can integrate similar subtasks into batches for simultaneous processing. This batch processing method significantly reduces the total time consumed by the inference of multiple first tasks. Figure 9 It is a schematic diagram of the offloading process of the first task provided by the embodiments of the present application. As Figure 9As shown, the terminal device 1 601 locally processes the corresponding first offloading subtasks, including subtask 1 and subtask 2, and then uploads the output data of subtask 2 to the network unit 603. At the same time, the terminal device 2 602 locally processes the corresponding first offloading subtasks, including subtask 1, subtask 2, and subtask 3, and uploads the output data of subtask 3 to the network unit 603. The network unit 603 determines the second offloading subtasks corresponding to the terminal device 1 601, including subtask 3 and subtask 4, and additionally determines the second offloading subtask 4 corresponding to the terminal device 2 602, and processes them in sequence. During this process, the network unit 603 combines the subtask 4 from the terminal device 1 601 and the subtask 4 from the terminal device 2 602 into a batch for concurrent processing. Finally, the network unit sends the calculation result back to the corresponding terminal device.
[0140] In this way, the network unit determines the first offloading subtasks that need to be executed by the terminal device and the second offloading subtasks that need to be executed by the network unit in the first task based on the offloading policy, and issues them to the terminal device, improving the efficiency of the terminal device and the network unit to cooperate to complete the first task.
[0141] In another embodiment of the present application, the task information of the first task includes one or more of the following:
[0142] The identifier of the first task, task description, model parameters, and quality requirements;
[0143] The identifier, model parameters, and quality requirements of the subtasks of the first task;
[0144] The information of the terminal device includes one or more of the following:
[0145] The identifier of the terminal device;
[0146] The cell identifier of the cell where the terminal device is located;
[0147] The relevant information of the first computing resource of the terminal device.
[0148] In the embodiment of the present application, the task information included in the offloading request message may include information related to the first task. For example, it may include: the identifier (task ID / Index) of the first task; the task description of the first task, such as the application description or the complete AI model parameters or files used by the first task; the quality requirements of the first task, such as the Quality of Service (QoS) requirements, which may include millisecond-level or second-level latency requirements; or include the overall computing power requirements, such as the number of computing units, the number of bits or the number of floating-point operations (FLOPS), and the abstract value of the computing power level from 0 to 100.
[0149] In the embodiments of the present application, the terminal device side can split the first task into multiple subtasks. In this case, the task information can further include the subtask information of the first task. For example, it can include: the identifier (subtask ID / Index) of the subtask in the first task; the task description of the computing subtask, such as the type of the model layer or the complete AI model parameters or files used by the subtask. The quality requirements of the computing subtask can include latency requirements and overall computing power requirements. For specific details, reference can be made to the relevant description of the quality requirements of the foregoing first task.
[0150] In the embodiments of the present application, the information of the terminal device can include: the identifier of the terminal device (UE ID), which can be, for example, the UE network temporary identifier (Radio Network Temporary Identifier, RNTI) or the UE user permanent identifier (Subscription Permanent Identifier, SUPI). It should be noted that the identifier of the terminal device is unique, and the UE, network unit, and access network device have a consistent understanding of the identifier of the terminal device.
[0151] In the embodiments of the present application, the information of the terminal device can further include the identifier of the cell where the terminal device is located (CellID), such as the physical cell identifier (Physical Cell Identifier, PCI). It should be noted that the identifier of the cell is unique, and the UE, network unit, and access network device have a consistent understanding of the identifier of the cell.
[0152] In the embodiments of the present application, the information of the terminal device can further include the relevant information of the first computing resource of the terminal device. For example, it can be the computing power that the terminal device can provide, including the number of processing units, the number of floating-point operations, and an abstract value from 1 to 100 used to describe the computing power level.
[0153] It should be understood that the offloading request message can further include the relevant information of other terminal devices, which is not specifically limited herein.
[0154] In some embodiments, after receiving the offloading request message in the foregoing step S201, the task offloading method can further include:
[0155] S701, when the information of the terminal device includes the identifier of the terminal device, the network unit sends a cell identifier request message to the access network device. The cell identifier request message includes the identifier of the terminal device.
[0156] In the embodiments of the present application, the offloading request message sent by the terminal device only includes the identifier of the terminal device, such as UE RNTI, UE SUPI, etc. In this case, the network element may send a cell identifier request message to a node with an Access and Mobility Management Function (AMF) in the network. The cell identifier request message includes the identifier of the terminal device and is used to request the relevant information of the terminal device from the AMF.
[0157] S702, the network element receives a cell identifier response message, and the cell identifier response message includes the identifier of the cell where the terminal device is located.
[0158] It should be noted that this task offloading method can be applied to a task offloading system including a terminal device, a network element, and an AMF node.
[0159] In response to the cell identifier request message sent by the network element, the AMF sends a cell identifier response message to the network element. The cell identifier response message may include the relevant information of the terminal device, such as the identifier of the cell where the terminal device is located, the relevant information of the first computing resource corresponding to the terminal device, etc.
[0160] In this way, based on the offloading request message, the network element obtains the information of the terminal device and / or the information of the first task, and then determines the offloading strategy, improving the accuracy and reliability of the offloading strategy.
[0161] In another embodiment of the present application, the offloading request message is a computing layer message in the non-access stratum message.
[0162] As described above, the format of the offloading request message may be a NAS message or a Computing Layer message. In the embodiments of the present application, the format of the offloading request message may be a Computing Layer message included in the NAS message. It should be understood that the offloading request message may also be in other formats, which are not specifically limited herein.
[0163] In some embodiments, the relevant information of the terminal device and / or the task information of the first task included in the computing offloading request are encapsulated in the offloading task-related information field in the non-access stratum message or the computing layer message.
[0164] Among them, when the offloading request message is transmitted through the user plane, the non-access stratum message or the computing layer message is encapsulated based on the protocol or format of the service data adaption protocol (SDAP) layer.
[0165] When the offloading request message is transmitted through the user plane (UP), the offloading request message sent by the terminal device is forwarded to the network element through the UPF. As Figure 10 shown, in this case, the protocol stack of the offloading request message from bottom to top includes: the model physical layer (Port Physical Layer, PHY) 801, the media access control layer (media access control layer, MAC) 802, the radio link control sublayer (Radio Link Controlstructure, RLC), the packet data convergence protocol (Packet Data Convergence Protocol, PDCP) layer 803, the service data adaptation protocol layer 804, and the non-access layer / computing layer (NAS / Computing Layer) 806, which are encapsulated from bottom to top on the terminal device side. That is to say, the content of the NAS / Computing Layer layer is encapsulated in sequence based on the above protocol stack, and an offloading request message is generated after the encapsulation is completed. Among them, the content of the NAS / Computing Layer layer can be encapsulated in the format based on the NAS protocol, or in the format based on the Computing Layer layer message, or the Computing Layer message can be encapsulated in the NAS protocol, and the Computing Layer message carries the data content of the terminal device.
[0166] In some embodiments, the relevant information of the terminal device and / or the task information of the first task included in the computing offloading request are encapsulated in the offloading task-related information field in the non-access layer message or the computing layer message.
[0167] Among them, when the offloading request message is transmitted through the control plane (Control Plane, CP), the non-access layer message or the computing layer message is encapsulated based on the protocol or format of the radio resource control (Radio Resource Control, RRC) layer.
[0168] When the offloading request message is transmitted through the control plane, the offloading request message sent by the terminal device is forwarded to the network element through the access network device. As Figure 11 shown, in this case, the protocol stack of the offloading request message from bottom to top includes: the model physical layer 801, the media access control layer 802, the radio link control sublayer 803, the packet data convergence protocol 804, the radio resource control layer 807, and the non-access layer / computing layer 806, which are encapsulated from bottom to top on the terminal device side.
[0169] It should be noted that, as Figure 12The NAS / Computing Layer message example shown below. The first field occupies octet 1 and is the Extended protocol discriminator field 901; the second field occupies octet 2 and is the Security header type associated with a spare half octet / PDU session identity field 902 of the Protocol Data Unit (PDU); the third field occupies octet 2a * and is the Procedure transaction identity field 903; the fourth field occupies octet 4 and is the Message type field 904; the fifth field occupies octet n, where n is determined according to the actual data length, and is the Computing Task Offloading Request Related Info field 905. This field is used to store information of the terminal device in the offloading request message, and / or, information such as task information of the first task and other related information.
[0170] Thus, in the embodiment of the present application, the relevant information of the terminal device, and / or, the task information of the first task are encapsulated in the Computing Task Offloading Request Related Info field of the NAS / Computing Layer message, and based on whether the offloading request message is forwarded by the control plane or the data plane, the offloading request message is encapsulated in the format of the corresponding protocol, thereby improving the forwarding efficiency of the offloading request message, and further reducing the processing time of the first task.
[0171] In another embodiment of the present application, based on Figure 4 , a task offloading method as shown in Figure 13 is provided. This method can be applied to a task offloading system including a terminal device, an access network device, and a network unit. The task offloading method may further include:
[0172] S407, the network unit receives a data volume related prediction request message sent by the access network device, and the data volume related prediction request message is used to request a data volume related prediction message.
[0173] S408, the network unit transmits a data volume related prediction message to the access network device, and the data volume related prediction message includes the upload time of the terminal calculation result determined based on the first offloading subtask, the uplink and downlink data volumes, and the deadline of the first offloading subtask.
[0174] In an embodiment of the present application, the network unit encapsulates the future uplink and downlink data volumes (expressed in bytes or bits) related to the first task, or the upload data volume, the arrival times of the uplink and downlink data, including the upload time of the terminal calculation result, and may further include the calculation result of the first task calculated by the network unit based on the terminal calculation result, and the time of transmission through the downlink, and the deadline of the first task in a data volume-related prediction message and sends it to the access network device.
[0175] It should be noted that the data volume-related prediction message can be a response to the first resource prediction request (Ncomp_Event_Exposure_Subscribe Request) message of the access network device, or it can be sent to the network unit actively by the above-mentioned message corresponding to one or more terminal devices after the offloading policy is generated without relying on the subscription of the access network device. After determining the offloading policy, the network unit notifies the access network device to trigger the access network device to send a data volume-related prediction request message.
[0176] It should also be noted that the arrival time of the uplink and downlink data can be an absolute timestamp (Timestamp), or a relative time (Time Offset) based on the time when the message is sent, which is not specifically limited here.
[0177] In an embodiment of the present application, based on Figure 13 the task offloading method shown, the interaction process between the access network device and the network unit is as Figure 14 shown, where:
[0178] S4081, the network unit sends a first resource prediction (Ncomp_Event_Exposure_Subscribe Response) sub-message to the access network device.
[0179] S4082, the network unit sends a second resource prediction (Ncomp_Event_Exposure_Subscribe_Notify) sub-message to the access network device.
[0180] Among them, the first resource prediction sub-message indicates that the network unit accepts the data volume-related prediction request message.
[0181] In the embodiments of the present application, the network unit will, according to the requirements in the data volume-related prediction request message, provide relevant prediction information such as the uplink and downlink data volumes and arrival times of one or more terminal devices through the second resource prediction sub-message. For example, when the network unit generates an offloading policy for multiple terminal devices, it can predict that a terminal device will complete a first offloading subtask at a future moment and upload the terminal computing structure of the first offloading subtask, or predict that the network unit will process and complete a second offloading subtask at a future moment and send down the calculation result of the first task.
[0182] The second resource prediction sub-message sent by the network unit to the corresponding terminal device may include at least one of the following:
[0183] The burst data volume of the terminal device's uplink transmission, that is, the size of the calculation terminal result, in bits or bytes;
[0184] The arrival time of the calculation terminal result of the terminal device's uplink transmission, represented by a timestamp or relative time;
[0185] The burst data volume of the network unit's downlink transmission, that is, the size of the calculation result of the first task, in bits or bytes;
[0186] The arrival time of the calculation result of the first task of the network unit's downlink transmission, represented by a timestamp or relative time.
[0187] It should also be noted that, referring to the foregoing embodiments, the above first resource prediction sub-message and second resource prediction sub-message can be combined into one data volume-related prediction message for transmission.
[0188] In some embodiments, the data volume-related prediction request message includes at least one of the following:
[0189] The identifier of at least one terminal device;
[0190] The identifier of the cell where at least one terminal device is located;
[0191] An indicator for requesting uplink and downlink data volumes;
[0192] An indicator for requesting the upload time.
[0193] Among them, the identifier of the terminal device (UE ID) indicates that the access network device requests the prediction data corresponding to the corresponding terminal device, for example, including the upload time of the foregoing terminal calculation result, the uplink and downlink data volumes, and the deadline of the first offloading subtask, etc.
[0194] Among them, if the data volume-related prediction request message includes the identifier (Cell ID) of the cell where the terminal device is located and the identifier of the terminal device in this cell, it indicates that the access network device requests the prediction data corresponding to these terminal devices; if the first resource request message only includes the identifier of the cell and does not include the identifier of the terminal device, it indicates that the access network device subscribes to the prediction information corresponding to all current terminal devices with the first task under these cells.
[0195] Among them, the indicator for requesting the uplink and downlink data volume can be used to request the data volume of the uplink data transmitted by the terminal device and the data volume of the downlink data transmitted by the network unit.
[0196] Among them, the indicator for requesting the upload time can be used to request the arrival time of the uplink data transmitted by the terminal device and the arrival time of the downlink data transmitted by the network unit.
[0197] In the embodiments of the present application, the access network device can reserve radio resources for the network unit and the terminal device according to the data volume-related prediction message transmitted by the network unit.
[0198] In some embodiments, based on Figure 4 , a task offloading method as Figure 15 shown is provided, and this method can be applied to a task offloading system including a terminal device, an access network device, and a network unit.
[0199] S409. The terminal device receives the uplink data-related prediction request message transmitted by the access network device.
[0200] S410. The terminal device transmits an uplink data-related prediction message to the access network device, and the uplink data-related prediction message includes the upload time, upload data volume, and deadline of the first subtask to be offloaded determined based on the first subtask to be offloaded in the offloading policy.
[0201] In the embodiments of the present application, the prediction data corresponding to the terminal device can also be separately sent by each terminal device to the access network device, so that the access network device reserves radio resources in advance. It should be understood that the content included in the uplink data-related prediction request message and the uplink data-related prediction message refers to the foregoing embodiments.
[0202] It should be noted that the terminal device can send the uplink data-related prediction message through the RRC UE assistance information (Assistance Information), or send the uplink data-related prediction message through the buffer status report (Buffer Status reporting, BSR) MAC control element (Control Element).
[0203] It should also be noted that the deadline of the first offloading subtask can be a soft deadline or a hard deadline. The soft deadline must be met, and the hard deadline is preferably met.
[0204] It should also be noted that only the foregoing steps S407 to S408 may be executed, or only the foregoing steps S409 to S410 may be executed, or steps S407 to S410 may be executed, which is not specifically limited herein.
[0205] In this way, in the embodiment of the present application, the network unit sends a data volume-related prediction message to the access network device, and / or the terminal device sends an uplink data-related prediction message to the access network device. The data volume-related prediction message and the uplink data-related prediction message may both include the uplink and downlink data volumes, arrival times, deadlines, etc. of one or more terminal devices, so that the access network device can reserve corresponding radio resources in advance, improving the transmission efficiency of data between the network unit and the terminal device, and further improving the processing efficiency of the first task.
[0206] In another embodiment of the present application, based on Figure 4 , a task offloading method as Figure 16 shown is provided. This method can be applied to a task offloading system including a terminal device and a network unit. The task offloading method may further include:
[0207] S411, the terminal device transmits a calculation result message to the network unit; correspondingly, the network unit receives the calculation result message, and the calculation result message includes the terminal calculation result of the terminal device for the first offloading subtask.
[0208] The terminal calculation result is obtained based on the offloading policy and by determining and executing the first offloading subtask assigned to the terminal device in the first task.
[0209] In the embodiment of the present application, the terminal device may start to execute the first offloading subtask in the first task according to the foregoing offloading policy issued by the network unit, complete the calculation of the first offloading subtask before the deadline of the terminal device indicated by the network unit, and upload the calculated terminal calculation result or the calculation intermediate variable to the network unit.
[0210] Further, the network unit starts to execute the second offloading subtask based on the received terminal calculation result, completes it before the task deadline of the first task, and sends the calculation result of the first task to the terminal device.
[0211] In some embodiments, before the foregoing step S411, the task offloading method applied to the terminal device and the access network device may further include:
[0212] S1101, the terminal device stores the terminal calculation result in the task buffer area, triggering the terminal device to send a resource request message to the access network device. The resource request message includes the identifier of the terminal device and the amount of data to be uploaded.
[0213] S1102, the network element receives the resource response message sent by the access network device. The resource response message includes the uplink scheduling resources allocated for the terminal device.
[0214] In the embodiment of the present application, when the terminal device completes the calculation of the first offloading sub-task, the generated terminal calculation result is stored in the task buffer area, and a resource request message sent to the access network device is triggered.
[0215] Among them, the task buffer can be a Computing Task Delay Buffer, which can include a Data Radio Bearer (DRB), or a Computing Radio Bearer (CRB). The resource request message can be a BSR, and the terminal device can also include the terminal calculation result in the BSR and report it to the access network device.
[0216] In response to the resource request message, the access network device sends a resource response message carrying the uplink scheduling resources (UL grant) allocated for the terminal device to the terminal device.
[0217] In some embodiments, the task offloading method applied to the terminal device may further include: determining the order of sending the resource request message via the transceiver based on the priority of at least one terminal calculation result in the task buffer area.
[0218] In the embodiment of the present application, after the terminal device receives the uplink scheduling resources allocated by the access network device, when deciding which data in the task buffer area to upload during the execution of the logical channel priority, the terminal device will consider the priority of the terminal calculation results stored correspondingly in these task buffer areas. The smaller the value of the priority, the higher the priority of the terminal calculation result and the task buffer area where it is located, and the earlier the order in which the terminal device sends the terminal calculation result. Among them, the priority of the terminal calculation result and the task buffer area where it is located is determined according to the difference between the first deadline and the current time.
[0219] Thus, in the embodiments of the present application, the terminal device requests uplink scheduling resources from the access network device, and based on the priorities of the terminal calculation results in the task buffer on the uplink scheduling resources, sends the terminal calculation results to the network unit through the access network device. In this way, the efficiency of the terminal device for uploading data can be improved, and the terminal calculation results can be sent before the network unit starts to execute the corresponding second offloading subtask, avoiding the task completion time exceeding the deadline.
[0220] In another embodiment of the present application, based on the foregoing embodiment, the process for the network unit to determine the offloading strategy may include:
[0221] S1201, determine at least one first strategy and at least one second strategy according to the task information, the relevant information of the first computing resources of the terminal device, the relevant information of the second computing resources of the network unit, and the air interface resource information of the terminal device or the air interface resource information of the serving cell of the terminal device.
[0222] Among them, the first strategy is used to represent the allocation arrangement of the first offloading subtask and / or the second offloading subtask, and the second strategy is used to represent the batch processing arrangement of the second offloading subtask by the network unit.
[0223] S1202, determine the offloading strategy based on the latency violation time determined by the first strategy.
[0224] In the embodiments of the present application, in order to ensure that different computing tasks from different terminal devices can be completed within their respective latency requirements and returned to the terminal devices, the network unit needs to comprehensively consider the following factors to determine the offloading strategy and schedule different second offloading tasks (for batch processing):
[0225] In the scenario of multiple terminal devices, the first task, the corresponding latency requirement (QoS), etc. in the task information provided by each terminal device;
[0226] The occupancy of the air interface time-frequency resources, including the latency of transmitting the terminal calculation results or the intermediate calculation variables on the air interface uplink, which is determined according to the data volume of the first offloading subtask in the offloading strategy, the air interface resources of the network unit, and the air interface resources of the terminal device or the air interface resources of the serving cell of the terminal device; it also includes the latency of transmitting the calculation results on the air interface downlink, which is determined according to the data volume of the second offloading subtask, the air interface resources of the network unit, the air interface resources of the terminal device or the air interface resources of the serving cell of the terminal device;
[0227] The relevant information of the first computing resources of the terminal device and the relevant information of the second computing resources of the network unit are respectively used to determine the computing latency (the first computing time) of the first offloading subtask on the terminal device and the computing latency (the second computing time) of the second offloading subtask on the network unit.
[0228] In this way, the network element will generate a computing offloading policy (the first policy) and a computing scheduling policy (the second policy) based on the current or predicted information on the radio access resource occupancy / idle status of the relevant terminal device or cell obtained from the access network device, and by comprehensively considering the QoS requirements of the first task, the relevant information on the first computing resource, the relevant information on the second computing resource, etc.
[0229] It should be understood that each second offloading subtask needs to be completed within a given latency constraint. As mentioned above, the latency violation time for completing the first task includes the terminal device processing latency, the uplink and downlink transmission latencies, and the network element processing latency. To minimize the latency violation time for the first task, it is crucial to carefully design the offloading policy for each first task. In addition, considering that the terminal device may generate new first tasks at any time, the network element can improve the computing efficiency and reduce the latency violation time by waiting for more first tasks to be uploaded and combining more second offloading subtasks into one processing cycle.
[0230] In the embodiment of the present application, the network element sends the information related to the corresponding terminal device in the generated first policy and second policy to the corresponding terminal device. Among them, the information related to the terminal device in the second policy may include: the start time and end time of the second offloading subtask offloaded by the terminal device to the network element. It should be noted that the start time and the end time may be an absolute timestamp, or a relative time based on the message sending time. In another case, the start time does not actually represent the time when the network element starts to process the second offloading task, but represents the time before which the terminal device needs to complete the local computation and upload the generated terminal computation result to the network element.
[0231] In the embodiment of the present application, after receiving the offloading policy sent by the terminal device, the network element may send back an Acknowledge message indicating that it has received and will execute according to this offloading policy.
[0232] Furthermore, after receiving the offloading policy, the corresponding terminal device, in order to ensure that the required terminal computation result is uploaded to the network element before the start time indicated by the network element, will calculate in real time and store the terminal computation result in the task buffer, and upload it based on the order determined in the foregoing embodiment.
[0233] Therefore, in the embodiments of the present application, it is necessary to optimize the task allocation between the terminal device and the network unit; it is also necessary to carefully design the batch processing of the network unit to avoid processing a large number of first tasks in a single batch, which may extend the inference time of a single task, so as to balance the trade-off between resource utilization and latency constraints; and, on the premise that the first tasks are generated by multiple users, competing for radio resources or the second computing resources of the network unit with other users, all of these are factors that can affect the network unit to determine the first strategy and the second strategy.
[0234] Therefore, when the network unit determines the first strategy and the second strategy, the problems to be solved include:
[0235] (1) Latency violation rate: The high frequency of tasks arriving at the edge server and the strict latency constraints make it difficult to ensure that all tasks are completed within the required time range, resulting in latency violations.
[0236] (2) Efficiency of batch processing: The lack of a strategy to effectively utilize the parallel processing ability of the network unit for batch processing tasks leads to the computing efficiency not reaching the optimal level.
[0237] (3) Adaptability to task arrival and resource competition: The continuous arrival of tasks and the competition of multiple users for wireless resources exacerbate the latency violation rate, and an adaptive algorithm is required for model partitioning and offloading.
[0238] (4) Task arrival prediction: In order to design an effective batch processing algorithm and model partitioning strategy, it is necessary to accurately predict the situation of offloading tasks arriving at the network unit.
[0239] In the embodiments of the present application, after the terminal device generates a new first task, it encapsulates the task information of the first task and / or the relevant information of the terminal device in an offloading request message and sends it to the network unit. As described above, the offloading request message may include the AI model required for the newly generated first task, the latency requirement for completing the first task, the computing queue length of the terminal device, the uplink transmission queue length, or may also include the channel state information of the terminal device.
[0240] Furthermore, the network unit continuously collects the status information and channel state information of various terminal devices through its communication facilities. By integrating it with the computing resource utilization status of the network unit itself, it can coordinate the allocation of computing and wireless communication resources in the entire system, including all terminal devices and itself. The result of this process is to generate an offloading strategy for the first offloading tasks of the mobile device and the second offloading subtasks of the network unit.
[0241] Specifically, based on the foregoing embodiments, for each time slot, the network unit performs the following operations:
[0242] (1) The network unit collects environmental status information, including the information in the offloading request message reported by the terminal device and the radio interface resource information (real-time channel status information) uploaded by the access network device.
[0243] (2) The network unit uses the currently collected radio interface resource information to evaluate the channel carrying capacity of each terminal device, that is, the number of bits that each terminal device can upload per time slot per unit frequency bandwidth. The network unit can also calculate the continuous average channel carrying capacity of each mobile device by using the dynamic sliding average method.
[0244] (3) For the terminal device in the uplink transmission stage, the network unit or the access network device allocates the uplink transmission bandwidth according to the currently uploaded first task deadline and the previously evaluated average channel carrying capacity.
[0245] (4) If the network unit receives a new first task from the terminal device within the current time slot, it determines the offloading strategy for the newly generated first task based on information such as the delay requirement of the first task, the first computing resource usage of the terminal device, the uplink transmission channel resource usage, and the second computing resource usage of the network unit. Subsequently, the network unit transmits the offloading strategy to the corresponding terminal device through the downlink.
[0246] (5) When determining the offloading strategy for the newly generated first task, the network unit also orchestrates its computing resources and reserves them to predict the first tasks arriving at the network unit in the future. The network unit decides whether to start the inference of the tasks currently cached in it and which specific second offloading subtasks to process according to the current state of computing resource allocation.
[0247] Among them, for step (3), as Figure 17 shown, the detailed process of bandwidth resource allocation executed by the network unit or the access network device is as follows:
[0248] S1301, Start.
[0249] S1302, Initialize the bandwidth allocation list and the data packet queue, and sort them in ascending order according to the message deadline.
[0250] S1303, Determine whether the remaining data packet queue is empty.
[0251] If it is, continue to execute step S1308; if not, continue to execute step S1304.
[0252] S1304, Retrieve the data packet at the end of the queue and calculate the required bandwidth resources.
[0253] S1305, Check whether the remaining bandwidth is sufficient.
[0254] If so, proceed to step S1306; if not, proceed to step S1307.
[0255] S1306. Allocate the bandwidth required for transmitting the current data packet to the corresponding mobile device.
[0256] S1307. Allocate the remaining bandwidth to the mobile device associated with the current data packet.
[0257] S1308. End.
[0258] Among them, in step (4) of the network unit operation process, the offloading strategy for the newly generated first task will be determined. At this time, two types of state variables need to be continuously updated. These variables represent the current utilization rate of communication resources and the utilization rate of the computing resources of the network unit respectively.
[0259] The reserved state of channel resources is represented by a list. The length of the list is a hyperparameter indicating the number of future time slots to be considered. In time slot t, the i-th element of the list represents the bandwidth not reserved for other first tasks in time slot t + i.
[0260] Figure 18 Shows an example where the list contains six elements, indicating the network unit's consideration of the states of the next six transmission time slots. Each time slot has five units of bandwidth available for allocation. However, since the network unit predicts that a first task will start uplink transmission in the next two time slots, four units of bandwidth have been reserved for this, leaving only one unit of frequency band resource available in time slot t + 2.
[0261] For step (5) above, as Figure 19 shown, it may include the following steps:
[0262] S1401. Start.
[0263] S1402. Update the available frequency band resource list and the batch scheduling strategy of the network unit according to the changes in the current task transmission state and channel conditions.
[0264] S1403. Collect the newly generated first tasks from the terminal device.
[0265] S1404. Determine whether offloading strategies have been established for all new first tasks.
[0266] If so, proceed to step S1408; if not, proceed to step S1405.
[0267] S1405. Select the task with the earliest deadline from the first tasks for which offloading strategies have not been determined.
[0268] S1406. Determine the offloading strategy and batch processing arrangement for the first task, and then update the available frequency band resource list and the batch scheduling linked list of the server.
[0269] S1407. Output the offloading strategy of the newly generated task, the available frequency band resource list, and the batch scheduling linked list of the network unit.
[0270] S1408. End.
[0271] Among them, as Figure 20 shown, for step S1406, it may include:
[0272] S1501. Start.
[0273] S1502. Determine whether all offloading strategies have been explored.
[0274] If so, continue to execute step S1506; if not, continue to execute step S1503.
[0275] S1503. Select an unexplored offloading strategy, calculate the local processing completion time, and then determine the time to reach the cache of the network unit according to the current available bandwidth resource list.
[0276] S1504. Iterate through all nodes in the network unit batch scheduling linked list associated with the first task, and select the optimal association method.
[0277] S1505. Use the first task to associate with the nodes in the batch scheduling list of the network unit, calculate the delay violation time, and evaluate the value of the offloading strategy.
[0278] S1506. Select the offloading strategy with the highest value, and update the available bandwidth list and the batch scheduling linked list of the network unit
[0279] S1507. End.
[0280] That is to say, the network unit will explore the possible first strategy and second strategy for each first task, determine the corresponding delay violation time under different combinations of the first strategy and the second strategy, and finally select the strategy with the minimum delay violation time. The delay violation time takes into account the inference completion time of the current task and the usage of communication and computing resources. Similarly, when calculating the merit value of various offloading methods for the first task, the inference completion time of the current scheme and the occupancy of communication and computing resources are comprehensively considered.
[0281] In some embodiments, the process of determining the delay violation time determined by the first strategy may include:
[0282] S12021. Determine at least one first computing time for the offloading strategy in the first task completed by the terminal device based on the relevant information of the first computing resource and at least one first policy.
[0283] S12022. Determine the second computing time corresponding to the second offloading subtask in the first task completed by the network unit based on the relevant information of the second computing resource, the first policy, and the second policy.
[0284] S1203. Determine the delay violation time based on the first computing time, the second computing time, the radio resource information, and the processing time parameter of the first task.
[0285] In the embodiments of this application, select the first policy and the second policy, and determine the delay violation time based on the following model.
[0286] (1) DNN inference task model. In the embodiments of this application, take executing DNN as the first task as an example. First, determine the DNN inference task model. Assume that the DNN for inference can be divided into N consecutive subtasks, and use A n to represent the computing load of the nth subtask, and use B n to represent the size of its output data. Specifically, B0 represents the size of the initial data input to the DNN model. Below, take MobileNet-v2 as an example for illustration.
[0287] (2) DNN splitting model. Divide the continuous time line into a series of equally long time slots (slots), and each time slot has a duration Assume that the probability that the terminal device k generates a new first task in each time slot follows the Bernoulli distribution Each new first task must be completed within the specified time limit When the terminal device k generates the first task it will immediately notify the network unit through the control channel. Then the network unit evaluates the current state of the computing and communication resources to determine the offloading strategy of the task, and relays this offloading strategy back to the terminal device.
[0288] In the embodiments of this application, for a given first task composed of N consecutive subtasks, there are N + 1 possible partitioning methods. In addition, there is also the option of not processing the task. Therefore, each first task has a total of N + 2 different first policies. In the embodiments of this application, use to represent the offloading strategy of the task , and its value set can be {-1, 0,..., N}. If n ∈ N, n < N, it means that the first n subtasks are processed by the terminal device, and the remaining subtasks from n + 1 to N are sent to the network unit for processing. A value of -1 indicates that the task is discarded, 0 indicates that the first task is fully offloaded to the network unit, and N indicates that the first task is fully processed on the terminal device.
[0289] It should be noted that at any given time slot, the terminal device k may have unfinished tasks from earlier time slots. Some of these tasks are still being processed locally, while others are being uploaded. In the embodiments of this application, these tasks are treated as separate cases.
[0290] (3) Local computing model. When modeling the local computing state of the terminal device, determine the first computing time required for the terminal device to compute. Denote as the remaining computing amount of the first task at time t, and determine the first computing time according to Before the task is generated it is set to 0; at when the terminal device k creates the first task and executes the policy the calculation formula is:
[0291]
[0292] where is a random variable, the probability of being equal to 1 is and the probability of being equal to 0 is For the state transition is represented by the following formula:
[0293]
[0294] It should be noted that in the above formula (2), f k is the computing amount completed by the key device per time slot. When the terminal device k initializes the unfinished task (k, t″) before the task i.e., this task will be given priority. Conversely, for tasks that violate the delay constraint, they will be deleted.
[0295] (4) Uplink transmission model. In the uplink data transmission phase, define as the data to be transmitted for the task at time t. When let For subsequent time the state transition formula is represented as:
[0296]
[0297] where
[0298]
[0299] As described above, formula (4) gives the When the terminal device selects the full offloading strategy, i.e., the task generation time is set to the data volume B0. On the contrary, for the partial offloading strategy between 1 and N-1, the data volume is allocated after the local processing is completed in the subsequent interval The completion of the local processing is determined by the condition. describes the state change of the task during the upload transmission:
[0300]
[0301] Among them, is the data volume transmitted by the terminal device k within the time slot t. Formula (5) reflects that the task will only be uploaded after the terminal device has completed the calculation. The terminal device k preferentially transmits the task with the most urgent deadline and discards the task that exceeds the delay constraint.
[0302] In the embodiment of the present application, it is considered that multiple terminal devices using orthogonal frequency division compete for the bandwidth of the frequency band. Let p k,t represent the power transmitted by the terminal device k in the time slot t, represent the channel vector between the terminal device k and the M antennas of the system at an interval of t, and the power of the Gaussian white noise is Therefore, the uplink transmission rate of the terminal device in the time slot k is:
[0303]
[0304] Therefore, if the terminal device occupies the bandwidth in the time slot t, the transmitted data volume is
[0305] In the embodiment of the application, the tasks from multiple devices use the Earliest Deadline First (EDF) strategy to compete for the uplink bandwidth, and the task with the most urgent deadline is preferentially allocated the bandwidth in each time slot t.
[0306] (5) Batch processing model. Once the task is uploaded to the network unit, the network unit performs batch processing. The network unit uses the associated variable to represent the subtask Whether it is executed in batch b, the second calculation time is determined according to the execution time of the network unit for different subtasks. If then the task is included in batch b, and each subtask is associated with at most one batch:
[0307]
[0308] Among them, the offloading strategy that is not connected to any batch of the task will be discarded. For the set 1} contains all tasks related to batch b, and the processing time of this batch is the sum of the calculation times of all subtasks in it:
[0309] (8)
[0310] Among them, F n (x) is the calculation time required for the server to process the n-type subtask x. It should be noted that when the network unit executes the forward propagation process of the same model, the calculation efficiency is higher. F n (x) is related to the hardware resources and the corresponding DNN structure. In order to calculate F n (x), a lookup table can be constructed in advance through experimental data. Only when F n (x) < xF n (x), parallel processing will provide advantages. The start time of batch b is denoted as earlier than the arrival time of all task data in the cache:
[0311]
[0312] In the embodiments of the present application, in order to ensure that the task delay meets its delay constraint, the following inequality must hold:
[0313]
[0314] And, in order to ensure that batch b + 1 does not start before the end of batch b, the following inequality must be satisfied:
[0315]
[0316] (6) Delay conflict model. For the task the completion status at an interval of t is expressed as follows:
[0317] In some embodiments, if the task is completely processed on the terminal device, the first calculation time
[0318] In some embodiments, if some subtasks are offloaded to the network unit, the second computing time
[0319] In the embodiments of the present application, a set is defined which contains the tasks generated between t ′ and t. The delay violation rate from time 0 to time t is calculated by the following formula:
[0320]
[0321] where, represents the number of elements in the set and only the tasks generated before are considered, and the tasks that have not yet reached the deadline are not considered. For formula (12) to be valid, the condition must be satisfied.
[0322] (7) Problem model. In the embodiments of the present application, the optimization objective is to minimize the delay violation time. The variables involved include the splitting strategy of each first task the variable associating the second offloaded subtasks with batches and the start time of each batch related to the above variables The optimization problem can be expressed by the following formula:
[0323]
[0324] Based on formula (13), it can be seen that the problem to be solved in the embodiments of the present application is essentially NP-hard. Due to the mutual dependence between the task splitting strategy and batch association, it becomes more complex, and both of them will affect the overall delay violation rate.
[0325] Therefore, in the embodiments of the present application, to solve the problem represented by formula (13), the network unit needs to associate a radio resource pool (WRP) and maintain a batch scheduling linked list (BSLL) that describes the communication and computing resource characteristics of the network unit.
[0326] In the embodiments of the present application, for the radio resource pool, at each time t, the state of the WRP can be represented by a vector where the i-th element represents the bandwidth resource available at time t + i. Assuming that the transmission capacity of the channel changes slowly, the evaluation of the uplink channel capacity of the network unit for the terminal device k at the interval t is denoted as then it can be updated by the moving average method:
[0327]
[0328] Among them, ρ represents the smoothing factor.
[0329] In the embodiments of the present application, when a first task appears, its information and the local computing queue information of the user are reported to the network unit. In this way, the network unit can calculate the local computing completion time of all tasks. Combining with the channel capacity evaluation, resources can be reserved for the tasks and updated For example, after determining the offloading strategy, the network unit can calculate the task will complete the operation of the terminal device at time t + i and initiate an uplink transmission. Taking the transmission data volume as calculation, the required bandwidth is Then, for j > i, update to At the same time, the time when the task arrives at the cache of the network unit can be evaluated for subsequent batch scheduling.
[0330] For the batch scheduling linked list, in some embodiments, for the above-mentioned second calculation time, it can be determined based on the following steps:
[0331] S120221, according to the first task included in the offloading request message and the processing time parameter of the first task, determine the batch processing order of the second offloading subtask in the scheduling linked list;
[0332] S120222, based on the batch processing order of the second offloading subtask in the scheduling linked list, determine the second calculation time.
[0333] Currently, for the batch processing strategy of the network unit, a solution is that a fixed batch processing size is predetermined. Once the number of tasks in the cache of the network unit reaches the set threshold, batch processing will be started.
[0334] In the embodiments of the present application, the reservation status of the second computing resource of the network unit is represented by the data structure of a linked list, where the i-th node (batch node) of the linked list represents the configuration information of the i-th first task starting from the current time slot. The configuration information includes the start time, end time of the inference, and the task set participating in the inference. In each time slot, the network unit may adjust the task set within the linked list node. At the same time, the network unit will calculate the start and end times of the inference batch represented by the node according to the status of the first task set within the node and the latest evaluated channel carrying capacity.
[0335] In the embodiments of the present application, Figure 21An example is given where the task set of the second offloading subtask within a linked list node can be represented by (i, j), indicating the j-th subtask of the first task i. Batch node 1 1601 includes subtask 2 and subtask 3 of the first task 1 corresponding to terminal device 1, and subtask 3 of the first task 1 corresponding to terminal device 2, indicating that these tasks will be processed in the next inference cycle of the network unit. Batch node 2 1602 includes subtask 3 and subtask 4 corresponding to the terminal device, and these subtasks are scheduled for the next cycle. It should be noted that the first tasks within a batch node are those that have been generated but not yet started. These tasks may be in the local computing stage, the uplink transmission stage, or already waiting in the cache of the network unit for processing. Due to the significant impact of channel conditions on uplink transmission latency during these stages, the network unit cannot determine the exact time when a task arrives at the cache. Instead, it estimates the arrival time based on the current assessment of channel capacity, which in turn affects the calculation of the start and end times of the node. In each time slot, the linked list must be updated according to changes in the task transmission status, changes in the channel carrying capacity, and the emergence of new first tasks.
[0336] It should be noted that the start time is determined as the later of the above two times: the end time of the previous node in the sequence and the time when the last task in the current node arrives at the network unit. For the first batch node, if the network unit is currently busy, the end time of the previous batch node is replaced with the end time of the local busy state; if the network unit is idle, it is replaced with the current time. The end time is calculated as the start time of the inference plus the time required for the tasks in the inference set. Since the GPU of the network unit has the ability of parallel processing, the inference time is not a linear sum of the single-task inference times, but a function of the number of first tasks. The network unit must construct a lookup table in advance to map the number of first tasks to the inference time.
[0337] Among them, in time slot (slot) t, the network unit sorts the newly generated first tasks according to their deadlines, from the earliest to the latest. Then, it sequentially determines the offloading strategies for these tasks. For the i-th first task, the network unit checks each possible offloading method, and for each method, considers where to place the task in the batch scheduling linked list of the network unit. The network unit determines the optimal offloading strategy and batch arrangement. According to this arrangement and its impact on communication resources, the network unit updates the available frequency band resource list and the batch scheduling linked list of the network unit.
[0338] It should also be noted that Figure 21It illustrates the process by which a network unit determines a new task offloading strategy. When associating a task with the batch scheduling linked list of the network unit, a more efficient approach is to explore only two scenarios: associating the task with the last node of the list and creating a new node for the task.
[0339] In the embodiment of the present application, the network unit determines whether to initiate a new inference batch according to the batch scheduling linked list. When the network unit is busy, it does not start a new inference batch. However, when the network unit is idle, it checks the batch scheduling linked list. If there are two nodes with non-empty task sets, it starts a new inference batch to process the tasks of the first node in the list. Or, if there is a node with a non-empty task set, the network unit evaluates whether to start a new inference batch according to a threshold determined by the task deadline of the node and the number of tasks in the node.
[0340] Among them, the computing resource status of the network unit is represented by a linked list data structure, where the i-th node represents the configuration details of the i-th batch of inferences starting from the current time slot. These details include the start time, end time, and the task set involved. Whenever a new task pops up, the network unit must find the relevant batch and then perform an incremental update on the linked list.
[0341] It should be noted that once an offloading task is generated its information is reported to the network unit to complete the update of the WRP. The time when the task is uploaded and reaches the cache is determined as t + i. Subsequently, the network unit will consider the following batch association options for this task:
[0342] Discard the task: In this case, the node status in the BSLL will not be updated. However, the new task will violate the latency constraint.
[0343] Associate with a new batch: This means that the network unit will create a new inference batch for this task. Specifically, the task set of the new node in the linked list is Subsequently, the network unit will calculate the start time, duration, and end time of the new inference batch according to the above formulas (9), (11), and (8) respectively. Finally, the number of tasks violating the latency constraint is evaluated.
[0344] Associate with batch which requires associating the subtask with the -th node. Let the task set associated with node in the BSLL be denoted as After selecting this option, the network unit updates the status of all nodes in the linked list according to formulas (7)-(11) and evaluates the number of tasks violating the latency constraint.
[0345] In the embodiments of the present application, according to certain criteria, a batch processing strategy can be selected to update the BSLL.
[0346] Based on the foregoing embodiments, when a terminal device starts a new inference task, it reports the basic information of the task, covering its computing and caching status, to the network unit. Subsequently, the network unit uses the task offloading method provided in the embodiments of the present application to determine the offloading separation point for the task and related batch inference, and exhaustively explores all possible task splitting points and batch associations. For each combination of the splitting (related to the first strategy) and batch processing strategy (related to the second strategy), calculate the updated results of the WRP and BSLL, the number of tasks violating the latency constraint, and the latency violation time, etc. Based on this information, evaluate the superiority of the combined strategy and finally select the optimal strategy.
[0347] In the embodiments of the present application, for a task with N + 2 potential splitting methods, each method can be associated with batch association alternative options, where represents the number of nodes in the current BSLL. To reduce the exploration complexity, each splitting strategy only considers three batch association options: discard the task, associate with batch association, and associate with a new batch. Therefore, the total number of splitting strategies and batch association combinations explored for each task is 3(N + 2).
[0348] Then, turn the attention to the evaluation criteria for each combined splitting and batch inference association strategy. The pros and cons of a given strategy depend on two key aspects: the number of tasks violating the latency limit after implementation and the available idle amount in the cloud-edge computing and communication resources. For a specific splitting and batch processing inference association strategy, the network unit first calculates the user's local computing queue, WRP, and the batch processing scheduling linked list after execution. Then, it counts the number of tasks violating the latency constraint. The slack of the user's computing resources is estimated by the time required for the user to clean the computing queue. The slack of the communication resources is measured by the weighted sum of the remaining bandwidth in each slot of the WRP. At the same time, the slack of the network unit's computing resources is quantified by the batch inference completion time of the final node of the network unit. Finally, the advantage value of the combined strategy is expressed as the weighted sum of these factors, and all the involved weight parameters are hyperparameters.
[0349] In the embodiments of the present application, when generating K tasks, a sequential decision-making method is adopted to determine the splitting and batch association of one task before proceeding to the next task. Therefore, the number of combinations of the first strategy and the second strategy that need to be explored is 3K(N + 2), which significantly reduces the computational complexity compared with the exhaustive search number (3N + 6) K .
[0350] In the embodiments of the present application, after the server completes the splitting strategy for all tasks in the current slot, it will decide whether to start inference according to the status of the BSLL. Specifically, batch inference is started when the status of the BSLL and the server meets one of the following two conditions:
[0351] If the server is in an idle state and the BSLL contains more than two nodes, then as long as all the tasks contained in the first node have reached the server, batch inference will be started.
[0352] If the server is in an idle state and the BSLL has only one node, then when all associated tasks have reached the server and the interval between their earliest deadlines and the node completion time exceeds a predefined threshold, batch inference will be triggered. In other cases, no new inference will be started.
[0353] In some embodiments, the computing offloading strategy includes at least one of the following:
[0354] The identifier of the first task and the list of identifiers of the first offloading subtasks;
[0355] The identifier of the first task and the list of identifiers of the second offloading subtasks;
[0356] The identifier of the first task, the list of identifiers of the first offloading subtasks, and the list of identifiers of the second offloading subtasks;
[0357] The identifier of the first task and the initial identifier of the first offloading subtask;
[0358] The identifier of the first task and the initial identifier of the second offloading subtask.
[0359] In some embodiments, the computing scheduling strategy includes at least one batch node, and the representation form of the batch node is at least one of the following:
[0360] Batch start time;
[0361] Batch end time;
[0362] The identifier of the terminal device included in the batch and the corresponding second offloading subtask.
[0363] In some embodiments,
[0364] The offloading strategy includes at least one of the following:
[0365] Computing offloading strategy;
[0366] Information related to the terminal device in the computing scheduling strategy;
[0367] The information related to the terminal device in the computing scheduling strategy includes at least one of the following:
[0368] Start time of the first offloading subtask;
[0369] End time of the first offloading subtask.
[0370] It should be noted that the computing offloading strategy can be the strategy finally selected from the foregoing at least one first strategy, the computing scheduling strategy can be the strategy finally determined from the foregoing at least one second strategy, and the computing offloading strategy and the computing scheduling strategy constitute the offloading strategy.
[0371] In this way, the embodiment of the present application provides a task offloading method, formulates a method for minimizing the delay violation time, and effectively reduces the computational complexity. At the same time, the overall computing and communication resources of the system are characterized by a wireless resource pool and a batch scheduling linked list. On this basis, by evaluating different task splitting and batch scheduling strategies, more effective decisions can be made.
[0372] Fourthly, in another embodiment of the present application, a task offloading system is provided, including: at least one access node, at least one network unit in the foregoing embodiment, and at least one terminal device as described in any one of claims 22 to 26.
[0373] Next, in combination with Embodiment 1 to Embodiment 3, the task offloading method provided by the embodiment of the present application will be elaborated in detail. This method can be applied to the task offloading system including a terminal device, an access network device, and a network unit provided by the embodiment of the present application.
[0374] It should be noted that the terminal device can be called a mobile device or a UE; the network unit can be called a network computing node or a Computing NF; the access network device can be called a RAN.
[0375] Embodiment 1: Computing offloading strategy and computing scheduling strategy in a multi-UE scenario
[0376] In the embodiment of the present application, a UE with computing offloading service requirements will send a computing offloading request to a network computing node (Computing Network Function). The Computing NF will comprehensively consider the following factors to generate a computing offloading strategy and a computing scheduling strategy for each UE to meet the Computing QoS requirements of each computing task of each UE:
[0377] Computing QoS requirements of computing tasks and corresponding subtasks submitted by each UE in a multi-UE scenario;
[0378] Occupancy of radio interface time-frequency resources;
[0379] Computing resource status on the UE and Computing NF sides.
[0380] Furthermore, the Computing NF sends the generated computing offloading policy and computing scheduling policy information to the UE. The UE will execute the received computing offloading policy, complete the local computing task in a timely manner, and upload the generated intermediate variables to the Computing NF for subsequent computing processing.
[0381] The Computing NF may be a MEC, NWDAF, or a dedicated network function node providing computing services.
[0382] It should be noted that, as Figure 22 shown, the task offloading method may include the following processes:
[0383] S1701, the UE establishes connections with the Computing NF and the RAN. It may be a control plane connection, a user plane connection, or a computing plane connection.
[0384] S1702, the UE generates a computing task (computing task or the first task), and the computing task may be:
[0385] (1) AI third-party applications, such as image rendering, object recognition, video processing, etc.
[0386] (2) Communication-related applications, such as networks applying AI, such as UE traffic prediction and UE trajectory prediction.
[0387] Among them, each first task can be split into multiple subtasks. For example, if an inference task is completed by a multi-layer AI model, the subtask may be one or more layers of the model.
[0388] S1703, the UE sends an offloading request message to the Computing NF, and the message format may be:
[0389] NAS message;
[0390] Computing Layer message, where the Computing Layer is a new protocol layer for interacting with computing task-related information;
[0391] Computing Layer message included in the NAS message.
[0392] In the embodiments of the present application, the offloading request message may include one or more of the following contents:
[0393] Related to the first task:
[0394] 1) The identifier (task ID / Index) of the first task.
[0395] 2) The description of the first task, such as the application description.
[0396] 3) The complete AI model parameters or files used by the first task.
[0397] 4) The QoS requirements of the first task, such as the latency requirement (in milli-seconds or seconds), the overall computing power requirement, etc.
[0398] Related to the subtasks in the first task:
[0399] 1) The identifier (subtask ID / Index) of the subtasks in the first task
[0400] 2) The task description of the subtasks in the first task, such as the type of the model layer.
[0401] 3) The complete AI model parameters or files used by the subtasks in the first task.
[0402] 4) The QoS requirements of the subtasks in the first task, such as the latency requirement (in milli-seconds or seconds), the overall computing power requirement, etc.
[0403] It may also include the identifier of the terminal device, such as UE RNTI, UE SUPI, etc., and the UE, network element, and access network device have a consistent understanding of the identifier of the terminal device.
[0404] It may also include the identifier (Cell ID) of the cell where the terminal device is located, such as PCI, etc., which is used to obtain the radio access network resource information of a specific cell in the following step S1704.
[0405] It may also include the computing power provided by the UE, including the number of processing units, the number of floating-point operations, and the abstract value of 1-100 used to describe the computing power level, etc.
[0406] In another example, the UE only reports its own UE ID information, such as UE RNTI (assigned by the base station in the radio network) or UE SUPI (assigned by the core network in the radio network), and the Computing NF can query the AMF to obtain information such as the RAN and Cell where this UE ID is located.
[0407] Examples of the NAS / Computing Layer are as follows Figure 10 and Figure 11 as shown, where Figure 10 represents the protocol stack corresponding to the control plane, Figure 11 represents the protocol stack corresponding to the user plane. Examples of NAS / Computing Layer messages are as follows Figure 12 as shown.
[0408] S1704. According to the received Cell ID information, the Computing NF will receive cell radio resource usage information or prediction information from the RAN node. For specific reference, refer to the following Example 2.
[0409] S1705. The Computing NF will comprehensively consider the following factors and generate a computing offloading policy and a computing scheduling policy for the UE's task offloading request, specifically referring to the foregoing steps S1201 to S1203.
[0410] The calculated computing offloading policy and computing scheduling policy meet the corresponding QoS requirements:
[0411] The Computing QoS requirements of the first tasks and corresponding subtasks submitted by each UE in the multi-UE scenario;
[0412] The occupation situation of radio frequency resources in the air interface;
[0413] The computing resource situation on the UE and Computing NF sides.
[0414] Among them, the computing offloading policy indicates which computing subtasks of the UE will be completed locally by the UE and which computing subtasks will be offloaded to the Computing NF for processing. Its representation form refers to the content included in the computing offloading policy in the foregoing embodiments.
[0415] Among them, the computing scheduling policy indicates the start and end times of one or more computing tasks or computing subtasks of one or more UEs on the Computing NF. For example, the Computing NF may batch process different computing subtasks from different UEs, and the Computing NF may generate a computing scheduling policy similar to the following form:
[0416] - Batch 1:
[0417] Start time: Tstart1;
[0418] End time: Tend1;
[0419] The second offloading subtask in the batch:
[0420] UE#1, Task#1, Subtask#2, 3;
[0421] UE#3, Task#1, Subtask#3.
[0422] In S1706, the Computing NF sends the information related to the UE in the computing offloading policy and the computing scheduling policy generated in S1705 to the corresponding UE. Among them, the information related to the UE in the computing scheduling policy may include the start time Tstart and the end time Tend of the computing subtasks offloaded by the UE to the Computing NF.
[0423] The start and end times may be an absolute timestamp Timestamp, or a relative time Time Offset based on the time when the message is sent.
[0424] In another case, Tstart does not truly represent the time when the Computing NF starts to process the computing subtasks, but represents the time before which the UE needs to complete the local computing and upload the generated intermediate variables to the Computing NF.
[0425] In S1707, after receiving the offloading policy from the Computing NF, the UE may send back an Acknowledge message to indicate acceptance and execution according to the offloading policy.
[0426] In S1708, the Computing NF may predict the uplink and downlink data volumes, the uplink and downlink data arrival times (which may be an absolute timestamp Timestamp, or a relative time Time Offset based on the time when the message is sent), and the task deadline related to the future computing offloading tasks of UE#1#2#3... based on the computing offloading policy and the computing scheduling policy generated in the fifth step, and send the prediction results of multiple UEs related to the RAN node (within the coverage of the RAN node cell) to the RAN node. For specific reference, see Embodiment 2.
[0427] In S1709, after receiving the computing offloading policy and the computing scheduling policy of S1706, the UE actively reports to the RAN the uplink and downlink data volumes and the data arrival times (which may be an absolute timestamp Timestamp, or a relative time Time Offset based on the time when the message is sent), and the task deadline related to the future computing offloading tasks of UE#1. The UE may report through RRC UE Assistance Information or BSR MAC CE.
[0428] It should be noted that either step S1708 or step S1709 can be performed. In addition, the task deadline may be a soft deadline or a hard deadline. A hard deadline must be met; a soft deadline is best met.
[0429] S1710, the RAN node reserves air interface time and frequency resources in advance according to the UE uplink and downlink data volume and arrival time prediction information received in step S1708 and / or step S1709.
[0430] S1711, according to the computing offloading strategy received in S1706 and the information related to the UE in the computing scheduling strategy, the UE starts to execute the first offloading subtask locally, completes the first offloading subtask before the time Tstart indicated by the Computing NF, and uploads the terminal computing result (computation intermediate variable) output by the first offloading subtask to the Computing NF. Refer to the third embodiment and the aforementioned steps S1101 to S1102.
[0431] S1712, the Computing NF starts to execute the first offloaded subtask to the Computing NF based on the received computational intermediate variables.
[0432] S1713, the network unit completes the calculation of the second offloading subtask before Tend, and sends the output result to the UE.
[0433] Example 2: Information exchange between Computing NF and RAN
[0434] In Example 2.1, the Computing NF invokes the service (RAN event exposure service) provided by the RAN node, and the RAN node sends the current or future air interface resource occupancy / idleness status of one or more Cells and / or the current or future UL / DL packet delay of the relevant UE(s) to the Computing NF once or periodically.
[0435] Figure 6 Taking the subscription periodic air interface resource occupancy / idleness as an example, the signaling interaction process between the RAN node and the Computing NF in the above step S1704 is explained.
[0436] First, Computing NF calls the Event Exposure service provided by RAN and sends an Nran_Event_Exposure_Subscribe Request HTTPs message to RAN, which contains the following radio resource prediction information:
[0437] The identifier (Cell ID) of the cell where one or more terminal devices are located;
[0438] The identifier (UE ID) of one or more terminal devices;
[0439] Radio availability prediction indicator;
[0440] Periodically, the RAN node should periodically send a Notify message to provide the prediction result;
[0441] Prediction Window. The prediction in each Notify message is about the available / occupied radio resources within a certain time window.
[0442] Among them, the Prediction Window may be represented in the following forms:
[0443] A duration or time interval value, which means that the prediction result provided by each Notify message is about the average available / occupied radio resources within a time range from the time point of this Notify message to the future.
[0444] A start time value and an end time value. In the scenario of periodic Notify, the Prediction Window of each prediction is related to the time point of the periodically triggered Notify message.
[0445] In the scenario of periodic Notify, if there is no explicit indication of the Prediction Window, it means that the prediction information carried in each Notify message is about the prediction within a Periodicity time period from the current time point to the next Notify message.
[0446] Secondly, RAN replies with an Nran_Event_Exposure_Subscribe Response HTTPs message to indicate acceptance.
[0447] Finally, the RAN node will provide the measurement prediction results of the radio resource availability / occupation of the relevant cell and / or the measurement prediction results of the UL / DL packet delay of the relevant UE once or periodically as required, and send them to the Computing NF through the Nran_Event_Exposure_Notify HTTPs message. The radio resource availability / occupation may be represented by an integer value, such as 0-100, where 100 means that the air interface resources are completely idle / occupied.
[0448] The RAN node may generate and provide the average UL / DL packet delay value representing the average UL / DL packet delay of all UEs under the current cell.
[0449] In another example, the occupation / idle situation of radio resources can be represented by the occupation of PRBs per cell or per SSB, ranging from 0 to 100.
[0450] If it is a one-time request, it is completed by two Nran_Event_Exposure_Request and Nran_Event_Exposure_Response. The logic is similar and will not be elaborated here.
[0451] In Embodiment 2.2, the Computing NF will send the uplink and downlink data volumes and the uplink and downlink data arrival times (which may be an absolute timestamp Timestamp or a relative time Time Offset based on the time when the message is sent) related to the future computing offloading tasks of multiple UEs, and send the prediction results to the RAN node. It may be based on the subscription of the RAN node, or it may also be an unsolicited proactive transmission.
[0452] Figure 14 Taking the information interaction based on the RAN node subscription as an example, the interaction process between the RAN node and the Computing NF in step S1708 above is described.
[0453] First, the RAN calls the Event Exposure service provided by the Computing NF and sends an Ncomp_Event_Exposure_Subscribe Request HTTPs message to the Computing NF, which contains the following radio interface resource prediction information:
[0454] The identifier (Cell ID) of the cell where one or more terminal devices are located. If the UE ID is not provided simultaneously, it means subscribing to the prediction information of all UEs with downlink computing task data under the relevant cell;
[0455] The identifier (UE ID) of one or more terminal devices;
[0456] An indicator for requesting uplink and downlink data volumes;
[0457] An indicator for requesting the upload time. The indicator for requesting uplink and downlink data volumes may be the same indicator.
[0458] Then, the Computing NF replies with an Ncomp_Event_Exposure_Subscribe Response HTTPs message indicating acceptance.
[0459] Finally, the Computing NF node provides the prediction information of one or more relevant UEs as required through an Ncomp_Event_Exposure_Notify HTTPs message. For example, when the Computing NF generates a computing task scheduling policy for multiple UEs, it can predict that a UE will complete a local computing subtask and need to upload the computing result at a certain future moment, or predict that the Computing NF will complete an offloaded computing subtask and need to send down the computing result at a certain future moment. For example:
[0460] The identifier of the terminal device:
[0461] The uplink transmission quantity;
[0462] The arrival time of the uplink transmission data;
[0463] The downlink data volume;
[0464] The arrival time of the downlink transmission data.
[0465] If it is not based on the RAN subscription method, the Computing NF can also, like in Embodiment 1, after executing step S1708, actively send the prediction information of one or more UEs to the RAN after generating the computing scheduling policy.
[0466] Embodiment 3: Task Buffer Awareness UL Transmission of Terminal Devices
[0467] In S1711 of the first embodiment, after the UE receives the information related to the UE in the computing offloading policy and the computing scheduling policy, in order to ensure that the required computing intermediate variables are uploaded to the Computing NF before the time Tstart indicated by the Computing NF, the UE will calculate the task buffer for local computing and data upload (Computing Task DelayBuffer) in real time.
[0468] In one example, when the UE completes local computing, the generated terminal computing result reaches the UE's DRB, or the CRB computing bearer, and triggers the generation of the BSR buffer status report, the UE includes the current task buffer of the computing-related data in the BSR and reports it to the RAN node.
[0469] In another example, when the UE receives the uplink scheduling resource (UL grant) from the RAN node and decides which DRB / CRB data to upload during the execution of the logical channel priority process, the UE will also consider the current task buffer of the computing-related data. For the DRB / CRB carrying the uplink transmission of the computing-related data, the smaller the value of the task buffer, the higher the priority of the DRB / CRB.
[0470] Based on the foregoing embodiments, the offloading method of this application is deployed on an MEC server equipped with a GPU to process the first tasks sent by different terminal devices in parallel. K = 6 users are evenly distributed in a 200m × 200m area and are served by a base station equipped with M = 36 antennas. In time slot t, the channel vector from user k to the base station is denoted as h k,t = β k g k,t , where β k is the large-scale fading factor, and g k,t ∈C M×1 is the small-scale fading factor whose elements follow distribution. The path loss model follows the formula 31.48 + 21.5log10(d) + 19log 10 (f c ), where d is in meters and f c is in Hz. Under the conditions of a bandwidth of 20 MHz and a carrier frequency of 3.5 GHz, the users operate with a fixed transmit power of 1 dBm and a noise power spectral density of -174 dBm / Hz. The duration of each time interval Ts = 1 ms.
[0471] Among them, the latency violation rate is the main evaluation index. Assuming that the processing time of each subtask on the terminal device is twice the single batch processing time on the server, under different latency requirements and task generation probabilities, the baseline 1-5 and the method of this application are tested in 5000 time slots based on the MobileNet-V2 model.
[0472] Among them, baseline 1 and baseline 4 represent the Random splitting approach (RS or RP), baseline 2 and baseline 5 represent the Full-offloading approach (FO), and baseline 3 represents the Non-offloading approach (CL or NO). For baseline 3, the first offloaded subtask is a positive integer between 1 and n, where n is the number of subtasks included in the first task. For FO and RS, the considered baseline batch scheduling policy is: when the server is idle and there are cached tasks, it immediately starts the C subtasks with the most stringent latency requirements in the cache. In particular, when C = 1, it is equivalent to the server not performing batch processing. In our experiment, we combine C = 1 and C = 10 with FO and RS, plus NO, for a total of 5 baselines: baseline 1 (RP.batch size = 1, or RS.C = 1); baseline 2 (FO.batch size = 1 or FO.C = 1); baseline 3 (CL or NO); baseline 4 (RP.batch size = 10, or RS.C = 10); baseline 5 (FO.batch size = 10 or FO.C = 10).
[0473] As Figure 23 and Figure 25 shown, under the latency limit of 30 milliseconds, the task generation frequency of each time slot is changed to measure the latency violation rate. This shows that, compared with the baseline, the task offloading method of this application can handle a higher task generation frequency before serious latency violations occur. As Figure 24 and Figure 26 shown, when the fixed task generation probability of each time slot is 0.05, for different latency requirements, it shows that the method of this application is resilient to more stringent latency constraints.
[0474] Thus, the embodiment of this application provides a task offloading method that comprehensively considers communication and computing factors to improve the overall efficiency of multi-user computing offloading.
[0475] In the fifth aspect, to implement the above task offloading method, asFigure 27 As shown in the figure, an embodiment of the present application provides a device 1800 (network unit or terminal device), which may include at least one processor 1801 and at least one transceiver 1802 coupled to the at least one processor 1801. The transceiver 1802 may include at least one separate receiving circuit system and transmitting circuit system, or at least one integrated receiving circuit system and transmitting circuit system. The at least one processor 1801 may be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field-programmable gate array (FPGA), etc.
[0476] According to some embodiments of the present application, when the device 1800 is a network unit, the processor is configured to:
[0477] Receive an offloading request message via the transceiver, where the offloading request message is a non-access stratum message or a computing stratum message, and the offloading request information message includes relevant information of the terminal device and / or task information of the first task;
[0478] Determine an offloading strategy for the first task based on the offloading request message;
[0479] Transmit a task assignment message via the transceiver, where the task assignment message includes the offloading strategy.
[0480] In some embodiments, the processor is configured to:
[0481] Receive a radio resource response message via the transceiver, where the radio resource response message includes at least one of the current radio resources information of the terminal device, the predicted radio resources information of the terminal device, the current radio resources information of the serving cell of the terminal device, and the predicted radio resources information of the serving cell of the terminal device;
[0482] Determine an offloading strategy for the first task based on the offloading request message and the radio resource response message.
[0483] In some embodiments, the processor is configured to:
[0484] Determine a first offloading subtask assigned to the terminal device in the first task and / or a second offloading subtask assigned to the network unit in the first task based on the offloading strategy.
[0485] In some embodiments, the offloading request message is a computing stratum message in the non-access stratum message.
[0486] In some embodiments, the task information of the first task includes one or more of the following:
[0487] The identifier, task description, model parameters, and quality requirements of the first task;
[0488] The identifiers, model parameters, and quality requirements of the subtasks of the first task;
[0489] The information of the terminal device includes one or more of the following:
[0490] The identifier of the terminal device;
[0491] The cell identifier of the cell where the terminal device is located;
[0492] The relevant information of the first computing resource of the terminal device.
[0493] In some embodiments, the processor is configured to:
[0494] When the information of the terminal device includes the identifier of the terminal device, transmit a cell identifier request message via the transceiver, where the cell identifier request message includes the identifier of the terminal device;
[0495] Receive a cell identifier response message via the transceiver, where the cell identifier response message includes the identifier of the cell where the terminal device is located.
[0496] In some embodiments, the processor is configured to:
[0497] The relevant information of the terminal device included in the offloading request message and / or the task information of the first task is encapsulated in the offloading task-related information field in the non-access stratum message or the computing stratum message; where
[0498] When the offloading request message is transmitted through the user plane, the non-access stratum message or the computing stratum message is encapsulated based on the protocol or format of the SDAP layer.
[0499] In some embodiments, the processor is configured to:
[0500] The relevant information of the terminal device included in the offloading request message and / or the task information of the first task is encapsulated in the offloading task-related information field in the non-access stratum message or the computing stratum message; where
[0501] When the offloading request message is transmitted through the control plane, the non-access stratum message or the computing stratum message is encapsulated based on the protocol or format of the RRC layer.
[0502] In some embodiments, the processor is configured to:
[0503] Transmit an air interface resource request message via a transceiver. The air interface resource request message is used to request an air interface resource response message and indicate the content included in the air interface resource response message.
[0504] In some embodiments, the air interface resource request message includes at least one of the following:
[0505] The identifier of at least one terminal device;
[0506] The identifier of the cell where at least one terminal device is located;
[0507] A radio availability prediction indicator;
[0508] Periodic information;
[0509] A prediction window.
[0510] In some embodiments, when the periodic information included in the air interface resource request message is sent periodically, the air interface resource response message includes at least one of the following:
[0511] The respective air interface resource information of at least one terminal device within the prediction window;
[0512] The respective air interface resource information of at least one terminal device before the time corresponding to the next air interface resource response message;
[0513] The respective average air interface resource information of at least one cell corresponding to at least one terminal device within the prediction window;
[0514] The respective average air interface resource information of at least one cell corresponding to at least one terminal device before the time corresponding to the next air interface resource response message.
[0515] In some embodiments, the processor is configured to:
[0516] Transmit a data volume-related prediction message via a transceiver. The data volume-related prediction message includes the upload time of the terminal calculation result determined based on the first offloading subtask, the uplink and downlink data volumes, and the deadline of the first offloading subtask.
[0517] In some embodiments, the processor is configured to:
[0518] Receive a data volume-related prediction request message via a transceiver. The data volume-related prediction request message is used to request a data volume-related prediction message.
[0519] In some embodiments, the data volume-related prediction request message includes at least one of the following:
[0520] The identifier of at least one terminal device;
[0521] The identifier of the cell where at least one terminal device is located;
[0522] An indicator for requesting the uplink and downlink data volume;
[0523] An indicator for requesting the upload time.
[0524] In some embodiments, the processor is configured to:
[0525] Receive a calculation result message via a transceiver, where the calculation result message includes the terminal calculation result of the first offloading subtask by the terminal device.
[0526] In some embodiments, the processor is configured to:
[0527] Determine at least one first policy and at least one second policy according to the task information, the relevant information of the first computing resource of the terminal device, the relevant information of the second computing resource of the network unit, and the radio resource information of the terminal device or the radio resource information of the serving cell of the terminal device;
[0528] Wherein, the first policy is used to characterize the allocation arrangement of the first offloading subtask and / or the second offloading subtask, and the second policy is used to characterize the batch processing arrangement of the second offloading subtask by the network unit;
[0529] Determine an offloading policy based on the delay violation time determined by the first policy.
[0530] In some embodiments, the processor is configured to:
[0531] Determine at least one first computing time for the terminal device to complete the offloading policy in the first task based on the relevant information of the first computing resource and at least one first policy;
[0532] Determine a second computing time for the network unit to complete the second offloading subtask corresponding to the first task based on the relevant information of the second computing resource, the first policy, and the second policy;
[0533] Determine the delay violation time based on the first computing time, the second computing time, the radio resource information, and the processing time parameter of the first task.
[0534] In some embodiments, the processor is configured to:
[0535] Determine the batch processing order of the second offloading subtask in the scheduling linked list according to the first task included in the offloading request message and the processing time parameter of the first task;
[0536] Determine the second computing time based on the batch processing order of the second offloading subtask in the scheduling linked list.
[0537] In some embodiments, the computing offloading policy includes at least one of the following:
[0538] List of the identifier of the first task and the identifiers of the first offloading subtasks;
[0539] List of the identifier of the first task and the identifiers of the second offloading subtasks;
[0540] Identifier of the first task, list of identifiers of the first offloading subtasks, and list of identifiers of the second offloading subtasks;
[0541] Identifier of the first task and the initial identifier of the first offloading subtask;
[0542] Identifier of the first task and the initial identifier of the second offloading subtask.
[0543] In some embodiments, the computing scheduling policy includes at least one batch node, and the representation form of the batch node is at least one of the following:
[0544] Batch start time;
[0545] Batch end time;
[0546] Identifiers of the terminal devices included in the batch and the corresponding second offloading subtasks.
[0547] In some embodiments, the offloading policy includes at least one of the following:
[0548] Computing offloading policy;
[0549] Information related to the terminal device in the computing scheduling policy;
[0550] The information related to the terminal device in the computing scheduling policy includes at least one of the following:
[0551] Start time of the first offloading subtask;
[0552] End time of the first offloading subtask.
[0553] According to some embodiments of the present application, when the device 1800 is a terminal device, the processor is configured to:
[0554] Transmit an offloading request message via the transceiver, the offloading request message is a non-access stratum message or a computing stratum message, and the offloading request information message includes information of the terminal device and / or task information of the first task;
[0555] Receive a task assignment message via the transceiver, the task assignment message includes an offloading policy; the offloading policy is used to indicate the first offloading subtask assigned to the terminal device in the first task.
[0556] In some embodiments, the processor is configured to:
[0557] Transmit a calculation result message via a transceiver, the calculation result message including a terminal calculation result;
[0558] The terminal calculation result is obtained based on an offloading policy and by determining and executing a first offloading subtask assigned to the terminal device in a first task.
[0559] In some embodiments, the processor is configured to:
[0560] Transmit an uplink data-related prediction message via a transceiver, the uplink data-related prediction message including an upload time of the terminal calculation result, an upload data volume, and a deadline of the first offloading subtask determined based on the first offloading subtask in the offloading policy.
[0561] In some embodiments, the processor is configured to:
[0562] Store the terminal calculation result in a task buffer area, trigger the transmission of a resource request message via a transceiver, the resource request message including an identifier of the terminal device and the upload data volume;
[0563] Receive a resource response message via a transceiver, the resource response message including uplink scheduling resources allocated for the terminal device.
[0564] In some embodiments, the processor is configured to:
[0565] Determine an order for transmitting the resource request message via a transceiver based on the priorities of at least one terminal calculation result in the task buffer area.
[0566] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0567] It should be noted that in the embodiments of the present application, if the above task offloading method is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0568] Sixth aspect, to implement the above task offloading method, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, the steps in the task offloading method provided in the above embodiment are implemented.
[0569] Seventh aspect, an embodiment of the present application provides a storage medium, that is, a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the task offloading method provided in the above embodiment are implemented.
[0570] It should be noted here that: The descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0571] It should be understood that the term "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or posterior. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0572] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0573] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0574] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0575] In addition, each functional unit in the embodiments of this application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0576] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage media include various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.
[0577] Alternatively, if the above integrated units of this application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of this application essentially or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of this application. The foregoing storage media include various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.
[0578] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A network unit, comprising: Transceiver; and a processor coupled to the transceiver, wherein the processor is configured to: Receiving an offload request message via the transceiver, the offload request message being a non-access layer message or a computing layer message, the offload request information message including relevant information of the terminal device and / or task information of the first task; Determining an uninstallation strategy for the first task based on the uninstallation request message; A task assignment message is transmitted via the transceiver, the task assignment message including the offloading policy.
2. The network unit of claim 1, wherein the processor is configured to: Receiving an air interface resource response message via the transceiver, the air interface resource response message including at least one of current air interface resource information of the terminal device, predicted air interface resource information of the terminal device, current air interface resource information of a serving cell of the terminal device, and predicted air interface resource information of a serving cell of the terminal device; An offloading strategy for the first task is determined based on the offloading request message and the air interface resource response message.
3. The network unit according to claim 1 or 2, wherein the processor is configured to: Based on the unloading policy, a first unloading subtask in the first task allocated to the terminal device and / or a second unloading subtask in the first task allocated to the network unit are determined.
4. The network unit according to claim 1, wherein the task information of the first task comprises one or more of the following: The identification, task description, model parameters and quality requirements of the first task; Identification, model parameters and quality requirements of the subtasks of the first task; The terminal device information includes one or more of the following: The identification of the terminal device; The cell identifier of the cell where the terminal device is located; Relevant information of the first computing resource of the terminal device.
5. The network unit of claim 1, wherein the processor is configured to: The relevant information of the terminal device and / or the task information of the first task included in the uninstallation request message is encapsulated in the uninstallation task related information field in the non-access layer message or the computing layer message; wherein, In the case where the offload request message is transmitted through the user plane, the non-access layer message or the computing layer message is encapsulated based on the protocol or format of the SDAP layer.
6. The network unit of claim 1, wherein the processor is configured to: The relevant information of the terminal device and / or the task information of the first task included in the uninstallation request message is encapsulated in the uninstallation task related information field in the non-access layer message or the computing layer message; wherein, In the case where the offload request message is transmitted through the control plane, the non-access layer message or the computing layer message is encapsulated based on the protocol or format of the RRC layer.
7. The network unit of claim 2, wherein the processor is configured to: An air interface resource request message is transmitted via the transceiver, where the air interface resource request message is used to request the air interface resource response message and indicate content included in the air interface resource response message.
8. The network unit of claim 3, wherein the processor is configured to: A data volume related prediction message is transmitted via the transceiver, wherein the data volume related prediction message includes an upload time of a terminal calculation result determined based on the first offloading subtask, an uplink and downlink data volume, and a deadline of the first offloading subtask.
9. The network unit of claim 2, wherein the processor is configured to: Determine at least one first strategy and at least one second strategy according to the task information, relevant information of the first computing resource of the terminal device, relevant information of the second computing resource of the network unit, and air interface resource information of the terminal device or air interface resource information of the serving cell of the terminal device; in, The first strategy is used to characterize the allocation arrangement of the first offloading subtask and / or the second offloading subtask, and the second strategy is used to characterize the batch processing arrangement of the second offloading subtask by the network unit; The uninstallation strategy is determined based on the delay violation time determined by the first strategy.
10. A terminal device, comprising: Transceiver; and a processor coupled to the transceiver, wherein the processor is configured to: Transmitting an offload request message via the transceiver, wherein the offload request message is a non-access layer message or a computing layer message, and the offload request information message includes information of the terminal device and / or task information of the first task; A task assignment message is received via the transceiver, wherein the task assignment message includes an offloading policy; the offloading policy is used to indicate a first offloading subtask in the first task that is assigned to the terminal device.