Computing service method and apparatus, and first node and second node
By including performance parameters in the computing service request message, the first node decides whether to accept or reject the computing service request, thus solving the problem of poor computing service quality in the prior art and achieving higher quality computing services.
Patent Information
- Application Number
- PCT/CN2025/105250
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-15
AI Technical Summary
The quality of service for computing services in existing mobile communication systems is difficult to guarantee, mainly because the quality of service is poor due to the 5G QoS Identifier being used to control the quality of computing services.
By including performance parameters in the computing service request message, the first node can decide whether to accept or reject the computing service request based on these parameters, thereby optimizing the quality of the computing service.
By including performance parameters in the computing service request message, the first node can more accurately select a suitable computing node, thereby improving the quality of computing services.
Smart Images

Figure CN2025105250_15012026_PF_FP_ABST
Abstract
Description
Computing service method, apparatus, first node and second node
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410909436.6, filed in China on July 8, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application belongs to the field of communication technology, specifically relating to a computing service method, apparatus, first node, and second node. Background Technology
[0004] With the development of communication technology, some mobile communication systems support the provision of computing services, such as Discrete Fourier Transform (DFT), Fast Fourier Transform (FFT), Artificial Intelligence (AI) model training, and AI model inference. However, these technologies still rely on quality parameters of the mobile communication system's communication services (e.g., the 5G QoS Identifier (5QI)) to control the quality of computing services, which can easily lead to poor service quality. Summary of the Invention
[0005] This application provides a computing service method, apparatus, first node, and second node, which helps to ensure the quality of computing services.
[0006] Firstly, a computing service method is provided, the method comprising:
[0007] The first node receives a computing service request message from the second node, the computing service request message including the performance parameters of the computing service;
[0008] The first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
[0009] Secondly, a computing service device is provided, the device comprising:
[0010] The receiving module is configured to receive a computing service request message from the second node, the computing service request message including performance parameters of the computing service;
[0011] The sending module is used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
[0012] Thirdly, a computing service method is provided, the method comprising:
[0013] The second node sends a computing service request message to the first node, the computing service request message including the performance parameters of the computing service;
[0014] The second node receives a computing service response message from the first node, the computing service response message being used to indicate whether to accept or reject the computing service request message.
[0015] Fourthly, a computing service device is provided, the device comprising:
[0016] The sending module is used to send a computing service request message to the first node, the computing service request message including the performance parameters of the computing service;
[0017] The receiving module is configured to receive a computing service response message from the first node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
[0018] Fifthly, an apparatus for computing services is provided, the apparatus being configured to perform the steps of the method described in the first aspect, or to implement the steps of the method described in the third aspect.
[0019] In a sixth aspect, a first node is provided, the first node including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.
[0020] In a seventh aspect, a first node is provided, including a processor and a communication interface, wherein the communication interface is used to receive a computing service request message from a second node, the computing service request message including performance parameters of a computing service;
[0021] The communication interface is also used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
[0022] In an eighth aspect, a second node is provided, the second node including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the third aspect.
[0023] In a ninth aspect, a second node is provided, including a processor and a communication interface, wherein the communication interface is used to send a computing service request message to a first node, the computing service request message including performance parameters of a computing service;
[0024] The communication interface is also used to receive a computing service response message from the first node, the computing service response message being used to indicate whether to accept or reject the computing service request message.
[0025] In a tenth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the third aspect.
[0026] Eleventhly, a wireless communication system is provided, comprising: a first node and a second node, wherein the first node is configured to perform the steps of the computing service method as described in the first aspect, and the second node is configured to perform the steps of the computing service method as described in the third aspect.
[0027] In a twelfth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the third aspect.
[0028] In a thirteenth aspect, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method as described in the first aspect, or to implement the steps of the method as described in the third aspect.
[0029] In this embodiment, a first node receives a computing service request message from a second node, the computing service request message including performance parameters of the computing service; the first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message. That is, by carrying the performance parameters of the computing service in the computing service request message, the first node can determine whether to accept the computing service request message based on the performance parameters of the computing service, which helps to ensure the service quality of the computing service. Attached Figure Description
[0030] Figure 1 is a block diagram of a wireless communication system applicable to an embodiment of this application;
[0031] Figure 2 is a flowchart of a computing service method provided in an embodiment of this application;
[0032] Figure 3 is a flowchart of another computing service method provided in an embodiment of this application;
[0033] Figure 4 is a flowchart of another computing service method provided in an embodiment of this application;
[0034] Figure 5 is a flowchart of another computing service method provided in an embodiment of this application;
[0035] Figure 6 is a flowchart of another computing service method provided in an embodiment of this application;
[0036] Figure 7 is a structural diagram of a computing service device provided in an embodiment of this application;
[0037] Figure 8 is a structural diagram of another computing service device provided in an embodiment of this application;
[0038] Figure 9 is a structural diagram of the communication device provided in an embodiment of this application;
[0039] Figure 10 is a structural diagram of a network-side device provided in an embodiment of this application;
[0040] Figure 11 is a structural diagram of another network-side device provided in an embodiment of this application;
[0041] Figure 12 is a structural diagram of the terminal provided in an embodiment of this application. Detailed Implementation
[0042] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0043] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0044] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0045] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and the term NR is used in most of the following description; however, these technologies can also be applied to systems other than NR systems, such as 6th generation (6G) radio systems. th Generation 6G communication system.
[0046] Figure 1 shows a block diagram of a wireless communication system applicable to an embodiment of this application. The wireless communication system includes a terminal 11 and a network-side device 12. The terminal 11 can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the specific type of terminal 11 is not limited in this application embodiment. Network-side equipment 12 may include access network equipment or core network equipment, wherein access network equipment may also be referred to as Radio Access Network (RAN) equipment, radio access network function, or radio access network unit. Access network equipment may include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.The term "base station" can be referred to as Node B (NB), Evolved Node B (eNB), Next Generation Node B (gNB), New Radio Node B (NR Node B), Access Point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), Radio Base Station, Radio Transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home Evolved Node B, Transmit / Receive Point (TRP), or any other suitable term in the relevant field, as long as the same technical effect is achieved. The term "base station" is not limited to any specific technical terminology. It should be noted that this application embodiment only uses a base station in an NR system as an example for description and does not limit the specific type of base station.
[0047] Core network equipment, also known as core network nodes, core network functions, or core network elements, includes, but is not limited to, at least one of the following: Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (or L-NEF), and Binding Support. The core network functions include: BSF (Block Network Function), Application Function (AF), Location Management Function (LMF), Gateway Mobile Location Centre (GMLC), and Network Data Analytics Function (NWDAF). It should be noted that this application embodiment only uses core network equipment in the NR system as an example and does not limit the specific type of core network equipment. If the name of the core network equipment mentioned in this application embodiment changes in subsequent protocol versions (e.g., 6G), it will still be within the scope of protection of this application.
[0048] Optionally, the core network equipment can be implemented by one or more functional modules in a single device, or by multiple devices working together; this application does not specifically limit this. It is understood that the aforementioned functional modules can be network elements in hardware devices, software functional modules running on dedicated hardware, or virtualized functional modules instantiated on a platform (e.g., a cloud platform).
[0049] For ease of understanding, the following describes some aspects of the embodiments of this application:
[0050] I. Location Service Quality (QoS)
[0051] The performance requirements for different location service levels are described in Table 7.3.2.2-1 of protocol TS22.261.
[0052] Positioning accuracy is used to represent the distribution of positioning service performance errors. It is defined using a confidence level and a positioning error threshold, which is the percentage of the distance between the positioning result and the actual location that is within the positioning error threshold range. For example, a 95% confidence level for positioning accuracy <3m means that there is a 95% probability that the positioning result has a distance error of less than 3 meters from the actual location.
[0053] Location QoS is included in the location request, which includes both the location request from the location requester and the location request from the location result provider (e.g., the Location Management Function (LMF)). The LMF's location request is generated based on the location request from the location requester.
[0054] The location request from the aforementioned location requester (e.g., a Location Services (LCS) client or AF) may include the following parameters:
[0055] (1) QoS type / level (LCS QoS Class), including:
[0056] Best Effort: The most lenient QoS type for location services. If the location result does not meet other QoS requirements, the location result still needs to be fed back, but it needs to indicate that the requested QoS was not met. If no location result is obtained, the reason for failure is fed back.
[0057] Multiple QoS (QoS) type: This is a moderately stringent QoS type for location, meaning it includes QoS requirements corresponding to multiple QoS levels. If the location result does not meet the most stringent QoS requirement, LMF will initiate the location process again to try to meet the lower QoS requirements until one of the QoS requirements is met. If the most lenient QoS requirement is still not met, no location result will be reported, only the reason for the failure will be reported.
[0058] Assured type. The most stringent location QoS type. If the location result does not meet other QoS requirements, no location result will be provided; only the reason for the failure will be reported.
[0059] (2) Positioning accuracy, including horizontal positioning accuracy and / or vertical positioning accuracy;
[0060] (3) Response Time type: LMF needs to balance positioning accuracy and response time type.
[0061] No delay: The LMF should immediately report the initial location or the most recent location result of the target UE. If no location result is found, a failure message should be reported, and a location process can be triggered to respond to subsequent location requests.
[0062] Low latency: Response time is prioritized over accuracy. The LCS server should return the current location with minimal latency.
[0063] Delay-insensitive: Prioritizing accuracy over response time. LMF can delay the feedback of positioning results until the required positioning accuracy is met.
[0064] The location request for the aforementioned LMF may include the following parameters:
[0065] Horizontal accuracy includes accuracy and confidence level.
[0066] Vertical accuracy includes both accuracy and confidence level.
[0067] Response time is the delay between when the UE receives a location information request and when the location information is provided.
[0068] II. AI and Communication Application Scenarios
[0069] AI and communications are among the 6G application scenarios identified by the International Telecommunication Union Radiocommunication Sector (ITU-R). Typical use cases include IMT-2030 (6G) assisted autonomous driving, autonomous collaboration between devices in healthcare applications, cross-device / network compute offloading, creation and prediction of digital twins, and IMT-2030 (6G) assisted collaborative robots. These application scenarios will require support for high mobile network capacity and user experience data rates, as well as low latency and high reliability. Beyond communications, this application scenario is expected to include a set of new functionalities integrating AI and computing-related features into 6G systems, including data acquisition, preparation, and processing from various sources; distributed AI model training; model sharing and distributed inference across mobile communication systems; and compute resource orchestration.
[0070] Existing cloud services primarily define the static, long-term (e.g., monthly) quality of service between cloud service providers and users through Service Level Agreements (SLAs). These services mainly target enterprise users and have fewer direct consumer-facing users, especially mobile consumers.
[0071] The computing service method provided in this application will be described in detail below with reference to the accompanying drawings and through some embodiments and application scenarios.
[0072] Please refer to Figure 2, which is a flowchart of a computing service method provided in an embodiment of this application. The method can be executed by a first node, as shown in Figure 2, and includes the following steps:
[0073] Step 201: The first node receives a computing service request message from the second node, the computing service request message including the performance parameters of the computing service.
[0074] In this embodiment, the aforementioned first node may also be referred to as a computing management node, computing management function, computing control function, computing management and control function, computing service control function, or computing service management function, etc. Furthermore, the aforementioned first node may be a radio access network node or a core network node. The aforementioned first node may be a network node with at least one function, such as receiving computing service requests, processing computing service requests, scheduling computing resources, exchanging computing information, or processing computing data. The aforementioned first node may be a newly added network node; or it may be a node obtained by enhancing an existing network node, such as an enhanced AMF, enhanced SMF, or enhanced base station, etc.
[0075] The second node mentioned above may include, but is not limited to, at least one of a terminal, an AF (Application Function), and a Network Function (NF). When the second node is an AF, the computation service request message needs to be sent to the first node via a NEF (Network Function Entity). It should be noted that the second node can also be referred to as a computation request node.
[0076] The aforementioned computing services may include at least one of DFT, FFT, AI model training, and AI model inference. The AI model training or AI model inference may include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, localization, perception, and channel state information (CSI) feedback. The aforementioned computing services may include at least one computing task.
[0077] The performance parameters of the aforementioned computing service can be used to characterize the performance requirements that the computing service needs to meet. For example, the performance parameters of the aforementioned computing service may include, but are not limited to, at least one of the following: hardware requirements for the computing task, computing requirements, computing result requirements, and transmission requirements. The aforementioned hardware requirements may include, but are not limited to, at least one of the following: computing power type, minimum memory, and minimum storage. The aforementioned computing requirements may include, but are not limited to, at least one of the following: minimum number of operations or operands, minimum computing speed, minimum computing intensity, computing power consumption threshold, and computing energy efficiency threshold. The aforementioned computing result requirements may include, but are not limited to, at least one of the following: maximum failure rate and lower limit of AI model performance. The aforementioned transmission requirements may include, but are not limited to, at least one of the following: minimum transmission bandwidth and transmission latency budget.
[0078] Specifically, the first node can determine whether to accept the computing service request message based on the performance parameters of the aforementioned computing service. For example, the first node can select a computing node based on the computing service request message. For instance, the first node can select a computing node that meets the performance parameters of the aforementioned computing service, and can determine whether to accept the computing service request message based on the selection result. It is understood that when selecting a computing node, in addition to the performance parameters of the aforementioned computing service, other parameters can also be considered, such as the status of the computing node, the type of computing service, etc.
[0079] It should be noted that the number of compute nodes selected can be at least one, or it can be zero, meaning no suitable compute nodes were selected. For example, if no compute nodes exist that can meet the performance parameters of the computing service, the number of compute nodes selected is zero.
[0080] Step 202: The first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
[0081] For example, when the number of selected computing nodes can be at least one, the above computing service response message is used to indicate acceptance of the computing service request message; when the number of selected computing nodes is 0, the above computing service response message is used to indicate rejection of the computing service request message.
[0082] In this embodiment, a first node receives a computing service request message from a second node, the computing service request message including performance parameters of the computing service; the first node sends a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message. That is, by carrying the performance parameters of the computing service in the computing service request message, the first node can determine whether to accept the computing service request message based on the performance parameters of the computing service, which helps to ensure the service quality of the computing service.
[0083] Optionally, the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
[0084] Resource type; minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, representing the upper limit of the computing task's duration; computing power type; data type; minimum memory, representing the lower limit of memory required for the computing task; minimum storage, representing the lower limit of storage required for the computing task; minimum transmission bandwidth, representing the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
[0085] The resource types mentioned above can be used to reflect QoS classes, such as guaranteed, non-guaranteed, and latency-sensitive. These resource types can also be referred to as computing QoS classes, or computing and communication QoS classes, etc.
[0086] Optionally, the resource type includes at least one of the following:
[0087] Guaranteed Computing Rate (GCR), Guaranteed Computing Intensity (GCI), Non-Guaranteed Computing Rate (Non-GCR), Non-Guaranteed Computing Intensity (Non-GCI), Delay-Critical Guaranteed Computing Rate (DCR), Delay-Critical Guaranteed Computing Intensity (DCI).
[0088] It should be noted that the above-mentioned guaranteed computational speed can also be called guaranteed operational rate (GOR), and the above-mentioned guaranteed computational intensity can also be called guaranteed operational intensity (GOI).
[0089] For example, the computation speed characterized by at least one of the guaranteed computation speed, non-guaranteed computation speed, and latency-sensitive guaranteed computation speed can refer to the computation time complexity divided by the tolerable upper limit of computation time, wherein the computation time complexity can be represented, for example, by operands.
[0090] For example, the computational strength characterized by at least one of the guaranteed computational strength, non-guaranteed computational strength, and latency-sensitive guaranteed computational strength may refer to the computational speed divided by the bandwidth.
[0091] The aforementioned guaranteed computing speed can be understood as the computing resources associated with the guaranteed computing speed of a computing task being permanently allocated during the computing task. For example, computing resources such as CPU, GPU, or storage on the selected computing node associated with the guaranteed computing speed of the computing task are allocated to that computing task. The computing resources allocated to that computing task are dedicated to that computing task and are not shared with other computing tasks.
[0092] The aforementioned guaranteed computing strength can be understood as computing resources, or computing and communication resources, permanently allocated during the computing task to guarantee its computing strength. For example, when the computing strength is the value obtained by dividing the computing speed by the memory bandwidth, computing resources such as CPU, GPU, and memory related to the guaranteed computing strength are allocated to the computing task; when the computing strength is the value obtained by dividing the computing speed by the target bandwidth (i.e., the smaller of the memory bandwidth and the transmission bandwidth), CPU, GPU, memory, and network resources related to the guaranteed computing strength are allocated to the computing task. Network resources include air interface resources or wired transmission resources, etc. Air interface resources are related to the air interface transmission bandwidth of the computing task between the UE and the network, while wired transmission resources are related to the transmission bandwidth from the access network to the computing node of the computing task between the UE and the network.
[0093] The aforementioned latency-sensitive guarantee of computational speed can be understood as follows: if the latency of a data packet in a computation task exceeds the computational latency budget, the data packet is discarded. Furthermore, the latency-sensitive guarantee of computational speed requires that a first proportion (e.g., 98%) of data packets do not exceed the computational latency budget.
[0094] The aforementioned latency-sensitive guarantee of computational strength can also be understood as follows: if the latency of a data packet in the computation task exceeds the computational latency budget, the data packet is discarded. Furthermore, the latency-sensitive guarantee of computational speed requires that a second proportion (e.g., 98%) of the data packets cannot exceed the computational latency budget.
[0095] The aforementioned non-guaranteed computing speed can be understood as resources related to the computing speed of a computing task not being permanently allocated during that computing task. For example, computing resources such as CPU, GPU, or storage on the selected computing node that are related to the guaranteed computing speed of the computing task are not allocated to that computing task but are shared with other computing tasks.
[0096] The aforementioned non-guaranteed computational intensity can be understood as resources related to the computational intensity of a computational task not being permanently allocated during that task. For example, computational resources such as CPU, GPU, and memory, as well as network resources, on the selected computing node that are related to the guaranteed computational speed of the task, are not allocated to that task but are shared with other computational tasks.
[0097] It should be noted that the above resource types can determine the allocation of computing resources related to the guaranteed computational load at the computing task level, or the allocation of computing resources related to the guaranteed computational intensity at the computing task level, or the allocation of computing and communication resources. The aforementioned computing task level can be a single computing task level or a computing task group level. For example, if the task identifier (ID) of the large model service is A, then the above resource type determines the computing resource allocation for task ID A; as another example, if the task ID of AI model training for AI beam management is B, then the above resource type determines the computing resource allocation for task ID B; as yet another example, if the task ID of the large model service is A, and the task ID of AI image recognition is C, and tasks A and C form a task group, then the above resource type can determine the computing resource allocation at the task group level.
[0098] The aforementioned minimum number of operands can be used to represent the minimum number of operands required for a computational task.
[0099] The aforementioned minimum computation speed can be used to indicate the lower limit of the computation speed for a computational task. Here, computation speed can refer to the computational time complexity divided by the tolerable upper limit of computation time. For example, computational time complexity is typically represented by operations, and usually corresponds to a data type. For instance, 2 floating-point operations (TFLOPs) indicates that the time complexity of the computational task is 2 × 10⁻⁶. 12 Floating-point operands; if the maximum tolerable computation time for this task is 20ms, meaning the computation can be completed in a maximum of 20ms, then the computation speed is 2 × 10⁻⁶. 12 / (20×10 -3 = 100 floating-point operations per second (TFLOPS).
[0100] Optionally, the minimum computing speed includes at least one of the following: theoretical minimum computing speed, and actual minimum computing speed.
[0101] In some optional embodiments, the performance parameters may further include a first test case indication, which corresponds to the actual minimum computing speed and is used to indicate that the actual minimum computing speed is the minimum computing speed obtained based on the test cases indicated by the first test case indication.
[0102] In some optional embodiments, the actual minimum computing speed can be expressed by the theoretical minimum computing speed and computing efficiency; or it can be expressed by the ideal minimum computing speed and computing efficiency. The aforementioned ideal minimum computing speed can be understood as the minimum computing speed obtained under ideal conditions such as no task preemption, based on test cases.
[0103] The aforementioned minimum computational intensity can be used to indicate the lower limit of the computational intensity of a computational task. Here, computational intensity can refer to computational speed divided by bandwidth. It should be noted that in the process of providing computational services in a mobile communication system (e.g., 6G), bandwidth consists of multiple components, including the transmission bandwidth between the second node (e.g., UE) and the computational node, and the memory bandwidth of the computational node. One implementation is to define computational intensity in segments; for example, computational intensity is computational speed divided by memory bandwidth. Another implementation is to use the smallest of the aforementioned multiple bandwidth components as the bandwidth for calculating computational intensity, i.e., computational intensity is computational speed divided by min{memory bandwidth, transmission bandwidth from the second node to the computational node}, where min{memory bandwidth, transmission bandwidth from the second node to the computational node} represents the smaller of the transmission bandwidth and the memory bandwidth.
[0104] In some optional embodiments, the transmission bandwidth between the second node and the computing node can be further divided into the air interface bandwidth between the second node and the access network node, and the wired transmission bandwidth between the access network node and the computing node; or the transmission bandwidth between the second node and the computing node can be further divided into the bandwidth between the second node and the UPF, and the wired transmission bandwidth between the UPF and the computing node.
[0105] The aforementioned Computing Delay Budget (CDB) indicates at least one of the following: the upper limit of tolerable computation delay, the upper limit of tolerable transmission delay, and the upper limit of tolerable computation and transmission delay. The aforementioned delay upper limit can be understood as the maximum delay.
[0106] Optionally, the computation delay budget includes at least one of the following: an upper limit for computation delay, an upper limit for transmission delay, and an upper limit for both computation and transmission delay.
[0107] The above upper limit of computation latency can be understood as the upper limit of the latency that the computation task can tolerate when it is performed between the second node (such as the UE) and the computation node.
[0108] The upper limit of the above transmission delay can be understood as the upper limit of the sum of the transmission delay from the second node to the computing node and the transmission delay from the computing node to the computing receiving node.
[0109] The aforementioned upper limit for computation and transmission latency can be understood as the upper limit of the sum of the latency of the computation task between the second node (such as the UE) and the computation node, the transmission latency from the second node to the computation node, and the transmission latency from the computation node to the computation receiving node. Here, the aforementioned computation receiving node can be understood as the node that receives the computation response data.
[0110] For example, the latency represented by the above computational latency budget can be defined in two ways: one is the length of the time interval between the sending of the first data packet of a single computational task and the receiving of the last data packet of that task; the other is the length of the time interval between the sending of the first data packet of a group of computational tasks and the receiving of the last data packet. For example, for image recognition computational tasks, one approach is to treat single image recognition as a single computational task, while another approach is to treat multiple images (e.g., 100 images) as a group of computational tasks.
[0111] For example, the latency represented by the computational latency budget for AI model inference can include at least one of the following:
[0112] Total end-to-end inference latency: Specifically, this refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time elapsed before sending the first byte of the first computation task (or job) is denoted as T. ISThe last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS .
[0113] Total inference latency: Specifically, it refers to the total latency of multiple consecutive inference operations. The calculation method is as follows: the time before inference for the first computational task (or computational job) is denoted as T. ITS The time when all computational tasks (or computational jobs) finish inference is denoted as T. ITE Then the computational latency budget for AI model inference is T. ITE -T ITS .
[0114] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time t is the time before the second node sends the first byte of a computation task (or job). TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS .
[0115] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The time when the inference of this computational task ends is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0116] The failure rate mentioned above represents the ratio of the number of failed computation requests to the total number of requests within a unit of time. For example, the unit of time can be 1 second, 5 minutes, 1 day, etc. The failed requests can include incomplete computation requests and computation requests that completed after exceeding a latency threshold. The total number of requests can be the total number of requests received by the first node, or the total number of valid requests, where valid requests refer to computation requests accepted and processed by the first node.
[0117] The average window mentioned above is used to represent the average statistical time window of at least one of the computation speed, computation intensity, and computation delay budget.
[0118] Maximum time is used to represent the upper limit of the duration of a computation task, that is, the maximum time the computation task can last. For example, the maximum time can be represented by at least one of the start time, duration, end time, etc.
[0119] The aforementioned computing power types may include at least one of the following: a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a data processing unit (DPU), a smart network interface card (smartNIC), a tensor processing unit (TPU), and a neural network processing unit (NPU).
[0120] Optionally, the above-mentioned computing power types may also include at least one of the following:
[0121] Clock speed, for example, the lowest clock speed;
[0122] Number of cores, for example, minimum number of cores.
[0123] The data types mentioned above may include at least one of integers (e.g., int8, int4, etc.) and floating-point numbers (e.g., half-precision floating-point numbers, single-precision floating-point numbers, double-precision floating-point numbers, double-half-precision floating-point numbers, etc.).
[0124] The minimum memory mentioned above represents the lower limit of memory required for a computational task. For example, the minimum memory can be represented by at least one of memory size (e.g., 8GB), sustained memory bandwidth, and memory random access rate.
[0125] The minimum storage mentioned above represents the lower limit of storage required for a computing task. For example, the minimum storage can be represented by at least one of storage size (e.g., 1T) and storage bandwidth (e.g., the maximum input / output (IO) flow per unit time).
[0126] The aforementioned minimum transmission bandwidth is used to represent the lower limit of the bandwidth of the computing task, for example, at least one of the lower limit of the uplink bandwidth and the lower limit of the downlink bandwidth that the computing task can tolerate.
[0127] Optionally, the minimum transmission bandwidth includes at least one of the following: minimum uplink bandwidth, minimum downlink bandwidth, uplink bandwidth indication, downlink bandwidth indication, and target indication;
[0128] The target indication is used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different.
[0129] That is, the minimum transmission bandwidth is represented by at least one of the following: minimum uplink bandwidth, minimum downlink bandwidth, uplink bandwidth indication, downlink bandwidth indication, and target indication.
[0130] The uplink bandwidth indicator described above is used to indicate uplink bandwidth; for example, 0 represents uplink bandwidth. The downlink bandwidth indicator described above is used to indicate downlink bandwidth; for example, 1 represents downlink bandwidth.
[0131] The above target indication is used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different. For example, 0 indicates that the uplink bandwidth and downlink bandwidth are different, and 1 indicates that the uplink bandwidth and downlink bandwidth are the same.
[0132] The above-mentioned computational task arrival pattern is used to represent the arrival pattern of the computational job of the computational task.
[0133] Optionally, the computing task includes at least two computing jobs, and the computing task arrival mode includes at least one of the following:
[0134] Continuous arrival mode or single arrival mode is used to indicate that one computation job arrives at a time;
[0135] Fixed-cycle arrival mode is used to indicate the arrival of control calculation jobs according to a fixed cycle;
[0136] Poisson distribution arrival pattern, used to indicate the arrival of computational jobs controlled by Poisson distribution;
[0137] Peak arrival mode is used to indicate the arrival of σ computational jobs within a target period of Poisson distribution, where the duration of the target period is less than a preset duration, and σ is a positive integer.
[0138] Offline arrival mode, used to indicate that all computation jobs arrive at once.
[0139] The aforementioned continuous arrival mode or single arrival mode is used to indicate that a computation job arrives at a time. For example, the i-th computation job arrives after the (i-1)-th computation job is completed. If the (i-1)-th computation job is not completed or the computation delay budget threshold is not reached, the i-th computation job is not sent. i is a positive integer.
[0140] The fixed-period arrival pattern described above is used to indicate the arrival of computational jobs controlled according to a fixed period. For example, n computational jobs arrive at intervals of time T, where T represents the fixed period and n is a positive integer.
[0141] The aforementioned Poisson distribution arrival pattern is used to indicate the arrival of computational jobs controlled by the Poisson distribution, for example, according to The arrival of computational jobs is controlled, where k represents the number of computational jobs arriving within a target unit of time, and k is a positive integer; λ represents the average number of computational jobs arriving within the target unit of time, and λ is a positive integer. For example, the target unit of time could be 1 second or 2 seconds, etc.
[0142] The aforementioned peak arrival mode is used to control the arrival of σ computational jobs within a target period of a Poisson distribution, wherein the duration of the target period is less than a preset duration, for example, the preset duration could be 10 seconds or 5 seconds, etc. For example, σ>2 5 One calculation job per second.
[0143] It should be noted that the Poisson distribution arrival pattern has j short periods, where j is a positive integer. Each period contains a sudden surge of computational tasks, and the period lasts for a certain duration T. G For example, 5s to 10s, while maintaining a certain concurrency level σ. Among them, the arrival of computing jobs within the short period conforms to the fixed period arrival pattern.
[0144] It should also be noted that in this embodiment, the same computing task arrival mode or different computing task arrival modes can be used for computing jobs of different computing tasks. For multiple computing jobs of the same computing task, one computing task arrival mode or multiple computing task arrival modes can be used. For example, the computing task includes 20 computing jobs, of which 10 computing jobs use a single arrival mode and the other 10 computing jobs use an offline arrival mode.
[0145] The parameters corresponding to the above-mentioned computational task arrival modes, for example, for the fixed-period arrival mode, may include the period and the number of computational jobs arriving in each period; for the Poisson distribution arrival mode, the parameters may include the average number of computational jobs arriving per target unit time, the target unit time, etc.; for the peak arrival mode, the parameters may include the target period and σ, etc.
[0146] The training accuracy of the aforementioned AI model can be represented by the data type of the AI model parameters output, such as single-precision floating-point numbers or half-precision floating-point numbers.
[0147] The aforementioned lower limit of AI model performance may include at least one of the following: the lower limit of AI model training performance, and the lower limit of AI model inference performance.
[0148] For AI models, performance parameters typically differ across different scenarios. For example, performance parameters for image recognition and object detection include top-1 accuracy and average accuracy; performance parameters for semantic segmentation include mean intersection over union (MIOU); and performance parameters for speech recognition include word error rate (WER). For large models, the performance lower limit can be represented by the scores from benchmark tests of large models.
[0149] Optionally, the performance parameters of the computation task may further include a dataset indicator, which may correspond to the performance threshold of the AI model. The performance of the AI model corresponding to the dataset indicated by the dataset indicator must meet the lower performance threshold of the AI model. Commonly used datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, Criteo, etc.
[0150] The minimum throughput of the AI model inference mentioned above, for example, is usually the number of images inferred per second (images / s) for vision models, and the number of sentences inferred per second (sentences / s) or tokens / s for natural model models, where tokens refer to words, punctuation marks or other text units in the input text processed by the model.
[0151] The aforementioned computational power consumption threshold can, for example, be used to represent the maximum computational power consumption. For instance, computational power consumption can refer to the power consumption of the computing node within the aforementioned computational latency budget; or, computational power consumption can refer to P2-P1 or the ratio of P2 to P1, where P1 represents the power consumption of the computing node per unit time when powered on, and P2 represents the power consumption per unit time when processing computational tasks.
[0152] Computational energy efficiency thresholds can, for example, represent minimum computational energy efficiency. For instance, computational energy efficiency could be the computation speed per unit time divided by computational power consumption, or the throughput of an AI model inference divided by computational power consumption.
[0153] It should be noted that there can be conversion relationships between the different performance parameters of the above-mentioned computational tasks. That is, one set of performance parameters can be used to determine another set of performance parameters for the computational task. Therefore, the performance parameters of the above-mentioned computational tasks may only include a portion of the performance parameters listed above. Examples are provided below:
[0154] Example 1: The performance parameters of the above computational task include minimum computation speed. In this case, the first node can use the minimum computation speed as a selection parameter for the computation node.
[0155] Example 2: The performance parameters of the above computation task include the minimum number of operations or operands and the computation delay budget. If the computation delay budget only includes computation latency, then the first node can use these two parameters as selection parameters for the computation node; if the computation delay budget includes end-to-end latency, then the first node will use the data transmission latency and computation latency allocation information, along with these two parameters, as selection parameters for the computation node.
[0156] Example 3: The performance parameters of the above computing tasks include computing power type (e.g., CPU, clock speed, number of cores), memory, storage, and maximum time. The advantage of using these parameters as selection criteria for computing nodes is that they are relatively static and easy to choose.
[0157] Example 4: The performance parameters of the above computing task include minimum computing speed and minimum transmission bandwidth.
[0158] Example 5: The performance parameters of the above computational task include minimum computational intensity and computational delay budget.
[0159] It should also be noted that the performance parameters of the above-mentioned computational tasks may apply only to that specific computational task, meaning only that task needs to meet the above performance parameters; or they may apply to every computational task, meaning each task needs to meet the above performance parameters; or all computational tasks as a whole may meet the above performance parameters. The performance parameters of the above-mentioned computational task group may apply to each computational task within that group, meaning each task needs to meet the above performance parameters; or, all computational tasks in the group as a whole may meet the above performance parameters.
[0160] Optionally, one of the computational tasks is mapped to a Quality of Service (QoS) stream;
[0161] Alternatively, one of the computational tasks can be mapped to a set of QoS flows;
[0162] Alternatively, one of the computing tasks may be mapped to a Protocol Data Unit (PDU) session;
[0163] Alternatively, one of the computing tasks can be mapped to a set of PDU sessions;
[0164] Alternatively, one of the computing tasks can be mapped to a radio bearer (RB);
[0165] Alternatively, one of the computational tasks can be mapped to a set of RBs;
[0166] Alternatively, one of the computational tasks can be mapped to a logical channel (LC).
[0167] Alternatively, one of the computational tasks can be mapped to an LC set;
[0168] Alternatively, one of the computing tasks can be mapped to a physical layer resource;
[0169] Alternatively, a computing task may be mapped to a set of physical layer resources.
[0170] For example, a computing task may be mapped to a physical layer resource, such as a computing task mapped to a physical layer resource scheduled via Downlink Control Information (DCI).
[0171] A computing task is mapped to a set of physical layer resources. For example, a computing task is mapped to physical layer resources scheduled through multiple DCIs.
[0172] For example, when a compute task is mapped to a Quality of Service (QoS) flow, the performance parameters of the compute task correspond to the QoS flow. When a compute task is mapped to a set of QoS flows, the performance parameters of the compute task correspond to the QoS flow set. When a compute task is mapped to a PDU session, the performance parameters of the compute task correspond to the PDU session. When a compute task is mapped to a set of PDU sessions, the performance parameters of the compute task correspond to the PDU session set. When a compute task is mapped to a RB, the performance parameters of the compute task correspond to the RB. When a compute task is mapped to a set of RBs, the performance parameters of the compute task correspond to the RB set. When a compute task is mapped to an LC, the performance parameters of the compute task correspond to the LC. When a compute task is mapped to a set of LCs, the performance parameters of the compute task correspond to the LC set. When a compute task is mapped to a physical layer resource, the performance parameters of the compute task correspond to the physical layer resource. When a computing task is mapped to a set of physical layer resources, the performance parameters of the computing task correspond to the set of physical layer resources.
[0173] Optionally, the computing service request message may further include a computing service identifier for identifying the computing service.
[0174] For example, the aforementioned computing service identifier can be used to identify computing tasks such as one-dimensional DFT, two-dimensional FFT, AI model training, and AI model inference. Further, the aforementioned AI model training or AI model inference can include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, localization, perception, CSI feedback, etc.
[0175] It should be noted that when the computing service request message also includes a computing service identifier, the first node can select a computing node based on the computing service identifier and the performance parameters of the computing service. For example, the first node can select a computing node that supports the computing service identified by the computing service identifier and meets the performance parameters.
[0176] In this embodiment, the second node indicates the computing service it requests by carrying a computing service identifier in the computing service request message, which makes it easier for the second node to flexibly request different computing services.
[0177] Optionally, the method further includes:
[0178] The first node selects a computing node based on the computing service request message and the status information of at least one computing node.
[0179] The status information of the computing node includes at least one of the following: computing power type, computing load, available computing speed, available computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency.
[0180] In this embodiment, the first node selects a computing node based on the computing service request message and the status information of at least one computing node, which helps to more accurately select a computing node that meets the performance parameters of the computing service from the at least one computing node.
[0181] Optionally, the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
[0182] In this embodiment, the aforementioned computing task identifier is used to identify computing tasks. It should be noted that there can be a one-to-one mapping between computing services and computing tasks, or multiple computing services can be mapped to one computing task, or one computing service can be mapped to multiple computing tasks.
[0183] The aforementioned compute node identifier can be used to identify the selected compute node.
[0184] The aforementioned PDU session information can be used to determine the PDU session for data transmission related to the aforementioned computing services.
[0185] It is understood that, when the aforementioned computing service response message is used to indicate acceptance of the aforementioned computing service request message, the aforementioned computing service response message may include at least one of the following: computing task identifier, computing node identifier, and PDU session information.
[0186] Optionally, the PDU session information includes one of the following:
[0187] Instructions to establish a PDU session;
[0188] Modify PDU session indicator and PDU session identifier;
[0189] PDU session identifier, QoS flow identifier, and QoS rules.
[0190] The following examples illustrate different scenarios:
[0191] When a new PDU session needs to be established for data transmission related to the aforementioned computing services, the first node sends a PDU session establishment instruction to the second node. Upon receiving the PDU session establishment instruction, the second node can establish a PDU session, which can then be used for data transmission related to the aforementioned computing services.
[0192] When it is necessary to modify the PDU session for data transmission related to the aforementioned computing services, the first node sends a modification PDU session instruction and a PDU session identifier to the second node. Upon receiving the modification PDU session instruction and the PDU session identifier, the second node can modify the PDU session identified by the aforementioned PDU session identifier. The modified PDU session can then be used for data transmission related to the aforementioned computing services.
[0193] If a PDU session already exists in the multiplexing process to transmit data related to the aforementioned computing services, the first node sends a PDU session identifier, a QoS flow identifier, and QoS rules to the second node. Upon receiving the PDU session identifier, QoS flow identifier, and QoS rules, the second node can perform data transmission related to the aforementioned computing services based on the PDU session identifier, QoS flow identifier, and QoS rules.
[0194] Optionally, the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
[0195] In this embodiment, the aforementioned computing service request message may further include information indicating whether the requested computing node is a wireless access network node. Specifically, the indication of whether the requested computing node is a wireless access network node can be made explicitly or implicitly.
[0196] Understandably, when the compute service request message indicates that the requested compute node is a radio access network (RAN) node, the compute node selected by the first node must be a RAN node, or the first node should preferentially select a RAN node as the compute node. Selecting a RAN node as the compute node helps reduce compute service latency and better meets the needs of low-latency scenarios.
[0197] Optionally, the computing service request message includes at least one of the following:
[0198] The first indication information is used to indicate whether the requested computing node is a wireless access network node;
[0199] Latency type indicator, used to indicate the latency type of the computing service.
[0200] In one implementation, the first indication information can explicitly indicate whether the requested computing node is a wireless access network node. For example, a single bit can be used to indicate whether the requested computing node is a wireless access network node, where 0 indicates that the requested computing node is not a wireless access network node, and 1 indicates that the requested computing node is a wireless access network node.
[0201] In another embodiment, the latency type can be used to indicate whether the requested computing node is a radio access network (RAN) node. For example, if the latency type indicated by the latency type indicator is a preset latency type, it means that the requested computing node is a RAN node; otherwise, it means that the requested computing node is not a RAN node. For example, the preset latency type can be a low latency type.
[0202] Optionally, if the computing service is a preset type of computing service or the resource type of the computing service is delay-critical, the selected computing node is a wireless access network node.
[0203] or,
[0204] If the first indication information indicates that the requested computing node is a wireless access network node, the selected computing node is a wireless access network node.
[0205] or,
[0206] When the latency type indicated by the latency type indicator is a preset latency type, the selected computing node is a wireless access network node.
[0207] For example, the aforementioned preset type of computing service may include computing services for mobile network optimization, such as beam management, positioning, sensing, CSI feedback, etc. The aforementioned latency-sensitive type may include, but is not limited to, at least one of latency-sensitive guaranteed computing speed, latency-sensitive guaranteed computing strength, etc.
[0208] Optionally, the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
[0209] The representation of computing tasks and the identifier of computing nodes in this embodiment can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0210] It is understood that when the selected computing node is a radio access network node, the aforementioned computing service response message may include second indication information to indicate that the selected computing node is a radio access network node; otherwise, the aforementioned computing service response message does not include the second indication information.
[0211] Optionally, when the selected computing node is a radio access network node, the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
[0212] In this embodiment, the aforementioned radio bearer indication is used to indicate a radio bearer (RB). For example, the radio bearer indication may include a radio bearer identifier (ID). The aforementioned radio bearer may include a signal radio bearer (SRB), a data radio bearer (DRB), or a newly defined RB, etc., and the radio bearer can be used for the transmission of data related to computing services.
[0213] It should be noted that if it is necessary to create or modify an existing radio bearer, the first node can trigger the communication function in the radio access network node to send a radio bearer add or modify message.
[0214] The aforementioned physical layer channel indicator can be used to indicate a physical layer channel that can be used for data transmission related to computing services. The aforementioned physical layer resource indicator is used to indicate physical layer resources that can be used for data transmission related to computing services.
[0215] It should be noted that if computing service-related data, such as at least one of computing data and computing response data, is transmitted through a physical channel, it is beneficial to further reduce the transmission latency of computing service-related data.
[0216] Optionally, the radio bearer indication is used to indicate at least one of the following: signaling radio bearer (SRB), data radio bearer (DRB), and data plane (RB).
[0217] Optionally, when the selected computing node is a wireless access network node, the method further includes:
[0218] The first node sends first information to the radio access network node, the first information including third indication information, the third indication information being used to instruct the radio access network node to add or modify a radio bearer.
[0219] In this embodiment, when the selected computing node is a radio access network node, the first node sends first information to the radio access network node, and then the radio access network node can send a Radio Resource Control (RRC) reconfiguration message to the second node based on the first information to request the establishment or modification of radio bearers or the configuration of physical layer resources for computing services.
[0220] Optionally, the first information may also include performance parameters of the computing service.
[0221] For example, by including the performance parameters of the computing service in the first information, the wireless access network node can learn about the performance parameters of the computing service, and then the wireless access network node can provide computing services based on the performance parameters of the computing service, which helps to further ensure the quality requirements of the computing service.
[0222] Optionally, the method further includes:
[0223] The first node sends a first request message to the selected computing node, the first request message being used to request the creation or modification of a computing task;
[0224] The first request message includes at least one of the following:
[0225] Computation task identifier; priority indicator; preemption capability indicator; preemption capability indicator; migration capability indicator; guaranteed computation speed, used to indicate the computation speed that a computing node guarantees to provide to a computing task within an average window; guaranteed computation intensity, used to indicate the computation intensity that a computing node guarantees to provide to a computing task within an average window; maximum computation speed, used to indicate the upper limit of the maximum computation speed that a computing node can provide to a computing task; maximum computation intensity, used to indicate the upper limit of the maximum computation intensity that a computing node can provide to a computing task.
[0226] In this embodiment, the priority indicator is used to represent the importance of the computing resource request of the computing task.
[0227] The aforementioned preemption capability indicator is used to indicate whether a computing task can obtain computing resources that have been allocated to another computing task with lower priority.
[0228] The preemption capability indicator is used to indicate whether to drop a computational task in order to execute a computational task with a higher priority.
[0229] Migration capability indicators are used to indicate whether the migration of computing tasks from one computing node to another is supported.
[0230] Guaranteed computation speed is used to indicate the computation speed that a computing node guarantees to provide to the computing task within the average window.
[0231] Guaranteed computational intensity is used to indicate the computational intensity that a computing node guarantees to provide to the computing task within an average window.
[0232] In some optional embodiments, the aforementioned guarantee of computational speed may include a latency-sensitive guarantee of computational speed. The aforementioned guarantee of computational strength may include a latency-sensitive guarantee of computational strength.
[0233] Optionally, the method further includes:
[0234] The first node receives a first response message from the selected computing node;
[0235] The first response message includes status information of the selected computing node after it has established or modified the computing task.
[0236] In this embodiment, the status information of the computing node after establishing or modifying the computing task may include, but is not limited to, at least one of the following: computing load, computing speed, computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency.
[0237] For example, when the first node receives the first response message, it can determine whether the selected computing node can meet the performance parameters of the computing service based on the status information after the computing task is established or modified by the selected computing node. If it does, the first node can send a computing service response message to the second node, indicating that it accepts the computing service request message. If it does not meet the requirements, the first node can reselect a computing node, which helps to further ensure the service quality of the computing service.
[0238] Please refer to Figure 3, which is a flowchart of a computing service method provided in an embodiment of this application. This method can be executed by a second node, and as shown in Figure 3, it includes the following steps:
[0239] Step 301: The second node sends a computing service request message to the first node, the computing service request message including the performance parameters of the computing service;
[0240] Step 302: The second node receives a computing service response message from the first node, the computing service response message being used to indicate whether to accept or reject the computing service request message.
[0241] Optionally, the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
[0242] Resource type; minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, indicating the upper limit of the computing task's duration; computing power type; data type; minimum memory, indicating the lower limit of memory required for the computing task; minimum storage, indicating the lower limit of storage required for the computing task; minimum transmission bandwidth, indicating the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
[0243] Optionally, the computing service request message may also include a computing service identifier.
[0244] Optionally, the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
[0245] Optionally, the PDU session information includes one of the following:
[0246] Instructions to establish a PDU session;
[0247] Modify PDU session indicator and PDU session identifier;
[0248] PDU session identifier, QoS flow identifier, and QoS rules.
[0249] Specifically, when the PDU session information includes a PDU session establishment indication, the second node can send a PDU establishment request message to the third node and receive a PDU establishment response message returned by the third node; when the PDU session information includes a PDU session modification indication and a PDU session identifier, the second node can send a PDU modification request message to the third node and receive a PDU modification response message returned by the third node. For example, the third node may include at least one of an AMF and an SMF.
[0250] It should be noted that the specific procedures for PDU session establishment and PDU session modification in this embodiment can be found in relevant technologies, and will not be elaborated here.
[0251] Optionally, the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
[0252] Optionally, the computing service request message includes at least one of the following:
[0253] The first indication information is used to indicate whether the requested computing node is a wireless access network node;
[0254] Latency type indicator, used to indicate the latency type of the computing service.
[0255] Optionally, the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
[0256] Optionally, the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
[0257] Optionally, when the selected computing node is a wireless access network node, the method further includes:
[0258] The second node receives a Radio Resource Control (RRC) reconfiguration message sent by a Radio Access Network (RAN) node. The RRC reconfiguration message includes a target configuration for computing services. The target configuration includes at least one of the following: radio bearer add configuration, radio bearer modify configuration, and physical layer resource configuration.
[0259] The second node sends an RRC reconfiguration complete message to the radio access network node.
[0260] For example, the aforementioned wireless access network node may include a base station or a centralized unit (CU), etc.
[0261] In this embodiment, when the second node receives the target configuration for computing services, it can configure computing services based on the target configuration. For example, if the target configuration includes a wireless bearer addition configuration, a wireless bearer can be added based on the wireless bearer addition configuration, and the added wireless bearer can be used for data transmission related to computing services. If the target configuration includes a wireless bearer modification configuration, the wireless bearer can be modified based on the wireless bearer modification configuration, and the modified wireless bearer can be used for data transmission related to computing services. If the target configuration includes a physical layer resource configuration, data related to computing services can be transmitted based on the physical layer resources configured in the physical layer resource configuration.
[0262] It should be noted that transmitting computing service-related data via wireless bearer or physical layer can help reduce the transmission latency of computing services.
[0263] Optionally, the method further includes:
[0264] The second node sends a second request message to the third node. The second request message is used to request the establishment or modification of a PDU session. The second request message includes fourth indication information, which is used to indicate that the PDU session supports at least one QoS flow for computing services.
[0265] The second node receives a second response message from the third node, the second response message indicating acceptance of the second request message.
[0266] The aforementioned third node may include at least one of AMF and SMF.
[0267] The aforementioned at least one QoS flow used for computing services can be understood as satisfying the QoS parameters for computing services, or satisfying the QoS parameters for both computing and communication services.
[0268] For example, upon receiving the second request message, the third node may perform at least one of the following:
[0269] A fourth node, such as UPF, is selected based on the second request message, and a third request message is sent to the selected fourth node. The third request message is used to request the establishment or modification of an N4 session. The third request message includes third information for computing services or for computing services and communication services. The third information includes at least one of the following: rule ID, priority, packet detection information, forwarding rule, enforcement rule, and reporting rule.
[0270] Send an N2 Session Management (SM) message to a radio access network node (e.g., a base station), the N2 Session Management message including at least one of a target QoSI and a first Quality of Service Profile (QoS Profile), the first QoS Profile being used for computing services, or the first QoS Profile being used for both computing services and communication services;
[0271] The third node determines the computing nodes, for example, by selecting the computing nodes itself or by obtaining computing nodes from the first node. Optionally, the number of computing nodes determined may be more than one, depending on the number of packet filters and packet filter information provided by the terminal's computing service.
[0272] The rule identifier is used to uniquely identify the rule for the aforementioned computing service. The priority is used to determine the order in which detection information for all rules applying to the computing services is processed.
[0273] The aforementioned packet detection information includes at least one of the following: terminal Internet Protocol (IP) address, core network (CN) tunnel information, packet filter set, and target QoSI.
[0274] The packet filter mentioned above may include IP addresses, port numbers, etc., and is used to detect which data packets belong to the computing service flow. The target QoSI mentioned above may include a first QoSI or a second QoSI. The first QoSI represents at least one QoS parameter for the computing service, and the second QoSI represents at least one QoS parameter for both the computing and communication services. The first QoSI may represent at least one of the following: a first resource type, a first priority level, a first computing latency budget, a failure rate, an average window, a maximum number of operations or operands, a maximum computing speed, and a computing power type. The second QoSI may represent at least one of the following: a second resource type, a second priority level, a second computing latency budget, and bandwidth. It should be noted that the first QoSI mentioned above can also be called the computing QoSI, and the second QoSI mentioned above can also be called the computing and communication QoSI.
[0275] The aforementioned forwarding rules include a forwarding rule identifier and a definition of forwarding the corresponding data packet to the aforementioned selected computing node (such as the computing node's IP address).
[0276] The aforementioned execution rules include an execution rule identifier and a definition of the QoS operation to be executed. For example, for delay-sensitive GCI or delay-sensitive GCR, the corresponding 5QI type is delay-critical GBR.
[0277] The aforementioned reporting rules include a reporting rule identifier and a definition of the measurement operations to be performed. For example, measuring the latency, throughput, and data volume (throughput multiplied by time) from the UPF to the compute node.
[0278] It should be noted that the embodiments of this application do not limit the order in which the second node sends the second request message to the third node and the second node sends the computing service request message to the first node. For example, the second node may first send the second request message to the third node and then send the computing service request message to the first node; or, the terminal may first send the computing service request message to the first node and then send the second request message to the third node.
[0279] Optionally, the third node can also receive N4 session establishment or modification response messages from the fourth node.
[0280] In this embodiment, since the PDU session supports at least one QoS stream for computing services, transmitting computing service-related data based on the PDU session can further guarantee the quality of service for computing services.
[0281] Optionally, the second response message includes a first QoS rule, which is used for computing services or for both computing services and communication services.
[0282] For example, the third node can send the first QoS rule through the N1 SM container.
[0283] Optionally, the first QoS rule includes at least one of the following:
[0284] The target quality identifier (QoSI) for the QoS flow of the computing service includes a first QoSI or a second QoSI, wherein the first QoSI is used to represent at least one QoS parameter for the computing service, and the second QoSI is used to represent at least one QoS parameter for both the computing service and the communication service.
[0285] Packet filter set;
[0286] A priority indicator, which is used to indicate the priority of the first QoS rule.
[0287] For example, the packet filter described above may include IP addresses, port numbers, etc., to detect which packets belong to the computation service flow. The priority indicator described above can be used to determine the order in which packets are matched against multiple QoS rules.
[0288] Optionally, the first QoSI is used to represent at least one of the following: first resource type, first priority level, first computation latency budget, failure rate, average window, maximum number of operations or operands, maximum computation speed, and computing power type.
[0289] The aforementioned first resource type can also be referred to as a computing resource type. For example, the aforementioned first resource type may include at least one of the following: guaranteed computing speed, non-guaranteed computing speed, latency-sensitive guaranteed computing speed, guaranteed computing intensity, non-guaranteed computing intensity, and latency-sensitive guaranteed computing intensity.
[0290] In some optional embodiments, the first resource type determines the allocation of computing resources related to the QoS flow-level guaranteed computational load. Alternatively, the first resource type determines the allocation of computing resources related to the QoS flow-level guaranteed computational intensity. Here, one computational task may be mapped to one QoS flow, or one computational task may be mapped to multiple QoS flows, or multiple computational tasks may be mapped to one QoS flow.
[0291] The computational speed and computational intensity can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here. Optionally, the computational intensity in this embodiment can be a segmented computational intensity, for example, the computational speed divided by the memory bandwidth.
[0292] The first priority level mentioned above is used to determine the priority of computing resource scheduling for computing QoS streams.
[0293] The aforementioned first computational latency budget can be used to indicate the upper limit of the latency (i.e., the maximum computational latency) that can be tolerated when the computational task of the QoS flow (also referred to as the computational packet or computational packet set) is computed at the computing node.
[0294] For example, one definition of computation latency is the length of the time interval between the first data packet of a single computation task being sent and the last data packet of the computation task being received; another definition is the length of the time interval between the first data packet of a group of computation tasks being sent and the last data packet being received. For example, for image recognition computation tasks, one approach is to treat single image recognition as a single computation task, while another approach is to treat multiple images (e.g., 100 images) as a group of computation tasks. For details, please refer to the foregoing embodiments regarding the relevant explanation of computation latency budgeting.
[0295] For example, computational latency for AI model inference can include at least one of the following:
[0296] Total inference latency: Specifically, it refers to the total latency of multiple consecutive inference operations. The calculation method is as follows: the time before inference for the first computational task (or computational job) is denoted as T. ITS The time when all computational tasks (or computational jobs) finish inference is denoted as T. ITE Then the computational latency budget for AI model inference is T. ITE -T ITS .
[0297] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The time when the inference of this computational task ends is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0298] The failure rate and average window mentioned above can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0299] It should be noted that the average window within the first QoSI mentioned above applies only to GCR or GCI or delay critical GCR or delay critical GCI.
[0300] The maximum number of operands or operations mentioned above is used to represent the maximum number of operands or operations for a QoS flow. The maximum number of operands or operations within the first QoSI mentioned above applies only to GCRs or delay-critical GCRs, such as 2 TFLOPs.
[0301] The aforementioned maximum computation speed is used to indicate the maximum computation speed of the QoS flow. The maximum computation speed within the first QoSI applies only to GCI or delay-critical GCI. Optionally, the aforementioned maximum computation speed may include at least one of the theoretical maximum computation speed and the actual maximum computation speed.
[0302] The theoretical maximum computing speed included in the aforementioned maximum computing speed is used to indicate that the theoretical maximum computing speed of the computing node required by QoS flows is not higher than the theoretical maximum computing speed included in the aforementioned maximum computing speed. The theoretical peak operand count is calculated as processor clock speed × number of operations performed per clock cycle × total number of system cores. For example, a GPU with a clock speed of 840MHz, 256 computational logic units (meaning the processor performs 256 floating-point operations per clock cycle), and 3 cores has a theoretical peak operand count of 840MHz × 256 × 3 = 645.120G FLOPS (FP32), or 840MHz × 256 × 3 × 2 = 1290.24G FLOPS (FP16). Therefore, one method of representing the theoretical maximum computing speed of a computing node is through processor clock speed, number of cores, etc., while another method is to represent it through FLOPS or Operations Per Second (OPS).
[0303] The aforementioned maximum computing speed includes the actual maximum computing speed, indicating that the actual maximum computing speed of the computing node required by the QoS flow is not higher than the actual maximum computing speed included in the aforementioned maximum computing speed. The actual maximum computing speed is the maximum computing speed obtained through testing.
[0304] In some optional embodiments, the first QoSI described above can also be used to characterize a second test case indication, which corresponds to the actual maximum computing speed and is used to indicate that the actual maximum computing speed is the maximum computing speed obtained based on the test case indicated by the second test case indication.
[0305] For example, test cases can be at least one of the following: MobileNetVx (x can be any available version number, such as 1), one-dimensional DFT, one-dimensional FFT, two-dimensional FFT, matrix multiplication, sparse linear equations, dense linear equations, YOLOvy (y can be any available version number, such as YOLOv5), image recognition models (such as resnet50_v1.5), and large models (such as Llama3, Llama2). Test cases can be open-source software programs or custom software programs between the UE and the network. This means that the actual maximum computation speed required for QoS flows is the actual maximum computation speed that the computing node can achieve under the indicated test case conditions.
[0306] In some optional embodiments, the aforementioned actual maximum computing speed can be expressed as the theoretical maximum computing speed and computing efficiency, or as the ideal maximum computing speed and computing efficiency. The ideal maximum computing speed can be understood as the maximum computing speed obtained under ideal conditions such as no task preemption, based on test cases.
[0307] One definition of computational efficiency is the ratio of the actual maximum computational speed to the theoretical maximum computational speed under ideal conditions. In this case, the actual maximum computational speed required for the computational task is represented by the theoretical maximum computational speed and the computational efficiency.
[0308] Another definition of the aforementioned computational efficiency is the ratio of the maximum computational speed measured based on test cases to the theoretical maximum computational speed. In this case, the first QoSI is also used to characterize the second test case indication, and the actual maximum computational speed required for the QoS flow can be represented by the theoretical maximum computational speed and computational efficiency.
[0309] Another definition of the aforementioned computational efficiency is: the ratio of the computational speed measured based on test cases to the maximum computational speed under ideal conditions. The computational speed measured based on test cases is usually obtained from recent tests and can represent the current state of the computing node. The maximum computational speed under ideal conditions, measured based on test cases, refers to the best measured performance of the computing node. In this case, the aforementioned first QoSI is also used to characterize the second test case indication, and the actual maximum computational speed required for the QoS flow can be represented by the ideal maximum computational speed and computational efficiency.
[0310] The above-mentioned computing power types can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0311] For example, the first QoSI described above can be used to mark the index value of the forwarding processing parameters of the QoS flow of the computing service.
[0312] It should also be noted that the quality parameters represented by the first QoSI in this embodiment (i.e., at least one of the above-mentioned first resource type, first priority level, first computation delay budget, failure rate, average window, maximum number of operations or operands, maximum computation speed and computing power type) correspond to QoS flow, that is, the above-mentioned first QoSI is a QoS flow-level quality of service parameter.
[0313] Optionally, the second QoSI is used to represent at least one of the following: a second resource type, a second priority level, a second computational delay budget, and bandwidth.
[0314] The aforementioned second resource type can also be referred to as a computing and communication resource type. For example, the aforementioned second resource type may include at least one of the following: guaranteed computing speed, non-guaranteed computing speed, latency-sensitive guaranteed computing speed, guaranteed computing strength, non-guaranteed computing strength, and latency-sensitive guaranteed computing strength.
[0315] In some optional embodiments, the second resource type described above may determine the allocation of computational and communication resources related to the QoS flow-level guaranteed computational load, or the second resource type may determine the allocation of computational and communication resources related to the QoS flow-level guaranteed computational intensity. Here, one computational task may be mapped to one QoS flow, or one computational task may be mapped to multiple QoS flows, or multiple computational tasks may be mapped to one QoS flow.
[0316] The computational speed and computational intensity can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0317] Optionally, in this embodiment, the smallest of several bandwidth components, such as the transmission bandwidth between the second node (e.g., UE) and the computing node, and the memory bandwidth of the computing node, can be used as the bandwidth for computational intensity. That is, the computational intensity is the computational speed divided by min{memory bandwidth, transmission bandwidth between the second node and the computing node}, where min{memory bandwidth, transmission bandwidth between the second node and the computing node} represents the smaller of the transmission bandwidth and the memory bandwidth. In some optional embodiments, the transmission bandwidth between the second node and the computing node can be further divided into the air interface bandwidth between the second node and the access network node, and the wired transmission bandwidth between the access network node and the computing node; or the transmission bandwidth between the second node and the computing node can be further divided into the bandwidth between the second node and the UPF, and the wired transmission bandwidth between the UPF and the computing node.
[0318] The aforementioned second priority level is used to determine the priority of computation and communication resource scheduling for QoS flows used for computation services.
[0319] The aforementioned second computational delay budget represents the upper limit of the tolerable latency for the computation task of the QoS flow (also referred to as computation data packets or computation data packet sets) during computation and transmission. The tolerable latency during computation and transmission is the sum of the computation latency and the transmission latency. The transmission latency can refer to the sum of the transmission latency from the second node to the computation node and the transmission latency from the computation node to the computation receiving node.
[0320] For example, one definition of computation and transmission latency is the length of the time interval between the first data packet of a single computation task being sent and the last data packet of the computation task being received; another definition is the length of the time interval between the first data packet of a group of computation tasks being sent and the last data packet being received. For example, for image recognition computation tasks, one approach is to treat single image recognition as a single computation task, while another approach is to treat multiple images (e.g., 100 images) as a group of computation tasks. See the foregoing embodiments for details regarding computation latency budgeting.
[0321] For example, computation and transmission latency for AI model inference can include at least one of the following:
[0322] Total end-to-end inference latency: Specifically, this refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time elapsed before sending the first byte of the first computation task (or job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS .
[0323] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time t is the time before the second node sends the first byte of a computation task (or job). TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS .
[0324] The aforementioned bandwidth is used to represent at least one of the uplink bandwidth lower limit and downlink bandwidth lower limit of the QoS flow. For details regarding the bandwidth in this embodiment, please refer to the relevant descriptions in the foregoing embodiments; they will not be repeated here.
[0325] For example, the second QoSI described above can be used to mark the index value of the forwarding processing parameters of the QoS flow of computing and communication services.
[0326] It should be noted that the quality parameters represented by the second QoSI in this embodiment (i.e., at least one of the second resource type, the second priority level, the second computational delay budget, and the bandwidth) correspond to a QoS flow, that is, the above-mentioned second QoSI is a QoS flow-level quality of service parameter.
[0327] Optionally, the second request message further includes at least one of the following:
[0328] The number of packet filters used to calculate the number of packet filters for a service.
[0329] Optionally, the second response message includes a computing node identifier.
[0330] For example, the aforementioned computing node identifier may include an IP address or an internal network ID, etc.
[0331] Optionally, the method further includes:
[0332] The second node sends a second message to the computing node, the second message including computing data.
[0333] Optionally, the second information further includes at least one of the following:
[0334] Compute node identifier;
[0335] A receiving node indication is used to indicate the node that receives the computation response corresponding to the computation data.
[0336] It should be noted that the implementation method of this method can be found in the relevant description of the embodiment shown in Figure 2, and will not be repeated here.
[0337] The following examples illustrate this embodiment:
[0338] Example 1: The main idea of this example is to solve the problems of performance parameter identification and interaction when a mobile network provides computing services, and how to ensure the quality of service of computing services, based on computing service request messages containing computing service parameters. Furthermore, in this embodiment, the computing task is mapped to a PDU session and QoS flow; the first node can be a core network node, and the computing node can be a core network node or an edge computing node, etc.
[0339] For example, referring to Figure 4, the computing service method provided in this application embodiment includes the following steps:
[0340] Step 11: The second node sends a computing service request message to the first node.
[0341] The aforementioned computing service request message may include performance parameters of the computing service. For details on the performance parameters of the computing service, please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.
[0342] Step 12: The first node sends a computing task creation / modification request message to the computing node.
[0343] In the node-to-computing service request message, the first node can select a suitable computing node based on the computing service request message and the status information of at least one computing node. The status information of the computing nodes can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0344] Optionally, if a suitable computing node is selected, the first node can send a task creation / modification request message to the selected computing node. The description of the aforementioned task creation / modification request message can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0345] Step 13: The compute node sends a compute task creation / modification response message to the first node.
[0346] The aforementioned computation task creation / modification response messages can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0347] It should be noted that steps 12 and 13 above can be optional steps. For example, existing computing tasks can be used for calculation and processing.
[0348] Step 14: The first node sends a computing service response message to the second node.
[0349] In this step, the aforementioned computing service response message can be used to indicate whether or not to accept the aforementioned computing service request message. For details, please refer to the relevant description in the foregoing embodiments, which will not be repeated here.
[0350] It should be noted that Figure 4 shows the case where the above computing service response message indicates acceptance of the above computing service request message.
[0351] Step 15: The second node sends a PDU session establishment / modification request message to the third node.
[0352] Step 16: The third node sends a PDU session establishment / modification response message to the second node.
[0353] Steps 15 and 16 above can be found in the process of establishing or modifying a PDU session in related computing scenarios, and will not be elaborated here. It should be noted that steps 15 and 16 above can be optional steps. For example, existing PDU sessions can be used to transmit data related to computing services.
[0354] Step 17: The second node sends computation data to the computing node.
[0355] For example, the second node can transmit the aforementioned computational data based on a PDU session.
[0356] Step 18a: The computing node sends computing response data to the second node.
[0357] Step 18b: The computing node sends computing response data to the computing receiving node.
[0358] Understandably, in this case, the receiving node and the second node are different nodes.
[0359] Example 2: The main idea of this example is to map computing tasks to radio bearers or physical layer resources, rather than PDU sessions or QoS flows. Furthermore, the first node can be a core network node or a radio access network node, and the computing node is a radio access network node. The computing service method provided in this example can solve the problems of performance parameter identification and interaction when mobile networks provide computing services, as well as how to ensure the quality of service of computing services, especially for low-latency scenarios or scenarios where the radio access network node is a trusted node.
[0360] For example, referring to Figure 5, the computing service method provided in this application embodiment includes the following steps:
[0361] Step 21: The second node sends a computing service request message to the first node.
[0362] The aforementioned computing service request message may include performance parameters of the computing service. Specific details regarding these performance parameters can be found in the descriptions of the foregoing embodiments and will not be repeated here. Optionally, the aforementioned computing service request message may also indicate whether the requested computing node is a radio access network node.
[0363] Step 22: The first node sends a computing task creation / modification request message to the wireless access network node.
[0364] In the node-to-computing service request message, the first node can select a suitable computing node based on the computing service request message and the status information of at least one computing node, wherein the computing node is a radio access network node. The status information of the computing node can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0365] Optionally, if a suitable radio access network node is selected, the first node may send a task creation / modification request message to the selected radio access network node. The aforementioned task creation / modification request message can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0366] Step 23: The wireless access network node sends a computing task creation / modification response message to the first node.
[0367] The aforementioned computation task creation / modification response messages can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0368] It should be noted that steps 22 and 23 above can be optional steps. For example, existing computing tasks can be used for calculation and processing.
[0369] Step 24: The first node sends a computing service response message to the second node.
[0370] In this step, the aforementioned computing service response message can be used to indicate whether or not to accept the aforementioned computing service request message. For details, please refer to the relevant description in the foregoing embodiments, which will not be repeated here.
[0371] Optionally, in this example, the computing service response message may further include at least one of the following: a computing node is a radio access network node indication (i.e., second indication information), a radio bearer indication, a physical layer channel indication, and a physical layer resource indication.
[0372] It should be noted that Figure 5 shows the case where the above computing service response message indicates acceptance of the above computing service request message.
[0373] Step 25: The wireless access network node sends an RRC reconfiguration message to the second node.
[0374] For example, the RRC reconfiguration message may include a radio bearer add / modify request.
[0375] The RRC reconfiguration message mentioned above can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0376] Step 26: The second node sends an RRC reconfiguration complete message to the radio access network node.
[0377] Step 27: The second node sends computation data to the computing node.
[0378] For example, the second node can transmit the aforementioned computational data based on a PDU session.
[0379] Step 28a: The computing node sends computing response data to the second node.
[0380] Step 28b: The computing node sends computing response data to the computing receiving node.
[0381] Understandably, in this case, the receiving node and the second node are different nodes.
[0382] Example 3: The main idea of this example is to define the service quality parameters related to computing services as the first QoSI, and the service quality parameters related to computing and communication as the second QoSI. The flexible parameters and the precise values for each computing task require interaction through the computing service process of each computing task. It should be noted that the meanings of the first QoSI and the second QoSI can be found in the relevant descriptions of the foregoing embodiments, and will not be repeated here.
[0383] For example, referring to Figure 6, the computing service method provided in this application embodiment includes the following steps:
[0384] Step 31: The UE sends a PDU session establishment or modification request message to the third node. The PDU session establishment or modification request message contains fourth indication information, which indicates that the PDU session contains at least one QoS flow for computing services, that is, the PDU session supports at least one QoS flow for computing services.
[0385] Step 32: The third node sends an N4 session establishment or modification request message to the fourth node.
[0386] Specifically, to simplify the process, the third node corresponds to the AMF and SMF. The AMF receives the PDU session establishment or modification request message from the UE and selects the appropriate SMF based on information such as whether QoS flows for computing services are needed. The SMF selects the appropriate fourth node (e.g., UPF) based on the PDU session establishment or modification request message. Furthermore, the SMF can also select the appropriate computing node or obtain the appropriate computing node from the computing management node.
[0387] The N4 session establishment or modification request message includes packet detection, execution, and reporting rules for the computing service, etc. For details, please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.
[0388] Step 33: The fourth node sends an N4 session establishment or modification response message to the third node.
[0389] For example, the UPF sends an N4 session establishment message or an N4 session modification response message to the SMF.
[0390] Step 34: The third node sends N2 session management information to the radio access network node. The N2 session management information may include at least one of the following: target QoSI and calculated QoS profile.
[0391] The target QoSI can be used to mark the index value of the forwarding processing parameters of the QoS flow of the computing service. For example, the third node maps the QoS flow corresponding to the computing service to one data radio bearer, and maps the QoS flows of other non-computing services to another data radio bearer.
[0392] Step 35: The third node sends the first QoS rule and the number of node identifiers to the UE through the N1 session management container.
[0393] Step 36: The UE sends a computing service request message to the first node.
[0394] This step is the same as step 11 above, and will not be repeated here.
[0395] Step 37: The first node sends a computing task creation / modification request message to the computing node.
[0396] This step is the same as step 12 above, and will not be repeated here.
[0397] Step 38: The compute node sends a compute task creation / modification response message to the first node.
[0398] This step is the same as step 13 above, and will not be repeated here.
[0399] Step 39: The first node sends a computing service response message to the UE.
[0400] This step is the same as step 14 above, and will not be repeated here.
[0401] Step 40: The UE sends computation data to the computing node.
[0402] This step is the same as step 17 above, and will not be repeated here.
[0403] Step 41a: The computing node sends computing response data to the UE.
[0404] This step is the same as step 18a above, and will not be repeated here.
[0405] Step 41b: The computing node sends computing response data to the computing receiving node.
[0406] This step is the same as step 18b above, and will not be repeated here.
[0407] It should be noted that this example does not limit the execution order of steps 31 to 35 and steps 36 to 39. For example, steps 31 to 35 can be executed first, followed by steps 36 to 39; or steps 36 to 39 can be executed first, followed by steps 31 to 35.
[0408] The computing service method provided in this application transmits performance parameters (e.g., performance minimum requirements) corresponding to the computing service based on the computing service request message. This solves the problems of defining, identifying, transmitting, and using relevant parameters (especially performance parameters) when a mobile network provides computing services, thereby addressing the issue of ensuring the quality of service for computing services based on demand. This method is applicable to providing computing services to both AFs (Automatic Front-End) and UEs (User Equipment) and NFs (Network Functions). It has broader applicability, suitable for core network nodes acting as both computing management nodes and computing nodes, core network nodes acting as both computing management nodes and radio access network nodes acting as computing nodes, and radio access network nodes acting as both computing management nodes and computing nodes.
[0409] It should be noted that the computing service method provided in this application embodiment can be executed by a computing service device. This application embodiment uses the execution of the computing service method by a computing service device as an example to illustrate the computing service device provided in this application embodiment.
[0410] This application provides a computing service device. As an example, the computing service device may be a communication device or a component within a communication device, such as a chip. The communication device may be a terminal, a network-side device, or a server, etc. Exemplarily, the terminal may include, but is not limited to, the type of terminal 11 listed above, and the network-side device may include, but is not limited to, the type of network-side device 12 listed above. This application does not impose specific limitations.
[0411] The computing service device includes a receiving module, a transmitting module, and a processing module. These modules can be implemented in software or hardware. When implemented in hardware, the processing module can be implemented by a processor. For example, the processor can include general-purpose processors, special-purpose processors, such as a Central Processing Unit (CPU), microprocessor, Digital Signal Processor (DSP), Artificial Intelligence (AI) processor, Graphics Processing Unit (GPU), Application Specific Integrated Circuit (ASIC), Network Processor (NP), Field Programmable Gate Array (FPGA), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The receiving and transmitting modules can be implemented by a communication interface, which can include one or more of the following: transceiver, pins, circuits, bus, radio frequency unit, etc.
[0412] Specifically, referring to Figure 7, when the computing service device is a network-side device or a component of a network-side device, the computing service device 700 includes a receiving module 701, used to receive a computing service request message from a second node, the computing service request message including performance parameters of the computing service; and a sending module 702, used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
[0413] Optionally, the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
[0414] Resource type; minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, representing the upper limit of the computing task's duration; computing power type; data type; minimum memory, representing the lower limit of memory required for the computing task; minimum storage, representing the lower limit of storage required for the computing task; minimum transmission bandwidth, representing the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
[0415] Optionally, the resource type includes at least one of the following:
[0416] Guarantee computation speed, guarantee computation intensity, do not guarantee computation speed, do not guarantee computation intensity, latency-sensitive guarantee computation speed, latency-sensitive guarantee computation intensity.
[0417] Optionally, the minimum computing speed includes at least one of the following: theoretical minimum computing speed, and actual minimum computing speed.
[0418] Optionally, the computation delay budget includes at least one of the following: an upper limit for computation delay, an upper limit for transmission delay, and an upper limit for both computation and transmission delay.
[0419] Optionally, the minimum transmission bandwidth includes at least one of the following: minimum uplink bandwidth, minimum downlink bandwidth, uplink bandwidth indication, downlink bandwidth indication, and target indication;
[0420] The target indication is used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different.
[0421] Optionally, the computing task includes at least two computing jobs, and the computing task arrival mode includes at least one of the following:
[0422] Continuous arrival mode or single arrival mode is used to indicate that one computation job arrives at a time;
[0423] Fixed-cycle arrival mode is used to indicate the arrival of control calculation jobs according to a fixed cycle;
[0424] Poisson distribution arrival pattern, used to indicate the arrival of computational jobs controlled by Poisson distribution;
[0425] Peak arrival mode is used to indicate the arrival of σ computational jobs within a target period of Poisson distribution, where the duration of the target period is less than a preset duration, and σ is a positive integer.
[0426] Offline arrival mode, used to indicate that all computation jobs arrive at once.
[0427] Optionally, the lower limit of AI model performance includes at least one of the following: the lower limit of AI model training performance, and the lower limit of AI model inference performance.
[0428] Optionally, the performance parameters of the computing task may also include a dataset indicator, wherein the performance of the AI model corresponding to the dataset indicated by the dataset indicator must meet the lower limit of the AI model performance.
[0429] Optionally, one of the computational tasks is mapped to a Quality of Service (QoS) stream;
[0430] Alternatively, one of the computational tasks can be mapped to a set of QoS flows;
[0431] Alternatively, one of the computing tasks can be mapped to a Protocol Data Unit (PDU) session;
[0432] Alternatively, one of the computing tasks can be mapped to a set of PDU sessions;
[0433] Alternatively, one of the computing tasks can be mapped to a radio bearer (RB).
[0434] Alternatively, one of the computational tasks can be mapped to a set of RBs;
[0435] Alternatively, one of the computational tasks may be mapped to a logical channel LC;
[0436] Alternatively, one of the computational tasks can be mapped to an LC set;
[0437] Alternatively, one of the computing tasks can be mapped to a physical layer resource;
[0438] Alternatively, a computing task may be mapped to a set of physical layer resources.
[0439] Optionally, the computing service request message may further include a computing service identifier for identifying the computing service.
[0440] Optionally, the device further includes:
[0441] The processing module is used to select a computing node based on the computing service request message and the status information of at least one computing node;
[0442] The status information of the computing node includes at least one of the following: computing power type, computing load, available computing speed, available computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency.
[0443] Optionally, the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
[0444] Optionally, the PDU session information includes one of the following:
[0445] Instructions to establish a PDU session;
[0446] Modify PDU session indicator and PDU session identifier;
[0447] PDU session identifier, QoS flow identifier, and QoS rules.
[0448] Optionally, the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
[0449] Optionally, the computing service request message includes at least one of the following:
[0450] The first indication information is used to indicate whether the requested computing node is a wireless access network node;
[0451] Latency type indicator, used to indicate the latency type of the computing service.
[0452] Optionally, if the computing service is a preset type of computing service or the resource type of the computing service is a latency-sensitive type, the selected computing node is a wireless access network node.
[0453] or,
[0454] If the first indication information indicates that the requested computing node is a wireless access network node, the selected computing node is a wireless access network node.
[0455] or,
[0456] When the latency type indicated by the latency type indicator is a preset latency type, the selected computing node is a wireless access network node.
[0457] Optionally, the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
[0458] Optionally, when the selected computing node is a radio access network node, the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
[0459] Optionally, the radio bearer indication is used to indicate at least one of the following: signaling radio bearer (SRB), data radio bearer (DRB), and data plane (RB).
[0460] Optionally, when the selected computing node is a wireless access network node, the apparatus further includes:
[0461] Send first information to the wireless access network node, the first information including third indication information, the third indication information being used to instruct the wireless access network node to add or modify a wireless bearer.
[0462] Optionally, the first information may also include performance parameters of the computing service.
[0463] Optionally, the sending module is further configured to send a first request message to the selected computing node, the first request message being used to request the establishment or modification of a computing task;
[0464] The first request message includes at least one of the following:
[0465] Computation task identifier; priority indicator; preemption capability indicator; preemption capability indicator; migration capability indicator; guaranteed computation speed, used to indicate the computation speed that a computing node guarantees to provide to a computing task within an average window; guaranteed computation intensity, used to indicate the computation intensity that a computing node guarantees to provide to a computing task within an average window; maximum computation speed, used to indicate the upper limit of the maximum computation speed that a computing node can provide to a computing task; maximum computation intensity, used to indicate the upper limit of the maximum computation intensity that a computing node can provide to a computing task.
[0466] Optionally, the receiving module is further configured to receive a first response message from the selected computing node;
[0467] The first response message includes status information of the selected computing node after it has established or modified the computing task.
[0468] The computing service device provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG2 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0469] Referring to Figure 8, when the computing service device is a terminal or a component within a terminal, or when the computing service device is a network-side device or a component within a network-side device, the computing service device 800 includes a sending module 801, configured to send a computing service request message to a first node, the computing service request message including performance parameters of the computing service; and a receiving module 802, configured to receive a computing service response message from the first node, the computing service response message indicating whether to accept or reject the computing service request message.
[0470] Optionally, the performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, wherein the performance parameters of the computing task or group of computing tasks include at least one of the following:
[0471] Resource type; minimum number of operations or operands; minimum computing speed; minimum computing intensity; computing latency budget; maximum failure rate, which represents the ratio of failed requests to total requests per unit time; average window; maximum time, indicating the upper limit of the computing task's duration; computing power type; data type; minimum memory, indicating the lower limit of memory required for the computing task; minimum storage, indicating the lower limit of storage required for the computing task; minimum transmission bandwidth, indicating the lower limit of bandwidth required for the computing task; computing task arrival mode; parameters corresponding to the computing task arrival mode; AI model training accuracy; AI model performance lower limit; minimum throughput of AI model inference; computing power consumption threshold; computing energy efficiency threshold.
[0472] Optionally, the computing service request message may also include a computing service identifier.
[0473] Optionally, the computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information.
[0474] Optionally, the PDU session information includes one of the following:
[0475] Instructions to establish a PDU session;
[0476] Modify PDU session indicator and PDU session identifier;
[0477] PDU session identifier, QoS flow identifier, and QoS rules.
[0478] Optionally, the computing service request message is also used to indicate whether the requested computing node is a wireless access network node.
[0479] Optionally, the computing service request message includes at least one of the following:
[0480] The first indication information is used to indicate whether the requested computing node is a wireless access network node;
[0481] Latency type indicator, used to indicate the latency type of the computing service.
[0482] Optionally, the computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node.
[0483] Optionally, the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication and a physical layer resource indication.
[0484] Optionally, the receiving module is further configured to receive a Radio Resource Control (RRC) reconfiguration message sent by a radio access network node. The RRC reconfiguration message includes a target configuration for computing services. The target configuration includes at least one of the following: radio bearer add configuration, radio bearer modify configuration, and physical layer resource configuration.
[0485] The sending module is also used to send an RRC reconfiguration complete message to the radio access network node.
[0486] Optionally, the sending module is further configured to send a second request message to the third node, the second request message being used to request the establishment or modification of a PDU session, the second request message including fourth indication information, the fourth indication information being used to indicate that the PDU session supports at least one QoS flow for computing services;
[0487] The receiving module is further configured to receive a second response message from the third node, the second response message being used to indicate acceptance of the second request message.
[0488] Optionally, the second response message includes a first QoS rule, which is used for computing services or for both computing services and communication services.
[0489] Optionally, the first QoS rule includes at least one of the following:
[0490] The target quality identifier (QoSI) for the QoS flow of the computing service includes a first QoSI or a second QoSI, wherein the first QoSI is used to represent at least one QoS parameter for the computing service, and the second QoSI is used to represent at least one QoS parameter for both the computing service and the communication service.
[0491] Packet filter set;
[0492] A priority indicator, which is used to indicate the priority of the first QoS rule.
[0493] Optionally, the first QoSI is used to represent at least one of the following: first resource type, first priority level, first computation latency budget, failure rate, average window, maximum number of operations or operands, maximum computation speed, and computing power type.
[0494] Optionally, the second QoSI is used to represent at least one of the following: a second resource type, a second priority level, a second computational delay budget, and bandwidth.
[0495] Optionally, the second request message further includes at least one of the following:
[0496] The number of packet filters used to calculate the number of packet filters for a service.
[0497] Optionally, the second response message includes a computing node identifier.
[0498] Optionally, the sending module is further configured to send second information to the computing node, the second information including computing data.
[0499] Optionally, the second information further includes at least one of the following:
[0500] Compute node identifier;
[0501] A receiving node indication is used to indicate the node that receives the computation response corresponding to the computation data.
[0502] The computing service device provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG3 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0503] As shown in Figure 9, this application embodiment also provides a communication device 900, including a processor 901 and a memory 902. The memory 902 stores programs or instructions that can run on the processor 901. For example, when the communication device 900 is a first node, the program or instructions executed by the processor 901 implement the various steps of the above-described first node-side computing service method embodiment and achieve the same technical effect. When the communication device 900 is a second node, the program or instructions executed by the processor 901 implement the various steps of the above-described second node-side computing service method embodiment and achieve the same technical effect. To avoid repetition, this will not be described again here.
[0504] This application also provides a network-side device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method embodiment shown in FIG2 or 3. This network-side device embodiment corresponds to the above-described first node or second node-side method embodiment. All implementation processes and methods of the above-described method embodiments can be applied to this network-side device embodiment and can achieve the same technical effect.
[0505] Specifically, this application embodiment also provides a network-side device, which may be the computing service device shown in FIG7. As shown in FIG10, the network-side device 1000 includes: an antenna 1001, a radio frequency device 1002, a baseband device 1003, a processor 1004, and a memory 1005. The antenna 1001 is connected to the radio frequency device 1002. In the uplink direction, the radio frequency device 1002 receives information through the antenna 1001 and sends the received information to the baseband device 1003 for processing. In the downlink direction, the baseband device 1003 processes the information to be transmitted and sends it to the radio frequency device 1002, which processes the received information and then transmits it through the antenna 1001.
[0506] The method executed by the network-side device in the above embodiments can be implemented in the baseband device 1003, which includes a baseband processor.
[0507] The baseband device 1003 may include at least one baseband board, on which multiple chips are disposed, as shown in FIG10. One of the chips is, for example, a baseband processor, which is connected to the memory 1005 via a bus interface to call the program in the memory 1005 and execute the network device operation shown in the above method embodiment.
[0508] The network-side device may also include a network interface 1006, such as a Common Public Radio Interface (CPRI).
[0509] Specifically, the network-side device 1000 in this application embodiment further includes: instructions or programs stored in memory 1005 and executable on processor 1004. Processor 1004 calls the instructions or programs in memory 1005 to execute the methods executed by each module shown in FIG7 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0510] Specifically, this application embodiment also provides a network-side device. As shown in FIG11, the network-side device 1100 includes: a processor 1101, a network interface 1102, and a memory 1103. The network-side device may be the computing service device shown in FIG7 or FIG8. The network interface 1102 is, for example, a common public radio interface (CPRI).
[0511] Specifically, the network-side device 1100 in this application embodiment further includes: instructions or programs stored in memory 1103 and executable on processor 1101. Processor 1101 calls the instructions or programs in memory 1103 to execute the methods executed by the modules shown in FIG7 or FIG8 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0512] This application also provides a terminal, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in the method embodiment shown in FIG3. This terminal embodiment corresponds to the above-described terminal-side method embodiment, and all implementation processes and methods of the above-described method embodiments can be applied to this terminal embodiment and can achieve the same technical effect. The terminal may be the computing service device shown in FIG8. Specifically, FIG12 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application.
[0513] The terminal 1200 includes, but is not limited to, at least some of the following components: radio frequency unit 1201, network module 1202, audio output unit 1203, input unit 1204, sensor 1205, display unit 1206, user input unit 1207, interface unit 1208, memory 1209, and processor 1210.
[0514] Those skilled in the art will understand that the terminal 1200 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to the processor 1210 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The terminal structure shown in Figure 12 does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0515] It should be understood that, in this embodiment, the input unit 1204 may include a graphics processor 12041 and a microphone 12042. The graphics processor 12041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1206 may include a display panel 12061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1207 includes a touch panel 12071 and at least one of other input devices 12072. The touch panel 12071 is also called a touch screen. The touch panel 12071 may include a touch detection device and a touch controller. Other input devices 12072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0516] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1201 can transmit it to the processor 1210 for processing; in addition, the radio frequency unit 1201 can send uplink data to the network-side device. Typically, the radio frequency unit 1201 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0517] The memory 1209 can be used to store software programs or instructions, as well as various data. The memory 1209 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1209 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1209 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0518] Processor 1210 may include one or more processing units; optionally, processor 1210 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1210.
[0519] The radio frequency unit 1201 is used to receive a computing service request message from the second node, the computing service request message including performance parameters of the computing service;
[0520] The radio frequency unit 1201 is also configured to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message.
[0521] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the computing service method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.
[0522] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described computing service method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0523] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0524] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described computing service method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0525] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0526] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described computing service method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0527] This application also provides a wireless communication system, including a first node and a second node, wherein the first node can be used to execute the steps of the computing service method described above, and the second node can be used to execute the steps of the computing service method described above.
[0528] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0529] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0530] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
A computing service method, comprising: The first node receives a computing service request message from the second node, the computing service request message including the performance parameters of the computing service; The first node sends a computing service response message to the second node, the computing service response message being used to indicate whether to accept or reject the computing service request message. The method according to claim 1, characterized in that, The performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, and the performance parameters of the computing task or group of computing tasks include at least one of the following: Resource type; Minimum number of operands; Minimum computation speed; Minimum calculated strength; Calculate the delay budget; Maximum failure rate, which is used to represent the ratio of the number of failed requests to the total number of requests in a task per unit time. Average window; Maximum time, used to represent the upper limit of the duration of the computation task; Computing power type; Data type; Minimum memory, used to represent the lower limit of memory required for a computing task; Minimum storage, used to represent the lower limit of storage required for a computational task; Minimum transmission bandwidth, used to represent the lower limit of bandwidth for a computing task; Calculate the task arrival pattern; Calculate the parameters corresponding to the task arrival mode; Artificial intelligence (AI) model training accuracy; Lower limit of AI model performance; Minimum throughput for AI model inference; Calculate the power consumption threshold; Calculate the energy efficiency threshold. The method according to claim 2, wherein, The resource type includes at least one of the following: Guarantee computation speed, guarantee computation intensity, do not guarantee computation speed, do not guarantee computation intensity, latency-sensitive guarantee computation speed, latency-sensitive guarantee computation intensity. The method according to claim 2 or 3, wherein, The minimum computing speed includes at least one of the following: theoretical minimum computing speed, and actual minimum computing speed. The method according to any one of claims 2 to 4, wherein, The computation delay budget includes at least one of the following: an upper limit for computation delay, an upper limit for transmission delay, and an upper limit for both computation and transmission delay. The method according to any one of claims 2 to 5, wherein, The minimum transmission bandwidth includes at least one of the following: minimum uplink bandwidth, minimum downlink bandwidth, uplink bandwidth indication, downlink bandwidth indication, and target indication; The target indication is used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different. The method according to any one of claims 2 to 6, wherein, The computation task includes at least two computation jobs, and the computation task arrival mode includes at least one of the following: Continuous arrival mode or single arrival mode is used to indicate that one computation job arrives at a time; Fixed-cycle arrival mode is used to indicate the arrival of control calculation jobs according to a fixed cycle; Poisson distribution arrival pattern, used to indicate the arrival of computational jobs controlled by Poisson distribution; Peak arrival mode is used to indicate the arrival of σ computational jobs within a target period of Poisson distribution, where the duration of the target period is less than a preset duration, and σ is a positive integer. Offline arrival mode, used to indicate that all computation jobs arrive at once. The method according to any one of claims 2 to 7, wherein, The lower limit of AI model performance includes at least one of the following: the lower limit of AI model training performance, and the lower limit of AI model inference performance. The method according to any one of claims 2 to 8, wherein, The performance parameters of the computing task also include a dataset indicator, wherein the performance of the AI model corresponding to the dataset indicated by the dataset indicator must meet the lower limit of the AI model performance. The method according to any one of claims 2 to 9, wherein, One of the computational tasks is mapped to a Quality of Service (QoS) stream; Alternatively, one of the computational tasks can be mapped to a set of QoS flows; Alternatively, one of the computing tasks can be mapped to a Protocol Data Unit (PDU) session; Alternatively, one of the computing tasks can be mapped to a set of PDU sessions; Alternatively, one of the computing tasks can be mapped to a radio bearer (RB). Alternatively, one of the computational tasks can be mapped to a set of RBs; Alternatively, one of the computational tasks may be mapped to a logical channel LC; Alternatively, one of the computational tasks can be mapped to an LC set; Alternatively, one of the computing tasks can be mapped to a physical layer resource; Alternatively, a computing task may be mapped to a set of physical layer resources. The method according to any one of claims 1 to 10, wherein, The computing service request message also includes a computing service identifier, used to identify the computing service. The method according to any one of claims 1 to 11, wherein, The method further includes: The first node selects a computing node based on the computing service request message and the status information of at least one computing node. The status information of the computing node includes at least one of the following: computing power type, computing load, available computing speed, available computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency. The method according to any one of claims 1 to 12, wherein, The computing service response message includes at least one of the following: computing task identifier, computing node identifier, and PDU session information. The method according to claim 13, wherein, The PDU session information includes the following: Instructions to establish a PDU session; Modify PDU session indicator and PDU session identifier; PDU session identifier, QoS flow identifier, and QoS rules. The method according to any one of claims 1 to 12, wherein, The computing service request message is also used to indicate whether the requested computing node is a wireless access network node. The method according to claim 15, wherein, The computing service request message includes at least one of the following: The first indication information is used to indicate whether the requested computing node is a wireless access network node; Latency type indicator, used to indicate the latency type of the computing service. The method according to claim 16, wherein, When the computing service is a preset type of computing service or the resource type of the computing service is a latency-sensitive type, the selected computing node is a wireless access network node; or, If the first indication information indicates that the requested computing node is a wireless access network node, the selected computing node is a wireless access network node. or, When the latency type indicated by the latency type indicator is a preset latency type, the selected computing node is a wireless access network node. The method according to any one of claims 1 to 12, 15 to 17, wherein, The computing service response message includes at least one of the following: second indication information, computing task identifier, and computing node identifier; the second indication information is used to indicate that the selected computing node is a wireless access network node. The method according to claim 18, wherein, When the selected computing node is a radio access network node, the computing service response message may further include a radio bearer indication, or the computing service response message may further include at least one of the following: a physical layer channel indication, a physical layer resource indication. The method according to claim 19, wherein, The radio bearer indication is used to indicate at least one of the following: signaling radio bearer (SRB), data radio bearer (DRB), and data plane (RB). The method according to claim 19 or 20, wherein, When the selected computing node is a wireless access network node, the method further includes: The first node sends first information to the radio access network node, the first information including third indication information, the third indication information being used to instruct the radio access network node to add or modify a radio bearer. The method according to claim 21, wherein, The first information also includes the performance parameters of the computing service. The method according to any one of claims 1 to 22, wherein, The method further includes: The first node sends a first request message to the selected computing node, the first request message being used to request the creation or modification of a computing task; The first request message includes at least one of the following: Calculate the task identifier; Priority indication; Capture capability indicator; Capability preemption indication; Migration capability indicator; Guarantee computation speed, used to instruct computing nodes to guarantee the computation speed provided to computing tasks within the average window; Guarantee computational intensity, used to instruct computing nodes to guarantee the computational intensity provided to computing tasks within the average window; Maximum computing speed indicates the upper limit of the maximum computing speed that a computing node can provide to a computing task; Maximum computational intensity indicates the upper limit of the maximum computational intensity that a computing node can provide to a computing task. The method according to claim 23, wherein, The method further includes: The first node receives a first response message from the selected computing node; The first response message includes status information of the selected computing node after it has established or modified the computing task. A computing service method, comprising: The second node sends a computing service request message to the first node, the computing service request message including the performance parameters of the computing service; The second node receives a computing service response message from the first node, the computing service response message being used to indicate whether to accept or reject the computing service request message. The method according to claim 25, wherein, The performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, and the performance parameters of the computing task or group of computing tasks include at least one of the following: Resource type; Minimum number of operands; Minimum computation speed; Minimum calculated strength; Calculate the delay budget; Maximum failure rate, which is used to represent the ratio of the number of failed requests to the total number of requests in a task per unit time. Average window; Maximum time indicates the upper limit of the duration of the computation task; Computing power type; Data type; Minimum memory, used to indicate the lower limit of memory required for a computing task; Minimum storage, used to indicate the lower limit of storage required for a computational task; Minimum transmission bandwidth, used to indicate the lower limit of bandwidth for computing tasks; Calculate the task arrival pattern; Calculate the parameters corresponding to the task arrival mode; AI model training accuracy; Lower limit of AI model performance; Minimum throughput for AI model inference; Calculate the power consumption threshold; Calculate the energy efficiency threshold. The method according to claim 25 or 26, wherein, The computing service request message also includes a computing service identifier. The method according to any one of claims 25 to 27, wherein, The computing service request message is also used to indicate whether the requested computing node is a wireless access network node. The method according to claim 28, wherein, The computing service request message includes at least one of the following: The first indication information is used to indicate whether the requested computing node is a wireless access network node; Latency type indicator, used to indicate the latency type of the computing service. The method according to any one of claims 25 to 29, wherein, When the selected computing node is a wireless access network node, the method further includes: The second node receives a Radio Resource Control (RRC) reconfiguration message sent by a Radio Access Network (RAN) node. The RRC reconfiguration message includes a target configuration for computing services. The target configuration includes at least one of the following: radio bearer add configuration, radio bearer modify configuration, and physical layer resource configuration. The second node sends an RRC reconfiguration complete message to the radio access network node. The method according to any one of claims 25 to 27, wherein, The method further includes: The second node sends a second request message to the third node. The second request message is used to request the establishment or modification of a PDU session. The second request message includes fourth indication information, which is used to indicate that the PDU session supports at least one QoS flow for computing services. The second node receives a second response message from the third node, the second response message indicating acceptance of the second request message. The method according to claim 31, wherein, The second response message includes a first QoS rule, which is used for computing services or for both computing services and communication services. The method according to claim 32, wherein, The first QoS rule includes at least one of the following: The target quality identifier (QoSI) for the QoS flow of the computing service includes a first QoSI or a second QoSI, wherein the first QoSI is used to represent at least one QoS parameter for the computing service, and the second QoSI is used to represent at least one QoS parameter for both the computing service and the communication service. Packet filter set; A priority indicator, which is used to indicate the priority of the first QoS rule. The method according to claim 33, wherein, The first QoSI is used to represent at least one of the following: first resource type, first priority level, first computation latency budget, failure rate, average window, maximum number of operations or operands, maximum computation speed, and computing power type; And / or, The second QoSI is used to represent at least one of the following: a second resource type, a second priority level, a second computational delay budget, and bandwidth. The method according to any one of claims 31 to 34, wherein, The second request message also includes at least one of the following: The number of packet filters used to calculate the number of packet filters for a service. The method according to any one of claims 31 to 35, wherein, The second response message includes a compute node identifier. The method according to any one of claims 25 to 36, wherein, The method further includes: The second node sends a second message to the computing node, the second message including computing data. The method according to claim 37, wherein, The second information also includes at least one of the following: Compute node identifier; A receiving node indication is used to indicate the node that receives the computation response corresponding to the computation data. A computing service device, comprising: The receiving module is configured to receive a computing service request message from the second node, the computing service request message including performance parameters of the computing service; The sending module is used to send a computing service response message to the second node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message. The apparatus according to claim 39, wherein, The performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, and the performance parameters of the computing task or group of computing tasks include at least one of the following: Resource type; Minimum number of operands; Minimum computation speed; Minimum calculated strength; Calculate the delay budget; Maximum failure rate, which is used to represent the ratio of the number of failed requests to the total number of requests in a task per unit time. Average window; Maximum time, used to represent the upper limit of the duration of the computation task; Computing power type; Data type; Minimum memory, used to represent the lower limit of memory required for a computing task; Minimum storage, used to represent the lower limit of storage required for a computational task; Minimum transmission bandwidth, used to represent the lower limit of bandwidth for a computing task; Calculate the task arrival pattern; Calculate the parameters corresponding to the task arrival mode; Artificial intelligence (AI) model training accuracy; Lower limit of AI model performance; Minimum throughput for AI model inference; Calculate the power consumption threshold; Calculate the energy efficiency threshold. The apparatus according to claim 39 or 40, wherein, The device further includes: The processing module is used to select a computing node based on the computing service request message and the status information of at least one computing node; The status information of the computing node includes at least one of the following: computing power type, computing load, available computing speed, available computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency. A computing service device, comprising: The sending module is used to send a computing service request message to the first node, the computing service request message including the performance parameters of the computing service; The receiving module is configured to receive a computing service response message from the first node, wherein the computing service response message is used to indicate whether to accept or reject the computing service request message. The apparatus according to claim 42, wherein, The performance parameters of the computing service include performance parameters of at least one computing task or at least one group of computing tasks, and the performance parameters of the computing task or group of computing tasks include at least one of the following: Resource type; Minimum number of operands; Minimum computation speed; Minimum calculated strength; Calculate the delay budget; Maximum failure rate, which is used to represent the ratio of the number of failed requests to the total number of requests in a task per unit time. Average window; Maximum time indicates the upper limit of the duration of the computation task; Computing power type; Data type; Minimum memory, used to indicate the lower limit of memory required for a computing task; Minimum storage, used to indicate the lower limit of storage required for a computational task; Minimum transmission bandwidth, used to indicate the lower limit of bandwidth for computing tasks; Calculate the task arrival pattern; Calculate the parameters corresponding to the task arrival mode; AI model training accuracy; Lower limit of AI model performance; Minimum throughput for AI model inference; Calculate the power consumption threshold; Calculate the energy efficiency threshold. The apparatus according to claim 42 or 43, wherein, The sending module is further configured to send a second request message to the third node. The second request message is used to request the establishment or modification of a PDU session. The second request message includes fourth indication information, which is used to indicate that the PDU session supports at least one QoS flow for computing services. The receiving module is further configured to receive a second response message from the third node, the second response message being used to indicate acceptance of the second request message. The apparatus according to claim 44, wherein, The second response message includes a first QoS rule, which is used for computing services or for both computing services and communication services. A first node includes a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the computing service method as claimed in any one of claims 1 to 24. A second node includes a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the computing service method as described in any one of claims 25 to 38. A readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the computing service method as claimed in any one of claims 1 to 24, or implement the steps of the computing service method as claimed in any one of claims 25 to 38. A computer program product, which is executed by at least one processor to implement the steps of the computing service method as claimed in any one of claims 1 to 24, or to implement the steps of the computing service method as claimed in any one of claims 25 to 38.
Citation Information
Patent Citations
AI service request processing method, device and equipment and readable storage medium
CN118200882A
Computing service implementation method and device, communication equipment and readable storage medium
CN118283712A
Network edge computing method and communication apparatus
WO2022067816A1