Information transmission method and apparatus, and communication device
By exchanging information between the terminal and network nodes, the problem of unmet computing service needs in communication scenarios is solved, and more efficient allocation of computing resources and terminal computing service capabilities are achieved.
Patent Information
- Application Number
- PCT/CN2025/105359
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-15
AI Technical Summary
The demand for computing services in communication scenarios is difficult to meet. The static nature of existing cloud computing services is mismatched with dynamic communication scenarios, resulting in difficulties in allocating computing resources and meeting demand.
Terminals and network nodes exchange computing service-related information through information transmission methods. Terminals provide computing service auxiliary information to help network nodes schedule and allocate resources, including receiving and sending information and auxiliary information related to computing services.
It enables better fulfillment of computing service needs in dynamic communication scenarios, improves the allocation efficiency of computing resources and the computing service capabilities of terminals.
Smart Images

Figure CN2025105359_15012026_PF_FP_ABST
Abstract
Description
Information transmission methods, devices and communication equipment
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410909481.1, filed in China on July 8, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application belongs to the field of communication technology, specifically relating to an information transmission method, apparatus, and communication equipment. Background Technology
[0004] In related technologies, cloud service technology is typically used to provide computing services (i.e., cloud computing services). The cloud computing services provided by cloud service providers are usually static computing services. With the development of communication technology, more and more communication scenarios are also facing the need for computing services. However, communication scenarios are usually dynamic, and the static computing services provided by cloud computing are not suitable for these scenarios, making it difficult to meet the computing service demands of communication scenarios. Summary of the Invention
[0005] This application provides an information transmission method, apparatus, and communication device that can solve the problem that the demand for computing services in communication scenarios is difficult to meet.
[0006] Firstly, an information transmission method is provided, executed by a terminal, the method comprising:
[0007] The terminal performs a first operation, which includes at least one of the following:
[0008] Receive first information from the first node, the first information being used to indicate information related to the network's computing services;
[0009] Send a second message to the first node, the second message being used to indicate auxiliary information related to the terminal's computing services.
[0010] Secondly, an information transmission method is provided, executed by the first node, the method comprising:
[0011] The first node performs a second operation, which includes at least one of the following:
[0012] Send first information to the terminal, the first information being used to indicate information related to the network's computing services;
[0013] Receive second information from the terminal, the second information being used to indicate auxiliary information related to the terminal's computing services.
[0014] Thirdly, an information transmission method is provided, executed by a second node, the method comprising:
[0015] The second node receives a fourth message from the terminal, which is used to request computing services;
[0016] The second node determines the third node based on the fourth information, and the third node is a computing node;
[0017] The second node sends a third message to the third node, the third message being used to request the creation of a computing task or the modification of a computing task.
[0018] Fourthly, an information transmission method is provided, executed by a third node, the method comprising:
[0019] The third node receives a third message from the second node, which is used to request the creation of a computing task or the modification of a computing task.
[0020] Fifthly, an information transmission device is provided, the device comprising at least one of the following:
[0021] The first receiving module is configured to receive first information from the first node, wherein the first information is used to indicate information related to the computing services of the network;
[0022] The first sending module is used to send second information to the first node, the second information being used to indicate auxiliary information related to the computing services of the terminal.
[0023] Sixthly, an information transmission device is provided, the device comprising at least one of the following:
[0024] The first sending module is used to send first information to the terminal, the first information being used to indicate information related to the network's computing services;
[0025] A first receiving module is configured to receive second information from a terminal, the second information being used to indicate auxiliary information related to the terminal's computing services.
[0026] In a seventh aspect, an information transmission device is provided, the device comprising:
[0027] A receiving module is used to receive fourth information from the terminal, the fourth information being used to request computing services;
[0028] The processing module is used to determine a third node based on the fourth information, wherein the third node is a computing node;
[0029] The first sending module is used to send a third message to the third node, the third message being used to request the establishment of a computing task or the modification of a computing task.
[0030] Eighthly, an information transmission device is provided, the device comprising:
[0031] The first receiving module is used to receive a third message from the second node, the third message being used to request the establishment of a computing task or the modification of a computing task.
[0032] Ninth aspect, an information transmission apparatus is provided, the apparatus being configured to perform the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect, or implement the steps of the method described in the third aspect, or implement the steps of the method described in the fourth aspect.
[0033] In a tenth aspect, a terminal is provided, the terminal including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.
[0034] Eleventhly, a terminal is provided, including a processor and a communication interface, wherein the communication interface is used for at least one of the following: receiving first information from a first node, the first information being used to indicate information related to computing services of the network; and sending second information to the first node, the second information being used to indicate auxiliary information related to computing services of the terminal.
[0035] In a twelfth aspect, a network-side device is provided, the network-side device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the second aspect, or implementing the steps of the method as described in the third aspect, or implementing the steps of the method as described in the fourth aspect.
[0036] In a thirteenth aspect, a network-side device is provided, including a processor and a communication interface, wherein the communication interface is used for at least one of the following: sending first information to a terminal, the first information being used to indicate information related to network computing services; and receiving second information from the terminal, the second information being used to indicate auxiliary information related to the terminal's computing services.
[0037] In a fourteenth aspect, a network-side device is provided, including a processor and a communication interface, wherein the communication interface is configured to: receive fourth information from a terminal, the fourth information being used to request computing services; the processor is configured to: determine a third node based on the fourth information, the third node being a computing node; the communication interface is further configured to: send a third message to the third node, the third message being used to request the establishment of a computing task or the modification of a computing task.
[0038] In a fifteenth aspect, a network-side device is provided, including a processor and a communication interface, wherein the communication interface is used to: receive a third message from a second node, the third message being used to request the establishment of a computing task or the modification of a computing task.
[0039] In a sixteenth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect, or the steps of the method described in the third aspect, or the steps of the method described in the fourth aspect.
[0040] In a seventeenth aspect, a wireless communication system is provided, comprising: a terminal and a network-side device, wherein the terminal is configured to perform the steps of the method described in the first aspect, and the network-side device is configured to perform the steps of the method described in the second aspect, or implement the steps of the method described in the third aspect, or implement the steps of the method described in the fourth aspect.
[0041] Eighteenthly, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method as described in the first aspect, or the steps of the method as described in the second aspect, or the steps of the method as described in the third aspect, or the steps of the method as described in the fourth aspect.
[0042] In a nineteenth aspect, a computer program / program product is provided, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to implement the steps of the method as described in the first aspect, or the steps of the method as described in the second aspect, or the steps of the method as described in the third aspect, or the steps of the method as described in the fourth aspect.
[0043] In this embodiment, the terminal performs a first operation, which includes at least one of the following: receiving first information from a first node, the first information indicating information related to network computing services; and sending second information to the first node, the second information indicating auxiliary information related to the terminal's computing services. Since the terminal receives information related to network computing services, it can obtain network computing services based on this information when a computing service requirement exists. Furthermore, since the terminal sends auxiliary information related to computing services, it can provide assistance for computing services. Therefore, this embodiment can effectively meet the computing service requirements of communication scenarios. Attached Figure Description
[0044] Figure 1 is a schematic diagram of a network structure applicable to the embodiments of this application;
[0045] Figure 2 is a flowchart of an information transmission method provided in an embodiment of this application;
[0046] Figure 3 is a flowchart of an information transmission method provided in an embodiment of this application;
[0047] Figure 4 is a flowchart of an information transmission method provided in an embodiment of this application;
[0048] Figure 5 is a flowchart of an information transmission method provided in an embodiment of this application;
[0049] Figure 6 is a flowchart of Example 1;
[0050] Figure 7 is a flowchart of Example 2;
[0051] Figure 8 is a flowchart of Example 3;
[0052] Figure 9 is a structural diagram of an information transmission device provided in an embodiment of this application;
[0053] Figure 10 is a structural diagram of an information transmission device provided in an embodiment of this application;
[0054] Figure 11 is a structural diagram of an information transmission device provided in an embodiment of this application;
[0055] Figure 12 is a structural diagram of an information transmission device provided in an embodiment of this application;
[0056] Figure 13 is a structural diagram of a communication device provided in an embodiment of this application;
[0057] Figure 14 is a structural diagram of a terminal provided in an embodiment of this application;
[0058] Figure 15 is a structural diagram of a network-side device provided in an embodiment of this application;
[0059] Figure 16 is a structural diagram of another network-side device provided in an embodiment of this application. Detailed Implementation
[0060] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0061] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0062] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0063] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and the term NR is used in most of the following description; however, these technologies can also be applied to systems other than NR systems, such as 6th generation (6G) radio systems. th Generation 6G communication system.
[0064] Figure 1 shows a block diagram of a wireless communication system applicable to an embodiment of this application. The wireless communication system includes a terminal 11 and a network-side device 12. The terminal 11 can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the specific type of terminal 11 is not limited in this application embodiment. Network-side equipment 12 may include access network equipment or core network equipment, wherein access network equipment may also be referred to as Radio Access Network (RAN) equipment, radio access network function, or radio access network unit. Access network equipment may include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.The term "base station" can be referred to as Node B (NB), Evolved Node B (eNB), Next Generation Node B (gNB), New Radio Node B (NR Node B), Access Point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), Radio Base Station, Radio Transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home Evolved Node B, Transmit / Receive Point (TRP), or any other suitable term in the relevant field, as long as the same technical effect is achieved. The term "base station" is not limited to any specific technical terminology. It should be noted that this application embodiment only uses a base station in an NR system as an example for description and does not limit the specific type of base station.
[0065] Core network equipment, also known as core network nodes, core network functions, or core network elements, includes, but is not limited to, at least one of the following: Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (or L-NEF), and Binding Support. The core network functions include: BSF (Block Network Function), Application Function (AF), Location Management Function (LMF), Gateway Mobile Location Centre (GMLC), and Network Data Analytics Function (NWDAF). It should be noted that this application embodiment only uses core network equipment in the NR system as an example and does not limit the specific type of core network equipment. If the name of the core network equipment mentioned in this application embodiment changes in subsequent protocol versions (e.g., 6G), it will still be within the scope of protection of this application.
[0066] Optionally, the core network equipment can be implemented by one or more functional modules in a single device, or by multiple devices working together; this application does not specifically limit this. It is understood that the aforementioned functional modules can be network elements in hardware devices, software functional modules running on dedicated hardware, or virtualized functional modules instantiated on a platform (e.g., a cloud platform).
[0067] Before describing the embodiments of this application, the relevant technologies are briefly introduced below:
[0068] Technology 1: 5G System Messaging
[0069] The existing 3rd Generation Partnership Project (3GPP) standard defines a Master Information Block (MIB) and multiple System Information Blocks (SIBs) to provide system information (such as frame number, subcarrier spacing, Public Land Mobile Network (PLMN), Tracking Area Code (TAC), and cell selection information). User Equipment (UE) first obtains the necessary information through the MIB and SIB to successfully access the network and transmit data. Furthermore, system information can be sent and updated periodically so that newly entering UEs can obtain and update their system information.
[0070] Technology 2: UE Assistance Information (UAI)
[0071] The relevant technologies define the UAI procedure, through which the UE can provide the network with Discontinuous Reception (DRX) configuration, Carrier Aggregation (CA) configuration, Radio Resource Control (RRC) status, etc., for energy saving, heat dissipation, measurement and other optimization configurations.
[0072] Technology 3: Artificial Intelligence (AI) and Communication Application Scenarios
[0073] AI and communications are among the 6G application scenarios identified by the Radio Communication Division of the International Telecommunication Union (ITU-R). Typical use cases include IMT-2030 (6G) assisted autonomous driving, autonomous collaboration between devices in medical assistance applications, cross-device / network computation offloading, creation and prediction of digital twins, and IMT-2030 (6G) assisted collaborative robots.
[0074] These application scenarios will require support for high mobile network capacity and high user experience data rates, as well as low latency and high reliability. Beyond communications, this application scenario is expected to include a suite of new functionalities integrating artificial intelligence and computing into 6G systems, including data acquisition, preparation, and processing from various sources; distributed AI model training; model sharing and distributed inference across mobile communication systems; and computing resource orchestration.
[0075] In related technologies, cloud computing services are mainly defined by service level agreements (SLAs) to define the static, long-term (e.g., monthly) quality of service between cloud service providers and users. Users are mainly enterprise users, with fewer direct-to-consumer users, especially mobile terminal consumers.
[0076] Compared to cloud computing services in related technologies, 6G resources are more dynamic, and 6G computing resources can be centralized in the core network or distributed in the radio access network. Cloud computing technologies in related technologies are not suitable for the high dynamism of resource states and the distributed nature of computing resources, making it difficult to meet the computing service demands of communication scenarios.
[0077] In view of this, embodiments of this application provide an information transmission method, an information transmission device, and a communication equipment to solve the problem that the demand for computing services in communication scenarios is difficult to meet in related technologies.
[0078] The information transmission method provided in this application will be described in detail below with reference to the accompanying drawings and through some embodiments and application scenarios.
[0079] Figure 2 shows a flowchart of an information transmission method provided in an embodiment of this application. As shown in Figure 2, the information transmission method includes the following steps:
[0080] Step 201: The terminal performs a first operation, which includes at least one of the following:
[0081] Receive first information from the first node, the first information being used to indicate information related to the network's computing services;
[0082] Send a second message to the first node, the second message being used to indicate auxiliary information related to the terminal's computing services.
[0083] In this embodiment of the application, the computing service may include AI services, and may also include other computing services other than AI services.
[0084] The first node can be a core network functional node (such as AMF) or a radio access network node (such as a base station).
[0085] Information related to computing services can be understood or replaced with computing service information, and the same descriptions thereafter can be understood in the same way, so it will not be elaborated on further.
[0086] By receiving the first information from the first node, the terminal can obtain information related to the network's computing services. Therefore, when there is a need for computing services, the terminal can obtain the network's computing services based on this information.
[0087] By sending second information to the first node, the terminal can provide the first node with auxiliary information related to computing services. Based on this information, the first node can better understand the terminal's status and better schedule and allocate computing resources.
[0088] In this embodiment, the terminal performs a first operation, which includes at least one of the following: receiving first information from a first node, the first information indicating information related to network computing services; and sending second information to the first node, the second information indicating auxiliary information related to the terminal's computing services. Since the terminal receives information related to network computing services, it can obtain network computing services based on this information when a computing service requirement exists. Furthermore, since the terminal sends auxiliary information related to computing services, it can provide assistance for computing services. Therefore, this embodiment can effectively meet the computing service requirements of communication scenarios.
[0089] The following is an explanation of the first piece of information.
[0090] In some embodiments, the first information includes at least one of the following:
[0091] The first indication is used to indicate whether the target area supports computing services;
[0092] The second indication is used to indicate the types of resources supported by the network, including types of computing resources or types of computing and communication resources;
[0093] The third indicator is used to indicate the maximum computing speed supported by the network;
[0094] The fourth indicator is used to indicate the maximum transmission bandwidth supported by the network;
[0095] The fifth indicator is used to indicate the maximum computational intensity supported by the network;
[0096] The sixth indicator is used to indicate the type of computing power supported by the network;
[0097] The seventh instruction is used to indicate the types of computing services supported by the network.
[0098] The first indication can be called a computing service indication, and the target area can be understood as a specific geographical region. That is, the first indication is used to indicate whether a specific geographical region supports computing services. The geographical region can be a cell, an access network notification area (RAN-based notification area, RNA), a tracking area (TA), etc. If the network only supports AI services, the first indication can also be called an AI service indication.
[0099] The second indication can be called a resource type indication, which includes the type of computing resources (which can be called computing resource type) or the type of computing and communication resources (which can be called computing and communication resource type).
[0100] Optionally, the resource type includes at least one of the following:
[0101] The first type, the resources corresponding to the first type are those that can guarantee computing speed;
[0102] The second type, the resources corresponding to the second type are those that can guarantee computing power;
[0103] The third type, the resources corresponding to the third type are resources that do not guarantee computing speed;
[0104] The fourth type refers to resources that do not guarantee computational intensity.
[0105] The fifth type refers to resources that are latency-sensitive and guarantee computing speed.
[0106] The sixth type refers to resources that are latency-sensitive and guarantee computational intensity.
[0107] In other words, potential resource types include guaranteed computing rate (GCR) or guaranteed operational rate (GOR), guaranteed computing intensity (GCI) or guaranteed operational intensity (GOI), non-guaranteed computing rate (Non-GCR or Non-GOR), non-guaranteed computing intensity (Non-GCI or Non-GOI), delay-sensitive guaranteed computing rate (delay-critical GCR or delay-critical GOR), and delay-sensitive guaranteed computing intensity (delay-critical GCI or delay-critical GOI).
[0108] Computational speed refers to the computational time complexity divided by the tolerable upper limit of computation time. For example, computational time complexity is usually represented by operations, and the computation typically corresponds to a data type. Data types can be at least one of integers (such as int8, int4, etc.) and floating-point numbers (such as half-precision floating-point numbers, single-precision floating-point numbers, double-precision floating-point numbers, and semi-precision floating-point numbers). For instance, 2 TFLOPs (floating-point operations) indicates that the time complexity of the computation is 2 × 10⁻⁶. 12 Floating-point operands. If the tolerable computation time limit is 20ms, meaning the computation can be completed in a maximum of 20ms, then the computation speed is 2 × 10⁻⁶. 12 / (20×10 -3 = 100 TFLOPS (floating-point operations per second).
[0109] Computational intensity refers to computational speed divided by bandwidth. As mentioned earlier, in the process of providing computing services in 6G, bandwidth consists of multiple parts, including the transmission bandwidth between the computing request node (such as UE) and the computing node, and the memory bandwidth of the computing node. One type of computational intensity is defined in segments, such as computational speed divided by memory bandwidth. Another type of computational intensity takes the minimum of the relevant bandwidths as the numerator, that is, computational speed divided by min{memory bandwidth, transmission bandwidth between UE and computing node}. The transmission bandwidth between UE and computing node can be further divided into the air interface bandwidth between UE and access network node, and the wired transmission bandwidth between access network node and computing node; or, the transmission bandwidth between UE and computing node can be further divided into the bandwidth between UE and user plane function node (such as UPF) node, and the wired transmission bandwidth between UPF and computing node.
[0110] The aforementioned guaranteed computing speed can be understood as the computing resources associated with the guaranteed computing speed of a computing task being permanently allocated during that computing task. For example, computing resources such as the Central Processing Unit (CPU), Graphics Processing Unit (GPU), or storage on the selected computing node, associated with the guaranteed computing speed of the computing task, are allocated to that computing task. The computing resources allocated to this computing task are dedicated to that task and are not shared with other computing tasks.
[0111] The aforementioned guaranteed computing strength can be understood as computing resources, or computing and communication resources, permanently allocated during the computing task to guarantee its computing strength. For example, when the computing strength is the value obtained by dividing the computing speed by the memory bandwidth, computing resources such as CPU, GPU, and memory related to the guaranteed computing strength are allocated to the computing task; when the computing strength is the value obtained by dividing the computing speed by the target bandwidth (i.e., the smaller of the memory bandwidth and the transmission bandwidth), CPU, GPU, memory, and network resources related to the guaranteed computing strength are allocated to the computing task. Network resources include air interface resources or wired transmission resources, etc. Air interface resources are related to the air interface transmission bandwidth of the computing task between the UE and the network, while wired transmission resources are related to the transmission bandwidth from the access network to the computing node of the computing task between the UE and the network.
[0112] The aforementioned latency-sensitive guarantee of computational speed can be understood as follows: if the latency of a data packet in a computation task exceeds the computational latency budget, the data packet is discarded. Furthermore, the latency-sensitive guarantee of computational speed requires that a first proportion (e.g., 98%) of data packets do not exceed the computational latency budget.
[0113] The aforementioned latency-sensitive guarantee of computational strength can also be understood as follows: if the latency of a data packet in the computation task exceeds the computational latency budget, the data packet is discarded. Furthermore, the latency-sensitive guarantee of computational speed requires that a second proportion (e.g., 98%) of the data packets cannot exceed the computational latency budget.
[0114] The aforementioned non-guaranteed computing speed can be understood as resources related to the computing speed of a computing task not being permanently allocated during that computing task. For example, computing resources such as CPU, GPU, or storage on the selected computing node that are related to the guaranteed computing speed of the computing task are not allocated to that computing task but are shared with other computing tasks.
[0115] The aforementioned non-guaranteed computational intensity can be understood as resources related to the computational intensity of a computational task not being permanently allocated during that task. For example, computational resources such as CPU, GPU, and memory, as well as network resources, on the selected computing node that are related to the guaranteed computational speed of the task, are not allocated to that task but are shared with other computational tasks.
[0116] The third indication can be understood as an indication of the maximum computing speed, which can be used in conjunction with GCI or delay critical GCI or GCR or delay critical GCR.
[0117] Optionally, the maximum computing speed includes at least one of the following: theoretical maximum computing speed and actual maximum computing speed. The theoretical maximum computing speed is the maximum computing speed supported by the network. The theoretical peak number of operations is the maximum theoretical computing speed of all available computing nodes in the network. Typically, the maximum theoretical computing speed of a given computing node is obtained by multiplying the processor's clock speed, the number of operations performed by the processor per clock cycle, and the total number of cores in the system.
[0118] The actual maximum computing speed is the maximum computing speed that the network supports. The actual number of operations is the maximum computing speed obtained through testing.
[0119] Optionally, the third instruction includes at least one of the following:
[0120] The theoretical maximum computation speed;
[0121] Information used to indicate the target test case;
[0122] First computational efficiency, which is the ratio of the actual maximum computational speed to the theoretical maximum computational speed;
[0123] The second computational efficiency is the ratio of the maximum computational speed measured based on the target test case to the theoretical maximum computational speed.
[0124] The third computational efficiency is the ratio of the maximum computational speed measured based on the target test case to the maximum computational speed under ideal conditions.
[0125] The actual maximum computing speed is determined by at least one of the theoretical maximum computing speed, the target test case, the first computing efficiency, the second computing efficiency, and the third computing efficiency.
[0126] The information described above for indicating the target test cases can be called the twelfth indication or test case indication. That is, the test case indication identifies which test case was used to achieve the actual maximum computational speed. For example, a test case can be at least one of the following: MobileNetVx (where x can be any available version number, such as 1), one-dimensional Discrete Fourier Transform (DFT), one-dimensional Fast Fourier Transform (FFT), two-dimensional FFT, matrix multiplication, sparse linear equations, dense linear equations, YOLOvy (where y can be any available version number, such as YOLOv5), image recognition models (such as resnet50_v1.5), large models (such as Llama3, Llama2), etc. Test cases can be open-source software programs or custom software programs between the UE and the network. This means that the actual maximum computational speed supported by the network is the actual maximum computational speed that the node can achieve under the indicated test case conditions. Therefore, the test case indication can represent the actual maximum computational speed supported by the network.
[0127] Computational efficiency:
[0128] A computational efficiency is equal to the ratio of the actual maximum computational speed under ideal conditions to the theoretical maximum computational speed, i.e., the first computational efficiency. Therefore, the actual maximum computational speed supported by the network can be represented by the theoretical maximum computational speed and the first computational efficiency.
[0129] Another type of computational efficiency is the ratio of the maximum computational speed measured based on test cases to the theoretical maximum computational speed, which is called the second computational efficiency. Therefore, the actual maximum computational speed supported by the network can be represented by the theoretical maximum computational speed, the test case indication, and the second computational efficiency.
[0130] Another type of computational efficiency is the ratio of the computational speed measured based on test cases to the maximum computational speed under ideal conditions, also measured based on test cases; this is known as the third computational efficiency. Therefore, the actual maximum computational speed supported by the network can be represented by the maximum computational speed under ideal conditions, the test case indicators, and the third computational efficiency.
[0131] It should be noted that the computing speed measured based on test cases is usually obtained from recent tests and can represent the current state of the computing node; the maximum computing speed under ideal conditions measured based on test cases refers to the best measured performance of the computing node.
[0132] The fourth indication can be understood as the indication information of the maximum transmission bandwidth, which is used to define the maximum uplink bandwidth and / or maximum downlink bandwidth supported by the network.
[0133] Optionally, since uplink bandwidth and downlink bandwidth may be related to channel quality, the fourth indication includes at least one of the following:
[0134] Information used to indicate uplink or downlink bandwidth;
[0135] Used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different;
[0136] The first list includes at least one channel quality value and the uplink bandwidth corresponding to the at least one channel quality value;
[0137] The second list includes at least one channel quality value and the downlink bandwidth corresponding to the at least one channel quality value;
[0138] The maximum transmission bandwidth is determined based on at least one of the uplink bandwidth or the downlink bandwidth, the information indicating whether the uplink bandwidth and the downlink bandwidth are the same or different, the first list, and the second list.
[0139] The information mentioned above used to indicate uplink or downlink bandwidth can be called the thirteenth indication, or uplink or downlink bandwidth indication, for example, 0 represents uplink bandwidth and 1 represents downlink bandwidth.
[0140] The information mentioned above used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different can be called the fourteenth indication or the uplink and downlink bandwidth same indication. For example, 0 indicates that the uplink bandwidth and downlink bandwidth are different, and 1 indicates that the uplink bandwidth and downlink bandwidth are the same.
[0141] The first list can be understood as a list of maximum uplink bandwidths, consisting of channel quality information and the corresponding uplink bandwidth. For example, the uplink signal-to-noise ratio (SINR) is used to represent channel quality, and the channel is divided into N segments based on the range of SINR values, with each segment corresponding to one uplink bandwidth.
[0142] The second list can be understood as a list of maximum downlink bandwidths, consisting of channel quality information and the corresponding downlink bandwidth. For example, the channel quality is represented by the downlink signal-to-interference-plus-noise ratio (SINR), and it is divided into N segments according to the range of SINR values, with each segment corresponding to one downlink bandwidth.
[0143] The fifth indicator can be understood as information indicating the maximum computational strength, used to define the maximum computational strength supported by the network. According to the aforementioned definition of computational strength, if the minimum bandwidth is the air interface bandwidth, then the computational strength is also related to channel quality. It can be represented by channel quality information and the corresponding computational strength.
[0144] The sixth indicator can be understood as a computing power type indicator. Potential computing power types may include at least one of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), Data Processing Unit (DPU), Smart Network Interface Card (SmartNIC), Tensor Processing Unit (TPU), and Neural Network Processing Unit (NPU).
[0145] The seventh instruction can be understood as a computing service type instruction, which is used to define the types of computing services that the network can support.
[0146] Optionally, if the computing service type supported by the network includes AI services, the computing service type is determined by at least one of the following: AI service method identifier, AI model identifier, AI model training latency, AI model training accuracy, AI model inference latency, and AI model inference accuracy.
[0147] The AI service mode identifier is used to identify at least one of image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, channel state information (CSI) feedback, and large model chat.
[0148] Specifically, one way to indicate AI service types is by defining the service method, such as using identifiers to represent image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, localization, perception, CSI feedback, or large-scale model chat, etc. Another way to indicate AI service types is by defining the AI model, such as using identifiers to represent inception_v3, resnet101_v1, yolo_v3, fasterrcnn-vgg16, deeplab_v3, or Llama3, etc.
[0149] AI service types can also be defined by AI model training latency, AI model training accuracy, AI model inference latency, and AI model inference accuracy.
[0150] For AI models, performance parameters (such as training performance parameters and inference performance parameters) typically differ across different scenarios. For instance, image recognition and object detection performance parameters correspond to top-1 accuracy and average accuracy, semantic segmentation performance parameters correspond to mean intersection over union (MIOU), and speech recognition performance parameters correspond to word error rate (WER). For large models, performance can be represented by scores from benchmark tests.
[0151] AI performance can also be indicated by test datasets; that is, by indicating which dataset yields performance no lower than the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, or Criteo, etc.
[0152] AI model inference latency can include at least one of the following situations:
[0153] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS ;
[0154] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0155] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0156] The above is a brief explanation of the information provided by First Information.
[0157] The second piece of information will be explained below.
[0158] In some embodiments, the second information includes at least one of the following:
[0159] The eighth indicator is used to indicate the probability that the terminal requests computing services;
[0160] The ninth instruction is used to indicate the terminal's potential computing service requirements;
[0161] The tenth instruction is used to indicate whether the terminal's computing service function is enabled or disabled;
[0162] The eleventh instruction is used to indicate the computing services that the terminal can provide.
[0163] The second piece of information can be understood as computational auxiliary information.
[0164] If the UE supports requesting computing services from the network, the computing assistance information can include the probability that the UE will request computing services from the network (i.e., the eighth indication). For example, when the UE has sufficient power, it can use its own GPU and other resources to complete the computing without requesting computing services from the network. In this case, the computing assistance information can indicate that the probability of the UE requesting services from the network is 0.
[0165] If the UE supports requesting computing services from the network, the computing assistance information may include computing service information that the UE may potentially (or prefer) request (i.e., the ninth indication).
[0166] Optionally, the potential computing service requirements include at least one of the following:
[0167] The resource type of the potential request, used to indicate the type of resource requested by the UE;
[0168] The maximum computation speed of a potential request is used to indicate the maximum computation speed requested by the UE.
[0169] The computing power type of the potential request is used to define the computing power type requested by the UE.
[0170] The computation service type of the potential request is used to define the type of computation service requested by the UE.
[0171] If the UE supports providing computing services to the network or other computing request nodes, the computing assistance information may include whether the UE's computing service function is enabled or disabled (i.e., the tenth indication).
[0172] If the UE supports providing computing services to the network or other computing requesting nodes, the computing assistance information may include the computing services that the UE can provide (i.e., the eleventh indication).
[0173] Optionally, the available computing services include at least one of the following:
[0174] The resource types that can provide computing services are used to indicate the resource types supported by the UE;
[0175] The maximum computing speed that can provide computing services is used to indicate the maximum computing speed supported by the UE;
[0176] The computing power type that can provide computing services is used to define the types of computing power that the UE can support;
[0177] The computing service type that can provide computing services is used to define the types of computing services that the UE can support.
[0178] It should be noted that when the terminal provides computing assistance information through the second information, the terminal can act as a computing node, and the network-side device (such as the first node) can act as a computing request node.
[0179] Furthermore, considering that the air interface bandwidth transmitted over the network is determined by the network's resource scheduling for the terminal, when the terminal provides computing services, it is difficult to provide computing services that guarantee computing strength or latency-sensitive computing strength without network assistance.
[0180] The above is a related explanation of the second piece of information.
[0181] The embodiments of this application can be applied to UEs in an idle state, that is, idle UEs obtain network computing service information by receiving first information from a first node.
[0182] The following provides relevant implementation methods using an idle UE as an example.
[0183] In some embodiments, the first information is sent via system information;
[0184] The receipt of first information from the first node includes at least one of the following:
[0185] When the terminal has a computing service requirement, it receives first information from the first node;
[0186] When the power consumption of the terminal meets the requirements, it receives the first information from the first node.
[0187] In this embodiment, the system information is broadcast information, and the system information includes MIB and SIB.
[0188] In related technologies, system information is only communication-related and cannot provide computing service information. Under the assumption that 6G supports both computing services and communication, this application proposes a method for the network to send system information containing computing service information to assist the UE in determining whether to access a network that supports computing services and which computing services are available, in order to enable idle UEs to obtain information such as whether the network supports computing services and the types and performance of the supported computing services. For the terminal, it can determine whether to receive system information to obtain computing service-related information based on whether there is a need for computing services and information such as UE power consumption.
[0189] In some embodiments, the system information includes at least one of the following:
[0190] The fifteenth instruction is used to indicate that the target SIB carries information related to computing services;
[0191] The sixteenth instruction is used to instruct the network to support computing services.
[0192] One example method is to use a bit in the system information (such as MIB) of existing communication technologies to indicate that a certain SIB (such as SIB9) includes computing service-related information. Another example method is to use a bit in the system message of existing communication technologies to indicate a computing service indication (i.e., indicating that the network supports computing services) and another bit to indicate that a certain SIB includes other computing service-related information (which can be understood as at least one other than the "computing service indication").
[0193] In addition, the update cycle T of the computing service information can be predefined or pre-configured, with the first node sending the computing service information once every T time interval.
[0194] Alternatively, the first node determines whether to send computing service information based on changes in the computing service information. For example, the first node sends computing service information when the change in computing service information meets a threshold. Another example is when the types of resources supported by the network change. Yet another example is when the change in maximum computing speed exceeds a threshold.
[0195] In some embodiments, after receiving the first information from the first node, the method further includes:
[0196] Based on the first information, the terminal determines whether the computing services supported by the network meet the terminal's needs.
[0197] When the computing services supported by the network meet the needs of the terminal, the terminal sends a first message to the first node, the first message being used for random access.
[0198] In this implementation, after receiving the first information from the first node in the idle state, the terminal determines whether the computing service information supported by the network meets the requirements. If it does, the terminal sends the first message. The first message can be called the random access first message, which can be Message 1 (MSG1) of the existing random access procedure, or it can be a message such as MSG3.
[0199] Specifically, this can include the following situations:
[0200] Scenario 1: If the UE requires computing services, then the UE receives computing service information sent by the first node. If the network supports computing services in the geographical area where the UE is located, then the UE sends a first random access message.
[0201] Scenario 2: If the UE requires computing services, then the UE receives computing service information sent by the first node. If the network supports computing services in the geographical area where the UE is located, then the UE continues to receive other computing service information. If the other service information meets the UE's computing needs, then the UE sends a first random access message.
[0202] Scenario 3: If the network does not support computing services, or the computing services supported by the network do not meet the UE's needs, then the UE can continue to search for suitable cells.
[0203] In some embodiments, the first message includes a seventeenth indication, which indicates that the reason the terminal initiated random access is related to a computing service. The seventeenth indication can be understood as a computing indication, which indicates that the UE accesses the network because of a computing service.
[0204] For the first node, the first node receives the first message sent by the terminal and can determine whether to allow the UE to access the network based on the services supported by the UE, cell load, computing resource status (such as the number of active UEs supporting computing services, available computing resources, etc.).
[0205] In some embodiments, the method further includes:
[0206] The terminal receives a second message from the first node, the second message indicating that network access is permitted, or the second message indicating that network access is not permitted.
[0207] In this embodiment, the first node sends a second message to the terminal. The first message can be called the random access second message. This message can be MSG2 of the existing random access procedure, or it can be a message such as MSG4.
[0208] Specifically, it may include at least one of the following situations:
[0209] Scenario 1: If UE access is allowed, the first node sends a second message to establish an RRC connection.
[0210] Scenario 2: If UE access is not allowed, the first node may not send the second message, or the first node may send the second message to indicate that UE access is denied. The second message may also include the reason for denial (e.g., high load).
[0211] In scenario 1, the second message may also include a computation management node identifier and / or a computation node identifier. The computation management node identifier or the computation node identifier is used by the UE to subsequently send computation service requests and / or send computation data.
[0212] In this embodiment, the computing management node may be referred to as a computing management function, computing control function, computing management and control function, computing service control function, or computing service management function, etc. The computing management node can be a radio access network node or a core network node. The computing management node can be a network node responsible for receiving or processing computing service requests, computing resource scheduling, computing information interaction, computing data processing, or at least one of these functions. The computing management node can be an enhanced AMF or an enhanced SMF, or it can be an enhanced version of another network node or a newly defined network node.
[0213] Through the above implementation methods, the idle terminal can obtain the network's computing service information through the first information, and then determine whether to initiate random access. This enables the idle UE to access a suitable cell by obtaining the network's computing service information.
[0214] The embodiments of this application can also be applied to connected UEs, that is, connected UEs obtain network computing service information by receiving first information from the first node, and then determine whether to initiate a computing service request.
[0215] The following provides relevant implementation methods using a connected UE as an example.
[0216] In some embodiments, prior to receiving the first information from the first node, the method further includes:
[0217] The terminal sends third information to the first node, the third information indicating at least one of the following:
[0218] Does the terminal support requesting computing services from the network?
[0219] Does the terminal support providing computing services?
[0220] The third information can be understood as capability information, that is, the connected terminal can send capability information to the first node before the first node sends computing service information.
[0221] Specifically, the capability information may include at least one of the following:
[0222] An indication of requesting computing service capabilities from the network, used to indicate whether the UE supports requesting computing services from the network;
[0223] The indication that provides computing services indicates whether the UE supports providing computing services to the network or other computing requesting nodes.
[0224] Optionally, if the UE supports requesting computing services from the network, then the capability information (third information) may also include at least one of the following:
[0225] The computing service type indicates the type of computing service requested by the UE from the network. If only AI services are needed, it can also be called an AI service type. Examples include one-dimensional DFT, two-dimensional FFT, AI model training, and AI model inference. AI model training or AI model inference can further include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, CSI feedback, etc.
[0226] Computing power type identifies the type of computing power service that the UE may request from the network. Potential computing power types can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU.
[0227] Resource type is used to identify the type of computing service that the UE may request from the network. Potential resource types include guaranteed computing speed, guaranteed computing strength, non-guaranteed computing speed, non-guaranteed computing strength, latency-sensitive guaranteed computing speed, or latency-sensitive guaranteed computing strength.
[0228] Optionally, if the UE supports providing computing services to the network or other computing requesting nodes, then the capability information (third information) may also include at least one of the following:
[0229] Resource type indicator, used to indicate the types of resources supported by the UE. Resource type can also be called computing resource type, or computing and communication resource type. Potential resource types include guaranteed computing speed, guaranteed computing strength, non-guaranteed computing speed, non-guaranteed computing strength, latency-sensitive guaranteed computing speed, or latency-sensitive guaranteed computing strength, etc.
[0230] Maximum computing speed indicates the maximum computing speed supported by the UE. It can be used in conjunction with GCI or delay critical GCI or GCR or delay critical GCR. Optionally, the maximum computing speed includes the theoretical maximum computing speed and / or the actual maximum computing speed.
[0231] Computing power type defines the types of computing power that the UE can support. Potential computing power types can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU.
[0232] Computing service type is used to define the types of computing services that the UE can support.
[0233] For the first node, it can send computing service information (i.e., first information) to terminals that support the request for computing services based on the received capability information, without having to send computing service information to all terminals, which reduces signaling overhead. The first node can also determine which computing service information to send to the terminal based on at least one of the computing service type, computing power type, resource type, and the terminal's channel quality information indicated in the terminal capability information. In this way, the first node does not need to send information corresponding to all computing service types, computing power types, resource types, or UE channel quality supported by the network, which further reduces signaling overhead.
[0234] In some embodiments, the first information is sent via multicast, or the first information is sent via unicast.
[0235] As one implementation, the first node can group one or more terminals according to at least one of the computing service type, computing power type, resource type and channel quality information of the terminal indicated in the terminal capability information. Terminals in the same group receive the same computing service information. In this way, the first node can send the information via multicast instead of sending computing service information to each terminal individually, which can reduce signaling overhead.
[0236] As an alternative implementation, the first node can send the corresponding computing service information to each terminal individually.
[0237] In some embodiments, when the first information is sent via multicast, receiving the first information from the first node includes:
[0238] When the terminal has a computing service requirement, it receives first information from the first node.
[0239] In this implementation, if the UE can determine whether to receive computing services and information, it can decide whether to receive computing service information based on whether it needs computing services. If the UE needs computing services, it receives the computing service information. Alternatively, the UE receives computing service information according to network configuration. Furthermore, the UE can determine whether the computing service information supported by the network meets its requirements. If it does, it sends a computing service request message.
[0240] Specifically, it may include at least one of the following situations:
[0241] Scenario 1: If the UE requires computing services, then the UE receives computing service information sent by the network. If the network supports computing services in the geographical area where the UE is located, then the UE sends a computing service request message.
[0242] Scenario 2: If the UE requires computing services, then the UE receives computing service information sent by the network. If the network supports computing services in the geographical area where the UE is located, and the computing service type, computing power type, or resource type meets the UE's computing needs, then a computing service request is sent.
[0243] Scenario 3: If the network does not support computing services, or the computing services supported by the network do not meet the UE's needs, then the UE will not send a computing service request.
[0244] In addition to being sent via multicast or unicast, the first information can also be sent periodically. Specifically, the update period T of the computing service information is predefined or pre-configured, and the first node sends the computing service information once every T time interval. Furthermore, the first node can also determine whether to send the computing service information based on changes in the computing service information. For example, the first node sends the computing service information when the changes in the computing service information meet a threshold (such as changes in the types of resources supported by the network, or changes in the maximum computing speed exceeding a threshold).
[0245] In some embodiments, the first information is also used to indicate at least one of the following:
[0246] The second node is the computing management node;
[0247] The third node is a computation node.
[0248] By instructing the computing management node, the terminal can send a computing service request to the computing management node when it needs computing services; by instructing computing, the terminal can send computing data to the computing node when it needs computing services.
[0249] In some embodiments, the method further includes:
[0250] The terminal sends a fourth message to the second node, the fourth message being used to request computing services.
[0251] The fourth message can be understood as a computing service request message. The terminal, as a computing request node, sends a computing service request message to the computing management node.
[0252] Optionally, the fourth information includes at least one of the following:
[0253] Target identifier, used to identify the target computation task;
[0254] The eighteenth instruction is used to indicate the type of resource;
[0255] The nineteenth instruction is used to indicate the operand threshold or operand limit;
[0256] The twentieth indicator is used to indicate the calculation speed threshold;
[0257] The twenty-first instruction is used to indicate the calculation strength threshold;
[0258] Instruction No. 22, used to instruct on the calculation of delayed budgets;
[0259] Instruction number twenty-three is used to indicate the maximum failure rate;
[0260] The twenty-fourth instruction is used to indicate the average statistical time window of computation speed, computation intensity, or computation delay budget;
[0261] The 25th instruction is used to indicate the duration of the calculation task;
[0262] The twenty-sixth instruction is used to indicate the type of computing power;
[0263] The twenty-seventh instruction is used to indicate data types;
[0264] The twenty-eighth instruction is used to indicate the memory threshold required for a computational task;
[0265] The twenty-ninth instruction is used to indicate the storage space threshold required for a computing task;
[0266] The thirtieth instruction is used to indicate the transmission bandwidth threshold for computing tasks;
[0267] The thirty-first instruction is used to indicate the arrival mode or parameters of a computation task;
[0268] The thirty-second instruction is used to indicate the training accuracy of an AI model;
[0269] The thirty-third instruction is used to indicate the performance parameters of the AI model;
[0270] The thirty-fourth instruction is used to indicate the inference throughput threshold of the AI model;
[0271] The thirty-fifth instruction is used to indicate the power consumption threshold for calculation;
[0272] The thirty-sixth instruction is used to indicate the energy efficiency threshold for calculation.
[0273] Specifically, the computing service request message includes at least one of the following (1) to (6):
[0274] (1) Computation task identifier (or computation service identifier), used to identify computation tasks (or computation services). Examples include one-dimensional DFT, two-dimensional FFT, AI model training, and AI model inference. AI model training or AI model inference can also include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, localization, perception, CSI feedback, etc.
[0275] (2) Resource types, also known as computing resource types, or computing and communication resource types. Potential resource types include guaranteed computing speed, guaranteed computing strength, non-guaranteed computing speed, non-guaranteed computing strength, latency-sensitive guaranteed computing speed, and latency-sensitive guaranteed computing strength.
[0276] (3) Minimum number of operations: Defines the minimum number of operations required for the computation task.
[0277] (4) Minimum computation speed: Computation speed refers to the computation time complexity divided by the tolerable upper limit of computation time. For example, computation time complexity is usually represented by operands, and the computation usually corresponds to a data type. 2 TFLOPs means that the time complexity of the computation is 2 × 10^12 floating-point operands. If the tolerable upper limit of computation time is 20ms, that is, the computation can be completed in a maximum of 20ms, then the computation speed is 2 × 10^12 / (20 × 10^(-3)) = 100 TFLOPS.
[0278] Optionally, the minimum computing speed includes the theoretical minimum computing speed and / or the actual minimum computing speed. The actual minimum computing speed can be represented by a test case indicator, which indicates the test case used to obtain the actual minimum computing speed; that is, the actual minimum computing speed is the minimum computing speed obtained based on a specific test case. In one case, there exists a default test case, and the actual minimum computing speed is the minimum computing speed based on that default test case. Alternatively, if the computing request has no specific requirements for the test case, then it is the minimum computing speed measured by any test case. If the computing request has requirements for the test case, then the test case indicator can be designated as test case A, and the actual minimum computing speed refers to the actual minimum computing speed obtained based on test case A. Alternatively, the actual minimum computing speed required for the computing task can be represented by the theoretical minimum computing speed and computing efficiency. Alternatively, the actual minimum computing speed required for the task can be represented by the ideal minimum computing speed, test case indicators, and computing efficiency. Specifically, according to the aforementioned definition of computing efficiency (actual minimum computing speed / ideal minimum computing speed), the actual minimum computing speed can be obtained by multiplying the ideal minimum computing speed by the computing efficiency. Here, the ideal situation is, for example, that no other computing tasks on the computing node preempt or share computing resources.
[0279] (5) Minimum computational intensity. Computational intensity refers to computational speed divided by bandwidth. As mentioned earlier, in the process of providing computational services in 6G, bandwidth consists of multiple parts, including the transmission bandwidth between the computation request node (such as UE) and the computational node, and the memory bandwidth of the computational node. One type of computational intensity is defined in segments, such as computational speed divided by memory bandwidth. Another type of computational intensity is to take the minimum bandwidth among the aforementioned related bandwidths as the numerator, that is, computational speed divided by min{memory bandwidth, transmission bandwidth between UE and computational node}. The transmission bandwidth between UE and computational node can be further divided into air interface bandwidth between UE and access network node, and wired transmission bandwidth between access network node and computational node. Alternatively, the transmission bandwidth between UE and computational node can be further divided into bandwidth between UE and User Plane Function Node (UPF) node, and wired transmission bandwidth between UPF and computational node.
[0280] (6) Computing delay budget (CDB): Defines the maximum tolerable delay (i.e., maximum delay) when a computing task is computed between the computing request node (such as UE) and the computing node.
[0281] The upper limit of tolerable latency during computation refers to the computation latency of the computing node.
[0282] The upper limit of tolerable latency during transmission refers to the sum of the transmission latency from the requesting node to the computing node and the transmission latency from the computing node to the receiving node.
[0283] One definition is the length of the time interval between the first data packet of a single computational task being sent and the last data packet of the received computational task. Another definition is the length of the time interval between the first data packet of a group of computational tasks being sent and the last data packet being received. For example, for image recognition computational tasks, one approach is to treat single image recognition as a single computational task, while another approach is to treat multiple images (e.g., 100 images) as a group of computational tasks.
[0284] Specifically, for example, computational latency budgets for AI model inference include at least one of the following:
[0285] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS ;
[0286] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0287] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0288] (7) Maximum failure rate: The failure rate is the number of failed requests for the computation task per unit time (e.g., per second, per 5 minutes, per day, etc.) divided by the total number of valid requests per unit time. Failed requests include incomplete computation requests and computation requests that have exceeded the latency threshold. Valid requests refer to computation requests accepted and processed by the first node.
[0289] (8) Average window: Defines the average statistical time window for calculation speed, or calculation intensity, or calculation delay budget.
[0290] (9) Maximum time: Defines the maximum duration of the computation task. It can be represented by one or more of the start time, duration, and end time.
[0291] (10) The computing power type can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU. Optionally, it may also include at least one of the following: clock speed and number of cores.
[0292] (11) Data type: can be at least one of integers (such as int8, int4, etc.) and floating-point numbers (such as half-precision floating-point numbers, single-precision floating-point numbers, double-precision floating-point numbers, double-precision floating-point numbers, etc.).
[0293] (12) Minimum memory, used to define the minimum memory required for a computing task. It can be represented by at least one of the following: memory size (e.g., 8G), sustainable memory bandwidth, and memory random access rate.
[0294] (13) Minimum storage, used to define the minimum storage required for a computing task. It can be represented by at least one of the following: storage size (e.g., 1T) and storage bandwidth (maximum I / O flow per unit time).
[0295] (14) Minimum transmission bandwidth, used to define the lower limit of uplink bandwidth and / or downlink bandwidth that the computation task can tolerate. It can be represented by at least one of the following:
[0296] Uplink or downlink bandwidth indicator, for example, 0 represents uplink bandwidth and 1 represents downlink bandwidth;
[0297] The uplink and downlink bandwidth are the same, for example, 0 means that the uplink bandwidth and downlink bandwidth are different, and 1 means that the uplink bandwidth and downlink bandwidth are the same;
[0298] Minimum uplink bandwidth;
[0299] Minimum downlink bandwidth.
[0300] (15) Calculation task arrival mode or computation job arrival mode and parameters. A computation task may contain multiple computation jobs, and the potential modes and parameters include at least one of the following:
[0301] Consecutive arrival mode or single arrival mode: The i-th job (where i is a positive integer) arrives immediately after the (i-1)-th job is completed. Job i is not sent if job (i-1) is not completed or the delay budget threshold is not met.
[0302] Fixed-period arrival mode: Jobs arrive at a fixed period T, with n jobs arriving at a time (n is a positive integer);
[0303] Poisson distribution arrival pattern: Operations based on Where k is the number of jobs arriving per unit time (k is a positive integer), and λ (λ is a positive integer) is the average number of jobs arriving per unit time (e.g., per second);
[0304] Peak arrival pattern: In the Poisson distribution arrival pattern, there are j short periods, each period has a sudden surge in a large number of jobs, and the period lasts for a certain duration T. G (e.g., 5s-10s), and maintain a certain concurrency level σ (σ is a positive integer, e.g., σ>2). 5 (Number of jobs / second), jobs arriving within a short period conform to the fixed-period arrival pattern;
[0305] Offline arrival mode: All items arrive at once;
[0306] Mixed arrival mode: Composed of more than one of the above arrival modes.
[0307] (16) AI model training accuracy: One definition of training accuracy is the data type representation of AI model parameters output by single-precision floating-point numbers, half-precision floating-point numbers, etc.
[0308] (17) AI Performance Lower Bound: The tolerable lower bound for AI model training and / or AI model inference performance. For AI models, the performance parameters usually differ in different scenarios. For example, the performance of image recognition and object detection corresponds to top-1 accuracy and average accuracy, the performance of semantic segmentation corresponds to mean intersection over union (MIOU), and the performance of speech recognition corresponds to word error rate (WER). For large models, the performance lower bound can be represented by the score of large model benchmark tests.
[0309] Alternatively, the lower limit of AI performance can also be represented by the following information:
[0310] The test dataset indicator specifies which dataset yields performance at least equal to the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, and Criteo.
[0311] (18) Minimum AI model inference throughput. For visual models, this is usually the number of images inferred per second (images / s), while for natural research models, it is the number of sentences inferred per second (sentences / s) or tokens / s. Here, "tokens" refers to words, punctuation marks, or other text units in the input text processed by the model.
[0312] (19) Calculate power consumption threshold: for example, maximum computing power consumption. One approach is to calculate the node's power consumption within the aforementioned computing latency budget. Another approach is to denote the power consumption of the computing node per unit time when it is powered on as P1, and the power consumption per unit time when processing computing tasks as P2. Then the computing power consumption is P2-P1 or the ratio of P2 to P1.
[0313] (20) Computational energy efficiency threshold: For example, minimum computational energy efficiency. Computational speed per unit time divided by computational power consumption, or the AI model inference throughput at the level of computational power consumption.
[0314] It should be noted that there are conversion relationships between some of the above parameters. Therefore, if it is necessary to indicate the amount of computation (e.g., number of operands, computation speed, computational intensity, etc.), the information of the computing service can be represented by at least one of the following parameter combinations:
[0315] Combination 1: Minimum computation speed. The computation management node can use the minimum computation speed as a parameter for the computation node.
[0316] Combination 2: Minimum number of operands or operands and computation latency budget. If the computation latency budget only includes computation latency, then the compute management node can use these two parameters as compute node selection parameters. If the computation latency budget includes end-to-end latency, then the compute management node will use the data transmission latency and computation latency allocation information, along with these two parameters, as compute node selection parameters.
[0317] Combination 3: Computing power type (e.g., CPU, clock speed, number of cores), memory, storage, and maximum time. The advantage of using these parameters as the selection criteria for compute nodes is that they are relatively static and easy to choose.
[0318] Combination 4: Based on Combination 1, increase the minimum transmission bandwidth.
[0319] Combination 5: Minimum computational intensity and computational delay budget.
[0320] In some embodiments, the method further includes:
[0321] The terminal receives a fifth message from the second node, the fifth message indicating whether to accept the terminal's computing service request or not.
[0322] For the compute management node (i.e., the second node), it can select a suitable compute node based on the compute service request message and the compute node and status information contained in the request message (such as compute power type, compute load, available compute speed, available compute intensity, available memory, available storage, compute power consumption, or compute energy efficiency). If there are suitable compute nodes, and the number is greater than or equal to one, the compute management node can indicate whether to accept the compute service request from the requesting node through the fifth information. If there are no suitable nodes, the compute management node can indicate whether to reject the compute service request from the requesting node through the fifth information. The fifth information can be understood as a compute service response message, indicating whether to accept the compute request.
[0323] In some embodiments, the fifth piece of information is used to indicate acceptance of the terminal's computing service request;
[0324] The fifth piece of information includes at least one of the following:
[0325] Target identifier, used to identify the target computation task;
[0326] The identifier of the third node, wherein the third node is a computing node;
[0327] Protocol Data Unit (PDU) session information;
[0328] Wireless bearer indication;
[0329] Physical layer channel indication;
[0330] Physical layer resource indication.
[0331] The target identifier can be understood as the identifier of the computation task.
[0332] PDU session information can include the following:
[0333] Scenario 1: If a new PDU session needs to be created, the PDU session information can include an instruction to create a PDU session;
[0334] Scenario 2: If it is necessary to modify the PDU session, then it may also include modifying the PDU session indicator and the PDU session ID;
[0335] Scenario 3: If an existing PDU is reused, it can also include the PDU session ID, Quality of Service (QoS) flow ID, QoS rules, etc.
[0336] It should be noted that if the fifth information indicates the creation or modification of a PDU session, the UE can send a PDU creation / modification request to the first node according to the indication, and the first node will send a PDU creation / modification response.
[0337] For example, if the computing node is a radio access network node, the fifth piece of information may include a radio bearer indication, such as a radio bearer ID. The radio bearer may be a signalalling radio bearer (SRB), a data radio bearer (DRB), or a new type of radio bearer (RB) (e.g., a data plane RB). If it is necessary to create or modify an existing radio bearer, the computing management node may trigger the base station to send a radio bearer add / modify message.
[0338] If the computing node is a radio access network node, then the fifth piece of information may include a physical layer channel indication and / or a physical layer resource indication. If computing data and / or computing response data are transmitted via a physical channel, the transmission latency of the corresponding computing data is expected to be further reduced.
[0339] In some embodiments, the method further includes:
[0340] The terminal sends a sixth message to the third node, the sixth message including at least one of the following:
[0341] Calculate data;
[0342] The target identifier;
[0343] The identifier of the fourth node, which is the computing receiving node.
[0344] In this embodiment, the computation requesting node can send computation data to the computation node, including a computation task identifier and an indication of the computation receiving node.
[0345] The above are implementation examples of the method on the terminal side. The following describes implementation examples of the method on the first node side.
[0346] Figure 3 shows a flowchart of an information transmission method provided in an embodiment of this application. As shown in Figure 3, the information transmission method includes the following steps:
[0347] Step 301: The first node performs a second operation, which includes at least one of the following:
[0348] Send first information to the terminal, the first information being used to indicate information related to the network's computing services;
[0349] Receive second information from the terminal, the second information being used to indicate auxiliary information related to the terminal's computing services.
[0350] In some embodiments, the first information includes at least one of the following:
[0351] The first indication is used to indicate whether the target area supports computing services;
[0352] The second indication is used to indicate the types of resources supported by the network, including types of computing resources or types of computing and communication resources;
[0353] The third indicator is used to indicate the maximum computing speed supported by the network;
[0354] The fourth indicator is used to indicate the maximum transmission bandwidth supported by the network;
[0355] The fifth indicator is used to indicate the maximum computational intensity supported by the network;
[0356] The sixth indicator is used to indicate the type of computing power supported by the network;
[0357] The seventh instruction is used to indicate the types of computing services supported by the network.
[0358] In some embodiments, the second information includes at least one of the following:
[0359] The eighth indicator is used to indicate the probability that the terminal requests computing services;
[0360] The ninth instruction is used to indicate the terminal's potential computing service requirements;
[0361] The tenth instruction is used to indicate whether the terminal's computing service function is enabled or disabled;
[0362] The eleventh instruction is used to indicate the computing services that the terminal can provide.
[0363] In some embodiments, the method further includes:
[0364] The first node receives a first message from the terminal, the first message being used for random access;
[0365] The first node determines whether to allow the terminal to access the network based on at least one of the cell load and computing resource status.
[0366] In some embodiments, the first message includes a fifteenth indication, which indicates that the reason for the terminal initiating random access is related to a computing service.
[0367] In some embodiments, the method further includes:
[0368] The first node sends a second message to the terminal, the second message indicating that network access is permitted, or the second message indicating that network access is not permitted.
[0369] In some embodiments, where the second message is used to indicate permission to access the network, the second message is also used to indicate at least one of the following:
[0370] The second node is the computing management node;
[0371] The third node is a computation node.
[0372] In some embodiments, the first information is sent via at least one of system information and target system information block (SIB);
[0373] The system information includes at least one of the following:
[0374] The sixteenth instruction is used to indicate that the target SIB carries information related to computing services;
[0375] The seventeenth instruction is used to instruct the network-side device to support computing services.
[0376] In some embodiments, before sending the first information to the terminal, the method further includes:
[0377] The first node receives third information from the terminal, the third information indicating at least one of the following:
[0378] Does the terminal support requesting computing services from the network?
[0379] Does the terminal support providing computing services?
[0380] In some embodiments, the first information is sent via multicast, or the first information is sent via unicast.
[0381] In some embodiments, the method further includes at least one of the following:
[0382] Based on the third information, the first node determines the computing service information to be sent to the terminal;
[0383] The first node determines the computing service information to be sent to the terminal based on at least one of the terminal's computing service type, computing power type, resource type, and channel quality information.
[0384] In some embodiments, the first information is also used to indicate at least one of the following:
[0385] The second node is the computing management node;
[0386] The third node is a computation node.
[0387] In some embodiments, sending the first information to the terminal includes at least one of the following:
[0388] Based on a predefined or preconfigured target period, the first information is periodically sent to the terminal according to the target period;
[0389] When the changes in the network's computing service information meet preset conditions, the first information is sent to the terminal.
[0390] For related descriptions of the embodiments of this application, please refer to the related descriptions of the method embodiments in Figure 2, which can achieve the same technical effects. To avoid repetition, they will not be described again.
[0391] The above are method embodiments for the first node side. The following describes method embodiments for the second node side.
[0392] Figure 4 shows a flowchart of an information transmission method provided in an embodiment of this application. As shown in Figure 4, the information transmission method includes the following steps:
[0393] Step 401: The second node receives fourth information from the terminal, the fourth information being used to request computing services;
[0394] Step 402: The second node determines the third node based on the fourth information, and the third node is a computing node;
[0395] Step 403: The second node sends a third message to the third node, the third message being used to request the creation of a computing task or the modification of a computing task.
[0396] The second node can be understood as the computing management node. Upon receiving a computing service request, the computing management node can select a suitable computing node based on the request message and the related computing node and status information (computing power type, computing load, available computing speed, available computing intensity, available memory, available storage, computing power consumption, and computing energy efficiency) contained within the request message. If a suitable computing node is available, its number will be greater than or equal to one. If no suitable node is available, the computing management node can reject the computing service request from the requesting node.
[0397] For example, if the compute requesting node requests use cases for optimizing mobile networks, such as beam management, positioning, sensing, CSI feedback, etc., and these are delay-critical, then a potential approach is to select a radio access network node as the compute node.
[0398] In some embodiments, the fourth information includes at least one of the following:
[0399] Target identifier, used to identify the target computation task;
[0400] The eighteenth instruction is used to indicate the type of resource;
[0401] The nineteenth instruction is used to indicate the operand threshold or operand limit;
[0402] The twentieth indicator is used to indicate the calculation speed threshold;
[0403] The twenty-first instruction is used to indicate the calculation strength threshold;
[0404] Instruction No. 22, used to instruct on the calculation of delayed budgets;
[0405] Instruction number twenty-three is used to indicate the maximum failure rate;
[0406] The twenty-fourth instruction is used to indicate the average statistical time window of computation speed, computation intensity, or computation delay budget;
[0407] The 25th instruction is used to indicate the duration of the calculation task;
[0408] The twenty-sixth instruction is used to indicate the type of computing power;
[0409] The twenty-seventh instruction is used to indicate data types;
[0410] The twenty-eighth instruction is used to indicate the memory threshold required for a computational task;
[0411] The twenty-ninth instruction is used to indicate the storage space threshold required for a computing task;
[0412] The thirtieth instruction is used to indicate the transmission bandwidth threshold for computing tasks;
[0413] The thirty-first instruction is used to indicate the arrival mode or parameters of a computation task;
[0414] The thirty-second instruction is used to indicate the training accuracy of an AI model;
[0415] The thirty-third instruction is used to indicate the performance parameters of the AI model;
[0416] The thirty-fourth instruction is used to indicate the inference throughput threshold of the AI model;
[0417] The thirty-fifth instruction is used to indicate the power consumption threshold for calculation;
[0418] The thirty-sixth instruction is used to indicate the energy efficiency threshold for calculation.
[0419] In some embodiments, the method further includes:
[0420] The second node sends a fifth message to the terminal, the fifth message indicating whether to accept the terminal's computing service request or not.
[0421] In some embodiments, the fifth piece of information is used to indicate acceptance of the terminal's computing service request;
[0422] The fifth piece of information includes at least one of the following:
[0423] Target identifier, used to identify the target computation task;
[0424] The identifier of the third node, wherein the third node is a computing node;
[0425] Protocol Data Unit (PDU) session information;
[0426] Wireless bearer indication;
[0427] Physical layer channel indication;
[0428] Physical layer resource indication.
[0429] In some embodiments, the third message includes at least one of the following:
[0430] Target identifier, used to identify the target computation task;
[0431] The thirty-seventh instruction is used to indicate the priority of the target computing task;
[0432] The thirty-eighth instruction is used to indicate whether the target computing task can obtain the target computing resource, wherein the target computing resource is the computing resource that has been allocated to the second computing task, and the priority of the second computing task is lower than the priority of the target computing task;
[0433] The thirty-ninth instruction is used to indicate whether to discard the target computing task when a third computing task is available, wherein the priority of the third computing task is higher than that of the target computing task.
[0434] The fortieth instruction is used to indicate whether it is supported to migrate the target computing task from the third node to the fourth node, where the fourth node is a computing node;
[0435] The forty-first instruction is used to indicate the guaranteed computing speed provided to the target computing task within the average window;
[0436] The forty-second instruction is used to indicate the guaranteed computational intensity provided to the target computational task within the average window;
[0437] The forty-third instruction is used to indicate the maximum computing speed provided to the target computing task;
[0438] The forty-fourth instruction is used to indicate the maximum computational intensity provided to the target computational task.
[0439] Specifically, the computing management node sends a computing task creation / modification message to the selected computing node. The computing task creation / modification message includes at least one of the following:
[0440] A computation task identifier is used to uniquely identify a computation task. The aforementioned service identifiers can be mapped one-to-one to computation tasks, or multiple computation services can be mapped to a single computation task.
[0441] Priority indicators define the relative importance of the computing resource requests for this computing task;
[0442] The preemption capability indicator defines whether the computing task can obtain computing resources that have been allocated to another computing task with lower priority;
[0443] The preemption capability indicator is defined as whether to discard a computing task in order to execute a computing task with higher priority.
[0444] Migration capability indicator, defining whether the computing task is supported to migrate from one computing node to another;
[0445] Guarantee computation speed (GBR or delay critical GBR is required), instructing the compute nodes to guarantee the computation speed provided to the computation task within the average window;
[0446] Guarantee computational intensity (GBI or delay critical GBI is required), instructing the compute nodes to guarantee the computational intensity provided to the compute task within the average window;
[0447] Maximum computing speed indicates the upper limit of the maximum computing speed that a computing node can provide for this computing task;
[0448] Maximum computational intensity indicates the upper limit of the maximum computational intensity that a computing node can provide for this computing task.
[0449] For related descriptions of the embodiments of this application, please refer to the related descriptions of the method embodiments in Figures 2 and 3, which can achieve the same technical effects. To avoid repetition, they will not be described again.
[0450] The above are method embodiments for the second node side. The following describes method embodiments for the third node side.
[0451] Figure 5 shows a flowchart of an information transmission method provided in an embodiment of this application. As shown in Figure 5, the information transmission method includes the following steps:
[0452] Step 501: The third node receives a third message from the second node, the third message being used to request the creation of a computing task or the modification of a computing task.
[0453] In some embodiments, the third message includes at least one of the following:
[0454] Target identifier, used to identify the target computation task;
[0455] The thirty-seventh instruction is used to indicate the priority of the target computing task;
[0456] The thirty-eighth instruction is used to indicate whether the target computing task can obtain the target computing resource, wherein the target computing resource is the computing resource that has been allocated to the second computing task, and the priority of the second computing task is lower than the priority of the target computing task;
[0457] The thirty-ninth instruction is used to indicate whether to discard the target computing task when a third computing task is available, wherein the priority of the third computing task is higher than that of the target computing task.
[0458] The fortieth instruction is used to indicate whether it is supported to migrate the target computing task from the third node to the fourth node, where the fourth node is a computing node;
[0459] The forty-first instruction is used to indicate the guaranteed computing speed provided to the target computing task within the average window;
[0460] The forty-second instruction is used to indicate the guaranteed computational intensity provided to the target computational task within the average window;
[0461] The forty-third instruction is used to indicate the maximum computing speed provided to the target computing task;
[0462] The forty-fourth instruction is used to indicate the maximum computational intensity provided to the target computational task.
[0463] In some embodiments, the method further includes:
[0464] The third node receives sixth information from the terminal, the sixth information including at least one of the following:
[0465] Calculate data;
[0466] The target identifier;
[0467] The identifier of the fourth node, which is the computing receiving node.
[0468] In some embodiments, the method further includes:
[0469] The third node allocates and performs computations based on the computation task identifier (i.e., the target identifier), and then sends the computed data to the computation receiving node or terminal.
[0470] For related descriptions of the embodiments of this application, please refer to the descriptions of the method embodiments in Figures 2 to 4, which can achieve the same technical effects. To avoid repetition, they will not be described again.
[0471] The following provides several specific embodiments to illustrate the interactive process of the embodiments of this application.
[0472] Example 1: Random Access Method Based on Computing Service Information
[0473] The main idea of this embodiment is that the UE in idle state obtains network computing service information and then determines whether to initiate random access. This solves the problem of the UE being unable to obtain network computing service information and the problem of how to access a suitable cell. This embodiment involves the interaction between the terminal and the network; the network refers to the first node.
[0474] As shown in Figure 6, the steps include:
[0475] Step 1: The network sends computing service information (first information), the computing service information including at least one of the following (1) to (7):
[0476] (1) Calculation service indication, for example, represented by 1 bit, where 0 indicates not supported and 1 indicates supported. For example, the default geographic region identifier is the cell. Optionally, geographic regions can also be represented by RNA, TAC, etc.
[0477] (2) Resource type indication, for example, represented by 6 bits, with each bit corresponding to a resource type. 0 indicates not supported, and 1 indicates supported.
[0478] (3) Maximum computing speed, for example, can be represented by different bits to indicate the order of magnitude of the maximum computing speed, such as GFLOPS(10). 9 Floating-point operands), TFLOPS, etc. For example, it can be represented based on a basic computing speed (e.g., 1 GFLOPS), using N bits to represent the maximum computing speed. If the value corresponding to N bits is M, then the maximum computing speed indicates the maximum computing speed M * 1 GFLOPS supported by the network. Optionally, the maximum computing speed includes the theoretical maximum computing speed and / or the actual maximum computing speed. Specifically, it can include at least one of the following:
[0479] Theoretical maximum computation speed;
[0480] Actual maximum calculation speed;
[0481] Ideally, the actual maximum computation speed;
[0482] Test case instructions;
[0483] Computational efficiency.
[0484] (4) Maximum transmission bandwidth, including at least one of the following:
[0485] Uplink or downlink bandwidth indicator, for example, 0 represents uplink bandwidth and 1 represents downlink bandwidth;
[0486] The uplink and downlink bandwidth are the same, for example, 0 means that the uplink bandwidth and downlink bandwidth are different, and 1 means that the uplink bandwidth and downlink bandwidth are the same;
[0487] Maximum uplink bandwidth list: Consists of channel quality information and corresponding uplink bandwidth. For example, the uplink signal-to-interference-plus-noise ratio (SINR) represents channel quality, and is divided into N segments based on the range of SINR values, with each segment corresponding to one uplink bandwidth;
[0488] Maximum downlink bandwidth list: Composed of channel quality information and corresponding downlink bandwidth. For example, the channel quality is represented by the downlink signal-to-interference-plus-noise ratio (SINR), and divided into N segments according to the range of SINR values, with each segment corresponding to one downlink bandwidth;
[0489] (5) Maximum computational intensity. Optionally, it can also be represented by channel quality information and the corresponding computational intensity.
[0490] (6) Computing power type;
[0491] (7) AI service type. Optionally, it may also include AI model training latency, and / or AI model training accuracy, and / or AI model inference latency, and / or AI model inference accuracy. Wherein:
[0492] For AI models, performance parameters (AI model training, AI model inference) typically differ across different scenarios. For example, image recognition and object detection performance corresponds to top-1 accuracy and average accuracy, semantic segmentation performance corresponds to mean intersection over union (MIOU), and speech recognition performance corresponds to word error rate (WER). For large models, performance can be represented by scores from large model benchmark tests.
[0493] Alternatively, AI performance can also be represented by test dataset indicators:
[0494] The test dataset indicator specifies which dataset, when used, yields performance at least equal to the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, and Criteo.
[0495] AI model inference latency includes at least one of the following:
[0496] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS ;
[0497] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0498] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0499] In step 1, one example approach is to indicate, via a bit, that a particular SIB (e.g., SIB9) includes computing service system information within existing system information. Another example approach is to use a bit within existing communication system messages as an SIB indication of computing service information and other computing service information (at least one other than the "computing service indication").
[0500] The UE determines whether to receive the indicated SIB to obtain computing service system information based on whether it has computing service requirements and information such as UE power consumption.
[0501] Optionally, the computing service information update cycle T can be set (either predefined by the protocol or configured by the network). The network sends computing service information once every T.
[0502] Optionally, the network determines whether to send computing service information based on changes in the computing service information. For example, the network sends computing service information when the changes in the computing service information meet a threshold. This could be due to changes in the types of resources the network can support, or changes in the maximum computing speed exceeding a threshold.
[0503] Step 2 (optional): The UE determines whether to receive computing service information based on whether computing services are required. If computing services are required, the UE receives the computing service information and determines whether the computing services supported by the network meet the requirements. If they do, the UE sends a first random access message. Specifically, this may include at least one of the following:
[0504] If the UE requires computing services, then the UE receives computing service information sent by the network. If the network supports computing services in the geographical area where the UE is located, then the UE sends a first random access message.
[0505] If the UE requires computing services, it receives computing service information from the network. If the network supports computing services in the UE's geographical area, the UE continues to receive other computing service information. If other service information meets the UE's computing needs, it sends a first random access message.
[0506] If the network does not support computing services, or if the computing services supported by the network do not meet the UE's needs, then the UE continues to search for suitable cells.
[0507] Optionally, the random access first message includes a calculation indication. The calculation indication is used to indicate that the UE is accessing the network because of a computing service.
[0508] Step 3 (optional): The network receives the first random access message sent by the UE and determines whether to allow the UE to access the network based on the services supported by the UE (such as the computing indication in Step 2 indicating that the UE supports computing services), cell load, computing resource status (such as the number of active UEs supporting computing services, available computing resources, etc.). The network sends a second random access message to the UE. Specifically, this may include at least one of the following:
[0509] If UE access is permitted, a second random access message is sent to establish an RRC connection. Optionally, a computation management node identifier and / or a computation node identifier may also be included. The computation management node identifier or computation node identifier is used by the UE to subsequently send computation service requests and / or send computation data.
[0510] If UE access is not allowed, then a second random access message may not be sent. Alternatively, a second random access message may be sent to indicate rejection, and optionally, a reason for rejection (e.g., high load) may be included.
[0511] Example 2: Computing service request based on computing service information
[0512] Example 1 describes a random access procedure for an idle-state UE based on computing service information. The main idea of this example is that a connected-state UE obtains network computing service information and then determines whether to initiate a computing service request. This solves the problem that existing network UEs cannot obtain network computing service information, as well as the unnecessary computing service request messages caused by the inability to obtain such information, and the increased latency of the required computing service due to service rejection. In this example, the computing management node is the second node, the computing node is the third node, and the computing receiving node is the fourth node.
[0513] As shown in Figure 7, the steps include the following:
[0514] Step 1 (optional): The UE sends the first capability information to the first node.
[0515] The first node can be a core network function node (such as AMF) or a radio access network node (such as a base station).
[0516] The parameters of the first capability information describe the UE's ability to request computing services from the network, or the UE's ability to provide computing services. Specifically, the first capability information includes at least one of the following (1) to (2):
[0517] (1) A network request for computing service capability indication, used to indicate whether the UE supports requesting computing services from the network. If the UE supports requesting computing services from the network, it may also include at least one of the following:
[0518] The computing service type indicates the type of computing service requested by the UE from the network. If only AI services are needed, it can also be called an AI service type. Examples include one-dimensional DFT, two-dimensional FFT, AI model training, and AI model inference. AI model training or AI model inference can further include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, CSI feedback, etc.
[0519] Computing power type identifies the type of computing power service that the UE may request from the network. Potential computing power types can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU.
[0520] Resource type is used to identify the type of computing service that the UE may request from the network. Potential resource types include guaranteed computing speed, guaranteed computing strength, non-guaranteed computing speed, non-guaranteed computing strength, latency-sensitive guaranteed computing speed, and latency-sensitive guaranteed computing strength.
[0521] (2) Provide computing service indication to indicate whether the UE supports providing computing services to the network or other computing request nodes.
[0522] Optionally, if the UE supports providing computing services to the network or other computing requesting nodes, it may also include at least one of the following (a) to (d):
[0523] (a) Resource type indication, used to indicate the resource types supported by the UE. Resource types can also be called computing resource types, or computing and communication resource types. Potential resource types include guaranteed computing speed, guaranteed computing strength, non-guaranteed computing speed, non-guaranteed computing strength, latency-sensitive guaranteed computing speed, and latency-sensitive guaranteed computing strength.
[0524] Computational speed refers to the computational time complexity divided by the tolerable upper limit of computation time. For example, computational time complexity is often represented by operations, and the computation usually corresponds to a data type. 2 TFLOPs means the computational time complexity is 2 × 10⁻⁶. 12 Floating-point operands. If the tolerable computation time limit is 20ms, meaning the computation can be completed in a maximum of 20ms, then the computation speed is 2 × 10⁻⁶. 12 / (20×10 -3 = 100 TFLOPS.
[0525] Computational intensity refers to computational speed divided by bandwidth. As mentioned earlier, computational speed, in the process of providing computing services in 6G, bandwidth consists of multiple components, including the transmission bandwidth between the computation requesting node (such as the network) and the memory bandwidth of the computing node. One method of defining computational intensity is segmented, such as computational speed divided by memory bandwidth. Another method is to take the minimum of the aforementioned related bandwidths as the numerator, i.e., computational speed divided by min{memory bandwidth, transmission bandwidth between the computation requesting node and the computing node}.
[0526] Considering that the air interface bandwidth transmitted over the network is determined by the network's resource scheduling for the UE, when the UE provides computing services, without network assistance, it is difficult for the UE to provide computing services that guarantee computing strength or latency-sensitive computing strength.
[0527] (b) Maximum computing speed, indicating the maximum computing speed supported by the UE. It can be used in conjunction with GCI or delay critical GCI or GCR or delay critical GCR.
[0528] Optionally, the maximum computing speed includes the theoretical maximum computing speed and / or the actual maximum computing speed. Wherein:
[0529] The theoretical maximum computing speed is the maximum computing speed supported by the UE. The peak theoretical computing speed is calculated as the UE's processor clock speed × the number of operations performed per clock cycle × the total number of cores in the system. For example, the SD 888 GPU has a clock speed of 840MHz, 256 computational logic units (meaning the processor performs 256 floating-point operations per clock cycle), and 3 cores. Therefore, the theoretical peak computing speed of this processor is 840MHz × 256 × 3 = 645.120G FLOPS (FP32), or 840MHz × 256 × 3 × 2 = 1290.24G FLOPS (FP16). Therefore, one method is to represent the theoretical maximum computing speed of a computing node using processor clock speed, number of cores, etc., while another method is to represent it using FLOPS, OPS, etc.
[0530] The actual maximum computing speed is the maximum computing speed supported by the UE. The actual number of operations is the maximum computing speed obtained through testing.
[0531] Alternatively, the actual maximum computation speed can also be indicated by at least one of the following:
[0532] Test case indicators specify which test case was used to achieve the actual maximum computational speed. Test cases can be at least one of the following: MobileNetVx (where x can be any available version number, such as 1), one-dimensional DFT, one-dimensional FFT, two-dimensional FFT, matrix multiplication, sparse linear equations, dense linear equations, YOLOvy (where y can be any available version number, such as YOLOv5), image recognition models (such as ResNet50_v1.5), and large models (such as Llama3, Llama2). Test cases can be open-source software programs or custom software programs between the UE and the network. This means that the actual maximum computational speed supported by the UE is the actual maximum computational speed that the node can achieve under the indicated test case conditions. This is represented by the actual maximum computational speed and the test case indicators.
[0533] Computational efficiency can be categorized in two ways. One method equals the ratio of the actual maximum computing speed under ideal conditions to the theoretical maximum computing speed. Therefore, the theoretical maximum computing speed and computational efficiency together represent the actual maximum computing speed supported by the UE. Another method equals the ratio of the maximum computing speed measured based on test cases to the theoretical maximum computing speed. The computational speed measured based on test cases is typically obtained from recent tests and can represent the current state of the computing node. The ideal maximum computing speed measured based on test cases refers to the best measured performance of that computing node. Therefore, the actual maximum computing speed under ideal conditions, test case indicators, and computational efficiency together represent the actual maximum computing speed supported by the UE.
[0534] (c) Computing power type, used to define the types of computing power that the UE can support. Potential computing power types can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU.
[0535] (d) Computation Service Type: This defines the types of computing services that the UE can support. If only AI services are supported, it can also be called an AI Service Type. One way to define AI Service Types is by service method, such as using identifiers to represent image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, CSI feedback, large-scale model chat, etc. Another way is by AI model, such as using identifiers to represent inception_v3, resnet101_v1, yolo_v3, fasterrcnn-vgg16, deeplab_v3, Llama3, etc.
[0536] Optionally, it may also include AI model training latency, and / or AI model training accuracy, and / or AI model inference latency, and / or AI model inference accuracy. Wherein:
[0537] For AI models, performance parameters (AI model training, AI model inference) typically differ across different scenarios. For example, image recognition and object detection performance corresponds to top-1 accuracy and average accuracy, semantic segmentation performance corresponds to mean intersection over union (MIOU), and speech recognition performance corresponds to word error rate (WER). For large models, performance can be represented by scores from large model benchmark tests.
[0538] Alternatively, AI performance can also be represented by the following test dataset:
[0539] The test dataset indicator specifies which dataset, when used, yields performance at least equal to the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, and Criteo.
[0540] AI model inference latency includes at least one of the following:
[0541] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS ;
[0542] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0543] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0544] Step 2: The first node sends computing service information (first information), which includes at least one of the following (1) to (7):
[0545] (1) Calculation service indication, for example, represented by 1 bit, where 0 indicates not supported and 1 indicates supported. For example, the default geographic region identifier is the cell. Optionally, geographic regions can also be represented by RNA, TAC, etc.
[0546] (2) Resource type indication, for example, represented by 6 bits, with each bit corresponding to a resource type. 0 indicates not supported, and 1 indicates supported.
[0547] (3) Maximum computing speed, which can be represented by different bits to indicate the order of magnitude of the maximum computing speed, such as GFLOPS, TFLOPS, etc. For example, it can be represented based on a base computing speed (e.g., 1 GFLOPS), using N bits to represent the maximum computing speed. If the value corresponding to N bits is M, then the maximum computing speed indicates the maximum computing speed supported by the network, M*1GFLOPS. Optionally, the maximum computing speed includes the theoretical maximum computing speed and / or the actual maximum computing speed. Specifically, it can include at least one of the following:
[0548] Theoretical maximum computation speed;
[0549] Actual maximum calculation speed;
[0550] Ideally, the actual maximum computation speed;
[0551] Test case instructions;
[0552] Computational efficiency.
[0553] (4) Maximum transmission bandwidth, including at least one of the following:
[0554] Uplink or downlink bandwidth indicator, for example, 0 represents uplink bandwidth and 1 represents downlink bandwidth;
[0555] The uplink and downlink bandwidth are the same, for example, 0 means that the uplink bandwidth and downlink bandwidth are different, and 1 means that the uplink bandwidth and downlink bandwidth are the same;
[0556] Maximum uplink bandwidth list: Consists of channel quality information and corresponding uplink bandwidth. For example, the uplink signal-to-interference-plus-noise ratio (SINR) represents channel quality, and is divided into N segments based on the range of SINR values, with each segment corresponding to one uplink bandwidth;
[0557] Maximum downlink bandwidth list: Composed of channel quality information and corresponding downlink bandwidth. For example, the following uses the signal-to-interference-plus-noise ratio (SINR) to represent channel quality, divided into N segments according to the range of SINR values, with each segment corresponding to one downlink bandwidth.
[0558] (5) Maximum computational intensity. Optionally, it can also be represented by channel quality information and the corresponding computational intensity.
[0559] (6) Computing power type;
[0560] (7) AI service type. Optionally, it may also include AI model training latency, and / or AI model training accuracy, and / or AI model inference latency, and / or AI model inference accuracy. Wherein:
[0561] For AI models, performance parameters (AI model training, AI model inference) typically differ across different scenarios. For example, image recognition and object detection performance corresponds to top-1 accuracy and average accuracy, semantic segmentation performance corresponds to mean intersection over union (MIOU), and speech recognition performance corresponds to word error rate (WER). For large models, performance can be represented by scores from large model benchmark tests.
[0562] Alternatively, AI performance can also be represented by the following test dataset indicators.
[0563] The test dataset indicator specifies which dataset, when used, yields performance at least equal to the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, and Criteo.
[0564] AI model inference latency includes at least one of the following:
[0565] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T.IE -T IS ;
[0566] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0567] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0568] One example approach is for the first node to send the received first capability information of the UE to the UE that supports requesting computing services from the network.
[0569] Optionally, the first node may also determine the corresponding computing service capability to send based on the computing service type, computing power type, or resource type requested by the network from the received UE first capability information, or the UE's channel quality information, instead of sending all the computing service types, computing power types, resource types, or UE channel quality information supported by the network as in Embodiment 1.
[0570] Optionally, the first node can also group one or more UEs based on computing service type, computing power type, resource type, or UE channel quality information, etc., with UEs in the same group receiving the same computing service information. Therefore, the first node can send the information via multicast or other methods, without needing to send computing service information to each UE individually.
[0571] Optionally, the computation service information update cycle T can be set. The first node sends the computation service information once every T time interval.
[0572] Optionally, the first node determines whether to send computing service information based on changes in the computing service information. For example, the network sends computing service information when the changes in the computing service information meet a threshold. This could be due to changes in the types of resources the network can support, or changes in the maximum computing speed exceeding a threshold.
[0573] Step 3: If the UE can automatically determine whether to receive computing services and information, then the UE determines whether to receive computing service information based on whether it needs computing services. If the UE needs computing services, then it receives computing service information. Alternatively, the UE receives computing service information according to network configuration. It then determines whether the computing service information supported by the network meets the requirements. If it does, then it sends a computing service request message.
[0574] Specifically, it may include at least one of the following situations:
[0575] If the UE requires computing services, then the UE receives computing service information sent by the network. If the network supports computing services in the geographical area where the UE is located, then the UE sends a computing service request message;
[0576] If the UE requires computing services, it receives computing service information from the network. If the network supports computing service types, computing power types, or resource types in the geographical area where the UE is located that meet the UE's computing needs, it sends a computing service request.
[0577] If the network does not support computing services, or if the computing services supported by the network do not meet the UE's requirements, then the UE will not send a computing service request.
[0578] The UE, acting as a computing request node, sends a computing service request message to the computing management node. This computing request message includes at least one of the following (1) to (20):
[0579] (1) Computation service identifier, used to identify computing services. Examples include one-dimensional DFT, two-dimensional FFT, AI model training, and AI model inference. AI model training or AI model inference can further include image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, localization, perception, CSI feedback, etc.
[0580] (2) Resource types, also known as computing resource types, or computing and communication resource types. Potential resource types include guaranteed computing speed, guaranteed computing intensity, non-guaranteed computing speed, non-guaranteed computing intensity, latency-sensitive guaranteed computing speed, and latency-sensitive guaranteed computing intensity.
[0581] (3) Minimum number of operands or operands: Define the minimum number of operands or operands for this computation task.
[0582] (4) Minimum computation speed: Computation speed refers to the computation time complexity divided by the tolerable upper limit of computation time. For example, computation time complexity is usually represented by operations, and the computation usually corresponds to a data type. 2 TFLOPs (floating-point operations) means that the time complexity of the computation is 2 × 10^12 floating-point operations. If the tolerable upper limit of computation time is 20ms, that is, the computation can be completed in a maximum of 20ms, then the computation speed is 2 × 10^12 / (20 × 10^(-3)) = 100 TFLOPS (floating-point operations per second).
[0583] Optionally, the minimum computing speed includes the theoretical minimum computing speed and / or the actual minimum computing speed. The actual minimum computing speed can be represented by test case indicators and the actual minimum computing speed itself. Alternatively, it can be represented by the theoretical minimum computing speed, test case indicators, and computational efficiency.
[0584] (5) Minimum computational intensity. Computational intensity refers to computational speed divided by bandwidth. As mentioned earlier, in the process of providing computational services in 6G, bandwidth consists of multiple parts, including the transmission bandwidth between the computation request node (such as UE) and the computational node, and the memory bandwidth of the computational node. One type of computational intensity is defined in segments, such as computational speed divided by memory bandwidth. Another type of computational intensity is to take the minimum of the aforementioned related bandwidths as the numerator, that is, computational speed divided by min{memory bandwidth, transmission bandwidth between UE and computational node}. The transmission bandwidth between UE and computational node can be further divided into air interface bandwidth between UE and access network node, and wired transmission bandwidth between access network node and computational node. Alternatively, the transmission bandwidth between UE and computational node can be further divided into bandwidth between UE and User Plane Function Node (UPF) node, and wired transmission bandwidth between UPF and computational node.
[0585] (6) Computing delay budget (CDB): Defines the maximum tolerable delay (i.e., maximum delay) when a computing task is computed between the computing request node (such as UE) and the computing node.
[0586] The upper limit of tolerable latency during computation refers to the computation latency of the computing node.
[0587] The upper limit of tolerable latency during transmission refers to the sum of the transmission latency from the requesting node to the computing node and the transmission latency from the computing node to the receiving node.
[0588] One definition is the length of the time interval between the first data packet of a single computational task being sent and the last data packet of the received computational task. Another definition is the length of the time interval between the first data packet of a group of computational tasks being sent and the last data packet being received. For example, for image recognition computational tasks, one approach is to treat single image recognition as a single computational task, while another approach is to treat multiple images (e.g., 100 images) as a group of computational tasks.
[0589] Specifically, computational latency budgets for AI model inference include at least one of the following:
[0590] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS ;
[0591] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0592] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0593] (7) Maximum failure rate: The failure rate is the number of failed requests for the computation task per unit time (e.g., per second, per 5 minutes, per day, etc.) divided by the total number of valid requests per unit time. Failed requests include incomplete computation requests and computation requests that have exceeded the latency threshold. Valid requests refer to computation requests accepted and processed by the first node.
[0594] (8) Average window: Defines the average statistical time window for computation speed, computation intensity, or computation delay budget.
[0595] (9) Maximum time: Defines the maximum duration of the computation task. It can be represented by one or more of the start time, duration, and end time.
[0596] (10) The computing power type can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU. Optionally, it may also include at least one of the following: clock speed; number of cores.
[0597] (11) Data type: can be at least one of integers (such as int8, int4, etc.) and floating-point numbers (such as half-precision floating-point numbers, single-precision floating-point numbers, double-precision floating-point numbers, double-precision floating-point numbers, etc.);
[0598] (12) Minimum memory, used to define the minimum memory required for a computing task. It can be represented by at least one of the following: memory size (e.g., 8G), sustainable memory bandwidth, and memory random access rate.
[0599] (13) Minimum storage, used to define the minimum storage required for a computing task. It can be represented by at least one of the following: storage size (e.g., 1T) and storage bandwidth (maximum I / O flow per unit time).
[0600] (14) Minimum transmission bandwidth, used to define the lower limit of uplink bandwidth and / or downlink bandwidth that the computation task can tolerate. It can be represented by at least one of the following:
[0601] Uplink or downlink bandwidth indicator, for example, 0 represents uplink bandwidth and 1 represents downlink bandwidth;
[0602] The uplink and downlink bandwidth are the same, for example, 0 means that the uplink bandwidth and downlink bandwidth are different, and 1 means that the uplink bandwidth and downlink bandwidth are the same;
[0603] Minimum uplink bandwidth;
[0604] Minimum downlink bandwidth.
[0605] (15) Calculation task arrival mode or computation job arrival mode and parameters. A computation task may contain multiple computation jobs, and the potential modes and parameters include at least one of the following:
[0606] Consecutive arrival mode or single arrival mode: The i-th job (where i is a positive integer) arrives immediately after the (i-1)-th job is completed. Job i is not sent if job (i-1) is not completed or the delay budget threshold is not met.
[0607] Fixed-period arrival mode: Jobs arrive at a fixed period T, with n jobs arriving at a time (n is a positive integer).
[0608] Poisson distribution arrival pattern: Operations based on Where k is the number of jobs arriving per unit time (k is a positive integer), and λ (λ is a positive integer) is the average number of jobs arriving per unit time (e.g., per second).
[0609] Peak arrival pattern: In the Poisson distribution arrival pattern, there are j short periods, each period has a sudden surge in a large number of jobs, and the period lasts for a certain duration T. G (e.g., 5s-10s), and maintain a certain concurrency level σ (σ is a positive integer, e.g., σ>2). 5 (Number of jobs / second), jobs arriving within a short period conform to the fixed-period arrival pattern.
[0610] Offline arrival mode: All items arrive at once.
[0611] Mixed arrival mode: Composed of more than one of the above arrival modes.
[0612] (16) AI model training accuracy: One definition of training accuracy is the data type representation of AI model parameters output by single-precision floating-point numbers, half-precision floating-point numbers, etc.
[0613] (17) AI Performance Lower Bound: The tolerable lower bound for AI model training and / or AI model inference performance. For AI models, performance parameters typically differ across scenarios. For example, image recognition and object detection performance corresponds to top-1 accuracy and average accuracy, semantic segmentation performance corresponds to Mean Intersection Over Union (MIOU), and speech recognition performance corresponds to Word Error Rate (WER). For large models, the performance lower bound can be represented by the score of a large model benchmark test. Optionally, the AI performance lower bound can also be represented by at least one of the following:
[0614] The test dataset indicator specifies which dataset, when used, yields performance at least equal to the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, and Criteo.
[0615] (18) Minimum AI model inference throughput. For visual models, this is usually the number of images inferred per second (images / s), while for natural research models, it is the number of sentences inferred per second (sentences / s) or tokens / s. Here, "tokens" refers to words, punctuation marks, or other text units in the input text processed by the model.
[0616] (19) Calculate power consumption threshold: for example, maximum computing power consumption. One approach is to calculate the node's power consumption within the aforementioned computing latency budget. Another approach is to denote the power consumption of the computing node per unit time when it is powered on as P1, and the power consumption per unit time when processing computing tasks as P2. Then the computing power consumption is P2-P1 or the ratio of P2 to P1.
[0617] (20) Computational energy efficiency threshold: For example, minimum computational energy efficiency. Computational speed per unit time divided by computational power consumption, or the AI model inference throughput at the level of computational power consumption.
[0618] Among these parameters, there are conversion relationships between them. Therefore, if it is necessary to indicate the amount of computation (e.g., number of operands, computation speed, computational intensity, etc.), the information of the computing service can be represented by at least one combination of the following parameters:
[0619] Combination 1: Minimum computation speed. The computation management node can use the minimum computation speed as a parameter for the computation node;
[0620] Combination 2: Minimum number of operands or operands and computation latency budget. If the computation latency budget only includes computation latency, then the compute management node can use these two parameters as compute node selection parameters. If the computation latency budget includes end-to-end latency, then the compute management node will use the data transmission latency and computation latency allocation information, along with these two parameters, as compute node selection parameters.
[0621] Combination 3: Computing power type (e.g., CPU, clock speed, number of cores), memory, storage, and maximum time. The advantage of using these parameters as selection criteria for compute nodes is their relative static nature, making selection easier. A potential disadvantage is that it hinders the reuse of computing resources.
[0622] Combination 4: Based on Combination 1, increase the minimum transmission bandwidth;
[0623] Combination 5: Minimum computational intensity and computational delay budget.
[0624] Step 4 (Optional): The compute management node selects a suitable compute node based on the compute request information and the compute node and status information (compute type, compute load, available compute speed, available compute intensity, available memory, available storage, compute power consumption, compute energy efficiency) related to the parameters contained in the request message. If there are suitable compute nodes, their number is greater than or equal to one. If there are no suitable nodes, then in step 6, the compute management node rejects the compute request from the requesting node. If the service identifier indicates a use case for optimizing the mobile network, such as beam management, positioning, sensing, CSI feedback, etc., and is of a delay-critical type, then a potential approach is to select a radio access network node as the compute node.
[0625] Optionally, the computing management node sends a computing task creation / modification message to the selected computing node. The computing task creation / modification message includes at least one of the following (1) to (9):
[0626] (1) Computation task identifier, used to uniquely identify a computation task. The aforementioned service identifiers and computation tasks can be mapped one-to-one, or multiple computation services can be mapped to one computation task.
[0627] (2) Priority indicator, which defines the relative importance of the computing resource requests for the computing task.
[0628] (3) Preemption capability indicator, which defines whether the computing task can obtain computing resources that have been allocated to another computing task with lower priority.
[0629] (4) Preemption capability indicator, defined as whether to discard a computing task in order to execute a computing task with higher priority.
[0630] (5) Migration capability indicator, which defines whether the computing task is supported to migrate from one computing node to another.
[0631] (6) Guarantee computation speed (GBR or delay critical GBR is required), instructing the computing node to guarantee the computation speed provided to the computing task within the average window.
[0632] (7) Guarantee computation intensity (GBI or delay critical GBI is required), instructing the computing node to guarantee the computation intensity provided to the computing task within the average window.
[0633] (8) Maximum computing speed indicates the upper limit of the maximum computing speed that the computing node can provide for the computing task.
[0634] (9) Maximum computational intensity indicates the upper limit of the maximum computational intensity that a computing node can provide to the computing task.
[0635] Step 5 (optional): The compute node sends a compute task creation / modification response to the compute management node. This response message may also include updated status information of the compute node after the compute task is created or modified, such as compute load, compute speed, compute intensity, available memory, available storage, compute power consumption, and compute energy efficiency.
[0636] Step 6: The compute management node sends a compute service response message to the compute requesting node to indicate whether it accepts the compute request. If accepted, the message may also include the compute task identifier and / or the compute node identifier.
[0637] Optionally, it may also include PDU session information. Specifically, it may include at least one of the following:
[0638] If a new PDU session needs to be created, then an instruction to create a PDU session can also be included;
[0639] If it is necessary to modify the PDU session, it may also include modifying the PDU session indicator and the PDU session ID;
[0640] If an existing PDU is reused, it can also include PDU session ID, QoS flow ID, QoS rule, etc.
[0641] Optionally, if the computing node is a radio access network node, it may also include a radio bearer indication, including a radio bearer ID. The radio bearer may be an SRB, DRB, or a new type of RB (e.g., a data plane RB). If it is necessary to create or modify an existing radio bearer, the computing management node can trigger the base station to send a radio bearer add / modify message.
[0642] Optionally, if the computing node is a radio access network node, it may also include a physical layer channel indicator and / or a physical layer resource indicator. If computing data and / or computing response data are transmitted only through physical channels, the transmission latency of the corresponding computing data is expected to be further reduced.
[0643] Steps 7 and 8 (optional): If step 6 indicates to create a new PDU session or modify a PDU session, then the UE can send a PDU creation / modification request to the first node according to the indication, and the first node sends a PDU creation / modification response.
[0644] Step 9 (optional): The computing service requesting node sends computing data to the computing node based on the information in step 6, including the computing task identifier. Optionally, it may also include a computing receiving node indication.
[0645] Step 10 (optional): 10a. The computing node allocates computing resources and performs computation based on the computing task identifier, and then sends the computed data (i.e., computation response data) to the computing request node; 10b. The computing node allocates computing resources and performs computation based on the computing task identifier, and then sends the computed data (i.e., computation response data) to the computing receiving node.
[0646] This embodiment is applicable to core network nodes as computing management nodes and computing nodes, as well as core network nodes as computing management nodes and radio access network nodes as computing nodes, and radio access network nodes as computing management nodes and computing nodes.
[0647] Example 3: UE Calculation of Auxiliary Information
[0648] Example 1 describes a random access procedure for an idle-state UE based on computing service information, while Example 2 describes a connected-state UE obtaining network computing service information and determining whether to initiate a computing service request. The main idea of this example is that the UE actively sends its computing service information to the network so that the network can promptly understand the UE's status and better allocate and schedule computing resources. In this example, the network can be understood as the first node.
[0649] As shown in Figure 8, the steps include:
[0650] Step 1 (optional): The UE obtains network computing service information. As shown in Example 1 or Example 2, the computing service information can be obtained in idle state or connected state. This example will not be described in detail.
[0651] Step 2: The UE sends computational assistance information (second information) to the network, the computational assistance information including at least one of the following (a) to (d):
[0652] (a) If the UE supports requesting computing services from the network, then the probability of the UE requesting computing services from the network can be included. For example, when the UE has sufficient power, it can use the UE's GPU and other resources to complete the computation without requesting computing services from the network. Therefore, the probability of the UE requesting services from the network can be represented by the computational auxiliary information as 0.
[0653] (b) If the UE supports requesting computing services from the network, then the computing service information that the UE may potentially (prefer) request may specifically include at least one of the following (1) to (4):
[0654] (1) Resource type indication, used to indicate the type of resource requested by the UE. The resource type can also be called the computing resource type, or the computing and communication resource type. Potential resource types include guaranteed computing speed, guaranteed computing strength, non-guaranteed computing speed, non-guaranteed computing strength, latency-sensitive guaranteed computing speed, and latency-sensitive guaranteed computing strength.
[0655] Computational speed refers to the computational time complexity divided by the tolerable upper limit of computation time. For example, computational time complexity is often represented by operations, and the computation usually corresponds to a data type. 2 TFLOPs means the computational time complexity is 2 × 10⁻⁶. 12 Floating-point operands. If the tolerable computation time limit is 20ms, meaning the computation can be completed in a maximum of 20ms, then the computation speed is 2 × 10⁻⁶. 12 / (20×10 -3 = 100 TFLOPS.
[0656] Computational intensity refers to computational speed divided by bandwidth. As mentioned earlier, computational speed, in the process of providing computing services in 6G, bandwidth consists of multiple components, including the transmission bandwidth between the computation requesting node (such as the network) and the memory bandwidth of the computing node. One method of defining computational intensity is segmented, such as computational speed divided by memory bandwidth. Another method is to take the minimum of the aforementioned related bandwidths as the numerator, i.e., computational speed divided by min{memory bandwidth, transmission bandwidth between the computation requesting node and the computing node}.
[0657] (2) Maximum computing speed, used to indicate the maximum computing speed requested by the UE. It can be used in conjunction with GCI or delay critical GCI or GCR or delay critical GCR. Optionally, the maximum computing speed includes the theoretical maximum computing speed and / or the actual maximum computing speed. Wherein:
[0658] The theoretical maximum computing speed is the theoretical maximum computing speed requested by the UE.
[0659] The actual maximum computing speed is the actual maximum computing speed requested by the UE. The actual number of operations is the maximum computing speed obtained through testing. Alternatively, the actual maximum computing speed can also be indicated by the following test cases:
[0660] Test case indicators specify which test case was used to achieve the actual maximum computational speed. Test cases can be at least one of the following: MobileNetVx (where x can be any available version number, such as 1), one-dimensional DFT, one-dimensional FFT, two-dimensional FFT, matrix multiplication, sparse linear equations, dense linear equations, YOLOvy (where y can be any available version number, such as YOLOv5), image recognition models (such as ResNet50_v1.5), and large models (such as Llama3, Llama2). Test cases can be open-source software programs or custom software programs between the UE and the network. This means that the actual maximum computational speed requested by the UE is the actual maximum computational speed that the node can achieve under the indicated test case conditions. This is represented by the actual maximum computational speed and the test case indicators.
[0661] Computational efficiency can be categorized in two ways. One method equals the ratio of the actual maximum computational speed under ideal conditions to the theoretical maximum computational speed. This allows us to represent the actual maximum computational speed requested by the UE using both the theoretical maximum computational speed and computational efficiency. Another method equals the ratio of the maximum computational speed measured based on test cases to the theoretical maximum computational speed. The computational speed measured based on test cases is typically obtained from recent tests and represents the current state of the computing node. The ideal maximum computational speed measured based on test cases refers to the best measured performance of that computing node. Therefore, the actual maximum computational speed requested by the UE can be represented by the actual maximum computational speed under ideal conditions, test case indicators, and computational efficiency.
[0662] (3) Computing power type, used to define the computing power type requested by the UE. Potential computing power types can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU.
[0663] (4) Computation Service Type: This defines the type of computation service requested by the UE. If only AI services are supported, it can also be called an AI Service Type. One way to define an AI Service Type is by service method, such as using identifiers to represent image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, CSI feedback, large model chat, etc. Another way is by AI model, such as using identifiers to represent inception_v3, resnet101_v1, yolo_v3, fasterrcnn-vgg16, deeplab_v3, Llama3, etc. Optionally, it can also include AI model training latency, and / or AI model training accuracy, and / or AI model inference latency, and / or AI model inference accuracy. Where:
[0664] For AI models, performance parameters (AI model training, AI model inference) typically differ across different scenarios. For example, image recognition and object detection performance corresponds to top-1 accuracy and average accuracy, semantic segmentation performance corresponds to mean intersection over union (MIOU), and speech recognition performance corresponds to word error rate (WER). For large models, performance can be represented by scores from large model benchmark tests.
[0665] Alternatively, AI performance can also be represented by the following test dataset indicators.
[0666] The test dataset indicator specifies which dataset, when used, yields performance at least equal to the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, and Criteo.
[0667] AI model inference latency includes at least one of the following:
[0668] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS ;
[0669] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0670] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0671] (c) If the UE supports providing computing services to the network or other computing requesting nodes, then it may include enabling or disabling the UE's provision of computing services.
[0672] (d) If the UE supports providing computing services to the network or other computing requesting nodes, then it may include at least one of the following (1) to (4):
[0673] (1) Resource type indication, used to indicate the resource types supported by the UE. Resource types can also be called computing resource types, or computing and communication resource types. Potential resource types include guaranteed computing speed, guaranteed computing strength, non-guaranteed computing speed, non-guaranteed computing strength, latency-sensitive guaranteed computing speed, and latency-sensitive guaranteed computing strength.
[0674] Computational speed refers to the computational time complexity divided by the tolerable upper limit of computation time. For example, computational time complexity is often represented by operations, and the computation usually corresponds to a data type. 2 TFLOPs means the computational time complexity is 2 × 10⁻⁶. 12 Floating-point operands. If the tolerable computation time limit is 20ms, meaning the computation can be completed in a maximum of 20ms, then the computation speed is 2 × 10⁻⁶. 12 / (20×10 -3 = 100 TFLOPS.
[0675] Computational intensity refers to computational speed divided by bandwidth. As mentioned earlier, computational speed, in the process of providing computing services in 6G, bandwidth consists of multiple components, including the transmission bandwidth between the computation requesting node (such as the network) and the memory bandwidth of the computing node. One method of defining computational intensity is segmented, such as computational speed divided by memory bandwidth. Another method is to take the minimum of the aforementioned related bandwidths as the numerator, i.e., computational speed divided by min{memory bandwidth, transmission bandwidth between the computation requesting node and the computing node}.
[0676] Considering that the air interface bandwidth transmitted over the network is determined by the network's resource scheduling for the UE, when the UE provides computing services, without network assistance, it is difficult for the UE to provide computing services that guarantee computing strength or latency-sensitive computing strength.
[0677] (2) Maximum computing speed, indicating the maximum computing speed supported by the UE. It can be used in conjunction with GCI or delay critical GCI or GCR or delay critical GCR. Optionally, the maximum computing speed includes the theoretical maximum computing speed and / or the actual maximum computing speed.
[0678] The theoretical maximum computing speed is the maximum computing speed supported by the UE. The peak theoretical computing speed is calculated as the UE's processor clock speed × the number of operations performed per clock cycle × the total number of system cores. For example, the SD 888 GPU has a clock speed of 840MHz, 256 computational logic units (meaning the processor performs 256 floating-point operations per clock cycle), and 3 cores. Therefore, the theoretical peak computing speed of this processor is 840MHz × 256 × 3 = 645.120G FLOPS (FP32) or 840MHz × 256 × 3 × 2 = 1290.24G FLOPS (FP16). Therefore, one method is to represent the theoretical maximum computing speed of a computing node using processor clock speed, number of cores, etc., while another method is to represent it using FLOPS, OPS, etc.
[0679] The actual maximum computing speed is the maximum computing speed supported by the UE. The actual number of operations is the maximum computing speed obtained through testing.
[0680] Alternatively, the actual maximum computation speed can also be indicated by the following test dataset:
[0681] Test case indicators specify which test case was used to achieve the actual maximum computational speed. Test cases can be at least one of the following: MobileNetVx (where x can be any available version number, such as 1), one-dimensional DFT, one-dimensional FFT, two-dimensional FFT, matrix multiplication, sparse linear equations, dense linear equations, YOLOvy (where y can be any available version number, such as YOLOv5), image recognition models (such as ResNet50_v1.5), and large models (such as Llama3, Llama2). Test cases can be open-source software programs or custom software programs between the UE and the network. This means that the actual maximum computational speed supported by the UE is the actual maximum computational speed that the node can achieve under the indicated test case conditions. This is represented by the actual maximum computational speed and the test case indicators.
[0682] Computational efficiency can be categorized in two ways. One method equals the ratio of the actual maximum computing speed under ideal conditions to the theoretical maximum computing speed. Therefore, the theoretical maximum computing speed and computational efficiency together represent the actual maximum computing speed supported by the UE. Another method equals the ratio of the maximum computing speed measured based on test cases to the theoretical maximum computing speed. The computational speed measured based on test cases is typically obtained from recent tests and can represent the current state of the computing node. The ideal maximum computing speed measured based on test cases refers to the best measured performance of that computing node. Therefore, the actual maximum computing speed under ideal conditions, test case indicators, and computational efficiency together represent the actual maximum computing speed supported by the UE.
[0683] (3) Computing power type, used to define the types of computing power that the UE can support. Potential computing power types can be at least one of CPU, GPU, FPGA, DPU, SmartNIC, TPU, and NPU.
[0684] (4) Computation Service Type: This defines the types of computing services that the UE can support. If only AI services are supported, it can also be called an AI Service Type. One way to define AI Service Types is by service method, such as using identifiers to represent image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, CSI feedback, large-scale model chat, etc. Another way is by AI model, such as using identifiers to represent inception_v3, resnet101_v1, yolo_v3, fasterrcnn-vgg16, deeplab_v3, Llama3, etc.
[0685] Optionally, it may also include AI model training latency, and / or AI model training accuracy, and / or AI model inference latency, and / or AI model inference accuracy. Wherein:
[0686] For AI models, performance parameters (AI model training, AI model inference) typically differ across different scenarios. For example, image recognition and object detection performance corresponds to top-1 accuracy and average accuracy, semantic segmentation performance corresponds to mean intersection over union (MIOU), and speech recognition performance corresponds to word error rate (WER). For large models, performance can be represented by scores from large model benchmark tests.
[0687] Alternatively, AI performance can also be represented by the following test dataset indicators.
[0688] The test dataset indicator specifies which dataset, when used, yields performance at least equal to the stated AI performance lower limit. Commonly used AI datasets include ImageNet2012, Pascal VOC2012, LibriSpeech ASR Corpus, and Criteo.
[0689] AI model inference latency includes at least one of the following:
[0690] Total inference latency: Specifically, it refers to the total end-to-end latency of multiple consecutive inference operations. The calculation method is as follows: the time before the requesting node sends the first byte of the first computation task (or computation job) is denoted as T. IS The last byte received by the receiving node from all computing tasks (or computing jobs) is denoted as T. IE Then the computational latency budget for AI model inference is T. IE -T IS ;
[0691] End-to-end inference latency: Specifically, it refers to the difference between the time it takes to send a sample and the time it takes to receive a result. That is, the time elapsed before the requesting node sends the first byte of a computation task (or job) is denoted as t. TIS The last byte received by the receiving node for the computation task (or job) is denoted as t. TIE Then the computational latency budget for AI model inference is t. TIE -t TIS ;
[0692] Inference latency: Specifically, it refers to the difference between the start time and the end time of inference for a given sample. That is, the time taken before inference for a computational task (or job) is denoted as t. INS The last byte received by the receiving node for the computation task (or job) is denoted as t. INE Then the computational latency budget for AI model inference is t. INE -t INS .
[0693] The solution implemented in this application allows the UE to obtain computing service information provided by the network in either idle or connected mode, or the UE to provide computing assistance information. Obtaining the computing service information provided by the network solves the problem of how a UE with computing needs can access a suitable network and initiate computing service requests. The computing assistance information helps the UE save power or dissipate heat, and also helps the network better allocate and schedule computing resources.
[0694] The information transmission method provided in this application can be executed by an information transmission device. This application uses an information transmission device executing the information transmission method as an example to illustrate the information transmission device provided in this application.
[0695] This application provides an information transmission device. As an example, the information transmission device may be a communication device or a component within a communication device, such as a chip. The communication device may be a terminal, a network-side device, or a server, etc. Exemplarily, the terminal may include, but is not limited to, the type of terminal 11 listed above, and the network-side device may include, but is not limited to, the type of network-side device 12 listed above. This application does not impose specific limitations.
[0696] The information transmission device includes a receiving module, a transmitting module, and a processing module. These modules can be implemented in software or hardware. When implemented in hardware, the processing module can be implemented by a processor. For example, the processor can include general-purpose processors, special-purpose processors, such as a Central Processing Unit (CPU), microprocessor, Digital Signal Processor (DSP), Artificial Intelligence (AI) processor, Graphics Processing Unit (GPU), Application Specific Integrated Circuit (ASIC), Network Processor (NP), Field Programmable Gate Array (FPGA), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The receiving and transmitting modules can be implemented by a communication interface, which can include one or more of the following: transceiver, pins, circuits, bus, radio frequency unit, etc.
[0697] Specifically, referring to Figure 9, when the information transmission device is a terminal or a component within a terminal, the information transmission device 900 includes:
[0698] The first receiving module 901 is used to receive first information from the first node, the first information being used to indicate information related to the computing services of the network;
[0699] The first sending module 902 is used to send second information to the first node, the second information being used to indicate auxiliary information related to the computing service of the terminal.
[0700] Optionally, the first information includes at least one of the following:
[0701] The first indication is used to indicate whether the target area supports computing services;
[0702] The second indication is used to indicate the types of resources supported by the network, including types of computing resources or types of computing and communication resources;
[0703] The third indicator is used to indicate the maximum computing speed supported by the network;
[0704] The fourth indicator is used to indicate the maximum transmission bandwidth supported by the network;
[0705] The fifth indicator is used to indicate the maximum computational intensity supported by the network;
[0706] The sixth indicator is used to indicate the type of computing power supported by the network;
[0707] The seventh instruction is used to indicate the types of computing services supported by the network.
[0708] Optionally, the second information includes at least one of the following:
[0709] The eighth indicator is used to indicate the probability that the terminal requests computing services;
[0710] The ninth instruction is used to indicate the terminal's potential computing service requirements;
[0711] The tenth instruction is used to indicate whether the terminal's computing service function is enabled or disabled;
[0712] The eleventh instruction is used to indicate the computing services that the terminal can provide.
[0713] Optionally, the resource type includes at least one of the following:
[0714] The first type, the resources corresponding to the first type are those that can guarantee computing speed;
[0715] The second type, the resources corresponding to the second type are those that can guarantee computing power;
[0716] The third type, the resources corresponding to the third type are resources that do not guarantee computing speed;
[0717] The fourth type refers to resources that do not guarantee computational intensity.
[0718] The fifth type refers to resources that are latency-sensitive and guarantee computing speed.
[0719] The sixth type refers to resources that are latency-sensitive and guarantee computational intensity.
[0720] Optionally, the maximum computing speed includes at least one of the following:
[0721] Theoretical maximum computation speed;
[0722] Actual maximum computing speed.
[0723] Optionally, the third instruction includes at least one of the following:
[0724] The theoretical maximum computation speed;
[0725] Information used to indicate the target test case;
[0726] First computational efficiency, which is the ratio of the actual maximum computational speed to the theoretical maximum computational speed;
[0727] The second computational efficiency is the ratio of the maximum computational speed measured based on the target test case to the theoretical maximum computational speed.
[0728] The third computational efficiency is the ratio of the maximum computational speed measured based on the target test case to the maximum computational speed under ideal conditions.
[0729] The actual maximum computing speed is determined by at least one of the theoretical maximum computing speed, the target test case, the first computing efficiency, the second computing efficiency, and the third computing efficiency.
[0730] Optionally, the fourth instruction includes at least one of the following:
[0731] Information used to indicate uplink or downlink bandwidth;
[0732] Used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different;
[0733] The first list includes at least one channel quality value and the uplink bandwidth corresponding to the at least one channel quality value;
[0734] The second list includes at least one channel quality value and the downlink bandwidth corresponding to the at least one channel quality value;
[0735] The maximum transmission bandwidth is determined based on at least one of the uplink bandwidth or the downlink bandwidth, the information indicating whether the uplink bandwidth and the downlink bandwidth are the same or different, the first list, and the second list.
[0736] Optionally, if the computing service type supported by the network includes artificial intelligence (AI) services, the computing service type is determined by at least one of the following: AI service method identifier, AI model identifier, AI model training latency, AI model training accuracy, AI model inference latency, and AI model inference accuracy.
[0737] The AI service mode identifier is used to identify at least one of image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, channel state information (CSI) feedback, and large model chat.
[0738] Optionally, the potential computing service requirements include at least one of the following:
[0739] The type of resource that is being requested;
[0740] The maximum computational speed of potential requests;
[0741] The type of computing power requested;
[0742] The type of computing service requested.
[0743] Optionally, the available computing services include at least one of the following:
[0744] The types of resources that can provide computing services;
[0745] The maximum computing speed that can provide computing services;
[0746] The types of computing power that can provide computing services;
[0747] The types of computing services that can provide computing services.
[0748] Optionally, the first information is sent via system information;
[0749] The first receiving module is specifically used for at least one of the following:
[0750] When the terminal has a computing service requirement, it receives first information from the first node;
[0751] When the power consumption of the terminal meets the requirements, it receives the first information from the first node.
[0752] Optionally, the system information includes at least one of the following:
[0753] The fifteenth instruction is used to indicate that the target SIB carries information related to computing services;
[0754] The sixteenth instruction is used to instruct the network to support computing services.
[0755] Optionally, the device further includes:
[0756] The processing module is used to determine, based on the first information, whether the computing services supported by the network meet the needs of the terminal.
[0757] The second sending module is used to send a first message to the first node when the computing services supported by the network meet the needs of the terminal. The first message is used for random access.
[0758] Optionally, the first message includes a seventeenth indication, which indicates that the reason for the terminal initiating random access is related to the computing service.
[0759] Optionally, the device further includes:
[0760] The second receiving module is configured to receive a second message from the first node, the second message indicating that network access is permitted, or the second message indicating that network access is not permitted.
[0761] Optionally, the device further includes:
[0762] A third sending module is configured to send third information to the first node, the third information indicating at least one of the following:
[0763] Does the terminal support requesting computing services from the network?
[0764] Does the terminal support providing computing services?
[0765] Optionally, the first information is sent via multicast, or the first information is sent via unicast.
[0766] Optionally, when the first information is sent via multicast, the first receiving module is specifically used for:
[0767] When the terminal has a computing service requirement, it receives first information from the first node.
[0768] Optionally, the first information is also used to indicate at least one of the following:
[0769] The second node is the computing management node;
[0770] The third node is a computation node.
[0771] Optionally, the device further includes:
[0772] The fourth sending module is used to send fourth information to the second node, the fourth information being used to request computing services.
[0773] Optionally, the fourth information includes at least one of the following:
[0774] Target identifier, used to identify the target computation task;
[0775] The eighteenth instruction is used to indicate the type of resource;
[0776] The nineteenth instruction is used to indicate the operand threshold or operand limit;
[0777] The twentieth indicator is used to indicate the calculation speed threshold;
[0778] The twenty-first instruction is used to indicate the calculation strength threshold;
[0779] Instruction No. 22, used to instruct on the calculation of delayed budgets;
[0780] Instruction number twenty-three is used to indicate the maximum failure rate;
[0781] The twenty-fourth instruction is used to indicate the average statistical time window of computation speed, computation intensity, or computation delay budget;
[0782] The 25th instruction is used to indicate the duration of the calculation task;
[0783] The twenty-sixth instruction is used to indicate the type of computing power;
[0784] The twenty-seventh instruction is used to indicate data types;
[0785] The twenty-eighth instruction is used to indicate the memory threshold required for a computational task;
[0786] The twenty-ninth instruction is used to indicate the storage space threshold required for a computing task;
[0787] The thirtieth instruction is used to indicate the transmission bandwidth threshold for computing tasks;
[0788] The thirty-first instruction is used to indicate the arrival mode or parameters of a computation task;
[0789] The thirty-second instruction is used to indicate the training accuracy of an AI model;
[0790] The thirty-third instruction is used to indicate the performance parameters of the AI model;
[0791] The thirty-fourth instruction is used to indicate the inference throughput threshold of the AI model;
[0792] The thirty-fifth instruction is used to indicate the power consumption threshold for calculation;
[0793] The thirty-sixth instruction is used to indicate the energy efficiency threshold for calculation.
[0794] Optionally, the device further includes:
[0795] The third receiving module is used to receive fifth information from the second node, the fifth information being used to indicate whether to accept the computing service request from the terminal or not to accept the computing service request from the terminal.
[0796] Optionally, the fifth piece of information is used to indicate acceptance of the terminal's computing service request;
[0797] The fifth piece of information includes at least one of the following:
[0798] Target identifier, used to identify the target computation task;
[0799] The identifier of the third node, wherein the third node is a computing node;
[0800] Protocol Data Unit (PDU) session information;
[0801] Wireless bearer indication;
[0802] Physical layer channel indication;
[0803] Physical layer resource indication.
[0804] Optionally, the device further includes:
[0805] The fifth sending module is used to send sixth information to the third node, the sixth information including at least one of the following:
[0806] Calculate data;
[0807] The target identifier;
[0808] The identifier of the fourth node, which is the computing receiving node.
[0809] Referring to Figure 10, when the information transmission device is a network-side device or a component of a network-side device, the information transmission device 1000 includes at least one of the following:
[0810] The first sending module 1001 is used to send first information to the terminal, the first information being used to indicate information related to the network's computing services;
[0811] The first receiving module 1002 is used to receive second information from the terminal, the second information being used to indicate auxiliary information related to the terminal's computing services.
[0812] Optionally, the first information includes at least one of the following:
[0813] The first indication is used to indicate whether the target area supports computing services;
[0814] The second indication is used to indicate the types of resources supported by the network, including types of computing resources or types of computing and communication resources;
[0815] The third indicator is used to indicate the maximum computing speed supported by the network;
[0816] The fourth indicator is used to indicate the maximum transmission bandwidth supported by the network;
[0817] The fifth indicator is used to indicate the maximum computational intensity supported by the network;
[0818] The sixth indicator is used to indicate the type of computing power supported by the network;
[0819] The seventh instruction is used to indicate the types of computing services supported by the network.
[0820] Optionally, the second information includes at least one of the following:
[0821] The eighth indicator is used to indicate the probability that the terminal requests computing services;
[0822] The ninth instruction is used to indicate the terminal's potential computing service requirements;
[0823] The tenth instruction is used to indicate whether the terminal's computing service function is enabled or disabled;
[0824] The eleventh instruction is used to indicate the computing services that the terminal can provide.
[0825] Optionally, the device further includes:
[0826] The second receiving module is used to receive a first message from the terminal, the first message being used for random access;
[0827] The first processing module is used to determine whether to allow the terminal to access the network based on at least one of the cell load and computing resource status.
[0828] Optionally, the first message includes a fifteenth indication, which indicates that the reason for the terminal initiating random access is related to the computing service.
[0829] Optionally, the device further includes:
[0830] The second sending module is used to send a second message to the terminal, the second message being used to indicate that network access is allowed, or the second message being used to indicate that network access is not allowed.
[0831] Optionally, when the second message is used to indicate permission to access the network, the second message is also used to indicate at least one of the following:
[0832] The second node is the computing management node;
[0833] The third node is a computation node.
[0834] Optionally, the first information is sent via system information;
[0835] The system information includes at least one of the following:
[0836] The sixteenth instruction is used to indicate that the target SIB carries information related to computing services;
[0837] The seventeenth instruction is used to instruct the network-side device to support computing services.
[0838] Optionally, the device further includes:
[0839] A third receiving module is configured to receive third information from the terminal, the third information indicating at least one of the following:
[0840] Does the terminal support requesting computing services from the network?
[0841] Does the terminal support providing computing services?
[0842] Optionally, the first information is sent via multicast, or the first information is sent via unicast.
[0843] Optionally, the device further includes at least one of the following:
[0844] The second processing module is used to determine the computing service information to be sent to the terminal based on the third information;
[0845] The third processing module is used to determine the computing service information to be sent to the terminal based on at least one of the terminal's computing service type, computing power type, resource type, and channel quality information.
[0846] Optionally, the first information is also used to indicate at least one of the following:
[0847] The second node is the computing management node;
[0848] The third node is a computation node.
[0849] Optionally, the first sending module is specifically used for at least one of the following:
[0850] Based on a predefined or preconfigured target period, the first information is periodically sent to the terminal according to the target period;
[0851] When the changes in the network's computing service information meet preset conditions, the first information is sent to the terminal.
[0852] Referring to Figure 11, when the information transmission device is a network-side device or a component within a network-side device, the information transmission device 1100 includes:
[0853] The receiving module 1101 is used to receive fourth information from the terminal, the fourth information being used to request computing services;
[0854] Processing module 1102 is used to determine a third node based on the fourth information, wherein the third node is a computing node;
[0855] The first sending module 1103 is used to send a third message to the third node, the third message being used to request the establishment of a computing task or the modification of a computing task.
[0856] Optionally, the fourth information includes at least one of the following:
[0857] Target identifier, used to identify the target computation task;
[0858] The eighteenth instruction is used to indicate the type of resource;
[0859] The nineteenth instruction is used to indicate the operand threshold or operand limit;
[0860] The twentieth indicator is used to indicate the calculation speed threshold;
[0861] The twenty-first instruction is used to indicate the calculation strength threshold;
[0862] Instruction No. 22, used to instruct on the calculation of delayed budgets;
[0863] Instruction number twenty-three is used to indicate the maximum failure rate;
[0864] The twenty-fourth instruction is used to indicate the average statistical time window of computation speed, computation intensity, or computation delay budget;
[0865] The 25th instruction is used to indicate the duration of the calculation task;
[0866] The twenty-sixth instruction is used to indicate the type of computing power;
[0867] The twenty-seventh instruction is used to indicate data types;
[0868] The twenty-eighth instruction is used to indicate the memory threshold required for a computational task;
[0869] The twenty-ninth instruction is used to indicate the storage space threshold required for a computing task;
[0870] The thirtieth instruction is used to indicate the transmission bandwidth threshold for computing tasks;
[0871] The thirty-first instruction is used to indicate the arrival mode or parameters of a computation task;
[0872] The thirty-second instruction is used to indicate the training accuracy of an AI model;
[0873] The thirty-third instruction is used to indicate the performance parameters of the AI model;
[0874] The thirty-fourth instruction is used to indicate the inference throughput threshold of the AI model;
[0875] The thirty-fifth instruction is used to indicate the power consumption threshold for calculation;
[0876] The thirty-sixth instruction is used to indicate the energy efficiency threshold for calculation.
[0877] Optionally, the device further includes:
[0878] The second sending module is used to send fifth information to the terminal, the fifth information being used to indicate whether to accept the terminal's computing service request or not to accept the terminal's computing service request.
[0879] Optionally, the fifth piece of information is used to indicate acceptance of the terminal's computing service request;
[0880] The fifth piece of information includes at least one of the following:
[0881] Target identifier, used to identify the target computation task;
[0882] The identifier of the third node, wherein the third node is a computing node;
[0883] Protocol Data Unit (PDU) session information;
[0884] Wireless bearer indication;
[0885] Physical layer channel indication;
[0886] Physical layer resource indication.
[0887] Optionally, the third message includes at least one of the following:
[0888] Target identifier, used to identify the target computation task;
[0889] The thirty-seventh instruction is used to indicate the priority of the target computing task;
[0890] The thirty-eighth instruction is used to indicate whether the target computing task can obtain the target computing resource, wherein the target computing resource is the computing resource that has been allocated to the second computing task, and the priority of the second computing task is lower than the priority of the target computing task;
[0891] The thirty-ninth instruction is used to indicate whether to discard the target computing task when a third computing task is available, wherein the priority of the third computing task is higher than that of the target computing task.
[0892] The fortieth instruction is used to indicate whether it is supported to migrate the target computing task from the third node to the fourth node, where the fourth node is a computing node;
[0893] The forty-first instruction is used to indicate the guaranteed computing speed provided to the target computing task within the average window;
[0894] The forty-second instruction is used to indicate the guaranteed computational intensity provided to the target computational task within the average window;
[0895] The forty-third instruction is used to indicate the maximum computing speed provided to the target computing task;
[0896] The forty-fourth instruction is used to indicate the maximum computational intensity provided to the target computational task.
[0897] Referring to Figure 12, when the information transmission device is a network-side device or a component within a network-side device, the information transmission device 1200 includes:
[0898] The first receiving module 1201 is used to receive a third message from the second node, the third message being used to request the establishment of a computing task or the modification of a computing task.
[0899] Optionally, the third message includes at least one of the following:
[0900] Target identifier, used to identify the target computation task;
[0901] The thirty-seventh instruction is used to indicate the priority of the target computing task;
[0902] The thirty-eighth instruction is used to indicate whether the target computing task can obtain the target computing resource, wherein the target computing resource is the computing resource that has been allocated to the second computing task, and the priority of the second computing task is lower than the priority of the target computing task;
[0903] The thirty-ninth instruction is used to indicate whether to discard the target computing task when a third computing task is available, wherein the priority of the third computing task is higher than that of the target computing task.
[0904] The fortieth instruction is used to indicate whether it is supported to migrate the target computing task from the third node to the fourth node, where the fourth node is a computing node;
[0905] The forty-first instruction is used to indicate the guaranteed computing speed provided to the target computing task within the average window;
[0906] The forty-second instruction is used to indicate the guaranteed computational intensity provided to the target computational task within the average window;
[0907] The forty-third instruction is used to indicate the maximum computing speed provided to the target computing task;
[0908] The forty-fourth instruction is used to indicate the maximum computational intensity provided to the target computational task.
[0909] Optionally, the device further includes:
[0910] The second receiving module is configured to receive sixth information from the terminal, the sixth information including at least one of the following:
[0911] Calculate data;
[0912] The target identifier;
[0913] The identifier of the fourth node, which is the computing receiving node.
[0914] The information transmission device provided in this application embodiment can implement the various processes implemented in the method embodiments of Figures 2 to 5 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0915] As shown in Figure 13, this application embodiment also provides a communication device 1300, including a processor 1301 and a memory 1302. The memory 1302 stores a program or instructions that can run on the processor 1301. For example, when the communication device 1300 is a terminal, the program or instructions executed by the processor 1301 implement the various steps of the above-described terminal-side method embodiments and achieve the same technical effect. When the communication device 1300 is a network-side device, the program or instructions executed by the processor 1301 implement the various steps of the above-described first node-side, second node-side, or third node-side method embodiments and achieve the same technical effect. To avoid repetition, these will not be described again here.
[0916] This application also provides a terminal, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in the method embodiment shown in FIG2. This terminal embodiment corresponds to the above-described terminal-side method embodiment, and all implementation processes and methods of the above-described method embodiments can be applied to this terminal embodiment and can achieve the same technical effect. The terminal may be the information transmission device shown in FIG9. Specifically, FIG14 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application.
[0917] The terminal 1400 includes, but is not limited to, at least some of the following components: radio frequency unit 1401, network module 1402, audio output unit 1403, input unit 1404, sensor 1405, display unit 1406, user input unit 1407, interface unit 1408, memory 1409, and processor 1410.
[0918] Those skilled in the art will understand that the terminal 1400 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to the processor 1410 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The terminal structure shown in Figure 14 does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0919] It should be understood that, in this embodiment, the input unit 1404 may include a graphics processor 14041 and a microphone 14042. The graphics processor 14041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1406 may include a display panel 14061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1407 includes at least one of a touch panel 14071 and other input devices 14072. The touch panel 14071 is also called a touch screen. The touch panel 14071 may include a touch detection device and a touch controller. Other input devices 14072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0920] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1401 can transmit it to the processor 1410 for processing; in addition, the radio frequency unit 1401 can send uplink data to the network-side device. Typically, the radio frequency unit 1401 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0921] The memory 1409 can be used to store software programs or instructions, as well as various data. The memory 1409 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1409 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1409 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0922] Processor 1410 may include one or more processing units; optionally, processor 1410 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1410.
[0923] The radio frequency unit 1401 is used for at least one of the following:
[0924] Receive first information from the first node, the first information being used to indicate information related to the network's computing services;
[0925] Send a second message to the first node, the second message being used to indicate auxiliary information related to the terminal's computing services.
[0926] In this embodiment, since the terminal receives information related to network computing services, it can obtain network computing services based on this information when a computing service requirement exists. Furthermore, since the terminal sends auxiliary information related to computing services, it can provide assistance to the computing services. Therefore, this embodiment can effectively meet the computing service requirements of communication scenarios.
[0927] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the information transmission method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.
[0928] This application also provides a network-side device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method embodiment shown in FIG3, FIG4, or FIG5. This network-side device embodiment corresponds to the method embodiments of the first node side, second node side, or third node side described above. All implementation processes and methods of the above method embodiments can be applied to this network-side device embodiment and can achieve the same technical effect.
[0929] Specifically, this application embodiment also provides a network-side device, which may be the information transmission device shown in FIG10, FIG11, or FIG12. As shown in FIG15, the network-side device 1500 includes: an antenna 151, a radio frequency device 152, a baseband device 153, a processor 154, and a memory 155. The antenna 151 is connected to the radio frequency device 152. In the uplink direction, the radio frequency device 152 receives information through the antenna 151 and sends the received information to the baseband device 153 for processing. In the downlink direction, the baseband device 153 processes the information to be transmitted and sends it to the radio frequency device 152. The radio frequency device 152 processes the received information and transmits it through the antenna 151.
[0930] The method executed by the network-side device in the above embodiments can be implemented in the baseband device 153, which includes a baseband processor.
[0931] The baseband device 153 may include at least one baseband board, on which multiple chips are disposed, as shown in FIG15. One of the chips is, for example, a baseband processor, which is connected to the memory 155 via a bus interface to call the program in the memory 155 and execute the node operations shown in the above method embodiments.
[0932] The network-side device may also include a network interface 156, such as a Common Public Radio Interface (CPRI).
[0933] Specifically, the network-side device 1500 in this application embodiment further includes: instructions or programs stored in memory 155 and executable on processor 154. Processor 154 calls the instructions or programs in memory 155 to execute the methods executed by the modules shown in FIG10, FIG11 or FIG12 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0934] Specifically, this application also provides a network-side device. As shown in FIG16, the network-side device 1600 includes a processor 1601, a network interface 1602, and a memory 1603. The network-side device may be the information transmission device shown in FIG10, FIG11, or FIG12. The network interface 1602 is, for example, a Common Public Radio Interface (CPRI).
[0935] Specifically, the network-side device 1600 in this application embodiment further includes: instructions or programs stored in memory 1603 and executable on processor 1601. Processor 1601 calls the instructions or programs in memory 1603 to execute the methods executed by the modules shown in FIG10, FIG11 or FIG12 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0936] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described information transmission method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0937] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0938] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described information transmission method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0939] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0940] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described information transmission method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0941] This application also provides a communication system, including: a terminal and a first node, wherein the terminal can be used to execute the steps of the information transmission method described above, and the first node can be used to execute the steps of the information transmission method on the first node side as described above.
[0942] Optionally, the communication system further includes a second node, which can be used to perform the steps of the information transmission method described above for the second node side.
[0943] Optionally, the communication system further includes a third node, which can be used to perform the steps of the information transmission method described above for the third node side.
[0944] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0945] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0946] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. An information transmission method, comprising: The terminal performs a first operation, which includes at least one of the following: Receive first information from the first node, the first information being used to indicate information related to the network's computing services; Send a second message to the first node, the second message being used to indicate auxiliary information related to the terminal's computing services.
2. The method according to claim 1, wherein, The first information includes at least one of the following: The first indication is used to indicate whether the target area supports computing services; The second indication is used to indicate the types of resources supported by the network, including types of computing resources or types of computing and communication resources; The third indicator is used to indicate the maximum computing speed supported by the network; The fourth indicator is used to indicate the maximum transmission bandwidth supported by the network; The fifth indicator is used to indicate the maximum computational intensity supported by the network; The sixth indicator is used to indicate the type of computing power supported by the network; The seventh instruction is used to indicate the types of computing services supported by the network; or, The second information includes at least one of the following: The eighth indicator is used to indicate the probability that the terminal requests computing services; The ninth instruction is used to indicate the terminal's potential computing service requirements; The tenth instruction is used to indicate whether the terminal's computing service function is enabled or disabled; The eleventh instruction is used to indicate the computing services that the terminal can provide.
3. The method according to claim 2, wherein, The resource type includes at least one of the following: The first type, the resources corresponding to the first type are those that can guarantee computing speed; The second type, the resources corresponding to the second type are those that can guarantee computing power; The third type, the resources corresponding to the third type are resources that do not guarantee computing speed; The fourth type refers to resources that do not guarantee computational intensity. The fifth type refers to resources that are latency-sensitive and guarantee computing speed. The sixth type refers to resources that are latency-sensitive and guarantee computational intensity. or, The maximum computing speed includes at least one of the following: Theoretical maximum computation speed; Actual maximum calculation speed; or, The potential computing service requirements include at least one of the following: The type of resource that is being requested; The maximum computational speed of potential requests; The type of computing power requested; The type of computing service requested; or, The available computing services include at least one of the following: The types of resources that can provide computing services; The maximum computing speed that can provide computing services; The types of computing power that can provide computing services; The types of computing services that can provide computing services.
4. The method according to claim 3, wherein, The third instruction includes at least one of the following: The theoretical maximum computation speed; Information used to indicate the target test case; First computational efficiency, which is the ratio of the actual maximum computational speed to the theoretical maximum computational speed; The second computational efficiency is the ratio of the maximum computational speed measured based on the target test case to the theoretical maximum computational speed. The third computational efficiency is the ratio of the maximum computational speed measured based on the target test case to the maximum computational speed under ideal conditions. The actual maximum computing speed is determined by at least one of the theoretical maximum computing speed, the target test case, the first computing efficiency, the second computing efficiency, and the third computing efficiency.
5. The method according to claim 2, wherein, The fourth instruction includes at least one of the following: Information used to indicate uplink or downlink bandwidth; Used to indicate whether the uplink bandwidth and downlink bandwidth are the same or different; The first list includes at least one channel quality value and the uplink bandwidth corresponding to the at least one channel quality value; The second list includes at least one channel quality value and the downlink bandwidth corresponding to the at least one channel quality value; The maximum transmission bandwidth is determined based on at least one of the uplink bandwidth or the downlink bandwidth, the information indicating whether the uplink bandwidth and the downlink bandwidth are the same or different, the first list, and the second list.
6. The method according to claim 2, wherein, When the computing service types supported by the network include artificial intelligence (AI) services, the computing service type is determined by at least one of the following: AI service method identifier, AI model identifier, AI model training latency, AI model training accuracy, AI model inference latency, and AI model inference accuracy. The AI service mode identifier is used to identify at least one of image recognition, object detection, semantic segmentation, recommendation, natural language processing, speech recognition, optical character recognition, face recognition, beam management, positioning, perception, channel state information (CSI) feedback, and large model chat.
7. The method according to any one of claims 1 to 6, wherein, The first information is sent via system information; The receipt of first information from the first node includes at least one of the following: When the terminal has a computing service requirement, it receives first information from the first node; When the power consumption of the terminal meets the requirements, it receives the first information from the first node.
8. The method according to claim 7, wherein, The system information includes at least one of the following: The fifteenth instruction is used to indicate that the target system information block (SIB) carries information related to computing services; The sixteenth instruction is used to instruct the network to support computing services.
9. The method according to claim 7 or 8, wherein after receiving the first information from the first node, the method further comprises: Based on the first information, the terminal determines whether the computing services supported by the network meet the terminal's needs. When the computing services supported by the network meet the needs of the terminal, the terminal sends a first message to the first node, the first message being used for random access.
10. The method according to claim 9, wherein, The first message includes a seventeenth instruction, which indicates that the reason for the terminal to initiate random access is related to the computing service.
11. The method according to claim 9 or 10, further comprising: The terminal receives a second message from the first node, the second message indicating that network access is permitted, or the second message indicating that network access is not permitted.
12. The method according to any one of claims 1 to 6, wherein before receiving the first information from the first node, the method further comprises: The terminal sends third information to the first node, the third information indicating at least one of the following: Does the terminal support requesting computing services from the network? Does the terminal support providing computing services? 13. The method according to any one of claims 1 to 12, wherein, When the first information is sent via multicast, receiving the first information from the first node includes: When the terminal has a computing service requirement, it receives first information from the first node.
14. The method according to any one of claims 1 to 6, 12 to 13, wherein, The first information is also used to indicate at least one of the following: The second node is the computing management node; The third node is a computation node.
15. The method according to claim 14, further comprising: The terminal sends a fourth message to the second node, the fourth message being used to request computing services.
16. The method according to claim 15, wherein, The fourth piece of information includes at least one of the following: Target identifier, used to identify the target computation task; The eighteenth instruction is used to indicate the type of resource; The nineteenth instruction is used to indicate the operand threshold or operand limit; The twentieth indicator is used to indicate the calculation speed threshold; The twenty-first instruction is used to indicate the calculation strength threshold; Instruction No. 22, used to instruct on the calculation of delayed budgets; Instruction number twenty-three is used to indicate the maximum failure rate; The twenty-fourth instruction is used to indicate the average statistical time window of computation speed, computation intensity, or computation delay budget; The 25th instruction is used to indicate the duration of the calculation task; The twenty-sixth instruction is used to indicate the type of computing power; The twenty-seventh instruction is used to indicate data types; The twenty-eighth instruction is used to indicate the memory threshold required for a computational task; The twenty-ninth instruction is used to indicate the storage space threshold required for a computing task; The thirtieth instruction is used to indicate the transmission bandwidth threshold for computing tasks; The thirty-first instruction is used to indicate the arrival mode or parameters of a computation task; The thirty-second instruction is used to indicate the training accuracy of an AI model; The thirty-third instruction is used to indicate the performance parameters of the AI model; The thirty-fourth instruction is used to indicate the inference throughput threshold of the AI model; The thirty-fifth instruction is used to indicate the power consumption threshold for calculation; The thirty-sixth instruction is used to indicate the energy efficiency threshold for calculation.
17. The method according to claim 15 or 16, further comprising: The terminal receives a fifth message from the second node, the fifth message indicating whether to accept the terminal's computing service request or not.
18. The method according to claim 17, wherein, The fifth piece of information is used to indicate acceptance of the terminal's computing service request; The fifth piece of information includes at least one of the following: Target identifier, used to identify the target computation task; The identifier of the third node, wherein the third node is a computing node; Protocol Data Unit (PDU) session information; Wireless bearer indication; Physical layer channel indication; Physical layer resource indication.
19. The method according to any one of claims 14 to 18, the method further comprising: The terminal sends a sixth message to the third node, the sixth message including at least one of the following: Calculate data; Target identifier, used to identify the target computation task; The identifier of the fourth node, which is the computing receiving node.
20. An information transmission method, comprising: The first node performs a second operation, which includes at least one of the following: Send first information to the terminal, the first information being used to indicate information related to the network's computing services; Receive second information from the terminal, the second information being used to indicate auxiliary information related to the terminal's computing services.
21. The method according to claim 20, wherein, The first information includes at least one of the following: The first indication is used to indicate whether the target area supports computing services; The second indication is used to indicate the types of resources supported by the network, including types of computing resources or types of computing and communication resources; The third indicator is used to indicate the maximum computing speed supported by the network; The fourth indicator is used to indicate the maximum transmission bandwidth supported by the network; The fifth indicator is used to indicate the maximum computational intensity supported by the network; The sixth indicator is used to indicate the type of computing power supported by the network; The seventh instruction is used to indicate the types of computing services supported by the network; or, The second information includes at least one of the following: The eighth indicator is used to indicate the probability that the terminal requests computing services; The ninth instruction is used to indicate the terminal's potential computing service requirements; The tenth instruction is used to indicate whether the terminal's computing service function is enabled or disabled; The eleventh instruction is used to indicate the computing services that the terminal can provide.
22. The method according to claim 20 or 21, further comprising: The first node receives a first message from the terminal, the first message being used for random access; The first node determines whether to allow the terminal to access the network based on at least one of the cell load and computing resource status.
23. The method according to claim 22, wherein, The first message includes a fifteenth instruction, which indicates that the reason for the terminal to initiate random access is related to the computing service.
24. The method according to claim 22 or 23, further comprising: The first node sends a second message to the terminal, the second message indicating that network access is permitted, or the second message indicating that network access is not permitted.
25. The method according to claim 24, wherein, When the second message is used to indicate permission to access the network, the second message is also used to indicate at least one of the following: The second node is the computing management node; The third node is a computation node.
26. The method according to any one of claims 20 to 25, wherein, The first information is sent via system information; The system information includes at least one of the following: The fifteenth instruction is used to indicate that the target system information block (SIB) carries information related to computing services; The sixteenth instruction is used to instruct the network to support computing services.
27. The method according to claim 20 or 21, wherein before sending the first information to the terminal, the method further comprises: The first node receives third information from the terminal, the third information indicating at least one of the following: Does the terminal support requesting computing services from the network? Does the terminal support providing computing services? 28. The method of claim 27, further comprising at least one of the following: Based on the third information, the first node determines the computing service information to be sent to the terminal; The first node determines the computing service information to be sent to the terminal based on at least one of the terminal's computing service type, computing power type, resource type, and channel quality information.
29. The method according to any one of claims 20 to 21, 27 to 28, wherein, The first information is also used to indicate at least one of the following: The second node is the computing management node; The third node is a computation node.
30. The method according to any one of claims 20 to 29, wherein, Sending the first information to the terminal includes at least one of the following: Based on a predefined or preconfigured target period, the first information is periodically sent to the terminal according to the target period; When the changes in the network's computing service information meet preset conditions, the first information is sent to the terminal.
31. An information transmission method, comprising: The second node receives a fourth message from the terminal, which is used to request computing services; The second node determines the third node based on the fourth information, and the third node is a computing node; The second node sends a third message to the third node, the third message being used to request the creation of a computing task or the modification of a computing task.
32. The method according to claim 31, wherein, The fourth piece of information includes at least one of the following: Target identifier, used to identify the target computation task; The eighteenth instruction is used to indicate the type of resource; The nineteenth instruction is used to indicate the operand threshold or operand limit; The twentieth indicator is used to indicate the calculation speed threshold; The twenty-first instruction is used to indicate the calculation strength threshold; Instruction No. 22, used to instruct on the calculation of delayed budgets; Instruction number twenty-three is used to indicate the maximum failure rate; The twenty-fourth instruction is used to indicate the average statistical time window of computation speed, computation intensity, or computation delay budget; The 25th instruction is used to indicate the duration of the calculation task; The twenty-sixth instruction is used to indicate the type of computing power; The twenty-seventh instruction is used to indicate data types; The twenty-eighth instruction is used to indicate the memory threshold required for a computational task; The twenty-ninth instruction is used to indicate the storage space threshold required for a computing task; The thirtieth instruction is used to indicate the transmission bandwidth threshold for computing tasks; The thirty-first instruction is used to indicate the arrival mode or parameters of a computation task; The thirty-second instruction is used to indicate the training accuracy of an AI model; The thirty-third instruction is used to indicate the performance parameters of the AI model; The thirty-fourth instruction is used to indicate the inference throughput threshold of the AI model; The thirty-fifth instruction is used to indicate the power consumption threshold for calculation; The thirty-sixth instruction is used to indicate the energy efficiency threshold for calculation.
33. The method according to claim 31 or 32, further comprising: The second node sends a fifth message to the terminal, the fifth message indicating whether to accept the terminal's computing service request or not.
34. The method according to claim 33, wherein, The fifth piece of information is used to indicate acceptance of the terminal's computing service request; The fifth piece of information includes at least one of the following: Target identifier, used to identify the target computation task; The identifier of the third node, wherein the third node is a computing node; Protocol Data Unit (PDU) session information; Wireless bearer indication; Physical layer channel indication; Physical layer resource indication.
35. The method according to any one of claims 31 to 34, wherein, The third message includes at least one of the following: Target identifier, used to identify the target computation task; The thirty-seventh instruction is used to indicate the priority of the target computing task; The thirty-eighth instruction is used to indicate whether the target computing task can obtain the target computing resource, wherein the target computing resource is the computing resource that has been allocated to the second computing task, and the priority of the second computing task is lower than the priority of the target computing task; The thirty-ninth instruction is used to indicate whether to discard the target computing task when a third computing task is available, wherein the priority of the third computing task is higher than that of the target computing task. The fortieth instruction is used to indicate whether it is supported to migrate the target computing task from the third node to the fourth node, where the fourth node is a computing node; The forty-first instruction is used to indicate the guaranteed computing speed provided to the target computing task within the average window; The forty-second instruction is used to indicate the guaranteed computational intensity provided to the target computational task within the average window; The forty-third instruction is used to indicate the maximum computing speed provided to the target computing task; The forty-fourth instruction is used to indicate the maximum computational intensity provided to the target computational task.
36. An information transmission method, comprising: The third node receives a third message from the second node, which is used to request the creation of a computing task or the modification of a computing task.
37. The method of claim 36, wherein, The third message includes at least one of the following: Target identifier, used to identify the target computation task; The thirty-seventh instruction is used to indicate the priority of the target computing task; The thirty-eighth instruction is used to indicate whether the target computing task can obtain the target computing resource, wherein the target computing resource is the computing resource that has been allocated to the second computing task, and the priority of the second computing task is lower than the priority of the target computing task; The thirty-ninth instruction is used to indicate whether to discard the target computing task when a third computing task is available, wherein the priority of the third computing task is higher than that of the target computing task. The fortieth instruction is used to indicate whether it is supported to migrate the target computing task from the third node to the fourth node, where the fourth node is a computing node; The forty-first instruction is used to indicate the guaranteed computing speed provided to the target computing task within the average window; The forty-second instruction is used to indicate the guaranteed computational intensity provided to the target computational task within the average window; The forty-third instruction is used to indicate the maximum computing speed provided to the target computing task; The forty-fourth instruction is used to indicate the maximum computational intensity provided to the target computational task.
38. The method according to claim 36 or 37, further comprising: The third node receives sixth information from the terminal, the sixth information including at least one of the following: Calculate data; Target identifier, used to identify the target computation task; The identifier of the fourth node, which is the computing receiving node.
39. An information transmission device, the device comprising at least one of the following: The first receiving module is used to receive first information from the first node, the first information being used to indicate information related to the network's computing services; The first sending module is used to send second information to the first node, the second information being used to indicate auxiliary information related to the terminal's computing services.
40. The apparatus of claim 39, further comprising: The processing module is used to determine, based on the first information, whether the computing services supported by the network meet the needs of the terminal. The second sending module is used to send a first message to the first node when the computing services supported by the network meet the needs of the terminal. The first message is used for random access.
41. The apparatus of claim 40, further comprising: The second receiving module is configured to receive a second message from the first node, the second message indicating that network access is permitted, or the second message indicating that network access is not permitted.
42. The apparatus of claim 39, further comprising: A third sending module is configured to send third information to the first node, the third information indicating at least one of the following: Does the terminal support requesting computing services from the network? Does the terminal support providing computing services? 43. An information transmission device, the device comprising at least one of the following: The first sending module is used to send first information to the terminal, the first information being used to indicate information related to the network's computing services; A first receiving module is configured to receive second information from a terminal, the second information being used to indicate auxiliary information related to the terminal's computing services.
44. The apparatus of claim 43, further comprising: The second receiving module is used to receive a first message from the terminal, the first message being used for random access; The first processing module is used to determine whether to allow the terminal to access the network based on at least one of the cell load and computing resource status.
45. The apparatus of claim 44, further comprising: The second sending module is used to send a second message to the terminal, the second message being used to indicate that network access is allowed, or the second message being used to indicate that network access is not allowed.
46. The apparatus of claim 43, further comprising: A third receiving module is configured to receive third information from the terminal, the third information indicating at least one of the following: Does the terminal support requesting computing services from the network? Does the terminal support providing computing services? 47. The apparatus of claim 46, further comprising at least one of the following: The second processing module is used to determine the computing service information to be sent to the terminal based on the third information; The third processing module is used to determine the computing service information to be sent to the terminal based on at least one of the terminal's computing service type, computing power type, resource type, and channel quality information.
48. An information transmission device, comprising: A receiving module is used to receive fourth information from the terminal, the fourth information being used to request computing services; The processing module is used to determine a third node based on the fourth information, wherein the third node is a computing node; The first sending module is used to send a third message to the third node, the third message being used to request the establishment of a computing task or the modification of a computing task.
49. An information transmission device, comprising: The first receiving module is used to receive a third message from the second node, the third message being used to request the establishment of a computing task or the modification of a computing task.
50. A communication device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the information transmission method as claimed in any one of claims 1 to 19, or implementing the steps of the information transmission method as claimed in any one of claims 20 to 30, or implementing the steps of the information transmission method as claimed in any one of claims 31 to 35, or implementing the steps of the information transmission method as claimed in any one of claims 36 to 38.
51. A readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the information transmission method as claimed in any one of claims 1 to 19, or the steps of the information transmission method as claimed in any one of claims 20 to 30, or the steps of the information transmission method as claimed in any one of claims 31 to 35, or the steps of the information transmission method as claimed in any one of claims 36 to 38.
52. A computer program product comprising computer instructions which, when executed by a processor, implement the steps of the information transmission method as claimed in any one of claims 1 to 19, or the steps of the information transmission method as claimed in any one of claims 20 to 30, or the steps of the information transmission method as claimed in any one of claims 31 to 35, or the steps of the information transmission method as claimed in any one of claims 36 to 38.
Citation Information
Patent Citations
Terminal-cloud cooperative computing architecture, task scheduling device and task scheduling method
CN107087019A
Resource allocation method and device, computing equipment and computer readable storage medium
CN113069760A
Edge cloud application sensing method based on wireless ad hoc network environment
CN114466412A
Network attached reconfigurable computing device
US20180054359A1
Multi-tenant support on virtual machines in cloud computing networks
US20200097310A1