Communication method, communication apparatus, and communication system

By determining the quality of service flow and adding an identifier to the data packet, the end-to-end latency problem between the terminal device and the computing node is solved, improving communication efficiency and the processing power of the computing node, and reducing latency and jitter.

WO2026091922A1PCT designated stage Publication Date: 2026-05-07HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-09-12
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

In application scenarios where computing nodes are deployed, there is still no effective solution for ensuring end-to-end latency between terminal devices and computing nodes.

Method used

By obtaining the end-to-end latency and computation latency of the computing task, the corresponding quality of service flow is determined, and the quality of service flow identifier is added to the data packet, so as to achieve accurate control and scheduling of the computing task and ensure that the computing task is completed within the specified latency.

Benefits of technology

It achieves accurate control of end-to-end latency for computing tasks, improves communication efficiency and the processing power of computing nodes, and reduces latency and jitter.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025120939_07052026_PF_FP_ABST
    Figure CN2025120939_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A communication method, a communication apparatus, and a communication system. In the method, an end-to-end delay corresponding to a computing task comprises a transmission delay corresponding to the computing task, and also comprises a computing delay corresponding to the computing task. Thus, on the basis of the end-to-end delay and the computing delay corresponding to the computing task, the transmission delay of the computing task is accurately determined, a corresponding quality of service flow is determined on the basis of the transmission delay of the computing task, and the computing task is transmitted by means of the quality of service flow. The end-to-end delay corresponding to the computing task can be accurately controlled, and the end-to-end delay corresponding to the computing task is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

A communication method, communication device and communication system

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411516024.2, filed on October 28, 2024, entitled "A Communication Method, Communication Device and Communication System", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of communication technology, and in particular to a communication method, communication device and communication system. Background Technology

[0004] With the widespread deployment of networks, mobile internet has achieved tremendous success, gradually becoming one of the main ways for users to access the internet, leading to an explosive growth in mobile traffic. Traditional centralized anchor point deployment methods are increasingly unable to support this rapidly growing mobile traffic model. On the one hand, in networks with centralized anchor point gateway deployment, the increased traffic ultimately concentrates at the gateway and core equipment room, placing increasingly higher demands on backhaul network bandwidth, equipment room throughput, and gateway specifications. On the other hand, the long-distance backhaul network from the access network to the anchor point gateway and the complex transmission environment also result in significant latency and jitter in user packet transmission. The backhaul network mainly includes intermediate links from the core network (or backbone network) to the edge subnet.

[0005] Against this backdrop, the industry has proposed the concept of edge computing. Edge computing enables distributed local processing of business traffic by moving business processing capabilities to the network edge, avoiding excessive traffic concentration and thus significantly reducing the specification requirements for core data centers and centralized gateways. At the same time, edge computing also shortens the distance of the backhaul network, reduces end-to-end latency and jitter of user packets, and makes the deployment of ultra-low latency services possible.

[0006] On the other hand, with the rapid development of artificial intelligence (AI) services, generative AI services are gradually becoming more diverse and evolving into AI agents. Since AI services require significant computing power, power consumption, and storage resources, this places new challenges on terminal devices. Due to resource limitations on the edge side, cloud-edge-device computing collaboration is a new trend that can alleviate the computing power pressure on the edge side. Therefore, in one implementation, computing nodes with AI computing-related business processing capabilities can be deployed within a local data network (DN). In this case, the computing node can be an application server, a functional unit deployed within an application server, or it can be deployed within or near access network equipment.

[0007] In application scenarios where compute nodes are deployed, there is still no corresponding solution for ensuring end-to-end latency between terminal devices and compute nodes. Summary of the Invention

[0008] This application provides a communication method, communication device, and communication system to ensure end-to-end latency between terminal devices and computing nodes.

[0009] In a first aspect, embodiments of this application provide a communication method that can be applied to the network side, such as a network element with computing management functions, a module (e.g., a circuit, chip, or chip system) within the computing management function network element, or a logical node, logical module, or software capable of implementing all or part of the functions of the computing management function network element. The method includes: obtaining a first transmission delay between a terminal device and the computing node based on the end-to-end delay corresponding to the computing task and a first computing delay of the computing task on the computing node; wherein the end-to-end delay corresponding to the computing task refers to the delay between the terminal device sending data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node, and the first computing delay is used to indicate the delay for processing the computing task; and determining quality of service parameters in a quality of service file of a first quality of service (QoS) stream corresponding to the computing task based on the first transmission delay.

[0010] The first computational delay can also be replaced by a requirement related to the first computational delay, which can be used to determine or indicate the first computational delay. For example, the requirement related to the first computational delay may include a service level agreement (SLA) corresponding to the first computational delay.

[0011] The first transmission delay is equal to the sum of the uplink transmission delay of the terminal device sending the computing task to the computing node and the downlink transmission delay of the computing node sending the processing result of the computing task to the terminal device.

[0012] Based on the above scheme, the end-to-end latency corresponding to the computing task includes not only the transmission latency but also the computing latency. Therefore, by accurately determining the transmission latency of the computing task based on its end-to-end latency and computing latency, and by determining the corresponding quality of service flow based on the transmission latency of the computing task and transmitting the computing task through the quality of service flow, accurate control of the end-to-end latency corresponding to the computing task can be achieved, thus ensuring the end-to-end latency of the computing task.

[0013] In one possible implementation, the quality of service (QoS) delay in the QoS file corresponding to the computing task is determined based on the first transmission delay. This QoS delay is one type of QoS parameter. The QoS delay can be a packet delay budget (PDB), which indicates the uplink or downlink delay of data packets transmitted between the terminal device and the user plane network element (or computing node) via the first QoS stream for the computing task.

[0014] In one possible implementation, first information is sent to the terminal device and / or the computing node, the first information being used to map the computing task to the first quality of service stream for transmission.

[0015] Based on the above scheme, the terminal device can transmit the data corresponding to the computing task through the first quality of service stream according to the first information, for example, by adding the identifier of the first quality of service stream to the uplink data packet carrying the data. The computing node can transmit the processing result of the computing task through the first quality of service stream according to the first information, for example, by adding the identifier of the first quality of service stream to the downlink data packet carrying the processing result.

[0016] In one possible implementation, the first computation delay is sent to the access network device, the first computation delay being used by the access network device to schedule data packets for the computation task.

[0017] Based on the above scheme, the access network device can use the first computation delay of the computation task to determine the time when the access network device receives the processing result of the computation task from the computing node, thereby preparing scheduling resources in advance and improving communication efficiency.

[0018] In one possible implementation, the first computation delay is sent to the computing node, the first computation delay being used to indicate that the processing of the computing task is completed within the first computation delay.

[0019] Based on the above scheme, after receiving a computing task, the computing node completes the processing of the computing task and sends out the processing result within the first computing delay, which helps the computing node to complete the processing of the computing task in a timely manner.

[0020] In one possible implementation, a policy charging control rule is received from a policy control network element, the policy charging control rule including the identifier of the computing task and the end-to-end delay corresponding to the computing task.

[0021] Based on the above scheme, it is helpful to accurately obtain the end-to-end latency corresponding to the computing task, and thus help to correctly guarantee the end-to-end latency of the computing task between the terminal device and the computing node.

[0022] In one possible implementation, the end-to-end latency corresponding to the computing task is determined according to a local policy.

[0023] Based on the above scheme, it is helpful to accurately obtain the end-to-end latency corresponding to the computing task, and thus help to correctly guarantee the end-to-end latency of the computing task between the terminal device and the computing node.

[0024] In one possible implementation, the first computing latency corresponding to the computing task is obtained from a local network, network management network, network storage function network element, or unified data warehouse network element.

[0025] Based on the above scheme, it is helpful to accurately obtain the first computing latency corresponding to the computing task, and thus help to correctly guarantee the end-to-end latency of the computing task between the terminal device and the computing node.

[0026] In one possible implementation, a second transmission delay is obtained, which is the monitored or predicted uplink or downlink quality of service delay between the terminal device and the computing node, or the monitored or predicted uplink or downlink quality of service delay between the terminal device and the user plane network element; the first computing delay is determined based on the end-to-end delay corresponding to the computing task and the second transmission delay.

[0027] In one possible implementation, the detected second transmission delay can be obtained through a Quality of Service (QoS) monitoring mechanism. The core network element instructs the user plane network element to perform QoS monitoring for a specific QoS flow and report the monitored delay results (i.e., the second transmission delay).

[0028] In one possible implementation, the predicted second transmission delay is obtained through the QoS sustainability analysis mechanism of the network data analysis function (NWDAF) element. The NWDAF element reports the predicted delay result (i.e., the second transmission delay) to the core network element.

[0029] Based on the above scheme, it is helpful to accurately obtain the first computing latency corresponding to the computing task, and thus help to correctly guarantee the end-to-end latency of the computing task between the terminal device and the computing node.

[0030] In one possible implementation, the first computational delay is used to indicate the delay in processing the computational task under the first task splitting method.

[0031] In one possible implementation, the first task segmentation method is used to instruct the terminal device to convert data packets for the computing task (e.g., inference task) into multiple tokens.

[0032] Based on the above scheme, task splitting is supported, thereby reducing the time for computing nodes to process computing tasks and helping to ensure end-to-end latency of computing tasks between terminal devices and computing nodes.

[0033] In one possible implementation, a second computational delay of the computational task on the computing node is received, the second computational delay being used to indicate the latency predicted or obtained by the computing node for processing the computational task; a third transmission delay of the computational task between the terminal device and the computing node is obtained based on the end-to-end latency corresponding to the computational task and the second computational delay; and the quality of service parameters of the quality of service file of the second quality of service stream corresponding to the computational task are determined or the quality of service parameters of the quality of service file of the first quality of service stream are updated based on the third transmission delay.

[0034] Based on the above scheme, the transmission latency can be updated based on the updated computation latency, which helps to ensure the end-to-end latency of the computing task between the terminal device and the computing node.

[0035] In one possible implementation, capability information from a terminal device is received, indicating whether the terminal device has tokenization capabilities; based on the capability information, the computing node is selected. Tokenization capability refers to the terminal device converting data packets of a computing task (e.g., an inference task) into multiple tokens and sending these tokens to the computing node. If the terminal device does not support tokenization, it sends the raw data of the computing task to the computing node.

[0036] Based on the above scheme, it is helpful to select appropriate computing nodes. For example, when the capability information indicates that the terminal device does not have tokenization capability, a computing node that supports tokenization capability is selected. When the capability information indicates that the terminal device has tokenization capability, either a computing node that supports tokenization capability or a computing node that does not support tokenization capability can be selected.

[0037] In one possible implementation, a computing task refers to a computing model on a computing node that can be distinguished by the type of service it supports, such as supporting AI models, large language models, large model inference, image rendering, image recognition, human-computer interaction, and image-text question answering. It can also be distinguished by different computing services, such as deployment by a carrier, Wenxin Yiyan, Kimi, or other vendors. Furthermore, it can be distinguished by the service parameters of the different computing services it supports. Computing tasks can also be distinguished through one or more of the three methods mentioned above.

[0038] Secondly, embodiments of this application provide a communication method that can be applied to the network side, such as a computing node on the network side, a module (e.g., a circuit, chip, or chip system) in the computing node, or a logical node, logical module, or software capable of implementing all or part of the functions of the computing node. The method includes: receiving an uplink data packet of a computing task transmitted on a first quality of service (QoS) stream, the uplink data packet including an identifier of the first QoS stream corresponding to the computing task; and sending a downlink data packet of the computing task, the downlink data packet including the identifier of the first QoS stream, the downlink data packet being obtained by processing the uplink data packet within a first computing delay corresponding to the computing task; wherein, the QoS parameters in the QoS file corresponding to the identifier of the first QoS stream are determined based on the first transmission delay of the computing task, the first transmission delay being determined based on the end-to-end delay corresponding to the computing task and the first computing delay, the end-to-end delay corresponding to the computing task referring to the delay between the terminal device sending data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node, and the first computing delay indicating the delay for processing the computing task.

[0039] Based on the above scheme, the end-to-end latency corresponding to the computing task includes not only the transmission latency but also the computing latency. Therefore, by accurately determining the transmission latency of the computing task based on its end-to-end latency and computing latency, and by determining the corresponding quality of service flow based on the transmission latency of the computing task and transmitting the computing task through the quality of service flow, accurate control of the end-to-end latency corresponding to the computing task can be achieved, thus ensuring the end-to-end latency of the computing task.

[0040] In one possible implementation, after receiving the uplink data packet, the computing node can perform the following processing: computation, inference, rendering, analysis, or processing through AI business-related functions such as large language models.

[0041] In one possible implementation, first information is received, which is used to map the computing task to the first quality of service stream for transmission.

[0042] Based on the above scheme, the computing node can transmit the processing result of the computing task through the first quality of service stream according to the first information, for example, by adding the identifier of the first quality of service stream to the downlink data packet carrying the processing result.

[0043] In one possible implementation, a second computational delay of the computational task on the computing node is sent, the second computational delay being used to indicate the latency predicted or obtained by the computing node for processing the computational task.

[0044] In one possible implementation, an indication message is received, which instructs the computing node to send the latency predicted or acquired by the computing node for processing the computing task.

[0045] In one possible implementation, the second computation delay is determined based on a first parameter, a second parameter, and a third parameter corresponding to the computation task; wherein the first parameter indicates the time required to generate the first token of the computation task, the second parameter indicates the time interval between generating two adjacent tokens of the computation task, and the third parameter indicates the number of tokens generated for the computation task.

[0046] Based on the above scheme, the second calculation delay can be accurately determined.

[0047] In one possible implementation, the uplink data packet also includes an identifier of the computing task.

[0048] In one possible implementation, the downlink data packet also includes an identifier of the computing task.

[0049] Based on the above scheme, carrying the identifier of the computing task in the uplink data packet and / or downlink data packet helps to accurately indicate the computing task corresponding to the content carried in the data packet.

[0050] Thirdly, this application provides a communication method that can be applied to the terminal device side, such as the terminal device or the communication module in the terminal device, or the circuit or chip in the terminal device that is responsible for the communication function (such as a modem chip, also known as a baseband chip, or a system-on-chip (SoC) chip or system-in-package (SIP) chip that includes a modem core). The method includes: sending an uplink data packet of a computing task on a first quality of service (QoS) stream, the uplink data packet including an identifier of the first QoS stream corresponding to the computing task; receiving a downlink data packet, the downlink data packet including the identifier of the first QoS stream, the downlink data packet being obtained by a computing node processing the uplink data packet within a first computing delay corresponding to the computing task; wherein, the QoS parameters in the QoS file corresponding to the identifier of the first QoS stream are determined based on a first transmission delay of the computing task, the first transmission delay being determined based on the end-to-end delay corresponding to the computing task and the first computing delay, the end-to-end delay corresponding to the computing task referring to the delay between the terminal device sending data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node, the first computing delay being used to indicate the delay of processing the computing task.

[0051] Based on the above scheme, the end-to-end latency corresponding to the computing task includes not only the transmission latency but also the computing latency. Therefore, by accurately determining the transmission latency of the computing task based on its end-to-end latency and computing latency, and by determining the corresponding quality of service flow based on the transmission latency of the computing task and transmitting the computing task through the quality of service flow, accurate control of the end-to-end latency corresponding to the computing task can be achieved, thus ensuring the end-to-end latency of the computing task.

[0052] In one possible implementation, after receiving the uplink data packet, the computing node can perform the following processing: computation, inference, rendering, analysis, or processing through AI business-related functions such as large language models.

[0053] In one possible implementation, first information is received, which is used to map the computing task to the first quality of service stream for transmission.

[0054] Based on the above scheme, the terminal device can transmit the data corresponding to the computing task through the first quality of service stream according to the first information, for example, by adding the identifier of the first quality of service stream to the uplink data packet carrying the data.

[0055] In one possible implementation, capability information is sent, which indicates whether the terminal device has tokenization capability.

[0056] Based on the above scheme, the terminal device sends capability information, which can be used to assist in selecting computing nodes, thereby helping to choose a suitable computing node. For example, if the capability information indicates that the terminal device does not have tokenization capability, then a computing node that supports tokenization capability is selected; if the capability information indicates that the terminal device has tokenization capability, then either a computing node that supports tokenization capability or a computing node that does not support tokenization capability can be selected.

[0057] In one possible implementation, the uplink data packet also includes an identifier of the computing task.

[0058] In one possible implementation, the downlink data packet also includes an identifier of the computing task.

[0059] Based on the above scheme, carrying the identifier of the computing task in the uplink data packet and / or downlink data packet helps to accurately indicate the computing task corresponding to the content carried in the data packet.

[0060] Fourthly, this application provides a communication device that has the functions of the first aspect described above. For example, the communication device includes modules, units, or means corresponding to the operations involved in the first aspect. These modules, units, or means can be implemented by software, hardware, or a combination of software and hardware.

[0061] Fifthly, this application provides a communication device that has the functions of the second aspect above. For example, the communication device includes modules, units or means corresponding to the operations involved in the second aspect above. The modules, units or means can be implemented by software, or by hardware, or by a combination of software and hardware.

[0062] Sixthly, this application provides a communication device that has the functions of the third aspect above. For example, the communication device includes modules, units or means corresponding to the operations involved in the third aspect above. The modules, units or means can be implemented by software, or by hardware, or by a combination of software and hardware.

[0063] In a seventh aspect, this application provides a communication device including an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory stores part or all of the necessary computer program or instructions for implementing the functions described in the first aspect. The one or more processors are executable to carry out the computer program or instructions, causing the communication device to implement the methods in any possible design or implementation of the first aspect. The interface circuit is used to implement the communication functions within the communication device and / or the communication functions between the communication device and other devices or components.

[0064] The aforementioned communication device may be a computing management function network element, a module (e.g., a circuit, chip, or chip system) within a computing management function network element, or a logic node, logic module, or software capable of implementing all or part of the functions of a computing management function network element.

[0065] Eighthly, this application provides a communication device including an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory stores part or all of the necessary computer program or instructions for implementing the functions described in the second aspect above. The one or more processors are executable to carry out the computer program or instructions, causing the communication device to implement the methods in any possible design or implementation of the second aspect above. The interface circuit is used to implement the communication functions within the communication device and / or the communication functions between the communication device and other devices or components.

[0066] The aforementioned communication device may be a computing node, a module (e.g., a circuit, chip, or chip system) within a computing node, or a logic node, logic module, or software capable of implementing all or part of the functions of a computing node.

[0067] Ninthly, this application provides a communication device including an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory stores part or all of the computer program or instructions necessary to implement the functions described in the third aspect above. The one or more processors are executable to carry out the computer program or instructions, causing the communication device to implement the methods in any possible design or implementation of the third aspect above. The interface circuit is used to implement the communication functions within the communication device and / or the communication functions between the communication device and other devices or components.

[0068] In one possible design, the processor is used to communicate with other devices or components through the interface circuit.

[0069] In one possible design, the communication device may also include the memory.

[0070] The aforementioned communication device may be a terminal device, or a communication module in a terminal device, or a chip in a terminal device that is responsible for communication functions, such as a modem chip (also known as a baseband chip), or a SoC or SIP chip that includes a modem module.

[0071] In a tenth aspect, this application provides a computer-readable storage medium storing a computer program or instructions that, when executed, implement the method in any of the possible designs of the first to third aspects described above.

[0072] In one aspect, this application provides a computer program product comprising a computer program or instructions that, when executed, implement the method in any of the possible designs of the first to third aspects described above.

[0073] In a twelfth aspect, this application provides a communication system comprising at least two devices: a computing management function network element, a computing node, and a terminal device. The computing management function network element is used to execute the first aspect and any possible implementation method of the first aspect. The computing node is used to execute the second aspect and any possible implementation method of the second aspect. The terminal device is used to execute the third aspect and any possible implementation method of the third aspect. Attached Figure Description

[0074] Figure 1 is a schematic diagram of a 5G network architecture based on a service-oriented architecture;

[0075] Figure 2(a) is a flowchart illustrating a communication method provided in an embodiment of this application;

[0076] Figure 2(b) is a flowchart illustrating a communication method provided in an embodiment of this application;

[0077] Figure 3 is a flowchart illustrating a communication method provided in an embodiment of this application;

[0078] Figure 4 is a flowchart illustrating a communication method provided in an embodiment of this application;

[0079] Figure 5 is a possible exemplary block diagram of the communication device involved in the embodiments of this application;

[0080] Figure 6 is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0081] To address the challenges of wireless broadband technology and maintain the leading edge of the 3rd Generation Partnership Project (3GPP) network, the 3GPP standards group has developed a 5G network architecture. This architecture not only supports access to the 5G core network (CN) using radio access technologies defined by the 3GPP standards group (such as Long Term Evolution (LTE) and 5G Radio Access Network (RAN) technologies), but also supports access to the core network using non-3GPP access technologies through non-3GPP interworking functions (N3IWF) or next-generation packet data gateways (ngPDG).

[0082] Figure 1 is a schematic diagram of a service-oriented 5G network architecture. The 5G network architecture shown in Figure 1 may include access network equipment and core network equipment. Terminal devices access the DN through access network equipment and core network equipment. The core network equipment includes, but is not limited to, some or all of the following network elements: authentication server function (AUSF) network element, unified data management (UDM) network element, unified data repository (UDR) network element, network repository function (NRF) network element, network exposure function (NEF) network element, application function (AF) network element, policy control function (PCF) network element, access and mobility management function (AMF) network element, session management function (SMF) network element, user plane function (UPF) network element, and compute management function (CMF) network element.

[0083] Access network equipment, sometimes also called RAN nodes, RAN entities, or access nodes, is used to help terminal equipment achieve wireless access.

[0084] In one possible scenario, the access network device can be a base station, an evolved NodeB (eNodeB), an access point (AP), a transmission reception point (TRP), a next-generation NodeB (gNB), a base station in a future mobile communication system, or an access node in a wireless fidelity (WiFi) system. The access network device can be a macro base station, a micro base station, an indoor station, a relay node, or a donor node. Optionally, the access network device can also be a server, a wearable device, a vehicle, or an in-vehicle device. For example, the access network device in vehicle-to-everything (V2X) technology can be a roadside unit (RSU). All or part of the functions of the access network device in this application can also be implemented through software functions running on hardware, or through virtualization functions instantiated on a platform (e.g., a cloud platform). The access network device can also be equipped with communication modules, circuits, or chips that perform corresponding communication functions. The access network device can also be configured with program instructions for performing corresponding communication functions and corresponding program instructions. The access network device in this application may also be a logical node, logical module, or software that can implement all or part of the functions of the access network device.

[0085] In another possible scenario, multiple access network devices collaborate to assist terminal devices in achieving wireless access, with each access network device performing a portion of the base station's functions. For example, the access network devices can be a central unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU), etc. The CU and DU can be configured separately or included in the same network element, such as a baseband unit (BBU). The RU can be included in radio frequency equipment or radio frequency units, such as a remote radio unit (RRU), an active antenna unit (AAU), or a remote radio head (RRH).

[0086] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an open radio access network (ORAN) system, CU can also be called O-CU (open CU), DU can also be called O-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CU-UP, and RU can also be called O-RU. For ease of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples. Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through a software module, a hardware module, or a combination of software and hardware modules.

[0087] Terminal devices can also be called terminals, user equipment (UE), mobile stations, mobile terminals, etc. Terminal devices can be widely used in various scenarios, such as device-to-device (D2D), vehicle-to-everything (V2X) communication, machine-type communication (MTC), the Internet of Things (IoT), virtual reality, augmented reality, industrial control, autonomous driving, telemedicine, smart grids, smart furniture, smart offices, smart wearables, smart transportation, smart cities, etc. Terminal devices can be mobile phones, tablets, computers with wireless transceiver capabilities, wearable devices, vehicles, drones, helicopters, airplanes, ships, robots, robotic arms, smart home devices, transportation vehicles with wireless communication capabilities, communication modules, etc. The embodiments of this application do not limit the device form of the terminal device. Terminal devices typically contain communication modules, circuits, or chips that perform corresponding communication functions. Terminal devices can also be configured with program instructions for performing corresponding communication functions.

[0088] Access network equipment and terminal equipment can be fixed or mobile. They can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; they can also be deployed on water; and they can be deployed in the air on aircraft, balloons, and satellites. The embodiments of this application do not limit the application scenarios of the access network equipment and terminal equipment.

[0089] AMF network elements include functions such as performing mobility management or access authentication / authorization. They are also responsible for transmitting user policies between terminal devices and PCF network elements.

[0090] SMF network elements include functions such as performing session management, executing control policies issued by PCF network elements, selecting UPF network elements, or allocating Internet Protocol (IP) addresses to terminal devices.

[0091] UPF network elements include functions such as user plane data forwarding, session / flow-based billing statistics, and bandwidth limiting.

[0092] UDM network elements include functions such as managing contracted data or authorizing user access.

[0093] UDR includes functions for accessing data of various types, such as contract data, policy data, or application data.

[0094] NEF network elements are used to support the opening of capabilities and events, enabling third parties to indirectly interact with certain network elements within the 3GPP network.

[0095] AF (Application Provider) network elements convey application-side requests to the network side, such as QoS requirements or user state event subscriptions. AFs can be third-party functional entities or application services deployed by operators, such as IP Multimedia Subsystem (IMS) voice call services. AF network elements include AFs within the core network (i.e., operator-owned AFs) and third-party AFs (such as an enterprise's application server).

[0096] PCF network elements include policy control functions such as billing at the session and service flow levels, QoS bandwidth guarantee and mobility management, or terminal device policy decisions.

[0097] NRF network elements can be used to provide network element discovery functionality, providing network element information corresponding to the network element type based on requests from other network elements. NRF network elements also provide network element management services, such as network element registration, updates, deregistration, or network element status subscription and push.

[0098] The AUSF network element is responsible for authenticating users to determine whether to allow users or devices to access the network.

[0099] CMF network elements are responsible for establishing computing connections, managing computing plane sessions, and selecting computing nodes with service processing capabilities.

[0100] A Domain Provider (DN) is a network located outside of the carrier's network. A carrier's network can connect to multiple DNs, and various services can be deployed on a DN, providing data and / or voice services to terminal devices. For example, a DN might be the private network of a smart factory. Sensors installed in the workshop can act as terminal devices, and a control server for these sensors is deployed within the DN. The control server provides services to the sensors. Sensors can communicate with the control server, receive instructions from it, and transmit the collected sensor data back to the control server accordingly. Another example is a DN serving as an internal office network for a company. Employees' mobile phones or computers can act as terminal devices, accessing information and data resources on the company's internal office network.

[0101] In Figure 1, Nausf, Npcf, Nudr, Nudm, Naf, Namf, Nsmf, Nnef, and Nnrf are the service-based interfaces (SBIs) provided by AUSF, PCF, UDR, UDM, AF, AMF, SMF, NEF, and NRF, respectively, used to invoke the corresponding service-based operations. N1, N2, N3, N4, and N6 are interface sequence numbers, and their meanings are as follows:

[0102] 1) N1: The interface between the AMF network element and the terminal device, which can be used to transmit non-access stratum (NAS) signaling (such as QoS rules from the AMF network element) to the terminal device.

[0103] 2) N2: The interface between the AMF network element and the access network equipment, which can be used to transmit radio bearer control information from the core network side to the access network equipment.

[0104] 3) N3: The interface between the access network equipment and the UPF network element, mainly used to transmit uplink and downlink user plane data between the access network equipment and the UPF network element.

[0105] 4) N4: The interface between SMF network elements and UPF network elements. It can be used to transmit information between the control plane and the user plane, including the distribution of forwarding rules, QoS rules, traffic statistics rules, etc. from the control plane to the user plane, as well as the reporting of information from the user plane.

[0106] 5) N6: The interface between the UPF network element and the DN, used to transmit uplink and downlink user data streams between the UPF network element and the DN.

[0107] In the architecture shown in Figure 1, the various network functional elements are connected via a service-oriented bus and interact through service-oriented interfaces. The advantages of a service-oriented bus include improved network flexibility, openness, scalability, and intelligence, enabling support for diverse service scenarios and requirements. The service-oriented bus can be used to transmit various types of data and signaling. For example, it can be used to transmit latency-sensitive real-time signaling (e.g., service-oriented interface call signaling between network functional elements), latency-sensitive real-time data (e.g., real-time AI inference data), and non-real-time data (e.g., offline AI training data). Furthermore, when transmitting this data or signaling, the service-oriented bus couples the data or signaling together; that is, the service-oriented bus can simultaneously transmit real-time signaling, real-time data, and non-real-time data.

[0108] It should be noted that the term "network element" can be omitted when describing the above network elements (such as SMF network elements, UPF network elements, etc.). For example, an SMF network element can be abbreviated as SMF, a UPF network element as UPF, and so on. This abbreviated description is also used in Figure 1.

[0109] It is understood that the aforementioned network element or function can be a network component in a hardware device, a software function running on dedicated hardware, or a virtualized function instantiated on a platform (e.g., a cloud platform). Optionally, the aforementioned network element or function can be implemented by one device, multiple devices working together, or a functional module within a single device; this application embodiment does not specifically limit this.

[0110] The computing management function network element and user plane network element in this application can be the CMF network element and UPF network element in Figure 1, respectively, or they can be network elements in future communication networks that have the functions of the aforementioned CMF network element and UPF network element. This application does not limit them in this regard.

[0111] With the widespread deployment of networks, mobile internet has achieved tremendous success, gradually becoming one of the main ways for users to access the internet, leading to an explosive growth in mobile traffic. Traditional centralized anchor point deployment methods are increasingly unable to support this rapidly growing mobile traffic model. On the one hand, in networks with centralized anchor point gateway deployment, the increased traffic ultimately concentrates at the gateway and core equipment room, placing increasingly higher demands on backhaul network bandwidth, equipment room throughput, and gateway specifications. On the other hand, the long-distance backhaul network from the access network to the anchor point gateway and the complex transmission environment also result in significant latency and jitter in user packet transmission.

[0112] Against this backdrop, the industry has proposed the concept of edge computing. Edge computing enables distributed local processing of business traffic by moving business processing capabilities to the network edge, avoiding excessive traffic concentration and thus significantly reducing the specification requirements for core data centers and centralized gateways. At the same time, edge computing also shortens the distance of the backhaul network, reduces end-to-end latency and jitter of user packets, and makes the deployment of ultra-low latency services possible.

[0113] On the other hand, with the rapid development of AI services, generative AI services are gradually becoming more diverse and evolving into AI agents. Since AI services require significant computing power, power consumption, and storage resources, this presents new challenges for terminal devices. Due to resource limitations on the edge side, cloud-edge-device computing collaboration is a new trend that can alleviate the computing pressure on the edge side. Therefore, in one implementation, computing nodes with AI computing-related business processing capabilities can be deployed within the local data center (DN). In this case, the computing node can be an application server, a functional unit deployed within an application server, or it can be deployed within or near the access network equipment.

[0114] In application scenarios where compute nodes are deployed, there is still no corresponding solution for ensuring end-to-end latency between terminal devices and compute nodes.

[0115] To address the aforementioned issues, this application provides corresponding solutions.

[0116] The communication method and communication device will be further described below with reference to the accompanying drawings. It is understood that this application uses computing management function network elements, computing nodes, and terminal devices as examples of the execution entities in this interactive illustration, but this application does not limit the execution entities in the interactive illustration. For example, the method executed by the computing management function network element in this application can also be implemented by modules (e.g., circuits, chips, or chip systems) within the computing management function network element, or by logical nodes, logical modules, or software capable of implementing all or part of the computing management function network element's functions. Similarly, the method executed by the computing node in this application can also be implemented by modules (e.g., circuits, chips, or chip systems) within the computing node, or by logical nodes, logical modules, or software capable of implementing all or part of the computing node's functions. Likewise, the method executed by the terminal device in this application can also be implemented by a communication module within the terminal device, or by circuits or chips (such as modem chips (also known as baseband chips), or SoC chips including modem cores, or SIP chips) within the terminal device responsible for communication functions.

[0117] In this application, a computing node can be configured with one or more computing power applications and corresponding computing power resources for each application. Each computing power application supports processing corresponding computing tasks or large model inference requests. A computing node is also referred to as a computing unit, a far edge intelligent node (FEIN), or a computation execution function (CEF) network element. Optionally, a computing task can be uniquely identified by any one or more of the following: task identifier, address, or port number. Optionally, each computing task corresponds to one or more computing models. In one possible implementation, a computing task refers to a computing model on a computing node that can be distinguished based on the type of service it supports, such as supporting artificial intelligence models, large language models, large model inference, image rendering services, image recognition, human-computer interaction, or image-text question answering models; it can also be distinguished based on different computing services, such as deployment by operators, Wenxin Yiyan, Kimi, or other manufacturers; or it can be distinguished by the service parameters of different supported computing services. Computing tasks can also be distinguished by one or more of the three distinction methods mentioned above.

[0118] Figure 2(a) is a flowchart illustrating a communication method provided in an embodiment of this application. The method includes the following steps:

[0119] Step 201a: The computing management function network element obtains the first transmission delay of the computing task between the terminal device and the computing node based on the end-to-end delay corresponding to the computing task and the first computing delay of the computing task on the computing node.

[0120] In one possible implementation, the end-to-end latency for a computation task refers to the time between the terminal device sending the data corresponding to the computation task to the terminal device receiving the processing result of the computation task from the computing node. For example, an end-to-end latency of 500 milliseconds (ms) means that the time between the terminal device sending the data corresponding to the computation task to the computing node and the terminal device receiving the processing result of the computation task is 500 ms.

[0121] The computational task here refers to a task that requires related reasoning or computation. Examples include rendering an image, rendering a video, or a question-and-answer format. This application does not limit the specific type of computational task.

[0122] The end-to-end latency of a computing task comprises two parts: the transmission latency of the computing task and the computation latency of the computing task. This application can obtain the first transmission latency of the computing task between the terminal device and the computing node based on the end-to-end latency of the computing task and the first computation latency of the computing task on the computing node. Wherein, the first transmission latency + the first computation latency ≤ the end-to-end latency of the computing task.

[0123] The first computation delay is used to indicate the delay in processing the computation task. That is, after receiving the computation task, the computing node needs to process the computation task, obtain the processing result, and send out the processing result within the first computation delay.

[0124] The first transmission delay refers to the transmission delay of the computing task between the terminal device and the computing node. The first transmission delay consists of two parts: the uplink transmission delay of the terminal device sending the data corresponding to the computing task to the computing node, and the downlink transmission delay of the computing node sending the processing result of the computing task to the terminal device. In one application scenario, assuming the computing node is deployed near the access network device, the first transmission delay consists of the following delays: uplink transmission delay 1 for the computing task transmission between the terminal device and the access network device, uplink transmission delay 2 for the computing task transmission between the access network device and the computing node, downlink transmission delay 1 for the computing task transmission between the computing node and the access network device, and downlink transmission delay 2 for the computing task transmission between the access network device and the terminal device. In another application scenario, assuming the computing node is deployed within the DN, the first transmission delay consists of the following delays: uplink transmission delay 1 for the computing task transmitted between the terminal device and the access network device, uplink transmission delay 2 for the computing task transmitted between the access network device and the UPF network element, uplink transmission delay 3 for the computing task transmitted between the UPF network element and the computing node, downlink transmission delay 1 for the computing task transmitted between the computing node and the UPF network element, downlink transmission delay 2 for the computing task transmitted between the UPF network element and the access network device, and downlink transmission delay 3 for the computing task transmitted between the access network device and the terminal device.

[0125] The following example illustrates this. Assume the end-to-end latency for a computation task is 500ms, and the first computation latency on the computation node is 200ms. Then the first transmission latency is less than or equal to 300ms. If the uplink and downlink transmission latency of the computation task are the same, then both the uplink and downlink transmission latency for this computation task are 150ms.

[0126] Step 202a: The computing management function network element determines the quality parameters in the QoS profile of the first QoS flow corresponding to the computing task based on the first transmission delay.

[0127] The quality parameter includes at least the QoS latency of the first quality of service stream, which is used to indicate the uplink latency and / or downlink latency of computing tasks transmitted between wireless networks.

[0128] One possible interpretation is that the QoS delay is a portion of the aforementioned first transmission delay, meaning that the first transmission delay includes the QoS delay of the first quality of service stream.

[0129] In addition to quality parameters, the QoS file of the first QoS flow may also contain other information, such as the identifier of the first QoS flow, namely the QoS flow identity (QFI).

[0130] Based on the above scheme, the end-to-end latency corresponding to the computing task includes not only the transmission latency but also the computing latency. Therefore, by accurately determining the transmission latency of the computing task based on its end-to-end latency and computing latency, and by determining the corresponding quality of service flow based on the transmission latency of the computing task and transmitting the computing task through the quality of service flow, the end-to-end latency corresponding to the computing task can be controlled, thus ensuring the end-to-end latency of the computing task.

[0131] As one implementation method, step 201a above can be performed by other network elements (e.g., PCF network elements). If performed by other network elements, after determining the first transmission delay, the other network elements send the first transmission delay to the computing management function network element. As another implementation method, both steps 201a and 202a above can be performed by other network elements (e.g., PCF network elements). For ease of explanation, the following description uses the computing management function network element performing steps 201a and 202a above.

[0132] This application does not limit the implementation method of the computing management function network element to obtain the end-to-end latency corresponding to the computing task. For example, two different methods are given below.

[0133] Implementation Method 1: The computing management function network element receives policy and charging control (PCC) rules from the policy control network element (e.g., PCF network element). The PCC rules include the identifier of the computing task and the end-to-end delay corresponding to the computing task.

[0134] In method 2, the computing management function network element determines the end-to-end latency corresponding to the computing task based on the local policy.

[0135] The local policy can be pre-configured on the computing management function network element, or it can be received by the computing management network element from other network elements, such as being reported in advance by the terminal device to the computing management function network element.

[0136] This application does not limit the implementation method of the computing management function network element obtaining the first computing delay corresponding to the computing task. For example, two different methods are given below.

[0137] In method A, the computing management function network element obtains the first computing latency corresponding to the computing task from the local network, network management network, NRF network element, or UDR network element.

[0138] In method B, the computing management function network element obtains the second transmission delay and determines the task splitting method and the first computing delay based on the end-to-end delay corresponding to the computing task and the second transmission delay. The second transmission delay is the monitored or predicted uplink or downlink QoS delay between the terminal device and the computing node, or the monitored or predicted uplink or downlink QoS delay between the terminal device and the user plane network element. The first computing delay is used to indicate the delay for processing the computing task under this task splitting method.

[0139] When a computing task requires a task partitioning method, it refers to the need for terminal devices and computing nodes to jointly execute a computing task (such as an inference task). For example, if computing task #1 is not partitioned and is executed entirely by the computing node, the computation latency of computing task #1 on the computing node is 400ms. Under task partitioning method #1, the computation latency of computing task #1 on the computing node is 100ms. Under task partitioning method #2, the computation latency of computing task #1 on the computing node is 200ms. Under task partitioning method #3, the computation latency of computing task #1 on the computing node is 300ms.

[0140] In addition, there is another task partitioning method where the terminal device tokenizes the computing task. That is, the terminal device converts the data packet of the computing task (e.g., inference task) into multiple tokens, and then sends these tokens to the computing node. The computing node then processes these tokens to obtain the result. Based on this method, since the computing node does not need to perform the tokenization operation, the computation latency of the computing node executing the computing task can be reduced.

[0141] The second transmission delay can be monitored or predicted by one or more network elements such as access network equipment, UPF network element, or NWDAF network element.

[0142] Regarding implementation method B, an example is provided below. For instance, the end-to-end latency for computation task #1 is 500ms, and the second transmission latency is 100ms. This second transmission latency represents the monitored or predicted uplink QoS latency between the terminal device and the computing node. Assuming the uplink QoS latency between the terminal device and the computing node is the same as the downlink QoS latency, the round-trip latency between the terminal device and the computing node is calculated to be 200ms based on this second transmission latency (i.e., the sum of the uplink QoS latency and the downlink QoS latency between the terminal device and the computing node is 200ms). Therefore, the remaining latency available for processing the computation task is 300ms (i.e., 500ms minus 200ms). Assume the available task partitioning methods include task partitioning method #1, task partitioning method #2, and task partitioning method #3. The computation latency of task #1 on the computing node is 100ms under task partitioning method #1, 200ms under task partitioning method #2, and 400ms under task partitioning method #3. Since 100ms and 200ms are both less than 300ms, and 400ms is greater than 300ms, the computing management function network element can choose either task partitioning method #1 or task partitioning method #2. If task partitioning method #1 is selected, the computation latency of task #1 on the computing node under this selected method is determined to be 100ms, meaning the first computation latency of the task is equal to 100ms. Therefore, the first transmission latency between the terminal device and the computing node is determined to be 400ms (i.e., 500ms minus 100ms). If task splitting method #2 is selected, the computation latency of computation task #1 on the computing node under the selected task splitting method #2 is determined to be 200ms. That is, the first computation latency of the computation task is equal to 200ms. Therefore, the first transmission latency of the computation task between the terminal device and the computing node is determined to be 300ms (i.e., 500ms minus 200ms).

[0143] Regarding implementation method B, an example is provided below. For instance, the end-to-end latency for computation task #1 is 500ms, and the second transmission latency is 80ms. This second transmission latency represents the monitored or predicted uplink QoS latency between the terminal device and the user plane network element. Assuming the computing node is deployed within the local DN and the uplink transmission latency between the user plane network element and the computing node is 20ms, the uplink transmission latency between the terminal device and the computing node can be calculated as 100ms (i.e., 80ms plus 20ms). Assuming the uplink transmission latency between the terminal device and the computing node is the same as the downlink transmission latency, the round-trip latency between the terminal device and the computing node can be calculated as 200ms (i.e., the sum of the uplink and downlink transmission latency between the terminal device and the computing node is 200ms). Therefore, the remaining latency available for processing the computation task is 300ms (i.e., 500ms minus 200ms). Assume the available task partitioning methods include task partitioning method #1, task partitioning method #2, and task partitioning method #3. The computation latency of task #1 on the computing node is 100ms under task partitioning method #1, 200ms under task partitioning method #2, and 400ms under task partitioning method #3. Since 100ms and 200ms are both less than 300ms, and 400ms is greater than 300ms, the computing management function network element can choose either task partitioning method #1 or task partitioning method #2. If task partitioning method #1 is selected, the computation latency of task #1 on the computing node under this selected method is determined to be 100ms, meaning the first computation latency of the task is equal to 100ms. Therefore, the first transmission latency between the terminal device and the computing node is determined to be 400ms (i.e., 500ms minus 100ms). If task splitting method #2 is selected, the computation latency of computation task #1 on the computing node under the selected task splitting method #2 is determined to be 200ms. That is, the first computation latency of the computation task is equal to 200ms. Therefore, the first transmission latency of the computation task between the terminal device and the computing node is determined to be 300ms (i.e., 500ms minus 200ms).

[0144] As one implementation method, the computing management function network element can also send first information to the terminal device and / or computing node. This first information maps the data corresponding to the computing task to a first Quality of Service (QoS) stream for transmission. The terminal device can transmit the data corresponding to the computing task through the first QoS stream based on the first information, for example, by adding an identifier of the first QoS stream to the uplink data packet carrying the data, to ensure the QoS latency of the computing task, thereby ensuring the end-to-end latency of the computing task between the terminal device and the computing node. The computing node can also transmit the processing result of the computing task through the first QoS stream based on the first information, for example, by adding an identifier of the first QoS stream to the downlink data packet carrying the processing result, to ensure the QoS latency of the computing task, thereby ensuring the end-to-end latency of the computing task between the terminal device and the computing node.

[0145] As one implementation method, the computing management function network element can also send a first computing delay to the access network device. This first computing delay is used by the access network device to schedule the data corresponding to the computing task. That is, based on the first computing delay of the computing task, the access network device determines the time or time range at which it will receive the processing result of the computing task from the computing node, thus allowing it to prepare scheduling resources in advance and improve communication efficiency. For example, if the access network device knows that the first computing delay of computing task #1 is 100ms, after sending the uplink data packet of computing task #1 (containing the data corresponding to the computing task) to the computing node, it predicts that it will receive the downlink data packet of computing task #1 (containing the processing result of the computing task) from the computing node after (100+x)ms, thus allowing it to schedule the transmission resources for sending the downlink data packet in advance. Here, x represents the round-trip time for data transmission between the access network device and the computing node.

[0146] As one implementation method, the computing management function network element can also send a first computing delay to the computing node. The first computing delay is used to instruct the computing node to complete the processing of the computing task within the first computing delay. That is, after receiving the computing task, the computing node completes the processing of the computing task and sends the processing result within the first computing delay.

[0147] As one implementation method, the computing management function network element can also receive a second computing delay of the computing task on the computing node from the computing node. This second computing delay is used to indicate the delay (e.g., maximum delay or delay interval) predicted by the computing node for processing the computing task. Then, based on the end-to-end delay corresponding to the computing task and the second computing delay, the computing management function network element obtains a third transmission delay of the computing task between the terminal device and the computing node, and determines the quality parameters of the service quality file of the second service quality flow corresponding to the computing task or updates the quality parameters of the service quality file of the first service quality flow based on the third transmission delay. Based on this method, the computing management function network element receives a second computing delay and replaces the first computing delay with the second computing delay as the delay for the computing node to process the computing task. In other words, the second computing delay is determined as the delay for processing the computing task. Then, based on the received second computing delay, the computing management function network element re-acquires the transmission delay (i.e., the third transmission delay) between the terminal device and the computing node for the computing task. Based on the third transmission delay, it determines a new Quality of Service (QoS) flow (i.e., the second QoS flow) and the corresponding QoS file's quality parameters, or it retains the original QoS flow (i.e., the first QoS flow) and updates the quality parameters of the original QoS flow's quality file. Specifically, the computing management function network element replaces the first transmission delay with the third transmission delay as the transmission delay between the terminal device and the computing node. This third transmission delay consists of two parts: one part is the uplink transmission delay for the terminal device to send the data corresponding to the computing task to the computing node, and the other part is the downlink transmission delay for the computing node to send the processing result of the computing task to the terminal device.

[0148] For example, a computing node can determine the second computation delay of a computing task based on a first parameter, a second parameter, and / or a third parameter corresponding to the computing task. The first parameter indicates the time required to generate the first token of the computing task, the second parameter indicates the time interval between two adjacent tokens generated, and the third parameter indicates the number of tokens generated. For instance, if the first parameter indicates the time required to generate the first token of the computing task is T1, the second parameter indicates the time interval between two adjacent tokens generated is T2, and the third parameter indicates the number of tokens generated is N, then the second computation delay = T1 + (N-1)*T2.

[0149] As one implementation method, the computing node can receive indication information from the computing management function network element. This indication information instructs the computing node to send the predicted processing delay for the computing task (i.e., the second computing delay). Based on this indication information, the computing node obtains the second computing delay and sends it to the computing management function network element.

[0150] As one implementation method, the computing management function network element can also receive capability information from the terminal device. This capability information indicates whether the terminal device has tokenization capability. Tokenization refers to converting data packets of computing tasks (such as inference tasks) into multiple tokens. After tokenizing the computing task, the terminal device can send multiple tokens to the computing node. The computing node processes these tokens to obtain the corresponding processing results. The computing management function network element can select computing nodes based on the terminal device's capability information. For example, if the terminal device has tokenization capability, it can tokenize computing tasks. In this case, it is not necessary to restrict the computing node to have tokenization capability; therefore, it can select computing nodes with or without tokenization capability. Conversely, if the terminal device does not have tokenization capability, it cannot tokenize computing tasks. In this case, it is necessary to restrict the computing node to have tokenization capability, thus requiring the selection of computing nodes with tokenization capability. This method helps in selecting suitable computing nodes.

[0151] Figure 2(b) is a flowchart illustrating a communication method provided in an embodiment of this application. The method includes the following steps:

[0152] In step 201b, the terminal device sends uplink data packets of the computing task on the first quality of service stream. Correspondingly, the computing node receives the uplink data packets of the computing task transmitted on the first quality of service stream.

[0153] The uplink data packet includes data corresponding to the computation task. The uplink data packet also includes an identifier of the first Quality of Service (QoS) flow corresponding to the computation task. Optionally, the uplink data packet also includes an identifier of the computation task.

[0154] As one implementation method, before step 201b, the terminal device receives first information, which is used to map the computing task to a first quality of service stream for transmission. Based on the first information, the terminal device adds the identifier of the first quality of service stream corresponding to the computing task to the uplink data packet of the computing task; optionally, it also adds the identifier of the computing task itself.

[0155] If the compute node is deployed near the access network equipment, the uplink data packets for the compute task are sent by the terminal device and forwarded to the compute node via the access network equipment. If the compute node is deployed within the DN, the uplink data packets reach the compute node via the access network equipment and the UPF.

[0156] In step 202b, the compute node sends downlink data packets. Correspondingly, the terminal device receives the downlink data packets.

[0157] The downlink data packets for this computation task are obtained by the computing node processing the uplink data packets within the first computation delay corresponding to the computation task. In other words, the computing node processes the uplink data packets of the computation task within the first computation delay corresponding to the computation task to obtain the downlink data packets for the computation task.

[0158] The downlink data packet includes the processing result of the computation task. The downlink data packet also includes an identifier of the first quality of service flow. Optionally, the downlink data packet also includes an identifier of the computation task.

[0159] As one implementation method, before step 202b, the computing node receives first information, which is used to map the computing task to a first quality of service (QoS) stream for transmission. Based on the first information, the computing node adds the identifier of the first QoS stream corresponding to the computing task to the downlink data packet of the computing task; optionally, it also adds the identifier of the computing task. Alternatively, if the computing node does not receive the first information, it can also carry the identifier of the first QoS stream carried in the uplink data packet of the computing task in the downlink data packet of the computing task; optionally, it also adds the identifier of the computing task.

[0160] If the compute node is deployed near the access network equipment, the downlink data packets for the computing task are sent by the compute node and forwarded to the terminal device via the access network equipment. If the compute node is deployed within the DN, the downlink data packets reach the terminal device via the UPF network element and the forwarding of the access network equipment.

[0161] The quality parameters in the service quality file corresponding to the identifier of the first service quality flow are determined based on the first transmission delay of the computing task. The first transmission delay is determined based on the end-to-end delay and the first computing delay of the computing task. The relationship between the end-to-end delay, the first computing delay, and the first transmission delay of the computing task can be referred to the description of the embodiment in Figure 2(a) above.

[0162] Based on the above scheme, the end-to-end latency corresponding to the computing task includes not only the transmission latency but also the computing latency. Therefore, by accurately obtaining the transmission latency of the computing task based on its end-to-end latency and computing latency, and by determining the corresponding quality of service flow based on the transmission latency of the computing task and transmitting the computing task through the quality of service flow, accurate control of the end-to-end latency corresponding to the computing task can be achieved, thus ensuring the end-to-end latency of the computing task.

[0163] In one possible implementation, step 202a in this application can be replaced by: the computing management function network element determining the quality parameters in the service quality files of multiple service quality flows corresponding to the computing task based on the first transmission delay. Then, the computing management function network element can send the service quality files of these multiple service quality flows to the access network device. The access network device can dynamically select the corresponding service quality file based on the changes in the computing delay of the computing task, and control the service quality flow of the computing task based on the service quality file. For example, assume that the end-to-end delay for computing task #1 is 500ms, the first computing delay is 200ms, the first transmission delay is 300ms, and the uplink and downlink transmission delays between the terminal device and the computing node are equal, both being 150ms. Assume that the computing node is deployed within the local DN, and the data transmitted between the terminal device and the computing node needs to be routed through the access network device and the user plane network element. The computing management function network element determines the first QoS file for the first QoS flow, the second QoS file for the second QoS flow, and the third QoS file for the third QoS flow. The QoS parameters of the first QoS file include a first QoS delay, which indicates that the uplink or downlink delay of data packets for transmitting computing tasks between the terminal device and the user plane network element via the first QoS flow is 180ms. The QoS parameters of the second QoS file include a second QoS delay, which indicates that the uplink or downlink delay of data packets for transmitting computing tasks between the terminal device and the user plane network element via the second QoS flow is 120ms. The QoS parameters of the third QoS file include a third QoS delay, which indicates that the uplink or downlink delay of data packets for transmitting computing tasks between the terminal device and the user plane network element via the third QoS flow is 100ms. The computing management function network element can dynamically determine the latest computing delay of the computing task; that is, the computing delay for the computing node to process the computing task may be updated from the first computing delay to other computing delays. Then, the computing management function network element sends the latest computing latency to the access network device. The access network device can then select the appropriate Quality of Service (QoS) flow based on this latest latency and perform QoS control on the computing task or its processing result based on the QoS latency in the QoS file of the selected QoS flow. For example, at the first moment, the access network device selects to transmit the computing task or its processing result via the second QoS flow, therefore the QoS latency of this computing task between the terminal device and the user plane network element is 120ms; at the second moment, the access network device selects to transmit the computing task or its processing result via the third QoS flow, therefore the QoS latency of this computing task between the terminal device and the user plane network element is 100ms.

[0164] The embodiments of Figures 2(a) and 2(b) described above will be explained below with reference to the examples in Figures 3 and 4.

[0165] Figure 3 is a flowchart illustrating a communication method provided in an embodiment of this application. The method includes the following steps:

[0166] Step 300: The terminal device completes the attachment process.

[0167] During the attach process, the terminal device receives an attach response message from the AMF, which may include an identifier of the allowed computing tasks for the terminal device.

[0168] Step 301: The terminal device sends a compute plane connection request to the AMF. Correspondingly, the AMF receives the compute plane connection request.

[0169] The compute plane connection request includes an identifier of the terminal device. Optionally, the compute plane connection request also includes an identifier of the requested compute task. Optionally, the compute plane connection request also includes the end-to-end latency corresponding to the requested compute task. The requested compute task is included among the allowed compute tasks. The end-to-end latency refers to the time delay between the terminal device sending the data corresponding to the requested compute task and the terminal device receiving the processing result of the compute task.

[0170] For example, the requested computation tasks include task#1 and task#2, and the end-to-end latency of the requested computation tasks includes end-to-end latency #1 and end-to-end latency #2. Here, end-to-end latency #1 refers to the latency between the terminal device sending the information of task#1 and the terminal device receiving the processing result of task#1, and end-to-end latency #2 refers to the latency between the terminal device sending the information of task#2 and the terminal device receiving the processing result of task#2.

[0171] Step 302: AMF sends a computational plane connection request to CMF. Correspondingly, CMF receives the computational plane connection request.

[0172] This compute plane connection request is the same as the compute plane connection request in step 301, which means that the AMF will forward the compute plane connection request received from the terminal device to the CMF.

[0173] Step 303: The CMF sends a compute management association establishment request to the PCF. Correspondingly, the PCF receives the compute management association establishment request.

[0174] The computation plane management association establishment request includes the identifier of the requested computation task and the end-to-end latency of the request corresponding to the requested computation task.

[0175] Step 304: PCF determines the PCC rule.

[0176] In one possible implementation, the PCF determines the PCC rules based on the local configuration.

[0177] In one possible implementation, the PCF obtains the terminal's subscription data from the UDM. This subscription data includes identifiers of the allowed computing tasks of the terminal device and the allowed end-to-end latency corresponding to the allowed computing tasks. The allowed end-to-end latency refers to the time delay between the terminal device sending data corresponding to the allowed computing task and the terminal device receiving the processing result of that computing task.

[0178] PCF determines PCC rules based on the identifier of the allowed computing task of the terminal device and the allowed end-to-end latency corresponding to the allowed computing task, as well as the identifier of the requested computing task and the requested end-to-end latency corresponding to the requested computing task. The PCC rules include the identifier of the requested computing task and the end-to-end latency corresponding to the requested computing task.

[0179] For example, for any computing task in the requested computing task (taking the first computing task as an example), if the end-to-end latency of the request corresponding to the first computing task is greater than or equal to the allowed end-to-end latency corresponding to the first computing task, then the PCC rule includes the identifier of the first computing task and the latency requirement of the request corresponding to the identifier of the first computing task.

[0180] For example, for any computing task in the requested computing task (taking the first computing task as an example), if the end-to-end latency of the request corresponding to the first computing task is less than the allowed end-to-end latency corresponding to the first computing task, then the PCC rule includes the identifier of the first computing task and the allowed latency requirement corresponding to the identifier of the first computing task.

[0181] Step 305: The PCF sends a compute management association establishment response to the CMF. Correspondingly, the PCF receives the compute management association establishment response.

[0182] The computational plane management association establishes responses including PCC rules.

[0183] Alternatively, steps 301 to 303 can be omitted. Instead, the AF sends a compute plane connection request to the PCF via the NEF. This request includes one or more of the following: the terminal identifier, the identifier of the requested compute task, and the requested end-to-end latency corresponding to the compute task. This triggers the PCF to execute steps 304 and 305.

[0184] As another implementation method, steps 301 to 302 above can be omitted. Instead, the AF sends a computing plane connection request to the CMF via the NEF. The computing plane connection request includes one or more of the terminal identifier, the identifier of the requested computing task, and the end-to-end delay corresponding to the requested computing task, thereby triggering the CMF to execute step 303 and the PCF to execute steps 304 and 305.

[0185] Step 306: CMF determines the first computation delay on the computing node and the first transmission delay between the computing node and the terminal device for the computation task requested in the PCC rule.

[0186] For a computation task in a PCC rule, you can first determine the first computation delay of the computation task on the computing node, and then determine the first transmission delay of the computation task between the computing node and the terminal device. Alternatively, you can first determine the first transmission delay of the computation task between the computing node and the terminal device, and then determine the first computation delay of the computation task on the computing node.

[0187] The initial computational latency of the requested computation task on the computing node can be either the maximum computational latency of the task on the computing node or the computational latency of the task on the computing node under a certain task partitioning method. The computational latency of the same computational task differs on the computing node under different partitioning methods. Furthermore, the computational latency of the task on the computing node when partitioned is less than the computational latency of the task on the computing node when not partitioned (i.e., the maximum computational latency). This is because after partitioning, the computational task is jointly executed by the terminal device and the computing node, thus reducing the computational latency of the computing node. One possible interpretation is to decompose the computational task into two computational subtasks, with the terminal device and the computing node each executing one subtask. For example, if computation task #1 is not split, the computation latency (i.e., maximum computation latency) of computation task #1 on the computing node is 400ms. Under task splitting method #1, the computation latency of computation task #1 on the computing node is 100ms. Under task splitting method #2, the computation latency of computation task #1 on the computing node is 200ms. Under task splitting method #3, the computation latency of computation task #1 on the computing node is 300ms.

[0188] The first transmission delay refers to the transmission delay of the computing task between the terminal device and the computing node. The first transmission delay consists of two parts: one part is the uplink transmission delay of the terminal device sending the data corresponding to the computing task to the computing node, and the other part is the downlink transmission delay of the computing node sending the processing result of the computing task to the terminal device.

[0189] The following sections introduce different implementation methods for determining the computation latency of computing tasks in PCC rules on computing nodes and the transmission latency between computing nodes and terminal devices, based on different scenarios.

[0190] Application Scenario 1: The first computational latency of a computing task on a computing node refers to the maximum computational latency of the computing task on the computing node.

[0191] For this application scenario 1, there are two different implementation methods.

[0192] In method 1, the CMF obtains the first computational latency of the computational task on the computing node based on the task's identifier. Then, it determines the difference between the end-to-end latency corresponding to the computational task and the first computational latency on the computing node, using this difference as the first transmission latency between the computing node and the terminal device. Optionally, assuming the uplink transmission latency and downlink transmission latency of the computational task between the computing node and the terminal device are the same, then both the uplink and downlink transmission latency of the computational task between the computing node and the terminal device are half of the first transmission latency between the computing node and the terminal device.

[0193] For example, if the end-to-end latency of computation task #1 is 500ms, and the first computation latency of computation task #1 on the computing node is 300ms, then the first transmission latency of computation task #1 between the computing node and the terminal device is 200ms. If the uplink transmission latency of computation task #1 between the computing node and the terminal device is the same as the downlink transmission latency of computation task #1 between the computing node and the terminal device, then both the uplink and downlink transmission latency of computation task #1 between the computing node and the terminal device are 100ms.

[0194] In method 2, CMF obtains the first transmission delay of the computing task between the computing node and the terminal device based on the identifier of the computing task, and then determines the difference between the end-to-end delay of the computing task and the first transmission delay of the computing task between the computing node and the terminal device as the first computing delay of the computing task on the computing node.

[0195] For example, the end-to-end latency of computation task #1 is 500ms, and the first transmission latency of computation task #1 between the computing node and the terminal device is 200ms. Optionally, the uplink transmission latency of computation task #1 between the computing node and the terminal device is the same as the downlink transmission latency of computation task #1 between the computing node and the terminal device, that is, both are 100ms. Therefore, CMF determines that the first computation latency of computation task #1 between the computing node and the terminal device is 300ms.

[0196] Application Scenario 2: The computation latency of a computing task on a computing node refers to the maximum computation latency of a computing task on a computing node under a certain task splitting method.

[0197] For this application scenario 2, there are two different implementation methods.

[0198] In method 1, the CMF obtains the task partitioning method corresponding to the computing task based on the task identifier, and obtains the computing latency (i.e., the first computing latency) of the computing task on the computing node under the task partitioning method. Then, it determines the difference between the end-to-end latency of the computing task and the computing latency of the computing task on the computing node under the task partitioning method as the first transmission latency of the computing task between the computing node and the terminal device. Optionally, assuming that the uplink transmission latency of the computing task between the computing node and the terminal device is the same as the downlink transmission latency of the computing task between the computing node and the terminal device, then both the uplink transmission latency and the downlink transmission latency of the computing task between the computing node and the terminal device are half of the first transmission latency of the computing task between the computing node and the terminal device.

[0199] For example, if the end-to-end latency of computation task #1 is 500ms, the selected task partitioning method for computation task #1 is task partitioning method #1, and the computation latency of computation task #1 on the computation node under task partitioning method #1 is 100ms, then the first transmission latency of computation task #1 between the computation node and the terminal device is 400ms. If the uplink transmission latency of computation task #1 between the computation node and the terminal device is the same as the downlink transmission latency of computation task #1 between the computation node and the terminal device, then both the uplink and downlink transmission latency of computation task #1 between the computation node and the terminal device are 200ms.

[0200] In method 2, the CMF obtains the second transmission delay and determines the task splitting method and the first computation delay based on the end-to-end delay corresponding to the computation task and the second transmission delay. Then, based on the end-to-end delay corresponding to the computation task and the first computation delay, it determines the first transmission delay of the computation task between the computing node and the terminal device. The second transmission delay is either the monitored or predicted uplink or downlink service quality delay between the terminal device and the computing node, or the monitored or predicted uplink or downlink service quality delay between the terminal device and the computing node. The first computation delay is used to indicate the processing delay of the computation task under this task splitting method. The first transmission delay is calculated based on the first computation delay and the end-to-end delay corresponding to the computation task.

[0201] This implementation method 2 is the implementation method B described in the embodiment of Figure 2(a), and you can refer to the foregoing description for details.

[0202] Step 307: The CMF sends a compute plane connection configuration message to the compute node. Correspondingly, the compute node receives the compute plane connection configuration message.

[0203] The compute plane connection configuration message includes one or more of the following: a first rule, tunnel information of the access network device, address of the terminal device (e.g., IP address), task partitioning method corresponding to the requested compute task, or first compute latency corresponding to the requested compute task.

[0204] The first rule is used to indicate the correspondence between at least one QFI and at least one quintuple of information, and / or to indicate the correspondence between at least one QFI and the identifier of the requested computation task. The information used to indicate the correspondence between at least one QFI and the identifier of the requested computation task can also be referred to as first information.

[0205] In one possible implementation, for any two computation tasks in the requested computation tasks, if the difference between the first transmission delays of the two computation tasks is less than or equal to a first threshold, then the two computation tasks correspond to the same QFI.

[0206] Optionally, the tunnel information of the access network device is at the terminal device level.

[0207] In step 308, the CMF sends a computing plane connection configuration message to the access network device via the AMF. Correspondingly, the access network device receives the computing plane connection configuration message.

[0208] The compute plane connection configuration message includes one or more of the following: tunnel information of the compute node, a QoS profile, or the first compute latency of the requested compute task on the compute node. A QFI is used to uniquely identify a QoS flow. The QoS profile includes the QFI and QoS parameters, such as QoS latency, which indicates the uplink or downlink transmission latency of data packets for transmitting compute tasks via the first QoS flow between the terminal device and the user plane network element, or the QoS latency indicates the uplink or downlink transmission latency of data packets for transmitting compute tasks via the first QoS flow between the terminal device and the compute node (e.g., deployed within the access network equipment).

[0209] In one possible implementation, the access network device can calculate the time it will take to receive the processing result of the computing task from the computing node based on the first computing delay of the requested computing task on the computing node. This allows the access network device to prepare scheduling resources in advance, improving communication efficiency. For example, if the access network device knows that the first computing delay of computing task #1 is 100ms, after sending the uplink data packet of computing task #1 to the computing node, it predicts that it will receive the downlink data packet of computing task #1 from the computing node after (100+x)ms. Therefore, it can schedule the transmission resources for sending the downlink data packet in advance. Here, x represents the round-trip time for data transmission between the access network device and the computing node.

[0210] Optionally, the tunnel information of the computing node is at the terminal device granularity.

[0211] In step 309, the CMF sends a compute plane connection configuration message to the terminal device via the AMF. Correspondingly, the terminal device receives the compute plane connection configuration message.

[0212] The compute plane connection configuration message includes one or more of the following: QoS rules, first indication information, task splitting method corresponding to the requested compute task, or address of the compute node (e.g., IP address).

[0213] This QoS rule is used to indicate the correspondence between at least one QFI and at least one 5-tuple of information, and / or to indicate the correspondence between at least one QFI and the identifier of the requested computing task. The information used to indicate the correspondence between at least one QFI and the identifier of the requested computing task can also be referred to as first information.

[0214] In one possible implementation, for any two computation tasks in the requested computation tasks, if the difference between the first transmission delays of the two computation tasks is less than or equal to a first threshold, then the two computation tasks correspond to the same QFI.

[0215] The first indication information is used to instruct the terminal device to carry the identifier of the computing task in the header of the data packet of the computing task.

[0216] Optionally, if the CMF does not send the first rule mentioned above to the compute node, the compute plane connection configuration message also includes second indication information. This second indication information instructs the terminal device to add a reflective QoS indicator (RQI) to the uplink data packet of the compute task. Thus, after the compute node processes the uplink data packet and obtains the processing result, it places the downlink data packet carrying the processing result in the same QoS stream for transmission based on the RQI in the uplink data packet; that is, the QFI contained in the downlink data packet is the same as the QFI contained in the uplink data packet.

[0217] Step 310: The access network device sends an RRC configuration message to the terminal device. Correspondingly, the terminal device receives the RRC configuration message.

[0218] This RRC configuration message is used to configure air interface resources for the terminal device.

[0219] After completing the above configurations, the terminal device, access network device, and computing node can transmit uplink data packets and / or downlink data packets.

[0220] In the uplink direction, when a terminal device needs to send uplink data packets for a computing task, it determines the corresponding QFI based on QoS rules and the five-tuple information (or the identifier of the computing task corresponding to the uplink data packet), and adds the QFI to the header of the uplink data packet, thus mapping the uplink data packet to the corresponding QoS stream for transmission. The terminal device sends the uplink data packet to the access network device, and the access network device performs corresponding QoS processing on the uplink data packet according to the QoS parameters of the QoS file corresponding to the QoS stream. In one possible implementation, when the access network device and the computing node are directly connected, the access network device sends the uplink data packet to the computing node through the tunnel corresponding to the uplink data packet between the access network device and the computing node.

[0221] In the downlink direction, the compute node receives the uplink data packet, processes it accordingly to obtain the processing result, and carries the processing result in the downlink data packet. The compute node also determines the corresponding QFI based on the first rule and the five-tuple information corresponding to the downlink data packet (or the identifier of the computation task corresponding to the downlink data), and adds the QFI to the header of the downlink data packet, thus mapping the downlink data packet to the corresponding QoS stream for transmission. In one possible implementation, when the access network device and the compute node are directly connected, the compute node sends the downlink data packet to the access network device through the tunnel corresponding to the downlink data packet between the compute node and the access network device. In this downlink data packet, the header carries the QFI, and the body carries the processing result. The access network device determines the QoS file corresponding to the QFI based on the downlink data packet and performs corresponding QoS processing on the downlink data packet according to the QoS file. The access network device then sends the downlink data packet to the terminal device, thereby enabling the terminal device to receive the downlink data packet.

[0222] Based on the above scheme, the end-to-end latency corresponding to the computing task includes not only the transmission latency of the computing task but also the computing task itself. Therefore, by accurately determining the transmission latency of the computing task based on its end-to-end latency and computing latency, and by determining the corresponding quality of service flow based on the transmission latency of the computing task and transmitting the computing task through the quality of service flow, accurate control of the end-to-end latency corresponding to the computing task can be achieved, thus ensuring the end-to-end latency of the computing task.

[0223] Figure 4 is a schematic flowchart of a communication method provided in an embodiment of this application. The main difference between the embodiment in Figure 4 and the embodiment in Figure 3 is that the embodiment in Figure 4 adds a process of the computing node predicting the computation delay of the computing task to obtain the predicted computation delay (i.e., the second computation delay) and obtaining the updated transmission delay based on the second computation delay.

[0224] The method includes the following steps:

[0225] Steps 400 to 406 are the same as steps 300 to 306 in the embodiment of Figure 3.

[0226] In step 407, the CMF sends a compute plane connection configuration message to the compute node. Correspondingly, the compute node receives the compute plane connection configuration message.

[0227] The compute plane connection configuration message includes one or more of the following: a first rule, tunnel information of the access network device, address of the terminal device (e.g., IP address), task partitioning method corresponding to the requested compute task, or first compute latency corresponding to the requested compute task.

[0228] The first rule is used to indicate the correspondence between at least one QFI and at least one quintuple of information, and / or to indicate the correspondence between at least one QFI and the identifier of the requested computation task. The information used to indicate the correspondence between at least one QFI and the identifier of the requested computation task can also be referred to as first information.

[0229] In one possible implementation, for any two computation tasks in the requested computation tasks, if the difference between the transmission delays of the two computation tasks is less than or equal to a first threshold, then the two computation tasks correspond to the same QFI.

[0230] In addition, the computing plane connection configuration message also includes a third indication information, which is used to indicate the computation latency of the computing task predicted by the computing node.

[0231] Optionally, the tunnel information of the access network device is at the terminal device level.

[0232] Steps 408 to 410 are the same as steps 308 to 310 in the embodiment of Figure 3.

[0233] Step 411: The computing node predicts the computation latency of the requested computing task and obtains the second computation latency of the requested computing task.

[0234] The accuracy of the second computation delay obtained by the computing node in step 411 can be higher than that of the first computation delay in step 406.

[0235] For example, a computing node can determine the second computation delay of a computing task based on a first parameter, a second parameter, and / or a third parameter corresponding to the computing task. The first parameter indicates the time required to generate the first token of the computing task, the second parameter indicates the time interval between two adjacent tokens generated, and the third parameter indicates the number of tokens generated. For instance, if the first parameter indicates the time required to generate the first token of the computing task is T1, the second parameter indicates the time interval between two adjacent tokens generated is T2, and the third parameter indicates the number of tokens generated is N, then the second computation delay = T1 + (N-1)*T2.

[0236] Step 412: The computing node reports the second computation delay of the requested computing task.

[0237] As one implementation method, the computing node reports the second computation delay of the requested computing task to the access network device, so that the access network device can schedule downlink resources based on the second computation delay of the requested computing task. For example, the access network device determines the time to obtain the processing result of the computing task based on the second computation delay of the requested computing task, and then determines the time when the processing result of the computing task will arrive at the access network device, so as to schedule downlink resources in advance based on the time when the processing result of the computing task arrives at the access network device.

[0238] As another implementation method, the computing node reports the second computing delay of the requested computing task to the access network device. The access network device can then determine the transmission delay of the requested computing task (i.e., the transmission delay between the terminal device and the computing node, which consists of two parts: one part is the uplink transmission delay of the terminal device sending the data corresponding to the computing task to the computing node, and the other part is the downlink transmission delay of the computing node sending the processing result of the computing task to the terminal device) based on the second computing delay of the requested computing task and the end-to-end delay of the request corresponding to the computing task.

[0239] As another implementation method, the compute node reports the second computation delay of the requested compute task to the CMF. The CMF can update the transmission delay of the requested compute task based on the second computation delay, and thus update the QoS flow corresponding to the requested compute task. The updated transmission delay of the compute task determined by the CMF can also be called the third transmission delay. This third transmission delay consists of two parts: one is the uplink transmission delay of the terminal device sending the data corresponding to the compute task to the compute node, and the other is the downlink transmission delay of the compute node sending the processing result of the compute task to the terminal device.

[0240] As another implementation method, the computing node reports the second computing delay of the requested computing task to the AF. The AF can determine the transmission delay of the requested computing task based on the second computing delay of the requested computing task and the end-to-end delay of the request corresponding to the requested computing task. Then, through the QoS update process, the AF notifies the core network to update the transmission delay of the requested computing task (i.e., the transmission delay between the terminal device and the computing node, which consists of two parts: one part is the uplink transmission delay of the terminal device sending the data corresponding to the computing task to the computing node, and the other part is the downlink transmission delay of the computing node sending the processing result of the computing task to the terminal device).

[0241] In the embodiment shown in Figure 4, the computing node can also predict the computation latency of the computing task and calculate a second computation latency for the task. In another implementation, the access network device can also predict the computation latency of the computing task to obtain the second computation latency, and then report the second computation latency of the computing task to the core network element (e.g., UPF element, CMF element, or PCF element). Optionally, the access network device can also determine an updated QoS latency based on the second computation latency of the computing task and update the QoS file based on the updated QoS latency. This QoS latency is used to indicate the uplink or downlink latency of data packets transmitted between the terminal device and the UPF element (or computing node) via Quality of Service (QoS) flow for the computing task.

[0242] Based on the above scheme, the end-to-end latency corresponding to the computing task includes not only the transmission latency but also the computing latency. Therefore, by determining the transmission latency of the computing task based on its end-to-end latency and computing latency, and by determining the corresponding quality of service flow based on its transmission latency, the computing task can be transmitted through the quality of service flow. This enables accurate control of the end-to-end latency corresponding to the computing task, ensuring the end-to-end latency of the computing task.

[0243] Figure 5 illustrates a possible exemplary block diagram of the communication device involved in the embodiments of this application. As shown in Figure 5, the communication device 500 may include modules or units for implementing the method embodiments described above. In one possible design, the communication device 500 includes a processing unit 502 and a communication unit 503. Optionally, the communication device 500 may further include a storage unit 501 for storing device program code and / or data.

[0244] The communication device 500 can also be a network-side device in the above embodiments, such as a network-side computing management function network element, a module (e.g., circuit, chip or chip system) in the computing management function network element, or a logic node, logic module or software that can implement all or part of the computing management function network element functions.

[0245] For example, in one embodiment, the processing unit 502 is configured to obtain the first transmission delay of the computing task between the terminal device and the computing node based on the end-to-end delay corresponding to the computing task and the first computing delay of the computing task on the computing node; wherein, the end-to-end delay corresponding to the computing task refers to the delay between the terminal device sending the data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node, and the first computing delay is used to indicate the delay of processing the computing task; and determine the quality of service parameters in the quality of service file of the first quality of service stream corresponding to the computing task based on the first transmission delay.

[0246] In one possible implementation, the communication unit 503 is used to send first information to the terminal device and / or the computing node, the first information being used to map the computing task to the first quality of service stream for transmission.

[0247] In one possible implementation, the communication unit 503 is used to send the first computation delay to the access network device, the first computation delay being used by the access network device to schedule data packets for the computation task.

[0248] In one possible implementation, the communication unit 503 is used to send the first computation delay to the computing node, the first computation delay being used to indicate that the processing of the computing task is completed within the first computation delay.

[0249] In one possible implementation, the communication unit 503 is configured to receive policy charging control rules from the policy control network element, the policy charging control rules including the identifier of the computing task and the end-to-end delay corresponding to the computing task.

[0250] In one possible implementation, the processing unit 502 is further configured to determine the end-to-end latency corresponding to the computing task based on a local policy.

[0251] In one possible implementation, the processing unit 502 is further configured to obtain the first computing latency corresponding to the computing task from a local network, network management network, network storage function network element, or unified data warehouse network element.

[0252] In one possible implementation, the processing unit 502 is further configured to acquire a second transmission delay, wherein the second transmission delay is an uplink or downlink quality of service delay between the terminal device and the computing node that is monitored or predicted, or an uplink or downlink quality of service delay between the terminal device and the user plane network element that is monitored or predicted; and to determine the first computing delay based on the end-to-end delay corresponding to the computing task and the second transmission delay.

[0253] In one possible implementation, the first computational delay is used to indicate the delay in processing the computational task under the first task splitting method.

[0254] In one possible implementation, the communication unit 503 is configured to receive a second computation delay of the computation task on the computing node from the computing node, the second computation delay being used to indicate the predicted processing delay of the computation task by the computing node; the processing unit 502 is further configured to obtain a third transmission delay of the computation task between the terminal device and the computing node based on the end-to-end delay corresponding to the computation task and the second computation delay; and to determine the quality of service parameters of the quality of service file of the second quality of service stream corresponding to the computation task or update the quality of service parameters of the quality of service file of the first quality of service stream based on the third transmission delay.

[0255] In one possible implementation, the communication unit 503 is configured to receive capability information from the terminal device, the capability information being used to indicate whether the terminal device has tokenization capability; the processing unit 502 is further configured to select the computing node based on the capability information.

[0256] The communication device 500 can also be a network-side device in the above embodiments, such as a network-side computing node, a module (e.g., a circuit, chip, or chip system) in the computing node, or a logic node, logic module, or software that can implement all or part of the computing node's functions.

[0257] For example, in one embodiment, processing unit 502 is configured to receive, via communication unit 503, an uplink data packet of a computing task transmitted on a first quality of service (QoS) stream, the uplink data packet including an identifier of the first QoS stream corresponding to the computing task; and to send a downlink data packet of the computing task, the downlink data packet including the identifier of the first QoS stream, the downlink data packet being obtained by processing the uplink data packet within a first computing delay corresponding to the computing task; wherein, the QoS parameters in the QoS file corresponding to the identifier of the first QoS stream are determined based on the first transmission delay of the computing task, the first transmission delay being determined based on the end-to-end delay corresponding to the computing task and the first computing delay, the end-to-end delay corresponding to the computing task referring to the delay between the terminal device sending data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node, the first computing delay being used to indicate the delay of processing the computing task.

[0258] In one possible implementation, the processing unit 502 is further configured to receive first information via the communication unit 503, the first information being used to map the computing task to the first quality of service stream for transmission.

[0259] In one possible implementation, the processing unit 502 is further configured to send a second computational delay of the computational task on the computing node via the communication unit 503, the second computational delay being used to indicate the predicted processing delay of the computational task by the computing node.

[0260] In one possible implementation, the processing unit 502 is further configured to receive indication information via the communication unit 503, the indication information being used to instruct the computing node to send the latency predicted by the computing node for processing the computing task.

[0261] In one possible implementation, the processing unit 502 is further configured to determine the second computation delay based on the first parameter, the second parameter, and the third parameter corresponding to the computation task; wherein the first parameter is used to indicate the time required to generate the first token of the computation task, the second parameter is used to indicate the time interval between generating two adjacent tokens of the computation task, and the third parameter is used to indicate the number of tokens generated for the computation task.

[0262] In one possible implementation, the uplink data packet further includes an identifier of the computing task, and the downlink data packet further includes an identifier of the computing task.

[0263] The communication device 500 can be a terminal device-side device in the above embodiments, such as a terminal device or a communication module in a terminal device, or a circuit or chip in a terminal device that is responsible for communication functions.

[0264] For example, in one embodiment, processing unit 502 is configured to send an uplink data packet of a computing task on a first quality of service (QoS) stream via communication unit 503, the uplink data packet including an identifier of the first QoS stream corresponding to the computing task; and receive a downlink data packet, the downlink data packet including the identifier of the first QoS stream, the downlink data packet being obtained by a computing node processing the uplink data packet within a first computing delay corresponding to the computing task; wherein, the QoS parameters in the QoS file corresponding to the identifier of the first QoS stream are determined based on a first transmission delay of the computing task, the first transmission delay being determined based on the end-to-end delay corresponding to the computing task and the first computing delay, the end-to-end delay corresponding to the computing task referring to the delay between the terminal device sending data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node, the first computing delay being used to indicate the delay of processing the computing task.

[0265] In one possible implementation, the processing unit 502 is further configured to receive first information via the communication unit 503, the first information being used to map the computing task to the first quality of service stream for transmission.

[0266] In one possible implementation, the processing unit 502 is further configured to send capability information via the communication unit 503, the capability information being used to indicate whether the terminal device has tokenization capability.

[0267] In one possible implementation, the uplink data packet further includes an identifier of the computing task, and the downlink data packet further includes an identifier of the computing task.

[0268] In one possible design, when the communication device 500 is a terminal device or a communication module within a terminal device, the function of the processing unit 502 can be implemented by one or more processors. Specifically, the processor may include a modem chip, or a system-on-a-chip (SoC) chip or a SIP chip including a modem core. The function of the communication unit 503 can be implemented by transceiver circuitry.

[0269] In one possible design, when the communication device 500 is a circuit or chip responsible for communication functions in a terminal device, such as a modem chip or a system-on-a-chip (SoC) or SIP chip including a modem core, the function of the processing unit 502 can be implemented by a circuit system including one or more processors or processor cores in the aforementioned chip. The function of the communication unit 503 can be implemented by interface circuits or data transceiver circuits on the aforementioned chip.

[0270] It is understood that the division of units in the above-described device is merely a logical functional division. One function can correspond to one functional unit, or two or more functions can be integrated into one functional unit. In actual implementation, all or some units can be integrated onto a single physical entity, or distributed across different physical entities. Furthermore, the aforementioned functional units can be implemented in hardware, software, or a combination of both. Whether a function is executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for specific applications, but such implementations should not be considered beyond the scope of this application.

[0271] In one example, the functional unit in any of the above devices may be one or more integrated circuits configured to implement the above methods, such as: one or more application-specific integrated circuits (ASICs), or one or more central processing units (CPUs), one or more microcontroller units (MCUs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.

[0272] In one example, storage unit 501 may include random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory and / or registers, etc.

[0273] Figure 6 is a schematic diagram of the structure of a terminal device 600 provided in an embodiment of this application. This terminal device 600 can correspond to the terminal devices shown in Figures 1, 2(b), 3, or 4, and is used to implement the operation of the terminal devices in the above embodiments. As shown in Figure 6, the terminal device includes: one or more antennas 610, a radio frequency processing system 620, and a processor system 630.

[0274] In the downlink or sidelink direction, the RF processing system 620 receives RF signals through the antenna 610 and sends the RF-processed signals to the processor system 630 for further processing. In the uplink or sidelink direction, the processor system 630 processes the information from the terminal device side and sends it to the RF processing system 620, which then processes the signal and transmits it through the antenna 610.

[0275] In one example, the radio frequency (RF) processing system 620 serves as the communication interface for external communication of the terminal device and may include a radio frequency frontend (RFFE) 621 and an RF transceiver 622. The RFFE 621 is primarily used for one or more processing operations, such as shaping, passband selection, or gain adjustment, on the RF signals received by the antenna or those to be transmitted through the antenna. It may include one or more components such as RF switches, duplexers, filters, power amplifiers, antenna tuners, and low-noise amplifiers. The RFFE 621 can be a circuit system composed of multiple discrete components or integrated into one or more chips. The RF transceiver 622 processes the RF signals received by the RFFE into baseband / IF signals for further processing by the processor system 630, and processes the baseband / IF signals provided by the processor system 630 into RF signals for transmission to the RFFE 621. The baseband / IF signals transmitted between the RF transceiver 622 and the processor system 630 can be digital or analog signals. The radio frequency transceiver 622 can be implemented by one or more chips, which are commonly referred to as radio frequency chips (RFICs).

[0276] In one example, processor system 630 may include one or more processors for processing signals and executing one or more communication protocols. Optionally, processor system 630 may also include memory 636. In one example, the one or more processors include at least one baseband processor 631 (also known as a modem processor). Memory 636 is used to store data and / or computer program instructions. Optionally, processor system 630 may also include one or more application processors 632 for implementing processing of the terminal device operating system and application layer. Optionally, processor system 630 may also include one or more of a voice subsystem 633, a multimedia subsystem 634, or an interface circuit 635. The voice subsystem 633 is used to process voice signals, the multimedia subsystem 634 is used to handle multimedia-related operations, such as video encoding / decoding, image processing, etc., and the interface circuit 635 is used to enable communication with other terminal device components, such as display 640, input device 650, memory 660, etc. The above-mentioned components in processor system 630 can communicate with each other via a bus or communication interface circuit.

[0277] In one example, the processor system 630 can be packaged as a single processor chip, such as a SoC chip or a SIP chip. In another example, the processor system 630 can be a system composed of multiple chips; for example, the baseband processor 631 can be packaged as a single chip, or packaged with part or all of the circuitry of the radio frequency processing system into a single chip.

[0278] In one example, memory 636 can be on-chip memory, i.e., located on the processor system 630 chip. In another example, memory 660 can be off-chip memory, i.e. located outside the processor system 630 chip.

[0279] In one example, the baseband processor 631 may include one or more processor cores 6311 and interface circuitry 6314. The one or more processor cores 6311 are used to process signals and execute one or more communication protocols. Optionally, the baseband processor 631 may also include a memory 6312 for storing at least a portion of the corresponding computer program instructions and / or data. In one example, the one or more processor cores 6311 execute the computer program instructions stored in the memory 6312 to implement the relevant operations in the above method embodiments. In this disclosure, the memory 6312 storing the corresponding computer program instructions and / or data may mean that the memory 6312 stores all the corresponding computer program instructions and / or data for the processor core 6311 to execute; or it may mean that the memory 6312 stores a portion of the corresponding computer program instructions and / or data, which includes the computer program instructions and / or data currently required to be executed by the processor core 6311. The memory 6312 can store different portions of computer program instructions and / or data multiple times for the processor core 6311 to execute in order to implement the relevant operations in the above method embodiments. Interface circuit 6314 serves as a communication interface for communication with other components, such as transmitting signals with RF processing system 620, communicating with other subsystems and related components of processor system 630 via bus, such as transmitting data control signals with application processor 632, and transmitting data or computer program instructions with memory 636 or memory 660. Optionally, to reduce the load on the processor core, baseband signal processing circuit 6313 can also be provided to perform at least some baseband signal processing, including one or more of signal demodulation, modulation, encoding, or decoding.

[0280] In one example, the communication device provided in this application may be a terminal device 600, including a communication module of a processor system 630 and a radio frequency system 620, or a baseband processor 631.

[0281] The processor, processor system, application processor, baseband processor, processor circuit, or processor core mentioned above can be collectively referred to as a processor. The processor may include one or more of the following: central processing unit (CPU), digital signal processor (DSP), microprocessor unit (MPU), microcontroller unit (MCU), graphics processing unit (GPU), field programmable gate array (FPGA), artificial intelligence processor (AI processor), or neural processing unit (NPU).

[0282] The aforementioned memory may include one or more of the following storage media: random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), phase-change memory (PCM), resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), cache, register, read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), hard disk, etc. In one example, computer program instructions for executing the above embodiments may be stored on non-volatile memory, such as at least a portion of the aforementioned memory 660 (e.g., one or more of ROM, flash memory, EPROM, or hard disk). When the terminal device is running, the corresponding computer program instructions may be partially or wholly loaded onto a memory with a faster transfer speed than the processor, such as at least a portion of memory 636 and / or memory 6312 (e.g., one or more of RAM, SRAM, DRAM, PCM, RERAM, MRAM, FRAM, cache, or register), for the processor to execute in order to implement the steps in the above method embodiments.

[0283] In one example, the RF transceiver 622 and the RF front-end 621 can also be packaged in a single chip. In another example, the RF transceiver 622, the RF front-end 621, and the baseband processor 631 can also be packaged in a single chip.

[0284] This application provides a chip (or chip system) including a processor for implementing any of the above-described method embodiments.

[0285] This application provides a computer-readable storage medium storing a computer program or instructions that, when executed, implement any of the above-described method embodiments.

[0286] This application provides a computer program product, which includes a computer program or instructions that, when executed, implement any of the above-described method embodiments.

[0287] This application provides a communication system, including one or more of the computing management function network element, computing node, or terminal device in the above method embodiments.

[0288] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disks, portable hard disks, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Furthermore, the ASIC can reside in a first network element or a store-and-forward terrestrial function network element. Alternatively, the processor and storage medium can exist as discrete components in access network equipment or terminal equipment.

[0289] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. A computer program is a set of instructions that directs each step of an action of an electronic computer or other device with message processing capabilities. It is typically written in a programming language and runs on a target architecture. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video optical disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be volatile or non-volatile, or it can include both types of storage media.

[0290] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0291] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. In the textual description of this application, the character " / " generally indicates an "or" relationship between the preceding and following related objects; in the formulas of this application, the character " / " indicates a "division" relationship between the preceding and following related objects.

[0292] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The order of the process numbers described above does not imply the order of execution; the execution order of each process should be determined by its function and internal logic.

[0293] The terms "system" and "network" in this application embodiment are used interchangeably. "At least one" refers to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, "at least one of A, B, or C" includes A, B, C, AB, AC, BC, or ABC; "at least one of A, B, and C" can also be understood as including A, B, C, AB, AC, BC, or ABC. Furthermore, unless otherwise specified, the ordinal numbers such as "first" and "second" mentioned in this application embodiment are used to distinguish multiple objects and are not used to limit the order, sequence, priority, or importance of multiple objects.

[0294] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) that include computer-usable program code.

[0295] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0296] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0297] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0298] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A communication method, characterized in that, include: Based on the end-to-end latency corresponding to the computing task and the first computing latency of the computing task on the computing node, the first transmission latency of the computing task between the terminal device and the computing node is obtained; wherein, the end-to-end latency corresponding to the computing task refers to the latency between the terminal device sending the data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node, and the first computing latency is used to indicate the latency of processing the computing task; Based on the first transmission delay, determine the quality of service parameters in the quality of service file of the first quality of service stream corresponding to the computing task.

2. The method as described in claim 1, characterized in that, Also includes: Send first information to the terminal device and / or the computing node, the first information being used to map the computing task to the first quality of service stream for transmission.

3. The method as described in claim 1 or 2, characterized in that, Also includes: The first calculation delay is sent to the access network device, and the first calculation delay is used by the access network device to schedule the data packets of the calculation task.

4. The method according to any one of claims 1 to 3, characterized in that, Also includes: The first computation delay is sent to the computing node, and the first computation delay is used to indicate that the processing of the computing task is completed within the first computation delay.

5. The method according to any one of claims 1 to 4, characterized in that, Also includes: The system receives policy charging control rules from the policy control network element. The policy charging control rules include the identifier of the computing task and the end-to-end delay corresponding to the computing task.

6. The method according to any one of claims 1 to 4, characterized in that, Also includes: The end-to-end latency corresponding to the computing task is determined based on the local policy.

7. The method according to any one of claims 1 to 6, characterized in that, Also includes: The first computing latency corresponding to the computing task is obtained from local, network management, network storage function network elements or unified data warehouse network elements.

8. The method according to any one of claims 1 to 6, characterized in that, Also includes: The second transmission delay is obtained, which is the uplink or downlink service quality delay between the terminal device and the computing node as monitored or predicted, or the uplink or downlink service quality delay between the terminal device and the user plane network element as monitored or predicted. The first computation delay is determined based on the end-to-end delay corresponding to the computation task and the second transmission delay.

9. The method as described in claim 8, characterized in that, The first computation delay is used to indicate the delay in processing the computation task under the first task splitting method.

10. The method according to any one of claims 1 to 9, characterized in that, Also includes: The computing node receives a second computing latency of the computing task on the computing node, the second computing latency being used to indicate the latency predicted by the computing node for processing the computing task; Based on the end-to-end latency corresponding to the computing task and the second computing latency, the third transmission latency of the computing task between the terminal device and the computing node is obtained. Based on the third transmission delay, determine the service quality parameters of the service quality file of the second service quality stream corresponding to the computing task, or update the service quality parameters of the service quality file of the first service quality stream.

11. The method according to any one of claims 1 to 10, characterized in that, Also includes: Receive capability information from a terminal device, the capability information being used to indicate whether the terminal device has tokenization capability; Select the computing node based on the capability information.

12. A communication method, characterized in that, include: Receive uplink data packets of a computing task transmitted on a first quality of service stream, the uplink data packets including an identifier of the first quality of service stream corresponding to the computing task; Send downlink data packets for the computing task, the downlink data packets including the identifier of the first quality of service flow, the downlink data packets being obtained by processing the uplink data packets within the first computing delay corresponding to the computing task; The service quality parameters in the service quality file corresponding to the identifier of the first service quality flow are determined based on the first transmission delay of the computing task. The first transmission delay is determined based on the end-to-end delay corresponding to the computing task and the first computing delay. The end-to-end delay corresponding to the computing task refers to the delay between the terminal device sending the data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node. The first computing delay is used to indicate the delay of processing the computing task.

13. The method as described in claim 12, characterized in that, Also includes: Receive first information, which is used to map the computing task to the first quality of service stream for transmission.

14. The method as described in claim 12 or 13, characterized in that, Also includes: The computing task is sent a second computing delay on the computing node, the second computing delay being used to indicate the computing node's predicted processing delay for the computing task.

15. The method as described in claim 14, characterized in that, Also includes: The system receives an instruction message, which instructs the computing node to send the predicted latency for processing the computing task.

16. The method as described in claim 14 or 15, characterized in that, Also includes: The second computation delay is determined based on the first parameter, the second parameter, and the third parameter corresponding to the computation task; The first parameter indicates the time required to generate the first token of the computing task, the second parameter indicates the time interval between generating two adjacent tokens of the computing task, and the third parameter indicates the number of tokens generated for the computing task.

17. The method according to any one of claims 12 to 16, characterized in that, The uplink data packet also includes the identifier of the computing task, and the downlink data packet also includes the identifier of the computing task.

18. A communication method, characterized in that, include: Uplink data packets for a computing task are sent on a first quality of service flow, the uplink data packets including an identifier of the first quality of service flow corresponding to the computing task; The downlink data packet is received, the downlink data packet includes the identifier of the first quality of service flow, and the downlink data packet is obtained by the computing node processing the uplink data packet within the first computing delay corresponding to the computing task; The service quality parameters in the service quality file corresponding to the identifier of the first service quality flow are determined based on the first transmission delay of the computing task. The first transmission delay is determined based on the end-to-end delay corresponding to the computing task and the first computing delay. The end-to-end delay corresponding to the computing task refers to the delay between the terminal device sending the data corresponding to the computing task and the terminal device receiving the processing result of the computing task from the computing node. The first computing delay is used to indicate the delay of processing the computing task.

19. The method as described in claim 18, characterized in that, Also includes: Receive first information, which is used to map the computing task to the first quality of service stream for transmission.

20. The method as described in claim 18 or 19, characterized in that, Also includes: Send capability information, which indicates whether the terminal device has tokenization capability.

21. The method according to any one of claims 18 to 20, characterized in that, The uplink data packet also includes the identifier of the computing task, and the downlink data packet also includes the identifier of the computing task.

22. A communication device, characterized in that, Includes modules for performing the method of any one of claims 1 to 11, or the method of any one of claims 12 to 17, or the method of any one of claims 18 to 21.

23. A communication device, characterized in that, The device includes a processor and an interface circuit, the processor being configured to communicate with other devices via the interface circuit to implement the method of any one of claims 1 to 11, or the method of any one of claims 12 to 17, or the method of any one of claims 18 to 21.

24. A computer program product, characterized in that, The computer program product includes instructions that, when executed, implement the method of any one of claims 1 to 11, or the method of any one of claims 12 to 17, or the method of any one of claims 18 to 21.

25. A computer-readable storage medium, characterized in that, The storage medium stores a computer program or instructions that, when executed, implement the method of any one of claims 1 to 11, or the method of any one of claims 12 to 17, or the method of any one of claims 18 to 21.

26. A communication system, characterized in that, Includes at least two of the following devices: computing management function network element, computing node, and terminal equipment; The computing management function network element is used to implement the method described in any one of claims 1 to 11; The computing node is used to implement the method according to any one of claims 12 to 17; The terminal device is used to implement the method according to any one of claims 18 to 21.

Citation Information

Patent Citations

  • Electronic device, method and storage medium for communication system

    CN118202702A

  • Electronic device and method for communication system and storage medium

    WO2023083138A1

  • Wireless communication method and apparatus, device, storage medium, and program product

    WO2023221059A1