Communication method and related device, and communication system

By working together with management and billing network elements, the shortcomings in computing resource statistics and billing in the network architecture are resolved, enabling accurate statistics and flexible billing of computing resources, thus improving the user experience.

WO2026158246A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2026-01-19
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

The existing network architecture lacks an effective statistical method for providing computing resources, resulting in inaccurate billing, inability to meet the computing resource needs of different terminal devices, and inability to charge reasonably based on the complexity of computing tasks and differences in models.

Method used

The management function network element sends the UE's identifier and computing information detection indication to the computing node. The computing node collects and reports computing resource information. The management function network element then passes this information to the billing function network element to generate billing information. Taking into account the differences between the UE, model, computing task, and computing node, the corresponding charging standards are set.

Benefits of technology

It enables accurate statistics and billing of computing resources, improves user experience, and provides a variety of pricing options to meet the needs of different computing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026073502_30072026_PF_FP_ABST
    Figure CN2026073502_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of computing power networks, and relates in particular to a communication method and a related device, and a communication system. The method comprises: a management function network element sends first information to a computing node, the first information comprising an identifier of a UE and indication information for computing information detection, the indication information for computing information detection being used to instruct the computing node to collect statistics on computing information of the UE, and the computing information of the UE being computing information for a target computing task of the UE executed by the computing node; the management function network element receives the computing information of the UE from the computing node; and the management function network element sends the computing information of the UE to a charging function network element, the computing information of the UE being used by the charging function network element to generate charging information of the UE. Embodiments of the present application enable statistics collection on information about computing resources used by a computing node to process a computing task, thereby facilitating subsequent charging for the used computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Communication methods, related equipment and communication systems

[0001] This application claims priority to Chinese Patent Application No. 202510096253.1, filed on January 21, 2025, entitled "Communication Method, Related Devices and Communication System", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computing power networks, and in particular to a communication method, related equipment and communication system. Background Technology

[0003] With the development of wireless communication networks and artificial intelligence technologies, the demand for inference computing power in multimodal interactive artificial intelligence (AI) applications has surged. Due to limitations in computing power, power, and memory, terminal devices (user equipment, UE) face new challenges in developing computing power applications. Industry experts believe that cloud-edge-device collaboration can alleviate the computing power pressure on the device side.

[0004] The existing network architecture is designed with routing methods tailored to communication requirements. For example, after uplink data packets are sent to the user plane function (UPF) network element, the UPF element forwards the packets to the corresponding server based on their destination IP address. The core network processes data packets, only responsible for packet routing and forwarding. A concrete analogy can be drawn from existing applications (APPs) providing generative AI services. Users experience this service through an APP on the UE (User Equipment). For example, the UE's APP sends generative AI tasks to the server: "Generate a picture of a golden shaded cat on the grass," or "Find an ancient poem describing spring scenery." When the remote server receives the task, it selects an appropriate large language model for inference based on the task, obtains a response, and then sends it back to the UE. For example, the response could be multiple images or the text of a relevant ancient poem.

[0005] It's worth noting that generative AI services require significant computing resources. While some well-known OTT providers possess their own servers and large amounts of GPUs / CPUs for inference on large language models, many smaller vendors or user devices lack sufficient computing power for this task. Therefore, in future network development, operators or other open platforms can provide computing nodes / units. These computing units / nodes can deploy large models of multiple computing applications to support data processing. For example, an object requiring inference can send data related to that task to the computing node, which will process it and return the result. However, in scenarios where the network provides computing resources, there are no corresponding statistical methods to visualize the computing resources used to process tasks, making it impossible to bill based on the computing resources used. Summary of the Invention

[0006] This application provides a communication method, related equipment, and communication system. Using this application, it is possible to statistically analyze the computing resources used by computing nodes to process computing tasks, which facilitates subsequent billing of the computing resources used.

[0007] Firstly, embodiments of this application provide a communication method. This method should be applicable to management function network elements. The management function network element can be a session management function (SMF) network element or a compute management function (CMF) network element. The method includes:

[0008] The management function network element sends first information to the computing node. This first information includes the UE's identifier and computing information detection indication information. The computing information detection indication information is used to instruct the computing node to collect the UE's computing information. The UE's computing information is the computing information used by the computing node to execute the UE's target computing task. The management function network element receives the UE's computing information from the computing node.

[0009] In one possible implementation, the management function network element sends the UE's calculation information to the charging function network element, and the UE's calculation information is used by the charging function network element to generate the UE's charging information.

[0010] In one possible implementation, by receiving computing information from the UE at the computing node, the management function network element can monitor or obtain information about the computing resources used by the UE, which can then be used to subsequently send this information to other subscribing network elements. For example, a third-party application function (AF) network element can request computing power from the network for the UE. Optionally, the AF network element will request information about the computing resources used by the UE from the network. These two requests can be a single message or different messages; this application does not limit this. After obtaining the UE's computing information, the management function network element can then further send it to the AF network element.

[0011] The UE's computation information can be information about the computing resources used by the computing node or the first model on the computing node when processing the UE's target computing task.

[0012] As can be seen, by sending the UE's identifier and computing information detection indication information to the computing node through the management function network element, the computing node can collect the UE's computing information when processing the UE's target computing task, that is, the information of the computing resources used when processing the UE's target computing task, and report the UE's computing information to the management function network element. The management function network element sends the UE's computing information to the charging function network element, such as sending the UE's computing information to the charging function (CHF) network element, so that the target network element can generate the UE's charging information based on the UE's computing information, thereby realizing the charging of the computing resources used when processing the UE's target computing task.

[0013] In conjunction with the first aspect, in a feasible implementation, the computational information indicates one or more of the following:

[0014] The number of the first tokens input when a compute node processes a target compute task;

[0015] The number of second tokens output by the computing node when processing the target computing task;

[0016] The amount of computation used by a computing node when processing a target computing task;

[0017] The time taken by a computing node to process the target computing task;

[0018] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0019] It can be seen that by reporting one or more of the above information to the billing function network element, the billing function network element can obtain information on the computing resources used when processing the target computing task of the UE, so as to charge for the computing resources used by the UE.

[0020] In conjunction with the first aspect, in a feasible implementation, the indication information for computational information detection includes the UE's identification information and / or IP address and / or the identification information of the first model.

[0021] It can be seen that by carrying the UE's identification information and / or IP address and / or the first model's identification information in the indication information of the computing information detection, the computing node can statistically analyze the computing resources used when processing the UE's computing tasks, or the computing node can statistically analyze the computing resources used when the first model processes the UE's computing tasks.

[0022] In conjunction with the first aspect, in a feasible implementation, the method of this embodiment further includes:

[0023] The management function unit sends second information to the billing function network element. The second information includes one or more of the following: the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier.

[0024] As can be seen, since the charging standards for computing resources are different for different UEs, by reporting the UE's identifier to the billing function network element, the billing function network element can easily charge for the computing resources used when processing the target computing task of the UE.

[0025] Since different models have different functions and can have different charging standards, by reporting the identifier of the first model to the billing function network element, the billing function network element can charge for the computing resources used by the first model when processing the target computing task of the UE.

[0026] Since different computing tasks require different models or computing resources, the charging standards vary depending on the computing task. By reporting the identifier of the target computing task to the billing function network element, the billing function network element can charge for the computing resources used when processing the target computing task of the UE.

[0027] Because different compute nodes have varying inference speeds and performance, their pricing standards also differ. By reporting the identifier of the compute node to the billing function network element, it becomes easier for the billing function network element to subsequently charge for the computing resources used by the compute node when processing the UE's target computing tasks.

[0028] It's important to understand that the billing function network element can be pre-configured with a charging standard, which corresponds to one or more of the UE's identifier, the first model's identifier, the target computing task's identifier, and the computing node's identifier. Therefore, reporting one or more of these identifiers to the billing function network element facilitates subsequent billing of the computing resources used by the computing node when processing the UE's target computing task. This also provides users with multiple charging standards to choose from, enhancing the user experience.

[0029] In conjunction with the first aspect, in a feasible implementation, the first information also includes reporting rules, which instruct the computing node on how to report the UE's computing information to the management function network element. The reporting methods for the UE's computing information include one or more of the following: periodic reporting, reporting based on a computing information threshold, and reporting based on the number of computing tasks.

[0030] Among them, the reporting based on the number of computing tasks is used to instruct computing nodes to report the UE's computing information to the management function network element when they complete a single or multiple target computing tasks.

[0031] It can be seen that by setting the periodic reporting of computing resource information used by computing nodes when processing the target computing task of the UE, the billing function network element can perform billing periodically, avoiding excessively frequent reporting of billing resources. By setting the reporting of computing resource information used by computing nodes when processing the target computing task of the UE based on computing information thresholds, the management function network element can notify the billing function network element when the computing resource information used by computing nodes when processing the target computing task of the UE exceeds the computing information threshold. By setting the reporting of computing resource information used by computing nodes when processing the target computing task of the UE based on the number of computing tasks, the computing node reports the computing resource information used only after processing the target computing task of the UE.

[0032] In conjunction with the first aspect, in a feasible implementation, the method of this embodiment further includes:

[0033] The management function network element sends third information to the UPF network element. This third information is used by the UPF network element to forward data packets corresponding to the target computing task from the UE to the computing node. In this way, the computing node can process data packets corresponding to the target computing task from the UE to execute the UE's target computing task and collect information on the computing resources used when processing the UE's target computing task.

[0034] In conjunction with the first aspect, in a feasible implementation, the third information includes one or more of the following: tunnel information between the UPF network element and the computing node, data packet detection rules, and corresponding forwarding actions; the data packet detection rules are used to detect data packets of the target computing task, and the tunnel information or forwarding actions are used to forward the data packets of the UE's target computing task to the computing node.

[0035] In conjunction with the first aspect, in a feasible implementation, the packet detection rules include the identification information or IP address of the computing node or the identification information of the first model. Specifically, the destination IP address of the data packets for the UE's target computing task is either the IP address of the computing node or the IP address of the first model.

[0036] It can be seen that by carrying the IP address of the computing node or the IP address of the first model in the packet detection rules, the UPF network element is able to detect the packets of the target computing task.

[0037] In conjunction with the first aspect, in a feasible implementation, the method of this embodiment includes:

[0038] The management function network element determines the computing node based on one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the UE's location information. The computing node meets one or more of the following conditions:

[0039] The model is deployed as indicated by the model identifier;

[0040] The computing task that can execute the target computing task is the computing task type subscribed to by the UE;

[0041] The computational load is not lower than the computational load subscribed by the UE;

[0042] The UE's location information is within the service area of ​​the computing node.

[0043] It can be seen that the management function network element selects computing nodes based on the UE's subscription information and / or the UE's location information, so that the selected computing nodes can meet the UE's computing needs.

[0044] In conjunction with the first aspect, in a feasible implementation, the method of this embodiment further includes:

[0045] The management function network element receives the first rule from the policy control function (PCF) network element. The first rule includes one or more of the following: the model identifier subscribed by the UE, the type of computing task subscribed by the UE, the computing load subscribed by the UE, and the indication information that the UE is allowed to use computing services.

[0046] It can be seen that by sending the UE's subscribed model identifier, the UE's subscribed computing task type, and the UE's subscribed computing load to the management function network element, the UE can use one or more of the indication information of computing services, which makes it easier for the management function network element to select the computing node that processes the UE's computing task or meets the UE's computing needs.

[0047] In conjunction with the first aspect, in a feasible implementation, the method of this embodiment further includes:

[0048] The management function network element receives a computing plane connection request or computing plane update request from the UE. The management function network element receives the first PCC rule from the PCF network element, including: the management function network element receives the first PCC rule from the PCF network element based on the computing plane connection request or computing plane update request.

[0049] or,

[0050] The first PCC rule is sent by the PCF network element to the management function network element based on the computing plane connection request or computing plane update request from the application function (AF) network element.

[0051] It can be seen that if the UE or AF changes the computing requirements involved in the computing application, an update request can be initiated, which makes it easier for the management function network element to manage computing services based on the update requirements.

[0052] Secondly, embodiments of this application provide another communication method. This method is applied to a computing node. The method includes:

[0053] The computing node receives first information from the management function network element, which includes the UE's identifier and computing information detection indication information; the computing node determines the UE's computing information based on the UE's identifier and computing information detection indication information; the UE's computing information is the computing information for the computing node to process the UE's target computing task; the computing node reports the UE's computing information to the management function network element.

[0054] In one possible implementation, the UE's computational information is used to generate the UE's billing information.

[0055] In one possible implementation, the UE's computational information is used to generate monitoring information for the UE's computational services.

[0056] The UE's computation information refers to the information on the computing resources used by the computing node and / or the first model of the computing node to process the UE's target computing task.

[0057] As can be seen, by sending the UE's identifier and computing information detection indication information to the computing node through the management function network element, the computing node processing the UE's target computing task can collect the UE's computing information, that is, the information of the computing resources used when processing the UE's target computing task, and report the UE's computing information to the management function network element. Subsequently, the management function network element can send the UE's computing information to the billing function network element, and the billing function network element can generate the UE's billing information based on the UE's computing information, thus realizing the billing of the computing resources used when processing the UE's target computing task.

[0058] In conjunction with the second aspect, in a feasible implementation, the computational information indicates one or more of the following:

[0059] The number of the first tokens input when a compute node processes a target compute task;

[0060] The number of second tokens output by the computing node when processing the target computing task;

[0061] The amount of computation used by a computing node when processing a target computing task;

[0062] The time taken by a computing node to process the target computing task;

[0063] The memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization used by the computing node when processing the target computing task.

[0064] It can be seen that by reporting one or more of the above information to the management function network element, the management function network element can then report one or more of the above information to the billing function network element, thereby facilitating the billing function network element to charge for the computing resources used when processing the UE's target computing task.

[0065] In conjunction with the second aspect, in a feasible implementation, the indication information for computational information detection includes the UE's identification information and / or IP address and / or the identification information of the first model. The computing node determines the UE's computational information based on the UE's identification and the indication information for computational information detection, including:

[0066] The computing node determines the data packet of the target computing task, the source IP address of the data packet of the target computing task is the IP address of the UE and / or the data packet of the target computing task carries the identification information of the first model and / or the data packet of the target computing task includes the identification information of the UE; the computing node obtains the data packet of the target computing task, determines the number of first tokens input to the computing node or the first model of the computing node and / or the number of second tokens output and / or determines the computing load used by the computing node, the computing load used by the computing node is determined based on the number of first tokens and / or the number of second tokens.

[0067] It can be seen that by determining the data packet of the target computing task and determining the number of first tokens input to the computing node or the first model of the computing node and / or the number of second tokens output when processing the data packet of the target computing task through the first model, it is convenient for the subsequent billing function network element to generate the UE's billing information based on the number of first tokens and / or the number of second tokens.

[0068] In conjunction with the second aspect, in a feasible implementation, the method of this embodiment further includes:

[0069] The computing node determines the first model based on the semantic information of the data packet of the target computing task; or, it determines the first model based on the model ID or the target computing task type in the data packet of the target computing task; or it determines the first model based on the model's IP address or the port number assigned to the model in the data packet; the computing node inputs the first token into the first model for processing to obtain the second token.

[0070] It can be seen that determining the first model by using the relevant information of the data packets of the target computing task enables the computing node to select the appropriate model to process the UE's target computing task, which helps to improve the accuracy of the processing results obtained from processing the target computing task.

[0071] In conjunction with the second aspect, in a feasible implementation, the first information also includes reporting rules, which instruct the computing node on how to report the UE's computing information to the management function network element.

[0072] The methods for reporting UE calculation information include one or more of the following: periodic reporting, reporting based on calculation information thresholds, and reporting based on the number of calculation tasks.

[0073] Among them, the reporting based on the number of computing tasks is used to instruct computing nodes to report the UE's computing information to the management function network element when they complete a single or multiple target computing tasks.

[0074] It can be seen that by setting the periodic reporting of computing resource information used by computing nodes when processing the target computing task of the UE, the billing function network element can perform billing periodically, avoiding excessively frequent reporting of billing resources. By setting the reporting of computing resource information used by computing nodes when processing the target computing task of the UE based on computing information thresholds, the management function network element can notify the billing function network element when the computing resource information used by computing nodes when processing the target computing task of the UE exceeds the computing information threshold. By setting the reporting of computing resource information used by computing nodes when processing the target computing task of the UE based on the number of computing tasks, the computing node reports the computing resource information used only after processing the target computing task of the UE.

[0075] Thirdly, embodiments of this application provide another communication method applied to a charging function network element. The method includes:

[0076] The charging function network element receives the UE's computation information from the management function network element; the UE's computation information is the computation information of the computing node executing the UE's target computation task; the charging function network element generates the UE's charging information based on the UE's computation information.

[0077] The UE's computation information can be information about the computing resources used by the computing node and / or the first model on the computing node to process the UE's target computation task.

[0078] It can be seen that the billing function network element can generate the UE's billing information based on the UE's computing information, and realize the billing of the computing resources used by the computing node when processing the UE's target computing task.

[0079] In conjunction with the third aspect, in a feasible implementation, the computational information indicates one or more of the following:

[0080] The number of the first tokens input when a compute node processes a target compute task;

[0081] The number of second tokens output by the computing node when processing the target computing task;

[0082] The amount of computation used by a computing node when processing a target computing task;

[0083] The time taken by a computing node to process the target computing task;

[0084] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0085] As can be seen, by having the management function network element report one or more of the above information to the billing function network element, the billing function network element can obtain information on the computing resources used when processing the target computing task of the UE, and thus charge for the computing resources used by the UE.

[0086] In conjunction with the third aspect, in a feasible implementation, the method of this embodiment further includes:

[0087] The charging function network element sends a policy counter status notification or a spending limit notification to the PCF network element based on the UE's calculation information or charging information. The notification is used by the PCF network element to determine the service policy for the UE.

[0088] Among them, the service policies for UE include one or more of the following: changing the UE's usage type, changing the model requirements when processing the target computing task of the UE, the computing volume requirements, the required QoS information, computing latency requirements, and / or computing load requirements.

[0089] The required QoS information includes one or more of the following: latency requirements, packet loss rate requirements, and stability requirements.

[0090] It can be seen that, based on the UE's computing information or billing information, the PCF network element sends policy indicator status notifications or expenditure limit notifications, which in turn sends updated rules to the management function network element, so that the management function network element selects a low-performance computing node or the first model for processing the UE's target computing task.

[0091] In conjunction with the third aspect, in a feasible implementation, the method of this embodiment further includes:

[0092] The billing function network element receives second information from the management function network element. The second information includes one or more of the following: the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier.

[0093] Based on the above information, the billing function network element determines the charging standard for the target computing task processed by the UE, and then determines the UE's billing information.

[0094] As can be seen, since the charging standards for computing resources are different for different UEs, by reporting the UE's identifier to the billing function network element, the billing function network element can easily charge for the computing resources used when processing the target computing task of the UE.

[0095] Since different models have different functions and can have different charging standards, by reporting the identifier of the first model to the billing function network element, the billing function network element can charge for the computing resources used by the first model when processing the target computing task of the UE.

[0096] Since different computing tasks require different models or computing resources, the charging standards vary depending on the computing task. By reporting the identifier of the target computing task to the billing function network element, the billing function network element can charge for the computing resources used when processing the target computing task of the UE.

[0097] Because different compute nodes have varying inference speeds and performance, their pricing standards also differ. By reporting the identifier of the compute node to the billing function network element, it becomes easier for the billing function network element to subsequently charge for the computing resources used by the compute node when processing the UE's target computing tasks.

[0098] It's important to understand that the billing function network element can be pre-configured with a charging standard that corresponds to multiple identifiers among the UE, the first model, the target computing task, and the computing node. Therefore, reporting these correspondences to the billing function network element facilitates subsequent billing of the computing resources used by the computing node when processing the UE's target computing task. This also provides users with multiple charging standards to choose from, enhancing the user experience.

[0099] Fourthly, embodiments of this application provide a communication method applied to a UPF network element. The method includes:

[0100] The UPF network element receives third information from the management function network element; this third information is used to instruct the UPF network element to forward data packets from the UE to the computing node. The UPF network element receives data packets from the user equipment UE; when it is determined based on the third information that the data packets from the UE are data packets for the target computing task, the UPF network element forwards the data packets from the UE to the computing node based on the third information.

[0101] It can be seen that by forwarding the data of the UE's target computing task to the computing node through the UPF network element, the computing node can process the data packets corresponding to the UE's target computing task to execute the UE's target computing task, and collect information on the computing resources used when processing the UE's target computing task.

[0102] In conjunction with the fourth aspect, in a feasible implementation, the third information includes tunnel information between the UPF network element and the computing node, or the IP address, packet detection rules, and forwarding actions of the computing node. The method in this embodiment further includes:

[0103] The UPF network element determines whether a data packet from the UE is for the target computing task based on data packet detection rules; the UPF network element forwards the data packet from the UE to the computing node based on third-party information, including:

[0104] UPF network elements forward data packets from the UE to the computing node based on tunnel information between the UPF network element and the computing node, the IP address of the computing node, and / or forwarding actions.

[0105] It can be seen that, firstly, it is determined whether the data packet from the UE is the data packet of the target computing task. Only when it is determined that the data packet from the UE is the data packet of the target computing task is the data packet from the UE forwarded to the computing node based on the third information. This can forward the data packet from the UE to the appropriate computing node for processing, which is beneficial to improving the accuracy of the results obtained by the computing node in processing the data packet of the target computing task.

[0106] Fifthly, embodiments of this application provide a communication method applied to a UE. The method includes:

[0107] The UE sends a computing plane connection request or a computing plane update request to the management function network element;

[0108] Specifically, the computing plane connection request or computing plane update request is used by the PCF network element to send a first rule to the management function network element. The first rule includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use. The UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use are used to determine the computing node. This computing node is used to process the UE's target computing task, collect the UE's computing information, and report the UE's computing information to the management function network element. The UE's computing information is used by the charging function network element to generate the UE's charging information.

[0109] The UE's computation information can be information about the computing resources used by the computing node and / or the first model of the computing node to process the UE's target computing task.

[0110] It can be seen that when the computing node processes the target computing task of the UE, it collects the UE's computing information, that is, the information of the computing resources used when processing the target computing task of the UE, and reports the UE's computing information to the management function network element. The management function network element sends the UE's computing information to the billing function network element. The billing function network element can generate the UE's billing information based on the UE's computing information, thus realizing the billing of the computing resources used when processing the UE's target computing task.

[0111] In conjunction with the fifth aspect, in a feasible implementation, the computational information indicates one or more of the following:

[0112] The number of the first tokens input when a compute node processes a target compute task;

[0113] The number of second tokens output by the computing node when processing the target computing task;

[0114] The amount of computation used by a computing node when processing a target computing task;

[0115] The time taken by a computing node to process the target computing task;

[0116] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0117] As can be seen, by having the management function network element report one or more of the above information to the billing function network element, the billing function network element can obtain information on the computing resources used when processing the target computing task of the UE, and thus charge for the computing resources used by the UE.

[0118] Sixthly, embodiments of this application provide a communication method applied to an AF network element, the method comprising:

[0119] The AF network element sends a compute plane connection request or compute plane update request to the PCF network element;

[0120] Specifically, the computing plane connection request or computing plane update request is used by the PCF network element to send a first rule to the management function network element. The first rule includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use. The UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use are used to determine the computing node. This computing node is used to process the UE's target computing task, collect the UE's computing information, and report the UE's computing information to the management function network element. The UE's computing information is used by the charging function network element to generate the UE's charging information.

[0121] The UE's computation information can be information about the computing resources used by the computing node and / or the first model on the computing node to process the UE's target computation task.

[0122] It can be seen that when the computing node processes the target computing task of the UE, it collects the UE's computing information, that is, the information of the computing resources used when processing the target computing task of the UE, and reports the UE's computing information to the management function network element. The management function network element sends the UE's computing information to the billing function network element. The billing function network element can generate the UE's billing information based on the UE's computing information, thus realizing the billing of the computing resources used when processing the UE's target computing task.

[0123] In conjunction with the sixth aspect, in a feasible implementation, the computational information indicates one or more of the following:

[0124] The number of the first tokens input when a compute node processes a target compute task;

[0125] The number of second tokens output by the computing node when processing the target computing task;

[0126] The amount of computation used by a computing node when processing a target computing task;

[0127] The time taken by a computing node to process the target computing task;

[0128] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0129] As can be seen, by having the management function network element report one or more of the above information to the billing function network element, the billing function network element can obtain information on the computing resources used when processing the target computing task of the UE, and thus charge for the computing resources used by the UE.

[0130] In a seventh aspect, embodiments of this application provide a management function network element, including units or modules for implementing the method provided in the first aspect or any possible implementation of the first aspect.

[0131] Eighthly, embodiments of this application provide a computing node, including units or modules for implementing the methods provided in any two possible implementations of the second or first aspect.

[0132] Ninthly, embodiments of this application provide a billing function network element, including units or modules for implementing the method provided in the third aspect or any possible implementation of the third aspect.

[0133] In a tenth aspect, embodiments of this application provide a UPF network element, including units or modules for implementing the methods provided in the fourth aspect or any possible implementation of the fourth aspect.

[0134] Eleventhly, embodiments of this application provide a UE, including units or modules for implementing the method provided in the fifth aspect or any possible implementation of the fifth aspect.

[0135] In the twelfth aspect, embodiments of this application provide an AF network element, including units or modules for implementing the methods provided in the sixth aspect or any possible implementation of the sixth aspect.

[0136] In a thirteenth aspect, embodiments of this application provide a communication device, including a processor and a memory. The memory is used to store program code. The processor is used to invoke the program code stored in the memory to execute the method provided in the first aspect or any possible implementation of the first aspect, or the method provided in the second aspect or any possible implementation of the second aspect, or the method provided in the third aspect or any possible implementation of the third aspect, or the method provided in the fourth aspect or any possible implementation of the fourth aspect, or the method provided in the fifth aspect or any possible implementation of the fifth aspect, or the method provided in the sixth aspect or any possible implementation of the sixth aspect.

[0137] In a fourteenth aspect, embodiments of this application provide a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform a method provided by a first aspect or any possible implementation of the first aspect, or a method provided by a second aspect or any possible implementation of the second aspect, or a method provided by a third aspect or any possible implementation of the third aspect, or a method provided by a fourth aspect or any possible implementation of the fourth aspect, or a method provided by a fifth aspect or any possible implementation of the fifth aspect, or a method provided by a sixth aspect or any possible implementation of the sixth aspect.

[0138] In a fifteenth aspect, embodiments of this application provide a computer program product that, when run on a computer, causes the computer to perform a method as provided in the first aspect or any possible implementation of the first aspect, or a method as provided in the second aspect or any possible implementation of the second aspect, or a method as provided in the third aspect or any possible implementation of the third aspect, or a method as provided in the fourth aspect or any possible implementation of the fourth aspect, or a method as provided in the fifth aspect or any possible implementation of the fifth aspect, or a method as provided in the sixth aspect or any possible implementation of the sixth aspect.

[0139] In a sixteenth aspect, embodiments of this application also provide a communication method applied to a communication system, the communication system including a management function network element, a computing node, a charging function network element, and a UE, the method comprising:

[0140] The management function network element sends first information to the computing node. The first information includes the UE's identifier and the indication information for computing information detection. The indication information for computing information detection is used to instruct the computing node to collect the UE's computing information. The UE's computing information is the computing information used by the computing node to execute the UE's target computing task.

[0141] The computing node determines the UE's computing information based on the UE's identifier and the indication information detected by the computing information detection, and reports the UE's computing information to the management function network element;

[0142] The management function network element sends the UE's calculation information to the charging function network element;

[0143] The billing function network element generates the UE's billing information based on the UE's calculation information.

[0144] In a seventeenth aspect, embodiments of this application also provide a communication system, which includes a management function network element, a computing node, a charging function network element, and a UE;

[0145] The management function network element is used to send first information to the computing node. The first information includes the UE's identifier and the indication information for computing information detection. The indication information for computing information detection is used to instruct the computing node to collect the UE's computing information. The UE's computing information is the computing information for the computing node to execute the UE's target computing task.

[0146] The computing node is used to determine the UE's computing information based on the UE's identifier and the indication information detected by the computing information, and to report the UE's computing information to the management function network element;

[0147] The management function network element is also used to send the UE's calculation information to the charging function network element;

[0148] The billing function network element is used to generate UE billing information based on UE's calculation information.

[0149] It is understood that the beneficial effects of the embodiments described in the seventh to seventeenth aspects can be referred to the beneficial effects of the foregoing methods, and will not be repeated here. Attached Figure Description

[0150] Figure 1 is a schematic diagram of a communication system architecture provided in an embodiment of this application;

[0151] Figure 2 is a flowchart illustrating a communication method provided in an embodiment of this application;

[0152] Figure 3 is a flowchart illustrating another communication method provided in an embodiment of this application;

[0153] Figure 4 is a flowchart illustrating another communication method provided in an embodiment of this application;

[0154] Figure 5 is a flowchart illustrating another communication method provided in an embodiment of this application;

[0155] Figure 6 is an interactive flowchart of a communication method provided in an embodiment of this application;

[0156] Figure 7 is an interactive flowchart of another communication method provided in an embodiment of this application;

[0157] Figure 8 is an interactive flowchart of another communication method provided in an embodiment of this application;

[0158] Figure 9 is a schematic diagram of the structure of a management function network element provided in an embodiment of this application;

[0159] Figure 10 is a schematic diagram of the structure of a computing node provided in an embodiment of this application;

[0160] Figure 11 is a schematic diagram of the structure of a billing function network element provided in an embodiment of this application;

[0161] Figure 12 is a schematic diagram of the structure of a UPF network element provided in an embodiment of this application;

[0162] Figure 13 is a schematic diagram of the structure of a communication device provided in an embodiment of this application. Detailed Implementation

[0163] The terms “first,” “second,” “third,” and “fourth,” etc., used in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order.

[0164] "Multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating three possible relationships. For example, A and / or B means: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0165] The embodiments of this application will now be described with reference to the accompanying drawings.

[0166] Referring to Figure 1, Figure 1 is a schematic diagram of a communication system architecture provided in an embodiment of this application. This communication system is a schematic diagram of the network architecture of 5th Generation Mobile Networks (5G). It includes a radio access network (RAN), which can be represented as two parts: RAN equipment and a core network (CN). The RAN equipment is used to provide network access functionality for authorized user equipment (UE) in a specific area and can use transmission tunnels of different qualities according to the UE's level and service requirements. For example, the RAN equipment can manage radio resources, provide access services to the UE, and thus complete the forwarding of control information and / or data information between the UE and the CN.

[0167] To facilitate understanding of the embodiments of this application, an application scenario of the embodiments of this application will be described in detail first with reference to FIG1.

[0168] 1. User equipment (UE): This can be referred to as terminal equipment, terminal, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication equipment, user agent, or user device. Terminal equipment can also be a cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, vehicle-mounted device, drone, wearable device, terminal equipment in a 5G network, or terminal equipment in an evolved public land mobile network (PLMN), etc., and this application embodiment does not limit this.

[0169] 2. Access Network (AN): Provides network access for authorized users in a specific area and can use transmission tunnels of different qualities depending on the user's level and service requirements. Access networks can employ different access technologies. Current access network technologies include: radio access network technologies used in 3G systems, radio access network technologies used in 4G systems, or next-generation radio access network (NG-RAN) technologies (such as those used in 5G systems).

[0170] An access network that uses wireless communication technology to implement access network functions can be called a radio access network (RAN). A RAN manages radio resources, provides access services to terminals, and facilitates the forwarding of control signals and user data between terminals and the core network.

[0171] Wireless access network equipment can be, for example, a base station (NodeB), an evolved NodeB (eNB or eNodeB), a next-generation Node base station (gNB) in a 5G mobile communication system, a base station in a future mobile communication system, or an access point (AP) in a Wi-Fi hotspot system. It can also be a wireless controller in a cloud radio access network (CRAN) scenario, or it can be a relay station, access point, vehicle-mounted equipment, drone, wearable device, or network equipment in a 5G network or an evolved PLMN. This application does not limit the specific technology or equipment form used in the wireless access network equipment.

[0172] 3. Access Management Network Element: Primarily used for mobility management and access management, responsible for transmitting user policies between user equipment and PCF network elements, etc. It can be used to implement other functions of the Mobile Management Entity (MME) besides session management. For example, access authorization (authentication) function.

[0173] In 5G communication systems, the access management network element can be an access and mobility management function (AMF) network element. In future communication systems, the access management network element can still be an AMF network element, or it can have other names; this application does not limit this.

[0174] 4. Management Function Network Elements: These are mainly used for session management, allocation and management of Internet Protocol (IP) addresses for user equipment, selection of manageable user plane functions, policy control and billing function interface endpoints, and downlink data communication. Management function network elements include SMF (Software-Defined Function) network elements or CMF (Computational Function) network elements. SMF network elements manage regular sessions, while CMF network elements manage computing sessions.

[0175] In 5G communication systems, management function network elements can be either SMF (Supervisory Management Function) or CMF (Computational Management Function) network elements. In future communication systems, management function network elements can still be either SMF or CMF network elements, or they can have other names; for example, a CMF network element could be an enhancement of an existing SMF network element. This application does not impose any limitations on this. Management function network elements include either SMF or CMF network elements. SMF network elements manage regular sessions, while CMF network elements manage computational sessions.

[0176] 5. User plane network elements: used for packet routing and forwarding, QoS processing of user plane data, completion of user plane data forwarding, session / flow-level billing statistics, bandwidth limiting functions, etc.

[0177] In the system shown in Figure 1, the user plane network elements include a first UPF network element and a second UPF. The first UPF network element supports the UE's regular sessions, which are used for data transmission between the UE and the DN (e.g., EAS). That is, the first UPF network element forwards regular service flows between the UE and the DN. This session can be a PDU session. The first UPF network element can be used to implement user plane-related functions such as packet routing and transmission, packet inspection, service usage reporting, quality of service (QoS) processing, legality monitoring, uplink packet inspection, and downlink packet storage. For example, during the establishment of a PDU session, the SMF network element selects the first UPF network element for the session. This first UPF network element can be called the session anchor, such as the PDU session anchor (PSA). It can also be called the anchor UPF network element or the remote session anchor.

[0178] The second UPF network element supports the UE's computing session, which is used for data transmission between the UE and computing nodes. Specifically, the second UPF network element forwards computing service flows between the UE and the computing nodes. The second UPF network element can be used to implement user plane related functions such as packet routing and transmission, packet inspection, service usage reporting, QoS processing, legality monitoring, uplink packet inspection, and downlink packet storage.

[0179] In 5G communication systems, user plane network elements can be UPF network elements. In future communication systems, user plane network elements can still be UPF network elements, or they can have other names; this application does not limit this.

[0180] 6. Data network element: A network used to provide data transmission.

[0181] In 5G communication systems, data network elements can be data network (DN) elements. In future communication systems, data network elements may still be DN elements, or they may have other names; this application does not limit this. DN elements, also known as packet data networks (PDNs), are typically networks located outside the operator's network, such as third-party networks. A DN element may include one or more EASs, which provide local services to the UE by transmitting data with the UE.

[0182] 7. Policy control network element: A unified policy framework used to guide network behavior, providing policy rule information to control plane functional network elements (such as AMF, SMF, etc.).

[0183] In 4G communication systems, this policy control network element can be a Policy and Charging Rules Function (PCRF) network element. In 5G communication systems, this policy control network element can be a PCF network element. In future communication systems, this policy control network element can still be a PCF network element, or it can have other names; this application does not limit its scope.

[0184] 8. Data management network element: used to handle user equipment identification, access authentication, registration, and mobility management, etc.

[0185] In 5G communication systems, this data management network element can be a unified data management (UDM) network element; in 4G communication systems, this data management network element can be a home subscriber server (HSS) network element. In future communication systems, the data management network element can still be a UDM network element, or it can have other names; this application does not limit this.

[0186] 9. Network Exposure Function (NEF) Element: Used to securely expose services and capabilities provided by the 3rd Generation Partnership Project (3GPP) network functions to the outside world.

[0187] 10. Application Function (AF) Network Element: Provides a specific application layer service to the UE. When providing services to the UE, the AF has requirements for QoS and charging policies and needs to notify the network. Simultaneously, the AF also needs to obtain application-related information from the core network. The AF can possess all the functions defined in the technical specification (TS) 23.901R-15, as well as related functions for application services. That is, in the user plane architecture, the application server (AS) and the UE communicate in the user plane via the UE-RAN-UPF-AS path. The AF can also communicate with other network function (NF) network elements in the 5G core network (5GC) in the control plane architecture via the NEF. For example, it can communicate with the PCF network element via the NEF network element. If the AF is deployed by the 5GC operator, the AF network element can also communicate directly with other NF network elements in the 5GC in the control plane architecture without going through the NEF network element, such as directly communicating with the PCF network element.

[0188] 11. Network Data Analysis Function (NWDAF) Network Element: It can be used to collect data from network elements, AF and operation administration and maintenance (OAM) management system, and analyze the data through machine learning, artificial intelligence and other solutions, and feed it back to network elements, AF, etc. to optimize network or service configuration, thereby providing better network quality and service experience.

[0189] 12. Authentication server function (AUSF) network element: mainly responsible for authenticating users to determine whether to allow users or devices to access the network.

[0190] 13. Service Communication Proxy (SCP) element: It can be used for indirect communication between NFs. Service requests from NFs can be proxied by the SCP.

[0191] 14. Billing Function Network Element: Primarily responsible for providing users with calculated quotas or traffic configurations, and generating user billing invoices based on user traffic consumption or calculation information. In 5G communication systems, this billing function network element can be a CHF network element. In future communication systems, this billing function network element may still be a CHF network element, or it may have other names; this application does not impose any limitations.

[0192] 15. Computing Node: A node in the network used to process computing tasks. This computing node can be independent of other network devices in Figure 1, such as being located after the second UPF network element, or being located after the RAN device and directly connected to the RAN device; this computing node can also be deployed within the network devices in Figure 1, such as being deployed in the second UPF network element, or being deployed in the RAN device.

[0193] It should be understood that the network architecture described above for the embodiments of this application is merely an example, and the network architecture applicable to the embodiments of this application is not limited thereto. Any network architecture capable of realizing the functions of the above-described network elements is applicable to the embodiments of this application.

[0194] It should also be understood that the AMF, SMF, UPF, NEF, PCF, UDM, NWDAF, CMF, NRF, AUSF, and SCP network elements shown in Figure 1 can be understood as network elements in the core network used to implement different functions, such as network slices that can be combined as needed. These core network elements can be independent devices or integrated into the same device to implement different functions. This application does not limit the specific form of the above network elements.

[0195] It should also be understood that the above naming is defined only for the convenience of distinguishing different functions and should not constitute any limitation on this application. This application does not exclude the possibility of using other naming in 5G networks and other future networks. For example, in 6G networks or future networks, some or all of the above-mentioned network terms may be used, or other names may be used. The interface names between the various network elements in Figure 1 are just examples, and the interface names in specific implementations may be other names, which this application does not specifically limit. In addition, the names of the messages (or signaling) transmitted between the above-mentioned network elements are also just examples and do not constitute any limitation on the function of the messages themselves.

[0196] Referring to Figure 2, which is a flowchart illustrating a communication method provided in an embodiment of this application, this method is applied to a management function network element. This management function network element can be the SMF network element in the system shown in Figure 1, or a network element in Figure 1 specifically used for computational management, such as a CMF network element. As shown in Figure 2, the method includes:

[0197] S201. The management function network element sends first information to the computing node; wherein, the first information includes the UE's identifier and indication information for computing information detection. The indication information for computing information detection is used to instruct the computing node to collect the UE's computing information.

[0198] In one example, the UE's computation information is the computation information of the computing node executing the UE's target computation task.

[0199] In one example, the UE's computation information may be information about the computing resources used by the computing node and / or the first model of the computing node when processing the UE's target computing task.

[0200] In one example, the indication information for computational information detection includes the UE's identification information and / or the UE's IP address and / or the identification information of the first model. By carrying the UE's identification information and / or the UE's IP address and / or the identification information of the first model in the indication information for computational information detection, the computing node can statistically analyze the computing resources used when processing the UE's computational tasks, or the computing node can statistically analyze the computing resources used when the first model processes the UE's computational tasks.

[0201] Optionally, the identification information of the first model may be the IP address of the first model, the ID of the first model, the port number assigned to the first model, and / or the ID of the task that the first model can process.

[0202] S202, The management function network element receives the computing information from the UE of the computing node.

[0203] In one feasible implementation, the computational information indicates one or more of the following:

[0204] The number of the first tokens input when a compute node processes a target compute task;

[0205] The number of second tokens output by the computing node when processing the target computing task;

[0206] The amount of computation used by a computing node when processing a target computing task;

[0207] The time taken by a computing node to process the target computing task;

[0208] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0209] Memory access bandwidth refers to the maximum amount of data read from or written to main memory per unit of time. When a single SIM card cannot deploy an instance, instances need to be deployed across SIM cards. Communication between SIM cards is required when processing computing tasks; interconnect bandwidth refers to the transmission bandwidth between SIM cards. Inference runtime information includes the number of SIM cards used by the compute node to process the UE's target computing tasks and the average throughput per SIM card. Deployment resource information includes the number of instances used by the compute node to process the UE's target computing tasks and the number of SIM cards required for each instance.

[0210] Model computing power utilization rate is the ratio of computing power consumed by model computation to hardware computing resources. Hardware computing power utilization rate is the ratio of hardware computing power consumed by computation to the hardware computing power of the computing node.

[0211] It should be noted that during the processing of the target computing task, the computing node receives the data packet of the target computing task. In one possible implementation, the computing node calls a first model to process the data packet. In a specific example, the computing node converts the relevant data in the data packet of the target computing task into a first token, then inputs the first token into the first model for processing. The first model outputs a second token, which is the processing result of the computing node calling the first model to process the data packet of the target computing task. For example, if the target computing task is a text-to-image task, and the data packet of the target computing task carries the text "Please generate a photo of a little girl playing on the ground," this text is the relevant data in the data packet. The computing node tokenizes the text "Please generate a photo of a little girl playing on the ground" to obtain the first token.

[0212] In another specific example, the target computation task is a text-to-image task. The data packet for the target computation task carries the first token corresponding to the text "Please generate a photo of a little girl playing on the ground." In this case, the relevant data in the data packet for the target computation task is this first token. In this scenario, the compute node does not need to perform a tokenization operation.

[0213] It should be noted that a token is the smallest discrete unit in text or sequence data. In natural language processing, a token can be a word, a subword (such as a letter, syllable, or subword fragment), or a character.

[0214] It can be seen that by reporting one or more of the above information to the billing function network element, the billing function network element can obtain information on the computing resources used when processing the target computing task of the UE, thereby realizing the billing of the computing resources used by the UE.

[0215] In one feasible implementation, the first information also includes reporting rules, which instruct the computing node on how to report the UE's computing information to the management function network element. The reporting methods for the UE's computing information include one or more of the following: periodic reporting, reporting based on a computing information threshold, and reporting based on the number of computing tasks.

[0216] Among them, the reporting based on the number of computing tasks is used to instruct the computing node to report the UE's computing information to the management function network element when it completes one or more target computing tasks.

[0217] It can be seen that by setting the periodic reporting of computing resource information used by computing nodes when processing the target computing task of the UE, the billing function network element can perform billing periodically, avoiding excessively frequent reporting of computing information. By setting the reporting of computing resource information based on a computing information threshold, the management function network element can notify the billing function network element when the computing resource information used by the computing node exceeds the threshold. By setting the reporting of computing resource information based on the number of computing tasks, the computing node only reports the computing resource information used after completing the target computing task of the UE.

[0218] By setting a threshold for reporting computing resources used by computing nodes when processing the target computing task of the UE, the management function network element sends a notification when the computing resources used by the computing nodes when processing the target computing task of the UE exceed the computing information threshold.

[0219] By setting the information on the computing resources used by the computing node when processing the target computing task of the UE based on the number of computing tasks, the computing node can report the information on the computing resources used only after processing the target computing task of the UE.

[0220] S203. The management function network element sends the UE's calculation information to the charging function network element.

[0221] The UE's calculation information is used by the charging function network element to generate the UE's charging information.

[0222] In one feasible implementation, the management function unit also sends second information to the charging function network element, the second information including one or more of the following: the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier.

[0223] It should be noted that since the charging standards for computing resources differ for different UEs, reporting the UE's identifier to the billing function network element facilitates the billing function network element in charging for the computing resources used when processing the UE's target computing task.

[0224] Since different models have different functions and can have different charging standards, by reporting the identifier of the first model to the billing function network element, the billing function network element can charge for the computing resources used by the first model when processing the target computing task of the UE.

[0225] Since different computing tasks require different models or computing resources, the charging standards vary depending on the computing task. By reporting the identifier of the target computing task to the billing function network element, the billing function network element can charge for the computing resources used when processing the target computing task of the UE.

[0226] Because different compute nodes have varying inference speeds and performance, their pricing standards also differ. By reporting the identifier of the compute node to the billing function network element, it becomes easier for the billing function network element to subsequently charge for the computing resources used by the compute node when processing the UE's target computing tasks.

[0227] It's important to understand that the billing function network element can be pre-configured with a charging standard that corresponds to one or more of the UE's identifier, the first model's identifier, the target computing task's identifier, and the computing node's identifier. Therefore, reporting the correspondence between the UE's identifier, the first model's identifier, the target computing task's identifier, and the computing node's identifier to the billing function network element facilitates subsequent billing of the computing resources used by the computing node when processing the UE's target computing task. This also provides users with multiple charging standards, allowing them to choose according to their needs, thus improving the user experience.

[0228] In one possible implementation, by receiving computing information from the UE (User Equipment) at the computing node, the management function network element can monitor or obtain information about the computing resources used by the UE, which can then be sent to other subscribing network elements. For example, a third-party AF (Automatic Feedback) network element can request computing power from the network for the UE. Optionally, the AF network element may request information about the computing resources used by the UE from the network. These two requests can be a single message or different messages; this application does not limit this. After obtaining the UE's computing information, the management function network element can then send it to the AF network element.

[0229] In one feasible implementation, the management function network element sends third information to the UPF network element. This third information is used by the UPF network element to forward data packets corresponding to the target computing task from the UE to the computing node. This method enables the computing node to process data packets corresponding to the target computing task from the UE to execute the UE's target computing task and to collect information on the computing resources used in processing the UE's target computing task.

[0230] The third information includes one or more of the following: tunnel information between the UPF network element and the computing node, data packet detection rules, and corresponding forwarding actions; the data packet detection rules are used to detect data packets of the target computing task, and the tunnel information or forwarding actions are used to forward the data packets of the UE's target computing task to the computing node.

[0231] It should be noted that there is a tunnel between the management function network element and the computing node. After the management function network element detects the data packet of the target computing task, it forwards the data packet to the computing node through the tunnel.

[0232] In one feasible implementation, the packet detection rules include the identifier or IP address of the computing node, or the identifier information of the first model. Specifically, the destination IP address of the data packet for the UE's target computing task is either the IP address of the computing node or the identifier information of the first model carried by the data packet for the UE's target computing task.

[0233] It can be seen that by carrying the IP address of the computing node or the identification information of the first model in the packet detection rules, the UPF network element is able to detect the packets of the target computing task.

[0234] In one feasible implementation, the method of this embodiment includes:

[0235] The management function network element determines the computing node based on one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the UE's location information. The computing node meets one or more of the following conditions:

[0236] The model is deployed with the model identifier indicated by the UE's subscription;

[0237] The computing task that can execute the target computing task is the computing task type subscribed to by the UE;

[0238] The computational load is not lower than the computational load subscribed by the UE;

[0239] The UE's location information is within the service area of ​​the computing node.

[0240] It can be seen that the management function network element selects computing nodes based on the UE's subscription information and / or UE's location information, so that the selected computing nodes can meet the UE's computing needs, that is, the selected computing nodes can handle the computing tasks requested by the UE.

[0241] It should be noted that the subscription information of the UE mentioned above can be pre-configured in the management network element, or obtained by the PCF network element from the UDM network element and sent to the management function network element by carrying it in the first rule (such as the policy and charging control (PCC) rule).

[0242] In one feasible implementation, the method of this embodiment further includes:

[0243] The management function network element receives a computing plane connection request or computing plane update request from the UE. The management function network element receives the first PCC rule from the PCF network element, including: the management function network element receives the first PCC rule from the PCF network element based on the computing plane connection request or computing plane update request.

[0244] or,

[0245] The first PCC rule is sent by the PCF network element to the management function network element based on the computing plane connection request or computing plane update request from the AF network element.

[0246] It should be noted that when a UE needs to request computing resources from the network side, the UE sends a computing plane connection request to the management function network element through the AMF network element. After receiving the UE's computing plane connection request, the management function network element obtains the UE's subscription information. Optionally, the UE's subscription information can be sent from the PCF network element to the management function network element, or it can be obtained by the management function network element from its local configuration. After obtaining the UE's subscription information, the management function network element selects a computing node to provide computing resources for the UE and sends the first information to that computing node.

[0247] When a UE needs more computing resources from the network side, or when a UE needs to specify a particular computing model or request a specific computing service, the UE sends a computing plane update request to the management function network element through the AMF network element. The computing plane update request carries at least one of the following: the computing resources requested by the UE, the identifier of the requested computing model, and the identifier of the requested computing service, as well as the UE's ID. The management function network element reselects a computing node for the UE based on at least one of the following: the computing resources requested by the UE, the identifier of the requested computing model, and the identifier of the requested computing service.

[0248] The aforementioned compute plane connection request and compute plane update request can be triggered by the AF network element. The AF network element sends a compute plane connection request or compute plane update request to the PCF network element. Upon receiving the compute plane connection request or compute plane update request from the AF network element, the PCF network element obtains the UE's subscription data from the UDM network element and sends the UE's subscription information to the management function network element so that the management function network element can select a compute node for the UE.

[0249] As can be seen, by sending the UE's identifier and computing information detection indication information to the computing node through the management function network element, the computing node can collect the UE's computing information when processing the UE's target computing task, that is, the information of the computing resources used when processing the UE's target computing task, and report the UE's computing information to the management function network element. The management function network element sends the UE's computing information to the target network element, such as sending the EU's computing information to the billing function network element, so that the target network element can generate the UE's billing information based on the UE's computing information, thus realizing the billing of the computing resources used when processing the UE's target computing task.

[0250] Referring to Figure 3, which is a flowchart illustrating another communication method provided in an embodiment of this application, this method is applied to the computing node in Figure 1. As shown in Figure 3, the method includes:

[0251] S301. The computing node receives first information from the management function network element, the first information including the UE's identifier and the indication information for computing information detection.

[0252] In one example, the UE's identifier could be the user's phone number, the UE's physical address, or a subscription permanent identifier (SUPI).

[0253] S302. The computing node determines the UE's computing information based on the UE's identifier and the indication information detected by the computing information.

[0254] In one example, the UE's computation information is the computation information of the computing node for processing the UE's target computation task.

[0255] In another example, the UE's computation information can be information about the computing resources used by the computing node and / or the first model of the computing node to process the UE's target computing task.

[0256] In one feasible implementation, the computation information detection indication information includes the UE's identification information and / or the UE's IP address and / or the identification information of the first model. The computing node determines the UE's computation information based on the UE's identification and the computation information detection indication information, including:

[0257] The computing node determines the data packet of the target computing task, wherein the source IP address of the data packet of the target computing task is the IP address of the UE and / or the data packet of the target computing task carries the identification information of the first model, or the data packet of the target computing task carries the identification information of the UE; the computing node obtains the data packet of the target computing task to determine the number of first tokens input into the first model and / or the number of second tokens output by the first model; and / or determines the computing load used by the computing node, wherein the computing load used by the computing node is determined based on the number of first tokens and / or the number of second tokens.

[0258] In one feasible implementation, the method of this embodiment further includes:

[0259] The computing node determines the first model based on the semantic information of the data packet of the target computing task; or, it determines the first model based on the ID of the first model or the type of the target computing task in the data packet of the target computing task; or it determines the first model based on the IP address of the first model in the data packet or the port number assigned to the first model; the computing node inputs the first token into the first model for processing to obtain the second token.

[0260] Specifically, after receiving a data packet, the computing node determines whether the source IP address of the data packet is the UE's IP address or whether the data packet carries the identification information of the first model. If the source IP address of the data packet is the UE's IP address and / or the data packet carries the identification information of the first model, the computing node determines that the data packet is the data packet for the UE's target computing task. Optionally, the identification information of the first model includes the IP address of the first model, the ID of the first model, the port number allocated to the first model, and / or the ID of the task that the first model can process.

[0261] The computing node determines the first model of the data packets for processing the target computing task of the UE in the following way:

[0262] Method 1: The computing node determines the type of the target computing task based on the semantic information (prompt) of the data packet, and then determines the first model based on the type of the target computing task. For example, if the semantic information of the target computing task data packet includes "Please generate a photo of a little girl playing on the ground", the computing node determines that the target computing task is a text-to-image task, and then determines that the first model is a text-to-image model.

[0263] Method 2: The computing node determines the first model based on the ID of the first model carried in the data packet of the target computing task.

[0264] Method 3: The computing node determines the first model based on the type or ID of the target computing task carried in the data packet, wherein the first model is capable of executing the target computing task.

[0265] Method 4: The computing node determines the first model based on the IP address of the first model carried in the data packet of the target computing task or the port number assigned to the first model.

[0266] After determining the first model, the computing node converts the relevant data in the data packet of the target computing task into a first token, and then inputs the first token into the first model for processing. The first model outputs a second token, which is the processing result of the computing node calling the first model to process the data packet of the target computing task. For example, if the target computing task is a text-to-image task, and the data packet of the target computing task carries the text "Please generate a photo of a little girl sitting on the grass playing," this text is the relevant data in the data packet of the target computing task. The computing node tokenizes the text "Please generate a photo of a little girl sitting on the grass playing," obtaining the first token. The text-to-image model of the computing node processes the first token and outputs the second token, which includes the image of the little girl sitting on the grass playing.

[0267] During the first model of the computing node processing the target computing task of the UE, the computing node collects computing information from the UE, wherein the computing information of the UE indicates one or more of the following:

[0268] The number of the first tokens input when the computing node processes the target computing task of the UE;

[0269] The number of second tokens output by the computing node when processing the target computing task of the UE;

[0270] The amount of computation used by the computing node when processing the target computing task of the UE;

[0271] The time taken by the computing node to process the target computing task of the UE;

[0272] The memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization used by the computing node when processing the target computing task of the UE.

[0273] Memory access bandwidth refers to the maximum amount of data read from or written to main memory per unit of time. When a single SIM card cannot deploy an instance, instances need to be deployed across SIM cards. Communication between SIM cards is required when processing computing tasks; interconnect bandwidth refers to the transmission bandwidth between SIM cards. Inference runtime information includes the number of SIM cards used by the compute node to process the UE's target computing tasks and the average throughput per SIM card. Deployment resource information includes the number of instances used by the compute node to process the UE's target computing tasks and the number of SIM cards required for each instance.

[0274] Model computing power utilization rate is the ratio of computing power consumed by model computation to hardware computing resources. Hardware computing power utilization rate is the ratio of hardware computing power consumed by computation to the hardware computing power of the computing node.

[0275] In one example, the computational load is the number of floating-point operations. Specifically, the computational load used by the computing node to process the UE's target computational task can be calculated as follows:

[0276] In one example, the computational cost used by the compute node to process the UE's target computational task is (a+b)*X, where a and b are the number of tokens input to the first model and the number of tokens output by the first model when executing the UE's computational task, respectively, and X is the floating-point operation cost of the first model for processing a single sample. a and b are the number of the first token and the second token, respectively. In another example, the computational cost used by the compute node to process the UE's target computational task is (a+b)*N*D, where N is the number of layers in the first model, and D is the number of dimensions or channels.

[0277] S303. The computing node reports the UE's computing information to the management function network element.

[0278] In one possible implementation, the UE's computational information is used to generate the UE's billing information.

[0279] In one possible implementation, the UE's computational information is used to generate monitoring information for the UE's computational services.

[0280] In one feasible implementation, the first information also includes reporting rules, which instruct the computing node on how to report the UE's computing information to the management function network element. Therefore, the computing node reports the UE's computing information to the management function network element based on the reporting rules. It should be noted that the specific details of the method for reporting the UE's computing information can be found in the relevant description of the embodiment corresponding to Figure 2, and will not be repeated here.

[0281] It can be seen that by sending the UE's identifier and computing information detection indication information to the computing node through the management function network element, the computing node processing the UE's target computing task can collect the UE's computing information, that is, the information of the computing resources used when processing the UE's target computing task, and report the UE's computing information to the management function network element. Subsequently, the management function network element can send the UE's computing information to the billing function network element, and the billing function network element can generate the UE's billing information based on the UE's computing information, thus realizing the billing of the computing resources used when processing the UE's target computing task.

[0282] Referring to Figure 4, which is a flowchart illustrating another communication method provided in an embodiment of this application, this method is applied to the charging function network element in Figure 1. As shown in Figure 4, the method includes:

[0283] S401. The charging function network element receives the calculation information of the UE from the management function network element; the calculation information of the UE is the calculation information of the computing node executing the target computing task of the UE.

[0284] The UE's computation information can be information about the computing resources used by the computing node and / or the first model of the computing node to process the UE's target computing task.

[0285] S402. The billing function network element generates the UE's billing information based on the UE's calculation information.

[0286] In one feasible implementation, the computational information indicates one or more of the following:

[0287] The number of the first tokens input when the computing node processes the target computing task of the UE;

[0288] The number of second tokens output by the computing node when processing the target computing task of the UE;

[0289] The amount of computation used by the computing node when processing the target computing task of the UE;

[0290] The time taken by the computing node to process the target computing task of the UE;

[0291] The computing node uses one or more of the following when processing the target computing task of the UE: memory, memory access bandwidth, interconnect bandwidth, inference operation information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0292] Memory access bandwidth refers to the maximum amount of data read from or written to main memory per unit of time. When a single SIM card cannot deploy an instance, instances need to be deployed across SIM cards. Communication between SIM cards is required when processing computing tasks; interconnect bandwidth refers to the transmission bandwidth between SIM cards. Inference runtime information includes the number of SIM cards used by the compute node to process the UE's target computing tasks and the average throughput per SIM card. Deployment resource information includes the number of instances used by the compute node to process the UE's target computing tasks and the number of SIM cards required for each instance.

[0293] Model computing power utilization rate is the ratio of computing power consumed by model computation to hardware computing resources. Hardware computing power utilization rate is the ratio of hardware computing power consumed by computation to the hardware computing power of the computing node.

[0294] In one example, the UE's computation information indicates the number of first tokens input by the computing node when processing the target computation task, or the number of first tokens output by the computing node when processing the target computation task. The charging function network element obtains the price of a single token based on the UE's identifier, and determines the UE's charging information based on the price of a single token and the number of first tokens and / or the number of second tokens. For example, the UE's charging information includes the product of the number of first tokens and the price of a single token, or the product of the number of second tokens and the price of a single token, or the maximum value or the sum of the product of the number of first tokens and the price of a single token and the product of the number of second tokens and the price of a single token.

[0295] In another example, the UE's computation information indicates the amount of computation used by the computing node to process the UE's target computing task, such as the number of floating-point operations. The billing function network element obtains the unit price of the computation, such as the unit price of 1G of floating-point operations. The billing function network element determines the UE's billing information based on the unit price of the computation and the amount of computation used by the computing node to process the UE's target computing task. For example, the UE's billing information includes the product of the unit price of the computation and the amount of computation used by the computing node to process the UE's target computing task.

[0296] In another example, the UE's computation information indicates the time taken by the compute node to process the UE's target computation task. It's important to note that the time taken by the compute node to process the UE's target computation task refers to the time from when the compute node receives the data packet containing the UE's target computation task to when it sends out the processing result of the target computation task; in other words, it's the processing time of the compute node for the data packet containing the UE's target computation task. The charging function network element obtains the price per unit time for using the compute node, and determines the UE's charging information based on the price per unit time and the time taken by the compute node to process the UE's target computation task. For example, the UE's charging information may include the product of the price per unit time and the time taken by the compute node to process the UE's target computation task.

[0297] In another example, the UE's computation information indicates the memory used by the compute node to process the UE's target computation task. The billing function network element obtains the memory pricing standard and determines the UE's billing information based on the memory pricing standard and the memory used by the compute node to process the UE's target computation task. For example, the UE's billing information includes the unit price of memory and the product of the memory used by the compute node to process the UE's target computation task.

[0298] In another example, the UE's computation information indicates the memory access bandwidth used by the compute node to process the UE's target computation task. Memory access bandwidth refers to the maximum amount of data read from or written to main memory per unit of time. The memory access process includes: performing the first filling of the KVCache during the pre-filling phase, and performing a write operation to the KVCache while generating the first output; during the decoding phase, the data in the KVCache needs to be retrieved, and the key and value calculated from the output token obtained in the previous round are concatenated and rewritten into the KVCache.

[0299] The billing function network element obtains the charging standard for memory access bandwidth. Based on the charging standard for memory access bandwidth and the memory access bandwidth used by the compute node to process the UE's target computing task, the billing function network element determines the UE's billing information. For example, the UE's billing information includes the product of the unit price of memory access bandwidth and the memory access bandwidth used by the compute node to process the UE's target computing task.

[0300] In another example, the UE's computation information indicates the interconnect bandwidth used by the compute node to process the UE's target computation task. When a single SIM cannot deploy an instance, the instance needs to be deployed across SIMs. Communication between SIMs is required when processing computation tasks; therefore, smaller interconnect bandwidth results in greater time-to-first-token (TTFT) and inference latency, and lower throughput (qps) and throughput (#tokens / s). The billing function network element obtains the pricing standard for interconnect bandwidth and determines the UE's billing information based on the interconnect bandwidth pricing standard and the interconnect bandwidth used by the compute node to process the UE's target computation task. For example, the UE's billing information includes the product of the unit price of interconnect bandwidth and the interconnect bandwidth used by the compute node to process the UE's target computation task.

[0301] In another example, the UE's computation information indicates the inference runtime information when the compute node processes the UE's target computation task, including the number of cards used by the compute node to process the UE's target computation task and the average throughput per card. The billing function network element obtains the corresponding charging standard and determines the UE's billing information based on the charging standard and the inference runtime information. For example, the UE's billing information includes: Inference runtime cost ($ / token) = Throughput (tokens / s) × Number of cards × Price per card per second ($).

[0302] For example, if an inference system deployed with 4 SIM cards in the cloud has a throughput of 200 tokens / s, then the average throughput per SIM card is 50 tokens / s. Multiplying this by the price per SIM card ($0.0004 / s), the inference cost is approximately $8 / M tokens. As the calculation shows, increasing the throughput of a given SIM card can significantly reduce inference costs.

[0303] In another example, the UE's computation information indicates the deployment resource information when the compute node processes the UE's target computation task, including the number of instances used by the compute node to process the UE's target computation task and the number of cards required for each instance. The billing function network element obtains the corresponding charging standard and determines the UE's billing information based on this charging standard and the deployment resource information. For example, the UE's billing information includes: Deployment resource cost ($) = Number of instances × Number of cards required per instance × Price per card.

[0304] For example, assuming a peak concurrency of 64 for a user scenario, and the maximum concurrency that each instance of the inference service can handle is 32, then two instances are needed to ensure that user requests can be processed simultaneously without queuing. Each instance requires 4 GPUs to run, and each GPU costs $1500. Therefore, ensuring that user requests can be processed immediately without queuing requires a deployment cost of $12000. As can be seen from the formula, increasing the maximum concurrency of the inference system can effectively reduce deployment costs while meeting the SLO requirements of the inference service.

[0305] In another example, the UE's computation information indicates the model FLOPs utilization (MFU) when the compute node processes the UE's target computation task. Here, MFU is the ratio of the computing power consumed by the model computation to the hardware computing resources. The billing function network element obtains the prices corresponding to different MFU ranges and determines the UE's billing information based on the MFU when the compute node processes the UE's target computation task and the prices corresponding to different MFU ranges. The UE's billing information includes the price corresponding to the MFU range to which the compute node's MFU for processing the UE's target computation task belongs.

[0306] In another example, the UE's computing information indicates the hardware computing power utilization rate of the computing node when processing the UE's target computing task. Here, the hardware computing power utilization rate is the ratio of the hardware computing power consumed to the computing node's hardware computing power. The billing function network element obtains the prices corresponding to different hardware computing power utilization rate ranges, and determines the UE's billing information based on the hardware computing power utilization rate of the computing node when processing the UE's target computing task and the prices corresponding to different hardware computing power utilization rate ranges. The UE's billing information includes the price corresponding to the hardware computing power utilization rate range to which the computing node's hardware computing power utilization rate for processing the UE's target computing task belongs.

[0307] In one feasible implementation, for the same type of computational information, different rate standards corresponding to different computational information tiers are configured in the billing function network element. For example, when the number of tokens does not exceed 300, the unit price of a token is 0.01 yuan; when the number of tokens exceeds 300, the unit price of the first 300 tokens is 0.01 yuan, and the unit price of tokens exceeding 300 is 0.012 yuan. The billing function network element determines the rate standard of the reported computational information based on the reported computational information and the rate standard corresponding to the tier of the computational information, and then generates billing information based on the rate standard and the reported billing information.

[0308] It's important to note that different computational tasks within a UE vary in difficulty, and the computational resources used by the computing node also differ. Computing nodes use more computational resources to handle complex tasks than simple ones. For example, in the first computational task, "What is the antonym of 'pretty'?", the computing node's model output is "ugly," with 5 first tokens as input and 1 second token as output. In the second computational task, "Generate an annual meeting speech," the computing node's model input is 5 first tokens, but the number of second tokens output is greater than 1. Therefore, the number of second tokens output by the model can distinguish the difficulty of the computational model being processed by the computing node.

[0309] The UE's billing information includes the UE's identifier. Optionally, the UE's billing information may also include one or more of the following:

[0310] The number of the first tokens input when the computing node processes the target computing task of the UE;

[0311] The number of second tokens output by the computing node when processing the target computing task of the UE;

[0312] The amount of computation used by the computing node when processing the target computing task of the UE;

[0313] The time taken by the computing node to process the target computing task of the UE;

[0314] The computing node uses one or more of the following when processing the target computing task of the UE: memory, memory access bandwidth, interconnect bandwidth, inference operation information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0315] Memory access bandwidth refers to the maximum amount of data read from or written to main memory per unit of time. When a single SIM card cannot deploy an instance, instances need to be deployed across SIM cards. Communication between SIM cards is required when processing computing tasks; interconnect bandwidth refers to the transmission bandwidth between SIM cards. Inference runtime information includes the number of SIM cards used by the compute node to process the UE's target computing tasks and the average throughput per SIM card. Deployment resource information includes the number of instances used by the compute node to process the UE's target computing tasks and the number of SIM cards required for each instance.

[0316] Model computing power utilization rate is the ratio of computing power consumed by model computation to hardware computing resources. Hardware computing power utilization rate is the ratio of hardware computing power consumed by computation to the hardware computing power of the computing node.

[0317] In another feasible implementation, when the same billing information corresponds to multiple billing methods, the management function network element will report an identifier for indicating the billing method when reporting the UE's calculation information to the billing function network element. This billing method can be the billing method subscribed by the UE. The billing function network element determines the billing method based on the identifier for indicating the billing method and uses this billing method to bill the reported calculation information.

[0318] Optionally, the UE's billing information may also include the corresponding charging standard.

[0319] It should be noted that the UE's billing information can also be called a charging data record (CDR).

[0320] In one feasible implementation, the method of this embodiment further includes:

[0321] The billing function network element sends a policy indicator status notification or expenditure limit notification to the PCF network element based on the UE's calculation information or billing information. The policy indicator status notification or expenditure limit notification is used by the PCF network element to determine the service policy for the UE.

[0322] The service policies for the UE include one or more of the following: changing the UE's usage type, changing the UE's requirements for the model, changing the requirements for computational load, requesting QoS information, requesting computational latency, and / or requesting computational load. QoS information includes latency, packet loss rate, and stability.

[0323] Specifically, in one example, after receiving the UE's calculation information reported by the computing node, the billing function network element determines whether the UE's calculation information exceeds the calculation information threshold, where the calculation information threshold is determined based on the UE's information; if the UE's calculation information exceeds the calculation information threshold, it sends a policy indicator status notification or expenditure limit notification to the PCF network element.

[0324] In another example, after determining the UE's billing information, the billing function network element determines the UE's balance based on this information. If the UE's balance is lower than a balance threshold, it sends a policy indicator status notification or expenditure limit notification to the PCF network element. The policy indicator status notification or expenditure limit notification is used by the PCF network element to determine the service policy for the UE, such as changing the UE's usage type, changing the model for processing the UE's target computing task, the computational load requirements, QoS requirements, computational latency requirements, and / or computational load requirements. This causes the PCF network element to send an update rule to the management function network element. This update rule carries one or more of the updated UE usage type, updated model information, updated computational load requirements, updated QoS requirements, updated computational latency, and updated computational load requirements. This causes the management function network element to reselect the computing node or first model for processing the UE's target computing task. Specifically, the performance of the reselected computing node may be lower than that of the original computing node; for example, the former's computational load may be lower than the latter's, or the former's computational latency may be higher than the latter's. Alternatively, the performance of the reselected first model may be lower than that of the original first model. It can be seen that, based on the UE's computing information, the PCF network element sends the UE's computing information or billing information to the PCF network element, and sends policy indicator status notifications or expenditure limit notifications to the PCF network element. This causes the PCF network element to send updated rules to the management function network element, which in turn causes the management function network element to select a low-performance computing node or the first model for processing the UE's target computing task.

[0325] In one feasible implementation, the method of this embodiment further includes:

[0326] The charging function network element receives second information from the management function network element. The second information includes one or more of the following: the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier. The UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier are used by the charging function network element to determine the charging standard when the computing node processes the UE's target computing task.

[0327] It can be seen that the billing function network element can generate UE billing information based on the UE's computing information, realizing billing for the computing resources used by computing nodes when processing the UE's target computing tasks. Based on the UE's computing information or billing information, it sends policy indicator status notifications or expenditure limit notifications to the PCF network element, causing the PCF network element to send updated rules to the management function network element, so that the management function network element can select a low-performance computing node or the first model for processing the UE's target computing tasks.

[0328] Referring to Figure 5, which is a flowchart illustrating another communication method provided in an embodiment of this application, this method is applied to a UPF network element, specifically the second UPF network element in Figure 1. As shown in Figure 5, the method includes:

[0329] S501, the UPF network element receives third information from the management function network element; the third information is used to instruct the UPF network element to forward the data packet of the target computing task from the UE to the computing node.

[0330] S502, the UPF network element receives data packets from the user equipment (UE); when it determines that the data packet from the UE is a data packet for the target computing task based on third information, the UPF network element forwards the data packet from the UE to the computing node based on the third information.

[0331] In one feasible implementation, the third information includes tunnel information between the UPF network element and the computing node, or the IP address, packet detection rules, and forwarding actions of the computing node. The method in this embodiment further includes:

[0332] The UPF network element determines whether a data packet from the UE is for the target computing task based on data packet detection rules; the UPF network element forwards the data packet from the UE to the computing node based on third-party information, including:

[0333] UPF network elements forward data packets from the UE to the computing node based on tunnel information between the UPF network element and the computing node, the IP address of the computing node, and / or forwarding actions.

[0334] The packet detection rules include the IP address of the computing node or the IP address of the first model.

[0335] Specifically, after receiving a data packet from the UE, the UPF network element determines whether the data packet belongs to the target computing task based on data packet detection rules. For example, it determines whether the destination address of the data packet is the IP address of the computing node or the IP address of the first model. If the destination address of the data packet is the IP address of the computing node or the IP address of the first model, the UPF network element determines that the data packet from the UE belongs to the target computing task of the UE. The UPF network element determines the tunnel between the UPF network element and the computing node based on the tunnel information between the UPF network element and the computing node, and forwards the data packet of the target computing task to the computing node through the tunnel, the IP address of the computing node, and / or forwarding actions.

[0336] It can be seen that, firstly, it is determined whether the data packet from the UE is the data packet of the target computing task. Only when it is determined that the data packet from the UE is the data packet of the target computing task is the data packet from the UE forwarded to the computing node based on the third information. This can forward the data packet from the UE to the appropriate computing node for processing, which is beneficial to improving the accuracy of the results obtained by the computing node in processing the data packet of the target computing task.

[0337] Referring to Figure 6, Figure 6 is an interactive flowchart illustrating a communication method provided in an embodiment of this application. This method is applied to the system shown in Figure 1. As shown in Figure 6, the method includes:

[0338] S600 and UE complete the attachment process.

[0339] Optionally, during the UE attach process, the AMF network element carries a task identifier or service type identifier in the attach response message sent back to the UE. The task identifier or service identifier indicates the tasks the UE is allowed to perform or the service types the UE is allowed to use, such as text-to-text services, text-to-image services, or interactive AI applications. During the UE attach process, the AMF network element obtains the UE's subscription data from the UDM network element. The UE's subscription data includes indication information indicating the task / service types the UE has subscribed to. Based on this indication information, the AMF network element can determine the task / service types that the UE can use.

[0340] S601, the UE sends a computing plane connection request to the AMF network element.

[0341] The compute plane connection request includes the UE's identifier. Optionally, the request may also include the service type requested by the UE.

[0342] In one example, the service type requested by the UE is a newly defined service type based on the computing service, such as a task / service type in the attach response message fed back by the AMF network element. In this case, the computing plane connection request is a newly defined connection request, and new signaling is required to represent this computing plane connection request.

[0343] In another example, the service type requested by the UE is determined based on the DNN / S-NSSAI assigned to the computing service. In other words, the server type requested by the UE is indicated by the DNN / S-NSSAI assigned to the computing service. In this case, the computing plane connection request reuses an existing session connection request; that is, the computing plane connection request is an existing session connection request that includes the DNN / S-NSSAI assigned to the computing service. This DNN / S-NSSAI is used to determine the service type requested by the UE, and can be seen as an identifier of the service type requested by the UE.

[0344] S602, the AMF network element forwards the computing plane connection request from the UE to the management function network element.

[0345] Specifically, when the AMF network element determines that the UE requests computing services based on the computing plane connection request represented by the new signaling or DNN / S-NSSAI, the AMF network element forwards the computing plane connection request from the UE to the management function network element.

[0346] Optionally, the AMF network element selects a management function network element based on the UE's location information or the DNN / S-NSSAI allocated for computing services. Optionally, when selecting a management function network element, the AMF will consider the service type requested by the UE carried in the computing plane connection request, such as whether the selected management function network element meets the latency requirements of the UE's requested service type. Before this, the AMF network element determines whether the UE has permission to initiate a computing plane connection request; if it determines that the UE has permission to initiate a computing plane connection request, the AMF network element performs the operation of selecting a management function network element. Specifically, the AMF network element determines whether the UE has permission to initiate a computing plane connection request based on whether the UE has subscribed to a computing session. If it determines that the UE has subscribed to a computing session, the AMF network element determines that the UE has permission to initiate a computing plane connection request; if it determines that the UE has not subscribed to a computing session, the AMF network element determines that the UE does not have permission to initiate a computing plane connection request.

[0347] In one possible implementation architecture, the UE can directly send a connection request for the computing plane to the management function network element.

[0348] S603. The management function network element sends a compute management (CM) policy negotiation creation request to the PCF network element.

[0349] S604, the PCF network element obtains the UE's subscription data from the UDM network element.

[0350] The UE's computational subscription data includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computational task type, and the UE's subscribed computational load. The UE's subscribed model includes a default model. The PCF network element generates a first rule based on the UE's computational subscription data. In one example, the first rule is a PCC rule.

[0351] S605, PCF network element sends CM policy negotiation creation response to management function network element.

[0352] The CM strategy negotiation creation response is used to respond to the CM strategy negotiation creation request, and the CM strategy negotiation creation response includes the first rule.

[0353] It should be noted that if the local configuration of the management function network element includes the UE's subscription data or processing rules, then the management function network element does not need to initiate S603, that is, S603-S605 do not need to be executed.

[0354] S606, the management function network element selects the first model or computing node for the UE based on the first rule or local configuration.

[0355] The first model or computing node can be used to execute the UE's target computing task, i.e., the service requested by the UE.

[0356] Specifically, when the compute plane connection request carries the service type requested by the UE, the management function network element determines, based on the service type requested by the UE, either a first model capable of providing the requested service or a first model capable of processing the target compute service from the models subscribed to by the UE, and determines the IP address of the first model or the IP address of the compute node deployed with the first model. When the compute plane connection request does not carry the service requested by the UE, the management function network element uses the default model from the models subscribed to by the UE as the first model, and determines the IP address of the first model or the IP address of the compute node deployed with the first model.

[0357] It should be noted that if there are multiple computing nodes with the first model deployed, the management function network element can select the computing node closest to the UE from the multiple computing nodes with the first model deployed as the computing node.

[0358] The management function network element feeds back the IP address of the first model or the IP address of the computing node to the UE.

[0359] S607, the management function network element sends third information to the UPF network element.

[0360] The third piece of information includes one or more of the following: tunnel information between the UPF network element and the computing node, packet detection rules, and QoS enforcement rules (QER). There is a correspondence between the packet detection rules, the IP address of the computing node, and the QER. It should be understood that the UPF network element here refers to the second UPF network element in Figure 1.

[0361] In one possible design, the compute nodes and UPF network elements are deployed together, or in other words, the UPF itself has the functionality of a compute node. In this case, there is no need to send third-party information.

[0362] In the uplink direction: Packet inspection rules are used to inspect packets for the UE's target computing service. Packet inspection rules include the IP address of the computing node or the IP address of the first model. The IP address of the computing node is used by the UPF network element to forward packets for the UE's target computing task to the computing node, and the IP address of the first model is used to forward packets for the target computing task to the computing node so that the first model of the computing node can process the packets for the target computing task. QER is used to indicate forwarding actions (e.g., forward), tunnel information, and / or forwarding actions are used to forward packets for the UE's target computing task to the computing node. In one example, the packet inspection rule 5-tuple information, such as the 5-tuple information for packets that need to be forwarded to the computing node, or the 5-tuple information for packets for the UE's target computing task, includes the UE's IP address (source IP address) and the IP address of the computing node or the IP address of the first model (destination IP address).

[0363] UPF network elements forward data packets for target computing tasks to computing nodes in the following ways: Method 1: The UPF network element forwards the data packet to the computing node based on the destination IP address (e.g., the computing node's IP address) carried in the UE's target computing task data packet. Method 2: The UPF network element knows the tunnel information between itself and the computing node, and forwards the UE's target computing service data packets to the computing node based on the tunnel between itself and the computing node. Method 3: After receiving the target computing service data packet, the UPF network element determines whether the five-tuple information of the data packet is the same as the five-tuple information in the data packet detection rules. If they are the same, it determines the computing node's IP address and QER based on the correspondence between the packet detection rules, the computing node's IP address, and the QER. The UPF network element then forwards the target computing service data packet to the computing node based on the computing node's IP address and QER.

[0364] In the downlink direction: Packet inspection rules are used to detect packets that need to be forwarded to RAN devices, such as those via a GTP-U tunnel between UPF network elements and RAN devices. Data forwarding between UPF network elements and RAN devices can be found in existing technologies and will not be described here.

[0365] S608, the management function network element sends the fourth information to the RAN equipment.

[0366] The fourth piece of information includes tunnel information between the RAN and UPF, QoS flow identifiers, and the corresponding QoS file. This QoS file includes QoS requirement parameters, such as packet delay budget.

[0367] After receiving the fourth information, the RAN device performs radio resource control (RRC) configuration.

[0368] S609, The management function network element sends the fifth information to the UE.

[0369] The fifth piece of information includes QoS rules, which are used by the UE to map data packets of the target computing task to a QoS stream. Optionally, the fifth piece of information may also include the IP address of the first model or the IP address of the computing node, used to inform the UE of the destination IP address for executing the target computing task.

[0370] S610, the management function network element sends the first information to the computing node.

[0371] The first information includes the UE's ID and indication information for computational information detection. Optionally, the indication information for computational quantity detection may also include the UE's IP address and information about the first model, such as the IP address of the first model, the ID of the first model, the port number assigned to the first model, and the service type of the first model, or one or more of these. Optionally, the first information may also include a session ID and / or reporting rules. The UE's ID and session ID are used to identify a specific session of a specific UE.

[0372] The reporting rules include reporting triggered by computing tasks, periodic reporting, on-demand reporting, and threshold reporting. For example, if the reporting rule is to trigger reporting by computing tasks, after the UE completes a computing task, the computing node will statistically process the computing information of the UE's target computing task and report the computing information to the management function network element.

[0373] It should be noted that task-triggered reporting is mostly used for sudden tasks, such as generative AI tasks. For management function network elements, it is impossible to determine when a sudden task will occur. Periodic reporting is for scenarios with periodic computing tasks, such as a robot sending uplink images at a frequency of 10Hz. In this case, the computing node can periodically report the computing power used to send the images. The reporting period can be the same as the computing task period, or it can be different, such as reporting once a day or once an hour. Threshold reporting refers to the billing function network element configuring a reporting threshold for the computing node. When the computing node uses more or less computing power than the reporting threshold when processing the target computing task, the computing node reports the computing power to the management function network element.

[0374] The present invention does not limit the order of S607-S610.

[0375] S611. The computing node receives data packets from the target computing task of the UE, processes them, and then sends downlink data packets to the UE.

[0376] In one example, the UE knows the IP address of the computing node where the first model resides. The UE sets the destination IP address of the data packet for the target computing task to the IP address of the computing node. The data packet for the target computing task carries identification information of the first model, such as the ID of the first model, the ID of the task executed by the first model, or the port number corresponding to the first model. After receiving the data packet for the target computing task, the computing node parses the identification information of the first model from the data packet, determines the first model based on the identification information, and uses the first model to process the data in the data packet for the target computing task, that is, to execute the UE's computing services.

[0377] In another example, the UE sets the destination address of the data packet of the target computing task to the IP address of the first model. After receiving the data packet of the target computing task, the computing node uses the first model to process the data in the data packet of the target computing task based on the destination address of the data packet of the target computing task, that is, to execute the computing service of the UE.

[0378] In another example, the computing node identifies the semantic information, or prompt, of the data packet of the target computing task. Based on the semantic information of the data packet of the target computing task, it determines the first model. For example, if the target computing service requested by the UE is a text-to-image task: "Generate an image containing a little girl sitting on the grass", the data packet of the target computing task carries the data corresponding to the text-to-image task: "Please generate an image containing a little girl sitting on the grass". The computing node determines that the target computing task requested by the UE is a text-to-image task based on the data corresponding to the text-to-image task, and determines the first model based on the text-to-image task. The first model can execute the text-to-image task and use the first model to process the data in the data packet of the target computing task, that is, to execute the UE's computing service.

[0379] In another example, the destination address of the data packet for the target computing task is the IP address of the computing node, and the source address is the IP address of the UE. The data packet for the target computing task carries the service type requested by the UE or the ID of the model that the UE needs to use. The computing node determines the first model, which can provide the service corresponding to the service type requested by the UE, or the ID of the first model is the ID of the model that the UE needs to use. The computing node uses the first model to process the data in the data packet of the target computing task, that is, to execute the UE's computing services.

[0380] It should be understood that the data packet of the target computation task carries the data corresponding to the target computation task. For example, if the target computation task is a text-to-image task, the data corresponding to the target computation task includes "generating an image containing a little girl sitting on the grass". The data packet of the uplink carries the data corresponding to the computation task, including "generating an image containing a little girl sitting on the grass" without tokenization, or it may include a tokenized "generating an image containing a little girl sitting on the grass".

[0381] When the data corresponding to the target computation task includes "generating an image containing a little girl sitting on the grass" without tokenization, the computing node tokenizes "generating an image containing a little girl sitting on the grass" to obtain a corresponding token, then inputs this token into the first model for processing, and obtains the token output by the first model. When the data corresponding to the target computation task includes a tokenized "generating an image containing a little girl sitting on the grass," the computing node directly inputs this token into the first model for processing, and obtains the token output by the first model. The token output by the first model is the result corresponding to the first model performing the text-to-image task. Optionally, the data packet sent to the UE carries the token output by the first model, or carries an image processed based on the token output by the first model.

[0382] S612. The computing node reports the UE's computing information to the management function network element.

[0383] Specifically, during the process of the computing node processing the data packets of the UE's target computing task, the computing node collects statistics on the UE's computing information. For a description of the UE's computing information, please refer to the relevant description in S302, which will not be repeated here.

[0384] Optionally, when the computing node reports the UE's computing information to the management function network element, it simultaneously sends the UE's ID and session ID to the management function network element to indicate that the reported computing information is used to process the target computing task of the UE. Optionally, the signaling transmitted between the computing node and the management function network element is session-granular. When the computing node reports the UE's computing information to the management function network element through this signaling, the charging function network element can determine the session corresponding to the reported computing information based on the signaling, and thus determine which UE the reported computing information corresponds to. In this case, the computing node does not need to report the UE's ID and session ID to the management function network element.

[0385] S613. The management function network element sends the UE's calculation information to the billing function network element.

[0386] S614. The billing function network element generates the UE's billing information based on the UE's calculation information.

[0387] Optionally, the management function network element may also send one or more of the following to the billing function network element: the UE ID, the first model identifier, the target computing task identifier, and the computing node identifier. The UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier are used by the billing function network element to determine the charging standard when the computing node processes the UE's target computing task.

[0388] It should be noted that the specific process by which the charging function network element generates the UE's charging information can be found in the relevant description of S402, and will not be repeated here. The beneficial effects of the embodiment corresponding to Figure 6 can be found in the relevant descriptions of the embodiments corresponding to Figures 2-5, and will not be repeated here.

[0389] Referring to Figure 7, Figure 7 is an interactive flowchart illustrating a communication method provided in an embodiment of this application. This method is applied to the system shown in Figure 1. As shown in Figure 7, the method includes:

[0390] S700 and UE complete the attachment process and the computing session creation process shown in Figure 6.

[0391] It should be noted that the implementation process of S700 can be found in the relevant description of the embodiment corresponding to Figure 6, and will not be described again here.

[0392] S701, the UE sends a computation plane update request to the AMF network element.

[0393] Specifically, when a UE needs more computing resources from the network side, or needs to execute a specific model, or needs to request a specified computing service, the UE sends a computing plane update request to the AMF network element. The computing plane update request includes one or more of the following: the identifier of the model requested by the UE, the type of computing service requested by the UE, and the computing information requested by the UE.

[0394] It should be noted that the calculation information requested by the UE includes one or more of the following:

[0395] The number of tokens input to the model, the number of tokens output by the model, computational cost, processing time, memory used, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0396] S702, the AMF network element forwards the compute plane update request from the UE to the management function network element.

[0397] S703, the management function network element sends a CM policy negotiation update request to the PCF network element.

[0398] The CM policy negotiation update request includes one or more of the following: the identifier of the model requested by the UE, the type of computing service requested by the UE, and the computing information requested by the UE.

[0399] S704, the PCF network element obtains the UE's subscription data from the UDM network element.

[0400] The UE's computational subscription data includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computational task type, and the UE's subscribed computational load. The UE's subscribed model includes a default model. The PCF network element generates a first rule based on the UE's computational subscription data. In one example, the first rule is a PCC rule.

[0401] S705 and PCF network elements send CM policy negotiation update responses to management function network elements.

[0402] The CM policy negotiation creation response is used to respond to the CM policy negotiation update request. Specifically, after the PCF network element obtains the UE's computational subscription data, it determines whether the UE can use the requested information based on the UE's computational subscription data. If it is determined that the UE can use the requested information, the CM policy negotiation update response includes the first rule; if it is determined that the UE cannot use the requested information, the CM policy negotiation update response does not include the first rule and stops the execution of subsequent processes.

[0403] S706, the management function network element selects the first model or computing node for the UE based on the first rule or local configuration.

[0404] It should be noted that the first model or computing node selected again can satisfy the UE's request, including the computing resources requested by the UE, the execution of a specific model requested by the UE, or the computing service specified by the UE.

[0405] It should be noted that the UE's computational subscription data includes the models that the UE is allowed to use. When the identifier of the model requested by the UE is not carried in the computational plane update request, the management function network element will select a default model as the first model for the UE. The newly selected first model can be the same as or different from the currently used model.

[0406] When the identifier of the model requested by the UE is carried in the computing plane update request, if the model indicated by the identifier of the model requested by the UE is included in the UE's computing subscription data and is among the models allowed to be used by the UE, then the first model selected by the management function network element is the model indicated by the identifier of the model requested by the UE; if the model indicated by the identifier of the model requested by the UE carried in the computing plane update request is not included in the models allowed to be used by the UE, then the management function network element selects a model that the UE can use as the first model. The newly selected first model may be the same as or different from the currently used model.

[0407] Similarly, the UE's computing subscription data includes the types of computing tasks that the UE is allowed to use. When the computing service type requested by the UE is not carried in the computing plane update request, the management function network element will select a default computing node as the new computing node for the UE, or select a default model as the first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0408] When the computing plane update request includes the type of computing service requested by the UE, if the requested computing service is included in the UE's computing subscription data and includes computing services permitted by the UE, then the computing node or first model selected by the management function network element can process the computing service requested by the UE. If the computing service requested by the UE in the computing plane update request is not included in the computing services permitted by the UE, then the management function network element selects either a computing node capable of processing computing services permitted by the UE as a new computing node, or selects a first model capable of processing computing services permitted by the UE as a new first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0409] Similarly, the UE's computing subscription data includes the computing load or computing resources that the UE is allowed to use. When the computing plane update request does not carry the computing resources or computing load requested by the UE, the management function network element will select a default computing node as the new computing node for the UE, or select a default model as the first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0410] When a computing load or computing resources requested by the UE are included in the computing plane update request, if the requested computing load or computing resources are included in the UE's computing subscription data and are permitted by the UE, then the computing node or first model selected by the management function network element can provide the requested computing load or computing resources. If the requested computing load or computing resources are not included in the permitted computing load or computing resources, then the management function network element selects either a computing node capable of providing the UE's computing load or computing resources as a new computing node, or selects a first model capable of providing the UE's requested computing load or computing resources as a new first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0411] S707, the management function network element sends third information to the UPF network element.

[0412] S708, the management function network element sends the fourth information to the RAN equipment.

[0413] S709, the management function network element sends the fifth information to the UE.

[0414] S710, the management function network element sends the first information to the computing node.

[0415] S711. The computing node receives data packets from the target computing task of the UE, processes them, and then sends downlink data packets to the UE.

[0416] S712. The computing node reports the UE's computing information to the management function network element.

[0417] S713, The management function network element sends the UE's calculation information to the charging function network element.

[0418] S714. The billing function network element generates the UE's billing information based on the UE's calculation information.

[0419] It should be noted that the specific implementation process of S706-S714 can be found in the relevant descriptions of S606-S614, and will not be repeated here. The beneficial effects of the embodiment corresponding to Figure 7 can be found in the relevant descriptions of the embodiments corresponding to Figures 2-5, and will not be repeated here.

[0420] Referring to Figure 8, Figure 8 is an interactive flowchart illustrating a communication method provided in an embodiment of this application. This method is applied to the system shown in Figure 1. As shown in Figure 8, the method includes:

[0421] S800 and UE complete the attachment process and the computing session creation process shown in Figure 6.

[0422] It should be noted that the implementation process of S800 can be found in the relevant description of the embodiment corresponding to Figure 6, and will not be described again here.

[0423] S801 and AF network elements send computing plane connection requests or computing plane update requests to PCF network elements.

[0424] When network computing resources are made available to external users, the AF network element can request a specific model, a designated computing service, or computing resources for the UE. In this request, the AF network element sends a computing plane connection request or a computing plane update request to the PCF network element. This request (computing plane connection request or computing plane update request) includes one or more of the following: the identifier of the model requested by the UE, the type of computing service requested by the UE, and the computing information requested by the UE.

[0425] It should be noted that the calculated information requested for the UE includes one or more of the following:

[0426] The number of tokens input to the model, the number of tokens output by the model, computational cost, processing time, memory used, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0427] S802, the PCF network element obtains the UE's subscription data from the UDM network element.

[0428] The UE's computational subscription data includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computational task type, and the UE's subscribed computational load. The UE's subscribed model includes a default model. The PCF network element generates a first rule based on the UE's computational subscription data. In one example, the first rule is a PCC rule.

[0429] S803, PCF network element sends the first rule to management function network element.

[0430] Specifically, after the PCF network element obtains the UE's computational subscription data, it determines whether the UE can use the AF network element to obtain the information requested by the UE based on the UE's computational subscription data. If it is determined that the UE can use the AF network element to obtain the information requested by the UE, the PCF network element sends the first rule to the management function network element; if it is determined that the UE cannot use the AF network element to obtain the information requested by the UE, the subsequent process will not be executed.

[0431] S804, the management function network element selects the first model or computing node for the UE based on the first rule or local configuration.

[0432] It should be noted that the first model or computing node selected again can meet the UE's request, including the computing resources requested by the UE, the specific model that the UE needs to execute, or the specific computing service that the UE needs to request.

[0433] It should be noted that the UE's computational subscription data includes the models that the UE is allowed to use. When the AF network element does not carry the identifier of the model requested by the UE in the computational plane connection request or computational plane update request, the management function network element will select a default model as the first model for the UE. The newly selected first model can be the same as or different from the currently used model.

[0434] When the AF element carries the identifier of the model requested by the UE in the computing plane connection request or computing plane update request, if the model indicated by the identifier of the model requested by the UE is included in the UE's computing subscription data and is among the models allowed by the UE, then the first model selected by the management function network element is the model indicated by the identifier of the model requested by the AF element; if the model indicated by the identifier of the model requested by the UE carried in the computing plane connection request or computing plane update request is not included in the models allowed by the UE, then the management function network element selects a model that the UE can use as the first model. The newly selected first model can be the same as or different from the currently used model.

[0435] Similarly, the UE's computing subscription data includes the types of computing tasks that the UE is allowed to use. When the computing plane connection request or computing plane update request does not carry the AF network element as the computing service type requested by the UE, the management function network element will select a default computing node as the new computing node for the UE, or select a default model as the first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0436] When the computing plane connection request or computing plane update request carries the computing service type requested by the AF network element for the UE, if the computing service requested by the AF network element for the UE is included in the UE's computing subscription data and includes computing services allowed by the UE, then the computing node or first model selected by the management function network element can process the computing service requested by the AF network element for the UE. If the computing service requested by the AF network element for the UE carried in the computing plane connection request or computing plane update request is not included in the computing services allowed by the UE, then the management function network element selects a computing node capable of processing computing services that the UE can use as a new computing node, or selects a first model capable of processing computing services that the UE can use as a new first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0437] Similarly, the UE's computing subscription data includes the computing load or computing resources that the UE is allowed to use. When the computing plane connection request or computing plane update request does not carry the computing resources or computing load requested by the AF network element for the UE, the management function network element will select a default computing node as the new computing node for the UE, or select a default model as the first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0438] When the computing plane connection request or computing plane update request carries the computing load or computing resources requested by the AF network element for the UE, if the computing load or computing resources requested by the AF network element for the UE are included in the UE's computing subscription data and include the computing load or computing resources allowed by the UE, then the computing node or first model selected by the management function network element can provide the computing load or computing resources requested by the AF network element for the UE. If the computing load or computing resources requested by the AF network element carried in the computing plane connection request or computing plane update request are not included in the computing load or computing resources requested by the UE from the AF network elements allowed by the UE, then the management function network element selects a computing node that can provide the computing load or computing resources that the UE can use as a new computing node, or selects a first model that can provide the computing load or computing resources requested by the AF network element for the UE as a new first model. The newly selected first model can be the same as or different from the currently used model, and the newly selected computing node can be the same as or different from the currently used computing node.

[0439] S805, the management function network element sends third information to the UPF network element.

[0440] S806, the management function network element sends the fourth information to the RAN equipment.

[0441] S807, the management function network element sends the fifth information to the UE.

[0442] S808, the management function network element sends the first information to the computing node.

[0443] S809. The computing node receives the data packet from the target computing task of the UE, processes it, and then sends the downlink data packet to the UE.

[0444] S810, the computing node reports the UE's computing information to the management function network element.

[0445] S811, The management function network element sends the UE's calculation information to the billing function network element.

[0446] S812, the billing function network element generates the UE's billing information based on the UE's calculation information.

[0447] It should be noted that the specific implementation process of S804-S812 can be found in the relevant descriptions of S606-S614, and will not be repeated here. The beneficial effects of the embodiment corresponding to Figure 7 can be found in the relevant descriptions of the embodiments corresponding to Figures 2-5, and will not be repeated here.

[0448] Referring to Figure 9, which is a structural schematic diagram of a management function network element provided in an embodiment of this application, the management function network element 900 includes:

[0449] The sending unit 901 is used to send first information to the computing node. The first information includes the UE's identifier and computing information detection indication information. The computing information detection indication information is used to instruct the computing node to count the UE's computing information. The UE's computing information is the computing information of the computing node executing the UE's target computing task.

[0450] The receiving unit 902 is used to receive computing information from the UE of the computing node;

[0451] The sending unit 901 is also used to send the UE's calculation information to the charging function network element, and the UE's calculation information is used by the charging function network element to generate the UE's charging information.

[0452] The UE's computation information may be information about the computing resources used by the computing node and / or the first model on the computing node when processing the UE's target computation task.

[0453] In one feasible implementation, the computational information indicates one or more of the following:

[0454] The number of the first tokens input when a compute node processes a target compute task;

[0455] The number of second tokens output by the computing node when processing the target computing task;

[0456] The amount of computation used by a computing node when processing a target computing task;

[0457] The time taken by a computing node to process the target computing task;

[0458] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0459] In one feasible implementation, the indication information for computational information detection includes the UE's identification information and / or the UE's IP address and / or the identification information of the first model.

[0460] In one feasible implementation, the sending unit 901 is further configured to send second information to the charging function network element, the second information including one or more of the following: the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier.

[0461] In one feasible implementation, the first information also includes reporting rules, which instruct the computing node on how to report the UE's computing information to the management function network element. The reporting methods for the UE's computing information include one or more of the following: periodic reporting, reporting based on a computing information threshold, and reporting based on the number of computing tasks.

[0462] Among them, the reporting based on the number of computing tasks is used to instruct the computing node to report the UE's computing information to the management function network element after completing a single or multiple target computing tasks.

[0463] In one feasible implementation, the sending unit 901 is further configured to send third information to the UPF network element, the third information being used by the UPF network element to forward the data packet corresponding to the target computing task from the UE to the computing node.

[0464] In one feasible implementation, the third information includes one or more of the following: tunnel information between the UPF network element and the computing node, data packet detection rules, and corresponding forwarding actions; the data packet detection rules are used to detect data packets of the target computing task, and the tunnel information or forwarding actions are used to forward the data packets of the UE's target computing task to the computing node.

[0465] In one feasible implementation, the packet detection rules include the IP address of the computing node or the identification information of the first model. Specifically, the destination IP address of the data packets for the UE's target computing task is either the IP address of the computing node or the IP address of the first model.

[0466] In one feasible implementation, the management function network element 900 includes:

[0467] The determining unit 903 is used to determine a computing node based on one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the UE's location information. The computing node satisfies one or more of the following conditions:

[0468] The model is deployed as indicated by the model identifier;

[0469] The computing task that can execute the target computing task is the computing task type subscribed to by the UE;

[0470] The computational load is not lower than the computational load subscribed by the UE;

[0471] The UE's location information is within the service area of ​​the computing node.

[0472] In one feasible implementation, the receiving unit 902 is further configured to receive a first rule from the PCF network element. The first rule includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the UE's indication information on allowing the use of computing services.

[0473] In one feasible implementation, the receiving unit 902 is further configured to receive a computing plane connection request or a computing plane update request from the UE. In terms of receiving the first PCC rule from the PCF network element, the receiving unit 902 is specifically configured to: receive the first PCC rule from the PCF network element based on the computing plane connection request or the computing plane update request.

[0474] or,

[0475] The first PCC rule is sent by the PCF network element to the management function network element based on the computing plane connection request or computing plane update request from the AF network element.

[0476] It is worth noting that the specific functional implementation of the management function network element 900 is described in the specific description of the embodiment shown in Figure 2. Each unit or module in the management function network element 900 can be individually or entirely merged into one or more other units or modules, or some of the units or modules can be further divided into multiple functionally smaller units or modules. This achieves the same operation without affecting the technical effects of the embodiments of this application. The aforementioned units or modules are based on logical functional division. In practical applications, the function of one unit (or module) is implemented by multiple units (or modules), or the function of multiple units (or modules) is implemented by one unit (or module).

[0477] Referring to Figure 10, which is a schematic diagram of a computing node provided in an embodiment of this application, the computing node 1000 includes:

[0478] The receiving unit 1001 is configured to receive first information from the management function network element, the first information including the UE's identifier and detection indication information of the calculation information;

[0479] The determining unit 1002 is used to determine the UE's computing information based on the UE's identifier and the indication information detected by the computing information; the UE's computing information is the computing information of the computing node for processing the UE's target computing task.

[0480] The sending unit 1003 is used to report the UE's calculation information to the management function network element. The UE's calculation information is used to generate the UE's billing information.

[0481] The UE's computation information can be information about the computing resources used by the computing node and / or the first model on the computing node to process the UE's target computation task.

[0482] In one feasible implementation, the computational information indicates one or more of the following:

[0483] The number of the first tokens input when a compute node processes a target compute task;

[0484] The number of second tokens output by the computing node when processing the target computing task;

[0485] The amount of computation used by a computing node when processing a target computing task;

[0486] The time taken by a computing node to process the target computing task;

[0487] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0488] In one feasible implementation, the indication information for computational information detection includes the UE's identification information and / or the UE's IP address and / or the identification information of the first model, and the determining unit 1002 is specifically used for:

[0489] The data packets of the target computing task are determined, wherein the source IP address of the data packets of the target computing task is the IP address of the UE and / or the data packets of the target computing task carry the identification information of the first model or the data packets of the target computing task carry the identification information of the UE; the data packets of the target computing task are obtained to determine the number of first tokens input into the first model and / or the number of second tokens output by the first model; and / or the computing load used by the computing node is determined based on the number of first tokens and / or the number of second tokens.

[0490] In one feasible implementation, the determining unit 1002 is further used for:

[0491] The first model is determined based on the semantic information of the data packet of the target computing task; or, the first model is determined based on the ID of the first model in the data packet of the target computing task or the type of the target computing task; or the first model is determined based on the IP address of the first model in the data packet or the port number assigned to the first model; the first token is input into the first model for processing to obtain the second token.

[0492] In one feasible implementation, the first information also includes reporting rules, which instruct the computing node on how to report the UE's computing information to the management function network element.

[0493] It is worth noting that the specific functional implementation of computing node 1000 is described in the specific description of the embodiment shown in Figure 3. Each unit or module in computing node 1000 can be individually or entirely merged into one or more other units or modules, or some of the units or modules can be further divided into multiple functionally smaller units or modules. This achieves the same operation without affecting the technical effects of the embodiments of this application. The aforementioned units or modules are based on logical functional division. In practical applications, the function of one unit (or module) is implemented by multiple units (or modules), or the function of multiple units (or modules) is implemented by one unit (or module).

[0494] Referring to Figure 11, which is a structural schematic diagram of a billing function network element provided in an embodiment of this application, as shown in Figure 9, the billing function network element 1100 includes:

[0495] The receiving unit 1101 is used to receive computing information from the UE of the management function network element; the computing information of the UE is the computing information of the computing node executing the target computing task of the UE;

[0496] The generation unit 1102 is used to generate the UE's billing information based on the UE's calculation information.

[0497] In one feasible implementation, the computational information indicates one or more of the following:

[0498] The number of the first tokens input when a compute node processes a target compute task;

[0499] The number of second tokens output by the computing node when processing the target computing task;

[0500] The amount of computation used by a computing node when processing a target computing task;

[0501] The time taken by a computing node to process the target computing task;

[0502] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0503] In one feasible implementation, the billing function network element 1100 also includes:

[0504] The sending unit 1103 is used to send a policy counter status notification or a spending limit notification to the PCF network element based on the UE's calculation information or billing information. The notification is used by the PCF network element to determine the service policy for the UE.

[0505] In one feasible implementation, the receiving unit 1101 is further configured to receive second information from the management function network element, the second information including one or more of the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier.

[0506] It is worth noting that the specific functional implementation of the billing function network element 1100 is described in the specific description of the embodiment shown in Figure 4. Each unit or module in the billing function network element 1100 can be individually or entirely merged into one or more other units or modules, or some of the units or modules can be further divided into multiple functionally smaller units or modules. This achieves the same operation without affecting the technical effect of the embodiments of this application. The above-mentioned units or modules are based on logical functional division. In practical applications, the function of one unit (or module) is implemented by multiple units (or modules), or the function of multiple units (or modules) is implemented by one unit (or module).

[0507] Referring to Figure 12, which is a schematic diagram of the structure of a UPF network element provided in an embodiment of this application, the UPF network element 1200 includes:

[0508] The receiving unit 1201 is used to receive third information from the management function network element; the third information is used to instruct the UPF network element to forward the data packet of the target computing task from the UE to the computing node; and to receive data packets from the user equipment UE.

[0509] When the determining unit 1202 determines that the data packet from the UE is the data packet for the target computing task based on the third information, the sending unit 1203 forwards the data packet from the UE to the computing node based on the third information.

[0510] In one feasible implementation, the third information includes tunnel information between the UPF network element and the computing node, or the IP address, packet detection rules, and forwarding actions of the computing node. The determining unit 1202 is also used for the UPF network element to determine whether the data packet from the UE is a data packet of the target computing task based on the packet detection rules.

[0511] Transmitting unit 1203 is specifically used for:

[0512] Data packets from the UE are forwarded to the computing node based on the tunnel information between the UPF network element and the computing node, the IP address of the computing node, and / or the forwarding action.

[0513] It is worth noting that the specific functional implementation of the UPF network element 1200 is described in the specific description of the embodiment shown in Figure 5. Each unit or module in the UPF network element 1200 can be individually or entirely merged into one or more other units or modules, or some of the units or modules can be further divided into multiple functionally smaller units or modules. This achieves the same operation without affecting the technical effects of the embodiments of this application. The aforementioned units or modules are based on logical functional division. In practical applications, the function of one unit (or module) is implemented by multiple units (or modules), or the function of multiple units (or modules) is implemented by one unit (or module).

[0514] This application provides a UE, including:

[0515] The sending unit is used to send computing plane connection requests or computing plane update requests to the management function network elements;

[0516] Specifically, the computing plane connection request or computing plane update request is used by the PCF network element to send a first rule to the management function network element. The first rule includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use. The UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use are used to determine the computing node. This computing node is used to process the UE's target computing task, collect the UE's computing information, and report the UE's computing information to the management function network element. The UE's computing information is used by the charging function network element to generate the UE's charging information.

[0517] The UE's computation information can be information about the computing resources used by the computing node and / or the first model on the computing node to process the UE's target computation task.

[0518] In one feasible implementation, the computational information indicates one or more of the following:

[0519] The number of the first tokens input when a compute node processes a target compute task;

[0520] The number of second tokens output by the computing node when processing the target computing task;

[0521] The amount of computation used by a computing node when processing a target computing task;

[0522] The time taken by a computing node to process the target computing task;

[0523] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0524] It should be understood that the UE includes not only a transmitting unit, but may also include a receiving unit and a unit for data processing.

[0525] This application provides an AF network element, including:

[0526] The sending unit is used to send computing plane connection requests or computing plane update requests to PCF network elements;

[0527] Specifically, the computing plane connection request or computing plane update request is used by the PCF network element to send a first rule to the management function network element. The first rule includes one or more of the following: the UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use. The UE's subscribed model identifier, the UE's subscribed computing task type, the UE's subscribed computing load, and the indication information of the computing services that the UE is allowed to use are used to determine the computing node. This computing node is used to process the UE's target computing task, collect the UE's computing information, and report the UE's computing information to the management function network element. The UE's computing information is used by the charging function network element to generate the UE's charging information.

[0528] The UE's computation information refers to the information on the computing resources used to process the UE's target computation task for the computing node and / or the first model on the computing node.

[0529] In one feasible implementation, the computational information indicates one or more of the following:

[0530] The number of the first tokens input when a compute node processes a target compute task;

[0531] The number of second tokens output by the computing node when processing the target computing task;

[0532] The amount of computation used by a computing node when processing a target computing task;

[0533] The time taken by a computing node to process the target computing task;

[0534] The computing node uses any one or more of the following when processing the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

[0535] It should be understood that an AF network element includes not only a transmitting unit, but also a receiving unit and a unit for data processing.

[0536] Based on the description of the above method embodiments and related device embodiments, please refer to FIG13, which provides a schematic diagram of the structure of a communication device 1300. The communication device 1300 shown in FIG13 includes a memory 1301, a processor 1302, a communication interface 1303, and a bus 1304. The memory 1301, the processor 1302, and the communication interface 1303 are interconnected through the bus 1304.

[0537] Optionally, the memory 1301 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM).

[0538] The memory 1301 is capable of storing programs. When the program stored in the memory 1301 is executed by the processor 1302, the processor 1302 and the communication interface 1303 are used to execute the various steps of the communication method of the embodiments shown in FIG2-FIG8.

[0539] The processor 1302 employs a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits to execute relevant programs to implement the functions required by the units in the management function network element 900, computing node 1000, billing function network element 1100, UPF network element 1200, UE, or AF network element in this application embodiment, or to execute the communication methods of the embodiments shown in Figures 2-8 of this application.

[0540] Processor 1302 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the communication method shown in Figures 2-8 of this application can be completed through integrated logic circuits in the hardware of processor 1302 or instructions in software form. Optionally, processor 1302 can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Processor 1302 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor is a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. Optionally, the software modules are located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory 1301. The processor 1302 reads the information in the memory 1301 and, in conjunction with its hardware, performs the functions required by the units included in the management function network element 900, computing node 1000, billing function network element 1100, UPF network element 1200, UE or AF network element in this application embodiment, or executes the communication method of the embodiments shown in Figures 2-8.

[0541] The communication interface 1303 uses transceiver devices, such as, but not limited to, transceivers, to enable communication between the communication device 1300 and other devices or communication networks.

[0542] Bus 1304 may include a pathway for transmitting information between various components of communication device 1300 (e.g., memory 1301, processor 1302, communication interface 1303).

[0543] It should be noted that although the communication device 1300 shown in Figure 13 only illustrates the memory, processor, and communication interface, those skilled in the art should understand that in specific implementations, the communication device 1300 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the communication device 1300 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the communication device 1300 may only include the devices necessary for implementing the embodiments of this application, and not necessarily all the devices shown in Figure 13.

[0544] This application also provides a chip, which includes a processor and a data interface. The processor reads instructions stored in a memory through the data interface to implement the communication method of this application.

[0545] Optionally, as one implementation, the chip further includes a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the communication method.

[0546] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of any of the above methods.

[0547] This application also provides a computer program product containing instructions. When the computer program product is run on a computer or processor, it causes the computer or processor to perform one or more steps of any of the methods described above.

[0548] Those skilled in the art will appreciate that the functionality described in conjunction with the various illustrative logic blocks, modules, and algorithmic steps disclosed herein can be implemented by hardware, software, firmware, or any combination thereof. If implemented in software, the functionality described by the various illustrative logic blocks, modules, and steps can be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may comprise a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium that includes any medium facilitating the transfer of a computer program from one place to another (e.g., based on a communication protocol). In this way, the computer-readable medium may substantially correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described herein. A computer program product may comprise a computer-readable medium.

[0549] By way of example and not limitation, such computer-readable storage media includes RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other media that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. Furthermore, any connection is properly referred to as computer-readable media. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that the computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but are specifically referring to non-temporary tangible storage media. As used herein, disks and optical discs include Compact Discs (CDs), Laser Discs, Optical Discs, Digital Versatile Discs (DVDs), and Blu-ray Discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.

[0550] Instructions can be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structures suitable for implementing the techniques described herein. Furthermore, in some aspects, the functions described in the various illustrative logic blocks, modules, and steps described herein are provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, the techniques can be fully implemented within one or more circuit or logic elements.

[0551] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Optionally, the coupling, direct coupling, or communication connection shown or discussed between them may be through some interfaces, indirect coupling or communication connection of devices or units, such as electrical, mechanical, or other forms.

[0552] Optionally, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0553] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of this application are generated.

[0554] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.

Claims

1. A communication method, characterized in that, The method is applied to a management function network element, and the method includes: Send first information to the computing node. The first information includes the identifier of the user equipment (UE) and the indication information for computing information detection. The indication information for computing information detection is used to instruct the computing node to collect the computing information of the UE. The computing information of the UE is the computing information of the computing node executing the target computing task of the UE. Receive computation information from the UE from the computing node; The calculation information of the UE is sent to the charging function network element, and the calculation information of the UE is used by the charging function network element to generate the charging information of the UE.

2. The method according to claim 1, characterized in that, The UE's computing information refers to the information on the computing resources used by the first model of the computing node when processing the UE's computing tasks.

3. The method according to claim 1 or 2, characterized in that... The calculated information indicates one or more of the following: The number of the first tokens input when the computing node processes the target computing task; The number of second tokens output by the computing node when processing the target computing task; The amount of computation used by the computing node when processing the target computing task; The time taken by the computing node to process the target computing task; The computing node uses any one or more of the following parameters to process the target computing task: memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and hardware computing power utilization.

4. The method according to any one of claims 1-3, characterized in that, The indication information for the computational information detection includes the IP address of the UE and / or the identification information of the first model.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Send second information to the billing function network element. The second information includes one or more of the following: the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier.

6. The method according to any one of claims 1-5, characterized in that, The first information also includes reporting rules, which indicate how the computing node reports the UE computing information to the management function network element.

7. The method according to claim 6, characterized in that, The methods for reporting the UE calculation information include one or more of the following: periodic reporting, reporting based on a calculation information threshold, and reporting based on the number of calculation tasks.

8. The method according to claim 7, characterized in that, The reporting based on the number of computing tasks is used to instruct the computing node to report the UE computing information to the management function network element when the target computing task is completed.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: The third information is sent to the user plane function network element, which is used by the user plane function network element to forward the data packet corresponding to the target computing task from the UE to the computing node.

10. The method according to claim 9, characterized in that, The third information includes one or more of the following: tunnel information between the user plane function network element and the computing node, packet detection rules, and corresponding forwarding actions; the packet detection rules are used to detect the packets of the target computing task, and the tunnel information or forwarding actions are used to forward the packets of the target computing task to the computing node.

11. The method according to claim 10, characterized in that, The packet detection rules include the IP address of the computing node or the identification information of the first model.

12. The method according to any one of claims 1-11, characterized in that, The method further includes: The computing node is determined based on one or more of the following: the model identifier subscribed by the UE, the computing task type subscribed by the UE, the computing load subscribed by the UE, and the location information of the UE. The computing node satisfies one or more of the following conditions: The model indicated by the aforementioned model identifier is deployed; It is capable of executing the computing task corresponding to the computing task type subscribed to by the UE for the target computing task; The computational load is not lower than the computational load subscribed by the UE; The location information of the UE is within the service area of ​​the computing node.

13. The method according to claim 12, characterized in that, The method further includes: The system receives a first rule from the policy control function (PCF) network element. The first rule includes one or more of the following: the model identifier subscribed by the UE, the computing task type subscribed by the UE, the computing load subscribed by the UE, and indication information indicating that the UE is allowed to use computing services.

14. The method according to claim 13, characterized in that, The method further includes: Receive a computing plane connection request or a computing plane update request from the UE; receiving a first rule from the Policy Control Function (PCF) network element includes: receiving a first rule from the PCF network element based on the computing plane connection request or computing plane update request. or, The first rule is that the PCF network element sends a computing plane connection request or computing plane update request to the management function network element based on the application function AF network element.

15. A communication method, characterized in that, The method is applied to a computing node, and the method includes: Receive first information from the management function network element, the first information including the identifier of the user equipment (UE) and indication information for computational information detection; The computation information of the UE is determined based on the UE's identifier and the indication information detected by the computation information; the computation information of the UE is the computation information of the computing node for processing the target computation task of the UE; The calculation information of the UE is reported to the management function network element, and the calculation information of the UE is used to generate the billing information of the UE.

16. The method according to claim 15, characterized in that, The UE's computation information refers to the information on the computing resources used by the first model of the computing node to process the UE's target computing task.

17. The method according to claim 15 or 16, characterized in that, The UE's calculation information indicates one or more of the following: The number of first tokens input into the first model when processing the target computation task. The number of second tokens output by the first model when processing the target computation task; The amount of computation used by the computing node when processing the target computing task; The time taken by the computing node to process the target computing task; The computing node uses memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and / or hardware computing power utilization to process the target computing task.

18. The method according to claim 17, characterized in that, The indication information for the computational information detection includes the IP address of the UE and / or the identification information of the first model. Determining the computational information of the UE based on the UE's identification and the indication information for the computational information detection includes: The data packet of the target computing task is determined, wherein the source IP address of the data packet of the target computing task is the IP address of the UE and / or the data packet of the target computing task carries the identification information of the first model; Obtain the data packet of the target computing task to determine the number of first tokens input into the first model and / or the number of second tokens output by the first model; and / or determine the amount of computing power used by the computing node, which is determined based on the number of first tokens and / or the number of second tokens.

19. The method according to claim 18, characterized in that, The method further includes: The first model is determined based on the semantic information of the data packets of the target computing task; or, the first model is determined based on the identifier of the first model in the data packets of the target computing task or the type of the target computing task; or the first model is determined based on the IP address of the first model in the data packets or the port number assigned to the first model. The first token is input into the first model for processing to obtain the second token.

20. The method according to any one of claims 15-19, characterized in that, The first information also includes reporting rules, which instruct the computing node on how to report the UE's computing information to the management function network element.

21. A communication method, characterized in that, The method is applied to a billing function network element, and the method includes: Receive computing information from the user equipment (UE) of the management function network element; the computing information of the UE is the computing information of the computing node executing the target computing task of the UE; The billing information for the UE is generated based on the UE's calculation information.

22. The method according to claim 21, characterized in that, The UE's computation information refers to the information on the computing resources used by the first model of the computing node to process the UE's computational tasks.

23. The method according to claim 21 or 22, characterized in that The calculated information indicates one or more of the following: The number of the first tokens input when processing the target computation task. The number of second tokens output when processing the target computation task; The amount of computation used by the computing node when processing the target computing task; The time taken by the computing node to process the target computing task; The computing node uses memory, memory access bandwidth, interconnect bandwidth, inference runtime information, deployment resource information, model computing power utilization, and / or hardware computing power utilization to process the target computing task.

24. The method according to claims 21-23, characterized in that, The method further includes: Based on the UE's calculation information, the policy control function network element sends a policy indicator status notification or expenditure limit notification, which is used by the policy control function network element to determine the service policy for the UE.

25. The method according to any one of claims 21-24, characterized in that, The method further includes: The system receives second information from a management function network element, the second information including one or more of the following: the UE identifier, the first model identifier, the target computing task identifier, and the computing node identifier.

26. A management function network element, characterized in that, The management function network element includes a unit or module that implements any one of claims 1-14.

27. A computing node, characterized in that, The management function network element includes a unit or module that implements any one of claims 15-20.

28. A billing function network element, characterized in that, The management function network element includes a unit or module that implements any one of claims 21-27.

29. A communication device, characterized in that, The method includes a processor and a memory, wherein the memory is used to store program code, and the processor is used to execute the program code to implement the method according to any one of claims 1-25.

30. A communication method, characterized in that, The method is applied to a communication system, which includes management function network elements, computing nodes, charging function network elements, and user equipment (UE). The method includes: The management function network element sends first information to the computing node. The first information includes the UE's identifier and computing information detection indication information. The computing information detection indication information is used to instruct the computing node to collect the UE's computing information. The UE's computing information is the computing information of the computing node executing the UE's target computing task. The computing node determines the computing information of the UE based on the UE's identifier and the indication information of the computing information detection, and reports the UE's computing information to the management function network element; The management function network element sends the UE's calculation information to the charging function network element. The billing function network element generates the UE's billing information based on the UE's calculation information.

31. A communication system, characterized in that, The communication system includes management function network elements, computing nodes, billing function network elements, and user equipment (UE). The management function network element is used to send first information to the computing node. The first information includes the identifier of the UE and the indication information for computing information detection. The indication information for computing information detection is used to instruct the computing node to collect the computing information of the UE. The computing information of the UE is the computing information of the computing node executing the target computing task of the UE. The computing node is used to determine the computing information of the UE based on the UE's identifier and the indication information of the computing information detection, and to report the UE's computing information to the management function network element; The management function network element is also used to send the UE's calculation information to the billing function network element. The billing function network element is used to generate the billing information of the UE based on the UE's calculation information.

32. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-25.

33. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-25.