Ai model-based task allocation method in communication system, and related product
By introducing a task allocation method for AI models into the communication system and aligning device computing power with model IDs, the problem of underutilization of device computing power is solved, achieving high efficiency and low latency in distributed training.
Patent Information
- Application Number
- PCT/CN2025/102422
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-19
- Filing Date
- 2025-06-20
- Publication Date
- 2026-01-22
AI Technical Summary
Existing AI mechanisms cannot fully utilize the computing power of various devices in the field of communication, resulting in insufficient release and utilization of computing power, which affects the efficiency and cost of distributed training.
By introducing an AI model-based task allocation method into the communication system, and using model IDs to align computing power between devices, distributed training and task segmentation are achieved, reducing inference latency.
It achieves alignment of computing power between devices, improves the efficiency of task partitioning, and reduces inference latency and task decomposition instruction overhead.
Smart Images

Figure CN2025102422_22012026_PF_FP_ABST
Abstract
Description
AI-based task allocation methods and related products in communication systems
[0001] This application claims priority to Chinese Patent Application No. 202410982471.0, filed on July 19, 2024, with the China National Intellectual Property Administration, entitled “Task Allocation Method and Related Products Based on AI Model in Communication Systems”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the fields of artificial intelligence (AI) and communication technology, and in particular to a task allocation method based on an AI model in a communication system and related products. Background Technology
[0003] In recent years, large-scale AI models, exemplified by the chat generative pre-trained transformer (ChatGPT), have gained increasing popularity due to their outstanding natural language and multimodal understanding capabilities and their contribution to generative AI. However, the sheer number of parameters and training data required for ChatGPT-like models leads to a surge in computational demands, making distributed cluster training a common choice for large-scale models. Systematic analysis and targeted optimization of the network communication characteristics of distributed training for large-scale models are crucial for improving training speed, reducing training time, and lowering training costs.
[0004] On the other hand, the training and fine-tuning of AI models rely on real-time data. How to utilize distributed terminals to obtain high-quality real-time data is a key issue of concern in the AI field. In addition, considering the trend of large models being jointly deployed in the cloud, pipeline, and edge environments, the computing power, real-time status, and willingness to participate in training / inference of terminals can be coordinated in real time through wireless networks.
[0005] Existing AI mechanisms cannot fully leverage the computing power of various devices (such as terminals, base stations, and servers) in the field of communications, and the computing power of these devices cannot be fully released and utilized. Summary of the Invention
[0006] This application discloses a communication method, device, storage medium, and program product that can achieve alignment of computing power between devices, enable distributed training, and reduce the overall latency of inference.
[0007] Firstly, embodiments of this application provide a communication method. This method can be applied to a terminal side, such as a terminal or a communication / processing module within the terminal, or circuits or chips in the terminal responsible for communication functions (e.g., modem chips, also known as baseband chips, or system-on-chip (SoC) chips or system-in-package (SIP) chips containing modem cores), or circuits or chips in the terminal responsible for processing functions (e.g., graphics processing unit (GPU)). Taking the application of this method to a terminal as an example, in this method, the terminal sends first information indicating a first model identification (ID). The first model ID indicates an artificial intelligence (AI) model deployed on the terminal. The terminal also receives second information indicating a second model ID. The second model ID is the model ID within the first model ID, and the second model ID is associated with a first task. Furthermore, the terminal executes the first task based on the AI model corresponding to the second model ID.
[0008] Using the above method, the terminal indicates its deployed AI model through the first model ID, so that access network devices or core network devices are aware of its model capabilities. This enables the alignment of computing power among devices, fully utilizing the computing power of different devices and providing conditions for task partitioning (allocation). The terminal executes the first task based on the AI model indicated by the second model ID, enabling distributed training. This results in shorter inference latency for the enabled task; moreover, indicating the task based on the second model ID saves on task decomposition indication overhead.
[0009] In one possible design, the first model ID corresponds to one or more of the following of the AI model deployed on the first communication device: function, number of input nodes, number of output nodes, version release time, and version number. This example, by constructing a mapping rule that associates model IDs with model characteristics, enables terminals, access network devices, core network elements, etc., to directly determine model functions and corresponding functional characteristics. By using the model ID generated through the mapping rule, model capabilities on different devices are aligned with minimal overhead.
[0010] In one possible design, the first model ID includes a first field and a second field. The first field and the second field respectively indicate one or more of the following: the function, the number of input nodes, the number of output nodes, the version release date, and the version number. The first field and the second field indicate different content. In this example, the model ID is displayed based on fields with specific meanings, which can intuitively express the corresponding model characteristics. The model ID generated according to rules aligns with the model capabilities of different devices with minimal overhead.
[0011] In one possible design, the first information indicates a first checksum, which is generated based on the first model ID. This example uses a checksum to align computing power between devices, and its short bit overhead allows terminals, access network devices, core network elements, etc., to complete model capability alignment with minimal overhead.
[0012] In one possible design, the terminal also receives a fourth piece of information indicating a third model ID. This third model ID indicates the AI model deployed by the second communication device. This example allows the terminal to obtain the model capabilities of the access network device, enabling the terminal to compare the received model ID with its own model capabilities, thereby providing feedback on its own capabilities, or supplementing differentiated capabilities, enabling the two devices to align their model capabilities within a short latency.
[0013] In one possible design, the first model ID includes the model ID in the third model ID. The terminal reports the model ID in the third model ID, which allows the access network device to be aware of the model capabilities possessed by the terminal, enabling the two devices to align their model capabilities within a shorter timeframe and reducing the latency of task segmentation.
[0014] In another possible design, the first model ID includes other model IDs besides the third model ID. The terminal reports model IDs other than the third model ID, enabling the two devices to align models within a shorter latency, thus reducing the latency of task segmentation.
[0015] In another possible design, the first model ID includes the model ID in the third model ID, as well as other model IDs besides the third model ID. The terminal can report not only the model ID in the third model ID, but also the model IDs outside the third model ID, enabling devices to fully interact with their respective model capabilities. This allows two devices to align their model capabilities within a short time, reducing the latency of task partitioning.
[0016] In one possible implementation, the terminal also receives a sixth piece of information indicating the number of in-degrees and out-degrees of the directed acyclic graph (DAG). This allows the terminal to perceive the sub-task requirements, assisting access network devices in performing some computations. The terminal can accurately obtain information related to the sub-tasks to be executed, helping other devices complete task inference with less latency.
[0017] Secondly, embodiments of this application provide a communication method. This method can be applied to the core network side, such as core network elements or modules within core network elements (e.g., circuits, chips, or chip systems), or logical nodes, logical modules, or software capable of implementing all or part of the core network element functions. Taking the application of this method to a core network element as an example, in this method, the core network element sends first information, which indicates a first model ID. The first model ID indicates an AI model deployed by the core network element. The core network element also receives second information, which indicates a second model ID. The second model ID is a model ID within the first model ID, and the second model ID is associated with a first task. Furthermore, the core network element executes the first task based on the AI model corresponding to the second model ID.
[0018] Using the above method, core network elements indicate their deployed AI models through a first model ID, enabling access network devices to understand their model capabilities. This achieves alignment of computing power among devices and provides conditions for task partitioning (allocation). Core network elements execute the first task based on the AI model indicated by the second model ID, enabling distributed training. This results in shorter inference latency for the enabled task; moreover, indicating the task based on the second model ID saves on task decomposition and indication overhead.
[0019] Some possible implementations and beneficial effects of the second aspect can be found in the first aspect mentioned above, and will not be elaborated further.
[0020] In one possible design, the first task includes a first subtask and a second subtask. The core network element also receives third information, which includes a first data packet and a second data packet. The first data packet and the second data packet have the same service function chaining (SFC) header. The first data packet corresponds to the first subtask, and the second data packet corresponds to the second subtask. This example allows the first and second subtasks to receive the same transmission guarantees in the core network, providing conditions for achieving lower inference latency.
[0021] In another possible design, the first task includes a third and a fourth subtask. The core network element also sends a fifth message indicating a first profile and a second profile. Both the first and second profiles are configured with the same 5G Quality of Service (QoS) identifier (5QI) value. The first profile corresponds to the third subtask, and the second profile corresponds to the fourth subtask. This example ensures that the access network device provides the same QoS guarantee for the third and fourth subtasks during scheduling, allowing terminals to receive data packets as simultaneously as possible and achieving lower inference latency.
[0022] Thirdly, this method can be applied to the network side, such as access network devices, modules (e.g., circuits, chips, or chip systems) within the access network devices, or logical nodes, logical modules, or software that can implement all or part of the functions of the access network devices. Taking the application of this method to an access network device as an example, in this method, the access network device receives first information, which indicates a first model ID. The first model ID indicates an AI model deployed by a first communication device. The access network device also sends second information, which indicates a second model ID. The second model ID is the model ID within the first model ID, and the second model ID is associated with the first task.
[0023] The possible implementations and beneficial effects of the third aspect can be referenced from the first aspect mentioned above, and will not be elaborated further.
[0024] In one possible design, the first task includes a first subtask and a second subtask. The access network device also sends third information, which includes a first data packet and a second data packet. The first data packet and the second data packet have the same Service Function Chain (SFC) header. The first data packet corresponds to the first subtask, and the second data packet corresponds to the second subtask.
[0025] In another possible design, the first task includes a third subtask and a fourth subtask. The access network device also receives fifth information, which indicates a first configuration file and a second configuration file. Both the first and second configuration files have the same 5QI value. The first configuration file corresponds to the third subtask, and the second configuration file corresponds to the fourth subtask. This ensures that the access network device provides the same QoS guarantee for the third and fourth subtasks during scheduling, allowing terminals to receive data packets as simultaneously as possible and achieving lower inference latency.
[0026] In one possible design, the access network device also sends a fourth piece of information indicating a third model ID, which in turn indicates the AI model deployed by the access network device. This example allows the terminal to obtain the model capabilities of the access network device, enabling the terminal to compare the received model ID with its own model capabilities, thereby providing feedback on its own capabilities or supplementing differentiated capabilities, enabling the two devices to align their model capabilities within a short latency.
[0027] Fourthly, this application provides a communication device that has the functions of the first aspect described above. For example, the communication device includes modules, units, or means that perform the operations involved in the first aspect. These modules, units, or means can be implemented by software, hardware, or a combination of software and hardware.
[0028] In one implementation, the communication device includes a communication unit for: transmitting first information indicating a first model ID, the first model ID indicating an AI model deployed on the terminal.
[0029] The communication unit is also configured to: receive second information indicating a second model ID, the second model ID being a model ID in the first model ID, and the second model ID being associated with the first task.
[0030] The processing unit is used to: execute the first task based on the AI model corresponding to the second model ID.
[0031] The possible implementations and beneficial effects of the fourth aspect can be referenced in the first aspect above, and will not be elaborated further.
[0032] Fifthly, this application provides a communication device that has the functions of the second aspect above. For example, the communication device includes modules, units, or means that perform the operations involved in the second aspect above. These modules, units, or means can be implemented by software, hardware, or a combination of software and hardware.
[0033] In one implementation, the communication device includes a communication unit for: transmitting first information indicating a first model ID, the first model ID indicating an AI model deployed in a core network element.
[0034] The communication unit is also configured to: receive second information indicating a second model ID, the second model ID being a model ID in the first model ID, and the second model ID being associated with the first task.
[0035] The processing unit is used to: execute the first task based on the AI model corresponding to the second model ID.
[0036] The possible implementations and beneficial effects of the fifth aspect can be referred to the first aspect above, and will not be elaborated further.
[0037] Sixthly, this application provides a communication device that has the functions of the third aspect above. For example, the communication device includes modules, units, or means that perform the operations involved in the third aspect above. These modules, units, or means can be implemented by software, hardware, or a combination of software and hardware.
[0038] In one implementation, the communication device includes a communication unit for: receiving first information indicating a first model ID. The first model ID indicates an AI model deployed by the first communication device.
[0039] The communication unit is further configured to: send second information indicating a second model ID. The second model ID is a model ID within the first model ID, and the second model ID is associated with the first task.
[0040] The possible implementations and beneficial effects of the sixth aspect can be referred to the first aspect above, and will not be elaborated further.
[0041] In a seventh aspect, this application provides a communication device including an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory stores part or all of the necessary computer program or instructions for implementing the functions described in the first aspect. The one or more processors are executable to carry out the computer program or instructions, causing the communication device to implement the methods in any possible design or implementation of the first aspect. The interface circuit is used to implement the communication functions within the communication device and / or the communication functions between the communication device and other devices or components.
[0042] In one possible design, the processor is used to communicate with other devices or components through the interface circuit.
[0043] In one possible design, the communication device may also include the memory.
[0044] The aforementioned communication device may be a terminal, or a communication / processing module in the terminal, or a chip in the terminal responsible for communication functions such as a modem chip (also known as a baseband chip) or a SoC or SIP chip containing a modem module, or a circuit or chip in the terminal responsible for processing functions (such as a GPU).
[0045] Eighthly, this application provides a communication device including an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory stores part or all of the necessary computer program or instructions for implementing the functions described in the second aspect above. The one or more processors are executable to carry out the computer program or instructions, causing the communication device to implement the methods in any possible design or implementation of the second aspect above. The interface circuit is used to implement the communication functions within the communication device and / or the communication functions between the communication device and other devices or components.
[0046] Ninthly, this application provides a communication device including an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory stores part or all of the computer program or instructions necessary to implement the functions described in the third aspect above. The one or more processors are executable to carry out the computer program or instructions, causing the communication device to implement the methods in any possible design or implementation of the third aspect above. The interface circuit is used to implement the communication functions within the communication device and / or the communication functions between the communication device and other devices or components.
[0047] In a tenth aspect, this application provides a communication system comprising the communication apparatus described in the fourth aspect, and further comprising the communication apparatus described in the sixth aspect. Optionally, the communication system may also include the communication apparatus described in the fifth aspect.
[0048] Alternatively, the communication system may include the communication device described in the fifth aspect, and also the communication device described in the sixth aspect. Optionally, the communication system may also include the communication device described in the fourth aspect.
[0049] Alternatively, the communication system may include the communication device described in the fourth aspect, and also the communication device described in the fifth aspect. Optionally, the communication system may also include the communication device described in the sixth aspect.
[0050] Eleventhly, this application provides a computer-readable storage medium storing computer-readable instructions, which, when read and executed by a computer, cause the computer to perform any of the possible designs in the first to third aspects described above.
[0051] In a twelfth aspect, this application provides a computer program product that, when read and executed by a computer, causes the computer to perform any one of the first to third aspects described above, or any possible design method of the first to third aspects.
[0052] It is understood that the apparatus described in the fourth, fifth, sixth, seventh, eighth, and ninth aspects, the system described in the tenth aspect, the computer storage medium described in the eleventh aspect, or the computer program product described in the twelfth aspect are all used to perform the method provided in any of the first, second, or third aspects. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0053] The accompanying drawings used in the embodiments of this application are described below.
[0054] Figure 1 is a schematic diagram of a communication system provided in an embodiment of this application;
[0055] Figures 2-4 are schematic diagrams of possible application frameworks in the communication system provided in the embodiments of this application;
[0056] Figure 5 is a flowchart illustrating a communication method provided in an embodiment of this application;
[0057] Figure 6 is a schematic diagram of a DAG provided in an embodiment of this application;
[0058] Figure 7 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;
[0059] Figure 8 is a schematic diagram of the structure of a terminal provided in an embodiment of this application. Detailed Implementation
[0060] The embodiments of this application are described below with reference to the accompanying drawings.
[0061] The technology provided in this application can be applied to various communication systems, such as fourth-generation (4G) communication systems (e.g., Long Term Evolution (LTE) systems), fifth-generation (5G) communication systems, wireless local area network (WLAN) systems, satellite communication systems, integrated systems of multiple systems, or future communication systems. Among these, 5G communication systems can also be referred to as new radio (NR) systems.
[0062] In a communication system, a network element can send signals to or receive signals from another network element. These signals can include information, signaling, or data. The term "network element" can also be replaced by an entity, network entity, device, communication equipment, communication module, node, communication node, etc. This application uses a network element as an example for description. For instance, a communication system may include at least one terminal and at least one access network device. The access network device can send downlink signals to the terminal, and / or the terminal can send uplink signals to the access network device. Furthermore, it is understood that if the communication system includes multiple terminals, these terminals can also exchange signals; that is, both the signal-sending network element and the signal-receiving network element can be a terminal.
[0063] Figure 1 illustrates a possible, non-limiting system diagram. As shown in Figure 1, the communication system 10 includes a radio access network (RAN) 100 and a core network (CN) 200. RAN 100 includes at least one RAN node (110a and 110b in Figure 1, collectively referred to as 110) and at least one terminal (120a-120j in Figure 1, collectively referred to as 120). RAN 100 may also include other RAN nodes, such as wireless relay devices and / or wireless backhaul devices (not shown in Figure 1). Terminal 120 is wirelessly connected to RAN node 110. RAN node 110 is wirelessly or wired connected to core network 200. The core network equipment in core network 200 and RAN node 110 in RAN 100 can be different physical devices, or they can be the same physical device integrating core network logical functions and radio access network logical functions.
[0064] RAN 100 can be a cellular system related to the 3rd Generation Partnership Project (3GPP), such as 4G, 5G mobile communication systems, or future-oriented evolution systems. RAN 100 can also be an open RAN (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. RAN 100 can also be a communication system that integrates two or more of the above systems.
[0065] RAN node 110, sometimes also referred to as access network equipment, RAN entity, or access node, constitutes part of the communication system and is used to help terminals achieve wireless access. Multiple RAN nodes 110 in communication system 10 can be of the same type or different types. In some scenarios, the roles of RAN node 110 and terminal 120 are relative. For example, network element 120i in Figure 1 can be a helicopter or drone, which can be configured as a mobile base station. For terminals 120j accessing RAN 100 through network element 120i, network element 120i is a base station; but for base station 110a, network element 120i is a terminal. RAN node 110 and terminal 120 are sometimes both referred to as communication devices. For example, network elements 110a and 110b in Figure 1 can be understood as communication devices with base station functions, and network elements 120a-120j can be understood as communication devices with terminal functions.
[0066] In one possible scenario, the RAN node can be a base station, an evolved NodeB (eNodeB), an access point (AP), a transmission reception point (TRP), a next-generation NodeB (gNB), a base station in a future mobile communication system, or an access node in a WiFi system. The RAN node can be a macro base station (as shown in Figure 1, 110a), a micro base station or indoor station (as shown in Figure 1, 110b), a relay node or donor node, or a radio controller in a CRAN scenario. Optionally, the RAN node can also be a server, wearable device, vehicle, or in-vehicle equipment. For example, the access network equipment in vehicle-to-everything (V2X) technology can be a roadside unit (RSU). All or part of the functions of the RAN node in this application can also be implemented through software functions running on hardware, or through virtualization functions instantiated on a platform (e.g., a cloud platform). The RAN node can also be equipped with communication modules, circuits, or chips that perform corresponding communication functions. The RAN node can also be configured with program instructions for performing corresponding communication functions, as well as corresponding program instructions. The RAN node in this application can also be a logical node, logical module, or software capable of implementing all or part of the RAN node's functions.
[0067] In another possible scenario, multiple RAN nodes collaborate to assist the terminal in achieving wireless access, with each RAN node performing a portion of the base station's functions. For example, RAN nodes can be central units (CUs), distributed units (DUs), CU-control plane (CPs), CU-user plane (UPs), or radio units (RUs), etc. CUs and DUs can be separate entities or included in the same network element, such as a baseband unit (BBU). RUs can be included in radio frequency equipment or radio frequency units, such as remote radio units (RRUs), active antenna units (AAUs), or remote radio heads (RRHs).
[0068] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an ORAN system, CU can also be called O-CU (open CU), DU can also be called O-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CU-UP, and RU can also be called O-RU. For ease of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples. Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software and hardware modules.
[0069] A terminal can be a device or module that accesses the aforementioned communication system and has corresponding communication functions. A terminal can also be called a terminal device, user equipment (UE), mobile station, mobile terminal, etc. Terminals can be widely used in various scenarios, such as device-to-device (D2D), vehicle-to-everything (V2X) communication, machine-type communication (MTC), Internet of Things (IoT), virtual reality, augmented reality, industrial control, autonomous driving, telemedicine, smart grids, smart furniture, smart offices, smart wearables, smart transportation, smart cities, etc. Terminals can be mobile phones, tablets, computers with wireless transceiver capabilities, wearable devices, vehicles, drones, helicopters, airplanes, ships, robots, robotic arms, smart home devices, transportation vehicles with wireless communication capabilities, communication modules, etc. The embodiments of this application do not limit the device form of the terminal. A terminal typically contains a communication module, circuit, or chip that performs the corresponding communication function. The terminal can also be configured with program instructions for performing the corresponding communication function.
[0070] The core network 200 is used to complete three major functions: registration, connection, and session management. It mainly includes network exposure function (NEF) network elements, policy control function (PCF) network elements, application function (AF) network elements, access and mobility management function (AMF) network elements, session management function (SMF) network elements, and user plane function (UPF) network elements.
[0071] NEF network element: exposes the services and capabilities of 3GPP network functions to AF, and also allows AF to provide information to 3GPP network functions. The corresponding interface is N33 interface.
[0072] PCF network element: performs policy management for charging and QoS policies;
[0073] AF network elements primarily transmit the application side's requirements to the network side;
[0074] AMF (Automatic Facilitation Module) elements primarily perform mobility management, access authentication / authorization, and other functions. They are also responsible for transmitting user policies between the UE and PCF (Programmable Module Function). The N1 interface is the signaling plane interface between the UE and AMF. Since the UE cannot directly interact with the core network, it needs to pass through the access network (AN) to transmit non-access stratum (NAS) information. The N2 interface is the signaling plane interface through which the AMF requests resources from the AN to allocate for Protocol Data Unit (PDU) sessions.
[0075] SMF network element: Completes session management functions such as UE IP address allocation, UPF selection, and billing and QoS policy control;
[0076] UPF network elements: As the interface with the data network, they perform functions such as user plane data forwarding, session / flow-level billing statistics, and bandwidth limiting. The N3 interface is the interface between the RAN and UPF, mainly used to transmit uplink and downlink user plane data between the 5G RAN and UPF.
[0077] Optionally, the communication system 10 also includes an Internet 300. The Internet 300 is communicatively connected to the core network 200. And / or, the Internet 300 is communicatively connected to the RAN 100.
[0078] To support AI technology in wireless networks, AI nodes may also be introduced into the network.
[0079] AI nodes can be deployed in one or more of the following locations within the communication system: access network nodes (RAN nodes), terminal devices, or core network devices. Alternatively, AI nodes can be deployed independently, for example, in a location other than any of the aforementioned devices, such as in the host or cloud server of an over-the-top (OTT) system. AI nodes can communicate with other devices in the communication system, which can be one or more of the following: network devices, terminal devices, or core network elements.
[0080] It is understood that this application does not limit the number of AI nodes. For example, when there are multiple AI nodes, these nodes can be divided based on function, such as different AI nodes being responsible for different functions.
[0081] It can also be understood that AI nodes can be independent devices, or they can be integrated into the same device to achieve different functions. Alternatively, they can be network elements in hardware devices, software functions running on dedicated hardware, or virtualization functions instantiated on a platform (e.g., a cloud platform). This application does not limit the specific form of the aforementioned AI nodes.
[0082] AI nodes can be AI network elements or AI modules.
[0083] Figure 2 illustrates a possible application framework in a communication system. As shown in Figure 2, network elements in the communication system are connected via interfaces (e.g., NG, Xn) or air interfaces. These network element nodes, such as core network equipment, access network nodes (RAN nodes), terminals, or one or more devices in operation administration and maintenance (OAM), are equipped with one or more AI modules (only one is shown in Figure 2 for clarity). An access network node can be a single RAN node or can include multiple RAN nodes, for example, including CU and DU. The CU and / or DU can also be equipped with one or more AI modules. A CU can also be split into CU-CP and CU-UP, with one or more AI modules configured in the CU-CP and / or CU-UP.
[0084] AI modules are used to implement corresponding AI functions. AI modules deployed in different network elements can be the same or different. The models of AI modules can achieve different functions depending on the parameter configurations. The models of AI modules can be configured based on one or more of the following parameters: structural parameters (e.g., at least one of the following: number of neural network layers, neural network width, inter-layer connections, neuron weights, neuron activation function, or biases in the activation function), input parameters (e.g., the type and / or dimension of the input parameters), or output parameters (e.g., the type and / or dimension of the output parameters). The biases in the activation function can also be referred to as the biases of the neural network.
[0085] In one example, the neural network mentioned above can be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), or a generative adversarial network (GAN).
[0086] Deep Neural Networks (DNNs) are artificial neural network architectures with multiple layers of nonlinear transformation units stacked in a hierarchical structure to form deep computational models. Compared to shallow neural networks, deep neural networks have more hidden layers, allowing the network model to capture more complex data structures and higher-level abstract features.
[0087] A CNN is a deep neural network with a convolutional structure. A CNN contains a feature extractor consisting of convolutional layers and subsampling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as performing convolution between a trainable filter and an input image or a convolutional feature map.
[0088] RNN is a type of recursive neural network that takes sequence data as input, recursively moves along the direction of sequence evolution, and connects all nodes (recurrent units) in a chain-like manner.
[0089] GAN is a deep learning model. It consists of a generator and a discriminator, and is trained through adversarial learning. Its purpose is to estimate the potential distribution of data samples and generate new data samples.
[0090] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or they can be deployed on the same node or device.
[0091] Figure 3 illustrates another possible application framework in a communication system. As shown in Figure 3, the communication system includes a RAN intelligent controller (RIC). For example, the RIC can be an AI module used to implement AI-related functions. RICs include near-real-time RICs (near-RT RICs) and non-real-time RICs (non-RT RICs). Non-real-time RICs primarily process non-real-time information, such as data that is not sensitive to latency, with latency in the order of seconds. Real-time RICs primarily process near-real-time information, such as data that is relatively sensitive to latency, with latency in the order of tens of milliseconds.
[0092] Near real-time (NRT) RICs are used for model training and inference. For example, they are used to train AI models and then use those models for inference. NRT RICs can obtain network-side and / or terminal-side information from RAN nodes (e.g., CUs, CU-CPs, CU-UPs, DUs, and / or RUs) and / or terminals. This information can be used as training data or inference data. NRT RICs can deliver inference results to RAN nodes and / or terminals. Inference results can be exchanged between CUs and DUs, and / or between DUs and RUs. For example, a NRT RIC delivers an inference result to a DU, which then forwards it to an RU.
[0093] Non-real-time RICs are also used for model training and inference. For example, they can be used to train AI models and then use those models for inference. Non-real-time RICs can obtain network-side and / or terminal-side information from RAN nodes (e.g., CUs, CU-CPs, CU-UPs, DUs, and / or RUs) and / or terminals. This information can be used as training data or inference data, and the inference results can be delivered to RAN nodes and / or terminals. Inference results can be exchanged between CUs and DUs, and / or between DUs and RUs; for example, a non-real-time RIC delivers inference results to a DU, which then forwards them to an RU.
[0094] Near real-time RICs and non-real-time RICs can also be configured as separate network elements. Near real-time RICs and non-real-time RICs can also be part of other devices. For example, near real-time RICs can be set in RAN nodes (e.g., CU, DU), while non-real-time RICs can be set in OAM, cloud servers, core network devices, or other network devices.
[0095] Referring to Figure 4, another architectural schematic diagram of the communication system provided in this application embodiment is shown. The communication system may include a first communication device 401, a second communication device 402, and a third communication device 403.
[0096] In one possible implementation, the first communication device 401 may be a terminal. The second communication device 402 may be an access network device, such as a base station. The third communication device 403 may be a core network element, such as an AMF element, SMF element, UPF element, etc.
[0097] In another possible implementation, the first communication device 401 can be a core network element, the second communication device 402 can be an access network device, and the third communication device 403 can be a terminal.
[0098] In the example above, the terminal and core network elements communicate with the access network equipment, respectively.
[0099] In another possible implementation, the first communication device can be a terminal, the second communication device 402 can be a core network element, and the third communication device 403 can be an access network device.
[0100] Alternatively, the first communication device may be an access network device, the second communication device 402 may be a core network element, and the third communication device 403 may be a terminal.
[0101] In this example, the terminal and access network equipment communicate with the core network elements respectively.
[0102] Of course, the first communication device can also be an access network device, the second communication device 402 can be a terminal, and the third communication device 403 can be a core network element. Alternatively, the first communication device can be a core network element, the second communication device 402 can be a terminal, and the third communication device 403 can be an access network device. In other words, the access network device and the core network element communicate with the terminal respectively, and this solution does not impose any restrictions on this.
[0103] Because current AI mechanisms cannot fully utilize the computing power of the aforementioned devices in the communication field, the computing power of each device cannot be fully utilized. Therefore, this application provides a communication method that can align the computing power between devices, providing conditions for task partitioning (allocation), enabling distributed training, and reducing the overall latency of inference.
[0104] Based on the architecture of the aforementioned communication system, the communication method and apparatus of this solution will be further described below with reference to the accompanying drawings. It is understood that this application uses a first communication device (terminal or core network element) and a second communication device (access network equipment) as examples to illustrate the execution of this interactive illustration, but this application does not limit the execution subject of the interactive illustration. For example, the method executed by the first communication device, such as a terminal or core network element, in this application can also be implemented by a communication / processing module in the terminal or core network element, or by a circuit or chip (such as a modem chip (also known as a baseband chip), or a SoC chip containing a modem core, or a SIP chip, or a GPU) responsible for communication / processing functions in the terminal or core network element; similarly, the method executed by the second communication device, such as an access network equipment, in this application can also be implemented by a module (e.g., a circuit, chip, or chip system) in the access network equipment, or by a logical node, logical module, or software capable of implementing all or part of the functions of the access network equipment.
[0105] Referring to Figure 5, a flowchart illustrating a communication method provided in an embodiment of this application is shown. Optionally, this method can be applied to the aforementioned communication system, such as the communication system shown in Figure 1. The communication method shown in Figure 5 may include steps 501-503. Steps 501-503 are as follows:
[0106] 501. The first communication device sends first information to the second communication device, the first information indicating a first model ID, the first model ID indicating an AI model deployed by the first communication device. Accordingly, the second communication device receives the first information.
[0107] The first communication device may be a terminal or a core network element. The second communication device may be an access network device, such as a base station.
[0108] This AI model could be, for example, an image recognition model, a text generation model, or the like.
[0109] The AI model deployed by the first communication device refers to the AI model possessed by the first communication device, or the AI model supported by the first communication device. In other words, the first communication device can use the AI model corresponding to the first model ID, meaning that the first communication device possesses the model capability of the AI model corresponding to the first model ID.
[0110] The following is an introduction to the first model ID.
[0111] In one possible implementation, the first model ID corresponds to one or more of the following of the AI model deployed on the first communication device: function, number of input nodes, number of output nodes, version release time, and version number. The function of the AI model refers to its purpose. This function could be, for example, image recognition (inputting a single image and outputting the objects detected in the image), or speech recognition (inputting a speech segment and outputting the converted text). The input nodes of the AI model correspond to its input parameters. The number of input nodes could be, for example, 256. Correspondingly, the AI model has 256 input parameters. The output nodes of the AI model correspond to its output parameters. The number of output nodes could be, for example, 64 or 128. Correspondingly, the AI model has 64 or 128 output parameters. The version release time could be, for example, 202401. The version number could be, for example, 01, 02, etc. The first model ID corresponds to one or more of the following: the function, number of input nodes, number of output nodes, version release time, and version number of the AI model deployed on the first communication device. This can be understood as the first model ID being associated with one or more of these attributes. Alternatively, the first model ID is obtained based on one or more of these attributes. This example, by constructing a mapping rule that associates model IDs with model characteristics, allows terminals, access network devices, core network elements, etc., to directly determine model functions and corresponding functional characteristics. By using the model ID generated through the mapping rule, model capabilities on different devices are aligned with minimal overhead.
[0112] The first model ID can be represented by one or more of the following: numbers, letters, symbols, graphics, etc. This scheme does not impose any restrictions on this.
[0113] As exemplarily shown in Table 1, this is a mapping table between model functions and model IDs, representing an embodiment of this application.
[0114] Table 1
[0115] In this example, the model ID corresponds to the model function.
[0116] For example, the first model ID corresponds to the function, number of input nodes, number of output nodes, and version release time of the AI model deployed on the first communication device. This first model ID can be represented as: 0001-256-128-202401, or 0001256128202401. Where 0001 indicates the function of the AI model (e.g., image recognition), 256 indicates the number of input nodes, 128 indicates the number of output nodes, and 202401 indicates the version release time. Another example is that the first model ID can be represented as 0005-256-64-202406. Where 0005 indicates the function of the AI model (e.g., speech recognition), 256 indicates the number of input nodes, 64 indicates the number of output nodes, and 202406 indicates the version release time. It is understood that the above representations of the first model ID are merely examples, and this solution does not limit its specific form.
[0117] In one possible implementation, the first model ID includes a first field and a second field, which respectively indicate one or more of the aforementioned functions, number of input nodes, number of output nodes, version release time, and version number. The first field and the second field indicate different contents.
[0118] In this example, the first and second fields indicate different content. The first field indicates the first part of the aforementioned function, number of input nodes, number of output nodes, release date, and version number, while the second field indicates the second part of the same information. These first and second parts are entirely different. For example, the first field indicates the number of input and output nodes, while the second field indicates the function and release date. Another example is that the first field indicates the function, and the second field indicates the version number. In this example, the model ID is displayed based on fields with specific meanings, intuitively expressing the corresponding model characteristics. The model ID, generated according to rules, aligns with the model capabilities of different devices with minimal overhead.
[0119] Table 2 shows the physical meaning correspondence of each field in a model ID, as exemplified in an embodiment of this application.
[0120] Table 2
[0121] Table 2 is just one example. Of course, the model ID may not include the above-mentioned reserved field #1 and / or reserved field #2, or may include other fields. This solution does not restrict this.
[0122] In one possible implementation, the first information indicates a first checksum, which is generated based on a first model ID. For example, a cyclic redundancy check (CRC) can be used to calculate the checksum based on the model ID. For instance, the checksum f(x) satisfies f(x) = x 8 +x 2 +x+1, where x represents the bits of the first model ID. For example, if the first model ID is 0001-256-128-20240115, the checksum is 15. This example uses a checksum to align computing power between devices. Based on its short bit overhead, it allows terminals, access network devices, core network elements, etc., to complete model capability alignment with less overhead.
[0123] The first model ID has been introduced above. The following section introduces several possible implementation methods for the first communication device to send the first information to the second communication device.
[0124] In one possible implementation, the first communication device obtains the first model ID based on the mapping relationship between model function and model ID. Then, the first communication device sends first information to the second communication device.
[0125] The mapping relationship between the model function and the model ID can be predefined, or it can be configured by a second communication device (access network device) or a third communication device. This predefinition includes, but is not limited to, predefined features in the protocol, factory default configurations, or unified regulations based on regional standards or operator specifications. The aforementioned configuration can refer to the corresponding communication device configuring the device through one or more of the following: radio resource control (RRC) messages, downlink control information (DCI), and media access control (MAC) control elements (MAC CE).
[0126] For example, the first communication device receives a mapping relationship between model functions and model IDs from a second or third communication device. For instance, this mapping relationship between model functions and model IDs is represented by the mapping table described above (as shown in Table 1). The first communication device generates a first model ID based on the mapping table sent by the second or third communication device. Here, the first communication device is a terminal, and the third communication device is a core network element. Alternatively, the first communication device is a core network element, and the third communication device is a terminal.
[0127] In this example, the first communication device sends first information to the second communication device so that the second communication device knows the model capabilities of the first communication device. This enables the alignment of computing power between devices and provides conditions for task partitioning.
[0128] The following describes how the second communication device interacts with the first communication device to enable it to perform model-related tasks.
[0129] In one possible implementation, the second communication device also sends a fourth message to the first communication device, which indicates a third model ID, indicating the AI model deployed by the second communication device.
[0130] In other words, the second communication device informs the first communication device of its own model capabilities, so that the terminal can obtain the model capabilities of the access network device. The terminal can then compare the received model ID with its own model capabilities and provide feedback on its own capabilities, or supplement differentiated capabilities, enabling the two devices to align their model capabilities within a short time delay.
[0131] In one possible implementation, the second communication device obtains the third model ID based on the mapping relationship between model function and model ID. Then, the second communication device sends the aforementioned fourth information to the first communication device.
[0132] The mapping relationship between the model function and the model ID can be predefined, or it can be configured by the first or third communication device. For a detailed description of this part, please refer to the foregoing records; it will not be repeated here.
[0133] For example, the second communication device receives a mapping relationship between model functions and model IDs from the first or third communication device. For instance, the mapping relationship between model functions and model IDs is represented by the mapping table described above (as shown in Table 1). The second communication device generates a third model ID based on the mapping table indicated by the first or third communication device. For example, the first communication device is a terminal, and the third communication device is a core network element. Alternatively, the first communication device is a core network element, and the third communication device is a terminal.
[0134] In one possible implementation, the first model ID includes the model ID in the third model ID. That is, the model capabilities reported by the first communication device are a subset of the model capabilities of the second communication device. The first communication device reports the model ID in the third model ID, which allows the second communication device to perceive the model capabilities possessed by the first communication device, enabling the two devices to align their model capabilities within a shorter time delay and reducing the latency of task partitioning.
[0135] In another possible implementation, the first model ID includes model IDs other than the third model ID. That is, the model capabilities reported by the first communication device are model capabilities that the second communication device does not possess. The first communication device can report model IDs other than the third model ID, enabling the two devices to align model capabilities within a shorter latency, thus reducing the latency of task partitioning.
[0136] In another possible implementation, the aforementioned first model ID includes the model ID in the third model ID, as well as other model IDs besides the third model ID. That is, the model capabilities reported by the first communication device include the model capabilities in the model capabilities of the second communication device, as well as model capabilities that the second communication device does not possess. The first communication device can report not only the model ID in the third model ID, but also model IDs other than the third model ID, enabling full interaction between devices to utilize their respective model capabilities. This allows the two devices to align their model capabilities within a shorter latency, reducing the latency of task partitioning.
[0137] 502. The second communication device sends second information to the first communication device, the second information indicating a second model ID, which is the model ID in the first model ID, and the second model ID is associated with the first task. Accordingly, the first communication device receives the second information.
[0138] The second model ID is a model ID within the first model ID, meaning it is one or more model IDs within the first model ID. This second model ID is associated with the first task, meaning the first task needs to be executed on the model corresponding to the second model ID. This first task could be, for example, a text-based destination localization task. For instance, the second communication device needs to complete a navigation task (e.g., a task requiring real-time video and target command text). The model deployed on the second communication device has in-video obstacle recognition capabilities, while the model deployed on the first communication device has text-based destination localization capabilities. Therefore, the second communication device decomposes the aforementioned task, enabling the second communication device to perform obstacle recognition in the video, and the first communication device to complete the text-based destination localization task. This task decomposition allows multiple devices to complete the task together, enabling distributed inference and shorter task inference latency.
[0139] This example uses the model ID to indicate the result of task decomposition, which saves on task decomposition indication overhead.
[0140] In one possible implementation, the second information further includes first data. This first data is data required to perform the first task. For example, the first task is a model training task, and the first data is data used for model training. Or, for instance, the first task is a text-based destination localization task, and the first data is text data used for destination localization.
[0141] Based on the model capabilities reported by the first communication device, the second communication device indicates the corresponding model ID to the first communication device, so that the first communication device can perform tasks based on the model corresponding to the model ID.
[0142] In one possible implementation, the second communication device is configured with a mapping relationship between model functions and model IDs. Based on the received first information, the second communication device can determine the function of the model deployed by the first communication device through this mapping relationship. Furthermore, the second communication device allocates tasks based on task requirements. For example, if the second communication device needs to complete a speech recognition task, and the first communication device has deployed a model with that speech recognition function, then the second communication device indicates the model ID corresponding to the speech recognition function to the first communication device.
[0143] When the second communication device is unaware of the mapping relationship between the model function and the model ID, in one possible implementation, the first communication device also sends the model function corresponding to the first model ID to the second communication device. Then, the second communication device allocates tasks based on task requirements.
[0144] In one possible implementation, the first task includes a first subtask and a second subtask. The second communication device also sends third information to the first communication device, which includes a first data packet and a second data packet, both having the same Service Function Chain (SFC) header. The first data packet corresponds to the first subtask, and the second data packet corresponds to the second subtask. Accordingly, the first communication device receives the third information.
[0145] In this example, the first communication device is a core network element. The first task could be, for example, a text-based destination location task. The first subtask is text-based keyword extraction; for example, if the text is "find the table in the bathroom," the keyword extraction would be to confirm "table" and "bathroom." The second subtask is keyword-based object location, i.e., determining the location of the table based on "table" and "bathroom." The data in the first data packet is the data required by the first communication device to perform the first subtask, such as providing similar keyword extraction examples for the first subtask. The data in the second data packet is the data required by the first communication device to perform the second subtask, such as an environmental map of the house, used by the first communication device to determine the location corresponding to the keywords. Optionally, the second communication device may only provide data for the first subtask or the second subtask, or may not provide any data (e.g., if the first communication device already possesses the corresponding data).
[0146] In this example, the second communication device sends a third message to the first communication device, in which the first data packet and the second data packet have the same SFC header. This enables the first communication device to provide the same service to different subtasks corresponding to the same task, thus achieving the same QoS guarantee.
[0147] In another possible implementation, the first task described above includes a third subtask and a fourth subtask. The second communication device receives fifth information from the first communication device, which indicates a first configuration file and a second configuration file, both configured with the same 5QI value. The first configuration file corresponds to the third subtask, and the second configuration file corresponds to the fourth subtask.
[0148] In this example, the first communication device is a core network element. The second communication device, based on the same 5QI value indicated by the fifth information, provides consistent service to the third and fourth subtasks during downlink data transmission, thereby achieving the same QoS guarantee. For a description of the third and fourth subtasks, please refer to the descriptions of the first and second subtasks above; they will not be repeated here.
[0149] In one possible implementation, the second communication device also sends a sixth message to the first communication device, indicating the number of in-degrees and out-degrees of the directed acyclic graph (DAG). The in-degree and out-degree represent the number of input (connected) nodes and the number of output (connected) nodes for a computation node, respectively. As shown in Figure 6, node 4 has an in-degree of 784 (see the illustration of weight w in Figure 6; the figure is only an example and not all nodes are shown), and an out-degree of 1. In this example, the second communication device indicates the in-degree and out-degree of the DAG to the first communication device, enabling the first communication device to perceive the sub-task requirements and assist the second communication device in performing some computations. This allows the terminal to accurately obtain information related to the sub-tasks to be executed, assisting other devices in completing task inference with a shorter latency. For example, the second communication device instructs the first communication device to calculate a vector with an in-degree of 784 and an out-degree of 1. After receiving this instruction, the first communication device can perform the calculation based on the data.
[0150] Furthermore, the sixth piece of information also includes raw data. This raw data is data related to the aforementioned sub-tasks (such as the first sub-task, the second sub-task, the third sub-task, or the fourth sub-task). For example, in a navigation task, when the first communication device needs to perform a text-based destination positioning task, it needs to be provided with raw text data information.
[0151] 503. The first communication device executes the first task based on the AI model corresponding to the second model ID.
[0152] The first communication device determines the AI model corresponding to the second model ID, and then performs the first task based on the AI model.
[0153] In one possible implementation, after completing the first task, the first communication device obtains the calculation result. Further, the first communication device sends the calculation result to the second communication device so that the second communication device can perform processing such as summarizing the calculation results.
[0154] In this embodiment, a first communication device sends first information to a second communication device, the first information indicating a first model ID, which in turn indicates an AI model deployed by the first communication device. Then, the second communication device sends second information to the first communication device, the second information indicating a second model ID, which is the model ID within the first model ID and is associated with a first task. Subsequently, the first communication device executes the first task based on the AI model corresponding to the second model ID. In this example, the first communication device indicates its model capabilities to the second communication device via the first model ID, and the second communication device then instructs the first communication device to use the AI model corresponding to the second model ID to execute the first task. This achieves alignment of computing power between devices, provides conditions for task partitioning, enables distributed training, and reduces the overall latency of inference.
[0155] The embodiment shown in Figure 5 above uses the first communication device as a terminal or core network element, and the second communication device as an access network device as an example. Alternatively, this application can also illustrate the interaction using the first communication device (terminal or access network device) and the second communication device (core network element) as the executing entities. For example, the method executed by the first communication device, such as a terminal or access network device, in this application can also be implemented by a communication / processing module in the terminal or access network device, or by a circuit or chip (such as a modem chip (also known as a baseband chip), or a SoC chip containing a modem core, or a SIP chip, or a GPU) responsible for communication / processing functions in the terminal or access network device; the method executed by the second communication device, such as a core network element, in this application can also be implemented by a module (e.g., a circuit, chip, or chip system) in the core network element, or by a logical node, logical module, or software capable of implementing all or part of the core network element functions. Specifically, the communication method may include:
[0156] The access network device or terminal sends first information to the core network element, the first information indicating a first model ID, which in turn indicates the AI model deployed by the access network device or terminal. Correspondingly, the core network element receives the first information.
[0157] Then, the core network element sends second information to the access network device or terminal. This second information indicates a second model ID, which is the model ID in the first model ID and is associated with the first task. Accordingly, the access network device or terminal receives the second information.
[0158] Then, the access network device or terminal performs the first task based on the AI model corresponding to the second model ID.
[0159] In another possible implementation, this application can also illustrate the interaction by using a first communication device (access network equipment or core network element) and a second communication device (terminal) as the executing entities. For example, the method executed by the first communication device, such as an access network equipment or core network element, in this application can also be implemented by a module (e.g., circuit, chip, or chip system) in the access network equipment or core network element, or by a logical node, logical module, or software that can implement all or part of the functions of the access network equipment or core network element; the method executed by the second communication device, such as a terminal, in this application can also be implemented by a communication / processing module in the terminal, or by a circuit or chip (e.g., a modem chip (also known as a baseband chip), or a SoC chip containing a modem core, or a SIP chip, or a GPU) in the terminal responsible for communication / processing functions. Specifically, the communication method may include:
[0160] The access network device or core network element sends first information to the terminal, which indicates a first model ID. The first model ID indicates the AI model deployed by the access network device or core network element. Accordingly, the terminal receives the first information.
[0161] Then, the terminal sends second information to the access network device or core network element. This second information indicates a second model ID, which is the model ID in the first model ID and is associated with the first task. Accordingly, the access network device or core network element receives the second information.
[0162] Then, the access network device or core network element executes the first task based on the AI model corresponding to the second model ID.
[0163] For details on this part, please refer to the description in the embodiment shown in Figure 5, which will not be repeated here.
[0164] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below. It is understood that the division of multiple units or modules in the various apparatus embodiments of this application is only a logical division based on function and is not intended to limit the specific structure of the apparatus. In specific implementations, some functional modules may be subdivided into more smaller functional modules, and some functional modules may be combined into a single functional module. However, regardless of whether these functional modules are subdivided or combined, the general flow executed by the apparatus is the same. For example, some apparatuses include a receiving unit and a transmitting unit. In some designs, the transmitting unit and the receiving unit can also be integrated into a communication unit, which can implement the functions implemented by the receiving unit and the transmitting unit. Typically, each unit corresponds to its own program code (or program instructions). When the program code corresponding to each unit runs on the processor, it causes the unit to be controlled by the processing unit to execute the corresponding flow and thus achieve the corresponding function.
[0165] This application also provides an apparatus for implementing any of the above methods. For example, a communication apparatus is provided that includes modules (or means) for implementing the steps performed by the first communication apparatus or the second communication apparatus in any of the above methods.
[0166] Figure 7 illustrates a possible exemplary block diagram of the communication device involved in the embodiments of this application. As shown in Figure 7, the communication device 700 may include modules or units for implementing the method embodiments described above. In one possible design, the communication device 700 includes a processing unit 702 and a communication unit 703. Optionally, the communication device 700 may further include a storage unit 701 for storing device program code and / or data.
[0167] The communication device 700 can be a terminal-side device as described in the above embodiments, such as a terminal or a communication module in a terminal, or a circuit or chip in a terminal that is responsible for communication functions.
[0168] For example, in one possible design, the communication unit 703 is used to: send first information indicating a first model ID, which indicates an AI model deployed by the first communication device.
[0169] The communication unit 703 is further configured to: receive second information indicating a second model ID, the second model ID being a model ID in the first model ID, and the second model ID being associated with the first task.
[0170] In one embodiment, the processing unit 702 is used to: perform a first task based on the AI model corresponding to the second model ID.
[0171] In one possible design, the first model ID corresponds to one or more of the following:
[0172] Function, number of input nodes, number of output nodes, release date, and version number.
[0173] In one possible design, the first model ID includes a first field and a second field, which respectively indicate one or more of the following: the function, the number of input nodes, the number of output nodes, the release date of the version, and the version number. The first field and the second field indicate different contents.
[0174] In one possible design, the first information indicates a first checksum, which is generated based on the first model ID.
[0175] In one possible design, the communication unit 703 is also used to: receive fourth information indicating a third model ID, which indicates an AI model deployed by the second communication device.
[0176] In one possible design, the first model ID includes: the model ID in the third model ID, and / or other model IDs besides the third model ID.
[0177] In one possible design, when the communication device 700 is a terminal or a communication module within a terminal, the function of the processing unit 702 can be implemented by one or more processors. Specifically, the processor may include a modem chip, or a system-on-a-chip (SoC) chip or a SIP chip containing a modem core. The function of the communication unit 703 can be implemented by transceiver circuitry.
[0178] In one possible design, when the communication device 700 is a circuit or chip in a terminal responsible for communication functions, such as a modem chip or a system-on-a-chip (SoC) or SIP chip containing a modem core, the function of the processing unit 702 can be implemented by a circuit system in the aforementioned chip that includes one or more processors or processor cores (for example, the first task is to perform object recognition based on the sensed echo signal, and the baseband chip can participate in executing the steps performed by the processing unit 702). The function of the communication unit 703 can be implemented by the interface circuit or data transceiver circuit on the aforementioned chip.
[0179] In one possible design, when the communication device 700 includes circuitry or chips responsible for communication functions in the terminal, such as a modem chip, a system-on-a-chip (SoC) chip containing a modem core, or a SIP chip, and also includes an AI chip, the function of the processing unit 702 can be implemented by a circuit system in the AI chip that includes one or more processors or processor cores. The function of the communication unit 703 can be implemented by the interface circuitry or data transceiver circuitry on the aforementioned modem chip, SoC chip containing a modem core, or SIP chip.
[0180] In one possible design, when the communication device 700 is a terminal or a processing module within a terminal, the functionality of the processing unit 702 can be implemented by one or more processors. Specifically, the processor may include a GPU, or a system-on-a-chip (SoC) or SIP chip containing a GPU. The functionality of the communication unit 703 can be implemented by transceiver circuitry.
[0181] In one possible design, when the communication device 700 is a circuit or chip in a terminal responsible for processing functions, such as a GPU or a system-on-a-chip (SoC) or SIP chip containing a GPU, the function of the processing unit 702 can be implemented by a circuit system in the aforementioned chip that includes one or more processors or processor cores. The function of the communication unit 703 can be implemented by interface circuitry or data transceiver circuitry on the aforementioned chip.
[0182] The communication device 700 can be a core network element in the above embodiments. For example, it can be a module (e.g., a circuit, chip, or chip system) in the core network element, or a logical node, logical module, or software that can implement all or part of the functions of the core network element.
[0183] Some possible implementations in this regard can refer to the terminal-side implementation methods mentioned above, and will not be elaborated further.
[0184] In one possible design, the first task includes a first subtask and a second subtask. The communication unit 703 is further configured to: receive third information, which includes a first data packet and a second data packet, the first data packet and the second data packet having the same Service Function Chain (SFC) header, the first data packet corresponding to the first subtask, and the second data packet corresponding to the second subtask.
[0185] The communication device 700 can be a network-side device in the above embodiments, such as a network-side access network device, a module (e.g., circuit, chip or chip system) in the access network device, or a logic node, logic module or software that can implement all or part of the functions of the access network device.
[0186] For example, in one embodiment, the communication unit 703 is configured to: receive first information indicating a first model ID, the first model ID indicating an AI model deployed by the first communication device.
[0187] The communication unit 703 is further configured to: send second information, the second information indicating a second model ID, the second model ID being a model ID in the first model ID, and the second model ID being associated with the first task.
[0188] In one possible design, the first task includes a first subtask and a second subtask. The communication unit 703 is also used to send third information, which includes a first data packet and a second data packet. The first data packet and the second data packet have the same Service Function Chain (SFC) header. The first data packet corresponds to the first subtask, and the second data packet corresponds to the second subtask.
[0189] In one possible design, the first task includes a third subtask and a fourth subtask, and the communication unit 703 is further configured to: send a fifth message indicating a first configuration file and a second configuration file, wherein the first configuration file and the second configuration file are configured with the same 5QI value, the first configuration file corresponds to the third subtask, and the second configuration file corresponds to the fourth subtask.
[0190] In one possible design, the communication unit 703 is also used to: send a fourth message indicating a third model ID, which in turn indicates an AI model deployed by the second communication device.
[0191] It is understood that the division of units in the above-described device is merely a logical functional division. One function can correspond to one functional unit, or two or more functions can be integrated into one functional unit. In actual implementation, all or some units can be integrated onto a single physical entity, or distributed across different physical entities. Furthermore, the aforementioned functional units can be implemented in hardware, software, or a combination of both. Whether a function is executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for specific applications, but such implementations should not be considered beyond the scope of this application.
[0192] In one example, the functional unit in any of the above devices may be one or more integrated circuits configured to implement the above methods, such as: one or more application-specific integrated circuits (ASICs), or one or more central processing units (CPUs), one or more microcontroller units (MCUs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0193] In one example, storage unit 701 may include random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory and / or registers, etc.
[0194] It should be understood that the division of modules in the above devices is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, modules in a communication device can be implemented by a processor calling software; for example, a communication device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each module in the device. The processor can be, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the modules in the device can be implemented as hardware circuits. The functionality of some or all units can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), and the functionality of some or all of the above units is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD), such as a field-programmable gate array (FPGA), which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the above units. All modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.
[0195] Referring to Figure 8, which is a structural schematic diagram of a terminal 800 provided in an embodiment of this application, the terminal 800 corresponds to the terminal shown in Figure 1 and is used to implement the operation of the terminal in the above embodiments. As shown in Figure 8, the terminal includes: one or more antennas 810, a radio frequency processing system 820, and a processor system 830.
[0196] In the downlink or sidelink direction, the RF processing system 820 receives RF signals through the antenna 810 and sends the RF-processed signals to the processor system 830 for further processing. In the uplink or sidelink direction, the processor system 830 processes the terminal-side information and sends it to the RF processing system 820, which then processes the signal and transmits it through the antenna 810.
[0197] In one example, the radio frequency (RF) processing system 820 serves as the communication interface for external communication of the terminal and may include an RF front end (RFFE) 821 and an RF transceiver 822. The RFFE 821 is primarily used for one or more processing operations, such as shaping, passband selection, or gain adjustment, on the RF signals received by the antenna or those to be transmitted through the antenna. It may include one or more components such as RF switches, duplexers, filters, power amplifiers, antenna tuners, and low-noise amplifiers. The RFFE 821 can be a circuit system composed of multiple discrete devices or integrated into one or more chips. The RF transceiver 822 processes the RF signals received by the RFFE into baseband / IF signals for further processing by the processor system 830, and processes the baseband / IF signals provided by the processor system 830 into RF signals for transmission to the RFFE 821. The baseband / IF signals transmitted between the RF transceiver 822 and the processor system 830 can be digital or analog signals. The RF transceiver 822 can be implemented by one or more chips, which are commonly referred to as RF ICs.
[0198] In one example, the processor system 830 may include one or more processors for processing signals and executing one or more communication protocols. Optionally, the processor system 830 may also include a memory 836. In one example, the one or more processors include at least one baseband processor 831 (also known as a modem processor). The memory 836 is used to store data and / or computer program instructions. Optionally, the processor system 830 may also include one or more application processors 832 for implementing processing of the terminal operating system and application layer. The application processor 832 may include, for example, a GPU. Optionally, the processor system 830 may also include one or more of a voice subsystem 833, a multimedia subsystem 834, or an interface circuit 835. The voice subsystem 833 is used to process voice signals, the multimedia subsystem 834 is used to handle multimedia-related operations, such as video encoding / decoding, image processing, etc., and the interface circuit 835 is used to implement communication with other terminal components, such as a display 840, an input device 850, a memory 860, etc. The above-mentioned components in the processor system 830 can communicate with each other via a bus or communication interface circuit.
[0199] In one example, the processor system 830 can be packaged as a single processor chip, such as a SoC chip or a SIP chip. In another example, the processor system 830 can be a system composed of multiple chips; for example, the baseband processor 831 can be packaged as a single chip, or packaged with part or all of the circuitry of the radio frequency processing system into a single chip.
[0200] In one example, memory 836 can be on-chip memory, i.e., located on the system-on-a-chip 830. In another example, memory 860 can be off-chip memory, i.e. located outside the system-on-a-chip 830.
[0201] In one example, the baseband processor 831 may include one or more processor cores 8311 and interface circuitry 8314. The one or more processor cores 8311 are used to process signals and execute one or more communication protocols. Optionally, the baseband processor 831 may also include a memory 8312 for storing at least a portion of the corresponding computer program instructions and / or data. In one example, the one or more processor cores 8311 execute the computer program instructions stored in the memory 8312 to implement the relevant operations (such as…) in the above method embodiments. In this disclosure, the memory 8312 storing the corresponding computer program instructions and / or data may mean that the memory 8312 stores all the corresponding computer program instructions and / or data for the processor core 8311 to execute; or it may mean that the memory 8312 stores a portion of the corresponding computer program instructions and / or data, which includes the computer program instructions and / or data currently needed to be executed by the processor core 8311. The memory 8312 can store different portions of computer program instructions and / or data multiple times for the processor core 8311 to execute in order to implement the relevant operations in the above method embodiments. The interface circuit 8314 serves as a communication interface for communication with other components, such as transmitting signals with the RF processing system 820, communicating with other subsystems and related components of the processor system 830 via a bus, such as transmitting data control signals with the application processor 832, and transmitting data or computer program instructions with the memory 836 or memory 860. Optionally, to reduce the load on the processor core, a baseband signal processing circuit 8313 can be provided to perform at least some baseband signal processing, including one or more of signal demodulation, modulation, encoding, or decoding.
[0202] In one example, the communication device provided in this application may be a terminal 800, a communication module including a processor system 830 and a radio frequency processing system 820, or a baseband processor 831.
[0203] The processor, processor system, application processor, baseband processor, processor circuit, or processor core mentioned above can be collectively referred to as a processor. The processor may include one or more of the following: central processing unit (CPU), digital signal processor (DSP), microprocessor unit (MPU), microcontroller unit (MCU), graphics processing unit (GPU), field programmable gate array (FPGA), artificial intelligence processor (AI processor), or neural processing unit (NPU).
[0204] The aforementioned memory may include one or more of the following storage media: random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), phase-change memory (PCM), resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), cache, register, read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), hard disk, etc. In one example, computer program instructions for executing the above embodiments may be stored in non-volatile memory, such as at least a portion of the aforementioned memory 860 (e.g., one or more of ROM, flash memory, EPROM, or hard disk). When the terminal is running, the corresponding computer program instructions may be partially or wholly loaded onto a memory with a faster transfer speed than the processor, such as at least a portion of memory 836 and / or memory 8312 (e.g., one or more of RAM, SRAM, DRAM, PCM, RERAM, MRAM, FRAM, cache, or register), for the processor to execute in order to implement the steps in the above method embodiments.
[0205] In one example, the RF transceiver 822 and the RF front-end 821 can also be packaged in a single chip. In another example, the RF transceiver 822, the RF front-end 821, and the baseband processor 831 can also be packaged in a single chip.
[0206] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of any of the above methods.
[0207] This application also provides a computer program product containing instructions. When the computer program product is run on a computer or processor, it causes the computer or processor to perform one or more steps of any of the methods described above.
[0208] It is understood that in this application, "instruction" can include direct instruction, indirect instruction, explicit instruction, and implicit instruction. When describing a certain instruction information to indicate A, it can be understood that the instruction information carries A, directly indicates A, or indirectly indicates A. In this application, the information indicated by the instruction information is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index, or indirectly indicating the information to be instructed by indicating other information, wherein there is an association between the other information and the information to be instructed. It is also possible to indicate only a part of the information to be instructed, while the other parts of the information to be instructed are known or agreed upon in advance. For example, the instruction of specific information can also be achieved by using the arrangement order of various information in advance (e.g., as specified by a protocol), thereby reducing the instruction overhead to a certain extent. The information to be instructed can be sent as a whole or divided into multiple sub-information to be sent separately, and the sending period and / or sending time of these sub-information can be the same or different. This application does not limit the specific sending method. The sending period and / or timing of these sub-information messages can be predefined, for example, according to a protocol, or configured by the transmitting device by sending configuration information to the receiving device.
[0209] In this application, "sending information" can be understood as one device sending information to another device, or it can also be understood as one logical module within a device sending information to another logical module. For example, "access network device sending information" can be understood as the access network device sending information to another device (such as a terminal), or it can be understood as logical module 1 in the access network device sending information to logical module 2 in the access network device.
[0210] In this application, "receiving information" can be understood as one device receiving information from another device, or it can also be understood as a logical module within a device receiving information from another logical module. For example, "access network device receiving information" can be understood as the access network device receiving information from another device (such as a terminal), or it can be understood as logical module 1 in the access network device receiving information from logical module 2 in the access network device.
[0211] In this application, phrases such as "sending information to... (e.g., a terminal)" or related illustrations in the accompanying drawings can be understood as indicating that the destination of the information is a terminal. This can include sending information directly or indirectly to a terminal. Similarly, phrases such as "receiving information from... (e.g., a terminal)," "receiving information from... (e.g., a terminal)," or "receiving information sent by (e.g., a terminal)," or related illustrations in the accompanying drawings, can be understood as indicating that the source of the information is a terminal. This can include receiving information directly or indirectly from a terminal. Information may undergo necessary processing between the source and destination, such as format changes, but the destination can understand the valid information from the source. Similar expressions in this application can be interpreted similarly and will not be elaborated further here.
[0212] The terms "system" and "network" in this application embodiment are used interchangeably. "At least one" refers to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, "at least one of A, B, or C" includes A, B, C, AB, AC, BC, or ABC; "at least one of A, B, and C" can also be understood as including A, B, C, AB, AC, BC, or ABC. Furthermore, unless otherwise specified, the ordinal numbers such as "first" and "second" mentioned in this application embodiment are used to distinguish multiple objects and are not used to limit the order, sequence, priority, or importance of multiple objects.
[0213] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.
[0214] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0215] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0216] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0217] The terms "comprising" and "having," and any variations thereof, used in this application as described below, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or optionally include other steps or units inherent to such processes, methods, products, or apparatus. It should be noted that in this application, words such as "exemplary" or "for example" are used to indicate illustrative, exemplary, or descriptive purposes. Any method or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other methods or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0218] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling, direct coupling, or communication connection shown or discussed between each other may be indirect coupling or communication connection through some interfaces, apparatuses, or units, and may be electrical, mechanical, or other forms.
[0219] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0220] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be read-only memory (ROM), random access memory (RAM), or magnetic media, such as floppy disks, hard disks, magnetic tapes, magnetic disks, or optical media, such as digital versatile discs (DVDs), or semiconductor media, such as solid-state disks (SSDs).
[0221] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A communication method characterized by comprising: The method comprises: sending first information, the first information indicating a first model identifier (ID), the first model ID indicating an artificial intelligence (AI) model deployed by a first communication device; receiving second information, the second information indicating a second model ID, the second model ID being a model ID in the first model ID, the second model ID being associated with a first task; performing the first task based on an AI model corresponding to the second model ID.
2. The method of claim 1, wherein, The first model ID corresponds to one or more of the following of the AI model deployed by the first communication device: function, number of input nodes, number of output nodes, version release time, version number.
3. The method of claim 2, wherein, The first model ID comprises a first field and a second field, the first field and the second field respectively indicating one or more of the function, the number of input nodes, the number of output nodes, the version release time, and the version number, the first field and the second field indicating different contents.
4. The method according to any one of claims 1 to 3, characterized in that, The first information indicates a first check code, the first check code being generated based on the first model ID.
5. The method according to any one of claims 1 to 4, characterized in that, The first task comprises a first subtask and a second subtask, and the method further comprises: receiving third information, the third information comprising a first data packet and a second data packet, the first data packet and the second data packet having the same service function chain (SFC) header, the first data packet corresponding to the first subtask, and the second data packet corresponding to the second subtask.
6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: receiving fourth information, the fourth information indicating a third model ID, the third model ID indicating an AI model deployed by a second communication device.
7. The method of claim 6, wherein, The first model ID comprises a model ID in the third model ID, and / or a model ID other than the third model ID.
8. A communication device, characterized by The apparatus comprises a module or unit for implementing the method of any of claims 1-7.
9. A computer-readable storage medium, characterized in that, The apparatus stores a computer program, which, when executed by a processor, causes the method of any of claims 1-7 to be implemented.
10. A computer program product comprising instructions which, when executed on a processor, cause the method of any of claims 1-7 to be implemented.
11. A communications device, characterized by The apparatus comprises an interface circuit and one or more processors, the one or more processors being coupled with a memory, the memory being configured to store a computer program or instructions, which, when executed by the one or more processors, cause the apparatus to implement the method of any of claims 1-7.
12. The apparatus of claim 11, wherein, The interface circuit is configured to implement a communication function within the apparatus and / or a communication function of the apparatus with other apparatuses or components. The interface circuit is configured to implement a communication function within the apparatus and / or a communication function of the apparatus with other apparatuses or components.
Citation Information
Patent Citations
Model configuration method and device
CN116541088A
Communication method, core network device, terminal, communication system and storage medium
CN118202675A
Information reporting method, apparatus and device, and storage medium
WO2021142609A1
Network model management method, session establishment or modification method, and apparatus
WO2021163895A1
Network-centric life cycle management of ai / ML models deployed in a user equipment (UE)
WO2023148010A1