Configuration method for distributed ai task, management node, and computing node
By constructing a generalized method for AI models in management nodes, and deploying sub-models in a distributed manner based on task information and functional module attribute information, the problems of limited storage and fixed strategies in management nodes are solved, thereby improving the performance and adaptability of AI services on computing nodes.
Patent Information
- Application Number
- PCT/CN2025/077133
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-02-13
- Publication Date
- 2025-11-27
AI Technical Summary
The limited number of AI models stored within the management node and the fixed processing strategies make it impossible to find a suitable model for the current AI task, affecting the performance of the computing node in providing AI services.
The management node receives task requests, determines sub-models based on task information and functional module attribute information, and distributes them across computing nodes to build target models suitable for handling tasks, thereby achieving the generalization of AI models.
It improves the performance and efficiency of computing nodes in providing AI services, ensures that computing nodes have sufficient AI resources and network status support, and enhances the functional adaptability and personalized service capabilities of sub-models.
Smart Images

Figure CN2025077133_27112025_PF_FP_ABST
Abstract
Description
A distributed AI task configuration method, a management node and a computing node
[0001] The present application claims priority from the Chinese patent application No. 202410376209.1 filed on March 28, 2024, and entitled "A distributed AI task configuration method, a management node and a computing node", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Embodiments of the present application relate to the field of communication, and in particular to a distributed AI task configuration method, a management node and a computing node. BACKGROUND
[0003] With the development of artificial intelligence (AI) technology, AI technology has derived intelligent applications in various industries, and has also derived powerful functions in the field of communication.
[0004] In the field of communication, on the one hand, a communication module using AI technology can replace a traditional communication module, thereby realizing more intelligent network functions. For example, AI4Net (AI for Network) can improve the performance, efficiency and user service experience of the network itself through AI. For example, AI technology is applied to the physical layer, which can realize intelligent beam management, channel prediction and resource allocation functions. For another example, AI technology is applied to the upper layer, which can realize terminal trajectory prediction, load balancing and energy saving functions. On the other hand, the computing, transmission capability and perceived information in the network can also support third-party AI applications of mobile terminals. For example, the calculation of third-party AI applications is offloaded to the computing node in the network, and the environment information and user behavior perceived by the network are used to provide better personalized services. In view of the potential scenarios and driving forces of the above two aspects, the deep integration of AI and communication has also become one of the important visions of the 6th generation mobile communication technology (6G) system: through the deep integration of AI functions, the mobile network provides endogenous AI services for network and third-party applications.
[0005] In the traditional technology, a plurality of computing nodes are distributed in the access network, and a management node managing the computing nodes can receive a task request from a base station or a third-party AI application, and select an AI model with a specific function and a computing node participating in the service for the task based on the task request. Then, the management node sends the selected AI model to each computing node, so that the computing node deploys the AI model, and processes the task data using the AI model, and then outputs the processing result of the AI task.
[0006] However, the number of AI models with specific functions that can be stored in the management node is limited, and the task processing strategy of the AI models with specific functions is fixed. The management node cannot necessarily find a suitable AI model for the current AI task for deployment in each computing node participating in the service, thereby affecting the performance of the AI service provided by each distributed computing node. SUMMARY
[0007] The present application provides a distributed AI task configuration method, a management node and a computing node, which are used to realize the generalization of AI models in the management node, so that the management node can instruct each computing node to deploy sub-models that can constitute a target model suitable for processing the task, thereby improving the performance of the AI service provided by each distributed computing node.
[0008] In a first aspect, the present application provides a distributed AI task configuration method, which can be executed by a management node managing a computing node, or by a component (for example, a processor, a chip or a chip system, etc.) of the management node. Taking the management node as an example, the management node receives a first task request, and the first task request includes task information of a first task. Then, the management node determines at least two sub-models and attribute information of the sub-models based on the task information and attribute information of a plurality of functional modules stored in the management node. The sub-models include at least one functional module and / or a part of a functional module. Then, the management node sends the sub-models to be deployed in each of at least two computing nodes and the attribute information of the sub-models to the at least two computing nodes, so that the at least two sub-models are distributed and deployed in the at least two computing nodes. The attribute information of the sub-models is determined based on the attribute information of the functional modules constituting the sub-models, and the attribute information of the sub-models is used for the computing nodes to deploy the sub-models. Then, the management node sends configuration information of the sub-models deployed in the at least two computing nodes to the at least two computing nodes based on the task information, and the configuration information of the sub-models is used to configure model parameters used by the sub-models when participating in processing the first task and to establish a connection relationship between the sub-models and the sub-models deployed in different computing nodes.
[0009] In the aspect, the management node can determine at least two sub-models based on the task information of the first task and the stored plurality of function modules, and send attribute information and configuration information of the at least two sub-models to at least two computing nodes, so that the at least two sub-models deployed on the at least two computing nodes can jointly process the first task. Since the at least two sub-models are sub-models spliced by the management node based on the task information using the function modules, the management node realizes the generalization of the AI model through the stored function modules, without the management node pre-storing specific AI models for tasks. The management node determines a sub-model suitable for processing the task according to the task information, thereby improving the performance of the distributed computing nodes providing AI services.
[0010] In a possible implementation, the attribute information of the function module includes a function type of the function module and a requirement of the function module, and the requirement of the function module is used to indicate a requirement for data processing and / or data transmission when the function of the function module is implemented. The function type of the function module is used to indicate what AI processing function the function module implements and an adaptation condition for the function module to implement the AI processing function.
[0011] In the embodiment, the attribute information of the function module includes the requirement of the function module, which is beneficial for the management node to determine a sub-model capable of meeting the task requirement of the first task based on the requirement of the function module, thereby improving the performance of the computing node deploying the sub-model determined by the management node to provide AI services.
[0012] In a possible implementation, the task information of the first task includes a task type of the first task and a task requirement of the first task, and the task requirement of the first task is used to indicate a requirement for data processing and / or data transmission of the distributed AI processing service requested by the first task.
[0013] In the embodiment, the task information of the first task includes the task requirement of the first task, which is beneficial for the management node to determine a sub-model capable of meeting the task requirement of the first task based on the requirement of the function module, thereby improving the performance of the computing node deploying the sub-model determined by the management node to provide AI services.
[0014] In a possible implementation, the management node determines the at least two sub-models and the attribute information of the sub-models based on the task information and the attribute information of the plurality of function modules stored by the management node, including: the management node determines a target model supporting processing the first task based on the task information and the attribute information of the plurality of function modules, the target model is composed of at least one function module in the plurality of function modules, the task type of the first task is used to determine the model type of the target model, the performance requirement of the target model meets the task requirement of the first task, and the performance requirement of the target model is determined based on the requirement of the at least one function module; and then the management node determines the at least two sub-models and the attribute information of each sub-model based on the target model.
[0015] In this embodiment, the management node can filter out the function module that can meet the task requirement of the first task based on the task requirement of the first task and the performance requirement of the function module, and further ensure that the target model composed of the plurality of function modules can also meet the task requirement of the first task, thereby improving the matching rate of the management node in configuring the task and improving the performance of the AI processing service.
[0016] In a possible implementation, the attribute information of the sub-model includes AI resource requirement of the sub-model; and before the management node sends the to-be-deployed sub-model of each of the at least two computing nodes and the attribute information of the sub-model to the at least two computing nodes, the method further includes: the management node acquires resource state information of the at least one computing node, the resource state information is used to indicate the use state of AI resources of the computing node, and the AI resources include model resources, computing resources and data resources; the management node determines the at least two computing nodes based on the resource state information of the at least one computing node, and determines the to-be-deployed sub-model of each of the at least two computing nodes, and the AI resources of the computing node meet the AI resource requirement of the to-be-deployed sub-model of the computing node.
[0017] In this embodiment, the management node determines the to-be-deployed sub-model of each of the computing nodes based on the collected resource state of each of the computing nodes, which is beneficial to ensure that the computing node has sufficient AI resources to deploy the sub-model, reduce the probability of failure of the computing node in deploying the sub-model, and improve the efficiency of the management node in configuring the sub-model for each of the computing nodes.
[0018] In a possible implementation, the method further includes: receiving, by the management node, network state information from an access network device connected to the computing node, the network state information being used to indicate a network state of the communication device applying for the first task; determining, by the management node, configuration information of the sub-model to be deployed on the computing node based on the task information, the resource state information of the computing node, and the network state information; and wherein the network state information includes first network state information and / or second network state information, the first network state information being used to indicate a radio state of the access network device connected to the computing node, and the second network state information being used to indicate a radio state of the terminal device sending the first task request.
[0019] In the embodiment, when determining the configuration information of the sub-model to be deployed on the computing node, the management node not only considers the task information of the first task, but also considers the resource state information and the network state information of the computing node, which is beneficial to determine appropriate configurations for each sub-model, thereby improving the performance of the AI model and improving the performance of the computing node in providing AI services.
[0020] In a possible implementation, the attribute information of the function module further includes information of a plurality of network layers included in the function module, an input dimension of the function module, and an output dimension of the function module; and the attribute information of the sub-model includes information of a plurality of network layers included in the sub-model.
[0021] In a possible implementation, the configuration information of the sub-model includes parameters of enabled network layers, the parameters of the enabled network layers being used to indicate network layers enabled by the sub-model deployed on the computing node when participating in processing the first task.
[0022] In the embodiment, the management node can flexibly indicate which network layers in the sub-model are enabled through the parameters of the enabled network layers, thereby improving the efficiency of the management node in configuring the sub-model.
[0023] In a possible implementation, the configuration information of the sub-model further includes a connection relationship of the sub-model with a sub-model deployed on another computing node.
[0024] In the embodiment, the management node can indicate the connection relationship of the sub-model with the sub-model deployed on another computing node to the computing node, so as to ensure that the sub-models deployed on different computing nodes can interact with each other, which is beneficial to improve the efficiency of the management node in configuring the sub-model.
[0025] In a possible implementation, the management node sends configuration information of the sub-model deployed by the computing node to the at least two computing nodes based on the task information, including: the management node sends first configuration information of a first sub-model to a first computing node, the first configuration information including first model parameters, the first model parameters being used to indicate enabling a first network layer in the first sub-model, the first network layer being at least one network layer in a plurality of network layers indicated by attribute information of the first sub-model.
[0026] In the embodiment, the management node sends at least first model parameters to the first computing node, the first model parameters indicating part of the network layers in the plurality of network layers indicated by the attribute information of the first sub-model, which indicates that the attribute information of the first sub-model indicates part of the redundant network layers, i.e., the network layers that are not enabled when processing the first task. It can be understood that the management node configures the first sub-model with the network layers that need to be enabled and the redundant network layers through the attribute information of the first sub-model, and indicates the network layers that need to be enabled for the first task through the configuration information of the first sub-model. This is conducive to enhancing the function or capability of the sub-model based on the redundant network layers by the computing node, and improving the adaptability of the sub-model deployed in the computing node to the requirements of the first task.
[0027] In a possible implementation, the method further includes: the management node sends second configuration information of the first sub-model to the first computing node, the second configuration information including second model parameters, the second model parameters being used to indicate enabling a second network layer in the first sub-model, the second network layer being at least one network layer in the first computing node, and the second network layer being different from the first network layer.
[0028] In the embodiment, the management node enables the new network layer in the first sub-model by the first computing node through the new configuration information, so as to enhance the function of the first sub-model, which is conducive to improving the adaptability of the sub-model deployed in the computing node to the requirements of the first task.
[0029] In a possible implementation, the method further includes: the management node sends a target model and attribute information of the target model to the first computing node, the attribute information of the target model including information of network layers included in each of the at least two sub-models, and a connection relationship between the network layers of different sub-models in the at least two sub-models.
[0030] In the embodiment, the management node sends the target model and the attribute information of the target model to the computing node, which is conducive to enabling the computing node to autonomously modify the network layers enabled by the first sub-model based on the attribute information of the target model, so as to enhance the sub-model, and improve the service quality of the AI service by the sub-model in the computing node.
[0031] In a possible implementation, the method further includes: the management node sending configuration information of the participating node to the at least one participating node, the configuration information of the participating node being used to indicate a data interaction strategy between the participating node and the computing node.
[0032] In a possible implementation, the method further includes: the management node receiving auxiliary information of the first task, the auxiliary information including prediction information of the access network device on the network environment and / or the user behavior, and the auxiliary information being related to the task type of the first task.
[0033] In this embodiment, the access network device provides the management node with the auxiliary information of the first task, so that the management node sends the preprocessed auxiliary information to the computing node, and then prompts the computing node to take the auxiliary information as input data of the AI model, which is beneficial to enhancing the performance of the sub-model in processing the first task, improving the capability of the sub-model, and providing personalized AI services according to user preferences.
[0034] In a possible implementation, the method further includes: the management node receiving at least one backbone model and attribute information of the backbone model, the attribute information of the backbone model being used to indicate model functions and model parameters of the backbone model; and the management node determining a plurality of function modules and attribute information of each function module based on the at least one backbone model, one function module being used to implement a sub-function of the model functions of the backbone model, and different function modules being used to implement different sub-functions of the model functions of the backbone model or different precisions / granularities of the same sub-function.
[0035] In this embodiment, the management node can split the received backbone model into function modules according to functions for storage, instead of directly storing the backbone model, which is not only beneficial to saving storage resources of the management node, but also can facilitate the management node to generate sub-models by using the function modules, and improves the diversity of the sub-models that can be generated by the management node. For the deployment node, through interconnection between the function modules, various types of distributed AI paradigms can be flexibly supported.
[0036] In a possible implementation, the method further includes: the management node receiving at least one function module and attribute information of the function module.
[0037] In a second aspect, the present application provides a method for configuring a distributed AI task, which can be executed by a computing node managed by a management node or by a component (e.g., a processor, a chip, or a chip system, etc.) of the computing node. Taking a first computing node as an example, the first computing node is a node managed by the management node and participating in processing a distributed AI task. After the management node determines a first sub-model and attribute information of the first sub-model to be sent to the first computing node, the first computing node can receive the first sub-model and the attribute information of the first sub-model of the first task sent by the management node, and the attribute information of the first sub-model is used to indicate the structure and function of the first sub-model. Then, the first computing node deploys the first sub-model based on the attribute information of the first sub-model. Then, the first computing node receives configuration information of the first sub-model sent by the management node, and the configuration information of the first sub-model is used to configure model parameters used by the first sub-model when participating in processing the first task and a connection relationship between the first sub-model and sub-models deployed in different computing nodes. The first computing node configures the first sub-model based on the configuration information of the first sub-model and establishes the connection relationship between the first sub-model and the sub-models deployed in different computing nodes.
[0038] In a possible implementation, the attribute information of the first sub-model includes information of a plurality of network layers included in the first sub-model.
[0039] In a possible implementation, the configuration information of the first sub-model includes parameters of enabled network layers, and the parameters of the enabled network layers are used to indicate network layers enabled by the first sub-model deployed in the first computing node when participating in processing the first task.
[0040] In a possible implementation, the configuration information of the first sub-model further includes a connection relationship between the first sub-model and sub-models deployed in other computing nodes.
[0041] In a possible implementation, the configuration information of the first sub-model includes first configuration information, and the first configuration information includes first model parameters used to indicate that a first network layer in the first sub-model is enabled, and the first network layer is at least one network layer in the plurality of network layers indicated by the attribute information of the first sub-model. In this case, the first computing node deploys the first sub-model based on the attribute information of the first sub-model, including: the first computing node enables the first network layer in the first sub-model based on the first model parameters.
[0042] In a possible implementation, the method further includes: receiving, by the first computing node, second configuration information of the first sub-model, the second configuration information including second model parameters, the second model parameters being used to indicate enabling a second network layer in the first sub-model, the second network layer being at least one network layer in the first computing node, the second network layer being different from the first network layer; and enabling, by the first computing node, the second network layer in the first sub-model based on the second model parameters.
[0043] In a possible implementation, the method further includes: receiving, by the first computing node, the target model and attribute information of the target model, the attribute information of the target model including information of network layers included in each of the at least two sub-models, and connection relationships between network layers of different sub-models in the at least two sub-models.
[0044] In a possible implementation, the method further includes: sending, by the first computing node, resource state information of the first computing node, the resource state information being used to indicate a usage state of an AI resource of the first computing node, the AI resource including a model resource, a computing resource, and a data resource.
[0045] In a third aspect, an apparatus is provided, which can be the management node in the foregoing embodiments, or a chip in the management node. The apparatus can include a processing module and a transceiver module. When the apparatus is the management node, the processing module can be a processor, and the transceiver module can be a transceiver. The management node can further include a storage module, which can be a memory. The storage module is configured to store instructions, and the processing module is configured to execute the instructions stored in the storage module, so that the management node performs the method in any of the embodiments of the first aspect. When the apparatus is a chip in the management node, the processing module can be a processor, and the transceiver module can be an input / output interface, a pin, or a circuit, etc. The processing module executes the instructions stored in the storage module, so that the management node performs the method in any of the embodiments of the first aspect. The storage module can be a storage module (for example, a register, a cache, etc.) in the chip, or a storage module (for example, a read-only memory, a random access memory, etc.) outside the chip in the management node.
[0046] Fourthly, embodiments of this application provide an apparatus, which may be a computing node as described in the foregoing embodiments, or a chip within the computing node. The apparatus may include a processing module and a transceiver module. When the apparatus is a computing node, the processing module may be a processor, and the transceiver module may be a transceiver; the computing node may further include a storage module, which may be a memory; the storage module is used to store instructions, and the processing module executes the instructions stored in the storage module to cause the computing node to perform the method of the second aspect or any embodiment of the second aspect. When the apparatus is a chip within the computing node, the processing module may be a processor, and the transceiver module may be an input / output interface, pin, or circuit, etc.; the processing module executes the instructions stored in the storage module to cause the computing node to perform the method of the second aspect or any embodiment of the second aspect. The storage module may be a storage module within the chip (e.g., a register, cache, etc.), or a storage module located outside the chip within the computing node (e.g., a read-only memory, random access memory, etc.).
[0047] Fifthly, this application provides an apparatus, which may be an integrated circuit chip. The integrated circuit chip includes a processor. The processor is coupled to a memory for storing programs or instructions that, when executed by the processor, cause the apparatus to perform the methods described in any of the embodiments of the foregoing aspects.
[0048] Sixthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in any of the foregoing embodiments.
[0049] In a seventh aspect, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the methods described in any of the preceding embodiments.
[0050] Eighthly, embodiments of this application provide a communication system, which includes a management node that performs the first aspect and any embodiment thereof, a computing node that performs the second aspect and any embodiment thereof, and an access network device.
[0051] Ninthly, embodiments of this application provide a communication system, which includes a management node that performs the first aspect and any embodiment of the first aspect, a computing node that performs the second aspect and any embodiment of the second aspect, an access network device, and a terminal device. Attached Figure Description
[0052] FIG. 1 is an example diagram of a network system to which a method for configuring a distributed AI task provided by the present application is applicable;
[0053] FIG. 2 is a flowchart of a method for configuring a distributed AI task provided by the present application;
[0054] FIG. 3A is an example diagram of a target model provided by the present application;
[0055] FIG. 3B is another example diagram of a target model provided by the present application;
[0056] FIG. 4 is another flowchart of a method for configuring a distributed AI task provided by the present application;
[0057] FIG. 5 is another flowchart of a method for configuring a distributed AI task provided by the present application;
[0058] FIG. 6 is a schematic diagram of an apparatus provided by the present application;
[0059] FIG. 7 is another schematic diagram of an apparatus provided by the present application. DETAILED DESCRIPTION
[0060] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments.
[0061] The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the terms thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to include only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to the process, method, product or device.
[0062] It should be understood that the term "and / or" in this document merely describes an associated relationship between associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be single or multiple. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects. In addition, "at least one of the following" or similar expressions in this document are used to represent any combination of the listed items; for example, at least one of A, B, and (or) C can represent the following six cases: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, B and C exist simultaneously, and A and C exist simultaneously, where A, B, and C can be single or multiple.
[0063] For ease of understanding, the system architecture to which the distributed AI task configuration method proposed in this application is applicable will be introduced first as follows:
[0064] The distributed AI task configuration method proposed in this application can be applied to the 5th generation mobile communication (5G) system, the 6th generation mobile communication technology (6G) system, and other subsequent evolution systems, which are not limited by this application.
[0065] FIG. 1 is an example diagram of a network system to which the distributed AI task configuration method provided by this application is applicable. As shown in FIG. 1, the network system at least includes a terminal device, an access network device, a computing node, and a management node.
[0066] The terminal device can be referred to as a terminal, a user equipment (UE), a wireless terminal device, a mobile terminal (MT) device, a subscriber unit, a subscriber station, a mobile station (MS), a mobile, a remote station, an access point (AP), a remote terminal, an access terminal, a user terminal, a user agent, or a user device, etc. In addition, the terminal device can be a mobile phone, a tablet, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc. It should be understood that embodiments of the present application do not limit the specific technology and specific device form of the terminal device. In the embodiments and subsequent embodiments, the terminal device is taken as an example for introduction.
[0067] The access network device can be any kind of device with wireless transceiver function, and can be used to be responsible for functions related to air interface, such as wireless link maintenance function, wireless resource management function, and part of mobility management function. The access network device can also be configured with a baseband unit (BBU) with baseband signal processing function. The access network device can be a radio access network (RAN) device (or RAN node) currently serving the terminal device. At present, some common examples of the access network device are: Node B (NB), evolved Node B (eNB), next generation Node B (gNB) in 5G new radio (NR) system, node (such as xNodeB) in 6G system, transmission reception point (TRP), radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), home base station (such as home evolved NodeB or home Node B (HNB)), and the like. In addition, in a network structure such as cloud radio access network (CloudRAN) or open radio access network (ORAN), the access network device can be a device including a centralized unit (CU) (also referred to as a control unit) and / or a distributed unit (DU). The RAN device including the CU and the DU splits the protocol layers of the gNB in the NR system, and the functions of part of the protocol layers are placed in the CU for centralized control, and the functions of the remaining part or all of the protocol layers are distributed in the DUs and controlled by the CU. Multiple DUs can share one CU. The splitting of the CU and the DU can be according to the protocol stack.For example, one possible split is that radio resource control (RRC), service data adaptation protocol (SDAP) and packet data convergence protocol (PDCP) layers are deployed in the CU, and the rest of the radio link control (RLC) layer, media access control (MAC) layer and physical (PHY) layer are deployed in the DU. The CU and the DU are connected through an F1 interface, the CU is connected with the core network through an NG interface on behalf of the gNB, and the CU is connected with other gNBs (or other CUs) through an Xn interface on behalf of the gNB. In actual deployment of a traditional RAN device, in addition to the logical gNB composed of the CU and the DU, the RAN device also includes an RU. The RU is a hardware unit containing part of the PHY layer function and / or antenna device. Optionally, the RU can be configured to be independent of the antenna device (for example, an antenna line device (ALD), which can also be integrated with the antenna device. For example, in a 5G NR system, the aforementioned RU can be an active antenna unit (AAU), that is, a processing unit integrated with a remote radio unit (RRU) (or a remote radio head (RRH)) and an antenna device. In some deployments, the RAN device mentioned in the embodiments of the present application can be a device including a CU or a DU; or the RAN device is a device including a CU and a DU; or the RAN device is a device including a control plane CU node (central unit-control plane (CU-CP)), a user plane CU node (central unit-user plane (CU-UP)) and a DU node. For example, the RAN device can include a gNB-CU-CP, a gNB-CU-UP and a gNB-DU. In other deployments, multiple RAN nodes cooperate to assist the terminal to implement wireless access, and different RAN nodes respectively implement part of the functions of the base station. For example, the RAN node can be a CU, a DU, a CU-CP, a CU-UP or an RU, etc. The CU and the DU can be separately arranged or can be included in the same network element, such as a BBU. The RU can be included in a radio frequency device or a radio frequency unit, such as an RRU, an AAU or an RRH.In one possible design, the processing unit in the BBU for implementing baseband functions is referred to as a base band high (BBH) unit, and the processing unit in the RRU / AAU / RRH for implementing baseband functions is referred to as a base band low (BBL) unit. It should be understood that the CU (or CU-CP and CU-UP), DU, or RU can also have different names in different systems, but those skilled in the art can understand their meanings. For example, in an ORAN system, the CU can also be referred to as an O-CU (open CU), the DU can also be referred to as an O-DU (open DU), the CU-CP can also be referred to as an O-CU-CP, the CU-UP can also be referred to as an O-CU-UP, and the RU can also be referred to as an O-RU. Any of the CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module. It should be understood that the access network device in the embodiments of this application can be any of the above devices or a chip in the above devices, and the specific implementation is not limited here. In this embodiment and subsequent embodiments, the access network device is mainly taken as an example for introduction.
[0068] The computing node, also referred to as an in-network computing node, is a node in a wireless network responsible for processing AI tasks. The computing node is configured with various AI resources, such as AI model resources, AI computing resources, and AI data resources. Optionally, the computing node is deployed with an AI Server module that performs AI task processing. Optionally, the computing node is also deployed with a task management function (TMF) responsible for local scheduling of computing tasks. Exemplarily, the computing node can be a node C (Node C) or a node I (Node I). The node C is a computing node in a mobile network that provides perception, computing, and other services. The node I is a dedicated node in a mobile network that provides data collection, model training, and management. Generally, the computing node is connected with an access network device, for example, the computing node is connected with a gNB; or the computing node is connected with a CU; or the computing node is connected with a DU; or the computing node is connected with both the CU and the DU, which is not limited in this application. One computing node can be connected with multiple access network devices (for example, gNB, CU, or DU), and provide AI task processing functions for the connected multiple access network devices. One access network device can also be connected with multiple computing nodes, so that the multiple computing nodes collectively provide AI task processing functions for the access network device. When there are multiple distributed computing nodes in a mobile network, the computing nodes communicate with each other through an access network device. The computing node in the embodiments of this application can be independent of the access network device, or can be integrated with the access network device, which is not limited in this application. In this embodiment and subsequent embodiments, the computing node is mainly taken as an example for introduction.
[0069] The management node is a newly added intelligent face network element in a mobile network, and is configured to respond to AI services and perform orchestration and configuration of distributed AI tasks. For example, the management node can determine appropriate computing nodes for an AI task initiated by an access network device or a third-party AI application, and distribute the AI task to each computing node, so that each computing node performs the AI task in a distributed manner. For example, the management node can be an AI management function (AIMF). The management node can be connected with the access network device, and the management node communicates with each computing node through the access network device; or, the management node and the computing node have a communication interface, and the management node directly communicates with each computing node. The management node in the embodiments of the present application can be deployed in an access network, for example, the management node can be independent of the access network device, or can be integrated with the access network device; the management node can also be deployed in a core network, for example, the management node is a network element in the core network responsible for AI management. The present application does not limit the specific implementation form of the management node.
[0070] In the system architecture shown in FIG. 1, the management node can receive a task request from an access network device or a terminal device (for example, a user terminal on which a third-party AI application is deployed), and select an AI model with a specific function and a computing node participating in the service based on the task request. Then, the management node sends the selected AI model to each computing node, so that the computing node deploys the AI model, and processes task data using the AI model, and then outputs the processing result of the AI task.
[0071] However, the number of AI models with specific functions that can be stored in the management node is limited, and the task processing strategy of the AI model with specific function is fixed. The management node may not be able to find a suitable AI model for the current AI task for deployment in each computing node participating in the service, thereby affecting the performance of each distributed computing node providing AI services.
[0072] To this end, the present application provides a distributed AI task configuration method, a management node and a computing node, which are configured to realize the generalization of AI models in the management node, so that the sub-models deployed by each computing node indicated by the management node for a task can constitute a target model suitable for processing the task, thereby improving the performance of each distributed computing node providing AI services.
[0073] The main process of the distributed AI task configuration method provided by the present application will be introduced below in conjunction with FIG. 2.
[0074] As shown in FIG. 2, a flowchart of a distributed AI task configuration method provided by the present application is shown. The distributed AI task configuration method is illustrated by taking the interaction between the management node and the computing node as an example. Of course, the subject performing the actions of the management node in the method can also be a device or module in the management node, such as a chip, processor or processing unit in the management node, etc. The subject performing the actions of the computing node in the method can also be a device or module in the computing node, such as a chip, processor or processing unit in the computing node, etc. The present application embodiment does not make specific limitations. The processing performed by a single execution subject (e.g., the management node or the computing node) in the present application embodiment can also be divided into processing performed by multiple execution subjects, which can be logically and / or physically separated. For example, as shown in FIG. 2, the distributed AI task configuration method includes the following steps:
[0075] In step 201, the management node receives a first task request.
[0076] The first task request is a request for applying for an AI service. Based on the different initiators of the AI service application, the first task request can be from an access network device or from a third-party application of a terminal device. In an implementation, the first task request is from the access network device. For example, the access network device sends the first task request to the management node to request an AI application service on the RAN side; correspondingly, the management node receives the first task request from the access network device. In another implementation, the first task request is from the terminal device. For example, the terminal device is configured with a client of a third-party application, and the terminal device sends the first task request to the management node through the access network device to request an AI service of the third-party application; correspondingly, the management node receives the first task request from the terminal device through the access network device.
[0077] The first task request includes task information of the first task. Optionally, the task information of the first task includes a task type of the first task and a task requirement of the first task.
[0078] The task type of the first task refers to what type of AI task the first task is, i.e., the task type of the AI task. In an implementation, if the first task request comes from an access network device, the task type of the AI task is related to a physical layer application or a high layer application of the access network device. For example, the task type of the first task can be any one of the AI tasks of the physical layer, such as beam management, channel prediction, or resource allocation; or can indicate any one of the AI tasks of the high layer, such as terminal trajectory prediction, terminal load balancing, or terminal energy saving. In another implementation, if the first task request comes from a third-party application of a terminal device, the task type of the AI task is related to the application type of the third-party application. For example, the task type of the first task can indicate the AI tasks of the type of navigation path planning or congestion prediction. It should be understood that in actual applications, the first task can also be other types of AI tasks, and the specific type of the first task is not limited in the present application.
[0079] In addition, the task requirement of the first task is used to indicate the requirement (or demand) of the distributed AI processing service requested by the first task for data processing and / or data transmission. The task requirement of the first task can be a QoS requirement or a performance requirement. Optionally, the task requirement of the first task includes a performance requirement of an AI model (for example, the task processing accuracy that the model can provide, the size of the model, the robustness, personalization, or generalization of the model, etc.) and a service efficiency requirement (for example, response delay, energy consumption, etc.). Optionally, the task requirement of the first task also includes an overhead requirement (for example, storage overhead, computing overhead, transmission overhead, etc.) for providing the AI service.
[0080] In step 202, the management node determines at least two sub-models and attribute information of the sub-models based on the task information of the first task and attribute information of a plurality of functional modules stored by the management node.
[0081] The management node stores a plurality of function modules, one function module can implement a specific AI processing function, and different function modules implement different AI processing functions. One function module includes at least one network layer, and different network layers have different functions or parameters. The network layer refers to one layer of a deep learning model constructed layer by layer, such as a neural network or a Transformer. For example, an image feature extraction function module includes at least a convolution layer and may include a pooling layer (for example, a max-pooling layer or an average-pooling layer). For another example, a classification function module generally includes a fully connected layer with an activation function (for example, a softmax layer or a sigmoid layer). It should be understood that the foregoing neural network can be a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), or other neural networks, and the present application is not limited.
[0082] In addition, one function module is an AI model with a specific AI processing function, or one function module is a part of an AI model with a specific AI processing function. A plurality of function modules with different functions can constitute an AI model for processing a specific AI task; accordingly, an AI model capable of processing a specific AI task can be divided into a plurality of different function modules according to functions.
[0083] Exemplarily, as shown in FIG. 3A, an example diagram of an image classification AI model listed by the present application is shown. In this example, the image classification AI model includes 5 functional modules, which are a low-level feature extraction module, a complex feature extraction module, an abstract feature extraction module, a feature fusion processing module, and a classification processing module. Among them, the low-level feature extraction module includes 2 convolution layers with large size (i.e., layer a1 and layer a2), 1 max-pooling layer (i.e., layer b1), 2 convolution layers with small size (i.e., layer a3 and layer a4); the complex feature extraction module includes 1 average-pooling layer (i.e., layer c1) and 3 convolution layers (i.e., layer a5, layer a6 and layer a7); the abstract feature extraction module includes 1 max-pooling layer (i.e., layer b2), 2 convolution layers (i.e., layer a8 and layer a9), 1 average-pooling layer (i.e., layer c2) and 1 max-pooling layer with small size (i.e., layer b3); the feature fusion processing module includes 2 full connected layers (i.e., layer d1 and layer d2); the classification processing module includes (i.e., layer e1 and layer e2). In this example, the low-level feature extraction module, the complex feature extraction module, the abstract feature extraction module, the feature fusion processing module, and the classification processing module constitute an image classification AI model.
[0084] In addition, the management node also stores attribute information of each functional module. The attribute information of the functional module includes a functional type of the functional module and a requirement of the functional module. Among them, the functional type of the functional module is used to indicate what kind of AI processing function is implemented by the functional module and the adaptation condition of the functional module for implementing the AI processing function. In addition, the requirement of the functional module is used to indicate the requirement (or demand) of data processing and / or data transmission when implementing the function of the functional module. In some scenarios, the requirement of the functional module can be referred to as the performance requirement or QoS requirement (e.g., computation-transmission QoS requirement) of the functional module.
[0085] Exemplarily, if the management node stores a feature extraction functional module, the attribute information of the feature extraction functional module includes a functional type of the feature extraction functional module and a performance requirement of the feature extraction functional module. Among them, the functional type of the feature extraction functional module indicates a feature extraction function, and the performance requirement of the feature extraction functional module indicates the requirement (or demand) of performance indicators (or QoS indicators) such as training time consumption, generalization performance, robustness, storage overhead, computation overhead, and energy consumption when implementing the feature extraction function.
[0086] Optionally, the attribute information of the function module further includes information of each network layer included in the function module (for example, data dimension, quantization information, pruning information, and other network layer parameters of the network layer), input dimension of the function module (i.e., dimension of data input to the first network layer of the function module), and output dimension of the function module (i.e., dimension of data output by the last network layer of the function module), etc. Taking the complex feature extraction module in the example shown in FIG. 3A as an example, the complex feature extraction module includes one mean pooling layer (i.e., layer c1) and three convolution layers (i.e., layer a5, layer a6, and layer a7), and the attribute information of the complex feature extraction module includes data dimension, quantization information, pruning information, and other network layer parameters of layer c1, layer a5, layer a6, and layer a7. In addition, the attribute information of the complex feature extraction module further includes input dimension of the complex feature extraction module (i.e., dimension of data input to layer c1) and output dimension of the complex feature extraction module (i.e., dimension of data output by layer a7).
[0087] Optionally, the attribute information of the function module further includes compatibility of the computing node, which is used to indicate requirements of the function module on hardware or software of the computing node on which part or all of the network layers of the function module are deployed.
[0088] In this step, after receiving the task information of the first task, the management node can determine, based on the task information of the first task, what type of AI task the first task is and performance requirement (or QoS requirement) of the AI task, and then, based on the task information of the first task and the attribute information of the plurality of function modules stored by the management node, the management node screens at least one function module from the plurality of function modules to constitute at least two sub-models for processing the first task in a distributed manner, and determines attribute information of the sub-models based on the attribute information of the screened function modules.
[0089] The sub-model refers to a model allocated to a certain distributed computing node for processing task data. One sub-model can include at least one function module, or can include a part of one function module. The attribute information of the sub-model is used for the distributed computing node to deploy the sub-model. The attribute information of the sub-model is related to the attribute information of the function modules constituting the sub-model. The following will be introduced respectively:
[0090] In a possible implementation, one sub-model includes at least one functional module. The attribute information of the sub-model is related to the attribute information of the at least one functional module constituting the sub-model. For example, if sub-model 1 is constituted by functional module A and functional module B, and the sub-model 1 can implement function A and function B, the attribute information of the sub-model 1 includes the attribute information of functional module A and the attribute information of functional module B. Taking FIG. 3A as an example, sub-model 2 includes a complex feature extraction module and an abstract feature extraction module. The attribute information of the sub-model 2 includes the information of each network layer constituting the complex feature extraction module and the information of each network layer constituting the abstract feature extraction module, that is, the related parameters of the mean pooling layer (i.e., layer c1) and the convolution layer (i.e., layers a5, a6, and a7) in the complex feature extraction module, and the related parameters of the max pooling layer (i.e., layer b2), the convolution layer (i.e., layers a8 and a9), the mean pooling layer (i.e., layer c2), and the smaller-size max pooling layer (i.e., layer b3) in the abstract feature extraction module.
[0091] In another possible implementation, one sub-model includes a part of one functional module. The attribute information of the sub-model is related to all or part of the attribute information of the functional module constituting the sub-model. For example, sub-model 2 can be constituted by a part of functional module C, and sub-model 3 is constituted by another part of functional module C, and sub-model 2 and sub-model 3 jointly implement function C. In this example, the attribute information of sub-model 2 includes all the attribute information of functional module C or the attribute information related to the part of functional module C constituting sub-model 2, and the attribute information of sub-model 3 includes all the attribute information of functional module C or the attribute information related to the other part of functional module C constituting sub-model 3. Taking FIG. 3B as an example, sub-model 1 includes a part of the low-level feature extraction module. The attribute information of sub-model 1 can include the information of each network layer constituting the low-level feature extraction module and belonging to sub-model 1, for example, the information of layers a1, a2, b1, and a3; or the attribute information of sub-model 1 can include the information of all the network layers constituting the low-level feature extraction module, for example, the information of layers a1, a2, b1, a3, and a4.
[0092] It should be understood that, among the at least two sub-models determined by the management node based on the screened at least one functional module, a part of the sub-models can include at least one functional module, another part of the sub-models can include only a part of one functional module, and another part of the sub-models can include at least one functional module and a part of one functional module, which is not limited in the present application.
[0093] Further, the attribute information of the sub-models is used to indicate the structure and function of the sub-models. The structure of a sub-model refers to the multiple network layers included in the sub-model and the order between the multiple network layers. Specifically, the attribute information of a sub-model includes the information of each network layer included in the sub-model, and the information of a network layer includes the data input dimension and the data output dimension of the network layer. Optionally, the attribute information of a sub-model further includes the function of the network layers included in the sub-model, i.e., which network layers in the sub-model are part of which function module. Taking the sub-model 3 in the example shown in FIG. 3A as an example, the sub-model 3 includes the mean pooling layer (i.e., layer c2) and the max pooling layer (i.e., layer b3) in the abstract feature extraction module, the two fully connected layers (i.e., layer d1 and layer d2) in the feature fusion processing module, and the two classification layers (i.e., layer e1 and layer e2) in the classification processing module. The attribute information of the sub-model 3 includes the information of the aforementioned network layers (i.e., the data input dimension and the data output dimension of each network layer in the layers c2, b3, d1, d2, e1 and e2), and the function modules to which the network layers included in the sub-model 3 belong (i.e., the layers c2 and b3 belong to the abstract feature extraction module, the layers d1 and d2 belong to the feature fusion processing module, and the layers e1 and e2 belong to the classification processing module).
[0094] Further, the at least two sub-models determined by the management node are distributedly deployed in the at least two computing nodes, one sub-model is deployed in one computing node, and different sub-models are deployed in different computing nodes. The at least two sub-models deployed in the at least two computing nodes constitute the target model for processing the first task. For example, if the management node determines the sub-model 1, the sub-model 2 and the sub-model 3, and determines that the sub-model 1, the sub-model 2 and the sub-model 3 are deployed in the computing node 1, the computing node 2 and the computing node 3 respectively, the aforementioned sub-model 1, sub-model 2 and sub-model 3 jointly constitute the target model for processing the first task.
[0095] Specifically, the management node first determines the target model for processing the first task based on the task information of the first task and the attribute information of the multiple function modules, and then determines the at least two sub-models and the attribute information of each sub-model based on the target model, so that the at least two sub-models are distributedly deployed in the at least one computing node. For example, the management node screens at least one function module from the multiple function modules based on the task information of the first task and the attribute information of the multiple function modules to constitute the target model, so that the target model can process the first task; then, the management node splits the target model into at least two sub-models, so that the at least two sub-models are distributedly deployed in the at least two computing nodes, thereby realizing joint processing of the first task by the sub-models located in different computing nodes.
[0096] It should be noted that when the management node splits the target model into at least two sub-models, the management node can split the target model according to the granularity of the functional modules, for example, one sub-model includes at least one functional module; the management node can also split the target model according to the granularity of the network layers, for example, one sub-model includes at least one network layer of a functional module. For details, please refer to the foregoing description of the sub-model, which will not be repeated here.
[0097] It should also be noted that when the management node splits the target model into at least two sub-models, the sub-models determined by the management node can include partially redundant network layers, so as to improve the reliability and flexibility of the sub-models, and facilitate subsequent functional enhancement or dynamic adjustment of the sub-models. Specifically, the management node can divide the same network layer into different sub-models, i.e., different sub-models contain the same at least one network layer. For example, the management node determines that two sub-models having a connection relationship contain the same network layer of the same functional module. For example, as shown in FIG. 3A, the management node determines that the target model for processing the first task includes a low-level feature extraction module, a complex feature extraction module, an abstract feature extraction module, a feature fusion processing module, and a classification processing module. Then, the management node splits the target model into three sub-models, i.e., a sub-model 1 containing the low-level feature extraction module and the complex feature extraction module, a sub-model 2 containing the complex feature extraction module and the abstract feature extraction module, and a sub-model 3 containing the feature fusion module, the classification processing module, and a part of the abstract feature extraction module. Among them, the sub-model 1 and the sub-model 2 both contain four network layers (i.e., layers c1, a5, a6, and a7) of the complex feature extraction module, and the sub-model 2 and the sub-model 3 both contain the network layer c2 and the network layer b3 in the abstract feature extraction module. In the present embodiment, since the two sub-models having a connection relationship have the same network layer, when the two sub-models are respectively deployed on different computing nodes, the management node can enable the network layer of the sub-model based on the resource state of the computing node on which the sub-model is deployed, thereby facilitating the flexibility of the sub-model deployed on the computing node.
[0098] It should also be noted that when the management node splits the target model into at least two sub-models, the sub-models determined by the management node can not include redundant network layers, i.e., different sub-models do not contain the same network layer, so as to save the signaling overhead of transmitting the sub-models and the attribute information of the sub-models. For example, as shown in FIG. 3B, the management node determines that the target model for processing the first task includes a low-level feature extraction module, a complex feature extraction module, an abstract feature extraction module, a feature fusion processing module, and a classification processing module. Then, the management node splits the target model into three sub-models, i.e., a sub-model 1 including part of the network layer of the low-level feature extraction module, a sub-model 2 including the complex feature extraction module and the abstract feature extraction module, and a sub-model 3 including the feature fusion module and the classification processing module. Among them, layer a3 in the sub-model 1 is connected with layer a4 in the sub-model 2, layer b3 in the sub-model 2 is connected with layer d1 in the sub-model 3, and there is no same network layer between the sub-model 1, the sub-model 2, and the sub-model 3. In this embodiment, there is no same network layer between the two sub-models having the connection relationship, and when the management node transmits the aforementioned two sub-models and the attribute information of the sub-models to at least two computing nodes, the management node can save the signaling overhead of transmitting the sub-models and the attribute information of the sub-models.
[0099] Optionally, the attribute information of the sub-model further includes the network layer position of the sub-model relative to the target model. It can be understood that the network layers included in the sub-model are which network layers among the multiple network layers included in the target model; or it can be understood that the network layers included in the sub-model are connected with which network layers in the target model. For example, taking the sub-model 2 shown in FIG. 3A as an example, the network layer position of the sub-model 2 relative to the target model is {a4, c1; b3, d1}, which means that layer c1 in the sub-model 2 is connected with layer a4 in the target model, i.e., the output of layer a4 in the target model can be used as the input of layer c1 in the sub-model 2; layer b3 in the sub-model 2 is connected with layer d1 in the target model, i.e., the output of layer b3 in the sub-model 2 can be used as the input of layer d1 in the target model.
[0100] It should be noted that in the prior art, the management node generally stores a backbone model, i.e., a complete model for performing an AI task. Since the backbone model stored by the management node in the prior art is limited, and the management node can select a limited number of backbone models for the currently received AI task, the prior art may consider the matching of the model functions, but ignore the matching of the performance requirements of the task and the performance requirements of the model. In this embodiment, the management node can filter the functional modules that can meet the task requirements of the first task based on the task requirements of the first task and the performance requirements of the functional modules, and then ensure that the target model composed of the multiple functional modules can also meet the task requirements of the first task, thereby improving the matching rate of the management node in configuring the task and improving the performance of the AI processing service.
[0101] In addition, in addition to determining the at least two sub-models for constituting the target model, the management node also needs to determine in which computing nodes the at least two sub-models are respectively deployed, so as to ensure that the computing nodes receiving the sub-model and the attribute information of the sub-model have sufficient AI resources to deploy the sub-model and establish a connection relationship with the sub-models deployed in other computing nodes.
[0102] In a possible implementation, the management node obtains resource state information of the at least one computing node, the resource state information being used to indicate a use state of AI resources of the computing node. The AI resources include model resources (for example, AI models already deployed in the computing node), computing resources (for example, idle computing resources of the computing node), and data resources (for example, data used for training or inference in the computing node). Then, the management node determines the sub-model to be deployed in each computing node based on the resource state information of the at least one computing node. Optionally, the AI resources of the computing node determined by the management node to deploy the sub-model need to meet the AI resource requirement of the sub-model to be deployed. For example, if the management node determines to deploy the sub-model 1 in the computing node 1, the computing node 1 can provide the model resources for deploying the sub-model 1, for example, the AI model already deployed in the computing node 1 can be compatible with the sub-model 1; or the computing node 1 can provide the computing resources for deploying the sub-model 1, for example, the computing node for deploying the sub-model 1 needs 1G of computing resources, and the computing node 1 has at least 1G of idle computing resources for deploying the sub-model 1; or the computing node 1 can provide the data resources for deploying the sub-model 1, for example, the sub-model 1 needs to use training data of a specific size or a specific type during running, and the computing node 1 stores the training data of the specific size or the specific type.
[0103] After the management node determines the sub-model to be deployed in each computing node and the attribute information of the sub-model, the management node sends the sub-model corresponding to each computing node and the attribute information of the sub-model to the computing node. Specifically, the management node performs step 203.
[0104] In step 203, the management node sends the sub-model to be deployed in each computing node and the attribute information of the sub-model to the at least two computing nodes; correspondingly, the at least two computing nodes receive the sub-model to be deployed in each computing node and the attribute information of the sub-model from the management node.
[0105] For example, the management node determines the first sub-model and the second sub-model, and determines the first computing node deploying the first sub-model and the second computing node deploying the second sub-model, the management node sends the first sub-model and the attribute information of the first sub-model to the first computing node, and sends the second sub-model and the attribute information of the second sub-model to the second computing node.
[0106] At step 204, the at least two computing nodes respectively deploy the received sub-models.
[0107] Specifically, the at least two computing nodes deploy the received sub-models based on the attribute information of the received sub-models. For example, after the first computing node receives the first sub-model and the attribute information of the first sub-model, the first computing node deploys the first sub-model based on the attribute information of the first sub-model, and the second computing node deploys the second sub-model based on the attribute information of the second sub-model.
[0108] It should be noted that after the computing node deploys the sub-model, the sub-model needs to be further configured to jointly implement the distributed processing of the first task with the sub-models deployed by other computing nodes.
[0109] At step 205, the management node sends configuration information of the sub-model deployed by the computing node to the at least two computing nodes; correspondingly, the at least two computing nodes receive the configuration information of the sub-model deployed by the computing node from the management node.
[0110] The configuration information of the sub-model is used to configure the model parameters used by the sub-model when participating in processing the first task and establish the connection relationship between the sub-model and the sub-model deployed in different computing nodes. Optionally, the configuration information of the sub-model is also used to configure the interaction rules and transmission mode of the inter-node AI data.
[0111] Optionally, the configuration information of the sub-model includes the parameter of the enabled network layer, which is used to indicate the network layer enabled by the sub-model deployed in the computing node when participating in processing the first task. The network layers enabled by the at least two sub-models constitute the target model. The parameter of the enabled network layer includes: the information of the starting network layer of the sub-model (i.e. the network layer of the data input end in the multiple network layers enabled by the sub-model), the information of the exit network layer of the sub-model (i.e. the network layer of the data output end in the multiple network layers enabled by the sub-model), and the information of the connection between the starting network layer and the exit network layer in the sub-model.
[0112] For example, referring to FIG. 3A, the management node determines that sub-model 1, sub-model 2 and sub-model 3 are deployed in computing node 1, computing node 2 and computing node 3 respectively, and the network layers enabled by sub-model 1, sub-model 2 and sub-model 3 can exactly constitute the target model shown in FIG. 3A. In an example, if the management node determines that sub-model 1 enables layer a1, layer a2, layer b1, layer a3 and layer a4, determines that sub-model 2 enables layer c1, layer a5, layer a6, layer a7, layer b2, layer a8 and layer a9, and determines that sub-model 3 enables layer c2, layer b3, layer d1, layer d2, layer e1 and layer e2, the configuration information of sub-model 1 includes the parameters of each network layer enabled by sub-model 1, i.e., the parameters of layer a1, layer a2, layer b1, layer a3 and layer a4, and indicates that layer a1 is the starting network layer of sub-model 1 and layer a4 is the exit network layer of sub-model 1; the configuration information of sub-model 2 includes the parameters of each network layer enabled by sub-model 2, i.e., the parameters of layer c1, layer a5, layer a6, layer a7, layer b2, layer a8 and layer a9, and indicates that layer c1 is the starting network layer of sub-model 2 and layer a9 is the exit network layer of sub-model 2; and the configuration information of sub-model 3 includes the parameters of each network layer enabled by sub-model 3, i.e., the parameters of layer c2, layer b3, layer d1, layer d2, layer e1 and layer e2, and indicates that layer c2 is the starting network layer of sub-model 3 and layer e2 is the exit network layer of sub-model 3.
[0113] Optionally, the configuration information of the sub-model further includes the connection relationship between the sub-model and the sub-model deployed in other computing nodes. For example, the connection relationship between the sub-model and the sub-model deployed in other computing nodes includes the information of the next network layer (i.e., the network layer connected with the exit network layer of the sub-model and the computing node where the network layer is located), and / or the information of the previous network layer (i.e., the network layer connected with the starting network layer of the sub-model and the computing node where the network layer is located).
[0114] For example, still taking FIG. 3A as an example, sub-model 1 enables {a1, layer a2, layer b1, layer a3, layer a4}, sub-model 2 enables {layer c1, layer a5, layer a6, layer a7, layer b2, layer a8, layer a9}, and sub-model 3 enables {layer c2, layer b3, layer d1, layer d2, layer e1, layer e2}. Sub-model 1 deployed on computing node 1 is connected to sub-model 2 deployed on computing node 2, and sub-model 2 deployed on computing node 2 is connected to sub-model 3 deployed on computing node 3. In this example, the configuration information of sub-model 1 further includes information of computing node 2 and information of layer c1 in sub-model 2, indicating that the exit network layer (i.e., layer a4) of sub-model 1 is connected to the next network layer (i.e., layer c1). The configuration information of sub-model 2 further includes information of computing node 1 and information of layer a4 in sub-model 1, indicating that the start network layer (i.e., layer c1) of sub-model 2 is connected to the previous network layer (i.e., layer a4); and information of computing node 3 and information of layer c2 in sub-model 3, indicating that the exit network layer (i.e., layer a9) of sub-model 2 is connected to the next network layer (i.e., layer c2). The configuration information of sub-model 3 further includes information of computing node 2 and information of layer a9 in sub-model 2, indicating that the start network layer (i.e., layer c2) of sub-model 3 is connected to the previous network layer (i.e., layer a9).
[0115] Optionally, the configuration information of the sub-model is further used to configure a task processing strategy of the sub-model when participating in processing the first task. It can be understood that the task processing strategy of the sub-model when participating in processing the first task is jointly determined by the sub-model and other sub-models (for example, the sub-model deployed on other computing nodes in the at least two computing nodes determined by the management node).
[0116] For example, the configuration information of the sub-model includes at least one of the following task processing strategies:
[0117] A sub-model strategy, the sub-model strategy is used to indicate the model parameters used by the sub-model when participating in processing the first task. For example, the sub-model strategy includes a network layer enabling parameter, indicating which network layers in the plurality of network layers included in the sub-model are enabled; an exit point parameter, indicating which network layers in the network layers included in the sub-model perform an exit operation, and a task function or granularity corresponding to each exit point; and a data dimension parameter of the sub-model, indicating the data dimension of the input and / or output of the network layer included in the sub-model.
[0118] A data filtering strategy, the data filtering strategy is used to indicate the strategy of the sub-model for filtering the received data. The received data includes auxiliary data provided by the base station and / or the terminal device, training data or inference data preconfigured by the computing node, and task data of the first task. For example, the data filtering strategy includes a filtering strategy for the auxiliary data provided by the base station and / or the terminal device, a sampling rule for the training data, a filtering rule and a preprocessing rule for the auxiliary inference data, and the like.
[0119] The sub-model interaction strategy is used to indicate the interaction strategy between the sub-model and the sub-model deployed in other computing nodes. For example, the sub-model interaction strategy includes the model interaction strategy of the sub-model and other sub-models, the data interaction strategy of the sub-model and other sub-models, the regularization term interaction strategy of the sub-model and other sub-models, and the proportion configuration of various interaction strategies between the sub-models.
[0120] In this embodiment, the management node can determine the configuration information of the sub-model deployed by the computing node through any one of the following implementation manners:
[0121] In one type of implementation manner, the management node determines the configuration information of the sub-model deployed by each computing node based on the task information of the first task.
[0122] In another type of implementation manner, the management node can obtain the network state information of the initiator (for example, an access network device or a terminal device) of the first task, and then determines the configuration information of the sub-model deployed by the computing node based on the task information of the first task and the network state information of the initiator of the first task.
[0123] In one implementation, the management node receives first network state information from the access network device, the first network state information being used to indicate the air interface state of the access network device connected to the computing node. Then, the management node determines the configuration information of the sub-model to be deployed by the computing node based on the resource state information of at least one computing node and the first network state information.
[0124] In another implementation, the management node receives second network state information from the access network device, the second network state information being used to indicate the air interface state of the terminal device sending the first task request. Then, the management node determines the configuration information of the sub-model to be deployed by the computing node based on the resource state information of at least one computing node, the first network state information and the second network state information.
[0125] It should be understood that the configuration information of different sub-models determined by the management node is not completely the same. For example, if the management node determines two sub-models, i.e., a first sub-model and a second sub-model, which need to be respectively deployed on a first computing node and a second computing node, the management node sends, to the first computing node, configuration information of the first sub-model deployed on the first computing node, the configuration information of the first sub-model enabling at least one network layer in the first sub-model; and the management node sends, to the second computing node, configuration information of the second sub-model deployed on the second computing node, the configuration information of the second sub-model enabling at least one network layer in the second sub-model. Since the first sub-model and the second sub-model are different sub-models, the configuration information of the first sub-model is not completely the same as the configuration information of the second sub-model, and the network layer enabled in the first sub-model is different from the network layer enabled in the second sub-model, but the network layer enabled in the first sub-model and the network layer enabled in the second sub-model have a connection relationship.
[0126] In step 206, the computing node configures the sub-model and establishes a connection relationship between the sub-model and the sub-model deployed on the different computing nodes.
[0127] Specifically, the computing node configures, based on the received configuration information of the sub-model, a model parameter used by the sub-model when participating in processing the first task and establishes a connection relationship between the sub-model and the sub-model deployed on the different computing nodes.
[0128] For example, the computing node enables, based on the parameter of the network layer enabled in the configuration information, a network layer used by the sub-model when participating in processing the first task, and then the computing node establishes, based on the information of the next network layer and / or the information of the previous network layer of the sub-model, a connection relationship between the sub-model and the sub-model deployed on other computing nodes, so that the sub-models deployed on different computing nodes can interact with each other.
[0129] Optionally, the computing node can also configure, based on the configuration information of the sub-model, a task processing strategy of the sub-model when participating in processing the first task. For example, the computing node configures, based on a data screening strategy, a screening strategy of the sub-model for the received data. For another example, the computing node configures, based on a sub-model interaction strategy, an interaction strategy between the sub-model and the sub-model deployed on other computing nodes.
[0130] In this embodiment, the management node can determine the sub-models deployed into each computing node and the attribute information for deploying each sub-model based on the task information and the information of the plurality of function modules stored by the management node, and the management node can also send the configuration information of the sub-models to the computing nodes based on the task information to configure the task processing strategy of the sub-models when participating in processing the first task. Since the plurality of function modules stored by the management node can determine the sub-models for constituting the target model, and the management node can send the configuration information of the sub-models to the computing nodes for deploying each sub-model based on the task information to configure the task processing strategy of the sub-models when participating in processing the first task, the management node realizes the generalization of the AI model through the stored function modules, so that the management node instructs each computing node to deploy the sub-models that can constitute the target model suitable for processing the task, thereby improving the performance of the AI service provided by each distributed computing node.
[0131] The following will introduce the scenarios of the management node configuring the distributed AI task initiated by the access network device and the scenarios of the management node configuring the distributed AI task initiated by the third-party application respectively.
[0132] As shown in FIG. 4, it is a flowchart of one embodiment of the configuration method of the distributed AI task provided by the present application. In this embodiment, the first task is an AI task initiated by an access network device, in which case the management node, the computing node and the access network device will perform the following steps:
[0133] Step 401, the access network device sends a first task request to the management node; correspondingly, the management node receives the first task request from the access network device.
[0134] The first task request includes the task information of the first task. The task information of the first task includes the task type of the first task and the task demand of the first task. For the task type of the first task and the task demand of the first task, please refer to the foregoing step 201, which will not be described here.
[0135] In this embodiment, the first task is an AI task initiated by an access network device. For example, the first task can be any one of the AI tasks of the physical layer such as beam management, channel prediction or resource allocation; or can be any one of the AI tasks of the high layer such as terminal trajectory prediction, terminal load balancing or terminal energy saving. It should be understood that in actual application, the first task can also be other types of AI tasks, which will not be described here.
[0136] Step 402, the management node determines a target model based on the task type, the task demand and the attribute information of the plurality of function modules stored by the management node.
[0137] In a possible implementation, the management node determines a target model supporting processing the first task based on the task type, the task requirement, and attribute information of the plurality of functional modules. The target model is composed of at least one functional module in the plurality of functional modules, the task type of the first task is used to determine a model type of the target model, a performance requirement of the target model meets the task requirement of the first task, and the performance requirement of the target model is determined based on the requirement of the at least one functional module. For example, the management node determines the model type of the target model based on the task type of the first task, and then determines the target model composed of the plurality of at least one functional module meeting the task requirement of the first task based on the task requirement and the requirement of the plurality of functional modules. In this implementation, the management node can determine the target model meeting the task requirement of the first task based on the task type, the task requirement, and the requirement of the functional module, which is conducive to ensuring that the target model determined by the management node can meet the AI service requirement of the first task. In addition, the management node can compose different backbone models by splicing the plurality of functional modules for different tasks, which is conducive to providing AI services with different performance requirements for different tasks.
[0138] It should be noted that the plurality of functional modules stored by the management node can be preconfigured or generated by the management node based on a preconfigured backbone model. The following will be introduced respectively:
[0139] In a possible implementation, the management node receives at least one functional module and attribute information of the functional module, and the attribute information of the functional module is used to describe the functional type, the performance requirement, and the like of the functional module. For the functional module and the attribute information of the functional module, please refer to the related description in the foregoing step 202, which will not be repeated here.
[0140] In another possible implementation, the management node receives at least one backbone model and attribute information of the backbone model, then determines a plurality of function modules based on the at least one backbone model, and determines attribute information of each function module based on the attribute information of the backbone model. The backbone model is a pre-trained artificial intelligence / machine learning (AI / ML) model, or an initial AI / ML model. One backbone model can be split into a plurality of function modules according to functions, one function module is used to implement a subfunction of the functions of the backbone model, and different function modules are used to implement different subfunctions of the functions of the backbone model or different precisions / granularities of the same subfunction. In addition, the attribute information of the backbone model is used to indicate the functions and structure of the backbone model. In this implementation, the management node can split the received backbone model into function modules according to functions for storage, instead of directly storing the backbone model, which not only saves the storage resources of the management node, but also facilitates the management node to generate submodels by using the function modules, and improves the diversity of the submodels that can be generated by the management node. For the deployment node, various types of distributed AI paradigms can be flexibly supported through interconnection between the function modules.
[0141] In step 403, the management node obtains resource state information of at least two computing nodes.
[0142] The resource state information of the computing node is used to indicate the use state of the AI resource of the computing node. For details about the resource state information of the computing node, refer to step 202 in the foregoing description, which will not be repeated here.
[0143] In an example, the management node sends state reporting indication information to at least two computing nodes managed by the management node, the state reporting indication information being used to instruct the computing nodes to send the resource state information of the computing nodes to the management node.
[0144] In another example, the computing node is preconfigured with a resource state reporting rule. The computing node can periodically send the resource state information of the computing node to the management node, or send updated resource state information of the computing node to the management node when the resource state changes.
[0145] In step 404, the management node determines at least two submodels, attribute information of the submodels, and at least two computing nodes for deploying the at least two submodels based on the resource state information of the computing nodes and the target model.
[0146] Specifically, the management node determines how to split the target model into at least two sub-models, in which computing nodes the at least two sub-models need to be deployed, and attribute information of each of the at least two sub-models based on the received resource state information of the at least two computing nodes. In which, the AI resources of the computing nodes determined by the management node to deploy the sub-models need to meet the AI resource requirements of the sub-models to be deployed.
[0147] In an implementation, the management node first splits the target model into at least two sub-models, and then determines at least two computing nodes capable of deploying the at least two sub-models based on the received resource state information of the at least two computing nodes. For example, the management node determines to split the target model into three sub-models, and then selects three computing nodes from the at least two computing nodes based on the resource state information of the at least two computing nodes to deploy the three sub-models respectively. If the management node determines to deploy sub-model 1 to computing node 1, sub-model 2 to computing node 2, and sub-model 3 to computing node 3, the AI resources of computing node 1 need to meet the AI resource requirements of sub-model 1, the AI resources of computing node 2 need to meet the AI resource requirements of sub-model 2, and the AI resources of computing node 3 need to meet the AI resource requirements of sub-model 3.
[0148] In another implementation, the management node first selects multiple idle computing nodes from the at least two computing nodes based on the received resource state information of the at least two computing nodes, and then determines to split the target model into at least two sub-models based on the number of idle computing nodes, in which the number of sub-models is less than or equal to the number of idle computing nodes. For example, the management node selects two idle computing nodes (e.g., computing node 1 and computing node 2) from the at least two computing nodes based on the received resource state information of the at least two computing nodes, and then splits the target model into two sub-models (e.g., sub-model 1 and sub-model 2). If the management node determines to deploy sub-model 1 to computing node 1 and sub-model 2 to computing node 2, the AI resources of computing node 1 need to meet the AI resource requirements of sub-model 1, and the AI resources of computing node 2 need to meet the AI resource requirements of sub-model 2.
[0149] It should be noted that after determining the at least two sub-models based on the target model, the management node will also determine the attribute information of the sub-models based on the attribute information of the functional modules constituting the target model. For the attribute information of the functional modules and the attribute information of the sub-models, please refer to the foregoing step 202, which will not be described here.
[0150] Hereinafter, the management node is taken as an example to determine to split the target model into two sub-models, i.e., a first sub-model and a second sub-model, which are to be deployed on a first computing node and a second computing node, respectively. Then, the management node sends the sub-model to be deployed on each computing node to the two computing nodes, respectively. Specifically, the management node performs steps 405a and 405b.
[0151] In step 405a, the management node sends the first sub-model and attribute information of the first sub-model to the first computing node. Correspondingly, the first computing node receives the first sub-model and the attribute information of the first sub-model from the management node.
[0152] In step 405b, the management node sends the second sub-model and attribute information of the second sub-model to the second computing node. Correspondingly, the second computing node receives the second sub-model and the attribute information of the second sub-model from the management node.
[0153] Optionally, the management node sends the target model and attribute information of the target model to the first computing node. For example, the management node sends the target model, the attribute information of the target model, the first sub-model, and the attribute information of the first sub-model to the first computing node in one message. For another example, the management node sends the target model, the attribute information of the target model, the second sub-model, and the attribute information of the second sub-model to the second computing node in one message.
[0154] The attribute information of the target model includes information of network layers included in each of the at least two sub-models and connection relationships between network layers of different sub-models. When the first computing node obtains the attribute information of the target model and the attribute information of the first sub-model, the first computing node can determine which network layers in the target model the first sub-model is composed of. Similarly, when the second computing node obtains the attribute information of the target model and the attribute information of the second sub-model, the second computing node can determine which network layers in the target model the second sub-model is composed of.
[0155] In step 406a, the first computing node deploys the first sub-model.
[0156] Specifically, the first computing node deploys the first sub-model based on the attribute information of the first sub-model.
[0157] In step 406b, the second computing node deploys the second sub-model.
[0158] Specifically, the second computing node deploys the second sub-model based on the attribute information of the second sub-model.
[0159] In step 407, the access network device sends first network state information to the management node. Correspondingly, the management node receives the first network state information from the access network device.
[0160] The first network state information is used to indicate air interface states of the access network devices connected to the computing nodes. For example, an air interface state of a first access network device connected to a first computing node and an air interface state of a second access network device connected to a second computing node.
[0161] At step 408, the management node determines configuration information of the sub-models deployed by each computing node based on the resource state information of the at least two computing nodes and the first network state information.
[0162] The at least two computing nodes include the first computing node and the second computing node. For example, the management node determines configuration information of a first sub-model deployed by the first computing node and configuration information of a second sub-model deployed by the second computing node based on the resource state information of the first computing node, the resource state information of the second computing node, and the first network state information. The configuration information of the first sub-model is used to configure model parameters used by the first sub-model when participating in processing the first task, and the configuration information of the second sub-model is used to configure model parameters used by the second sub-model when participating in processing the first task. In addition, the configuration information of the first sub-model is also used to indicate a connection relationship between the first sub-model and the second sub-model deployed on the second computing node, and the configuration information of the second sub-model is also used to indicate a connection relationship between the second sub-model and the first sub-model deployed on the first computing node.
[0163] The configuration information of the sub-models is explained in the foregoing step 205, which will not be repeated here.
[0164] Optionally, the management node also determines configuration information of at least one participating node. The participating node is a node having a connection relationship with the computing node on which the sub-model is deployed. For example, the participating node is a node having a connection relationship with the first computing node on which the first sub-model is deployed and / or the second computing node on which the second sub-model is deployed. The participating node does not deploy a sub-model or the sub-model deployed by the participating node does not participate in processing the first task, but only forwards data between the first computing node and / or the second computing node. The configuration information of the participating node is used to indicate a data interaction strategy between the participating node and the computing node.
[0165] At step 409a, the management node sends the configuration information of the first sub-model to the first computing node. Correspondingly, the first computing node receives the configuration information of the first sub-model from the management node.
[0166] In a possible implementation, the configuration information of the first sub-model comprises first configuration information of the first sub-model. The first configuration information comprises first model parameters, the first model parameters being used to indicate enabling a first network layer in the first sub-model, the first network layer being at least one network layer in a plurality of network layers indicated by the attribute information of the first sub-model. It can be understood that the first network layer is at least one network layer enabled in the plurality of network layers deployed in the first computing node. The management node sends the first configuration information of the first sub-model to the first computing node in the first configuration process, to indicate which network layers and which connection relationships are enabled by the first sub-model.
[0167] Optionally, the management node sends, to the at least one participant node, configuration information of the participant node, the configuration information of the participant node being used to indicate a data interaction strategy between the participant node and the computing node.
[0168] Step 409b, the management node sends, to the second computing node, configuration information of the second sub-model; correspondingly, the second computing node receives, from the management node, the configuration information of the second sub-model.
[0169] Step 409b is similar to step 409a, and details can be referred to the related description in step 409a, which will not be repeated here.
[0170] Step 410a, the first computing node configures a task processing strategy of the first sub-model.
[0171] After receiving the configuration information of the first sub-model, the first computing node configures model parameters used when participating in processing the first task based on the configuration information of the first sub-model, and establishes a connection relationship between the first sub-model and the second sub-model deployed in the second computing node.
[0172] Step 410b, the second computing node configures a task processing strategy of the second sub-model.
[0173] After receiving the configuration information of the second sub-model, the second computing node configures model parameters used when participating in processing the first task based on the configuration information of the second sub-model, and establishes a connection relationship between the second sub-model and the first sub-model deployed in the first computing node.
[0174] Step 411a, the access network device sends, to the first computing node, task data of the first task.
[0175] Step 411b, the access network device sends, to the second computing node, task data of the first task.
[0176] It should be noted that the access network device can send first data requiring processing of the first sub-model to the first computing node and send second data requiring processing of the second sub-model to the second computing node, respectively, where the first data and the second data are both task data of the first task. The access network device can also send the first data and the second data to one of the computing nodes (for example, the first computing node) and forward the first data and the second data to the other computing node (for example, the second computing node) by the one computing node.
[0177] At step 412, the first computing node and the second computing node jointly process the first task by the first sub-model and the second sub-model and output a processing result of the first task.
[0178] For example, the first task can be beam management, channel prediction, and resource allocation of a physical layer, or terminal trajectory prediction, load balancing, and energy saving of a higher layer. Taking the first task as channel prediction of a physical layer as an example, the first computing node and the second computing node process the received task data of the first task by the first sub-model and the second sub-model, respectively, to obtain a processing result (for example, a channel prediction result of a physical layer) of the first task.
[0179] It should be noted that when the sub-model deployed by the computing node is insufficient to meet the requirements of the first task, the computing node can enhance the function or capability of the sub-model under the instruction of the management node, or the computing node autonomously determines to enhance the function or capability of the sub-model. The following will be introduced respectively:
[0180] In a possible implementation, the management node can send new configuration information to the computing node to enable the computing node to enhance the function or capability of the sub-model. Taking the first computing node as an example, if the first computing node is short of software resources and / or hardware resources after deploying the sub-model, resulting in that the first computing node cannot meet the task requirements of the first task when processing the task data of the first task by the first sub-model, the management node sends second configuration information to the first computing node. The second configuration information includes second model parameters, and the second model parameters are used to instruct to enable a second network layer in the first sub-model. The second network layer is at least one network layer in the first computing node, and the second network layer is different from the first network layer. The management node modifies the configuration of the first sub-model by sending the second configuration information of the first sub-model to the first computing node, so that the first computing node configures the first sub-model by using the second configuration information.
[0181] In an example, the second model parameter indicates all network layers of the first sub-model that need to be enabled, i.e., the second network layer includes the first network layer. For example, referring to FIG. 3A, the management node indicates the first model parameter through the first configuration information, and the first model parameter indicates that layers a1, a2, b1, a3, a4, c1, and a5 in the sub-model 1 are enabled. When the management node determines that the first computing node needs to enhance the sub-model 1, the management node indicates the second model parameter through the second configuration information, and the second model parameter indicates that layers a1, a2, b1, a3, a4, c1, a5, a6, and a7 in the sub-model 1 are enabled.
[0182] In another example, the second model parameter indicates network layers that are newly added to the first sub-model and need to be enabled, i.e., the sum of the first network layer and the second network layer is the final network layer of the first sub-model that is enabled. For example, referring to FIG. 3A, the management node indicates the first model parameter through the first configuration information, and the first model parameter indicates that layers a1, a2, b1, a3, a4, c1, and a5 in the sub-model 1 are enabled. When the management node determines that the first computing node needs to enhance the sub-model 1, the management node indicates the second model parameter through the second configuration information, and the second model parameter indicates that layers a6 and a7 in the sub-model 1 are enabled.
[0183] In this embodiment, the computing node can modify the network layer of the first sub-model enabled under the indication of the management node, thereby achieving enhancement of the sub-model, and facilitating improvement of flexibility and accuracy of the management node in configuring the sub-model for each computing node.
[0184] In another possible embodiment, if the computing node stores the attribute information of the target model and the target model, the computing node can determine the function or capability of the enhanced sub-model based on the attribute information of the target model. For example, if the software resource and / or hardware resource of the first computing node is in shortage after the sub-model is deployed, and the first computing node stores the target model and the attribute information of the target model, the first computing node determines the second network layer based on the attribute information of the target model and the first configuration information of the first sub-model, and enables the second network layer to enhance the first sub-model. The second network layer is explained above and will not be repeated here.
[0185] In this embodiment, the computing node can modify the network layer of the first sub-model enabled based on the attribute information of the target model, thereby achieving enhancement of the sub-model, and facilitating improvement of the service quality of the sub-model in the computing node in improving the AI service.
[0186] In step 413, the first computing node sends a first task response to the access network device.
[0187] The first task response includes a processing result of the first task.
[0188] Optionally, the processing result of the first task can be sent to the access network device by one of the at least two computing nodes (for example, the first computing node), or can be sent to the access network device by the at least two computing nodes respectively, which is not limited here.
[0189] In this embodiment, a distributed AI configuration process in the RAN AI scenario is provided. In the process, the management node can determine at least two sub-models capable of meeting the requirements of the first task based on the task requirements of the first task and the requirements of the functional modules, and deploy the at least two sub-models to at least two computing nodes. In addition, the management node also configures the at least two sub-models and arranges the tasks in the at least two sub-models based on the task requirements of the first task and other information. Since the foregoing process fully considers the requirements and characteristics of the AI task itself, the AI resources of the computing nodes, the network status of the access network device, and the requirements of the functional modules that the management node can provide, the management node achieves more reasonable distributed AI setting and task allocation. In addition, this embodiment provides an implementation architecture of distributed intelligence in mobile networks. By multi-level interconnection of models based on functional modules, a unified deployment framework for various distributed AI paradigms is provided, and the fusion of multiple distributed AI paradigms can be achieved by setting the proportion of different interaction modes, thereby integrating the advantages of multiple paradigms and improving the quality of AI services.
[0190] As shown in FIG. 5, it is a flowchart of one embodiment of a distributed AI task configuration method provided by the present application. In this embodiment, the first task is an AI task initiated by a third-party application in a terminal device. In this case, the management node, the computing node, the access network device and the terminal device will perform the following steps:
[0191] Step 501, the terminal device sends a first task request; correspondingly, the management node receives the first task request.
[0192] For example, the terminal device sends the first task request to the management node through the access network device; correspondingly, the management node receives the first task request sent by the terminal device through the access network device.
[0193] The first task request includes task information of the first task. The task information of the first task includes the task type of the first task and the task requirements of the first task. For the task type of the first task and the task requirements of the first task, please refer to the foregoing step 201, which will not be repeated here.
[0194] In this embodiment, the first task is an AI task initiated by a third-party application of the terminal device. For example, the first task is an AI task of the type of indicating a navigation path planning, congestion prediction, etc. It should be understood that in actual applications, the first task can also be other types of AI tasks, which are not described herein.
[0195] Optionally, the first task request further includes resource state information of the terminal device, used to indicate a size of an idle resource of the terminal device. So that the management node determines whether to need to deploy a sub-model in the terminal device.
[0196] In step 502, the management node determines a target model based on the task type, the task requirement, and attribute information of a plurality of functional modules stored by the management node.
[0197] In step 503, at least two computing nodes send resource state information of the computing nodes to the management node; correspondingly, the management node receives the resource state information of the computing nodes from the at least two computing nodes.
[0198] In step 504, the management node determines at least two sub-models, attribute information of the sub-models, and at least two computing nodes for deploying the at least two sub-models based on the resource state information of the at least two computing nodes and the target model.
[0199] In this embodiment, steps 502 to 504 are similar to steps 402 to 404 described above, and specific details are described in steps 402 to 404 described above, which are not described herein.
[0200] Optionally, if the terminal device has the capability of supporting the deployment of a sub-model, and the management node can obtain the size of the idle resource of the terminal device, the at least two computing nodes determined by the management node can include the terminal device, that is, the terminal device can serve as a computing node for deploying a certain sub-model (hereinafter referred to as a third sub-model). In this case, the size of the idle resource of the terminal device needs to meet the resource requirement of deploying the third sub-model.
[0201] In this embodiment, the management node determines a first sub-model deployed in a first computing node and a second sub-model deployed in a second computing node based on the target model. Optionally, the management node can also determine a third sub-model deployed in the terminal device.
[0202] In step 505a, the management node sends the first sub-model and attribute information of the first sub-model to the first computing node; correspondingly, the first computing node receives the first sub-model and the attribute information of the first sub-model from the management node.
[0203] At step 505b, the management node sends the second sub-model and the attribute information of the second sub-model to the second computing node. Correspondingly, the second computing node receives the second sub-model and the attribute information of the second sub-model from the management node.
[0204] In this embodiment, the steps 505a and 505b are similar to the steps 405a and 405b described above. For details, refer to the description of the steps 405a and 405b above, which will not be repeated here.
[0205] At step 505c, the management node sends the third sub-model and the attribute information of the third sub-model to the terminal device. Correspondingly, the terminal device receives the third sub-model and the attribute information of the third sub-model from the management node.
[0206] In this embodiment, the step 505c is an optional step.
[0207] At step 506a, the first computing node deploys the first sub-model.
[0208] Specifically, the first computing node deploys the first sub-model based on the attribute information of the first sub-model.
[0209] At step 506b, the second computing node deploys the second sub-model.
[0210] Specifically, the second computing node deploys the second sub-model based on the attribute information of the second sub-model.
[0211] At step 506c, the terminal device deploys the sub-model.
[0212] In this embodiment, the step 506c is an optional step.
[0213] At step 507, the access network device sends the first network state information to the management node. Correspondingly, the management node receives the first network state information from the access network device.
[0214] In this embodiment, the step 507 is similar to the step 407 described above. For details, refer to the description of the step 407 above, which will not be repeated here.
[0215] At step 508, the terminal device sends the second network state information to the management node. Correspondingly, the management node receives the second network state information from the terminal device.
[0216] The second network state information is used to indicate the air interface state of the terminal device that sends the first task request, such as the channel quality between the terminal device and the access network device.
[0217] At step 509, the management node determines the configuration information of each sub-model based on the resource state information of the at least two computing nodes, the first network state information, and the second network state information.
[0218] The at least two computing nodes include the first computing node and the second computing node. For example, the management node determines configuration information of the first sub-model deployed by the first computing node and configuration information of the second sub-model deployed by the second computing node based on the resource state information of the first computing node, the resource state information of the second computing node, the first network state information, and the second network state information. The configuration information of the first sub-model is used to configure model parameters used by the first sub-model when participating in processing the first task, and the configuration information of the second sub-model is used to configure model parameters used by the second sub-model when participating in processing the first task. In addition, the configuration information of the first sub-model is also used to indicate the establishment of a connection relationship between the first sub-model and the second sub-model deployed on the second computing node, and the configuration information of the second sub-model is also used to indicate the establishment of a connection relationship between the second sub-model and the first sub-model deployed on the first computing node.
[0219] The configuration information of the sub-model is described in the foregoing step 205, which will not be described here.
[0220] Optionally, the management node also determines configuration information of at least one participating node. The configuration information of the participating node is described in the foregoing step 409, which will not be described here.
[0221] In step 510a, the management node sends the configuration information of the first sub-model to the first computing node. Correspondingly, the first computing node receives the configuration information of the first sub-model from the management node.
[0222] In step 510b, the management node sends the configuration information of the second sub-model to the second computing node. Correspondingly, the second computing node receives the configuration information of the second sub-model from the management node.
[0223] In step 510c, the management node sends the configuration information of the third sub-model to the terminal device. Correspondingly, the terminal device receives the configuration information of the third sub-model from the management node.
[0224] In this embodiment, step 510c is an optional step.
[0225] Optionally, the access network device sends the assistance information of the first task to the first computing node and / or the second computing node. Correspondingly, the first computing node and / or the second computing node receives the assistance information of the first task sent by the access network device.
[0226] The auxiliary information is related to a task type of the first task, and the content of the auxiliary information is different when the task type of the first task is different. Optionally, the auxiliary information includes prediction information of the access network device on a network environment and / or user behavior, and / or perception information of the access network device on the network environment and / or user behavior. For example, the auxiliary information is user / air interface auxiliary information required by a third-party AI task. For example, the auxiliary information includes user portrait, behavior prediction, location / environment perception, and the like.
[0227] In this embodiment, the access network device provides the first computing node and / or the second computing node with the auxiliary information of the first task, so that the first computing node and / or the second computing node take the auxiliary information as input data of the AI model, which is beneficial to enhance the performance of the sub-model in processing the first task, improve the capability of the sub-model, and provide personalized AI services according to user preferences.
[0228] In step 511a, the first computing node configures a task processing strategy of the first sub-model.
[0229] After receiving the configuration information of the first sub-model, the first computing node configures model parameters used when participating in processing the first task based on the configuration information of the first sub-model, and establishes a connection relationship between the first sub-model and the second sub-model deployed on the second computing node. Optionally, the first computing node also establishes a connection relationship between the first sub-model and the third sub-model deployed on the terminal device.
[0230] In step 511b, the second computing node configures a task processing strategy of the second sub-model.
[0231] After receiving the configuration information of the second sub-model, the second computing node configures model parameters used when participating in processing the first task based on the configuration information of the second sub-model, and establishes a connection relationship between the second sub-model and the first sub-model deployed on the first computing node. Optionally, the second computing node also establishes a connection relationship between the second sub-model and the third sub-model deployed on the terminal device.
[0232] In step 511c, the terminal device configures a task processing strategy of the third sub-model.
[0233] In this embodiment, step 511c is an optional step.
[0234] After receiving the configuration information of the third sub-model, the terminal device configures model parameters used when participating in processing the first task based on the configuration information of the third sub-model, and establishes a connection relationship between the terminal device and the first sub-model deployed on the first computing node, and a connection relationship between the terminal device and the second sub-model deployed on the second computing node.
[0235] Step 512a, the terminal device sends task data of the first task to the first computing node.
[0236] Step 512b, the terminal device sends task data of the first task to the first computing node.
[0237] It should be noted that the terminal device can send first data requiring processing of the first sub-model to the first computing node and send second data requiring processing of the second sub-model to the second computing node, respectively, wherein the first data and the second data are both task data of the first task. The terminal device can also send the first data and the second data to one of the computing nodes (for example, the first computing node) and forward them to the other computing node (for example, the second computing node) by the computing node.
[0238] Step 513a, the first computing node and the second computing node jointly process the first task through the first sub-model and the second sub-model, and output a processing result of the first task.
[0239] For example, the first task can be an AI task of path planning, congestion prediction, etc. Taking path planning as an example, the first computing node and the second computing node process the received task data of the first task through the first sub-model and the second sub-model, respectively, to obtain a processing result of the first task (for example, a path planning result).
[0240] It should be noted that when the sub-model deployed by the computing node is insufficient to meet the requirements of the first task, the computing node can enhance the function or capability of the sub-model under the instruction of the management node, or the computing node can autonomously determine to enhance the function or capability of the sub-model. For details, please refer to the related description in step 412 above, which will not be repeated here.
[0241] Step 513b, the first computing node, the second computing node and the terminal device jointly process the first task through the first sub-model, the second sub-model and the third sub-model, and output a processing result of the first task.
[0242] In this embodiment, step 513b is an optional step. In the case that the management node instructs the terminal device to deploy the third sub-model, the terminal device participates in processing the first task. Since the management node indicates the data interaction rule between the terminal device and the computing node (for example, the first computing node or the second computing node) in the configuration information of the third sub-model, the terminal device can obtain intermediate data for processing the first task or the processing result of the first task.
[0243] Step 514, the first computing node sends a first task response to the terminal device.
[0244] The first task response includes the processing result of the first task.
[0245] In this embodiment, step 514 is an optional step. For example, when the terminal device has obtained the processing result of the first task in the process of participating in processing the first task, the first computing node can not perform step 514.
[0246] In this embodiment, a distributed AI configuration process in a third-party AI scenario is provided. The process fully considers the requirements and characteristics of the AI task itself, the AI resources of the computing node, the network status of the access network device, and the requirements of the functional modules that the management node can provide. Therefore, the management node implements more reasonable distributed AI settings and task allocation. In addition, this embodiment supports the deployment of sub-models by terminal devices, enabling terminals and in-network computing nodes to jointly complete the collaborative processing of third-party AI services, which helps to improve the flexibility and privacy protection of distributed AI applications. In addition, the computing node can also obtain auxiliary information perceived by the access network device, thereby improving the personalized service quality of the third-party AI application.
[0247] As shown in FIG. 6, it is a structural schematic diagram of another apparatus 60 provided by the present application. It should be understood that the management node or the computing node in the corresponding method embodiments of the foregoing FIG. 2, FIG. 4 or FIG. 5 can be based on the structure of the apparatus 60 shown in FIG. 6 in this embodiment. As shown in FIG. 6, the apparatus 60 can include a processor 601, a memory 603 and a communication interface 602. Wherein, the processor 601 is coupled with the memory 603, and the processor 601 is coupled with the communication interface 602.
[0248] Wherein, the foregoing communication interface 602 is connected with other apparatuses through a communication link. For example, when the apparatus 60 implements the function of the management node, the communication interface 602 can include the interface between the management node and the access network device, and also include the interface between the management node and at least one computing node. For another example, when the apparatus 60 implements the function of the computing node, the communication interface 602 can include the interface between the computing node and the access network device, and also include the interface between the computing node and other computing nodes.
[0249] The processor 601 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The processor 601 can refer to one processor or include multiple processors, and the specific implementation is not limited herein.
[0250] In addition, the memory 603 is mainly used for storing software programs and data. The memory 603 can exist independently and be connected to the processor 601. Alternatively, the memory 603 can be integrated with the processor 601, for example, integrated in one or more chips. The memory 603 can store program codes for implementing the technical solutions of the embodiments of the present application and be controlled to execute by the processor 601. The executed computer programs of various types can also be regarded as the driver of the processor 601. The memory 603 can include a volatile memory such as a random-access memory (RAM), and can also include a non-volatile memory such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or a combination thereof. The memory 603 can refer to one memory or include multiple memories. For example, the memory 603 is used to store various data. For example, when the apparatus 60 implements the function of a management node, the memory 603 stores at least one function module and attribute information of each function module. For another example, when the apparatus 60 implements the function of a computing node, the memory 603 is used to store a received sub-model, attribute information of the sub-model, and configuration information of the sub-model.
[0251] In one design, the apparatus 60 is configured to perform the method of the management node in the foregoing embodiment of FIG. 2. In the apparatus 60, the communication interface 602 is configured to receive a first task request, the first task request including task information of a first task; the processor 601 is configured to determine at least two sub-models and attribute information of the sub-models based on the task information and attribute information of a plurality of function modules stored by the management node, the sub-models including at least one function module and / or a part of a function module, the at least two sub-models being configured to be distributedly deployed in at least two computing nodes, the attribute information of the sub-models being determined based on attribute information of the function modules constituting the sub-models, and the attribute information of the sub-models being used for the computing nodes to deploy the sub-models; the communication interface 602 is configured to send, to the at least two computing nodes, the sub-models to be respectively deployed by the computing nodes and the attribute information of the sub-models; and the processor 601 is configured to send, to the at least two computing nodes, configuration information of the sub-models to be deployed by the computing nodes based on the task information, the configuration information of the sub-models being used to configure model parameters used by the sub-models when processing the first task and to establish a connection relationship between the sub-models and sub-models deployed in different computing nodes.
[0252] In one possible implementation, the attribute information of the function module includes a function type of the function module and a requirement of the function module, the requirement of the function module being used to indicate a requirement on data processing and / or data transmission when implementing a function of the function module.
[0253] In one possible implementation, the task information of the first task includes a task type of the first task and a task requirement of the first task, the task requirement of the first task being used to indicate a requirement of the distributed AI processing service requested by the first task request on data processing and / or data transmission.
[0254] In one possible implementation, the processor 601 is specifically configured to determine a target model supporting processing of the first task based on the task information and the attribute information of the plurality of function modules, the target model being constituted by at least one function module of the plurality of function modules, the task type of the first task being used to determine a model type of the target model, a performance requirement of the target model satisfying the task requirement of the first task, and the performance requirement of the target model being determined based on the requirement of the at least one function module; and determine the at least two sub-models and the attribute information of each sub-model based on the target model.
[0255] In one possible implementation, the processor 601 is further configured to obtain resource state information of the at least one computing node, the resource state information being used to indicate a usage state of AI resources of the computing node, the AI resources including model resources, computing resources and data resources; and determine the at least two computing nodes and the sub-model to be deployed by each computing node of the at least two computing nodes based on the resource state information of the at least one computing node, the AI resources of the computing node satisfying an AI resource requirement of the sub-model to be deployed by the computing node.
[0256] In a possible implementation, the communication interface 602 is further configured to receive network state information from an access network device connected to the computing node, the network state information being used to indicate a network state of a communication device applying for the first task; the processor 601 is further configured to determine configuration information of the sub-model to be deployed on the computing node based on the task information, the resource state information of the computing node, and the network state information; the network state information includes first network state information and / or second network state information, the first network state information is used to indicate a radio state of the access network device connected to the computing node, and the second network state information is used to indicate a radio state of a terminal device sending the first task request.
[0257] In a possible implementation, the attribute information of the function module further includes information of a plurality of network layers included in the function module, an input dimension of the function module, and an output dimension of the function module; and the attribute information of the sub-model includes information of a plurality of network layers included in the sub-model.
[0258] In a possible implementation, the configuration information of the sub-model includes parameters of enabled network layers, the parameters of the enabled network layers being used to indicate network layers enabled by the sub-model deployed on the computing node when participating in processing the first task.
[0259] In a possible implementation, the configuration information of the sub-model further includes a connection relationship between the sub-model and a sub-model deployed on another computing node.
[0260] In a possible implementation, the communication interface 602 is specifically configured to send first configuration information of the first sub-model to the first computing node, the first configuration information including first model parameters, the first model parameters being used to indicate enabling a first network layer in the first sub-model, the first network layer being at least one network layer in the plurality of network layers indicated by the attribute information of the first sub-model.
[0261] In a possible implementation, the communication interface 602 is further configured to send second configuration information of the first sub-model to the first computing node, the second configuration information including second model parameters, the second model parameters being used to indicate enabling a second network layer in the first sub-model, the second network layer being at least one network layer in the first computing node, and the second network layer being different from the first network layer.
[0262] In a possible implementation, the communication interface 602 is further configured to send the target model and attribute information of the target model to the first computing node, the attribute information of the target model including information of network layers included in each of the at least two sub-models, and a connection relationship between network layers of different sub-models in the at least two sub-models.
[0263] In a possible implementation, the communication interface 602 is further configured to send configuration information of the participating node to at least one participating node, the configuration information of the participating node being used to indicate a data interaction strategy between the participating node and the computing node.
[0264] In a possible implementation, the communication interface 602 is further configured to receive auxiliary information of the first task, the auxiliary information including prediction information of the access network device on a network environment and / or user behavior, and the auxiliary information being related to a task type of the first task.
[0265] It should be noted that the specific implementation and advantages of the embodiments can refer to the method of the management node in the above embodiments, which will not be described here.
[0266] In another design, the apparatus 60 is configured to perform the method of the first computing node in the corresponding embodiments of FIG. 4 or FIG. 5. In the apparatus 60, the communication interface 602 is configured to receive a first sub-model of the first task and attribute information of the first sub-model, the attribute information of the first sub-model being used to indicate a structure and a function of the first sub-model; the processor 601 is configured to deploy the first sub-model based on the attribute information of the first sub-model; the communication interface 602 is further configured to receive configuration information of the first sub-model, the configuration information of the first sub-model being used to configure model parameters used by the first sub-model when participating in processing the first task and a connection relationship between the first sub-model and sub-models deployed in different computing nodes; and the processor 601 is further configured to configure the first sub-model based on the configuration information of the first sub-model and establish the connection relationship between the first sub-model and the sub-models deployed in different computing nodes.
[0267] In a possible implementation, the attribute information of the first sub-model includes information of a plurality of network layers included in the first sub-model.
[0268] In a possible implementation, the configuration information of the first sub-model includes parameters of enabled network layers, the parameters of the enabled network layers being used to indicate the network layers enabled by the first sub-model deployed in the first computing node when participating in processing the first task.
[0269] In a possible implementation, the configuration information of the first sub-model further includes a connection relationship between the first sub-model and sub-models deployed in other computing nodes.
[0270] In a possible implementation, the configuration information of the first sub-model includes first configuration information, the first configuration information including first model parameters, the first model parameters being used to indicate an enabled first network layer in the first sub-model, the first network layer being at least one network layer of a plurality of network layers indicated by the attribute information of the first sub-model. The processor 601 is specifically configured to enable the first network layer in the first sub-model based on the first model parameters.
[0271] In a possible implementation, the communication interface 602 is configured to receive second configuration information of the first sub-model, the second configuration information comprising second model parameters, the second model parameters being used to indicate enabling a second network layer in the first sub-model, the second network layer being at least one network layer in the first computing node, and the second network layer being different from the first network layer; and the processor 601 is configured to enable the second network layer in the first sub-model based on the second model parameters.
[0272] In a possible implementation, the communication interface 602 is further configured to receive the target model and attribute information of the target model, the attribute information of the target model comprising information of network layers included in each of the at least two sub-models, and connection relationships between network layers of different sub-models in the at least two sub-models.
[0273] In a possible implementation, the communication interface 602 is further configured to send resource state information of the first computing node, the resource state information being used to indicate a usage state of an AI resource of the first computing node, the AI resource comprising a model resource, a computing resource, and a data resource.
[0274] It should be noted that the specific implementation and advantages of the embodiments can refer to the method of the first computing node in the above embodiments, and will not be described here.
[0275] As shown in FIG. 7, the present application further provides an apparatus 70. The apparatus 70 can be a management node or a computing node, or a component (for example, an integrated circuit, a chip, etc.) of the management node or the computing node. The apparatus 70 can also be other functional modules for implementing the method in the method embodiments of the present application.
[0276] The apparatus 70 can include a processing module 701 (or a processing unit). Optionally, it can also include an interface module 702 (or a transceiving unit or a transceiving module) and a storage module 703 (or a storage unit). The interface module 702 is configured to implement communication with other devices. The interface module 702 can be, for example, a transceiving module or an input / output module.
[0277] In a possible design, one or more modules in FIG. 7 can be implemented by one or more processors, or by one or more processors and memories; or by one or more processors and transceivers; or by one or more processors, memories, and transceivers, and the embodiments of the present application do not make any limitation in this regard. The processor, the memory, and the transceiver can be separately arranged, or integrated together.
[0278] The apparatus 70 is configured to implement the functions of the management node described in the embodiments of the present application. For example, the apparatus 70 includes modules or units or means corresponding to the steps involved in the management node described in the embodiments of the present application, which are implemented by software or by hardware, or by a combination of hardware and software. For details, further reference can be made to the corresponding description in the foregoing method embodiments. For details, reference can be made to the foregoing description in the method embodiments, which will not be repeated here.
[0279] Alternatively, the apparatus 70 is configured to implement the functions of the computing node described in the embodiments of the present application. For example, the apparatus 70 includes modules or units or means corresponding to the steps involved in the computing node described in the embodiments of the present application, which are implemented by software or by hardware, or by a combination of hardware and software. For details, further reference can be made to the corresponding description in the foregoing method embodiments. For details, reference can be made to the foregoing description in the method embodiments, which will not be repeated here.
[0280] In addition, the present application provides a computer program product including one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. For example, the method related to the management node in the foregoing FIG. 2, FIG. 4 or FIG. 5 is implemented. For another example, the method related to the computing node in the foregoing FIG. 2, FIG. 4 or FIG. 5 is implemented. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.). The computer readable storage medium can be any available medium that can be used by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital versatile disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)) and the like.
[0281] Furthermore, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method related to the management node in the foregoing FIG. 2, FIG. 4 or FIG. 5.
[0282] Furthermore, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method related to the computing node in the foregoing FIG. 2, FIG. 4 or FIG. 5.
[0283] It should be understood that, in various embodiments of the present application, the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0284] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
Claims
1. A method for configuring distributed AI tasks, applied to a management node, characterized in that, The method comprises the following steps: receiving a first task request, the first task request comprising task information of the first task; determining at least two sub-models and attribute information of the sub-models based on the task information and attribute information of a plurality of function modules stored in the management node, the sub-models comprising at least one of the function modules and / or a part of one of the function modules, the at least two sub-models being used for distributed deployment in at least two computing nodes, the attribute information of the sub-models being determined based on attribute information of the function modules constituting the sub-models, and the attribute information of the sub-models being used for the computing nodes to deploy the sub-models; sending, to the at least two computing nodes, the sub-models to be deployed by the computing nodes respectively and the attribute information of the sub-models; sending, to the at least two computing nodes, configuration information of the sub-models to be deployed by the computing nodes based on the task information, the configuration information of the sub-models being used for configuring model parameters used by the sub-models when participating in processing the first task and establishing a connection relationship between the sub-models and sub-models deployed in different computing nodes.
2. The method of claim 1, wherein, The attribute information of the function modules comprises a function type of the function modules and a requirement of the function modules, the requirement of the function modules being used for indicating a requirement for data processing and / or data transmission when implementing the function of the function modules.
3. The method according to claim 1 or 2, characterized in that, The task information of the first task comprises a task type of the first task and a task requirement of the first task, the task requirement of the first task being used for indicating a requirement for data processing and / or data transmission of the distributed AI processing service of the first task request.
4. The method of claim 3, wherein, The method further comprises the following steps before the step of sending, to the at least two computing nodes, the sub-models to be deployed by the computing nodes respectively and the attribute information of the sub-models: determining a target model supporting processing the first task based on the task information and the attribute information of the plurality of function modules, the target model being constituted by at least one of the function modules, the task type of the first task being used for determining a model type of the target model, a performance requirement of the target model meeting the task requirement of the first task, and the performance requirement of the target model being determined based on the requirement of the at least one function module; determining the at least two sub-models and the attribute information of each of the sub-models based on the target model.
5. The method according to any one of claims 1 to 4, characterized in that, The attribute information of the sub-models comprises an AI resource requirement of the sub-models. The method further comprises the following steps before the step of sending, to the at least two computing nodes, the sub-models to be deployed by the computing nodes respectively and the attribute information of the sub-models: obtaining resource state information of at least one computing node, the resource state information being used for indicating a use state of AI resources of the computing node, the AI resources comprising model resources, computing resources and data resources; determining the at least two computing nodes based on the resource state information of the at least one computing node and determining the sub-models to be deployed by each of the computing nodes in the at least two computing nodes, the AI resources of the computing node meeting the AI resource requirement of the sub-model to be deployed by the computing node.
6. The method of claim 5, wherein, The method further comprises: receiving network state information from an access network device connected to the computing node, the network state information being used to indicate a network state of a communication device applying for the first task; determining configuration information of a sub-model to be deployed by the computing node based on the task information, resource state information of the computing node, and the network state information; The network state information includes first network state information and / or second network state information, the first network state information being used to indicate an air interface state of an access network device connected to the computing node, and the second network state information being used to indicate an air interface state of a terminal device sending the first task request.
7. The method of claim 6, wherein, The attribute information of the function module further includes information of a plurality of network layers contained in the function module, an input dimension of the function module, and an output dimension of the function module; and the attribute information of the sub-model includes information of a plurality of network layers contained in the sub-model.
8. The method according to any one of claims 1 to 7, characterized in that, The configuration information of the sub-model includes parameters of enabled network layers, the parameters of the enabled network layers being used to indicate network layers enabled by the sub-model deployed in the computing node when participating in processing the first task.
9. The method of claim 8, wherein, The configuration information of the sub-model further includes a connection relationship of the sub-model with sub-models deployed in other computing nodes.
10. The method according to any one of claims 1 to 9, characterized in that, The sending of the configuration information of the sub-model deployed by the computing node to the at least two computing nodes based on the task information comprises: sending first configuration information of a first sub-model to a first computing node, the first configuration information including first model parameters used to enable a first network layer in the first sub-model, the first network layer being at least one network layer of a plurality of network layers indicated by attribute information of the first sub-model.
11. The method of claim 10, wherein, The method further comprises: sending second configuration information of the first sub-model to the first computing node, the second configuration information including second model parameters used to enable a second network layer in the first sub-model, the second network layer being at least one network layer in the first computing node, and the second network layer being different from the first network layer.
12. The method according to any one of claims 4 to 11, characterized in that, The method further comprises: sending the target model and attribute information of the target model to the first computing node, the attribute information of the target model including information of network layers contained in each of the at least two sub-models, and a connection relationship between network layers of different sub-models in the at least two sub-models.
13. The method according to any one of claims 1 to 12, characterized in that, The method further comprises: sending configuration information of the participating node to at least one participating node, the configuration information of the participating node being used to indicate a data interaction strategy between the participating node and the computing node.
14. The method according to any one of claims 1 to 13, characterized in that, The method further comprises: receiving auxiliary information of the first task, the auxiliary information including prediction information of a network environment and / or user behavior by an access network device, and the auxiliary information being related to a task type of the first task.
15. A distributed AI task configuration method applied to a first computing node, comprising: The method further comprises: receiving a first sub-model of a first task and attribute information of the first sub-model, the attribute information of the first sub-model being used to indicate a structure and function of the first sub-model; deploy the first sub-model based on the attribute information of the first sub-model; receive configuration information of the first sub-model, the configuration information of the first sub-model being used to configure model parameters used by the first sub-model when participating in processing the first task and a connection relationship between the first sub-model and sub-models deployed on different computing nodes; configure the first sub-model based on the configuration information of the first sub-model and establish the connection relationship between the first sub-model and the sub-models deployed on different computing nodes.
16. The method of claim 15, wherein, The attribute information of the first sub-model includes information of a plurality of network layers included in the first sub-model.
17. The method according to claim 15 or 16, characterized in that, The configuration information of the first sub-model includes parameters of enabled network layers, the parameters of the enabled network layers being used to indicate network layers enabled by the first sub-model deployed on the first computing node when participating in processing the first task.
18. The method of claim 17, wherein, The configuration information of the first sub-model further includes a connection relationship between the first sub-model and sub-models deployed on other computing nodes.
19. The method according to any one of claims 15 to 18, characterized in that, The configuration information of the first sub-model includes first configuration information, the first configuration information including first model parameters, the first model parameters being used to indicate enabling a first network layer in the first sub-model, the first network layer being at least one network layer of a plurality of network layers indicated by the attribute information of the first sub-model; The deploying the first sub-model based on the attribute information of the first sub-model includes: enabling the first network layer in the first sub-model based on the first model parameters.
20. The method of any one of claims 15 to 19, wherein, The method further includes: receiving second configuration information of the first sub-model, the second configuration information including second model parameters, the second model parameters being used to indicate enabling a second network layer in the first sub-model, the second network layer being at least one network layer in the first computing node, the second network layer being different from the first network layer; enabling the second network layer in the first sub-model based on the second model parameters.
21. The method of any one of claims 15 to 20, wherein, The method further includes: receiving a target model and attribute information of the target model, the attribute information of the target model including information of network layers included in each of at least two sub-models and a connection relationship between network layers of different sub-models in the at least two sub-models.
22. The method of any one of claims 15 to 21, wherein, The method further includes: sending resource state information of the first computing node, the resource state information being used to indicate a usage state of AI resources of the first computing node, the AI resources including model resources, computing resources and data resources.
23. A management node, c h a r a c t e r i z e d b y, comprise a processor and a memory; wherein the memory stores a computer program; the processor invokes the computer program to cause the management node to perform the method of any one of claims 1 to 14.
24. A computing node, characterized in that, comprise a processor and a memory; wherein the memory stores a computer program; the processor invokes the computer program to cause the computing node to perform the method of any one of claims 15 to 22.
25. A computer-readable storage medium storing instructions which, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 14; or, the method of any one of claims 15 to 22.
26. A network system, characterized by including: a management node performing the method of any one of claims 1 to 14, and, a management node performing the method of any one of claims 15 to 22.