System and method for parallel and split learning

By splitting the AI ​​model into an augmented tree structure and introducing dynamic dropout control, the problem of high training cost for large-scale AI models is solved, achieving efficient utilization of network resources and rapid training.

CN121464446APending Publication Date: 2026-02-03HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380100034.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

The training cost of existing large-scale AI models is high, and existing solutions cannot effectively take advantage of changes in network connectivity and adapt to interruptions, resulting in a decline in training performance.

Method used

The parallel and split learning (PSL) framework is adopted to split the AI ​​model into an augmented tree structure, with each level corresponding to a partition and each node being a copy of the partition. Dynamic dropout control is introduced to adjust the dropout rate according to network link information to adapt to network changes.

Benefits of technology

It improves training speed, reduces training latency, avoids overfitting, effectively utilizes network resources, and reduces training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121464446A_ABST
    Figure CN121464446A_ABST
Patent Text Reader

Abstract

The present invention provides a method and a system for generating and instantiating an enhanced tree model from an artificial intelligence (AI) model through a parallel and split learning (PSL) controller, and a method and a system for generating and instantiating the enhanced tree model from the AI model through a parallel and split learning (PSL) controller. The method receives an AI model comprising an input layer and an output layer, obtains computing resources for each network node, and obtains link information for numerous links connecting these nodes. According to the method, the AI model is divided into different partitions, so that an enhanced tree model is created, and the enhanced tree model is characterized in that root nodes correspond to top-level partitions, leaf nodes correspond to bottom-level partitions, and branch nodes correspond to middle-level partitions. The nodes are organized in a tree structure, each branch node is connected to a number of leaf nodes, or the root node is connected to a number of branch nodes. The method further includes instantiating a node in the enhanced tree model as an entity of the network based on matching computing resources, link information, and communication requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application is the first application for this technology. The present invention generally relates to the fields of artificial intelligence (AI) and machine learning (ML), and more particularly to optimizing distributed AI models using parallel and split learning techniques and augmented tree models to improve computational efficiency and reduce training latency. Background Technology

[0002] Currently, large-scale artificial intelligence (AI) models with millions or billions of parameters have demonstrated enormous potential in various AI applications, such as natural language processing and computer vision (e.g., ChatGPT). Many consider it the most promising AI technology of the future. However, all existing large-scale AI models are owned by tech giants (such as Google and Microsoft) and trained using intensive computing resources, making their use prohibitively expensive for ordinary companies or individuals. To reduce the cost of using large-scale AI models, a promising approach is to deploy them as distributed systems in a network environment (CN or RAN), where each network entity embeds a partition of the large-scale AI model, the resource requirements of which can be handled by the entity's computing resources.

[0003] AI models consist of two main components: an encoder and a classifier. In split-learning models involving multiple clients, the training process is performed sequentially to ensure that all clients contribute to the model's learning. Cluster-based parallelization is an advanced technique that improves efficiency by classifying devices into clusters and training them in parallel within them.

[0004] However, the current approach has limitations, such as the lack of server-side parallelization, failure to utilize server-side partitioning for parallelization, inability to adapt to changes in network connectivity, and the pre-configured dropout rate not taking into account connection differences. Furthermore, connection interruptions can lead to data loss, negatively impacting training performance.

[0005] Finally, all existing large-scale AI models are owned by tech giants (such as Google and Microsoft) and trained using intensive computing resources, making their use prohibitively expensive for ordinary companies or individuals. To reduce the cost of using large-scale AI models, a promising approach is to deploy them as distributed systems within a network environment (CN or RAN), where each network entity embeds a partition of the large-scale AI model, the resource requirements of which can be handled by the entity's computing resources.

[0006] Therefore, there is a need for systems and methods that eliminate or reduce one or more limitations of existing technologies.

[0007] The purpose of the background information is to disclose information that the applicant considers relevant to this invention. It is not intended to acknowledge, nor should it be construed, that any of the foregoing information constitutes prior art relative to this invention. Summary of the Invention

[0008] Embodiments of this invention provide a method and system for a reinforcement tree model for parallel and split learning (PSL). The method splits an AI model into several partitions, thereby creating a reinforcement tree structure where each level corresponds to a partition, each branch node corresponds to an AI enabler, and each leaf node corresponds to a client. By using split learning, an AI model including neural networks can be transformed into a distributed model, with partitions ranging from the bottom level (including the input layer) to the top level (including the output layer). The reinforcement tree model represents a proposed method for implementing a distributed AI model, allowing multiple copies of each partition level to be trained simultaneously. Each node is a copy of its corresponding partition and requires deployment by a network entity.

[0009] The objective of this invention is to provide systems and methods for parallel and split learning. The Parallel and Split Learning (PSL) framework allows for parallel training on all nodes at the same level in a boosting tree model, thereby accelerating training. Once all branch nodes at the same level have received their corresponding forward propagation (FP) and back propagation (BP) data from their associated child and parent nodes, respectively, they can perform FP and BP in parallel. Furthermore, dynamic dropout control is introduced, where network link information is collected and a dropout rate metric is configured. This allows for local customization of dropout decisions, specifying which of one or more neurons or links should be temporarily silenced or removed in each training iteration. This achieves dynamic dropout while maintaining the average dropout rate indicated by the dropout rate metric. Dynamic dropout improves the training performance of the boosting tree model by leveraging the disruption or variation of network links between network entities, thus naturally enabling random dropout and preventing overfitting.

[0010] According to another aspect of the present invention, a method is provided for generating and instantiating an augmentation tree model based on an original artificial intelligence (AI) model including an input layer and an output layer. The method can be executed by a parallel and split learning (PSL) controller. The method includes partitioning the original AI model into multiple partitions. The multiple partitions include a bottom-level partition including an input layer, a top-level partition including an output layer, and multiple intermediate-level partitions including one or more intermediate layers between the input and output layers. The partitioned original AI model forms an AI model. An augmentation tree model of the AI ​​model is also generated. The augmentation tree model includes multiple levels, each level corresponding to a partition of the original AI model. The multiple levels include a top-level level, a bottom-level level, and one or more intermediate levels. The top-level level includes the top-level partition as the root node. The bottom-level level includes multiple copies of the bottom-level partition as leaf nodes. Each intermediate level includes multiple copies of the corresponding intermediate-level partition as branch nodes. The root node is connected to multiple branch nodes of the intermediate level adjacent to the top-level level. Each branch node is connected to multiple leaf nodes or multiple branch nodes in a lower level. The root node, branch nodes, and leaf nodes form multiple nodes of the augmentation tree model. The method also includes instantiating one of the multiple nodes of the augmented tree model as one of the multiple entities in the network.

[0011] The embodiment also includes receiving computing resources from multiple nodes of the network.

[0012] The embodiment also includes receiving link information of multiple links connecting multiple nodes within the network.

[0013] In another embodiment, the partitioning of the AI ​​model is based on the computing resources of multiple network entities and the computing load requirements of each of the multiple partitions.

[0014] In another embodiment, the instantiation of one of the plurality of nodes of the augmented tree model into one of the plurality of entities of the network is based on matching the computing resources of the one of the plurality of entities with the computing load requirements of the one of the plurality of nodes and on matching the link information of the plurality of links connected to the one of the plurality of entities to the communication requirements of the one of the plurality of nodes.

[0015] In another embodiment, the plurality of leaf nodes connected to the same branch node have similar available computing resources. In other words, two nodes or computing entities can have similar available computing resources if the difference in computing resources between them is within a predetermined threshold.

[0016] In another embodiment, multiple leaf nodes can communicate with each other to transmit parameters of the multiple leaf nodes, or multiple branch nodes at the same level can communicate with each other to transmit parameters of the multiple branch nodes.

[0017] The embodiment also includes receiving a PSL request, the PSL request including partitioning requirements for the original AI model and augmentation tree model requirements.

[0018] In another embodiment, the PSL request also includes model parameters of the original AI model, including any one of the weights and biases of neurons and links, dropout rates associated with layers, and inter-layer dependency information of adjacent layers.

[0019] The embodiment also includes the following: One or more intermediate-level branch nodes or the root node receive forward propagation (FP) data from adjacent lower intermediate-level connected branch nodes or adjacent lower bottom-level connected leaf nodes. Alternatively, one or more intermediate-level branch nodes receive backward propagation (BP) data from adjacent higher intermediate-level connected branch nodes or adjacent higher top-level connected root nodes. Alternatively, a leaf node receives backward propagation (BP) data from adjacent higher intermediate-level connected branch nodes. Alternatively, a node synchronizes its parameters to a node at the same level as the node, which is one of the root node, the branch node, or the leaf node.

[0020] According to another aspect of the present invention, a method for dynamically dropping neuron or link updates during the training of a distributed AI model is provided, wherein the distributed AI model is divided into multiple partitions, and at least two adjacent partitions of the multiple partitions are deployed on different network entities interconnected by network links. The method is performed by a dynamic dropout controller. The method includes: receiving a dynamic dropout control (DRC) request including information about a target node, wherein the target node is one of multiple nodes in a partition of the distributed AI model; also sending a request to the network entity deploying the target node for interruption probability information of the links connecting the target node to the adjacent partitions of the target node; then receiving the interruption probability information from the network entity deploying the target node, and calculating a dropout rate metric for the target node based on the interruption probability information. The dropout rate metric is used to determine the average dropout rate to be applied to neurons on the input layer of the target node. The dropout rate metric is also sent to the network entity deploying the target node, wherein the dropout rate metric configures the target node to perform dynamic dropout during the training of the distributed AI model.

[0021] In another embodiment, the average discard rate includes a set of available values.

[0022] In another embodiment, the drop rate metric is associated with the link in the distributed AI model that connects the target node to the adjacent partition.

[0023] In another embodiment, the discard rate metric is calculated to meet the discard requirements of the distributed AI model.

[0024] In another embodiment, the distributed AI model includes an augmented tree model.

[0025] In another embodiment, the interruption probability information includes the interruption probability information of the links in the augmented tree model that connect the target node to the child nodes of the target node.

[0026] In another embodiment, the drop rate metric is associated with one of the links in the augmented tree model that connect the target node to the child nodes of the target node.

[0027] According to another aspect of the invention, an apparatus is provided for generating and instantiating augmented tree models based on an original artificial intelligence (AI) model. The apparatus includes a parallel and split learning (PSL) controller, the PSL controller including a processor and a tangible, non-transitory computer-readable memory for performing methods as defined in any of the methods described above.

[0028] According to another aspect of the invention, a system is provided for generating and instantiating augmented tree models based on original artificial intelligence (AI) models. The system includes one or more computers, each computer including a processor and a tangible, non-transitory computer-readable storage device. The computer-readable storage device includes instructions recorded thereon, which will be executed by the one or more computers of the system to perform methods as defined in any of the methods described above.

[0029] According to another aspect of the invention, a tangible, non-transitory computer-readable storage medium is also provided, on which instructions are recorded, which will be executed by at least one processor to perform a method as defined in any of the methods described above.

[0030] Embodiments of the present invention may provide technical advantages or benefits.

[0031] The implementation addresses key challenges in the field of parallel split learning. Through the development of reinforcement tree models, parallel split learning is enabled while incorporating new parameters such as inter-layer dependencies during the AI ​​model splitting process. Reinforcement tree models support the transformation of any AI model (including neural networks) into distributed AI models by partitioning the model into partitions arranged from the bottom input layer to the top output layer. This model utilizes physical network entities to realize its logical concepts.

[0032] The implementation introduces a dynamic method for customizing the dropout rate, which adjusts based on network connectivity. This strategy leverages network connectivity interruptions to naturally achieve random dropout, which subsequently improves training speed and helps avoid overfitting. Furthermore, the implementation network can adapt not only to boosted tree models but also to other distributed AI model variants. Importantly, the dynamic dropout technique extends to a wide variety of distributed AI models implemented within the network. Therefore, this invention effectively utilizes the distributed resources supporting the network for large-scale model training within the NET4AI network.

[0033] This embodiment provides a cost-effective method for training large AI models in networks, potentially accelerating model training in NET4AI. It is significant for industry standards, as split learning and distributed learning are being established as the basis for future standards by entities such as 3GPP. Therefore, this method is likely to gain a substantial share of the current large-scale model training market, dominated by AI giants.

[0034] Embodiments have been described above in conjunction with various aspects of the present invention, and these embodiments can be implemented based on these aspects. It will be understood by those skilled in the art that embodiments can be implemented in conjunction with the aspects described therein, but may also be implemented together with other embodiments of said aspects. It will be apparent to those skilled in the art that embodiments are mutually exclusive or contradictory. Some embodiments may be described in conjunction with one aspect, but may also be applicable to other aspects, as will be apparent to those skilled in the art. Attached Figure Description

[0035] Further features and advantages of the invention will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which:

[0036] Figure 1 The diagram illustrates AI model splitting within a parallel and split learning (PSL) framework according to an embodiment.

[0037] Figure 2 An augmented tree model in the PSL framework according to an embodiment is shown.

[0038] Figure 3 An example of a PSL framework according to an embodiment is shown.

[0039] Figure 4 A flowchart illustrating the generation and instantiation of an enhanced tree model according to an embodiment is shown.

[0040] Figure 5 The structure for training the augmented tree model PSL according to an embodiment is shown.

[0041] Figure 6 A comparison of convergence speed analysis between cluster-based SL training and augmented tree model PSL training is shown.

[0042] Figure 7 A flowchart of dynamic discard control according to an embodiment is shown.

[0043] Figure 8 An electronic device according to an embodiment is shown.

[0044] It should be noted that the same features are identified by the same reference numerals throughout the accompanying drawings. Detailed Implementation

[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0046] Typically, artificial intelligence (AI) models can be divided into two main components: an encoder and a classifier. The encoder, comprising the lower layers, is responsible for capturing fine-scale, low-order features and is usually implemented on devices or clients with limited processing power. In contrast, the classifier, comprising the upper layers, captures coarse-scale, high-order features and is usually implemented on the cloud or server side with more robust processing power.

[0047] Split learning (SL), as used in this article, refers to the fundamental technology for enabling the distributed deployment of large-scale AI models. For example, an original AI model can be split into an encoder (partition 1) and a classifier (partition 0). The encoder consists of an input layer and one or more hidden layers (layers between the input and output layers) that may follow the input layer. The classifier consists of the remaining hidden layers (if any) and the output layer.

[0048] Typically, the encoder captures low-order or coarse-scale features of the original AI model, while the classifier captures high-order or fine-scale features and tends to be object-specific. Layers that break down the original AI model into encoder and classifier are called cut layers. The encoder contains cut layers (as the output layer of the encoder), while the classifier contains layers following the cut layers (as the input layer of the classifier). The encoder and classifier of the AI ​​model can be instantiated in different entities (e.g., servers, devices). Intermediate data used for AI model training or inference (e.g., forward training / inference data, backpropagation data) is transmitted via network connections / links, interacting between the encoder and classifier.

[0049] In split-learning models involving multiple encoders or clients, the training process must proceed sequentially among the clients, typically following five key steps. First, encoder or client X provides intermediate results to the classifier (forward propagation). Second, the classifier propagates the gradients back via backpropagation. Third, encoder or client X synchronizes its weights with encoder or client Y to facilitate communication between the two. Fourth, encoder or client Y then sends its intermediate results (forward propagation) to the classifier. Finally, the classifier propagates the gradients back to encoder or client Y via backpropagation. This sequential process ensures that all encoders or clients contribute to the model's learning, ultimately resulting in a more comprehensive and accurate AI model.

[0050] The cluster-based parallelization used in this paper refers to an advanced technique employed in split-learning training to improve efficiency. This method involves classifying client devices into one or more clusters based on resource availability and network condition similarity. Within each cluster, devices undergo parallel training, while sequential training is performed across different clusters. As training progresses, earlier clusters update their partitioning parameters and share them with all devices in subsequent clusters. This allows later clusters to efficiently perform parallel training to incorporate updated information. Overall, this method optimizes the split-learning process, thereby improving performance and accelerating training time.

[0051] This invention provides a method and system for augmented tree models using parallel and split learning (PSL). This method and system can leverage the distributed resources of the supporting network to train very large models while minimizing training latency.

[0052] Regarding this invention, the reinforcement tree model is generated by splitting an AI model into multiple partitions, allowing the reinforcement AI model to be represented by a reinforcement tree structure, where each level of the tree corresponds to a partition. In the reinforcement tree model, each branch node corresponds to an AI enabler, and each leaf node corresponds to a client. Furthermore, server-side partitions can be deployed in a network, and parallelization can be enabled.

[0053] Typically, reinforcement trees have two types of logical links: the first is between nodes within the same partition (level) to synchronize their parameters; the second is between nodes in different partitions (levels) to transmit forward-propagation (FP) and back-propagation (BP) data. In this paper, FP data refers to the process of moving computations in the model from input to output to generate predictions. On the other hand, BP data refers to the methods used in neural networks to compute gradients, which are needed to calculate the weights to be used in the network. Parallel training, as used in this paper, refers to the process of training different nodes simultaneously, which makes the training process faster and more efficient.

[0054] In this embodiment, the relationship between the distributed AI model, the reinforcement tree model, and the network that implements them is as follows:

[0055] First, the transformation of any AI model (including neural networks) into a distributed AI model can be achieved by employing a split learning technique. This method involves dividing the AI ​​model into multiple partitions arranged sequentially from the bottom level (including the input layer of the AI ​​model) to the top level (including the output layer). The partitions of the distributed AI model are located at any of the zero to multiple layers between the bottom-level partitions and the top-level partitions. These partitions constitute logical or model concepts and require physical network entities to implement.

[0056] Second, the reinforcement tree model represents a proposed approach to instantiating distributed AI models. This is achieved by training multiple copies of each partition level in parallel. In the reinforcement tree model, all nodes—whether leaf nodes, root nodes, branch nodes, child nodes, or parent nodes—are copies of their corresponding partitions. Each node in the reinforcement tree model, like a partition in a distributed AI model, is a logical or model concept and requires a physical network entity to implement its deployment.

[0057] Finally, the network not only has the ability to implement reinforcement tree models, but also other variants of distributed AI models within its entities. It should be noted that the application of dynamic discarding is not limited to reinforcement tree models, but extends to all forms of distributed AI models implemented in the network.

[0058] Figure 1The diagram illustrates the splitting of an AI model within a Parallel and Split Learning (PSL) framework according to an embodiment. In Parallel and Split Learning (PSL), the original AI model 1000 can be split into multiple partitions. A low-level partition 1010 (e.g., partition 0) includes an input layer and possibly one or more hidden layers following the input layer of the original AI model; this low-level partition may correspond to the encoder in a conventional split learning (SL) model. A top-level partition 1040 (e.g., partition C) includes an output layer and possibly one or more hidden layers preceding the output layer of the original AI model; this top-level partition may correspond to the output portion of a classifier in a conventional SL model. Each intermediate partition (e.g., partition 1 1020 or partition 2 1030 between partition 0 1010 and partition C 1040) may include one or more hidden layers of the original AI model; this intermediate partition corresponds to a portion of the classifier in a conventional SL model. Depending on different splitting decisions, the number of intermediate partitions can range from 0 to the total number of hidden layers in the original model.

[0059] As used in this paper, in terms of the forward propagation (FP) direction, the adjacent partition before a given partition is called the preceding partition, and the adjacent partition after a given partition is called the following partition (e.g., for partition 1, partition 0 is the preceding partition and partition 2 is the following partition). The layer that splits the partitions and the following partitions is called the split layer 1050, which is contained in the partitions that serve as the output layer.

[0060] Unlike traditional SL models that may have only one cutting layer, the PSL framework can include one or more cutting layers to split the original AI model. Therefore, PSL can split the original AI model into finer-grained partitions, each requiring fewer computational resources than a coarse-grained classifier in traditional SL.

[0061] Embodiments of the present invention support parallel training of partitions within a PSL framework. In the PSL framework, the original AI model can be instantiated as a boosted tree model, which consists of partitions (of the original AI model) instantiated on distributed network entities.

[0062] The reinforcement tree model can be represented as a hierarchical reinforcement tree 2000, such as... Figure 2As shown. Each level of tree 2000 corresponds to a partition of the original AI model, and each node in that level corresponds to a copy of the partition. It should be noted that the copy of the partition should have the same model structure as the partition. In the augmented tree model, the root node 2040 can be a copy of partition C1040, or can be based on partition C1040; each leaf node 2010 can be a copy of partition 01010, or can be based on partition 01010; and each branch node 2020 or 2030 can be a partition corresponding to the level of the branch node, or can be based on the partition corresponding to the level of the branch node.

[0063] The root node 2040, or each branch node, can have one or more child nodes, which are copies of the preceding partitions. Each leaf node, or each branch node, has a parent node, which is a copy of the following partition. In the augmented tree model, all leaf nodes should have the same depth, which ensures that input data from any leaf node can traverse all partitions of the original AI model until the output layer of partition C is reached.

[0064] In the context of the augmented tree model, root node 2040 can be a copy of partition C 1040, which can be implemented on the server side. Similarly, leaf node 2010 is a copy of partition 0 1010, which can be instantiated on the device or client side. Branch nodes such as 2030 or 2020 are copies of other partitions 1020 or 1030, which can be implemented on various entities in the network.

[0065] In some embodiments, the network operation entity can also be used to instantiate leaf nodes. If one or more clients belong to the same category or cluster, the leaf nodes instantiated on the clients should all be connected back to the same parent node in the augmented tree model 2000.

[0066] In an embodiment, edges, lines, or connections in the enhanced tree model 2000 can exchange intermediate data between child nodes and parent nodes. This intermediate data includes forward propagation (FP) or inference data and backward propagation (BP) data.

[0067] In this embodiment, the intermediate data is instantiated via network connections or links. These network connections or links can also exist between nodes at the same level. The purpose of this is to synchronize or unify the updated weights and biases of each copy of a partition during parallel training.

[0068] Compared to traditional SL models that only support parallel training of client partitions (encoders), the augmented tree model in the PSL framework supports parallel training of client partitions (partition 0) and higher-level partitions. Thus, when combined with... Figure 5 As mentioned above, training efficiency can be further improved in the augmented tree model.

[0069] The embodiments include a method using a parallel and split learning (PSL) framework (i.e., a boosted tree model). This framework serves a dual or multiple purpose. It supports the generation and instantiation of boosted tree models, as well as the dynamic customization and configuration of instantiated partitions, which are represented as nodes in the boosted tree model. Furthermore, the method leverages the distributed resources of the supporting network to train large-scale models, thereby minimizing training latency.

[0070] Figure 3 An example of a PSL framework 3000 according to an embodiment is shown. The PSL framework 3000 includes a PSL controller 3010 and a partition customization function (PCF) 3070 on each entity of one or more nodes of an augmented tree model.

[0071] The PSL controller 3010 used herein may be a logical controller that can be instantiated centrally (e.g., in a cloud environment or on a server) or on a distributed network entity. In some embodiments, all or one of the functions that make up the PSL controller 3010 may be instantiated as internal functions of a network control module or function, such as the Service Control Function (SCF) in NET4AI, a system architecture designed to support computing services such as AI computing services.

[0072] refer to Figure 3 The PSL controller 3010 may include an augmented tree model control (ATC) function 3020. It should be noted that ATC 3020 can split the original AI model to generate an augmented tree model. This process is based on the characteristics of the original AI model and the availability of computing resources within the network. Furthermore, it determines the network's instantiation decision for the augmented tree model.

[0073] In an embodiment, the PSL controller 3010 may include a Dynamic Dropout Control (DRC) function 3030. Random dropout is commonly used in large model training to prevent overfitting. This technique randomly silences or “quiets” a subset of neurons in each layer during each batch of training, thereby interrupting the forward (FP) and backward (BP) propagation of these neurons. Within the PSL framework 300, unstable network connectivity (e.g., using wireless connectivity) can naturally facilitate random dropout by losing FP or BP data from the front or back partitions. To achieve this network connectivity-based random dropout, the DRC 3030 may set a dropout rate metric and send this metric to each entity of the instantiated nodes of the augmented tree model.

[0074] In this embodiment, the PCF 3070 can be instantiated in each entity of the nodes that instantiate the augmented tree model (e.g., in each of 2020, 2030, and 2040). It should be noted that the PCF is not needed in the leaf nodes. The PCF 3070 can receive drop rate metrics from the DRC 3030 and then locally customize real-time drop decisions based on network conditions and the received drop rate metrics. This is in... Figure 7 Further details are provided below.

[0075] refer to Figure 3 The PSL framework 3000 includes the following three interfaces. The ATC-DRC interface (A2D interface) 3040 used in this paper refers to the communication channel between the Augmented Tree Model Control (ATC) function 3020 and the Dynamic Dropout Control (DRC) function 3030 within the Parallel and Split Learning (PSL) framework. It allows these two components to exchange information, such as dropout rates and model splitting details, to improve the overall effectiveness and efficiency of the learning process.

[0076] As used herein, the ATC-to-Network Entity Interface (A2N Interface) 3050 refers to the communication path connecting the ATC 3020 to network entities (including servers and clients in some embodiments) where partitions of the augmented tree model are instantiated. This interface allows the ATC to provide instructions regarding model partitions and control their instantiation on network entities. In some embodiments, the A2N Interface 3050 may be implemented as an interface between a data plane network entity and a network controller. Thus, messages related to augmented tree model generation and embedding are transmitted via the A2N Interface.

[0077] The DRC-to-Network Entity Interface (D2N Interface) 3060 used herein facilitates communication between the Dynamic Drop Control (DRC) function 3030 and network entities. These entities, in some embodiments, include clients and instantiate branch nodes or root nodes of the augmented tree model within a parallel and split-learning (PSL) framework. The D2N Interface 3060 allows DRC to efficiently assign drop rate metrics to entities embodying nodes in the augmented tree model. This connection facilitates skillful management of random drop processes across the entire network entity. In some embodiments, the D2N Interface 3060 is implemented via a connection between the data plane network entity and the network controller. Therefore, messages related to dynamic drop control are transmitted through the D2N Interface 3060 to ensure a well-coordinated and efficient drop process.

[0078] Figure 4A flowchart illustrating the process of generating and instantiating an augmented tree model according to an embodiment is shown. In step 4060, an augmented tree model user 4010 provides or sends a PSL request to ATC 4020 to trigger the generation and instantiation of the augmented tree model. In this embodiment, the augmented tree model user 4010 may be an internal network function (e.g., the task control function of the NET4AI service) or a third-party application. The PSL request may include original AI model information, partition requirements, augmented tree model requirements, etc.

[0079] In embodiments, the original AI model information typically includes structural information of the original AI model and (optionally) model parameters of the original AI model. For example, the model parameters of the original AI model may include (1) weights and biases on neurons and links of the original AI model, (2) dropout rates associated with each layer of the original AI model, and (3) inter-layer dependency information between adjacent layers (e.g., metrics associated with each layer indicating whether the layer can be selected as a cut layer), which can be obtained from prior knowledge of the original AI model. In some embodiments where model parameters are not provided in the original AI model information, the ATC may determine the model parameters of the original AI model based on predefined knowledge or algorithms.

[0080] In this embodiment, the partitioning requirements include requirements and constraints for splitting the original AI model into partitions, such as (but not limited to) the maximum or minimum number of partitions, and the maximum or minimum size. In other words, the maximum or minimum number of layers in partition 0, partition C, and other partitions, and the maximum or minimum number of neurons in each layer of the partition.

[0081] In this embodiment, the requirements for the augmented tree model may include requirements and constraints for the generation and instantiation of the augmented tree model, such as (but not limited to) the maximum number of copies of the corresponding partition in each level; the minimum or maximum number of clients in the cluster; the maximum number of child nodes of a branch node or root node; the minimum computing resource requirements for instantiating nodes by an entity, server, or device; and the network location or address of the root node instantiated by the server. It should be noted that the requirements for the augmented tree model may also include performance requirements for the augmented tree model, such as (but not limited to) the maximum convergence time threshold.

[0082] After step 4060, upon receiving the PSL request, the ATC 4020 (via the A2N interface) interacts with all available devices, entities, and servers 4030 in the network to collect information on available computing resources and links in the network, which is used to generate and instantiate the augmented tree model.

[0083] In this embodiment, the entity information collection process may include client information collection, network entity information collection, and server information collection. For example, the process typically begins with client information collection, such as... Figure 4 Steps 4070 and 4080 are shown in the diagram.

[0084] refer to Figure 4 For client information collection, ATC 4030 sends one or more client information requests to all or one or more available devices 4050 in step 4070. Upon receiving a client information request, each device 4050 sends the client information back to ATC 4020 in step 4080. In this embodiment, the client information includes information about available computing resources on the device, as well as the device's physical location or network address.

[0085] refer to Figure 4 For network information collection, ATC 4020 sends one or more network entity information requests to all available network entities 4040 in step 4090. Upon receiving a network entity information request, each network entity 4040 sends its network entity information back to ATC 4020 in step 4100. In this embodiment, the network entity information includes information about available computing resources on the network entity, and the physical location or network address of the network entity.

[0086] In some embodiments where the network entity can sense information about the connected network links (e.g., probability of interruption, available bandwidth, wireless channel path loss), the information about the connected network links of the entity is also included in the network entity information sent to the ATC.

[0087] refer to Figure 4 For server information collection, ATC 4020 sends one or more server information requests to all or one or more available servers 4030 in step 4110. Upon receiving a server information request, each server 4030 sends server information back to ATC 4020 in step 4120. In this embodiment, the server information includes information about available computing resources on the server, as well as the server's physical location or network address.

[0088] In an embodiment of the client clustering process, in step 4130, ATC 4020 classifies clients into different clusters based on the client information received in step 4080 (using a predefined clustering algorithm in ATC). It should be noted that ATC 4020 ensures that clients within the same cluster have similar computing capabilities within a predefined threshold.

[0089] In an embodiment of the augmented tree model generation process, in step 4140, ATC 4020 generates an augmented tree model (with information) and determines an augmented tree instantiation decision based on the received client information, network entity information, server information, and PSL request using a predefined algorithm or one or more AI schemes.

[0090] Therefore, the information of the generated augmented tree model 2000 can include hierarchical augmented tree structure information (i.e., the ID of each node in the tree, the height of the tree, the number of nodes in each level, and the parent and child node IDs of each node in the tree). Furthermore, the information of the generated augmented tree model can include information about the partitions associated with each level of the tree (i.e., information about the partition structure, partition parameters including weights and biases on neurons and links within the partition, and the dropout rate associated with each level in the partition).

[0091] Furthermore, the enhancement tree instantiation decision determined in step 4140 associates the ID or network address or location of each device, network entity, or server with the ID of the enhancement tree node (a copy of the partition) to be instantiated on the device, network entity, or server.

[0092] In an embodiment of the augmented tree model instantiation process, after obtaining the augmented tree instantiation decision for solving the joint optimization problem described below (represented by equation (1)), one or more root node instantiation decisions, one or more branch node instantiation decisions, and one or more leaf node instantiation decisions are transmitted.

[0093] refer to Figure 4 In step 4150, ATC 4020 (via the A2N interface) sends a root node instantiation decision to the server 4030 that instantiated the root node. In this embodiment, the root node instantiation decision includes information about partition C, and the ID, network address, or location of the network entity of one or more child nodes of the instantiated root node. Based on the received root node instantiation decision, the server instantiates a copy of partition C and configures an interface or connection to the network entity of the child node of the instantiated root node. In some embodiments, the server 4030 may send an acknowledgement (ACK) message to ATC 4020 after successfully instantiating a copy of partition C and configuring the interface or connection.

[0094] In step 4160, the ATC 4020 (via the A2N interface) sends a branch node instantiation decision to the network entity that instantiates the associated branch node. In an embodiment, the branch node instantiation decision includes information about the partition associated with the branch node's hierarchy, and the IDs or network addresses / locations of clients, network entities, or servers of the instantiated child node, parent node, and other branch nodes in the same hierarchy. Based on the received branch node instantiation decision, the network entity instantiates a copy of the partition and configures interfaces or connections to clients, network entities, or servers of the instantiated child node, parent node, and other nodes in the same hierarchy. In some embodiments, the network entity may send an ACK message to the ATC after successfully instantiating a copy of the partition and configuring the interface or connection.

[0095] In step 4170, the ATC 4020 (via the A2N interface) sends a leaf node instantiation decision to the client that instantiated the leaf node. In this embodiment, the leaf node instantiation decision includes information about partition 0, the IDs or network addresses or locations of the parent node of the instantiated leaf node and other network entities or clients of the other leaf nodes, and the ID of the cluster to which the client is associated. Based on the received leaf node instantiation decision, the client instantiates a copy of partition 0 and configures interfaces or connections to the parent node of the instantiated leaf node and other network entities or clients of the other leaf nodes. In some embodiments, the client may send an ACK message to the ATC after successfully instantiating a copy of partition 0 and configuring any interfaces or connections.

[0096] Subsequently, regarding the processing of the reinforcement tree model instantiation result ACK, in step 4180, ATC 4020 can send the reinforcement tree model instantiation result to reinforcement tree model user 4010 to indicate that the reinforcement tree model instantiation was successful (in other words, to confirm that the reinforcement tree model instantiation is complete). In an embodiment, the reinforcement tree model instantiation result may include information about the generated reinforcement tree model and the reinforcement tree instantiation decision.

[0097] As described above, in some embodiments, ATC can generate a boosting tree model and determine the boosting tree instantiation decision by solving a joint optimization problem (represented by equation (1)) that aims to minimize the overall training latency of the boosting tree model. For example, the joint optimization problem can be expressed as:

[0098]

[0099] So that,

[0100]

[0101] In optimization problems, Indicator of the reinforcement tree model Hierarchy Indicates a cluster of nodes that share the same parent node. Instruction cluster The first in 1 node Instruction cluster The maximum allowed number of child nodes at the parent node level The number of nodes constrained. In an embodiment, decision parameters may include... (include ), representing the computational load at partition 0 and other partitions of the original AI model, respectively, and furthermore This represents the available computing resources on a network entity or client.

[0102] In an embodiment, constraints may include those related to each partition. Associated communication load Available communication bandwidth of the link between a node and its parent node. The constraints include inter-layer dependency constraints (obtainable from the PSL request) and computational compatibility between a node and its child nodes (represented by the first constraint in the optimization problem). According to the computational compatibility constraint, a branch node with child nodes should have more computational resources compatible with the computational resources of all its child nodes. For example, for a branch node with four child nodes, if each leaf node uses one time unit to process one FP or BP, then the branch node must be able to process four FPs or BPs in one time unit to prevent latency from parallel training of the augmented tree model.

[0103] In an embodiment, the computing resources of a node or computing entity may include the amount, size, or volume of computing resources, such as processing power or speed, network speed, I / O speed, storage speed or amount, etc. Similarly, it can be said that if the difference in computing resources between two nodes or computing entities is within a predetermined threshold between them, then the two nodes or computing entities may have similar available computing resources.

[0104] In real-world scenarios, the joint optimization problem above may take different forms, as long as it includes the following features or parameters, with the goal of minimizing the overall training latency of the boosting tree model.

[0105] The computational load at each partition of the AI ​​model (e.g.) and ) and the available computing resources on each network entity or client of one or more partitions implementing the AI ​​model (represented as These are two important optimization variables that need to be determined to achieve the optimization objective. Since the original AI model is predetermined, each model partitioning or splitting decision can be mapped to the computational load of a set of partitions. Therefore, the computational load at each partition is equivalent to the partitioning or splitting decision of the AI ​​model.

[0106] In other words, the necessary parameters for formulating and solving the joint optimization problem include the communication load associated with each partition (e.g., ), and the available communication bandwidth of the link between the node and its parent node (e.g., And inter-layer dependency constraints, which are implicit in the available partitioning or splitting choices of the AI ​​model.

[0107] Constraints associated with the joint optimization problem include that branch nodes with one or more child nodes should have more computational resources compatible with the computational resources of all or more child nodes. For a branch node with four child nodes, if each leaf node uses one time unit to process one FP or BP, then the branch node must be able to process four FPs or BPs in one time unit to avoid latency in parallel training of the boosting tree model. Another constraint is that the number of nodes in the cluster is constrained by the maximum allowed number of child nodes at its parent node level, represented by the second constraint in the optimization problem.

[0108] like Figure 5 As shown, the Parallel and Split Learning (PSL) framework 5000 accelerates training by allowing parallel training on all nodes at the same level in the boosting tree model. This contrasts with cluster-based parallel training for split learning. As previously mentioned, for cluster-based parallel training for split learning, each cluster can perform one training iteration in parallel. That is, all clients simultaneously perform forward-propagation (FP) and backward-propagation (BP) of the training iteration, and then synchronize the encoder parameters (e.g., by averaging the encoder parameters across all clients). Training is sequential between clusters: once a cluster completes its parallel training iteration, it updates the synchronized encoder parameters to all devices in subsequent clusters. The cluster then performs its parallel training iteration and repeats the process until the split learning (SL) model converges. Examples can be used as follows... Figure 5 The Parallel and Split Learning (PSL) framework shown in the diagram works effectively.

[0109] like Figure 5 As depicted, except for the root node at the top level, all branch nodes at the same level can perform FP and BP in parallel, provided they receive the corresponding FP and BP data from their associated child and parent nodes, respectively. Furthermore, all leaf nodes connected to different clusters can perform FP and BP in parallel, and after one training iteration, all branch nodes synchronize their local partitioning parameters with other branch nodes at the same level, and all leaf nodes synchronize their local partitioning parameters with other leaf nodes (indicated by double-headed dashed arrows).

[0110] Figure 6 The training time consumption of cluster-based parallel training (SL) and augmented tree model parallel training (in other words, a convergence speed analysis of cluster-based parallel SL training and augmented tree model parallel training) was compared. Both models used the same client-side clustering results and partitioning decisions to ensure a fair comparison.

[0111] like Figure 6As shown, the parameter synchronization time in the augmented tree model parallel training 6010 is longer than that in the cluster-based parallel SL training 6000. This difference stems from the exchange of additional parameter synchronization data between branch nodes. However, due to sequential training across different clusters, the cluster-based parallel SL training 6000 requires three rounds of FP / BP time to complete the training of all clients classified into three clusters. In contrast, the augmented tree model parallel training 6010 only requires one round of FP / BP time because all clients in the three clusters are trained simultaneously. Therefore, the cluster-based parallel SL training 6000 consumes more training time than the augmented tree model parallel training.

[0112] Figure 7 A flowchart of dynamic discard control according to an embodiment is shown, wherein dynamic discard control is performed through the following process or steps.

[0113] First, in step 7040, a dynamic drop request is initiated, which includes the ID, network location, or address of the target network entity or server to which dynamic drop is to be applied. In this embodiment, the node instantiated in the augmented tree model of the target network entity or server is called the target node. In step 7040, the ATC 7010 or the dynamic drop user sends the dynamic drop request to the DRC 7020 via the A2D interface. In some embodiments, the dynamic drop user, which may be an internal network function or a third-party application, may also send the dynamic drop request to the DRC 7020.

[0114] The next stage (step 7050) is network link information collection. Upon receiving a dynamic drop request, in step 7050, DRC 7020 (via the D2N interface) sends a network link information request to each target network entity or server 7030 to perform dynamic drop. In step 7060, each target network entity / server that received the request sends the network link information back to DRC 7020. In some embodiments, the network link information includes the statistical probability of network link interruption for network entities connecting the target network entity or server to child nodes of the instantiated target node.

[0115] The third step (step 7070) involves configuring the drop rate metric. Based on the network link information received from each target network entity or client, the DRC 7020 calculates the drop rate metric for each target node. This metric is used to determine the average drop rate applied by the node at the input layer, and can be a single value or a set of available values ​​used as the average drop rate. It should be noted that each physical link corresponds to one drop rate metric. For multiple child nodes sharing the same physical link, a corresponding drop rate metric is assigned to that physical link. The drop rate metric is calculated to ensure the overall AI model drop requirements (known features of the original AI model) and to ensure that the corresponding link can support the drop rate. For example, a link with a statistical outage probability of 0.3 cannot support a drop rate greater than 0.7 indicated in the drop rate metric. In some embodiments, a predefined algorithm or scheme can be used to calculate the drop rate metric. After determining the drop rate metric for each target node, in step 7070, the DRC 7020 (via the D2N interface) sends the associated drop rate metric to each target entity or server.

[0116] The fourth process (step 7080) is dynamic drop-off execution. In step 7080, given the received drop-off rate metric and the sensed real-time physical link status, the PCF instantiated on each target network entity or server 7030 can locally customize drop-off decisions. This decision specifies which neurons or links should be temporarily silenced or deleted in each training iteration, thereby achieving dynamic drop-off while ensuring the average drop-off rate indicated by the drop-off rate metric.

[0117] Finally, in step 7090, a dynamic drop instantiation acknowledgment (ACK) occurs. In some embodiments, in step 7090, the PCF 7030 performing the dynamic drop may send a dynamic drop instantiation ACK to the DRC 7020 to indicate that the dynamic drop was successfully performed on the target network entity or server. After receiving the dynamic drop instantiation ACK, the DRC 7020 may forward the dynamic drop instantiation ACK to the ATC or dynamic drop user who sent the dynamic drop request 7010.

[0118] Compared to traditional random dropout with a fixed dropout rate at each layer, dynamic dropout leverages the interruption or change of network links between network entities / servers that instantiate boosting tree model nodes to naturally achieve random dropout, thereby preventing overfitting (the original intention of random dropout) and improving the training performance of boosting tree models (if traditional random dropout is applied in boosting tree model training, the accidental loss of FP and BP data from un-silenced neurons or links will reduce training performance).

[0119] Embodiments of this invention can address practical problems in the field of parallel split learning implementation. For example, the augmented tree model used herein is designed to support parallel split learning while considering new parameters such as inter-layer dependencies when splitting the AI ​​model. Furthermore, as used herein, the method for dynamically customizing the dropout rate can be adjusted based on the network connectivity status.

[0120] These embodiments present a number of technical advantages or benefits. Primarily, they allow for the use of distributed resources supporting the network to train a large number of models in the NET4AI network. Subsequently, they leverage network connectivity disruptions between partitions to naturally implement random dropouts, thereby accelerating training and helping to prevent overfitting.

[0121] From a business perspective, these methods offer a cost-effective strategy for training large models in networks and accelerate AI model training in NET4AI. As for industry standards, split learning and distributed learning are integral components of future standards from entities like 3GPP. In other words, split learning and distributed learning will form the cornerstone of future standards from entities like 3GPP in processing standards. The commercial appeal of these methods lies in their cost-effectiveness in training large models in networks and their potential to accelerate AI model training in NET4AI. These methods may seize market share from the large-scale model training market currently controlled by AI giants.

[0122] Figure 8 An apparatus, such as electronic device 800, according to one embodiment is shown, which can perform any or all operations of the methods and features of one or more aspects of the present invention, explicitly or implicitly described herein.

[0123] As shown in the figure, device 800 may include a processor 802, such as a central processing unit (CPU) or a dedicated processor (e.g., a graphics processing unit (GPU) or other such processor unit), a memory 803, a non-transient mass storage 804, an input-output (I / O) interface 809, and a network interface 806, all of which are communicatively coupled via a bidirectional bus 805. The I / O interface 809 can be connected to various I / O devices 810 as needed for each configuration of the electronic device 800. Similarly, the network interface 806 can be connected to various networks, such as network 807.

[0124] Depending on certain aspects, any or all of the depicted elements may be utilized, or only a subset of these elements may be used. Furthermore, the components of electronic device 800 may include multiple instances of certain elements, such as multiple processors, memories, or transceivers. Additionally, elements of the hardware device may be directly coupled to other elements without requiring bidirectional bus 805. Furthermore, in addition to processor 802 and memory 803, other electronic devices such as integrated circuits or ASICs may be employed to perform the required logical operations.

[0125] Memory 803 may include any type of non-transitory memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), read-only memory (ROM), or any combination thereof. Mass storage element 804 may include any type of non-transitory storage device, such as solid-state drive, hard disk drive, disk drive, optical disk drive, USB drive, or any computer program product for storing data and machine-executable program code. According to some aspects, memory 803 or mass storage 804 may record statements and instructions executable by processor 802 thereon for performing any of the methods described herein.

[0126] Embodiments of the present invention can be implemented using electronic hardware, software, or a combination thereof. Some embodiments can be implemented by one or more computer processors that execute program instructions stored in memory. In some embodiments, such as using one or more field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs) to rapidly perform processing operations, the present invention is implemented in part or in whole in hardware.

[0127] The actions associated with the methods described herein can be implemented as coded instructions in a computer program product. In other words, a computer program product is a computer-readable medium on which software code or instructions are recorded to perform the methods when the computer program product is loaded into memory and executed on a processor of a computing device.

[0128] Furthermore, each operation of the method can be executed on any real or virtual computing device such as a personal computer, server, tablet computer, or smartphone, based on one or more or a portion of one or more program elements, modules, or objects generated by any programming language such as C++ or Java. Additionally, each operation, or the file or object implementing each operation, can be executed by a dedicated hardware or circuit module designed for this purpose.

[0129] Obviously, the above embodiments of the present invention are examples and various variations are possible. These current or future variations should not be considered as departing from the spirit and scope of the invention, and it will be apparent to those skilled in the art that all such modifications should be included within the scope of the appended claims.

Claims

1. A method for generating and instantiating augmented tree models based on an original artificial intelligence (AI) model including an input layer and an output layer, the method being executed by a parallel and split learning PSL controller, the method comprising: The original AI model is divided into multiple partitions, including a bottom-level partition, a top-level partition, and multiple intermediate-level partitions. The bottom-level partition includes the input layer, the top-level partition includes the output layer, and the multiple intermediate-level partitions include one or more intermediate layers between the input layer and the output layer. The original AI model after division forms an AI model. The augmentation tree model that generates the AI ​​model includes multiple levels, each level corresponding to a partition of the original AI model. The multiple levels include a top level, a bottom level, and one or more intermediate levels. The top level includes the top level partition as the root node. The bottom level includes multiple copies of the bottom level partition as leaf nodes. Each intermediate level includes multiple copies of the corresponding intermediate level partition as branch nodes. The root node is connected to multiple branch nodes of the intermediate level adjacent to the top level. Each branch node is connected to multiple leaf nodes or multiple branch nodes in the lower level. The root node, the branch nodes, and the leaf nodes form multiple nodes of the augmentation tree model. Instantiate one of the plurality of nodes in the augmented tree model as one of the plurality of entities in the network.

2. The method according to claim 1, further comprising: Receive computing resources from the plurality of nodes in the network.

3. The method according to any one of claims 1 to 2, further comprising: Receive link information for multiple links connecting the multiple nodes within the network.

4. The method according to any one of claims 1 to 3, characterized in that, The partitioning of the AI ​​model is based on the computing resources of multiple network entities and the computing load requirements of each partition in the multiple partitions.

5. The method according to any one of claims 1 to 3, characterized in that, The instantiation of one of the plurality of nodes in the augmented tree model into one of the plurality of entities in the network is based on matching the computing resources of the one of the plurality of entities with the computing load requirements of the one of the plurality of nodes and matching the link information of the plurality of links connected to the one of the plurality of entities to the communication requirements of the one of the plurality of nodes.

6. The method according to any one of claims 3 to 5, characterized in that, The multiple leaf nodes connected to the same branch node have similar available computing resources.

7. The method according to any one of claims 1 to 6, characterized in that, Multiple leaf nodes can communicate with each other to transmit parameters of the multiple leaf nodes, or Multiple branch nodes at the same level can communicate with each other to transmit parameters of the multiple branch nodes.

8. The method according to any one of claims 1 to 7, further comprising: Receive a PSL request, which includes the partitioning requirements of the original AI model and the augmented tree model requirements.

9. The method according to claim 8, characterized in that, The PSL request also includes model parameters of the original AI model, which include any one of the following: weights and biases of neurons and links, dropout rates associated with layers, and inter-layer dependency information of adjacent layers.

10. The method according to any one of claims 1 to 9, further comprising: The branch nodes of one or more intermediate levels or the root node receive forward propagation FP data from the connected branch nodes of adjacent lower intermediate levels or the connected leaf nodes of adjacent lower bottom levels; or The branch nodes of one or more intermediate levels receive backpropagated BP data from the connected branch nodes of adjacent higher intermediate levels or the connected root nodes of adjacent higher top levels. or Leaf nodes receive backpropagated BP data from adjacent higher intermediate level connected branch nodes; or The node synchronizes the parameters of the node to the node at the same level as the node, wherein the node is one of the root node, the branch node, or the leaf node.

11. A method for dynamically discarding neuron or link updates during the training of a distributed AI model, said distributed AI model being split into multiple partitions, at least two adjacent partitions of said multiple partitions being deployed on different network entities interconnected by network links, said method being executed by a dynamic discard controller, said method comprising: Receive a Dynamic Drop Control (DRC) request that includes information about a target node, which is one of multiple nodes in a partition of the distributed AI model; Send a request to the network entity where the target node is deployed for interruption probability information of the links connecting the target node to the adjacent partition of the target node; Receive the interruption probability information from the network entity that deploys the target node; The dropout rate index of the target node is calculated based on the interruption probability information. The dropout rate index is used to determine the average dropout rate of the neurons applied on the input layer of the target node. The drop rate metric is sent to the network entity that deploys the target node, and the drop rate metric configures the target node to perform dynamic dropping during the training of the distributed AI model.

12. The method according to claim 11, characterized in that, The average discard rate includes a set of available values.

13. The method according to any one of claims 11 to 12, characterized in that, The drop rate metric is associated with the link in the distributed AI model that connects the target node to the adjacent partition.

14. The method according to any one of claims 11 to 13, characterized in that, The discard rate metric is calculated to meet the discard requirements of the distributed AI model.

15. The method according to any one of claims 11 to 12, characterized in that, The distributed AI model includes an augmented tree model.

16. The method according to claim 15, characterized in that, The interruption probability information includes the interruption probability information of the links in the augmented tree model that connect the target node to the child nodes of the target node.

17. The method according to any one of claims 15 to 16, characterized in that, The drop rate metric is associated with one of the links in the augmented tree model that connect the target node to the child nodes of the target node.

18. An apparatus for generating and instantiating an augmented tree model based on an original artificial intelligence (AI) model, the apparatus comprising: A parallel and split learning PSL controller includes a processor and a tangible non-transient computer-readable memory for performing the method of any one of claims 1 to 17.

19. A system for generating and instantiating augmented tree models based on original artificial intelligence (AI) models, the system comprising: One or more computers, each computer including a processor and a tangible non-transitory computer-readable memory, said tangible non-transitory computer-readable memory for implementing a parallel and split learning PSL controller that performs the method of any one of claims 1 to 17.

20. A tangible, non-transitory computer-readable storage medium comprising instructions recorded thereon, the instructions being executed by at least one processor to perform the method of any one of claims 1 to 17.