Parallel training and reasoning method of AI model and related system
By identifying repeated and non-repeated feature vectors in the parallel training and inference process, using differentiated processing and constructing a communication matrix, the problem of long communication time in parallel training and inference of AI models is solved, and end-to-end performance is improved.
Patent Information
- Application Number
- CN202410343509.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-23
- Publication Date
- 2025-09-23
AI Technical Summary
During the parallel training and inference process of AI models, the communication between multiple computing nodes takes too long, making it difficult to meet business needs. Especially when the number of parameters is large, the communication volume is large and the operation cannot be concealed, affecting end-to-end performance.
By identifying repeated and non-repeated feature vectors in parallel training and inference processes, differentiated processing is used to reduce communication volume. Local updates of repeated feature vectors are used to construct communication matrices or query vectors to reduce many-to-many communication and improve end-to-end performance.
Without losing accuracy, the communication volume between multiple computing nodes is reduced, the latency is shortened, and the training and reasoning efficiency of AI models is improved.
Smart Images

Figure CN120688654A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a parallel training method for an AI model, a parallel reasoning method for an AI model, a distributed system, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the development of AI technology, especially the continuous evolution of machine learning (ML) and deep learning (deeplearning), various AI models have emerged. Among them, AI models can be trained with large amounts of data. Once the training is complete, the trained AI models can be applied for inference. Both the training and inference stages involve a large amount of data processing. Given the large number of parameters in AI models, AI models usually use parallel methods for data processing.
[0003] Specifically, AI models can be deployed on multiple computing cards, such as neural network processing units (NPUs). AI models can process data in parallel through multiple NPUs. Data exchange is often required between multiple NPUs, which consumes a lot of communication time and makes it difficult to meet business needs. Summary of the Invention
[0004] This application provides a parallel training method for AI models. This method takes into account the repetitiveness of input data during different iterations of model training and differentiates between repeated and non-repeated feature vectors to reduce communication between multiple computing nodes. By reducing communication volume without sacrificing accuracy, end-to-end performance is improved. This application also provides a parallel reasoning method for AI models, a distributed system, a computer-readable storage medium, and a computer program product.
[0005] In a first aspect, the present application provides a method for parallel training of an AI model. The method can be applied to a distributed system. The distributed system includes a central processing unit and multiple computing nodes. Each of the multiple computing nodes deploys an AI model. The multiple computing nodes each store at least one feature vector. The multiple computing nodes include a first computing node and a second computing node. The feature vector stored by the first computing node is different from the feature vector stored by the second computing node. The first computing node and the second computing node are used to train the AI model in parallel through multiple iterations, where each iteration of the multiple iterations includes updating the parameters of the AI model once.
[0006] Specifically, the central processing unit obtains the input data of at least one computing node among the multiple computing nodes in this iteration, and the input data includes a batch of training data in the data set, and the training data includes features and labels corresponding to the features. Then, the central processing unit identifies the repeated feature vectors and non-repeated feature vectors of at least one computing node in this iteration based on the input data of at least one computing node in this iteration and the input data of at least one computing node in the previous n iterations. At least one computing node saves the repeated feature vectors of at least one computing node in this iteration during the previous n iterations. Then, the first computing node can obtain the non-repeated feature vectors of the first computing node in this iteration from the second computing node, and update the parameters of the AI model based on the non-repeated feature vectors of the first computing node in this iteration, the repeated feature vectors of the first computing node in this iteration saved by the first computing node, and the labels.
[0007] This method takes into account the repeatability of input data in different iterative processes during model training. For example, the input data of adjacent iterative processes in the recommendation scenario has strong repeatability. Based on the input data of this iteration and the input data of the previous n iterations, the repeated feature vectors and non-repeated feature vectors of this iteration are identified, and the repeated feature vectors and non-repeated feature vectors are differentiated. Non-repeated feature vectors can be obtained through communication between multiple computing nodes. Since repeated feature vectors have been saved in the current computing node during the previous n iterations, there is no need for unmasked communication operations. This can reduce communication between multiple computing nodes (such as communication between multiple cards). Without losing accuracy, the end-to-end performance can be improved by reducing the communication volume.
[0008] In some possible implementations, the first computing node may input the non-repeated feature vector of the first computing node in this iteration and the repeated feature vector of the first computing node in this iteration saved by the first computing node into the AI model for calculation, and obtain the gradient of the feature vector input to the AI model at the first computing node based on the calculation result and label of the AI model. Among them, the feature vector input to the AI model includes the first feature vector that will be reused in the next n iterations. The first computing node can obtain the gradient of the first feature vector at the second computing node through reduction communication (all reduce). The first computing node updates the parameters of the AI model based on the gradient of the first feature vector at the first computing node and the gradient of the first feature vector at the second computing node.
[0009] In this method, repeated feature vectors can be updated locally and wait for reuse in the next n iterations. Therefore, the two many-to-many (all to all) operators of repeated feature vectors, such as the gradient all-to-all communication in this iteration and the feature vector all-to-all communication in the next iteration, can be equivalently replaced by an all-reduce communication operator, realizing hot update of gradients based on repeated feature vectors. While ensuring the training accuracy of the AI model, the communication volume between multiple cards is reduced, thereby improving end-to-end performance.
[0010] In some possible implementations, when the AI model is a recommendation model, the first computing node may accumulate the gradient of the first eigenvector at the second computing node and the gradient of the first eigenvector at the first computing node, and reversely update the first eigenvector based on the accumulated first gradient. The updated first eigenvector can better represent the feature, thereby improving recommendation accuracy.
[0011] In some possible implementations, the feature vector input to the AI model includes a second feature vector that will not be reused in the next n iterations. Accordingly, the first computing node can also send the gradient of the second feature vector at the first computing node via many-to-many communication to a source computing node that pre-stores the second feature vector, such as the computing card that stored the second feature vector during the initialization phase, also referred to as the source card. The source computing node updates the second feature vector based on the gradient of the second feature vector at the first computing node.
[0012] In this method, the computing node only performs unmasked all-to-all communication for the second eigenvector that will not be reused in the next n iterations. For the first eigenvector that will be reused in the next n iterations, the gradient can be updated locally. This reduces communication between multiple cards and improves end-to-end performance.
[0013] In some possible implementations, the central processing unit identifies the first feature vector based on multiple iterations of input data collected from the dataset to obtain a recognition result. The central processing unit transmits the recognition result to the first computing node. This lays the foundation for subsequent distinguishing between repeated and non-repeated feature vectors.
[0014] In some possible implementations, the central processor may further construct a communication matrix based on the unique feature vectors of at least one computing node in the current iteration and the storage locations of the unique feature vectors. The central processor sends the communication matrix to the at least one computing node. The at least one computing node includes a first computing node. Accordingly, the first computing node may obtain the unique feature vectors of the first computing node in the current iteration from a second computing node through many-to-many communication based on the communication matrix.
[0015] This method constructs a communication matrix to instruct computing nodes to obtain non-repeated feature vectors from other computing nodes, eliminating the need to obtain repeated feature vectors through many-to-many communication, reducing unmaskable communication operations, shortening latency, and improving end-to-end performance.
[0016] In some possible implementations, the central processing unit may further construct a query vector based on the repeated feature vectors of at least one computing node in the current iteration and the storage location of the repeated feature vectors, and send the query vector to the at least one computing node. The at least one computing node may include a first computing node. The first computing node may search for the repeated feature vectors of the current iteration stored by the first computing node based on the query vector. This method constructs the query vector so that the computing node can quickly retrieve the feature vectors and use them for model training, thereby improving training efficiency.
[0017] In some possible implementations, the computing nodes include neural network processors (NPUs), graphics processing units (GPUs), or tensor processing units (TPUs). This method supports parallel training of computing nodes of different forms, with high availability and compatibility.
[0018] In some possible implementations, the distributed system includes one or more central processing units. When the distributed system includes a first central processing unit and a second central processing unit, the first central processing unit is connected to the first computing node, and the second central processing unit is connected to the second computing node.
[0019] This method supports parallel training of AI models in distributed systems with a single-machine multi-card architecture or a multi-machine multi-card architecture. It can meet the training needs of AI models of different scales and has high availability.
[0020] In some possible implementations, the feature vector is a dense vector. This dense vector can be a low-dimensional dense vector obtained by encoding high-dimensional sparse features. By converting high-dimensional sparse features into low-dimensional dense features, it is possible to facilitate matrix operations in the AI model, thereby improving the training efficiency of the AI model.
[0021] In a second aspect, the present application provides a parallel reasoning method for an AI model. The method is applied to a distributed system. The distributed system includes a central processing unit (CPU) and multiple computing nodes. Each of the multiple computing nodes deploys an AI model, and each of the multiple computing nodes stores at least one feature vector. The multiple computing nodes include a first computing node and a second computing node. The feature vector stored by the first computing node is different from the feature vector stored by the second computing node. The first computing node and the second computing node are used for parallel reasoning using the AI model.
[0022] Specifically, the central processing unit can obtain the input data of at least one computing node among multiple computing nodes in this reasoning, and the input data includes a batch of reasoning data in the data set, and the reasoning data includes features. Then, the central processing unit identifies the repeated feature vectors and non-repeated feature vectors of at least one computing node in this reasoning based on the input data of at least one computing node in this reasoning and the input data of at least one computing node in the previous n reasonings. Next, the first computing node obtains the non-repeated feature vector of the first computing node in this reasoning from the second computing node, and inputs the non-repeated feature vector of the first computing node in this reasoning and the repeated feature vector of the first computing node in this reasoning saved by the first computing node into the AI model for reasoning to obtain the reasoning result. The reasoning result includes a label corresponding to the feature.
[0023] This method identifies repeated feature vectors and non-repeated feature vectors of this inference through the input data of this inference and the input data of the previous n inferences, so as to differentiate between repeated feature vectors and non-repeated feature vectors. For example, non-repeated feature vectors are obtained through all-to-all, repeated feature vectors are obtained locally, and then repeated feature vectors and non-repeated feature vectors are integrated to complete the inference. While ensuring the accuracy of inference, the communication volume between multiple cards is reduced, thereby improving end-to-end performance.
[0024] In some possible implementations, the central processor may further construct a communication matrix based on the unique feature vectors of at least one computing node in the current reasoning and the storage locations of the unique feature vectors. The central processor then sends the communication matrix to the at least one computing node. The at least one computing node may include a first computing node. Accordingly, the first computing node may obtain the unique feature vectors of the first computing node in the current reasoning from a second computing node through many-to-many communication based on the communication matrix.
[0025] This method constructs a communication matrix to instruct computing nodes to obtain non-repeated feature vectors from other computing nodes, eliminating the need to obtain repeated feature vectors through many-to-many communication, reducing unmaskable communication operations, shortening latency, and improving end-to-end performance.
[0026] In a third aspect, the present application provides a distributed system. The distributed system includes a central processing unit and multiple computing nodes, each of the multiple computing nodes deploys an AI model, the multiple computing nodes respectively store at least one feature vector, the multiple computing nodes include a first computing node and a second computing node, the feature vector stored by the first computing node is different from the feature vector stored by the second computing node, the first computing node and the second computing node are used to train the AI model in parallel through multiple iterations, and one of the multiple iterations includes updating a parameter of the AI model once;
[0027] The central processing unit is configured to obtain input data of at least one computing node among the multiple computing nodes in this iteration, where the input data includes a batch of training data in the data set, and the training data includes features and labels corresponding to the features;
[0028] The central processing unit is configured to identify repeated feature vectors and non-repeated feature vectors of the at least one computing node in the current iteration based on input data of the at least one computing node in the current iteration and input data of the at least one computing node in the previous n iterations, wherein the at least one computing node stores the repeated feature vectors of the at least one computing node in the current iteration during the previous n iterations;
[0029] The first computing node is used to obtain the non-repeating feature vector of the first computing node in this iteration from the second computing node, and update the parameters of the AI model according to the non-repeating feature vector of the first computing node in this iteration, the repeated feature vector of the first computing node in this iteration saved by the first computing node, and the label.
[0030] In some possible implementations, the first computing node is specifically configured to:
[0031] Inputting the non-repeated feature vector of the first computing node in the current iteration and the repeated feature vector of the first computing node in the current iteration stored by the first computing node into the AI model for calculation, and obtaining the gradient of the feature vector input to the AI model at the first computing node based on the calculation result of the AI model and the label, wherein the feature vector input to the AI model includes the first feature vector that will be reused in the next n iterations;
[0032] Obtaining the gradient of the first eigenvector at the second computing node through reduction communication;
[0033] Update the parameters of the AI model according to the gradient of the first eigenvector at the first computing node and the gradient of the first eigenvector at the second computing node.
[0034] In some possible implementations, when the AI model is a recommendation model, the first computing node is specifically configured to:
[0035] The gradient of the first eigenvector at the second computing node and the gradient of the first eigenvector at the first computing node are accumulated, and the first eigenvector is reversely updated according to the accumulated first gradient.
[0036] In some possible implementations, the feature vector input to the AI model includes a second feature vector that will not be reused in subsequent n iterations, and the first computing node is further configured to:
[0037] Sending the gradient of the second eigenvector at the first computing node to a source computing node that pre-stores the second eigenvector through many-to-many communication;
[0038] The source computing node is used to update the second eigenvector according to the gradient of the second eigenvector at the first computing node.
[0039] In some possible implementations, the central processing unit is further configured to:
[0040] Identify the first feature vector based on multiple iterations of input data collected from the data set to obtain a recognition result;
[0041] Send the recognition result to the first computing node.
[0042] In some possible implementations, the central processing unit is further configured to:
[0043] Constructing a communication matrix according to the non-repeated eigenvectors of the at least one computing node in this iteration and the storage locations of the non-repeated eigenvectors;
[0044] sending the communication matrix to the at least one computing node, the at least one computing node including the first computing node;
[0045] The first computing node is specifically configured to:
[0046] According to the communication matrix, the non-repeated feature vector of the first computing node in this iteration is obtained from the second computing node through many-to-many communication.
[0047] In some possible implementations, the central processing unit is further configured to:
[0048] Constructing a query vector based on the repeated feature vector of the at least one computing node in this iteration and a storage location of the repeated feature vector, and sending the query vector to the at least one computing node, where the at least one computing node includes the first computing node;
[0049] The first computing node is further configured to:
[0050] The repeated feature vectors of this iteration stored in the first computing node are searched according to the query vector.
[0051] In some possible implementations, the computing node includes a neural network processor NPU, a graphics processor GPU, or a tensor processing unit TPU.
[0052] In some possible implementations, the distributed system includes one or more central processing units. When the distributed system includes a first central processing unit and a second central processing unit, the first central processing unit is connected to the first computing node, and the second central processing unit is connected to the second computing node.
[0053] In some possible implementations, the feature vector is a dense vector.
[0054] In a fourth aspect, the present application provides a distributed system. The distributed system includes a central processing unit and multiple computing nodes, each of the multiple computing nodes deploys an AI model, the multiple computing nodes respectively store at least one feature vector, the multiple computing nodes include a first computing node and a second computing node, the feature vector stored by the first computing node is different from the feature vector stored by the second computing node, and the first computing node and the second computing node are used for parallel reasoning using the AI model;
[0055] The central processing unit is configured to obtain input data of at least one computing node among the multiple computing nodes in this inference, wherein the input data includes a batch of inference data in the data set, and the inference data includes features;
[0056] The central processing unit is further configured to identify repeated feature vectors and non-repeated feature vectors of the at least one computing node in the current reasoning based on input data of the at least one computing node in the current reasoning and input data of the at least one computing node in the previous n reasonings;
[0057] The first computing node is used to obtain the non-repeating feature vector of the first computing node in this reasoning from the second computing node, and input the non-repeating feature vector of the first computing node in this reasoning and the repeated feature vector of the first computing node in this reasoning saved by the first computing node into the AI model for reasoning to obtain an inference result, where the inference result includes a label corresponding to the feature.
[0058] In some possible implementations, the central processing unit is further configured to:
[0059] Constructing a communication matrix according to the non-repeated feature vectors of the at least one computing node in the current reasoning and the storage locations of the non-repeated feature vectors;
[0060] sending the communication matrix to the at least one computing node, the at least one computing node including the first computing node;
[0061] The first computing node is specifically configured to:
[0062] According to the communication matrix, the non-repeated feature vector of the first computing node in this reasoning is obtained from the second computing node through many-to-many communication.
[0063] In a fifth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct a computing device or a computing device cluster to execute the method described in any implementation of the first aspect or the second aspect above.
[0064] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or the computing device cluster to execute the method described in any one of the implementations of the first or second aspect above.
[0065] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical methods of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments.
[0067] Figure 1 A schematic diagram of the structure of a recommended model provided for this application;
[0068] Figure 2 A schematic diagram of the architecture of a distributed system provided for this application;
[0069] Figure 3 A flowchart of a parallel training method for an AI model provided in this application;
[0070] Figure 4 A schematic diagram of multi-card communication provided in this application;
[0071] Figure 5 A flowchart of a parallel reasoning method for an AI model improved for this application;
[0072] Figure 6 A flowchart of a parallel training method for an AI model in a model training scenario provided in this application;
[0073] Figure 7 A schematic diagram of multi-card communication provided in this application;
[0074] Figure 8 A schematic diagram of using data repeatability for acceleration in a model training scenario provided in this application. DETAILED DESCRIPTION
[0075] The terms "first" and "second" in the embodiments of this application are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.
[0076] First, some technical terms involved in the embodiments of this application are introduced.
[0077] Artificial intelligence (AI), also known as machine intelligence, is the technology that correctly interprets external data, learns knowledge from that data, and flexibly applies that knowledge to achieve specific goals and tasks. Data-driven AI technologies include, but are not limited to, machine learning (ML) and deep learning. Machine learning is a type of algorithm that automatically analyzes data to identify patterns and uses these patterns to predict unknown data. Deep learning is an algorithm that uses artificial neural networks as a framework to learn how to represent data.
[0078] An AI model (also referred to as a model in some cases) refers to a mathematical model built based on AI technology, such as a model built using a machine learning algorithm or a deep learning algorithm. AI models can be applied to a variety of scenarios, including but not limited to search scenarios and recommendation scenarios. In the search scenario, a user provides a description of the content that the user intends to obtain, and the system (such as a search system) returns results that match the description to the user. In the recommendation scenario, the system (such as a recommendation system) actively recommends content to the user. Since there is no restriction on the information provided by the user, the recommendation scenario is more diverse. For example, the recommendation scenario can include personalized recommendations on the homepages of major applications (applications, APPs) and associated recommendations on the selection page. Associated recommendations include recommendations based on associated items and recommendations based on associated users. In e-commerce applications, recommendations based on associated items can be recommendations to users who purchase a product about what other users who purchase the product also purchase, or recommendations to users who browse the product about what other users who browse the product also browse. Recommendations based on associated users can be recommendations to users about what products their friends have purchased.
[0079] For ease of description, the following mainly uses the AI model as an example of the recommendation model in the recommendation system.
[0080] A recommendation model is an AI model used to implement content recommendations, such as video recommendations, text recommendations, or product recommendations. Common recommendation models may include collaborative filtering (Model-based Collaborative Filtering, collaborative filtering), dense neural network (DNN) and other models. Collaborative filtering predicts content that users may be interested in by analyzing the similarities between users or things and recommends this content to users. DNN is a basic structure widely used in neural networks. In this neural network, every neuron in each layer is connected to all neurons in the previous layer. The fully connected form enables DNN to have powerful representation capabilities when processing complex data, and the output layer is usually used as the output of tasks such as classification or regression.
[0081] Figure 1 A structural diagram of a recommendation model is shown, which can generally include an embedding layer and a dense neural network. Among them, the embedding layer is used to convert a high-dimensional sparse vector into a low-dimensional dense vector, which is also called embedded data, embedding data, or embedding, to "represent" an object. The object can refer to all things that can be recommended, such as goods, movies, music, news, etc. The embedding can express certain characteristics of the corresponding object, and the distance between vectors can also reflect the similarity between objects. High-dimensional sparse vectors may include but are not limited to identifier (ID) type features, such as user ID, item ID (itemID), document ID or video ID. Each ID corresponds to a low-dimensional dense vector of a preset size, which is the embedding data. Since the number of IDs is usually large, the above vectors can occupy more than 99% of the model volume. Furthermore, the recommendation system can also include non-ID type features. For example, in the text recommendation scenario, non-ID type features can include real vector features such as Latent Dirichlet Allocation (LDA). LDA is a topic model that can output the topic of each document in a document set in the form of a probability distribution. The above non-ID type features can be combined with the embedding data corresponding to the ID type features, such as aggregation (aggregate) or concatenation (concatenate, concat), and then input into the DNN. The prediction task of the output layer is used to predict whether the user clicks on the text and whether the user purchases the product. Correspondingly, the output layer can also output predicted click-through rate, purchase rate and other indicators.
[0082] During the training process, the recommendation system can concatenate the embedding data corresponding to the ID-type features in the training data and the non-ID-type features and input them into the DNN to predict whether the user clicks on the text or whether the user purchases the product, as well as the labels in the training data. The loss value is calculated through the loss function, and then the gradient is calculated based on the loss value. The gradient can be the gradient of the loss value with respect to the weight in the DNN, or the gradient of the loss value with respect to the embedding data in the embedding layer. The recommendation system can update parameters based on the gradient, such as updating the weights in the DNN and updating the embedding data in the embedding layer. Among them, the process from calculating the loss value to updating a parameter (such as the weight in the DNN or the embedding data in the embedding layer) can be regarded as an iteration. Specifically, an iteration can include the following processes: forward calculation of the loss value, reverse calculation of the gradient, and updating the parameters of the AI model based on the gradient.
[0083] In the recommendation model described above, the parameters in the embedding layer often occupy the vast majority of the recommendation model's volume, but the embedding layer's computational complexity is relatively low. However, the parameters in the DNN only occupy a small portion of the recommendation model's volume, but they account for the vast majority of the model's computational complexity. Therefore, a heterogeneous architecture can be used to train the recommendation model described above. A heterogeneous architecture refers to an architecture that uses different instruction sets for computation. For example, a heterogeneous architecture can use the instruction set of a central processing unit (CPU) and the instruction set of a neural network processing unit (NPU). The CPU has a large memory but low computing power, while the NPU has a small memory (also known as video memory) but high computing power.
[0084] The parameters of AI models, such as recommendation models, can be stored in memory for AI computing. AI computing refers to the calculations performed to build an AI model or use it for inference, and typically involves a large amount of highly parallel computing. Using deep learning models as an example, AI computing can include matrix operations. However, storing a large number of parameters in the memory of a single compute node is difficult. Using a single compute node to train an AI model can lead to issues such as insufficient memory.
[0085] To solve the problem of insufficient memory, when training AI models with large parameter scale (such as large-scale recommendation models), feature parallel strategy can be used for parallel training. Feature parallelism can be to split the feature vectors input to the AI model for training and distribute them to different computing nodes for model training. Figure 1 Taking the recommendation model as an example, the embedding data corresponding to the ID feature can be split and distributed to different NPUs. Different NPUs can obtain the embedding data allocated to them and input the embedding data into the DNN for model training. Among them, the computing node can be a computing card. The computing card can include but is not limited to an NPU, a graphics processing unit (GPU), or a tensor processing unit (TPU). For ease of description, the following example uses a computing card as an NPU.
[0086] To improve training efficiency, multi-level caching can also be introduced. Multi-level caching refers to setting up different levels of cache to cache data in the AI model. For example, the multi-level cache can include the NPU cache and the CPU cache. The CPU cache can include the full amount of embedding data, and each computing card can store part of the embedding data. Part of the embedding data used by each NPU during model training comes from other NPUs. Therefore, during the model training process, it is necessary to first perform gradient updates on the NPU that stores the original embedding data, and then distribute the feature vectors.
[0087] Specifically, the CPU side can send the information required for communication to the NPU side, and the NPU side performs multi-card communication to obtain the embedding data required by each NPU. Then, the NPU can pass the embedding data into the AI model for calculation to obtain the gradient. The NPU returns the gradient of the embedding data to the NPU where each embedding data originally resides (for the sake of convenience, it can also be called the original card) through many-to-many (all to all, a2a) communication. The original card accumulates the gradient of the corresponding embedding data and updates it to complete one iterative training.
[0088] However, the above method involves a large number of communication operations that cannot be overlaid. These operations cannot be performed in parallel with computation, requiring subsequent operations to wait for the completion of the communication. For example, the NPU needs to obtain embedding data through all-to-all communication and then pass it into the AI model for calculation. Furthermore, after obtaining the gradient of the embedding data, it also needs to obtain the gradient of the embedding data calculated by other NPUs through all-to-all communication before updating the embedding data. This results in a long communication delay, which severely impacts end-to-end performance.
[0089] In view of this, the present application provides a parallel training method for an AI model. The method can be applied to a distributed system. The distributed system includes a central processing unit (CPU) and multiple computing nodes. The computing node can be a computing card, wherein the computing card can be an NPU, a GPU, or a TPU. Each of the multiple computing nodes deploys an AI model. The multiple computing nodes respectively store at least one feature vector, which can be a low-dimensional dense vector, such as embedding data corresponding to an ID-type feature. The multiple computing nodes include a first computing node and a second computing node. The feature vector stored in the first computing node is different from the feature vector stored in the second computing node. The first computing node and the second computing node are used to train the AI model in parallel through multiple iterations, and one of the multiple iterations includes updating the parameters of the AI model once.
[0090] Specifically, the central processing unit obtains input data of at least one computing node in the current iteration of the plurality of computing nodes. The input data may include a batch of training data in the data set. The training data includes features and labels corresponding to the features. Figure 1 The example of the recommended scenario illustrates that the training data may include ID-type features, non-ID-type features, and labels. The central processing unit then identifies the repeated feature vectors and non-repeated feature vectors of at least one computing node in this iteration based on the input data of at least one computing node in this iteration and the input data of at least one computing node in the previous n iterations. Among them, at least one computing node saves the repeated feature vectors of at least one computing node in this iteration during the previous n iterations. The first computing node can obtain the non-repeated feature vectors of the first computing node in this iteration from the second computing node, and update the parameters of the AI model based on the non-repeated feature vectors of the first computing node in this iteration, the repeated feature vectors of the first computing node in this iteration saved by the first computing node, and the labels in the training data.
[0091] This method takes into account the repeatability of input data in different iterative processes during model training. For example, the input data of adjacent iterative processes in the recommendation scenario has strong repeatability. Based on the input data of this iteration and the input data of the previous n iterations, the repeated feature vectors and non-repeated feature vectors of this iteration are identified, and the repeated feature vectors and non-repeated feature vectors are differentiated. Non-repeated feature vectors can be obtained through communication between multiple computing nodes. Since repeated feature vectors have been saved in the current computing node during the previous n iterations, there is no need for unmasked communication operations. This can reduce communication between multiple computing nodes (such as communication between multiple cards). Without losing accuracy, the end-to-end performance can be improved by reducing the communication volume.
[0092] The parallel training method of the AI model of the present application can be applied to scenarios where the data used in different iterative processes during model training have strong repeatability. Among them, the data set including time series features (also referred to as time series features) is usually highly repeatable. During model training, the input data of different iterative processes sampled from the data set with strong repeatability also have strong repeatability. Therefore, the training method of the AI model of the present application can be applied to scenarios where models are trained in parallel based on time series features, including but not limited to search and recommendation scenarios. The method can be executed by a distributed system. The distributed system includes a central processing unit and multiple computing cards. The distributed system can adopt a single-machine multi-card architecture or a multi-machine multi-card architecture. The single-machine multi-card or multi-machine multi-card can be represented as x-machine x-card. It should be noted that the "machine" in the above x-machine x-card represents the host, and the "card" represents a device such as an accelerator card. Among them, the accelerator card can include but is not limited to a GPU, an NPU or a TPU.
[0093] In order to make the technical solution of the present application clearer and easier to understand, a distributed system with a single-machine multi-card architecture is used as an example for explanation below.
[0094] See also Figure 2 The hardware structure diagram of a distributed system shown in FIG. 1 may be a server 10 with a single-machine multi-card architecture. The server 10 includes a CPU on the host side and multiple NPUs on the device side. The CPU on the host side is also connected to a memory, which may be a dual in-line memory module (DIMM). The DIMM may be a double data rate (DDR) type, for example, the memory may be a DDR4 DIMM. Figure 2In the example, the host side can include 4 CPUs and 4 DDR4 DIMM groups. Each CPU is connected to 1 DDR4 DIMM group, and each DDR4 DIMM group includes 8 DDR4 DIMMs. Multiple CPUs on the host side can be connected to form a hydra mesh.
[0095] Optionally, the host side further includes an interface, such as one or more of a Serial Advanced Technology Attachment (SATA) interface, a next-generation non-volatile memory express (NVMe) interface, and a Gigabit Ethernet (GE) interface. The host side may also include a memory. The memory may include a SATA-supported memory or an NVMe-supported memory, such as a SATA-supported mechanical hard drive or an NVMe-supported solid-state drive (SSD).
[0096] The CPU on the host side and multiple NPUs on the device side can be connected through a bus. Figure 2 The example of the device side including 8 NPUs is used for illustration. In other possible implementations of the embodiment of the present application, the device side may also include more accelerator cards, such as a larger number of accelerator cards or more types of accelerator cards.
[0097] In other possible implementations of the embodiments of the present application, the distributed system may also be a multi-machine multi-card architecture, for example, a computing device cluster formed by multiple single-machine multi-card computing devices. Specifically, the distributed system may include multiple CPUs and multiple computing cards (such as NPUs). Among them, the multiple CPUs may include a first CPU and a second CPU. The first CPU and the second CPU may be CPUs of different hosts. Among them, the first CPU may be connected to a first computing card, and the second CPU may be connected to a second computing card. The first computing card may be one or more computing cards on the device side of the computing device where the first CPU is located, and the second computing card may be one or more computing cards on the device side of the computing device where the second CPU is located.
[0098] Based on the aforementioned distributed system, the present application also provides a parallel training method for an AI model. The method can be executed by a distributed system. The distributed system includes a CPU and multiple computing nodes, and each of the multiple computing nodes deploys an AI model. Multiple computing nodes respectively store at least one feature vector. The feature vector can be a dense vector obtained by encoding sparse features. Compared with sparse features, the dimension of dense vectors is greatly reduced, so it is also called a low-dimensional dense vector. For example, sparse features can be ID-type features in recommendation scenarios, including user ID and item ID, and the feature vector can include embedding data obtained by encoding the above-mentioned ID-type features. The multiple computing nodes include a first computing node and a second computing node, and the feature vector stored in the first computing node is different from the feature vector stored in the second computing node. The first computing node and the second computing node are used to train the AI model in parallel through multiple iterations, and one iteration of the multiple iterations includes updating the parameters of the AI model once. As Figure 3 As shown, the method includes:
[0099] S302: The CPU obtains input data of at least one computing node among the multiple computing nodes in this iteration.
[0100] Multiple computing nodes are used to train AI model data in parallel. The parallel method can be feature parallelism, which means splitting the feature vectors (such as embedding data) used to train the AI model and distributing them to different computing nodes. Based on this, the input data of multiple computing nodes in the same iteration can be different. In specific implementation, the CPU can sample data from the data set and then divide the sampled data, for example, by the hash value of the data, to obtain the input data of at least one computing node in this iteration.
[0101] In a training scenario, the input data includes a batch of training data from a dataset. This training data includes features and labels corresponding to the features. Still using the example of a recommendation scenario, features can include ID features, such as item IDs and user IDs, while labels indicate a user's decision regarding an item. For example, a label can indicate that a user purchased an item or not.
[0102] S304 . The CPU identifies repeated feature vectors and non-repeated feature vectors of at least one node in this iteration based on input data of at least one computing node in this iteration and input data of at least one computing node in previous n iterations.
[0103] For a compute node, the repeated feature vector for this iteration can be a feature vector used by the compute node in both this iteration and the previous n iterations. The non-repeated feature vector for this iteration can be a repeated feature vector used by the compute node in this iteration but not used in the previous n iterations. n can be a positive integer. For example, n can be 1, meaning that the non-repeated feature vector can be a feature vector that was not used in the previous iteration and is used in the current iteration. The repeated feature vector can also be determined based on the input data and repeated feature vectors for this iteration, for example, the remaining feature vectors in the feature vector corresponding to the input data, excluding the repeated feature vector.
[0104] Specifically, in the initialization phase, at least one feature vector stored by the computing node may be a feature vector assigned to each computing node during initialization, and after at least one iteration, at least one feature vector stored by the computing node may be a repeated feature vector retained after the iteration. The computing node can directly obtain the repeated feature vector locally without having to obtain the repeated feature vector through an unconcealable communication operation such as all-to-all communication. Based on this, the CPU can compare the input data of the computing node in this iteration with the input data of the computing node in the previous n iterations, for example, comparing the keys of the input data of different iterations, thereby identifying repeated feature vectors and non-repeated feature vectors, so as to differentiate between the repeated feature vectors and the non-repeated feature vectors.
[0105] S306: The first computing node obtains the non-repeated feature vector of the first computing node in this iteration from the second computing node.
[0106] S308. The first computing node updates the parameters of the AI model according to the non-repeating feature vector of the first computing node in this iteration, the repeated feature vector of the first computing node in this iteration saved by the first computing node, and the label.
[0107] Specifically, the first computing node can receive the communication matrix sent by the CPU, and then obtain the non-repeating feature vectors of the first computing node in this iteration from the second computing node through many-to-many communication based on the communication matrix. The communication matrix records, in matrix form, the non-repeating feature vectors of the computing node in this iteration that each computing node needs to obtain, as well as the storage location of the non-repeating feature vectors. Accordingly, each computing node can determine the computing node storing the above feature vectors through the communication matrix, and then obtain the above non-repeating feature vectors from the computing node storing the above non-repeating feature vectors through many-to-many communication.
[0108] The communication matrix can be constructed by the CPU based on the unique eigenvectors and their storage locations and sent to each computing node. This is described in detail below.
[0109] Specifically, the CPU parses the non-repeating feature vectors and determines the computing nodes and storage locations where the non-repeating feature vectors are stored. The CPU can maintain the metadata of the feature vectors stored on the device side. For example, the CPU can use a hash table to store the key of the feature vector stored by the computing node on the device side, the computing node that stores the feature vector, and the storage location of the feature vector in the computing node. The key of the feature vector can be a feature identifier (ID), such as an item identifier in a recommendation system. Accordingly, the CPU can perform a hash query through the hash table based on the key of the non-repeating feature vector to determine the computing node and storage location where the non-repeating feature vector is stored.
[0110] The CPU then constructs a communication matrix based on the computing nodes and storage locations that store non-repeating feature vectors. Specifically, the communication matrix may include multiple column vectors, each corresponding to a computing node. For example, the first column vector corresponds to NPU 0, the second column vector corresponds to NPU 1, and so on. The N+1th column vector corresponds to NPU N. The first column vector may store non-repeating feature vectors obtained from NPU 1 to NPU N and their storage locations. For example, the first element of the first column vector is 0, and the second to N+1 elements are the non-repeating feature vectors obtained from NPU 1 to NPU N and their storage locations, respectively.
[0111] Furthermore, the CPU can also determine the repeated feature vectors of at least one computing node in this iteration. The CPU can construct a query vector based on the repeated feature vectors of at least one computing node in this iteration and the storage location of the repeated feature vectors. The query vector can be used to query the repeated feature vectors so that the repeated feature vectors can be used to perform model training in this iteration. Among them, the computing node stores the repeated feature vectors of this iteration, and there is no need to obtain the above-mentioned repeated feature vectors from other computing nodes through all-to-all communication, which reduces the communication volume between multiple cards, especially the communication volume that cannot be covered, and improves end-to-end performance. Furthermore, the query vector can also be used to query non-repeating feature vectors. Among them, the CPU can construct a query vector based on the computing node and storage location storing the repeated feature vectors in this iteration and the computing node and storage location storing the non-repeating feature vectors in this iteration. Accordingly, the CPU can send a communication matrix to each computing node on the device side through host-to-device communication. Furthermore, when the CPU constructs the query vector, the CPU can also send the query vector to each computing node.
[0112] The first computing node may input the non-repeating feature vectors of the first computing node in this iteration and the repeated feature vectors of the first computing node in this iteration saved by the first computing node into the AI model. The first computing node may fuse the non-repeating feature vectors and the repeated feature vectors, and the fusion method may include but is not limited to splicing and concatenation (abbreviated as concat). For example, the first computing node may use the non-repeating feature vectors and the repeated feature vectors as inputs to the concat function, and then execute the concat function to obtain a feature matrix, and the first computing node may input the feature matrix into the AI model. Then, the first computing node may obtain the gradient (abbreviated as grad or grads) of the feature vector input to the AI model at the first computing node based on the calculation results and labels of the AI model. The gradient can be represented by a vector, the direction of the vector is consistent with the direction of the maximum directional derivative, and the modulus of the vector is the maximum value of the directional derivative. The first computing node may update the weights of the AI model based on the gradient. For example, the first computing card may update the weights of the AI model based on the gradient using the back propagation (BP) algorithm. Furthermore, in the recommendation scenario, in order to better represent the feature vector, the first computing node also updates the feature vector input to the AI model based on the gradient.
[0113] The feature vectors input to the AI model include a first feature vector that will be reused in the next n iterations. Furthermore, the feature vectors input to the AI model may also include a second feature vector that will not be reused in the next n iterations. The first computing node may perform differentiated processing on the feature vector that will be reused in the next n iterations and the second feature vector that will not be reused in the next n iterations.
[0114] For the first eigenvector that will be reused in the next n iterations, the first computing node can obtain the gradient of the first eigenvector at the second computing node through reduction communication (all reduce). Then, the first computing node updates the parameters of the AI model based on the gradient of the first eigenvector at the first computing node and the gradient of the first eigenvector at the second computing node. For example, when the AI model does not recommend a model, the first computing node can accumulate the gradient of the first eigenvector at the second computing node and the gradient of the first eigenvector at the first computing node, and reversely update the first eigenvector based on the accumulated first gradient.
[0115] For the second eigenvector that will not be reused in the next n iterations, the first computing node can send the gradient of the second eigenvector at the first computing node to the source computing node (such as the original card) that pre-stores the second eigenvector via many-to-many communication. Accordingly, the source computing node updates the second eigenvector based on the gradient of the second eigenvector at the first computing node. Similar to the first computing node, the source computing node can accumulate the gradient of the second eigenvector at the first computing node and the gradient of the second eigenvector at the source node, and then reversely update the second eigenvector based on the accumulated second gradient.
[0116] For ease of understanding, the following is an example. Figure 4 As shown, the CPU can send the key of the input data of this iteration to the NPU, for example, sending the key of a batch of samples to NPU 0 to NPUN. Among them, the key of a batch of samples may include the key of the sample processed by NPU 0, the key of the sample processed by NPU 1...the key of the sample processed by NPU N, which are respectively recorded as keys_0, keys_1...keys_N. Among them, NPU 0 to NPU N can obtain the keys of all samples processed by each NPU through all-to-all communication, and determine the repeated feature vectors stored locally and the feature vectors that need to be sent to other NPUs through the hash table of the NPU. NPU 0 to NPU N can obtain the non-repeated feature vectors required by each iteration in this iteration through all-to-all communication. Taking NPU0 as an example, NPU 0 can perform all-to-all communication through the communication matrix, so as to obtain the feature vector processed by NPU 0 stored in NPU N from NPU N, which is recorded as Embs_0. Similarly, NPU N can perform all-to-all communication through the communication matrix, thereby obtaining the feature vector stored in NPU 0 and processed by NPU N from NPU 0, which is recorded as Embs_N.
[0117] It should be noted that the CPU can collect input data for multiple iterations from a dataset. For example, while sending the key for the input data for the current iteration to the compute node, the CPU can collect input data for the next iteration. Based on the input data for the multiple iterations collected from the dataset, the CPU can identify a first eigenvector that will be reused in the next iteration and a second eigenvector that will not be reused in the next iteration, and obtain a recognition result. This recognition result indicates the eigenvectors that will be reused in the next iteration and the eigenvectors that will not be reused in the next iteration. Accordingly, upon receiving the recognition result, the compute node can differentiate between the eigenvectors that will be reused in the next iteration and the eigenvectors that will not be reused in the next iteration. For example, for the eigenvectors that will be reused in the next iteration, the gradients on multiple cards can be collected through allreduce communication, and the eigenvectors can be updated. The updated eigenvectors can be retained on the current compute card, awaiting reuse or call in the next iteration. For the eigenvectors that will not be reused in the next iteration, the compute node can transmit the gradients of these eigenvectors back to the compute node that originally stored them through all-to-all communication on the communication matrix, and the eigenvectors can be updated through the optimizer.
[0118] Based on the above description, the parallel training method of the AI model of the present application identifies the repeated feature vectors and non-repeated feature vectors of this iteration through the input data of this iteration and the input data of the previous n iterations, so as to perform differentiated processing on the repeated feature vectors and the non-repeated feature vectors. For example, a communication matrix is constructed for the non-repeated feature vectors to conduct communication between multiple cards, and the repeated feature vectors are processed locally. Under the premise of ensuring the accuracy of the AI model, the communication volume between multiple cards is reduced, thereby improving end-to-end performance.
[0119] Figure 3 The embodiment shown details the parallel training method of the AI model. Based on the trained AI model, the AI model can also be used for parallel reasoning. The parallel reasoning method of the AI model is described in detail below.
[0120] In the model inference scenario, the computing node does not need to calculate the gradient, nor does it need to update the feature vector based on the gradient. The computing node can input the non-repeated feature vectors received by all to all and the retained repeated feature vectors into the AI model for inference. Furthermore, the computing node can retain the feature vectors that will be reused in the next inference. For example, the CPU can identify the feature vectors that will be reused in the next inference based on the input data of the current inference and the input data of the next inference after receiving the input data of the next inference, and then send the recognition results to the computing node, so that the computing node can retain the feature vectors received by all to all (such as embedding data) and reuse them in the next inference.
[0121] See also Figure 5 A flowchart of a parallel reasoning method for an AI model is shown, which can be applied to a distributed system. The distributed system includes a CPU and multiple computing nodes, each of which deploys an AI model. The multiple computing nodes each store at least one feature vector. The multiple computing nodes include a first computing node and a second computing node. The feature vector stored in the first computing node is different from the feature vector stored in the second computing node. The first computing node and the second computing node are used for parallel reasoning based on the AI model. The method includes the following steps:
[0122] S502: The CPU obtains input data of at least one computing node among the multiple computing nodes in this inference.
[0123] The input data includes a batch of inference data from a dataset, which includes features. These features can include ID-based features, such as user IDs and item IDs. Furthermore, these features can also include non-ID-based features. For example, in a text recommendation scenario, non-ID-based features can include real-valued vector features such as LDA. LDA is a topic model that outputs the topic of each document in a document collection as a probability distribution.
[0124] S504 , the CPU identifies repeated feature vectors and non-repeated feature vectors of at least one computing node in the current inference based on input data of at least one computing node in the current inference and input data of at least one computing node in the previous n inferences.
[0125] Specifically, the CPU may compare the input data of at least one computing node in the current inference with the input data of at least one computing node in the previous n inferences, thereby obtaining repeated feature vectors and non-repeated feature vectors of at least one computing node in the current inference. The feature vectors may be low-dimensional dense vectors, such as embedding data obtained by encoding high-dimensional sparse features.
[0126] It should be noted that the specific implementation of the CPU determining the repeated feature vectors and non-repeated feature vectors for this inference can be referred to the description of the model training related content, which will not be repeated here.
[0127] S506: The first computing node obtains a non-repeated feature vector of the first computing node in this reasoning from the second computing node.
[0128] Specifically, the first computing node can receive the communication matrix sent by the CPU, and obtain the non-repeated feature vectors of the first computing node in this reasoning from the second computing node through all-to-all communication according to the storage location of the non-repeated feature vectors indicated by the communication matrix. The specific implementation of the communication matrix and obtaining the non-repeated feature vectors according to the communication matrix can be found in Figure 3 The description of the relevant contents of the embodiment will not be repeated here.
[0129] S508. The first computing node inputs the non-repeated feature vector of the first computing node in this reasoning and the repeated feature vector of the first computing node in this reasoning saved by the first computing node into the AI model for reasoning to obtain the reasoning result.
[0130] Specifically, the first computing node can fuse the non-repeated feature vectors of the first computing node in the current inference with the repeated feature vectors of the first computing node in the current inference saved by the first computing node, for example, by concatenation or splicing, to obtain a feature matrix. The first computing node can then input the feature matrix into the AI model to perform matrix operations and obtain an inference result.
[0131] The inference result includes a label corresponding to the feature. For example, in a product recommendation scenario, the label can be a buy or no buy label. Furthermore, the inference result can also include the probability of the user purchasing the product. When the probability of the user purchasing the product is greater than or equal to a threshold, the label can be buy. When the probability of the user purchasing the product is less than the threshold, the label can be no buy. For another example, in a text recommendation scenario, the label can be browse or no browse. Furthermore, the inference result can also include the probability of the user browsing the article. When the probability of the user browsing the article is greater than or equal to a threshold, the label can be browse. When the probability of the user browsing the article is less than the threshold, the label can be no browse.
[0132] Based on the above description, the parallel reasoning method of the AI model provided in this application identifies the repeated feature vectors and non-repeated feature vectors of this reasoning through the input data of this reasoning and the input data of the previous n reasonings, so as to differentiate the repeated feature vectors and non-repeated feature vectors. For example, non-repeated feature vectors are obtained through all to all, repeated feature vectors are obtained locally, and then the repeated feature vectors and non-repeated feature vectors are integrated to complete the reasoning. While ensuring the accuracy of reasoning, the communication volume between multiple cards is reduced, and the end-to-end performance is improved.
[0133] Next, the parallel training method of the AI model of this application is described in detail in combination with specific application scenarios.
[0134] See also Figure 6The flowchart of a parallel training method for an AI model in a model training scenario is shown. The method is executed by a distributed system that includes a CPU and multiple NPUs, where each NPU stores a different feature vector, for example, different embedding data. The method may include the following steps:
[0135] Step 1: The CPU obtains the input data of each NPU in this iteration.
[0136] Specifically, the CPU can start a process for each NPU to obtain the input data of the corresponding NPU in this iteration. The input data of the NPU in this iteration can be the samples data of this batch (or current batch), which is recorded as batch samples.
[0137] Step 2: The CPU determines whether the key of the input data of this iteration matches the key of the input data of the previous iteration, thereby identifying unique feature vectors. If the key of the input data of this iteration matches the key of the input data of the previous iteration, Step 3 is executed.
[0138] The CPU compares the key of the input data of this iteration with the key of the input data of the previous iteration to determine whether the key of the input data of this iteration matches the key of the input data of the previous iteration. If the key of the input data of this iteration matches the key of the input data of the previous iteration, it indicates that the input data (feature vector) is a repeated feature vector. If the key of the input data of this iteration does not match the key of the input data of the previous iteration, it indicates that the input data is a non-repeating feature vector. For non-repeating feature vectors, the CPU can execute Step 3.
[0139] Step 3: The CPU parses the non-repeating feature vectors and obtains the NPU and storage location where the non-repeating feature vectors are stored.
[0140] The CPU can obtain the NPU storing the non-repeated feature vector and the storage location of the feature vector in the NPU by looking up the table according to the key of the non-repeated feature vector. To improve query efficiency, the CPU can use a hash search method to obtain the NPU storing the non-repeated feature vector and the storage location.
[0141] Step 4: The CPU constructs a communication matrix based on the NPUs that store non-repeated feature vectors and their storage locations.
[0142] The specific implementation of CPU constructing communication matrix can be referred to Figure 3 The description of the relevant contents of the embodiment will not be repeated here.
[0143] Furthermore, the CPU may construct a query vector according to the NPU storing the repeated feature vectors in this iteration and the storage location, and send the query vector to the NPU.
[0144] Step 5: The NPU queries for unique feature vectors.
[0145] The NPU can query the corresponding feature vector based on the key of the unique feature vector, so as to facilitate the transmission of the unique feature vector between multiple cards through all-to-all communication. Furthermore, the NPU can also query the duplicate feature vectors of this iteration based on the query vector.
[0146] Step 6: NPUs perform all-to-all communication based on the communication matrix to transmit non-repeated feature vectors between NPUs.
[0147] like Figure 4 As shown in the figure, after querying non-repeating feature vectors, NPU 0 to NPU N can transfer non-repeating feature vectors between multiple cards through all-to-all communication. For example, NPU 0 can receive the non-repeating feature vector Embs_0 transmitted by NPU 1 to NPU N and need to be processed by NPU 0, and transmit the non-repeating feature vectors Embs_1, Embs_2, ..., Embs_N required by NPU 1 to NPU N to NPU 1 to NPU N.
[0148] Step 7: NPU concatenates the non-repeated feature vectors and the retained repeated feature vectors.
[0149] Step 8: The NPU inputs the concatenated feature matrix into the AI model for model training.
[0150] like Figure 4 As shown, NPU 0 to NPU N can concatenate their respective retained repeated feature vectors with the non-repeated feature vectors received from other NPUs, and then input the concatenated feature matrix into the AI model, and the AI model can perform matrix operations based on the feature matrix.
[0151] Step 9: The NPU obtains the gradient based on the input and output of the AI model.
[0152] Specifically, the NPU can determine the derivatives of the AI model's function in each direction based on the AI model's input and output, and thus determine the gradient. The direction of the gradient is consistent with the direction of the maximum directional derivative, and the modulus of the gradient is the maximum value of the directional derivative.
[0153] Step 10: The CPU determines whether the key of the feature vector input to the AI model matches the key of the input data for the next iteration. If so, it executes Step 11; if not, it executes Step 12.
[0154] The way in which the CPU determines whether the key of the feature vector input to the AI model is matched with the key of the input data of the next iteration is similar to Step 2 and will not be repeated here.
[0155] Step 11: NPU uses all-reduce communication to collect the gradient of the feature vector in multiple NPUs.
[0156] For the feature vector that will be reused in the next iteration (reusable feature vector), the NPU can collect the gradient of the feature vector through allreduce. Figure 7 As shown, NPU 0 to NPU N can generate the gradient of the feature vectors processed by different NPUs, such as NPU 0 can generate the gradient of the feature vectors processed by NPU 0, NPU 1...NPU N, which are recorded as Grads_00, Grads_10,...Grads_N0. For repeated feature vectors (reusable feature vectors), NPU 0 can collect the gradient of the feature vector in other NPUs through allreduce to obtain the gradient of the feature vector. For example, the current NPU can accumulate the gradient of the repeated feature vector collected from other NPUs with the gradient of the repeated feature vector in the current NPU to obtain the gradient of the feature vector. NPU 0 to NPU N can transmit the gradient of the repeated feature vector through allreduce communication, and accumulate the gradient of the feature vector in each NPU to obtain the gradient of the feature vector, specifically Grads_0, Grads_1...Grads_N.
[0157] Step 12: The NPU uses all-to-all communication to return the gradient of the feature vector to the NPU that originally stored the feature vector.
[0158] For feature vectors that will not be reused in the next iteration, the NPU can return the gradient of the feature vector to the NPU that originally stored the feature vector through all to all. Figure 4As shown, NPU 0 to NPU N can each output the gradients of feature vectors processed by different NPUs, such as Grads_0, Grads_1, ..., Grads_N. For feature vectors that will not be reused, NPU 0 to NPU N can transmit the gradients of the feature vectors through all-to-all communication. For example, NPU 0 can receive Grads_1 ..., Grads_N sent by NPU 1 to NPU N, and send Grads_1 ..., Grads_N to NPU 1 to NPU N respectively.
[0159] Step 13: NPU updates the feature vector based on the gradient in reverse.
[0160] like Figure 4 or Figure 7 As shown, the NPU can input the accumulated gradients into the optimizer, which updates the feature vectors. After the feature vectors that will be reused in the next iteration are updated, they can be input into the AI model together with new feature vectors (New Embs), such as the non-repeating feature vectors for the next iteration, for the next iteration. The non-repeating feature vectors for the next iteration can be obtained through a forward all-to-all process.
[0161] The parallel training method of the AI model of this application can reuse the embedding data of repeated keys between adjacent batches through all-reduce communication and update them locally, avoiding the two all-to-all communication operations required for each batch to be updated remotely. Only when these keys are not reused in a batch in the future, they are transmitted back to the remote end for update. Figure 8As shown, the input data of the previous iteration is batch 0 (denoted as B0). Batch 0 can include the feature vectors corresponding to keys1 and keys2. The NPUs transmit non-repeating feature vectors between multiple cards through forward all-to-all communication, and then input them into the AI model for calculation, such as performing matrix operations. The gradient is obtained based on the operation results, and then the gradient is transmitted between multiple cards through reverse all-to-all communication. Since batch 0 and the input data batch1 of this iteration have repeated keys, that is, repeated feature vectors, the NPU can update the gradient of the repeated feature vector locally. Based on this, the reverse all-to-all communication and forward all-to-all communication for repeated feature vectors can be replaced by all-reduce communication. In the next iteration, some feature vectors from this iteration can be reused, and the non-repeating feature vectors for the next iteration can be obtained through forward all-to-all communication. The reused repeated feature vectors and the above non-repeating feature vectors are then concatenated and input into the AI model for calculation. Then, the gradients of the non-repeating feature vectors, such as the gradient of the feature vector corresponding to keys4, are transmitted between multiple cards through reverse all-to-all communication. The gradients of the reused repeated feature vectors, such as the gradient of the feature vector corresponding to keys5, are collected through all-reduce. This can shorten communication time and accelerate model training.
[0162] This method identifies the repeated feature vectors and non-repeated feature vectors of this iteration based on the input data of this iteration and the next iteration. The repeated feature vectors can be updated locally and wait for reuse in the next iteration. Therefore, the two all-to-all operators of the repeated feature vectors can be equivalently replaced by an all-reduce communication operator, realizing hot update of gradients based on repeated keys in adjacent batches. While ensuring the training accuracy of the AI model, the communication volume between multiple cards is reduced, thereby improving end-to-end performance.
[0163] It should be noted that the model data processing method of this application can also be applied to model reasoning scenarios, such as inference scenarios for large-scale recommendation services. Figure 6 In the inference scenario, there is no need to calculate gradients or update feature vectors based on gradients. The NPU can simply retain and reuse the feature vectors received through all-to-all.
[0164] Based on the aforementioned parallel training method and parallel reasoning method of AI models, this application provides a distributed system. The distributed system includes a CPU and multiple computing nodes, and the computing nodes can be computing cards, including but not limited to NPUs, GPUs, or TPUs. Figure 2 The following example uses a distributed system including a CPU and multiple NPUs.
[0165] Each of the plurality of computing nodes deploys an AI model, and each of the plurality of computing nodes stores at least one feature vector. The plurality of computing nodes includes a first computing node and a second computing node, wherein the feature vector stored in the first computing node is different from the feature vector stored in the second computing node. The first computing node and the second computing node are configured to train the AI model in parallel through multiple iterations, wherein each of the multiple iterations includes updating a parameter of the AI model.
[0166] Specifically, the CPU is used to obtain input data of at least one computing node among the multiple computing nodes in the current iteration, where the input data includes a batch of training data in a data set, and the training data includes features and labels corresponding to the features. The CPU is also used to identify repeated feature vectors and non-repeated feature vectors of at least one computing node in the current iteration based on the input data of the at least one computing node in the current iteration and the input data of the at least one computing node in the previous n iterations, where the at least one computing node stores the repeated feature vectors of the at least one computing node in the current iteration during the previous n iterations.
[0167] The first computing node is used to obtain the non-repeating feature vector of the first computing node in this iteration from the second computing node, and update the parameters of the AI model according to the non-repeating feature vector of the first computing node in this iteration, the repeated feature vector of the first computing node in this iteration saved by the first computing node, and the label.
[0168] In some possible implementations, the first computing node is specifically configured to:
[0169] Input the non-repeated feature vector of the first computing node in this iteration and the repeated feature vector of the first computing node in this iteration saved by the first computing node into the AI model for calculation. Obtain the gradient of the feature vector input to the AI model at the first computing node based on the calculation result and the label of the AI model. The feature vector input to the AI model includes the first feature vector that will be reused in the next n iterations.
[0170] Obtaining the gradient of the first eigenvector at the second computing node through reduction communication;
[0171] Update the parameters of the AI model according to the gradient of the first eigenvector at the first computing node and the gradient of the first eigenvector at the second computing node.
[0172] In some possible implementations, when the AI model is a recommendation model, the first computing node is specifically configured to:
[0173] The gradient of the first eigenvector at the second computing node and the gradient of the first eigenvector at the first computing node are accumulated, and the first eigenvector is reversely updated according to the accumulated first gradient.
[0174] In some possible implementations, the feature vector input to the AI model includes a second feature vector that will not be reused in subsequent n iterations, and the first computing node is further configured to:
[0175] Sending the gradient of the second eigenvector at the first computing node to a source computing node that pre-stores the second eigenvector through many-to-many communication;
[0176] The source computing node is used to update the second eigenvector according to the gradient of the second eigenvector at the first computing node.
[0177] In some possible implementations, the CPU is further configured to:
[0178] Identify the first feature vector based on multiple iterations of input data collected from the data set to obtain a recognition result;
[0179] Send the recognition result to the first computing node.
[0180] In some possible implementations, the CPU is further configured to:
[0181] Constructing a communication matrix according to the non-repeated eigenvectors of the at least one computing node in this iteration and the storage locations of the non-repeated eigenvectors;
[0182] sending the communication matrix to the at least one computing node, the at least one computing node including the first computing node;
[0183] The first computing node is specifically configured to:
[0184] According to the communication matrix, the non-repeated feature vector of the first computing node in this iteration is obtained from the second computing node through many-to-many communication.
[0185] In some possible implementations, the CPU is further configured to:
[0186] Constructing a query vector based on the repeated feature vector of the at least one computing node in this iteration and a storage location of the repeated feature vector, and sending the query vector to the at least one computing node, where the at least one computing node includes the first computing node;
[0187] The first computing node is also used to:
[0188] The repeated feature vectors of this iteration stored in the first computing node are searched according to the query vector.
[0189] In some possible implementations, the computing node includes a neural network processor NPU, a graphics processor GPU, or a tensor processing unit TPU.
[0190] In some possible implementations, the distributed system includes one or more central processing units. When the distributed system includes a first central processing unit and a second central processing unit, the first central processing unit is connected to the first computing node, and the second central processing unit is connected to the second computing node.
[0191] In some possible implementations, the feature vector is a dense vector.
[0192] Another distributed system provided herein can be used for parallel reasoning of AI models. The distributed system includes a CPU and multiple computing nodes, each of which deploys an AI model, and each of the multiple computing nodes stores at least one feature vector. The multiple computing nodes include a first computing node and a second computing node. The feature vector stored by the first computing node is different from the feature vector stored by the second computing node, and the first computing node and the second computing node are used for parallel reasoning using the AI model.
[0193] The CPU is used to obtain input data of at least one computing node among the multiple computing nodes in this reasoning, where the input data includes a batch of reasoning data in a data set, and the reasoning data includes features; the CPU is also used to identify repeated feature vectors and non-repeated feature vectors of the at least one computing node in this reasoning based on the input data of the at least one computing node in this reasoning and the input data of the at least one computing node in the previous n reasonings.
[0194] The first computing node is used to obtain the non-repeating feature vector of the first computing node in this reasoning from the second computing node, and input the non-repeating feature vector of the first computing node in this reasoning and the repeated feature vector of the first computing node in this reasoning saved by the first computing node into the AI model for reasoning to obtain an inference result, where the inference result includes a label corresponding to the feature.
[0195] In some possible implementations, the CPU is further configured to:
[0196] Constructing a communication matrix according to the non-repeated feature vectors of the at least one computing node in the current reasoning and the storage locations of the non-repeated feature vectors;
[0197] sending the communication matrix to the at least one computing node, the at least one computing node including the first computing node;
[0198] The first computing node is specifically used for:
[0199] According to the communication matrix, the non-repeated feature vector of the first computing node in this reasoning is obtained from the second computing node through many-to-many communication.
[0200] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device (for example, a distributed system) to execute the parallel training method of the above-mentioned AI model or the parallel reasoning method of the AI model.
[0201] The embodiment of the present application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the above-mentioned parallel training method of the AI model, or the parallel reasoning method of the AI model. The embodiment of the present application also provides a computer program product containing instructions. When the computer program product is run on at least one computing device, the at least one computing device executes the above-mentioned parallel training method of the AI model or the parallel reasoning method of the AI model.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A parallel training method for an AI model, characterized in that: Applied to a distributed system, the distributed system includes a central processing unit and multiple computing nodes, each of the multiple computing nodes deploys an AI model, the multiple computing nodes respectively store at least one feature vector, the multiple computing nodes include a first computing node and a second computing node, the feature vector stored by the first computing node is different from the feature vector stored by the second computing node, the first computing node and the second computing node are used to train the AI model in parallel through multiple iterations, one of the multiple iterations includes updating a parameter of the AI model once, and the method includes: The central processing unit obtains input data of at least one computing node among the multiple computing nodes in this iteration, where the input data includes a batch of training data in the data set, and the training data includes features and labels corresponding to the features; The central processing unit identifies, based on input data of the at least one computing node in the current iteration and input data of the at least one computing node in the previous n iterations, repeated feature vectors and non-repeated feature vectors of the at least one computing node in the current iteration, wherein the at least one computing node stores the repeated feature vectors of the at least one computing node in the current iteration during the previous n iterations; The first computing node obtains the non-repeating feature vector of the first computing node in this iteration from the second computing node, and updates the parameters of the AI model according to the non-repeating feature vector of the first computing node in this iteration, the repeated feature vector of the first computing node in this iteration saved by the first computing node, and the label.
2. The method according to claim 1, characterized in that The first computing node updates the parameters of the AI model according to the non-repeated feature vector of the first computing node in the current iteration, the repeated feature vector of the first computing node in the current iteration saved by the first computing node, and the label, including: The first computing node inputs a non-repeated feature vector of the first computing node in the current iteration and a repeated feature vector of the first computing node in the current iteration stored by the first computing node into the AI model for calculation, and obtains a gradient of the feature vector input to the AI model at the first computing node based on the calculation result of the AI model and the label, where the feature vector input to the AI model includes a first feature vector that will be reused in the next n iterations; The first computing node obtains the gradient of the first eigenvector at the second computing node through reduction communication; The first computing node updates the parameters of the AI model according to the gradient of the first eigenvector at the first computing node and the gradient of the first eigenvector at the second computing node.
3. The method according to claim 2, characterized in that When the AI model is a recommendation model, the first computing node updates the parameters of the AI model according to the gradient of the first eigenvector at the first computing node and the gradient of the first eigenvector at the second computing node, including: The first computing node accumulates the gradient of the first eigenvector at the second computing node and the gradient of the first eigenvector at the first computing node, and reversely updates the first eigenvector according to the accumulated first gradient.
4. The method according to claim 2 or 3, characterized in that The feature vector input to the AI model includes a second feature vector that will not be reused in the next n iterations, and the method further includes: The first computing node sends the gradient of the second eigenvector at the first computing node to a source computing node that pre-stores the second eigenvector through many-to-many communication, and the source computing node updates the second eigenvector according to the gradient of the second eigenvector at the first computing node.
5. The method according to any one of claims 2 to 4, characterized in that The method further comprises: The central processing unit identifies the first feature vector based on multiple iterations of input data collected from the data set to obtain a recognition result; The central processing unit sends the recognition result to the first computing node.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The central processing unit constructs a communication matrix according to the non-repeated feature vectors of the at least one computing node in this iteration and the storage location of the non-repeated feature vectors; The central processing unit sends the communication matrix to the at least one computing node, the at least one computing node including the first computing node; The first computing node obtains, from the second computing node, a non-repeated feature vector of the first computing node in this iteration, including: The first computing node obtains the non-repeated feature vector of the first computing node in this iteration from the second computing node through many-to-many communication according to the communication matrix.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: The central processor constructs a query vector based on the repeated feature vector of the at least one computing node in this iteration and the storage location of the repeated feature vector, and sends the query vector to the at least one computing node, wherein the at least one computing node includes the first computing node; The first computing node searches for repeated feature vectors of this iteration stored by the first computing node according to the query vector.
8. The method according to any one of claims 1 to 7, characterized in that The computing node includes a neural network processor NPU, a graphics processor GPU or a tensor processing unit TPU.
9. The method according to any one of claims 1 to 8, characterized in that The distributed system includes one or more central processing units. When the distributed system includes a first central processing unit and a second central processing unit, the first central processing unit is connected to the first computing node, and the second central processing unit is connected to the second computing node.
10. The method according to any one of claims 1 to 9, characterized in that The feature vector is a dense vector.
11. A parallel reasoning method for an AI model, characterized in that: Applied to a distributed system, the distributed system includes a central processing unit and multiple computing nodes, each of the multiple computing nodes deploys an AI model, the multiple computing nodes respectively store at least one feature vector, the multiple computing nodes include a first computing node and a second computing node, the feature vector stored in the first computing node is different from the feature vector stored in the second computing node, the first computing node and the second computing node are used for parallel reasoning based on the input feature vector, the method includes: The central processing unit obtains input data of at least one computing node among the multiple computing nodes in current inference, where the input data includes a batch of inference data in the data set, and the inference data includes features; The central processing unit identifies repeated feature vectors and non-repeated feature vectors of the at least one computing node in the current reasoning based on input data of the at least one computing node in the current reasoning and input data of the at least one computing node in the previous n reasonings; The first computing node obtains the non-repeating feature vector of the first computing node in this reasoning from the second computing node, and inputs the non-repeating feature vector of the first computing node in this reasoning and the repeated feature vector of the first computing node in this reasoning saved by the first computing node into the AI model for reasoning to obtain an inference result, which includes a label corresponding to the feature.
12. The method according to claim 11, characterized in that The method further comprises: The central processing unit constructs a communication matrix according to the non-repeated feature vector of the at least one computing node in the current reasoning and the storage location of the non-repeated feature vector; The central processing unit sends the communication matrix to the at least one computing node, the at least one computing node including the first computing node; The first computing node obtains, from the second computing node, a non-repeated feature vector of the first computing node in current reasoning, including: The first computing node obtains the non-repeated feature vector of the first computing node in this reasoning from the second computing node through many-to-many communication according to the communication matrix.
13. A distributed system, characterized in that: The distributed system includes a central processing unit and multiple computing nodes, each of the multiple computing nodes deploys an AI model, the multiple computing nodes respectively store at least one feature vector, the multiple computing nodes include a first computing node and a second computing node, the feature vector stored in the first computing node is different from the feature vector stored in the second computing node, the first computing node and the second computing node are used to train the AI model in parallel through multiple iterations, and one iteration of the multiple iterations includes updating a parameter of the AI model once; The central processing unit is configured to obtain input data of at least one computing node among the multiple computing nodes in this iteration, where the input data includes a batch of training data in the data set, and the training data includes features and labels corresponding to the features; The central processing unit is configured to identify repeated feature vectors and non-repeated feature vectors of the at least one computing node in the current iteration based on input data of the at least one computing node in the current iteration and input data of the at least one computing node in the previous n iterations, wherein the at least one computing node stores the repeated feature vectors of the at least one computing node in the current iteration during the previous n iterations; The first computing node is used to obtain the non-repeating feature vector of the first computing node in this iteration from the second computing node, and update the parameters of the AI model according to the non-repeating feature vector of the first computing node in this iteration, the repeated feature vector of the first computing node in this iteration saved by the first computing node, and the label.
14. The system according to claim 13, wherein: The first computing node is specifically configured to: Inputting the non-repeated feature vector of the first computing node in the current iteration and the repeated feature vector of the first computing node in the current iteration stored by the first computing node into the AI model for calculation, and obtaining the gradient of the feature vector input to the AI model at the first computing node based on the calculation result of the AI model and the label, wherein the feature vector input to the AI model includes the first feature vector that will be reused in the next n iterations; Obtaining the gradient of the first eigenvector at the second computing node through reduction communication; Update the parameters of the AI model according to the gradient of the first eigenvector at the first computing node and the gradient of the first eigenvector at the second computing node.
15. The system according to claim 14, wherein: When the AI model is a recommendation model, the first computing node is specifically configured to: The gradient of the first eigenvector at the second computing node and the gradient of the first eigenvector at the first computing node are accumulated, and the first eigenvector is reversely updated according to the accumulated first gradient.
16. The system according to claim 14 or 15, characterized in that The feature vector input to the AI model includes a second feature vector that will not be reused in the next n iterations, and the first computing node is further used to: Sending the gradient of the second eigenvector at the first computing node to a source computing node that pre-stores the second eigenvector through many-to-many communication; The source computing node is used to update the second eigenvector according to the gradient of the second eigenvector at the first computing node.
17. The system according to any one of claims 14 to 16, characterized in that The central processing unit is also used for: Identify the first feature vector based on multiple iterations of input data collected from the data set to obtain a recognition result; Send the recognition result to the first computing node.
18. The system according to any one of claims 13 to 17, characterized in that The central processing unit is also used for: Constructing a communication matrix according to the non-repeated eigenvectors of the at least one computing node in this iteration and the storage locations of the non-repeated eigenvectors; sending the communication matrix to the at least one computing node, the at least one computing node including the first computing node; The first computing node is specifically configured to: According to the communication matrix, the non-repeated feature vector of the first computing node in this iteration is obtained from the second computing node through many-to-many communication.
19. The system according to any one of claims 13 to 18, characterized in that The central processing unit is also used for: Constructing a query vector based on the repeated feature vector of the at least one computing node in this iteration and a storage location of the repeated feature vector, and sending the query vector to the at least one computing node, where the at least one computing node includes the first computing node; The first computing node is further configured to: The repeated feature vectors of this iteration stored in the first computing node are searched according to the query vector.
20. The system according to any one of claims 13 to 19, characterized in that The computing node includes a neural network processor NPU, a graphics processor GPU or a tensor processing unit TPU.
21. The system according to any one of claims 13 to 20, characterized in that The distributed system includes one or more central processing units. When the distributed system includes a first central processing unit and a second central processing unit, the first central processing unit is connected to the first computing node, and the second central processing unit is connected to the second computing node.
22. The system according to any one of claims 13 to 21, characterized in that The feature vector is a dense vector.
23. A distributed system, characterized in that: The distributed system includes a central processing unit and multiple computing nodes, each of the multiple computing nodes deploys an AI model, the multiple computing nodes respectively store at least one feature vector, the multiple computing nodes include a first computing node and a second computing node, the feature vector stored in the first computing node is different from the feature vector stored in the second computing node, and the first computing node and the second computing node are used for parallel reasoning based on the input feature vector; The central processing unit is configured to obtain input data of at least one computing node among the multiple computing nodes in this inference, wherein the input data includes a batch of inference data in the data set, and the inference data includes features; The central processing unit is further configured to identify repeated feature vectors and non-repeated feature vectors of the at least one computing node in the current reasoning based on input data of the at least one computing node in the current reasoning and input data of the at least one computing node in the previous n reasonings; The first computing node is used to obtain the non-repeating feature vector of the first computing node in this reasoning from the second computing node, and input the non-repeating feature vector of the first computing node in this reasoning and the repeated feature vector of the first computing node in this reasoning saved by the first computing node into the AI model for reasoning to obtain an inference result, where the inference result includes a label corresponding to the feature.
24. The system according to claim 23, wherein: The central processing unit is also used for: Constructing a communication matrix according to the non-repeated feature vectors of the at least one computing node in the current reasoning and the storage locations of the non-repeated feature vectors; sending the communication matrix to the at least one computing node, the at least one computing node including the first computing node; The first computing node is specifically configured to: According to the communication matrix, the non-repeated feature vector of the first computing node in this reasoning is obtained from the second computing node through many-to-many communication.
25. A computer-readable storage medium, characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 1 to 12.
26. A computer program product, characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 1 to 12.