Model training method, device, system and related equipment

By dividing the device into multiple communication domains and performing local gradient updates in each round of training and performing global gradient fusion at multiple rounds, the problems of low training efficiency and high communication resource consumption in existing AI model training are solved, and efficient AI model training is achieved.

CN117312839BActive Publication Date: 2025-05-16HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211148350.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-06-29
Filing Date
2022-09-20
Publication Date
2025-05-16
Estimated Expiration
2042-09-20

AI Technical Summary

Technical Problem

The existing AI model is trained in a distributed environment, and interacts with gradient data through the ring-allreduce method, resulting in low training efficiency and high communication resource consumption.

Method used

By dividing devices into multiple communication domains, devices within each communication domain interact independently and fuse gradient data, global gradient fusion and update only occurs when multiple rounds are spaced apart.

Benefits of technology

It improves the overall training efficiency of AI models, reduces the consumption of communication resources required during the training process, and ensures a high level of model training effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117312839B_ABST
    Figure CN117312839B_ABST
Patent Text Reader

Abstract

A model training method is provided, including: obtaining an AI model to be trained and determining multiple communication domains; in each round of AI model training, using local gradient data corresponding to each communication domain to update the AI ​​model, wherein the local gradient data corresponding to each communication domain is obtained by gradient fusion based on gradient data respectively generated by multiple devices in the communication domain, and when multiple rounds of AI model training are performed, using all gradient data to update the AI ​​model trained in each communication domain, wherein all gradient data is obtained by gradient fusion based on gradient data in multiple communication domains. In this way, since the AI ​​model is updated using the gradient data generated by all devices training the AI ​​model only after multiple rounds of training, this can alleviate the problem that the overall training progress of the AI ​​model is lowered due to the low progress of some communication domains in training the AI ​​model over a period of time, thereby improving the overall training efficiency of the AI ​​model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on June 29, 2022, with application number 202210760755.6 and application name “Methods, devices and systems for deep learning”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and in particular to a model training method, device, system and related equipment. Background Art

[0003] With the development of artificial intelligence (AI), the scale of AI models is also gradually increasing. For example, the number of parameters of the Pangu model in the field of natural language processing (NLP) can be as high as 200 billion, and the amount of training sample data can be as high as 40 terabytes. Correspondingly, the larger the number of model parameters and sample data, the higher the computing power required to train the AI ​​model.

[0004] At present, the huge computing power required for AI model training can be solved by performing distributed training on AI models. In specific implementation, the training samples of the AI ​​model can be equally divided into multiple sample subsets, and each sample subset and the AI ​​model are assigned to a device, so that each device uses a sample subset to iteratively train the AI ​​model and generate a gradient for updating the AI ​​model. Then, different devices interact with each other to generate their own gradient data and perform gradient fusion, calculate the global gradient data (that is, the data obtained after gradient fusion of the gradient data generated by all devices), and then update the parameters in the AI ​​model that each device is responsible for training according to the global gradient data. Based on the above method, the parameters in the AI ​​model are iteratively updated for multiple rounds, and the training of the AI ​​model is finally completed.

[0005] In actual application scenarios, different devices usually use the ring-allreduce method to exchange gradient data, that is, the flow of gradient data exchanged between different devices can form a ring, and by performing multiple data exchanges and gradient data fusion between different devices, each device can obtain global gradient data. However, this method of exchanging gradient data usually leads to low training efficiency of AI models and high consumption of communication resources. Summary of the invention

[0006] A model training method, device, system, storage medium, and computer program product are provided to improve the training efficiency of AI models and reduce the communication resources consumed in training AI models.

[0007] In a first aspect, an embodiment of the present application provides a model training method, which can be executed by a corresponding model training device. Specifically, the model training device obtains an AI model to be trained and determines multiple communication domains. The AI ​​model can be, for example, an AI model with a large model parameter amount or sample data amount, such as a Pangu model, etc. Each determined communication domain includes multiple devices. For example, all devices for training the AI ​​model can be divided into multiple communication domains, etc.; in each round of distributed training of the AI ​​model using devices in multiple communication domains, the model training device uses the local gradient data corresponding to each communication domain to update the AI ​​model trained in the communication domain, wherein the local gradient data corresponding to each communication domain is obtained by gradient fusion according to the gradient data respectively generated by the multiple devices in the communication domain, and when the AI ​​model is distributedly trained using the devices in the multiple communication domains at intervals of multiple rounds (the number of rounds intervals can be a fixed number of rounds or a random number of rounds), the model training device uses all the gradient data to update the AI ​​model trained in each communication domain, and the all the gradient data is obtained by gradient fusion according to the gradient data respectively generated by the multiple devices in the multiple communication domains, so as to update the global AI model using the global gradient data.

[0008] Since the AI ​​model is updated using the gradient data generated by all devices training the AI ​​model after every multiple rounds of training, and in each round of model training, each communication domain uses the gradient data generated by multiple devices within it to update the AI ​​model trained by the communication domain. This can alleviate the problem that the overall training progress of the AI ​​model is lowered due to the low progress of some communication domains in training the AI ​​model over a period of time, thereby improving the overall training efficiency of the AI ​​model. In addition, the model training device will use the global gradient data to update the AI ​​model after every multiple rounds of training, which can ensure that the training effect of the AI ​​model can reach a high level. On this basis, since in each round of model training, devices in different communication domains do not need to exchange gradient data, this can effectively reduce the communication resources required to train the AI ​​model.

[0009] In one possible implementation, during each round of training, the AI ​​model can be independently trained and updated for each of the multiple communication domains. Taking one of the target communication domains as an example, during each round of training, the multiple devices in the target communication domain exchange their generated gradient data, and perform gradient fusion based on the gradient data exchanged between the multiple devices to generate local gradient data corresponding to the target communication domain, thereby using the local gradient data corresponding to the target communication domain to update the AI ​​model trained by the target communication domain. For other communication domains, the AI ​​models they are responsible for can also be trained in a similar manner. In this way, each communication domain can achieve gradient updates of the AI ​​model in a local range by exchanging gradient data internally, and the efficiency of updating the AI ​​model between different communication domains may not be affected by other communication domains, thereby improving the overall training efficiency of the AI ​​model.

[0010] In a possible implementation, when multiple devices in the target communication domain interact with each other's generated gradient data, the specific method may be to obtain the version number of the activation operation and the version number of the interaction operation corresponding to the target communication domain, wherein the activation operation is used to trigger the interaction of gradient data between different devices in the target communication domain, and the interaction operation refers to the operation of interacting gradient data between different devices in the target communication domain, so that when the version number of the activation operation is greater than or equal to the version number of the interaction operation, the multiple devices in the target communication domain interact with each other's generated gradient data. In this way, each communication domain can avoid some communication domains from executing the interaction gradient data too many times by limiting the number of times the interaction gradient data is executed, thereby avoiding asynchronous conflicts between multiple communication domains.

[0011] In a possible implementation manner, the physical connection between the multiple devices in the target communication domain is a ring connection, such as a connection based on a HCCS ring mode.

[0012] In one possible implementation, when determining multiple communication domains, multiple devices with higher affinity can be classified into the same communication domain. In specific implementation, the model training device can obtain the device topology relationship, which is used to indicate the connection relationship between multiple devices for training the AI ​​model, and then divide the multiple devices used for training the AI ​​model according to the device topology relationship to obtain multiple communication domains, wherein the communication rate between different devices in each communication domain is higher than the communication rate between devices in different communication domains. In this way, by classifying devices with higher communication rates into the same communication domain, the efficiency of gradient data interaction between different devices in the communication domain during subsequent model training can be improved, thereby improving the overall training efficiency of the AI ​​model.

[0013] In one possible implementation, the user can configure the communication domain to which each device belongs for all devices for training the AI ​​model. In a specific implementation, the model training device can generate a first configuration interface, which is used to present the identifiers of multiple devices used to train the AI ​​model to the user, so that the user can configure the communication domain to which each device belongs on the first configuration interface, so that the model training device can respond to the user's first configuration operation, determine the communication domain to which each device in the multiple devices used to train the AI ​​model belongs, and thereby divide the multiple communication domains. In this way, the user can configure multiple communication domains to facilitate the user to intervene in the training of the AI ​​model and achieve a better model training effect.

[0014] In a possible implementation, before training the AI ​​model, the model training device may also generate a second configuration interface, which is used to present a variety of interaction strategies to the user, each of which is used to indicate a way of interacting gradient data between multiple devices in a communication domain, such as allgather, allreduce, ring-allreduce, having-doubling allreduce strategies, etc., so as to determine the way of interacting gradient data between multiple devices in each communication domain in response to the user's second configuration operation for the multiple interaction strategies. Among them, multiple devices in different communication domains can use the same interaction strategy to interact with gradient data, or different interaction strategies can be used to interact with gradient data, etc. In this way, the user can manually configure the interaction strategy in each communication domain so that the user can intervene in the training of the AI ​​model, such as configuring the most appropriate interaction strategy according to the characteristics of the devices in each communication domain, etc., so as to achieve a better model training effect.

[0015] In a possible implementation, devices in different communication domains are located in the same computing node, or devices in multiple communication domains are located in different computing nodes.

[0016] In one possible implementation, the devices in each communication domain include a processor, a chip, or a server, etc., so that distributed training of AI models can be achieved based on devices of different granularities.

[0017] In a second aspect, an embodiment of the present application provides a model training device, comprising: an acquisition module for acquiring an AI model to be trained; a determination module for determining multiple communication domains, each of the multiple communication domains including multiple devices; an update module for updating the AI ​​model trained by each communication domain using local gradient data corresponding to the communication domain in each round of distributed training of the AI ​​model using the devices in the multiple communication domains, the local gradient data corresponding to each communication domain being obtained by gradient fusion based on gradient data respectively generated by the multiple devices in the communication domain; and, when the AI ​​model is distributedly trained using the devices in the multiple communication domains at intervals of multiple rounds, the AI ​​model trained by each communication domain is updated using global gradient data, the global gradient data being obtained by gradient fusion based on gradient data respectively generated by the multiple devices in the multiple communication domains.

[0018] In a possible implementation, the update module is used to: interact with gradient data generated by each of the multiple devices in a target communication domain, where the target communication domain is one of the multiple communication domains; the target communication domain performs gradient fusion based on the gradient data interacted between the multiple devices to generate local gradient data corresponding to the target communication domain; the target communication domain uses the local gradient data corresponding to the target communication domain to update the AI ​​model trained by the target communication domain.

[0019] In a possible implementation, the update module is used to: obtain the version number of the activation operation and the version number of the interaction operation corresponding to the target communication domain, the activation operation is used to trigger the interaction of gradient data between different devices in the target communication domain, and the interaction operation is an operation for interacting gradient data between different devices in the target communication domain; when the version number of the activation operation is greater than or equal to the version number of the interaction operation, multiple devices in the target communication domain interact with each other's generated gradient data.

[0020] In a possible implementation manner, the physical connection between the multiple devices in the target communication domain is a ring connection.

[0021] In one possible implementation, the determination module is used to: obtain a device topology relationship, wherein the device topology relationship indicates a connection relationship between multiple devices used to train the AI ​​model; and divide the multiple devices used to train the AI ​​model according to the device topology relationship to obtain the multiple communication domains, wherein the communication rate between different devices in each communication domain is higher than the communication rate between devices in different communication domains.

[0022] In one possible implementation, the determination module is used to: generate a first configuration interface, the first configuration interface being used to present to a user the identifications of multiple devices used to train the AI ​​model; and determine, in response to the user's first configuration operation, the communication domain to which each of the multiple devices used to train the AI ​​model belongs.

[0023] In one possible implementation, the determination module is also used to: generate a second configuration interface before training the AI ​​model, the second configuration interface being used to present a plurality of interaction strategies to the user, each of the plurality of interaction strategies being used to indicate a manner of interacting gradient data between a plurality of devices in a communication domain; and determine a manner of interacting gradient data between the plurality of devices in each communication domain in response to a second configuration operation of the user for the plurality of interaction strategies.

[0024] In a possible implementation, devices in different communication domains are located in the same computing node, or devices in the multiple communication domains are located in different computing nodes respectively.

[0025] In a possible implementation, the device in each communication domain includes a processor, a chip, or a server.

[0026] Since the model training device provided in the second aspect corresponds to the model training method provided in the first aspect, the technical effects of the second aspect and each implementation method thereof can refer to the corresponding first aspect and each implementation method thereof, and will not be elaborated here.

[0027] In a third aspect, an embodiment of the present application provides a model training system, characterized in that the model training system includes multiple devices, and the model training system is used to execute the model training method in the above-mentioned first aspect and any implementation method of the first aspect.

[0028] In a fourth aspect, an embodiment of the present application provides a computing device, the computing device comprising a processor and a memory; the memory is used to store instructions, and the processor executes the instructions stored in the memory so that the computing device executes the model training method in the above-mentioned first aspect or any possible implementation of the first aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The computing device may also include a bus. The processor is connected to the memory via a bus. The memory may include a readable memory and a random access memory.

[0029] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein a program or instruction is stored in the computer-readable storage medium, and when the computer-readable storage medium is run on at least one computer, the at least one computer executes the model training method in the above-mentioned first aspect and any implementation of the first aspect.

[0030] In a sixth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on at least one computer, enables at least one computer to execute the model training method in the first aspect and any implementation of the first aspect.

[0031] In addition, the technical effects brought about by any one of the implementation methods in the second to sixth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0033] Figure 1 A schematic diagram of four devices interacting with each other according to an embodiment of the present application;

[0034] Figure 2 A schematic diagram of the architecture of an exemplary model training system provided in an embodiment of the present application;

[0035] Figure 3 A flowchart of a model training method provided in an embodiment of the present application;

[0036] Figure 4 A schematic diagram of the topological structure between NPU1 to NPU8 for training AI models;

[0037] Figure 5 A schematic diagram of an exemplary configuration interface provided in an embodiment of the present application;

[0038] Figure 6 A schematic diagram of the interaction of gradient data between four processors;

[0039] Figure 7 A schematic diagram of another exemplary configuration interface provided in an embodiment of the present application;

[0040] Figure 8 A schematic diagram showing processor 2 notifying other processors of interactive gradient data;

[0041] Fig. 9A schematic diagram of an exemplary server architecture provided in an embodiment of the present application;

[0042] Fig.10 A schematic diagram of a process for performing distributed training on the Pangu model provided in an embodiment of the present application;

[0043] Fig.11 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0044] Fig.12 A schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In actual application, when the number of parameters in the AI ​​model to be trained and the amount of sample data used to train the AI ​​model are large, the limited computing power of a single device may be difficult to complete the training of the AI ​​model alone. Therefore, the AI ​​model can be trained by integrating the computing power of multiple devices in a distributed training manner. Among them, the device for training the AI ​​model can be a processor-level device, such as a neural-network processing unit (NPU), a graphics processing unit (GPU), etc. Alternatively, the device for training the AI ​​model can be a chip-level device, such as multiple chips connected to a host, etc. Alternatively, the device for training the AI ​​model can be a server-level device, such as multiple independent servers, etc. Among them, when the multiple devices for training the AI ​​model are processor-level devices or chip-level devices, the multiple processors can be located in the same server (the server can constitute a computing node) or in different servers. When the multiple devices for training the AI ​​model are server-level devices, the multiple devices can be located in the same data center (the data center can be regarded as a computing node), or the multiple devices can be located in different data centers, that is, the AI ​​model can be distributedly trained across data centers.

[0046] In the process of iterative training of AI models, multiple devices usually use a ring-allreduce method to exchange the gradient data generated by each device in each round of AI model training, and update the AI ​​model parameters by generating new gradient data after gradient fusion of the gradient data obtained by the interaction.

[0047] Take the example of using 4 devices to train an AI model. Figure 1In the device 1 to device 4 shown, in each round of iterative training of the AI ​​model, each device uses a sample subset to train the AI ​​model and generates corresponding gradient data. Then, the device 1 to device 4 can divide the trained gradient data into 4 slices according to the number of devices. Among them, the gradient data of device 1 can be divided into slices a1, b1, c1, and d1, the gradient data of device 2 can be divided into slices a2, b2, c2, and d2, the gradient data of device 3 can be divided into slices a3, b3, c3, and d3, and the gradient data of device 4 can be divided into slices a4, b4, c4, and d4. Then, in the first interaction process between device 1 and device 4, device 1 sends slice a1 to device 2, device 2 sends slice b2 to device 3, device 3 sends slice c3 to device 4, and device 4 sends slice d4 to device 1. Then, each device can perform gradient fusion on the slices of gradient data stored in itself and the slices received to generate new gradient data. For example, device 1 can perform gradient fusion on the slice d4 sent by device 4 and the slice d1 stored by it to generate a new slice D1 of gradient data, and use D1 to cover slice d1. For example, assuming that slice d4 is {3,5,4,2,7,2} and slice d1 is {1,5,6,9,11,21}, then D1 obtained by gradient fusion of slices d4 and d1 can be {2,5,5,6,9,12} (the values ​​at corresponding positions are added and the average value is calculated, and the average value is rounded up). Similarly, device 2 can generate a new slice B2 of gradient data and use B2 to cover slice b2; device 3 can generate a new slice C2 of gradient data and use C2 to cover slice c2; device 4 can generate a new slice C4 of gradient data and use C4 to cover slice c4.

[0048] Then, device 1 interacts with device 4 for the second time. Device 1 sends slice D1 to device 2, device 2 sends slice A2 to device 3, device 3 sends slice B3 to device 4, and device 4 sends slice C4 to device 1. In addition, each device uses its own stored gradient data slices to perform gradient fusion with the received slices to generate new gradient data slices and replace the previously stored gradient data slices. Through multiple interactions between device 1 and device 4, each device has a gradient data slice, which is obtained by gradient fusion using the corresponding slices in device 1 and device 4, as shown in FIG. Figure 1 shown.

[0049] Then, devices 1 to 4 continue to interact, and share the gradient fusion slices stored in each device with other devices, so that each device can obtain the gradient fusion result obtained after gradient fusion based on the gradient data in all devices, such as Figure 1As shown, each device can use the gradient fusion result to update the parameters in the AI ​​model. In this way, multiple devices can complete a round of training process for the AI ​​model.

[0050] Since the speed of gradient data exchange between different devices is usually different, for example, due to load or resource specifications, some devices have a high latency in sending / receiving gradient data, which reduces the overall efficiency of gradient data exchange between multiple devices, thereby affecting the training efficiency of AI models. In addition, the frequent exchange of gradient data between multiple devices will also result in a high consumption of communication resources required for training AI models.

[0051] Based on this, an embodiment of the present application provides a model training method, which can be executed by a corresponding model training device to improve the training efficiency of the AI ​​model. In specific implementation, the model training device obtains the AI ​​model to be trained and determines multiple communication domains, each of which includes multiple devices; the model training device uses the local gradient data corresponding to each communication domain to update the AI ​​model trained by the communication domain in each round of training the AI ​​model using the multiple communication domains, wherein the local gradient data corresponding to each communication domain is obtained by gradient fusion based on the gradient data respectively generated by the multiple devices in the communication domain; and, in each interval of multiple rounds of training the AI ​​model using the multiple communication domains, the model training device uses the global gradient data to update the AI ​​model trained by each communication domain, wherein the global gradient data is obtained by gradient fusion based on the gradient data respectively generated by the multiple devices in the multiple communication domains.

[0052] Since the AI ​​model is updated using the gradient data generated by all devices training the AI ​​model only after multiple rounds of training, and in each round of model training, each communication domain independently uses the gradient data generated by multiple devices within it to update the AI ​​model trained by the communication domain. This can alleviate the problem of the overall training progress of the AI ​​model being slowed down due to the low progress of some communication domains in training the AI ​​model over a period of time, thereby improving the overall training efficiency of the AI ​​model.

[0053] For ease of understanding, taking the iterative training of the AI ​​model from device 1 to device 4 as an example, the model training device can classify device 1 and device 2 into communication domain 1, and device 3 and device 4 into communication domain 2. In each round of AI model training, device 1 and device 2 in communication domain 1 train the AI ​​model respectively and obtain corresponding gradient data. The model training device can perform gradient fusion on the gradient data in communication domain 1, and use the generated local gradient data to update the AI ​​model trained by device 1 and device 2. At the same time, the model training device will also use the local gradient data generated in communication domain 2 to update the AI ​​model trained by device 3 and device 4 during this round of training.

[0054] Assume that in the first round of AI model training, it takes 40 seconds for communication domain 1 to complete AI model training and updating, and 60 seconds for communication domain 2 to complete AI model training and updating; in the second round of AI model training, it takes 55 seconds for communication domain 1 to complete AI model training and updating, and 40 seconds for communication domain 2 to complete AI model training and updating, and then the AI ​​model is globally updated based on the local gradient data generated by the two communication domains 2 in the second round of training, assuming that it takes 10 seconds. Since communication domain 1 and communication domain 2 are independent of each other in the two rounds of model training, and the model training time of sub-communication domain 1 is 95 seconds (i.e., 40 seconds + 55 seconds), and the time of sub-communication domain 2 is 100 seconds (60 seconds + 40 seconds), the overall training time of the AI ​​model is 110 seconds (i.e., 100 seconds + 10 seconds), which is less than the 125 seconds (i.e., 60 seconds + 55 seconds + 10 seconds) generated by the existing ring-allreduce method of training AI models, thereby improving the overall training efficiency of the AI ​​model.

[0055] In addition, the model training device will use the global gradient data to update the AI ​​model every multiple rounds, which can ensure that the training effect of the AI ​​model can reach a high level. On this basis, since the gradient data does not need to be exchanged between devices in different communication domains during each round of model training, this can effectively reduce the communication resources required for training the AI ​​model.

[0056] Exemplarily, the model training device for executing the model training method can be deployed in Figure 2 The system architecture shown in Figure 1. Figure 2 The system architecture shown may include a deep learning framework 201 , a computing architecture 202 , firmware and drivers 203 , and a hardware layer 204 .

[0057] The deep learning framework 201 can integrate resources such as development components and pre-trained models, shield users from the perception of the underlying complex hardware, and provide users with services for rapid development of AI models. Exemplarily, the deep learning framework 201 can be, for example, a TensorFlow framework, a PyTorch framework, or a MindSpore framework, or can be other types of deep learning frameworks, which are not limited.

[0058] The computing architecture 202 is used to provide an open programming interface, support users to quickly build AI applications and services based on AI models, and call multiple processors in the hardware layer 204 to realize the parallelization capability of AI model training. Furthermore, the computing architecture 102 can also realize functions such as graph-level and operator-level compilation optimization and automatic tuning of AI models. Exemplarily, the computing architecture 202 can be, for example, a neural network computing architecture (compute architecture for neural networks, CANN), or other applicable architectures.

[0059] Firmware and driver 203 are used to respond to the call of computing architecture 202 to hardware layer 204, and use multiple processors in hardware layer 204 to perform corresponding data processing operations, such as using multiple processors in hardware layer 204 to parallelize the training of AI models.

[0060] The hardware layer 204 includes multiple processors, such as Figure 2 Processors 1 to 8 in the system also include other devices such as memory, network card, etc. ( Figure 2 The processor included in the hardware layer 204 may include, for example, a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a data processing unit (DPU), or other types of processors, which are not limited thereto.

[0061] Figure 2 The system architecture shown is only an exemplary description. In actual application, the model training device can also be deployed in other types of system architectures to implement distributed training of AI models. For example, in other possible system architectures, the hardware layer 204 can include multiple servers, that is, the AI ​​model can be distributedly trained at the server level.

[0062] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, various non-limiting implementation methods in the embodiments of the present application are exemplarily described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained based on the above content belong to the scope of protection of the present application.

[0063] like Figure 3As shown, it is a flow chart of a model training method in an embodiment of the present application, which can be applied to Figure 2 In practical application, this method can also be applied to other applicable system architectures. Figure 2 Taking the system architecture shown in the figure as an example for illustrative description, the method may specifically include:

[0064] S301: Obtain the AI ​​model to be trained.

[0065] In actual application scenarios, when a user develops an AI application on the deep learning framework 201, he can provide the deep learning framework 201 with an AI model for implementing the AI ​​application, so that the deep learning framework 201 can provide the AI ​​model to the model training device in the computing architecture 202 to trigger the training of the AI ​​model.

[0066] As an implementation example, a user can write a training script (an executable file written in a specific format) on the deep learning framework 201, and the training script can integrate the file of the AI ​​model built by the user on the deep learning framework 201. Then, the deep learning framework 201 can provide the training script to the model training device in the computing architecture 202, so that the model training device can parse the AI ​​model from the training script and perform distributed training on the AI ​​model according to the model training logic indicated by the training script.

[0067] As another implementation example, the deep learning framework 201 can provide a configuration interface to the user, which can present multiple AI models that have been built, so that the deep learning framework 201 can determine the AI ​​model to be trained selected by the user according to the user's selection operation for the AI ​​model. Furthermore, the configuration interface can also present a variety of deep learning algorithms that can be used to train the AI ​​model, so that the user can select a deep learning algorithm on the configuration interface and configure corresponding parameters based on the selected deep learning algorithm, such as learning rate, loss function, etc. Then, the deep learning framework 201 can provide the AI ​​model, deep learning algorithm and configured parameters selected by the user to the model training device, so that the model training device can perform distributed training on the AI ​​model based on the deep learning algorithm and the configured parameters. Of course, the model training device can also obtain the AI ​​model to be trained in other ways, and this embodiment does not limit this.

[0068] S302: Determine a plurality of communication domains, each of the plurality of communication domains including a plurality of devices.

[0069] In this embodiment, the model training device can use N processors in the hardware layer 204 to train the AI ​​model, where N is a positive integer (such as N is 8, 16, etc.). In addition, before training the AI ​​model, the model training device can first divide the N processors for training the AI ​​model into multiple sets, each set includes at least two processors, and the processors in each set can constitute a communication domain. For example, the model training device can divide 8 processors into 2 communication domains, each communication domain includes 4 processors, etc.

[0070] In the process of training AI models, processors in different communication domains can train AI models independently. For example, when the processor in communication domain 1 completes a round of AI model training, it can directly execute the next round of AI model training without waiting for the processor in communication domain 2 to complete a round of AI model training. In addition, the processors in each communication domain can exchange the gradient data generated by each round of AI model training through allreduce, ring-allreduce, etc.

[0071] For ease of understanding, this embodiment provides the following two implementation examples for determining multiple communication domains:

[0072] In the first implementation example, the model training device can classify multiple devices with higher affinity into the same communication domain according to the affinity between the devices. In specific implementation, the model training device can obtain the device topology relationship between the N processors in the hardware layer 204, and the device topology relationship can be used to indicate the connection relationship between the N processors, so that the model training device can divide the N processors used to train the AI ​​model according to the device topology relationship to obtain multiple communication domains, and the communication rate between different processing in each communication domain is higher than the communication rate between processors in different communication domains. In actual application, the model training device can classify multiple processors that are physically connected as a ring connection into a communication domain according to the device topology relationship, so that multiple communication domains can be divided. Exemplarily, between processors in the same communication domain, for example, a physical connection can be established based on the Huawei cache coherence system (HCCS) ring connection method.

[0073] For example, suppose the N processors used to train the AI ​​model include Figure 4As shown in FIG. 1 , NPU1 to NPU4 are connected in full mesh mode, NPU5 to NPU8 are also connected in full mesh mode, and NPU1 and NPU5 can be connected through the CPU. Then, the model training device can determine to classify NPU1 to NPU4 into communication domain 1 and NPU5 to NPU8 into communication domain 2 according to the topological structure between NPU1 to NPU8. Generally, the communication rate between NPU1 to NPU4 is higher than the communication rate between NPUs across communication domains.

[0074] In the second implementation example, the model training device can generate a configuration interface, for example, Figure 5 The configuration interface shown includes the identifiers of M processors that can be used to train the AI ​​model (such as processor names, etc.), where M is a positive integer greater than or equal to N, so that the model training device can present the configuration interface to the user through the deep learning framework 201, so that the user can select the N processors used to train the AI ​​model this time from the M processors presented, and further configure the communication domain to which each processor selected by the user belongs. Accordingly, the model training device can execute the initialization process of the communication domain, which can specifically be to determine the communication domain to which each of the N processors used to train the AI ​​model belongs in response to the user's configuration operation, thereby dividing into multiple communication domains and determining the size of each communication domain. Among them, the number of processors included in each communication domain can be the same or different.

[0075] For example, suppose Figure 5 The configuration interface shown shows 16 processors for the user to select, and based on the user's selection operation for the processor, it is determined to select processors 1 to 8 to train the AI ​​model. Then, the user can create two communication domains on the configuration interface, namely communication domain 1 and communication domain 2, and specify the communication domains to which processors 1 to 8 belong on the configuration interface. In this way, the model training device can determine the processors included in each communication domain according to the user's configuration of the communication domain to which each processor belongs, thereby obtaining multiple communication domains.

[0076] Of course, the above-mentioned implementation methods of determining the communication domain are only some exemplary explanations. In actual application, the model training device can also determine multiple communication domains through other methods. For example, after the model training device determines the multiple processors selected by the user, the processors located in the same server can be classified into one communication domain, etc. This embodiment does not limit this.

[0077] S303: In each round of training the AI ​​model using multiple communication domains, the local gradient data corresponding to each communication domain is used to update the AI ​​model trained by the communication domain. The local gradient data corresponding to each communication domain is obtained by gradient fusion based on the gradient data generated by multiple processors in the communication domain.

[0078] After determining multiple communication domains, the model training device can use processors in multiple communication domains to perform distributed training on the AI ​​model.

[0079] In specific implementation, the model training device can allocate an AI model and a subset of training samples for training the AI ​​model to each processor. Different processors are allocated the same AI model but different subsets of training samples, and each subset of training samples includes at least one training sample. Then, in each round of training, each processor can train the AI ​​model using the allocated subset of training samples, and generate gradient data based on the difference between the inference result of the AI ​​model based on the training sample and the actual result. The gradient data is used to perform gradient updates on the parameters in the AI ​​model. Since each processor trains the AI ​​model based on part of the training samples (that is, a subset of training samples), different processors can exchange their generated gradient data and perform gradient fusion, and perform gradient data on the parameters in the AI ​​model on each processor based on the results of the gradient fusion, so as to achieve the effect of training the AI ​​model using multiple subsets of training samples.

[0080] In this embodiment, during each round of model training, gradient data is not exchanged between all processors. Each processor only exchanges gradient data within its own communication domain and performs gradient fusion and model parameter update, so that the model training processes in different communication domains do not interfere with each other. Figure 4 Taking the AI ​​model training of NPU1 to NPU8 as an example, in each round of model training, NPU1 to NPU4 only exchange gradient data in communication domain 1, and do not exchange gradient data with NPU5 to NPU8 in communication domain 2. Similarly, NPU5 to NPU8 only exchange gradient data in communication domain 2, and do not exchange gradient data with NPUs in communication domain 1. In this way, after completing gradient data interaction, gradient fusion, and model parameter update, each communication domain can directly execute the next round of model training process without waiting for other communication domains to complete a round of AI model training.

[0081] Among them, multiple processors in each communication domain can exchange gradient data based on any strategy. For ease of understanding and explanation, the following is an exemplary explanation using one of the multiple communication domains (hereinafter referred to as the target communication domain) as an example. The processors in the remaining communication domains can exchange data with each other in a similar manner. The new gradient data generated by gradient fusion based on the gradient data of all processors in each communication domain is called local gradient data. Exemplarily, multiple processors in the target communication domain can exchange data based on the following implementation methods.

[0082] In the first implementation example, multiple processors in the target communication domain can interact with gradient data based on any one of the strategies of allgather, allreduce, ring-allreduce, and having-doubling allreduce.

[0083] Taking the interactive gradient data based on the allreduce strategy as an example, assume that the target communication domain includes 4 processors, namely processor 1, processor 2, processor 3 and processor 4, and the gradient data on these 4 processors are gradient data a, gradient data b, gradient data c and gradient data d, as shown in Figure 6 As shown. Then, in the first interaction, processor 1 can exchange gradient data with processor 2, and processor 3 and processor 4 can exchange gradient data. At this time, processor 1 and processor 2 can generate gradient data M by performing gradient fusion on gradient data a and gradient data b; processor 3 and processor 4 can generate gradient data N by performing gradient fusion on gradient data c and gradient data d. In the second interaction, processor 1 can exchange gradient data with processor 3, specifically, processor 1 sends gradient data M to processor 3, and processor 3 sends gradient data N to processor 1. At the same time, processor 2 and processor 4 exchange gradient data. At this time, processor 1 and processor 3 can generate gradient data X by performing gradient fusion on gradient data M and gradient data N, and the gradient data X is the data generated by performing gradient fusion on gradient data a, gradient data b, gradient data c and gradient data d. In addition, processor 2 and processor 4 can also generate gradient data X by performing gradient fusion on gradient data M and gradient data N. In this way, after two interactions, each processor can obtain gradient data X generated by performing gradient fusion on gradient data of all processors in the target communication domain.

[0084] Furthermore, the strategy of interactive gradient data used in each communication domain can be configured by the user. For example, the model training device can present the following to the user through the deep learning framework 201: Figure 7The configuration interface shown in FIG. 1 is used so that the user can configure the interaction strategy for each communication domain on the configuration interface. Specifically, Figure 7 As shown, the configuration interface can provide multiple interaction strategy candidates for each communication domain, such as allgather, allreduce, ring-allreduce, half-time allreduce, etc., so that the user can configure an interaction strategy for each communication domain from multiple candidates, and the interaction strategies adopted by different communication domains may be the same or different, which is not limited in this embodiment.

[0085] In the second implementation, the speeds at which different processors in the target communication domain train the AI ​​model may be different. Therefore, when different processors in the target communication domain exchange gradient data, the processor that has completed the AI ​​model training can exchange gradient data first without waiting for all other processors in the target communication domain to complete the AI ​​model training, thereby improving the efficiency of exchanging gradient data between multiple processors in the target communication domain.

[0086] Still taking the example of 4 processors included in the target communication domain, processors 1 to 4 train the AI ​​model in parallel. Assuming that processor 2 completes the training of the AI ​​model first in the target communication domain, processor 2 can generate an activation message and use the activation message to notify processors 1, 3, and 4 to start exchanging gradient data. In actual application, based on the physical connection and communication rules between processors, processor 2 can first send an activation message to processor 1 to notify processor 1 to start exchanging gradient data, and then send an activation message to processor 4, and processor 1 sends an activation message to processor 2 to notify processor 2 to start exchanging gradient data, such as Figure 8 As shown. In this way, if processor 1 completes the AI ​​model training second, processor 2 and processor 3 can directly exchange gradient data (and perform gradient data fusion). Then, if processor 3 completes the AI ​​model training third, processor 2 can exchange gradient data with processor 3 again. And, when processor 4 also completes the AI ​​model training, processor 2 exchanges gradient data with processor 4 again. In this way, processor 2 can obtain the gradient data generated by all processors in the target communication domain. Finally, processor 2 can send the local gradient data generated based on the gradient data of the four processors to the remaining processors, so as to use the local gradient data to update the parameters of the AI ​​model on each processor, as shown Figure 8 shown.

[0087] Furthermore, since each communication domain independently performs the training of the AI ​​model and the process of gradient data fusion during multiple rounds of training of the AI ​​model, each communication domain can avoid excessive execution of the interactive gradient data by some communication domains by limiting the number of interactive gradient data, thereby avoiding asynchronous conflicts between multiple communication domains. In specific implementation, the processor that first completes the AI ​​model training in the target communication domain can generate an activation message, which is used to notify the remaining processors in the target communication domain to start interactive gradient data, and the activation message includes the version number of the activation operation, which is used to trigger the interactive gradient data between different processors in the target communication domain. In addition, the processor that first completes the AI ​​model training can obtain the version number of the currently executed interactive operation, which is the operation of interactive gradient data between different devices in the target communication domain, so that the processor can compare the version number of the activation operation with the version number of the interactive operation. In addition, when the version number of the activation operation is greater than or equal to the version number of the interactive operation, the processor starts to interact with other processors for gradient data; otherwise, the interactive operation of gradient data is not performed between multiple processors.

[0088] In the third implementation example, multiple processors in the target communication domain can exchange gradient data through shared memory. In specific implementation, multiple processors in the target communication domain can be configured with shared memory, and multiple processors can access the shared memory. In this way, after each processor in the target communication domain completes a round of training for the AI ​​model and generates gradient data, the gradient data can be sent to a specified area in the shared memory, so that the shared memory can store gradient data generated by multiple processors. In this way, each processor can access the gradient data generated by all processors in the target communication domain from the shared memory, and by performing gradient fusion on these gradient data, local gradient data can be obtained.

[0089] It should be noted that in each round of AI model training, multiple processors in each communication domain can exchange gradient data and generate local gradient data in accordance with the above method. In addition, the above-mentioned methods for exchanging gradient data in the communication domain are only some examples. In other embodiments, multiple devices in each communication domain can also exchange gradient data in other ways, which is not limited in this embodiment.

[0090] S304: When the AI ​​model is distributedly trained using devices in multiple communication domains at intervals of multiple rounds, the AI ​​model trained in each communication domain is updated using global gradient data, wherein the global gradient data is obtained by gradient fusion based on gradient data generated by multiple processors in all communication domains.

[0091] Since each communication domain uses a partial subset of training samples to train the AI ​​model, the reasoning performance (such as reasoning accuracy, etc.) of the AI ​​model trained by each communication domain is usually difficult to reach the reasoning performance of the AI ​​model trained based on the full set of training samples. To this end, in this embodiment, after completing multiple rounds of training of the AI ​​model, gradient data can be exchanged between multiple communication domains to update the AI ​​model on each processor based on the gradient data generated by all processors. Specifically, the gradient data generated by all processors can be gradient fused, and the new gradient data generated by gradient fusion (hereinafter referred to as global gradient data) is used to update the parameters in the AI ​​model on each processor. In this way, the reasoning performance of the AI ​​model finally trained can usually reach the reasoning performance of the AI ​​model trained based on the full set of training samples.

[0092] In one possible implementation, each communication domain trains an AI model separately, and the processor in each communication domain can count the current number of iterations for the AI ​​model during each round of training of the AI ​​model. If the current number of iterations is an integer multiple of the T value, not only do the multiple processors in the communication domain exchange gradient data in the manner described above and generate local gradient data corresponding to the communication domain through gradient fusion, but the communication domain also exchanges local gradient data with other communication domains so that each communication domain can obtain the local gradient data generated by all communication domains. In this way, by performing gradient fusion on the local gradient data generated by all communication domains, global gradient data can be obtained, and the global gradient data can be used to update the parameters of the AI ​​model in each communication domain.

[0093] The way in which the local gradient data is exchanged between multiple communication domains is similar to the way in which the gradient data is exchanged between multiple processors in each communication domain. For example, the local gradient data can be exchanged between multiple communication domains based on any one of the strategies of allgather, allreduce, ring-allreduce, and half-allreduce, or the local gradient data can be exchanged in sequence by multiple communication domains in the order of completing (m*T) rounds of model training (m is a positive integer), or the local gradient data can be exchanged based on a shared storage area, etc., which is not limited in this embodiment.

[0094] In the above implementation, different communication domains exchange local gradient data, while in another possible implementation, multiple communication domains may directly exchange gradient data generated by each processor.

[0095] For example, when the number of times the processor in each communication domain iterates the training of the AI ​​model is an integer multiple of the T value, a processor in each communication domain can summarize the gradient data generated by each processor in the communication domain to obtain the gradient data set corresponding to the communication domain, and the gradient data set includes the gradient data generated by all processors in the communication domain, so that the processors responsible for summarizing the gradient data in multiple communication domains can exchange their respective gradient data sets. Alternatively, when the number of times the processor in each communication domain iterates the training of the AI ​​model is an integer multiple of the T value, all processors participating in the AI ​​model training directly exchange their respective generated gradient data, etc. In this way, each communication domain can obtain the gradient data generated by the processors in all communication domains, so that by gradient fusion of all gradient data, global gradient data can be obtained, so that the global gradient data can be used to update the parameters of the AI ​​model in each communication domain.

[0096] Among them, in the above implementation, an exemplary explanation is given by taking the example of multiple communication domains exchanging gradient data (or local gradient data) every (T-1) rounds. In other embodiments, the number of model training times between each interaction of gradient data between multiple communication domains may not be a fixed value. For example, in the process of distributed training of AI models, when the number of iterative training of the AI ​​model in each communication domain reaches 1000 times, the multiple communication domains exchange gradient data (or local gradient data) for the first time, and the number of model training times between them is 1000; then, when the number of iterative training of the AI ​​model in each communication domain reaches 1900 times, the multiple communication domains exchange gradient data (or local gradient data) for the second time, and the number of model training times between them is 900; when the number of iterative training of the AI ​​model in each communication domain reaches 2700 times, the multiple communication domains exchange gradient data (or local gradient data) for the third time, and the number of model training times between them is 800; when the number of iterative training of the AI ​​model in each communication domain reaches 3400 times, the multiple communication domains exchange gradient data (or local gradient data) for the fourth time, and the number of model training times between them is 700, and so on.

[0097] It should be noted that in this embodiment, the device in the communication domain is specifically a processor for illustrative description. In other embodiments, the device in the communication domain may also be a chip or a server. The specific implementation process of distributed training of the AI ​​model can be understood by referring to the relevant description of this embodiment and will not be repeated here.

[0098] In this embodiment, since the AI ​​model is updated with the gradient data generated by all processors training the AI ​​model only after each multiple rounds of model training, and in each intermediate round of model training, each communication domain separately uses the gradient data generated by multiple processors inside it to update the AI ​​model trained by the communication domain, this can alleviate the impact of the low progress of the AI ​​model training in some communication domains on the overall training progress of the AI ​​model, that is, it can improve the overall training efficiency of the AI ​​model. For example, the progress of set 1 in the first round is reduced by 3 seconds, and the progress of set 2 in the second round is reduced by 5 seconds. The overall progress will not be reduced to 8 seconds, but the slowest progress, that is, 5 seconds. In addition, the use of global gradient data to update the AI ​​model after each multiple rounds of model training can ensure that the training effect of the AI ​​model reaches a high level. On this basis, since the processors in different communication domains do not need to exchange gradient data during each intermediate round of model training, this can effectively reduce the communication resources required to train the AI ​​model.

[0099] Next, we will introduce the specific implementation process of distributed training of AI models in combination with specific application scenarios. Figure 1 The system architecture described above can be deployed in a server that includes 4 CPUs and can be connected to 8 NPU chips. Fig. 9 As shown, the NPU1 to NPU8 in the server can be used to implement distributed training of the Pangu model (an AI model). In other embodiments, the distributed training of the Pangu model can also be implemented based on the NPU chips in multiple servers. The training method is similar to the implementation method of distributed training of the Pangu model using multiple NPUs in one server, which can be understood by reference.

[0100] exist Fig. 9 In the server shown, each CPU can support 8 4th generation double data rate 4dual inline memory modules (DDR4 DIMMs), and CPU1 to CPU4 can be fully meshed. The CPUs in the server can provide a bandwidth capacity of 90GB / s (gigabytes per second), where each CPU can provide a unidirectional bandwidth of 30GB / s and a bidirectional bandwidth of 60GB / s.

[0101] Among the 8 NPU chips connected to the server, NPU1 to NPU4 can be fully interconnected and can be located on one NPU motherboard, and NPU5 to NPU8 can be fully interconnected and can be located on another NPU motherboard. In addition, there is a connection between the 8 NPU chips connected to the server and the CPU, for example, they can be connected based on a peripheral component interconnect express (PCIE) bus, etc. ( Fig. 9 Only part of the connection between NPU and CPU is shown in the figure), so that NPU1 to NPU4 can exchange data with NPU5 to NPU8 through the CPU in the server. Each NPU motherboard can provide a bandwidth capacity of 90GB / s, of which each NPU can provide a unidirectional bandwidth of 30GB / s and a bidirectional bandwidth of 60GB / s. Fig. 9 The server shown in FIG. 1 can realize distributed training of the Pangu model. The distributed training process is as follows: Fig.10 The user can provide a training script to the server, which may include a file of the Pangu model, and specify to use NPU1 to NPU8 to train the Pangu model, and define NPU1 to NPU4 as belonging to communication domain 1, and define NPU5 to NPU8 as belonging to communication domain 2.

[0102] In this way, the CPU on the host side of the server can parse the Pangu model to be trained from the training script, and determine the multiple NPUs used for distributed training of the Pangu model and the communication domain to which each NPU belongs.

[0103] Then, the CPU can extract the computational graph according to the training script. The computational graph includes multiple nodes, and there are edges connecting different nodes. The nodes in the computational graph are used to indicate the calculations defined in the training script, and the edges between the nodes are used to indicate the dependencies between different calculations. The extracted computational graph can be saved to a flash card (trans-flashcard)

[0104] Then, the CPU can compile the computation graph in the flash card, generate an intermediate representation (IR), and provide the IR to the compiler. The compiler can define one or more operator libraries, such as Fig.10The neural network (NN) operator library, Huawei collective communication library (HCCL) operator library, etc. are shown. Exemplarily, the NN operator library may include convolutional layer operators, pooling layer operators, loss functions, etc.; the HCCL operator library may include operators for defining data communication methods, such as allreduce operators, allgather operators, etc.

[0105] In this way, the CPU can use the compiler to determine the operators that need to be executed sequentially for the distributed training of the Pangu model, generate corresponding device instructions based on them, and send the device instructions to the NPU on the device side.

[0106] NPU1 and NPU8 on the device side can execute the corresponding operators in a loop and perform gradient updates on the Pangu model based on the device instructions issued by the host side until the iteration termination condition is met, thereby realizing distributed training of the Pangu model. In the process of distributed training of the Pangu model, NPU1 to NPU4 in communication domain 1 and NPU5 to NPU8 in communication domain 2 train the Pangu model separately, and communication domain 1 and communication domain 2 interactively train the gradient data generated by the Pangu model at each interval (T-1) round of model training to realize global gradient updates of the Pangu model. For the specific training process, please refer to the aforementioned Figure 3 The description of the relevant aspects of the illustrated embodiment will not be repeated here.

[0107] Finally, after completing the distributed training for the Pangu model, the device side can send the training results to the host side. The training results may include, for example, the trained Pangu model, the attribute information of the Pangu model (such as reasoning accuracy), etc.

[0108] Combined with the above Figures 1 to 10 , describes in detail the model training method provided by this application, and will be combined with Figure 11 to Figure 12 , respectively describe the model training device and computing equipment provided by this application.

[0109] With the same inventive concept as the above method, the present application embodiment also provides a model training device. Fig.11 , is a structural diagram of a model training device provided in an embodiment of the present application, Fig.11 The model training device 1100 shown in the figure may be, for example, the above Figure 3 The model training device mentioned in the embodiment shown. Fig.11 As shown, the model training device 1100 includes:

[0110] An acquisition module 1101 is used to acquire an AI model to be trained;

[0111] A determination module 1102 is configured to determine a plurality of communication domains, each of the plurality of communication domains including a plurality of devices;

[0112] An updating module 1103 is configured to update the AI ​​model trained by each communication domain using local gradient data corresponding to each communication domain in each round of distributed training of the AI ​​model using the devices in the multiple communication domains, the local gradient data corresponding to each communication domain being obtained by gradient fusion based on gradient data respectively generated by the multiple devices in the communication domain; and, when the AI ​​model is distributedly trained using the devices in the multiple communication domains at intervals of multiple rounds, update the AI ​​model trained by each communication domain using global gradient data, the global gradient data being obtained by gradient fusion based on gradient data respectively generated by the multiple devices in the multiple communication domains.

[0113] In a possible implementation, the updating module 1103 is used to:

[0114] Multiple devices in a target communication domain exchange gradient data generated by each device, wherein the target communication domain is one of the multiple communication domains;

[0115] The target communication domain performs gradient fusion according to the gradient data exchanged between the multiple devices to generate local gradient data corresponding to the target communication domain;

[0116] The target communication domain updates the AI ​​model trained by the target communication domain using the local gradient data corresponding to the target communication domain.

[0117] In a possible implementation, the updating module 1103 is used to:

[0118] Obtaining a version number of an activation operation and a version number of an interaction operation corresponding to the target communication domain, wherein the activation operation is used to trigger the interaction of gradient data between different devices in the target communication domain, and the interaction operation is an operation for interacting gradient data between different devices in the target communication domain;

[0119] When the version number of the activation operation is greater than or equal to the version number of the interaction operation, the plurality of devices in the target communication domain interact with each other to generate their own gradient data.

[0120] In a possible implementation manner, the physical connection between the multiple devices in the target communication domain is a ring connection.

[0121] In a possible implementation, the determining module 1102 is configured to:

[0122] Acquire a device topology relationship, where the device topology relationship indicates a connection relationship between multiple devices used to train the AI ​​model;

[0123] According to the device topology relationship, multiple devices used to train the AI ​​model are divided to obtain the multiple communication domains, wherein the communication rate between different devices in each communication domain is higher than the communication rate between devices in different communication domains.

[0124] In a possible implementation, the determining module 1102 is configured to:

[0125] Generate a first configuration interface, wherein the first configuration interface is used to present to a user the identifiers of multiple devices used to train the AI ​​model;

[0126] In response to a first configuration operation by a user, a communication domain to which each of the multiple devices used to train the AI ​​model belongs is determined.

[0127] In a possible implementation manner, the determining module 1102 is further configured to:

[0128] Before training the AI ​​model, generating a second configuration interface, the second configuration interface is used to present multiple interaction strategies to the user, each of the multiple interaction strategies is used to indicate a way of interacting gradient data between multiple devices in the communication domain;

[0129] In response to a second configuration operation of the user for the plurality of interaction strategies, a manner of exchanging gradient data between the plurality of devices in each communication domain is determined.

[0130] In a possible implementation, devices in different communication domains are located in the same computing node, or devices in the multiple communication domains are located in different computing nodes respectively.

[0131] In a possible implementation, the device in each communication domain includes a processor, a chip, or a server.

[0132] Fig.11 The data processing device 100 shown corresponds to Figure 3 The data processing device in the illustrated embodiment, and therefore the specific implementation of each functional module in the data processing device 100 and the technical effects thereof, can be found in the relevant description of the aforementioned embodiment, and will not be elaborated here.

[0133] The present application also provides a computing device, such as Fig.12As shown, the computing device 1200 may include a communication interface 1210 and a processor 1220. Optionally, the computing device 1200 may also include a memory 1230. The memory 1230 may be disposed inside the computing device 1200 or outside the computing device 1200. Figure 3 Each action performed by the data processing device in the embodiment shown can be implemented by the processor 1220. The processor 1220 can obtain the AI ​​model to be trained and multiple communication domains through the communication interface 1210, and is used to implement Figure 3 In the implementation process, each step of the processing flow can be completed by the hardware integrated logic circuit in the processor 1220 or the instructions in the form of software. Figure 3 For the sake of brevity, it will not be described here. The program code executed by the processor 1220 to implement the above method can be stored in the memory 1230. The memory 1230 is connected to the processor 1220, such as a coupling connection.

[0134] Some features of the embodiments of the present application may be completed / supported by the processor 1220 executing program instructions or software codes in the memory 1230. The software components loaded on the memory 1230 may be summarized in terms of function or logic.

[0135] Any communication interface involved in the embodiments of the present application may be a circuit, a bus, a transceiver or any other device that can be used for information exchange. For example, the communication interface 1210 in the computing device 1200, illustratively, the other device may be a device connected to the computing device 1200, etc.

[0136] The processor involved in the embodiments of the present application may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0137] The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, modules or modules, which can be electrical, mechanical or other forms, and is used for information exchange between devices, modules or modules.

[0138] The processor may operate in conjunction with a memory. The memory may be a non-volatile memory, such as a hard disk or a solid-state drive, or a volatile memory, such as a random access memory. The memory is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0139] The specific connection medium between the communication interface, processor and memory is not limited in the embodiments of the present application. For example, the memory, processor and communication interface may be connected via a bus. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0140] Based on the above embodiments, the embodiments of the present application further provide a computer storage medium, in which a software program is stored, and when the software program is read and executed by one or more processors, the method performed by the model training device provided in any one or more of the above embodiments can be implemented. The computer storage medium may include: a U disk, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and other media that can store program codes.

[0141] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices, systems, storage media or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0142] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0143] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0145] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances. This is just a way of distinguishing objects with the same attributes when describing the embodiments of this application.

[0146] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the scope of the embodiments of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A model training method, characterized in that: The method comprises: Get the AI ​​model to be trained; determining a plurality of communication domains, each of the plurality of communication domains comprising a plurality of devices; In each round of distributed training of the AI ​​model using the devices in the multiple communication domains, the AI ​​model trained by the communication domain is updated using the local gradient data corresponding to each communication domain, and the local gradient data corresponding to each communication domain is obtained by gradient fusion based on the gradient data respectively generated by the multiple devices in the communication domain; When the AI ​​model is distributedly trained using the devices in the multiple communication domains for multiple rounds, the AI ​​model trained separately in each communication domain is updated using global gradient data, and the global gradient data is obtained by gradient fusion based on the gradient data respectively generated by the multiple devices in the multiple communication domains.

2. The method according to claim 1, characterized in that The updating of the AI ​​model trained by each communication domain using the local gradient data corresponding to the communication domain includes: Multiple devices in a target communication domain exchange gradient data generated by each device, wherein the target communication domain is one of the multiple communication domains; The target communication domain performs gradient fusion according to the gradient data exchanged between the multiple devices to generate local gradient data corresponding to the target communication domain; The target communication domain updates the AI ​​model trained by the target communication domain using the local gradient data corresponding to the target communication domain.

3. The method according to claim 2, characterized in that The multiple devices in the target communication domain exchange the gradient data generated by each other, including: Obtaining a version number of an activation operation and a version number of an interaction operation corresponding to the target communication domain, wherein the activation operation is used to trigger the interaction of gradient data between different devices in the target communication domain, and the interaction operation is an operation for interacting gradient data between different devices in the target communication domain; When the version number of the activation operation is greater than or equal to the version number of the interaction operation, the plurality of devices in the target communication domain interact with each other to generate their own gradient data.

4. The method according to claim 2, characterized in that: The physical connection between the multiple devices in the target communication domain is a ring connection.

5. The method according to claim 1, characterized in that The determining of multiple communication domains includes: Acquire a device topology relationship, where the device topology relationship indicates a connection relationship between multiple devices used to train the AI ​​model; According to the device topology relationship, multiple devices used to train the AI ​​model are divided to obtain the multiple communication domains, wherein the communication rate between different devices in each communication domain is higher than the communication rate between devices in different communication domains.

6. The method according to claim 1, characterized in that The determining of multiple communication domains includes: Generate a first configuration interface, wherein the first configuration interface is used to present to a user the identifiers of multiple devices used to train the AI ​​model; In response to a first configuration operation by a user, a communication domain to which each of the multiple devices used to train the AI ​​model belongs is determined.

7. The method according to claim 1, characterized in that Before training the AI ​​model, the method further includes: generating a second configuration interface, wherein the second configuration interface is used to present a plurality of interaction strategies to a user, each of the plurality of interaction strategies being used to indicate a manner in which gradient data is exchanged between a plurality of devices in a communication domain; In response to a second configuration operation of the user for the plurality of interaction strategies, a manner of exchanging gradient data between the plurality of devices in each communication domain is determined.

8. The method according to claim 1, characterized in that Devices in different communication domains are located in the same computing node, or devices in the multiple communication domains are located in different computing nodes respectively.

9. The method according to any one of claims 1 to 8, characterized in that: Devices in each communication domain include processors, chips, or servers.

10. A model training device, characterized in that: The device comprises: The acquisition module is used to obtain the AI ​​model to be trained; A determination module, configured to determine a plurality of communication domains, each of the plurality of communication domains comprising a plurality of devices; An update module is used to update the AI ​​model trained by each communication domain using local gradient data corresponding to each communication domain in each round of distributed training of the AI ​​model using the devices in the multiple communication domains, the local gradient data corresponding to each communication domain being obtained by gradient fusion based on gradient data respectively generated by the multiple devices in the communication domain; and when the AI ​​model is distributedly trained using the devices in the multiple communication domains at intervals of multiple rounds, the AI ​​model trained by each communication domain is updated using global gradient data, the global gradient data being obtained by gradient fusion based on gradient data respectively generated by the multiple devices in the multiple communication domains.

11. The device according to claim 10, characterized in that The update module is used to: Multiple devices in a target communication domain exchange gradient data generated by each device, wherein the target communication domain is one of the multiple communication domains; The target communication domain performs gradient fusion according to the gradient data exchanged between the multiple devices to generate local gradient data corresponding to the target communication domain; The target communication domain updates the AI ​​model trained by the target communication domain using the local gradient data corresponding to the target communication domain.

12. The device according to claim 11, characterized in that The update module is used to: Obtaining a version number of an activation operation and a version number of an interaction operation corresponding to the target communication domain, wherein the activation operation is used to trigger the interaction of gradient data between different devices in the target communication domain, and the interaction operation is an operation for interacting gradient data between different devices in the target communication domain; When the version number of the activation operation is greater than or equal to the version number of the interaction operation, the plurality of devices in the target communication domain interact with each other to generate their own gradient data.

13. The device according to claim 11, characterized in that The physical connection between the multiple devices in the target communication domain is a ring connection.

14. The device according to claim 10, characterized in that The determining module is used to: Acquire a device topology relationship, where the device topology relationship indicates a connection relationship between multiple devices used to train the AI ​​model; According to the device topology relationship, multiple devices used to train the AI ​​model are divided to obtain the multiple communication domains, wherein the communication rate between different devices in each communication domain is higher than the communication rate between devices in different communication domains.

15. The device according to claim 10, characterized in that The determining module is used to: Generate a first configuration interface, wherein the first configuration interface is used to present to a user the identifiers of multiple devices used to train the AI ​​model; In response to a first configuration operation by a user, a communication domain to which each of the multiple devices used to train the AI ​​model belongs is determined.

16. The device according to claim 10, characterized in that The determining module is further used for: Before training the AI ​​model, generating a second configuration interface, the second configuration interface is used to present multiple interaction strategies to the user, each of the multiple interaction strategies is used to indicate a way of interacting gradient data between multiple devices in the communication domain; In response to a second configuration operation of the user for the plurality of interaction strategies, a manner of exchanging gradient data between the plurality of devices in each communication domain is determined.

17. The device according to claim 10, characterized in that Devices in different communication domains are located in the same computing node, or devices in the multiple communication domains are located in different computing nodes respectively.

18. The device according to any one of claims 10 to 17, characterized in that Devices in each communication domain include processors, chips, or servers.

19. A model training system, characterized in that: The model training system includes multiple devices, and the model training system is used to execute the method as described in any one of claims 1 to 9.

20. A computing device, characterized in that including a processor and a memory; The processor is configured to execute instructions stored in the memory so that the computing device performs the method according to any one of claims 1 to 9.

21. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed on at least one computing device, enable the at least one computing device to perform the method according to any one of claims 1 to 9.

22. A computer program product comprising instructions, characterized in that When the method is executed on at least one computing device, the at least one computing device is caused to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Model training method, content generation method and related device

    CN111265881A

  • Method based on distributed system training model, equipment and program product

    CN113656175A