Computing method, computing system, computing device cluster and storage medium

By storing the intermediate results of the forward computing task to the processing unit and performing the recomputation task during the AI model training process, the problem of wasting computing power resources of the acceleration device is solved, and more efficient computing performance is achieved.

CN120297340APending Publication Date: 2025-07-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410046110.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the process of AI model training, the acceleration device needs to re-execute forward calculation before performing backward calculations, resulting in wasted computing resources and affecting operating performance.

Method used

After the host responds to the training request, the intermediate execution result of the forward computing task is stored in the processing unit, and the processing unit performs the recalculation task, reducing the computing burden of the acceleration device, and using the computing and storage functions of the processing unit to realize near data processing.

Benefits of technology

It effectively saves the computing power resources of the acceleration equipment, reduces the computing overhead and data transmission volume, and improves the overall computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297340A_ABST
    Figure CN120297340A_ABST
Patent Text Reader

Abstract

The invention discloses a computing method, a computing system, a computing device cluster and a storage medium, and relates to the technical field of AI (artificial intelligence), the method is applied to the computing system, the system comprises a host, an acceleration device and a processing unit, and the method comprises the steps that the host triggers a training process for an AI model in response to a training request for the AI model, and in the process, the AI model is trained; the acceleration device executes the forward calculation task and stores an intermediate execution result of the forward calculation task to the processing unit, and the processing unit executes the re-calculation task to obtain an execution result of the forward calculation task, so that the acceleration device obtains the execution result of the forward calculation task from the processing unit to execute the backward calculation task. Therefore, the computing power resource of the acceleration equipment is effectively saved, and the computing overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence (AI) technology, and particularly relates to a calculation method, a calculation system, a cluster of computing devices, and a storage medium. Background Art

[0002] With the rapid development of AI technology, a series of acceleration devices have emerged to provide computing power for calculating matrices, vectors, etc., so as to accelerate the calculation process of AI models. Currently, during the training process of AI models, recomputation can effectively reduce the memory occupation of acceleration devices.

[0003] In related technologies, an implementation method of using recomputation during the training process of AI models is, for example: taking some layers of the AI model as checkpoints. When the acceleration device performs forward calculation, it transfers the activation values (activation, or intermediate execution results) of the checkpoints to the host memory. Before the acceleration device performs backward calculation, it transfers the activation values of the checkpoints from the host memory and re-executes the forward calculation based on these activation values to obtain the activation values of each layer in the AI model, so as to save the calculation amount of the activation values of some layers.

[0004] However, although the above method saves the calculation amount of the activation values of some layers, the acceleration device still needs to re-execute the forward calculation before performing the backward calculation, which wastes the computing power resources of the acceleration device and thus affects the running performance of the acceleration device. Summary of the Invention

[0005] Embodiments of this application provide a calculation method, a calculation system, a cluster of computing devices, and a storage medium, which can save the computing power resources of acceleration devices and improve the running performance of acceleration devices during the training process of AI models.

[0006] In a first aspect, this application provides a calculation method, which is applied to a scenario where recomputation is used to reduce the memory occupation of acceleration devices during the training process of an AI model. Here, this application does not limit the type of the AI model or the type of the acceleration device. Schematically, this method is applied to a calculation system, which includes a host, an acceleration device, and a processing unit. The method includes:

[0007] The host responds to a training request for an artificial intelligence AI model and sends a forward calculation task of the AI model to the acceleration device;

[0008] The acceleration device executes the forward calculation task and stores the intermediate execution results of the forward calculation task in the processing unit;

[0009] The host sends a recomputation task of the AI model to the processing unit, and the recomputation task instructs to execute the forward computation task to obtain the execution result of the forward computation task;

[0010] The processing unit executes the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task;

[0011] The host sends a backward computation task of the AI model to the acceleration device;

[0012] The acceleration device obtains the execution result of the forward computation task from the processing unit and executes the backward computation task based on the execution result of the forward computation task.

[0013] In the above process, the host triggers the training process of the AI model in response to a training request for the AI model. In this process, the acceleration device executes the forward computation task and stores the intermediate execution result of the forward computation task in the processing unit. The processing unit executes the recomputation task to obtain the execution result of the forward computation task, so that the acceleration device obtains the execution result of the forward computation task from the processing unit to execute the backward computation task. In this way, the computing power resources of the acceleration device are effectively saved, the computing overhead is reduced, and moreover, since the data volume of the execution result of the forward computation task is smaller than the data volume of the intermediate execution result, this method can effectively reduce the cross-device data transmission volume.

[0014] In some embodiments, the processing unit includes a computing subunit and a storage subunit;

[0015] The processing unit executes the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task, including:

[0016] The computing subunit executes the recomputation task based on the intermediate execution result of the forward computation task in the storage subunit, obtains the execution result of the forward computation task, and stores the execution result of the forward computation task in the storage subunit.

[0017] In some embodiments, the computing subunit and the storage subunit are integrated on the same chip. In this way, since the processing unit is integrated on one chip, in-memory processing for the recomputation task is achieved, which can not only save the computing power resources of the acceleration device, but also increase the memory access bandwidth, reduce data movement, and improve the overall computing efficiency.

[0018] In some embodiments, the computing subunit and the storage subunit are integrated in the same die of the chip. In this way, since the processing unit is integrated in the same die of one chip, in-memory processing for the recomputation task is achieved, which can not only save the computing power resources of the acceleration device, but also increase the memory access bandwidth, reduce data movement, and improve the overall computing efficiency.

[0019] In some embodiments, the computing system is a distributed computing system, and the processing unit is a computing node in the distributed computing system. In this way, since the processing unit is a computing node in the computing system, near-memory processing for re-computation tasks in a distributed computing scenario is achieved, which can not only save the computing power resources of the acceleration device, but also increase the memory access bandwidth, reduce data migration, and improve the overall computing efficiency.

[0020] In some embodiments, the method further includes:

[0021] The host sends the task information of the re-computation task to the processing unit, and the task information indicates the task identifier and data identifier of the re-computation task;

[0022] The processing unit executes the re-computation task based on the intermediate execution result of the forward computation task in the processing unit, including: the processing unit executes the re-computation task based on the task information and the intermediate execution result of the forward computation task in the processing unit.

[0023] In some embodiments, the method further includes: when the training request involves a re-computation task and cross-device data transmission, the host generates the task information. In this process, the host has the function of identifying whether the operator involves a re-computation task and cross-device data transmission. In this way, the automatic generation of task information can be realized without the upper-layer user being aware.

[0024] In a second aspect, the present application provides a computing method, which is applied to a processing unit in a computing system. The system further includes a host and an acceleration device. The host is used to send the forward computation task and backward computation task of the AI model to the acceleration device in response to a training request for the AI model; the method includes:

[0025] Receiving the re-computation task of the AI model sent by the host, where the re-computation task instructs to execute the forward computation task to obtain the execution result of the forward computation task;

[0026] Executing the re-computation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task, so that the acceleration device executes the backward computation task based on the execution result of the forward computation task. The intermediate execution result of the forward computation task is obtained by the acceleration device executing the forward computation task and stored by the acceleration device in the processing unit.

[0027] In some embodiments, the processing unit includes a computing subunit and a storage subunit. Among them, executing the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task includes:

[0028] Through the computing subunit, execute the recomputation task based on the intermediate execution result of the forward computation task in the storage subunit to obtain the execution result of the forward computation task, and store the execution result of the forward computation task in the storage subunit.

[0029] In some embodiments, the computing subunit and the storage subunit are integrated on the same chip.

[0030] In some embodiments, the computing subunit and the storage subunit are integrated in the same die of the chip.

[0031] In some embodiments, the system is a distributed computing system, and the processing unit is a computing node in the distributed computing system.

[0032] In some embodiments, the method further includes:

[0033] Receiving the task information of the recomputation task sent by the host, where the task information indicates the task identifier and data identifier of the recomputation task;

[0034] The executing the recomputation task based on the intermediate execution result of the forward computation task in the processing unit includes: executing the recomputation task based on the task information and the intermediate execution result of the forward computation task in the processing unit.

[0035] In a third aspect, the present application provides a computing method, which is applied to an acceleration device in a computing system. The system further includes a host and a processing unit. The method includes:

[0036] Receiving the forward computation task of the AI model sent by the host in response to a training request for the AI model;

[0037] Executing the forward computation task, and storing the intermediate execution result of the forward computation task in the processing unit. The processing unit is used to receive the recomputation task of the AI model sent by the host, and execute the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task;

[0038] Receiving the backward computation task of the AI model sent by the host;

[0039] Obtain the execution result of the forward calculation task from the processing unit, and execute the backward calculation task based on the execution result of the forward calculation task.

[0040] In a fourth aspect, the present application provides a computing system, which includes a host, an acceleration device, and a processing unit, and the system is used to implement the computing method provided in the foregoing first aspect or any possible implementation manner of the first aspect.

[0041] In a fifth aspect, the present application provides a processing unit, which includes a computing subunit and a storage subunit, and the processing unit is used to implement the steps executed by the processing unit in the computing method provided in the foregoing first aspect or any possible implementation manner of the first aspect.

[0042] In a sixth aspect, the present application provides an acceleration device, which includes a processor and a memory, and the processor is used to execute at least one piece of program code stored in the memory, so that the acceleration device implements the steps executed by the acceleration device in the computing method provided in the foregoing first aspect or any possible implementation manner of the first aspect.

[0043] In a seventh aspect, the present application provides a computing device cluster, which includes at least one computing device, and each computing device includes a processor and a memory;

[0044] The processor of the at least one computing device is used to execute the instructions stored in the memory of the at least one computing device, so that the computing device cluster implements the computing method provided in the foregoing first aspect or any possible implementation manner of the first aspect.

[0045] In an eighth aspect, the present application provides a computer-readable storage medium, which is used to store at least one piece of program code. When the at least one piece of program code is executed by a computing device cluster, the computing device cluster implements the computing method provided in the foregoing first aspect or any possible implementation manner of the first aspect.

[0046] In a ninth aspect, the present application provides a computer program product. When the computer program product runs on a computing device cluster, the computing device cluster implements the computing method provided in the foregoing first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a schematic diagram of an application scenario of a related technology and the present application;

[0048] Figure 2 is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0049] Figure 3 It is a schematic structural diagram of an acceleration device provided by an embodiment of the present application;

[0050] Figure 4 It is a schematic structural diagram of a processing unit provided by an embodiment of the present application;

[0051] Figure 5 It is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application;

[0052] Figure 6 It is a schematic diagram of the runtime architecture of a computing system provided by an embodiment of the present application;

[0053] Figure 7 It is a flowchart of a computing method provided by an embodiment of the present application;

[0054] Figure 8 It is a schematic diagram of a computing method provided by an embodiment of the present application;

[0055] Figure 9 It is a schematic diagram of another computing method provided by an embodiment of the present application;

[0056] Figure 10 It is a schematic diagram of yet another computing method provided by an embodiment of the present application;

[0057] Figure 11 It is a schematic structural diagram of a computing device provided by an embodiment of the present application;

[0058] Figure 12 It is a schematic structural diagram of another computing device provided by an embodiment of the present application. Detailed implementation manners

[0059] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings. It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the relevant data such as AI models and computing tasks involved in the present application are obtained under full authorization.

[0060] For the convenience of understanding, the following first explains the key terms and key concepts involved in the present application.

[0061] An artificial intelligence (AI) model is a type of mathematical algorithm model that uses machine learning concepts to solve practical problems. Generally, an AI model includes a large number of parameters and calculation formulas (or calculation rules).

[0062] An operator (OP) refers to a computing unit or computing function that runs on a computing device. In the field of deep learning, a neural network layer and even the entire model are composed of operators, and these operators correspond to the computing logic in the neural network layer. For example, a convolution layer is an operator; the process of summing weights in a fully-connected layer (FClayer) is an operator.

[0063] An acceleration device, also known as an accelerator, acceleration card, or acceleration chip, is a type of specialized hardware accelerator or computer system designed to accelerate AI applications, especially neural networks, machine vision, and machine learning, etc. For example, it provides computing power for performing calculations on matrices, vectors, etc. to accelerate the calculation of an AI model. Schematically, the acceleration device is, for example, a graphics processing unit (GPU), a neural network processing unit (XPU), an intelligent processing unit (IPU), a tensor processing unit (TPU), a domain specific architecture (DSA) chip, etc., and the present application is not limited thereto.

[0064] Recomputation means that during the training process of an AI model, when performing the forward computation task, the intermediate execution results (or called activation values) of each layer are not stored. Instead, before performing the backward computation task, the forward computation task is executed again to obtain the intermediate execution results of each layer for the backward computation task. This method can reduce the memory occupancy of the end-to-end training branches. For example, when using an acceleration device to train an AI model, it can effectively reduce the memory occupancy of the acceleration device during the model training process. In some embodiments, recomputation is combined with checkpoint (a way to save key processes / results, generally used for fault recovery and resuming training from a breakpoint) to further save the computational amount of the activation values of some layers in the AI model. Schematically, some layers of the AI model are used as checkpoints. When the acceleration device performs the forward computation task, the intermediate execution results of the checkpoints are transferred to the host memory. Before the acceleration device performs the backward computation task, the intermediate execution results of the checkpoints are transferred from the host memory, and based on these intermediate execution results, the forward computation task is executed again to obtain the intermediate execution results of each layer in the AI model, that is, to obtain the execution result of the forward computation task. For example, if the AI model includes five layers and the first layer and the third layer are used as checkpoints, when the acceleration device performs the forward computation task, the output results of the first layer and the third layer (that is, the intermediate execution results of the forward computation task) are transferred to the host memory. Before the acceleration device performs the backward computation, the forward computation task is executed based on the output results of the first layer and the third layer to obtain the output results of each layer in the AI model, that is, to obtain the execution result of the forward computation task, and then the backward computation task is executed based on the execution result of the forward computation task.

[0065] Near data processing (NDP) means arranging the computational tasks of data on the computing device closest to the storage location of the data. Among them, the distance between the storage location of the data and the computing device can be defined as the sum of the time for the data to be transferred to the computing device and the time for the computing device to process the data. In other words, the computing device closest to the storage location of a certain data can be the computing device that can obtain and process the data fastest among multiple computing devices. Through this near data processing method, unnecessary data copying can be significantly reduced, and the computational processing efficiency of the data can be improved.

[0066] The application scenarios and implementation environments of the present application are introduced below.

[0067] The present application is applied to the scenario of using recomputation during the training process of an AI model to reduce the memory occupancy of the acceleration device. Among them, the type of the AI model and the type of the acceleration device are not limited in the present application. The following refers to Figure 1 to introduce the application scenario of the present application.

[0068] Figure 1 It is a schematic diagram of the application scenarios of a related technology and the present application.

[0069] As Figure 1 shown in Fig. (a), in the related technology, the training process of the AI model is realized through multiple computing tasks, such as CompTask-0, CompTask-1, etc. Taking CompTask-0 as the forward computing task and CompTask-1 as the backward computing task as an example, when the acceleration device executes CompTask-0, it stores the intermediate execution result of this task into the target memory, or offloads it to the target memory. Among them, the intermediate execution result is, for example, the activation value of the checkpoint, and the target memory is, for example, the host memory. Then, before the acceleration device executes CompTask-1, it obtains the previously stored intermediate execution result from the target memory and executes the recomputation task based on the intermediate execution result, that is, re-executes the forward computing task to obtain the execution result of the forward computing task (that is, the activation value of each layer in the AI model, or the execution result of the recomputation task). After that, the acceleration device executes CompTask-1 based on the execution result of the forward computing task. It can be seen that in the related technology, there is redundant computing in the acceleration device and the computing overhead is relatively large.

[0070] As Figure 1 shown in Fig. (b), the present application provides an implementation method of adopting recomputation during the training process of the AI model, which can effectively save the computing resources of the acceleration device. In the present application, when the acceleration device executes CompTask-0, it stores the intermediate execution result of this task into the processing unit with computing and storage functions. The processing unit executes the recomputation task based on the intermediate execution result of the forward computing task, that is, re-executes the forward computing task to obtain the execution result of the forward computing task (that is, the activation value of each layer in the AI model, or the execution result of the recomputation task). Then, before the acceleration device executes CompTask-1, it obtains the execution result of the forward computing task from the processing unit and executes CompTask-1. It can be seen that the solution provided by the present application can effectively save the computing resources of the acceleration device, reduce the computing overhead, and moreover, since the data volume of the execution result of the forward computing task is smaller than the data volume of the intermediate execution result, this solution can effectively reduce the cross-device data transfer volume.

[0071] Next, refer to Figure 2 to introduce the implementation environment of the present application. Figure 2 It is a schematic diagram of an implementation environment provided by an embodiment of the present application. As Figure 2As shown, the implementation environment includes a computing system, which includes a host 100, an acceleration device 200, and a processing unit 300. The host 100, the acceleration device 200, and the processing unit 300 are directly or indirectly connected through a wireless network or a wired network.

[0072] The host 100 refers to a device for running an AI model and can provide AI services for users. In the embodiment of this application, the host 100 can control the acceleration device 200 and the processing unit 300 to execute the computing tasks involved in the AI model training process. For example, in response to a training request for an AI model, the host 100 sends a forward computing task and a backward computing task to the acceleration device 200, and sends a recomputation task to the processing unit 300, etc. This process can also be understood as loading the computing tasks involved in the AI model training process into the acceleration device 200 and the processing unit 300 for execution. In addition, the number of hosts 100 can be one or more, and this application does not limit this.

[0073] The acceleration device 200 is used to provide computing power for the training process of the AI model to accelerate the training process of the AI model. For example, the acceleration device 200 is a GPU, XPU, IPU, TPU, DSA chip, etc., and this application is not limited thereto. Schematically, the acceleration device 200 receives the computing tasks of the AI model (such as forward computing tasks, backward computing tasks, etc.) sent by the host 100, calls the corresponding operators, processes the input data of the operator, obtains the output data, that is, obtains the execution result of the computing task, and returns the execution result to the host 100. Among them, when the acceleration device 200 executes the forward computing task, it stores the intermediate execution result of the forward computing task in the processing unit 300, and before executing the backward computing task, it obtains the execution result of the forward computing task from the processing unit 300, and based on the execution result of the forward computing task, executes the backward computing task and returns the execution result of the backward computing task to the host 100. In addition, the number of acceleration devices 200 can be one or more, and this application does not limit this.

[0074] The processing unit 300 has computing and storage functions, and is used to provide storage space and computing power for the recomputation tasks involved in the AI model training process, so as to save the computing power resources of the acceleration device 200. Schematically, before the acceleration device 200 executes the post-forward computing task, the processing unit 300 receives the recomputation task of the AI model sent by the host 100, calls the corresponding operator, processes the input data of the operator to obtain the output data, that is, obtains the execution result of the recomputation task, and returns the execution result to the acceleration device 200 so that the acceleration device 200 can execute the post-forward computing task. In the embodiments of the present application, there are multiple implementation manners of the processing unit 300, which will be introduced in detail in the subsequent content and will not be elaborated here. Moreover, the number of the processing units 300 can be one or more, and the present application does not make any limitation thereto.

[0075] The above-mentioned host 100, acceleration device 200, and processing unit 300 can be integrated in one computing device or can be separately arranged, and the present application does not make any limitation thereto. Schematically, taking the case where the host 100, acceleration device 200, and processing unit 300 can be integrated in one computing device as an example, the host 100, acceleration device 200, and processing unit 300 are communicatively connected through a peripheral component interconnect express (PCIe) link, and data interaction is performed among the host 100, acceleration device 200, and processing unit 300 through the PCIe link. Alternatively, data interaction among the host 100, acceleration device 200, and processing unit 300 is achieved through NVIDIA Link (NvLink), Compute Express Link (CXL), Universal Chiplet Interconnect Express (UCIe), Huawei Cache Coherent System (HCCS), Cache Coherent Interconnect for Accelerators (CCIX), etc., and the present application is not limited thereto.

[0076] For example, the computing device can be an independent physical server, or a server cluster or distributed file system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. Taking the computing device as a cloud server as an example, the computing device can also be referred to as a cloud platform (i.e., the abbreviation of the cloud computing platform), which refers to services based on hardware resources and software resources, providing computing, network, and storage capabilities. Through the network "cloud", huge data computing is processed remotely and then returned to the user, with characteristics such as large scale, distribution, virtualization, high availability, scalability, on-demand service, and security. The cloud platform can achieve the rapid distribution and release of configurable computing resources with a relatively small management cost or a relatively low interaction complexity between the user and the service provider.

[0077] In addition, the networks involved above include but are not limited to any combination of data center networks, storage area networks (SAN), local area networks (LAN), metropolitan area networks (MAN), wide area networks (WAN), mobile, wired or wireless networks, private networks or virtual private networks. In some implementation manners, technologies and / or formats including hyper text markup language (HTML), extensible markup language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as secure sockets layer (SSL), transport layer security (TLS), virtual private network (VPN), and internet protocol security (IPsec) can be used to encrypt all or part of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0078] The structure of the acceleration device 200 in the above computing system will be introduced below. Refer to Figure 3 , Figure 3This is a schematic structural diagram of an acceleration device provided by an embodiment of the present application. As Figure 3 shown, the acceleration device 200 includes a communication interface 201, at least one AI processing core 202, a processor 203, a memory 204, and a bus 205. Among them, the communication interface 201, at least one AI processing core 202, the processor 203, and the memory 204 are communicatively connected to each other through the bus 205.

[0079] The communication interface 201 is used to provide program instructions and / or data for at least one AI processing core 202. The communication interface 201 includes a PCIe communication interface, other general peripheral interfaces, etc., which are not limited in this application. For example, when the acceleration device 200 is used as an acceleration card of the host 100, data exchange is achieved with the host 100 through the PCIe communication interface. Another example is that the acceleration device 200 realizes communication between the acceleration device 200 and other devices or communication networks through the peripheral interface.

[0080] The AI processing core 202 is used to implement the functions of the acceleration device as shown above Figure 2 shown. Schematically, the AI processing core adopts the Da Vinci architecture, which realizes high throughput, large computing power, and low power consumption, and is suitable for processing common calculations necessary for neural networks in deep learning, such as matrix multiplication, etc., which is not limited in this application.

[0081] The processor 203 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or an integrated circuit used to control the execution of the program of the solution of the present application. The processor 203 can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The number of the processors 203 can be one or multiple. Taking the multi-core processor as an example, multiple cores can be divided into a control CPU dedicated to controlling the overall operation of the acceleration device 200 and an AI CPU dedicated to undertaking non-matrix complex calculations according to functions. The number of CPU cores occupied by the two types of tasks can be dynamically allocated by software according to the actual operation situation of the system, which is not limited in this application.

[0082] The memory 204 can be a high bandwidth memory (HBM), a double data rate (DDR) memory, a static random-access memory (SRAM), a read-only memory (ROM), or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0083] The bus 205 can include paths for transmitting information between various components of the acceleration device 200 (e.g., the communication interface 201, at least one AI processing core 202, the processor 203, the memory 204).

[0084] It should be noted that the above Figure 3 shows only a hardware structure diagram of a chip that can be configured as the above acceleration device 200 provided by this application. In some embodiments, the chip may further include other components to achieve more functions. For example, the chip further includes a task scheduler (TS) for efficiently allocating and scheduling computing tasks on the AI processing cores, etc. This application is not limited thereto.

[0085] Next, the structure of the processing unit 300 in the above computing system will be introduced. Refer to Figure 4 , Figure 4 is a schematic structural diagram of a processing unit provided by an embodiment of this application. As shown in Figure 4As shown, the processing unit 300 includes a computing subunit 301 and a storage subunit 302. Among them, the computing subunit 301 is used to implement the computing function of the processing unit 300. Schematically, the computing subunit 301 includes a processor, such as a CPU, an ASIC, or an integrated circuit for controlling the execution of the program of the solution of the present application. The processor can be a single-core processor or a multi-core processor. The number of processors can be one or multiple, and the present application is not limited thereto. The storage subunit 302 is used to implement the storage function of the processing unit 300. Schematically, the storage subunit 302 includes a memory. The memory can be a ROM or other types of static storage devices that can store static information and instructions, a RAM, or other types of dynamic storage devices that can store information and instructions, or an EEPROM, a CD-ROM, or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0086] In the embodiment of the present application, in the processing unit 300, the distance between the computing subunit 301 and the storage subunit 302 meets the conditions. For example, the computing subunit 301 and the storage subunit 302 are respectively a processor and a memory with a qualified physical deployment distance. Among them, the condition content can be set according to actual needs, and the present application does not limit this. It should be understood that in the present application, the distance between the computing subunit 301 and the storage subunit 302 meeting the conditions can also be understood as that the processing unit 300 has the ability of near-data processing. In this way, by using the processing unit 300 to execute the recomputation task, unnecessary data copying can be significantly reduced, and the computing and processing efficiency of the data can be improved.

[0087] In some embodiments, the computing subunit 301 and the storage subunit 302 are integrated on the same chip. That is, through chip packaging and board assembly, etc., the computing subunit 301 and the storage subunit 302 are integrated to achieve processing near memory (PNM), increase the memory access bandwidth, reduce data migration, and improve the overall computing efficiency.

[0088] In other embodiments, the computing subunit 301 and the storage subunit 302 are integrated in the same die of the chip. That is, during the chip manufacturing process, the storage and computing are integrated in the same die, enabling the memory itself to have a certain computing ability and achieving processing in memory (PIM).

[0089] In some other embodiments, when the computing system 200 is a distributed computing system, the processing unit 300 is a computing node in the distributed computing system. For example, in a distributed computing system, the host, DPU, and XPU are interconnected peer-to-peer and are all computing nodes in the distributed computing system. The memory uses a unified addressing format, or in other words, memory pooling. The processing unit 300 can be a host with sufficient computing power resources, or an XPU in an idle state, etc. The present application is not limited thereto.

[0090] In addition, the present application also provides a computing device cluster, which includes at least one computing device. Each computing device includes a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster realizes the functions of the foregoing computing system 200. Refer to Figure 5 , Figure 5 is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application. As Figure 5 shown, the computing device cluster includes a plurality of computing devices 500. Each computing device 500 includes a memory 501, a processor 502, a communication interface 503, and a bus 504. The memory 501, the processor 502, and the communication interface 503 are communicatively connected through the bus 504. The same instructions for executing the computing method provided by the present application can be stored in the memory 501 of the plurality of computing devices 500 in the computing device cluster. In some embodiments, the memory 501 of the plurality of computing devices 500 in the computing device cluster can also store partial instructions for executing the computing method provided by the present application respectively. In other words, the combination of the plurality of computing devices 500 is used to jointly execute the computing method provided by the present application. In some embodiments, the plurality of computing devices 500 in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. The present application is not limited thereto.

[0091] Next, the computing method provided by the present application will be introduced. Based on the foregoing introduction, it can be known that the present application provides an implementation method of using recomputation during the AI model training process, which can effectively save the computing resources of the acceleration device. Next, first refer to Figure 6 , and take the runtime architecture of the foregoing computing system 200 as an example to briefly describe the computing method provided by the present application. Figure 6 is a schematic diagram of the runtime architecture of a computing system provided by an embodiment of the present application. As Figure 6As shown in the figure, in the runtime environment, the upper layer 601 includes an application layer (user app) and an AI model framework layer (AI model framework). The application layer distributes operators, compiles graphs through the AI model framework layer, and then sends them down to the runtime layer 602 (implemented through the host). The runtime layer 602 includes a backend interface unit (backend interface), an operator library, and a driver layer. Among them, the backend interface unit is used to adapt the backend device interface for the operators sent down by the upper layer 601. The operator library is used to identify the attributes of the called operators. For example, it identifies whether the operator involves recomputation and cross-device data transmission, etc. If the operator involves recomputation and cross-device data transmission, corresponding task information is generated for the operator. This task information is used to indicate the task identifier and data identifier of the operator, and is also used to identify that the operator involves recomputation and cross-device data transmission. The driver layer is used to send the corresponding computation tasks to the acceleration device according to the called operators, and send the task information of the recomputation tasks to the acceleration device and the processing unit respectively, so that both the acceleration device and the processing unit can achieve data transmission through this task information. In this process, the acceleration device executes the forward computation task and stores the intermediate execution result of the forward computation task in the processing unit. Before the acceleration device executes the backward computation task, the driver layer sends the recomputation task to the processing unit. The processing unit executes the recomputation task based on the stored intermediate execution result of the forward computation task to obtain the execution result of the forward computation task. Thus, when the acceleration device executes the backward computation task, it can obtain the execution result of the forward computation task from the processing unit, that is, obtain the result after recomputation on demand. It can be seen that the software and hardware combined computing architecture provided by this application can effectively save the computing power resources of the acceleration device.

[0092] Next, taking the interaction between the host, acceleration device, and processing unit in the computing system as an example, the computing method provided by this application will be introduced. Refer to Figure 7 , Figure 7 which is a flowchart of a computing method provided by an embodiment of this application. As Figure 7 shown, this computing method includes the following steps 701 to 708.

[0093] 701. In response to a training request for an AI model, when the training request involves a recomputation task and cross-device data transmission, the host generates task information for the recomputation task, and this task information indicates the task identifier and data identifier of the recomputation task.

[0094] In the embodiments of the present application, before the recalculation task instruction is executed, the forward calculation task is performed to obtain the execution result of the forward calculation task. For details, reference can be made to the foregoing introduction and will not be elaborated herein. Cross-device data transmission means that when the acceleration device executes the forward calculation task of the AI model, the intermediate execution result of the forward calculation task is stored on other devices except the acceleration device, and when the backward calculation task is executed, the execution result of the forward calculation task is obtained from other devices. That is to say, this process involves cross-device data transmission between the acceleration device and other devices.

[0095] Schematically, in response to a training request for an AI model, the host performs operator compilation based on the model framework of the AI model, that is, generates operators of the AI model. For example, operator generation is divided into three processes: input tensor description, weight data conversion, and output tensor description. Among them, in the input tensor description, information such as the input dimension and memory size of each operator is calculated, and the form of the operator input data is defined; in the weight data conversion, data format, shape conversion, data compression, etc. are performed on the weight parameters used by the operator; in the output tensor description, information such as the output dimension and memory size of the operator is calculated. After the host performs operator compilation based on the model framework of the AI model, it adapts the operator to the backend device interface (that is, adapts the upper-layer operator to the underlying hardware device), and identifies the attributes of the operator, that is, identifies whether the operator involves a recalculation task and cross-device data transmission. Among them, when it is recognized that the operator involves a recalculation task and cross-device data transmission, task information of the recalculation task is generated. For example, the task information is an identifier, including the task identifier and data identifier of the recalculation task. The present application does not limit the form of the task information. In other words, the task information of the recalculation task is used to identify the operator corresponding to the recalculation task and the data processed by the operator. In this process, the host has the function of identifying whether the operator involves a recalculation task and cross-device data transmission. In this way, automatic generation of task information can be achieved without the upper-layer user being aware. Of course, the upper-layer user can also carry information in the training request to indicate that the current training involves a recalculation task and cross-device data transmission. The present application does not limit this.

[0096] 702. The host sends the task information of the recalculation task to the processing unit.

[0097] In the embodiments of the present application, the process of the host sending the task information to the processing unit can also be understood as the process of registering the recalculation task in the processing unit. Schematically, the host determines the processing unit for executing the recalculation task from the computing system, sends the task information to the processing unit, and creates a task list in the processing unit so that the processing unit can execute the corresponding task based on the task list.

[0098] In some embodiments, the processing unit provided in the present application includes a computing subunit and a storage subunit. In this step, the host sends task information of the recomputation task to the processing unit, including: the host sends task information of the recomputation task to the computing subunit. Additionally, based on the foregoing introduction to the processing unit, there are multiple implementation manners for the processing unit. For example, the computing subunit and the storage subunit are a processor and a memory whose physical deployment distances meet the conditions, or the computing subunit and the storage subunit are integrated on the same chip, or the computing subunit and the storage subunit are integrated in the same die of the chip, or in the case where the computing system is a distributed computing system, the processing unit is a computing node in the distributed computing system, etc. The present application does not limit this.

[0099] In some embodiments, the host determines the processing unit for executing the recomputation task based on the distance between the acceleration device and each processing unit in the computing system. For example, the processing unit closest to the acceleration device is used as the processing unit for executing the recomputation task. In this way, the data transmission path can be shortened and the data transmission efficiency can be improved. Among them, the distance between the acceleration device and the processing unit can be the physical deployment distance. For example, the computing system includes CPU0 and CPU1. CPU0 and CPU1 are respectively connected to 8 acceleration devices, and CPU0 and CPU1 are respectively connected to 8 memory modules. The 8 memory modules connected to CPU0 form processing unit 0, and the 8 memory modules connected to CPU1 form processing unit 1. Taking the acceleration device 0 connected to CPU0 as an example for executing the calculation task of the AI model, the host determines that processing unit 0 is the processing unit for executing the recomputation task based on the distances between acceleration device 0 and processing unit 0 and processing unit 1, and sends the task information of the recomputation task to this processing unit. The distance between the acceleration device and the processing unit can also be the distance in the address space. For example, in the computing system, CPU0 is connected to 8 acceleration devices, and these 8 acceleration devices are connected to a distributed memory pool with unified addressing. Taking the acceleration device 0 as an example for executing the calculation task of the AI model, the host determines that the address space xxx1 - xxx2 is the closest to the acceleration device 0 based on the address space range of the memory pool, takes the address space xxx1 - xxx2 as the storage subunit of the processing unit, and takes CPU0 as the computing subunit of the processing unit. It should be understood that the above is only an example and does not constitute a limitation to the present application. In some embodiments, the processing unit for executing the recomputation task can be determined according to the service requirements.

[0100] 703. The host sends the forward calculation task of the AI model to the acceleration device.

[0101] In an embodiment of the present application, the host can synchronously execute step 702 and step 703. That is, when the host sends the task information of the recomputation task to the processing unit, it synchronously sends the forward calculation task of the AI model to the acceleration device. Of course, the host can also execute step 702 and step 703 sequentially, or, after executing step 701, the host first executes step 703 and then executes step 702. The present application does not make any limitation in this regard.

[0102] In some embodiments, the forward calculation task sent by the host to the acceleration device carries the task information of the recomputation task. For example, taking the task information as the identifier, the host adds the task information to the handle corresponding to the forward calculation task. In other embodiments, the host sequentially sends the task information of the forward calculation task and the recomputation task to the acceleration device. The present application does not make any limitation in this regard. It should be understood that since cross-device data transmission is involved between the acceleration device and the processing unit in the present application, the host sends the task information of the recomputation task to the acceleration device and the processing unit respectively, which can associate the calculation tasks executed by both the acceleration device and the processing unit, providing technical support for data transmission between the two parties.

[0103] 704. The acceleration device executes the forward calculation task and stores the intermediate execution result of the forward calculation task in the processing unit.

[0104] In an embodiment of the present application, the number of operators corresponding to the forward calculation task can be one or multiple. The present application does not make any limitation in this regard. Schematically, the acceleration device can store the intermediate execution result of the forward calculation task in the processing unit after the forward calculation task is executed, or can continuously store the intermediate execution result of the forward calculation task in the processing unit during the execution of the forward calculation task. The present application does not make any limitation in this regard.

[0105] In addition, based on the foregoing introduction, it can be known that the processing unit includes a calculation subunit and a storage subunit. In this step, the acceleration device executes the forward calculation task and stores the intermediate execution result of the forward calculation task in the storage subunit of the processing unit.

[0106] After the above steps 701 to 704, in response to a training request for the AI model, on the one hand, the host sends the task information of the recomputation task to the processing unit for executing the recomputation task, and on the other hand, the host sends the forward calculation task of the AI model to the acceleration device, so that the acceleration device executes the forward calculation task and stores the intermediate execution result of the forward calculation task in the processing unit. The execution processes of the recomputation task and the backward calculation task will be introduced below through steps 705 to 708.

[0107] 705. The host sends the recomputation task of the AI model to the processing unit.

[0108] Based on the foregoing step 702, it can be seen that the host has sent the task information of the recomputation task to the processing unit, that is, the host has registered the recomputation task in the processing unit. In this step 705, the host sends an execution instruction of the recomputation task to the processing unit to trigger the processing unit to execute the recomputation task.

[0109] 706. The processing unit executes the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task.

[0110] Based on the foregoing introduction, it can be seen that the processing unit includes a computing subunit and a storage subunit. In this step, through the computing subunit, based on the intermediate execution result of the forward computation task in the storage subunit, the recomputation task is executed to obtain the execution result of the forward computation task, and the execution result of the forward computation task is stored in the storage subunit.

[0111] In addition, based on the foregoing step 702, it can be seen that the host has sent the task information of the recomputation task to the processing unit, that is, the processing unit has received the task information of the recomputation task. Based on this, in this step 706, the processing unit executes the recomputation task based on the task information and the intermediate execution result of the forward computation task in the processing unit. Schematically, the computing subunit obtains the intermediate execution result of the forward computation task from the storage subunit based on the task information to execute the recomputation task.

[0112] 707. The host sends a backward computation task of the AI model to the acceleration device.

[0113] In the embodiment of the present application, step 706 and step 707 can be executed synchronously or sequentially, and the present application does not make any limitation in this regard.

[0114] 708. The acceleration device obtains the execution result of the forward computation task from the processing unit and executes the backward computation task based on the execution result of the forward computation task.

[0115] In the embodiment of the present application, the number of operators corresponding to the backward computation task can be one or multiple, and the present application does not make any limitation in this regard. Schematically, the acceleration device can obtain the execution result of the forward computation task from the processing unit before executing the backward computation task, or can continuously obtain the execution result of the forward computation task from the processing unit based on the data required for the backward computation task during the execution of the backward computation task, and the present application does not make any limitation in this regard.

[0116] In addition, based on the foregoing introduction, it can be seen that the processing unit includes a computing subunit and a storage subunit. In this step, the acceleration device obtains the execution result of the forward computation task from the storage subunit of the processing unit and executes the backward computation task based on the execution result of the forward computation task.

[0117] After the above steps 705 to 708, the host sends a recomputation task to the processing unit to trigger the processing unit to execute the recomputation task. That is, the recomputation task is offloaded to the processing unit for execution. Moreover, based on the foregoing introduction to the processing unit, the processing unit has the near-data processing ability. Therefore, the above process is to utilize the near-data processing (NDP) ability of the processing unit to execute the recomputation task in advance, so that the acceleration device can obtain the data required for the backward computation task from the processing unit as needed. This method can also be understood as a speculative execution method, which can utilize the NDP ability of the computing hardware in the computing system without awareness, saving the computing power resources of the acceleration device, reducing the data transmission volume between the host and the acceleration device, and thus improving the effective computing power utilization and cluster system efficiency of the computing system.

[0118] In summary, in the computing method provided in this application, the host responds to a training request for an AI model and triggers the training process for the AI model. In this process, the acceleration device executes the forward computation task and stores the intermediate execution result of the forward computation task in the processing unit. The processing unit executes the recomputation task to obtain the execution result of the forward computation task, so that the acceleration device obtains the execution result of the forward computation task from the processing unit to execute the backward computation task. In this way, the computing power resources of the acceleration device are effectively saved, and the computing overhead is reduced. Moreover, since the data volume of the execution result of the forward computation task is smaller than the data volume of the intermediate execution result, this method can effectively reduce the cross-device data transmission volume.

[0119] Next, refer to Figures 8 to 10 , taking different implementation manners of the processing unit as examples, the above Figure 7 shown computing method will be illustrated by examples.

[0120] Figure 8 is a schematic diagram of a computing method provided by an embodiment of this application. As Figure 8 shown, this method is applied to a computing system, which includes a host, an acceleration device, and a processing unit. Among them, the processing unit includes a computing subunit and a storage subunit, and the computing subunit and the storage subunit are integrated on the same chip A. For example, the processing unit is a high-performance CPU, and such a CPU not only has computing ability but also has the ability of multi-level memory and multi-memory channels. This application is not limited thereto.

[0121] Schematically, the computing method includes the following steps:

[0122] Step A1: In response to a training request for an AI model, when the training request involves recomputation tasks and cross-device data transmission, the host generates task information for the recomputation tasks, sends the task information for the recomputation tasks to the processing unit, and sends the forward computation task of the AI model to the acceleration device.

[0123] Step A2: The acceleration device executes the forward computation task and stores the intermediate execution result of the forward computation task in the storage subunit of the processing unit.

[0124] Step A3: After receiving the recomputation task sent by the host, the computation subunit of the processing unit obtains the intermediate execution result of the forward computation task from the storage subunit based on the task information, executes the recomputation task, and obtains the execution result of the forward computation task.

[0125] Step A4: After receiving the backward computation task sent by the host, the acceleration device obtains the execution result of the forward computation task from the storage subunit and executes the backward computation task.

[0126] In this way, since the processing unit is integrated on a chip, the NDP capability of the processing unit is utilized, and in-memory processing for recomputation tasks is achieved. This not only saves the computing power resources of the acceleration device, but also increases the memory access bandwidth, reduces data movement, and improves the overall computing efficiency.

[0127] Figure 9 It is a schematic diagram of another computing method provided by an embodiment of the present application. As Figure 9 shown, this method is applied to a computing system, which includes a host, an acceleration device, and a processing unit. Among them, the computing system is a distributed computing system. The processing unit includes a computation subunit and a storage subunit, and the processing unit is a computing node in the distributed computing system. For example, in the distributed computing system, the host, DPU, and XPU are peer-to-peer interconnected and are all computing nodes in the distributed computing system. The memory uses a unified addressing format. The processing unit can be a host CPU with sufficient computing power resources, or an idle XPU, DPU, etc. The present application is not limited thereto.

[0128] Schematically, the computing method includes the following steps:

[0129] Step B1: In response to a training request for an AI model, when the training request involves recomputation tasks and cross-device data transmission, the host generates task information for the recomputation tasks, sends the task information for the recomputation tasks to the processing unit, and sends the forward computation task of the AI model to the acceleration device. Among them, the processing unit is a computing node in the computing system with sufficient computing power resources or in an idle state.

[0130] Step B2: The acceleration device executes the forward calculation task and stores the intermediate execution result of the forward calculation task in the storage subunit of the processing unit.

[0131] Step B3: After receiving the recomputation task sent by the host, the calculation subunit of the processing unit obtains the intermediate execution result of the forward calculation task from the storage subunit based on the task information, executes the recomputation task, and obtains the execution result of the forward calculation task.

[0132] Step B4: After receiving the backward calculation task sent by the host, the acceleration device obtains the execution result of the forward calculation task from the storage subunit and executes the backward calculation task.

[0133] In this way, since the processing unit is a computing node in the computing system, the NDP capability of the processing unit is utilized, and in-memory processing for the recomputation task is achieved. This not only saves the computing power resources of the acceleration device, but also increases the memory access bandwidth, reduces data movement, and improves the overall computing efficiency.

[0134] Figure 10 It is a schematic diagram of another calculation method provided by an embodiment of the present application. As Figure 10 shown, this method is applied to a computing system, which includes a host, an acceleration device, and a processing unit. Among them, the processing unit includes a calculation subunit and a storage subunit. The calculation subunit and the storage subunit are integrated in a die of a chip. For example, the processing unit is a memory with in-memory processing (PIM) capability. The present application is not limited to this.

[0135] Schematically, the calculation method includes the following steps:

[0136] Step C1: In response to a training request for an AI model, when the training request involves a recomputation task and cross-device data transmission, the host generates the task information of the recomputation task, sends the task information of the recomputation task to the processing unit, and sends the forward calculation task of the AI model to the acceleration device.

[0137] Step C2: The acceleration device executes the forward calculation task and stores the intermediate execution result of the forward calculation task in the storage subunit of the processing unit.

[0138] Step C3: After receiving the recomputation task sent by the host, the calculation subunit of the processing unit determines the intermediate execution result of the forward calculation task in the storage subunit based on the task information, executes the recomputation task, and obtains the execution result of the forward calculation task.

[0139] Step C4: After receiving the backward calculation task sent by the host, the acceleration device obtains the execution result of the forward calculation task from the storage subunit and executes the backward calculation task.

[0140] In this way, since the processing units are integrated in the same die of a chip, the NDP capabilities of the processing units are utilized to achieve in-memory processing for recomputation tasks, which can not only save the computing power resources of the acceleration device, but also increase the memory access bandwidth, reduce data movement, and improve the overall computing efficiency.

[0141] In addition, referring to Figure 11 , the present application also provides a computing device configured in the processing unit of a computing system. This computing device is used to implement some or all of the functions of the aforementioned processing unit. Figure 11 FIG. is a schematic structural diagram of a computing device provided by an embodiment of the present application. As Figure 11 shown, the device includes a receiving module 1101 and an execution module 1102.

[0142] The receiving module 1101 is used to receive the recomputation task of the AI model sent by the host. The recomputation task instructs to execute the forward computation task to obtain the execution result of the forward computation task; the host is used to send the forward computation task and the backward computation task of the AI model to the acceleration device in response to a training request for the AI model.

[0143] The execution module 1102 is used to execute the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task, so that the acceleration device executes the backward computation task based on the execution result of the forward computation task. The intermediate execution result of the forward computation task is obtained by the acceleration device executing the forward computation task and stored by the acceleration device in the processing unit.

[0144] In some embodiments, the processing unit includes a computing subunit and a storage subunit. Among them, the execution module 1102 is configured in the computing subunit. The execution module 1102 is used to execute the recomputation task based on the intermediate execution result of the forward computation task in the storage subunit to obtain the execution result of the forward computation task, and store the execution result of the forward computation task in the storage subunit.

[0145] In some embodiments, the receiving module 1101 is further used to receive the task information of the recomputation task sent by the host. The task information indicates the task identifier and data identifier of the recomputation task; the execution module 1102 is used to execute the recomputation task based on the task information and the intermediate execution result of the forward computation task in the processing unit.

[0146] In addition, referring to Figure 12 , the present application also provides another computing device configured in the acceleration device of a computing system. This computing device is used to implement some or all of the functions of the aforementioned acceleration device. Figure 12 FIG. is a schematic structural diagram of another computing device provided by an embodiment of the present application. As Figure 12As shown in the figure, the device includes a receiving module 1201 and an execution module 1202.

[0147] The receiving module 1201 is configured to receive the forward calculation task of the AI model sent by the host in response to the training request for the AI model.

[0148] The execution module 1202 is configured to execute the forward calculation task and store the intermediate execution result of the forward calculation task into the processing unit. The processing unit is configured to receive the recalculation task of the AI model sent by the host and execute the recalculation task based on the intermediate execution result of the forward calculation task in the processing unit to obtain the execution result of the forward calculation task.

[0149] The receiving module 1201 is further configured to receive the backward calculation task of the AI model sent by the host.

[0150] The execution module 1202 is further configured to obtain the execution result of the forward calculation task from the processing unit and execute the backward calculation task based on the execution result of the forward calculation task.

[0151] Through the above computing device, the host triggers the training process for the AI model in response to the training request for the AI model. In this process, the acceleration device executes the forward calculation task and stores the intermediate execution result of the forward calculation task into the processing unit. The processing unit executes the recalculation task to obtain the execution result of the forward calculation task, so that the acceleration device obtains the execution result of the forward calculation task from the processing unit to execute the backward calculation task. In this way, the computing power resources of the acceleration device are effectively saved and the computing overhead is reduced. Moreover, since the data volume of the execution result of the forward calculation task is smaller than that of the intermediate execution result, this method can effectively reduce the cross-device data transmission volume.

[0152] It should be noted that: when the computing device provided in the above embodiment performs calculations, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the computing device provided in the above embodiment and the embodiment of the computing method belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.

[0153] In this application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various examples, the first processing unit can be referred to as the second processing unit, and similarly, the second processing unit can be referred to as the first processing unit. The first processing unit and the second processing unit can both be processing units, and in some cases, they can be separate and different processing units.

[0154] In this application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units.

[0155] The above description is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art in the technical field disclosed by this application can easily think of various equivalent modifications or substitutions within the technical scope disclosed by this application, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0156] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of program structure information. This program structure information includes one or more program instructions. When the program instructions are loaded and executed on a computing device, the processes or functions in the embodiments of this application are generated in whole or in part.

[0157] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. This program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0158] As mentioned above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent substitutions on some of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A calculation method, characterized in that, Applied to a computing system, the system includes a host, an acceleration device, and a processing unit, and the method includes: In response to a training request for an artificial intelligence (AI) model, the host sends a forward calculation task of the AI model to the acceleration device; The acceleration device executes the forward calculation task and stores intermediate execution results of the forward calculation task in the processing unit; The host sends a recomputation task of the AI model to the processing unit, and the recomputation task instructs to execute the forward calculation task to obtain an execution result of the forward calculation task; Based on the intermediate execution results of the forward calculation task in the processing unit, the processing unit executes the recomputation task to obtain the execution result of the forward calculation task; The host sends a backward calculation task of the AI model to the acceleration device; The acceleration device obtains the execution result of the forward calculation task from the processing unit and executes the backward calculation task based on the execution result of the forward calculation task.

2. The method according to claim 1, characterized in that, The processing unit includes a computing subunit and a storage subunit; Based on the intermediate execution results of the forward calculation task in the processing unit, the processing unit executes the recomputation task to obtain the execution result of the forward calculation task, including: The computing subunit executes the recomputation task based on the intermediate execution results of the forward calculation task in the storage subunit, obtains the execution result of the forward calculation task, and stores the execution result of the forward calculation task in the storage subunit.

3. The method according to claim 2, wherein The computing subunit and the storage subunit are integrated on the same chip.

4. The method according to claim 3, characterized in that, The computing subunit and the storage subunit are integrated in the same die of the chip.

5. The method according to any one of claims 2 to 4, characterized in that The system is a distributed computing system, and the processing unit is a computing node in the distributed computing system.

6. The method according to any one of claims 1 to 5, characterized in that The method further includes: The host sends task information of the recomputation task to the processing unit, and the task information indicates a task identifier and a data identifier of the recomputation task; Based on the intermediate execution results of the forward calculation task in the processing unit, the processing unit executes the recomputation task, including: The processing unit executes the recomputation task based on the task information and the intermediate execution results of the forward calculation task in the processing unit.

7. The method according to claim 6, wherein The method further includes: When the training request involves a recomputation task and cross-device data transmission, the host generates the task information.

8. A calculation method, characterized in that, Applied to a processing unit in a computing system, the system further includes a host and an acceleration device, and the host is configured to send a forward calculation task and a backward calculation task of the AI model to the acceleration device in response to a training request for the AI model; the method includes: Receiving the recomputation task of the AI model sent by the host, where the recomputation task instructs to execute the forward calculation task to obtain an execution result of the forward calculation task; Execute the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task, so that the acceleration device executes the backward computation task based on the execution result of the forward computation task. The intermediate execution result of the forward computation task is obtained by the acceleration device executing the forward computation task and stored by the acceleration device in the processing unit.

9. The method according to claim 8, characterized in that, The processing unit includes a computing subunit and a storage subunit. Among them, executing the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task includes: Through the computing subunit, execute the recomputation task based on the intermediate execution result of the forward computation task in the storage subunit to obtain the execution result of the forward computation task, and store the execution result of the forward computation task in the storage subunit.

10. The method according to claim 9, characterized in that, The computing subunit and the storage subunit are integrated on the same chip.

11. The method according to claim 10, wherein The computing subunit and the storage subunit are integrated in the same die of the chip.

12. The method according to any one of claims 9 to 11, characterized in that The system is a distributed computing system, and the processing unit is a computing node in the distributed computing system.

13. The method according to any one of claims 8 to 12, characterized in that, The method further includes: Receive the task information of the recomputation task sent by the host, where the task information indicates the task identifier and data identifier of the recomputation task; Executing the recomputation task based on the intermediate execution result of the forward computation task in the processing unit includes: executing the recomputation task based on the task information and the intermediate execution result of the forward computation task in the processing unit.

14. A calculation method, characterized in that, Applied to an acceleration device in a computing system, the system further includes a host and a processing unit, and the method includes: Receive the forward computation task of the AI model sent by the host in response to a training request for the AI model; Execute the forward computation task and store the intermediate execution result of the forward computation task in the processing unit. The processing unit is used to receive the recomputation task of the AI model sent by the host, and execute the recomputation task based on the intermediate execution result of the forward computation task in the processing unit to obtain the execution result of the forward computation task; Receive the backward computation task of the AI model sent by the host; Obtain the execution result of the forward computation task from the processing unit, and execute the backward computation task based on the execution result of the forward computation task.

15. A computing system, characterized in that, The system includes a host, an acceleration device, and a processing unit, and the system is used to implement the computing method according to any one of claims 1 to 7.

16. A processing unit, characterized in that, The processing unit includes a computing subunit and a storage subunit, and the processing unit is used to implement the steps executed by the processing unit in the computing method according to any one of claims 1 to 7.

17. An acceleration device, characterized in that, The acceleration device includes a processor and a memory, and the processor is used to execute at least one program code stored in the memory, so that the acceleration device implements the steps executed by the acceleration device in the computing method according to any one of claims 1 to 7.

18. A cluster of computing devices, characterized in that, comprising at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster implements the computing method described in any one of claims 1 to 7.

19. A computer-readable storage medium, characterized in that, the computer-readable storage medium is used to store at least one segment of program code, and when the at least one segment of program code is executed by the computing device cluster, the computing device cluster implements the computing method described in any one of claims 1 to 7.

20. A computer program product, characterized in that, when the computer program product runs on the computing device cluster, the computing device cluster implements the computing method described in any one of claims 1 to 7.