A method and device for supporting multi-GPU Atomic instruction operations

By introducing Atomic instruction encoding and decoding modules within the GPU, it supports multiple types of Atomic instructions and custom data bit widths, solving the problem of limited Atomic instruction types and data bit widths supported by PCIe in the prior art, and improving the performance of GPU clusters.

CN119621289BActive Publication Date: 2025-06-06METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510157009.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Among the existing methods and devices that support multi-GPU Atomic instruction operation, PCIe supports few Atomic instruction types, only FetchAdd, Swap and CAS, and only three data bit widths: 32bit, 64bit, and 128bit, resulting in low efficiency in other application scenarios of data bit widths.

Method used

By introducing Atomic instruction encoding module and decoding module into the GPU, the instruction opcode and control information are encoded into PCIe TLP Prefix, and the address and data carried by the Atomic instructions are encoded into PCIe TLP, enabling the PCIe TLP Prefix function, supporting various types of Atomic instructions and custom data bit widths.

Benefits of technology

It realizes support for multiple types of Atomic instructions and customizes data bit width, improves the overall performance of the GPU cluster, and solves the problem of limited Atomic instructions type and data bit width in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621289B_ABST
    Figure CN119621289B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of GPU computing technology, and specifically to a method and device for supporting multi-GPU Atomic instruction operations. The method and device include: a user sends a task to a GPU according to demand, and the execution module of the GPU parses and executes the task and issues an instruction; an Atomic instruction encoding module determines whether it is an Atomic instruction according to the instruction opcode; if so, the instruction opcode control information is encoded into a PCIe TLP Prefix and transmitted to a destination GPU; if not, the instruction opcode is directly decoded by an Atomic instruction decoding module of the destination GPU; the instruction opcode is decoded from the PCIe TLP Prefix; if it is an Atomic instruction that needs to return the original data, the completion status and the original data will be returned; otherwise, the original data of the instruction address does not need to be read; a task completion packet is generated according to the execution result and returned to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of GPU computing technology, and in particular to a method and device for supporting multi-GPU Atomic instruction operations. Background Art

[0002] In the era of big data, computing power has become the driving force for the development of the digital economy. As the demand for data computing increases, GPUs have become an important part of computing infrastructure. GPU chips have a huge number of computing cores and a powerful instruction set. Their biggest advantage lies in parallel processing capabilities. In particular, through multi-GPU parallel execution, data processing speed can be significantly improved. They can be widely used in many fields such as data centers and artificial intelligence. In multi-GPU communication, Atomic instructions are often used to ensure the atomicity of operations performed between different GPUs, especially in scenarios involving shared memory or the need to update data synchronously. Atomic operations ensure that even if multiple GPUs operate on the same memory address at the same time, each operation can be executed correctly without interference from other operations. They are very useful in multi-GPU training and parallel computing, especially when it is necessary to accumulate gradients, update parameters, or implement certain synchronization mechanisms. For example, when training a deep neural network with multiple GPUs, each GPU may calculate a part of the gradient, and then these gradients need to be aggregated into a global gradient. Using Atomic instructions ensures that when updating the global gradient, the gradients calculated by different GPUs can be correctly accumulated without race conditions. PCIe (Peripheral Component Interconnect Express) is a universal serial expansion bus standard with the characteristics of high efficiency, flexibility, and scalability. It is widely used in connections between GPU chips.

[0003] However, in the existing methods and devices supporting multi-GPU Atomic instruction operations, the Atomic instructions supported by PCIe have the following problems: few types are supported, and only three types of Atomic instructions, namely FetchAdd (Fetch and Add), Swap (Unconditional Swap), and CAS (Compare and Swap), are supported. The implementation efficiency is low. The PCIe protocol only supports Atomic instructions with three data bit widths: 32 bits, 64 bits, and 128 bits. For application scenarios with other data bit widths, the implementation efficiency is low. Therefore, a method and device supporting multi-GPU Atomic instruction operations are provided. Summary of the invention

[0004] The purpose of the present invention is to provide a method and device supporting multi-GPU Atomic instruction operations, so as to solve the problem that the existing supported types are few, and only three types of Atomic instructions, namely FetchAdd (Fetch and Add), Swap (Unconditional Swap), and CAS (Compare and Swap), are supported in the above-mentioned background technology. The implementation efficiency is low, and the PCIe protocol only supports Atomic instructions with three data bit widths of 32 bits, 64 bits, and 128 bits. For application scenarios with other data bit widths, the implementation efficiency is low.

[0005] To achieve the above object, the present invention provides a method for supporting multi-GPU Atomic instruction operations, comprising the following steps:

[0006] S1. The user sends tasks to the GPU according to the requirements. The execution module inside the GPU parses and executes the tasks and issues various instructions.

[0007] S2, the Atomic instruction encoding module determines whether it is an Atomic instruction according to the instruction operation code;

[0008] S3. If it is an Atomic instruction, encode the instruction opcode control information into the PCIe TLP Prefix, encode the address and data carried by the Atomic instruction into the PCIe TLP, enable the PCIe TLP Prefix function, and transmit it to the destination GPU;

[0009] S4. If it is not an Atomic instruction, there is no need to enable the PCIe TLP Prefix function, and the Atomic instruction decoding module of the target GPU is directly decoded;

[0010] S5, the Atomic instruction decoding module of the target GPU decodes the instruction operation code from the PCIe TLP Prefix;

[0011] S6. All Atomic instructions are divided into two types: those that need to return original data and those that do not. If an Atomic instruction needs to return original data, it will return the completion status and original data.

[0012] S7. If it is an Atomic instruction that does not need to return the original data, there is no need to read the original data of the instruction address;

[0013] S8. If the execution module of the requesting GPU recognizes the return as Atomic, a task completion package is generated according to the execution result and returned to the user.

[0014] As a further improvement of the present technical solution, in S1, the user needs multiple GPUs to complete a task together, and sends the task to the GPU. After receiving the task, the device parses it and issues instructions according to the parsed task requirements.

[0015] As a further improvement of the present technical solution, in S2, the Atomic instruction and the supported Atomic types can be defined according to needs, including but not limited to logical operations, addition, subtraction, multiplication, self-increment, self-decrement, maximum value, minimum value, replacement, comparison operations, and support for custom data bit widths, including but not limited to 4bit, 8bit, 16bit, 32bit, 64bit, and 128bit data sizes.

[0016] As a further improvement of the technical solution, the S5 is used to obtain the address and data of the Atomic instruction from the PCIe TLP according to the instruction operation code, and transmit them to the local execution module at the same time.

[0017] As a further improvement of the technical solution, in S6, the specific steps of returning the completion status and original data are:

[0018] S3.1, the execution module first reads the original data at the corresponding position according to the instruction address;

[0019] S3.2, then write the calculated data to the instruction address according to the instruction operation code;

[0020] S3.3. Finally, return the original data and completion status.

[0021] As a further improvement of the technical solution, in S7, the original data that does not need to read the instruction address is specifically:

[0022] S4.1, the execution module directly obtains the calculation result according to the instruction operation code;

[0023] S4.2, write the calculation result to the instruction address;

[0024] S4.3. Finally, return to the completed status.

[0025] As a further improvement of the present technical solution, in S8, after the user receives the task completion package, it can be determined whether the current task is successfully completed and the next step can be executed.

[0026] On the other hand, the present invention provides a device supporting multi-GPU Atomic instruction operations, including an execution module, an Atomic instruction encoding module and an Atomic instruction decoding module. When the execution module executes the multi-GPU Atomic instructions, the steps of any of the above-mentioned methods for supporting multi-GPU Atomic instruction operations are implemented.

[0027] As a further improvement of the technical solution, the device supports Atomic instructions from the user side and also supports Atomic instructions from a remote GPU, thereby achieving synchronization of a GPU cluster.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. In the method and device for supporting multi-GPU Atomic instruction operations, the Atomic instructions are extensible. The Atomic instructions can be flexibly expanded according to scene requirements and are not limited by the PCIe protocol.

[0030] 2. The method and device for supporting multi-GPU Atomic instruction operations are highly efficient. Atomic instructions support flexible expansion, can meet the needs of different scenarios, and improve the overall performance of the GPU cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] See also Figure 1 As shown, this embodiment provides a method for supporting multi-GPU Atomic instruction operations, including the following steps:

[0034] S1. The user sends tasks to the GPU according to the requirements. The execution module inside the GPU parses and executes the tasks and issues various instructions.

[0035] In this example, the user needs multiple GPUs to complete a task together. The task is sent to the GPU, which parses the task after receiving it and issues instructions based on the parsed task requirements.

[0036] S2, the Atomic instruction encoding module determines whether it is an Atomic instruction according to the instruction operation code;

[0037] In this example, the supported Atomic types can be defined according to needs, including but not limited to logical operations, addition, subtraction, multiplication, increment, decrement, maximum value, minimum value, replacement, comparison operations, and support for custom data bit widths, including but not limited to 4-bit, 8-bit, 16-bit, 32-bit, 64-bit, and 128-bit data sizes.

[0038] S3. If it is an Atomic instruction, encode the instruction opcode control information into the PCIe TLP Prefix, encode the address and data carried by the Atomic instruction into the PCIe TLP, enable the PCIe TLP Prefix function, and transmit it to the destination GPU;

[0039] S4. If it is not an Atomic instruction, there is no need to enable the PCIe TLP Prefix function, and the Atomic instruction decoding module of the target GPU is directly decoded;

[0040] S5, the Atomic instruction decoding module of the target GPU decodes the instruction operation code from the PCIe TLP Prefix;

[0041] In this example, the S3 is used to obtain the address and data of the Atomic instruction from the PCIe TLP according to the instruction opcode, and transmit them to the local execution module at the same time.

[0042] S6. All Atomic instructions are divided into two types: those that need to return original data and those that do not. If an Atomic instruction needs to return original data, it will return the completion status and original data.

[0043] In this example, the specific steps to return the completion status and original data are:

[0044] S3.1, the execution module first reads the original data at the corresponding position according to the instruction address;

[0045] S3.2, then write the calculated data to the instruction address according to the instruction operation code;

[0046] S3.3. Finally, return the original data and completion status.

[0047] S7. If it is an Atomic instruction that does not need to return the original data, there is no need to read the original data of the instruction address;

[0048] In this example, the original data that does not need to read the instruction address is:

[0049] S4.1, the execution module directly obtains the calculation result according to the instruction operation code;

[0050] S4.2, write the calculation result to the instruction address;

[0051] S4.3. Finally, return to the completed status.

[0052] S8. If the execution module of the requesting GPU recognizes the return as Atomic, a task completion package is generated according to the execution result and returned to the user.

[0053] In this example, after the user receives the task completion package, he or she can determine whether the current task is successfully completed and perform the next step.

[0054] Embodiment 1:

[0055] In the initial state, the Host allocates the size and address of the GPU cluster synchronization space, such as starting from address 8_0000_0000, the 32-bit space is the synchronization space. This space is located inside GPU1.

[0056] If the host requires GPU0 and GPU1 to jointly execute a task, it will send part of the task to the two GPUs respectively.

[0057] The execution module inside the GPU will parse and execute the task after receiving it.

[0058] If GPU0 needs to synchronize execution results with GPU1, it issues an Atomic instruction to GPU1 that returns the original data, and the Atomic instruction carries the execution results of GPU0.

[0059] After the Atomic instruction encoding module of GPU0 recognizes the Atomic instruction, it enables the PCIe TLP Prefix function and encodes control information such as the instruction opcode into the PCIe TLP Prefix. The address and data carried by the Atomic instruction are encoded into the PCIe TLP and transmitted to GPU1.

[0060] The Atomic instruction decoding module of GPU1 decodes the instruction operation code from the PCIe TLP Prefix, obtains the address and data of the Atomic instruction from the PCIe TLP according to the instruction operation code, and transmits them to the execution module at the same time.

[0061] After receiving the Atomic instruction, the execution module inside GPU1 first reads the data at address 8_0000_0000, then calculates a result based on the instruction opcode, and writes the result to address 8_0000_0000. Finally, the completion status and original data are returned to GPU0.

[0062] The Atomic instruction encoding module of GPU1 enables the PCIe TLP Prefix function, and encodes control information such as instruction completion status into the PCIe TLP Prefix, encodes the original data into the PCIe TLP, and transmits it to GPU0.

[0063] The Atomic instruction decoding module of GPU0 decodes the instruction completion status from the PCIe TLP Prefix, obtains the Atomic original data from the PCIe TLP, and then returns it to the execution module.

[0064] After receiving the returned completion status and original data, the execution module of GPU0 can obtain the execution result of GPU1. Meanwhile, GPU1 also obtains the execution result of GPU0.

[0065] This cycle is executed until the task is completed, and the GPU generates a task completion packet and sends it to the Host.

[0066] After the host receives the task completion package, it can determine whether the current task is successfully completed and execute the next step.

[0067] Embodiment 2:

[0068] In the initial state, the Host allocates the synchronization space size and address, such as starting from address 600_0000_0000, the 16-bit space is the synchronization space. This space is located inside GPU0.

[0069] The host requires GPU0 and GPU3 to jointly execute a task, so it sends part of the task to the two GPUs respectively.

[0070] The execution module inside the GPU will parse and execute the task after receiving it.

[0071] GPU3 needs to inform GPU0 of the execution result, so it issues an Atomic instruction to GPU0 without returning the original data, and the Atomic instruction carries the execution result of GPU3.

[0072] After the Atomic instruction encoding module of GPU3 recognizes the Atomic instruction, it enables the PCIe TLP Prefix function, and encodes control information such as the instruction opcode into the PCIe TLP Prefix. The address and data carried by the Atomic instruction are encoded into the PCIe TLP and transmitted to GPU0.

[0073] The Atomic instruction decoding module of GPU0 decodes the instruction opcode from the PCIe TLP Prefix, obtains the address and data of the Atomic instruction from the PCIe TLP according to the instruction opcode, and transmits it to the execution module at the same time.

[0074] After receiving the Atomic instruction, the execution module inside GPU0 calculates a result based on the instruction opcode, writes the result to address 600_0000_0000, and then returns the completion status to GPU3.

[0075] The Atomic instruction encoding module of GPU0 enables the PCIe TLP Prefix function, and encodes control information such as instruction completion status into the PCIe TLP Prefix, and transmits it to GPU3.

[0076] The Atomic instruction decoding module of GPU3 decodes the instruction completion status from the PCIe TLP Prefix and then returns it to the execution module.

[0077] The task is completed after the execution module of GPU3 receives the returned completion status.

[0078] After receiving the execution result of GPU3 and its own execution result, the execution module of GPU0 generates a task completion packet and sends it to the Host.

[0079] After the host receives the task completion package, it can determine whether the current task is successfully completed and execute the next step.

[0080] Embodiment 3:

[0081] This embodiment provides a device that supports multi-GPU Atomic instruction operations, including an execution module, an Atomic instruction encoding module, and an Atomic instruction decoding module, characterized in that: when the execution module executes the multi-GPU Atomic instruction, the steps of any of the above-mentioned methods for supporting multi-GPU Atomic instruction operations are implemented.

[0082] In this example, the device supports Atomic instructions from the user side and also supports Atomic instructions from a remote GPU, thereby achieving synchronization of a GPU cluster.

[0083] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for supporting multi-GPU Atomic instruction operations, characterized in that: The following steps are involved: S1. The user sends tasks to the GPU according to the requirements. The execution module inside the GPU parses and executes the tasks and issues various instructions. S2, the Atomic instruction encoding module determines whether it is an Atomic instruction according to the instruction operation code; S3. If it is an Atomic instruction, encode the instruction opcode control information into the PCIe TLP Prefix, encode the address and data carried by the Atomic instruction into the PCIe TLP, enable the PCIe TLP Prefix function, and transmit it to the destination GPU; S4. If it is not an Atomic instruction, there is no need to enable the PCIe TLP Prefix function, and the Atomic instruction decoding module of the target GPU is directly decoded; S5, the Atomic instruction decoding module of the target GPU decodes the instruction operation code from the PCIe TLP Prefix; S6. All Atomic instructions are divided into two types: those that need to return original data and those that do not. If an Atomic instruction needs to return original data, it will return the completion status and original data. S7. If it is an Atomic instruction that does not need to return the original data, there is no need to read the original data of the instruction address; S8. If the execution module of the requesting GPU recognizes the return as Atomic, a task completion package is generated according to the execution result and returned to the user.

2. The method for supporting multi-GPU Atomic instruction operations according to claim 1, characterized in that: In S1, the user requires multiple GPUs to jointly complete a task and sends the task to the GPU.

3. The method for supporting multi-GPU Atomic instruction operations according to claim 1, characterized in that: In S2, supported Atomic types include but are not limited to logical operations, addition, subtraction, multiplication, increment, decrement, maximum value, minimum value, replacement, and comparison operations; and support for custom data bit widths, including but not limited to 4-bit, 8-bit, 16-bit, 32-bit, 64-bit, and 128-bit data sizes.

4. The method for supporting multi-GPU Atomic instruction operations according to claim 1, characterized in that: The S5 is used to obtain the address and data of the Atomic instruction from the PCIe TLP according to the instruction operation code, and transmit them to the local execution module at the same time.

5. The method for supporting multi-GPU Atomic instruction operations according to claim 1, characterized in that: In S6, the specific steps of returning the completion status and original data are: S3.1, the execution module first reads the original data at the corresponding position according to the instruction address; S3.2, then write the calculated data to the instruction address according to the instruction operation code; S3.

3. Finally, return the original data and completion status.

6. The method for supporting multi-GPU Atomic instruction operations according to claim 1, characterized in that: In S7, the original data that does not need to read the instruction address is specifically: S4.1, the execution module directly obtains the calculation result according to the instruction operation code; S4.2, write the calculation result to the instruction address; S4.

3. Finally, return to the completed status.

7. The method for supporting multi-GPU Atomic instruction operations according to claim 1, characterized in that: In S8, after receiving the task completion package, the user can determine whether the current task is successfully completed and execute the next step.

8. A device supporting multi-GPU Atomic instruction operations, comprising an execution module, an Atomic instruction encoding module and an Atomic instruction decoding module, characterized in that: When the execution module executes the multi-GPU Atomic instructions, the steps of the method for supporting multi-GPU Atomic instruction operations described in any one of claims 1 to 7 are implemented.

9. The device for supporting multi-GPU Atomic instruction operations according to claim 8, characterized in that: The device supporting multi-GPU Atomic instruction operation supports Atomic instructions from the user side and also supports Atomic instructions from a remote GPU for realizing synchronization of a GPU cluster.

Citation Information

Patent Citations

  • Processor fuzzy testing method supporting runtime instruction variation

    CN117033101A

  • Bus agent capable of supporting extended atomic operations and method therefor

    US20130346655A1