AI calculation method and system
The DMA engine on the disk directly interacts with the data between the disk and the GPU, solving the inefficiency problem caused by frequent data transfer in AI computing, and achieving more efficient data transmission and computing efficiency.
Patent Information
- Application Number
- CN202510183839.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
In AI computing, the data to be calculated and the data result need to be frequently moved between disk, system memory and GPU memory, resulting in a decrease in computing efficiency.
The DMA engine on the disk directly writes the data to be calculated from the disk to the GPU memory, and directly reads the result data from the GPU memory to the disk, avoiding data being transferred through the system memory.
This reduces the number of data transfers, improves the efficiency of data transmission, and thus improves the computing efficiency.
Smart Images

Figure CN120104534A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data communication technology, and in particular to an AI computing method and system. Background Art
[0002] like Figure 1 As shown in Figure 1, in AI (Artificial Intelligence) calculations, the general calculation process is:
[0003] 1. The CPU (Central Processing Unit) moves the data to be calculated from the disk (SSD) to the system memory (HostMemory);
[0004] 2. The CPU calls the DMA (Direct Memory Access) engine to move the data to be calculated from the system memory (Host Memory) to the GPU memory (Device Memory);
[0005] 3. The GPU (Graphics Processing Unit) starts calculating, and the GPU stores the calculation results in the GPU memory (DeviceMemory);
[0006] 4. The CPU calls the DMA engine to move the calculation results from the GPU memory (DeviceMemory) to the system memory (HostMemory);
[0007] 5. The CPU moves the calculation result data from the system memory (HostMemory) to the disk (SSD) to save the calculation result.
[0008] From the above process, it can be seen that in AI applications, the data to be calculated and the calculation result data need to be frequently moved between the disk (SSD)-system memory (HostMemory)-GPU memory (DeviceMemory). Excessive data movement operations greatly reduce the calculation efficiency. Summary of the invention
[0009] The technical problem to be solved by the present invention is to provide an AI computing method and system that can reduce the number of data movement times and improve computing efficiency.
[0010] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0011] An AI calculation method, comprising: a CPU calls a DMA engine on a disk to directly write data to be calculated from the disk to a GPU memory; a GPU performs AI calculations based on the data to be calculated and stores the calculated calculation result data in the GPU memory; and a CPU calls a DMA engine on a disk to directly read the calculation result data from the GPU memory to the disk and save it.
[0012] An AI computing system comprises a disk, a CPU and a GPU; the GPU is used to perform AI computing according to data to be calculated to obtain computing result data, and store the computing result data in GPU memory; the disk is used to cache the data to be calculated and save the computing result data, and the disk is configured with a DMA engine, the DMA engine is used to directly write the data to be calculated from the disk to the GPU memory and directly read the computing result data from the GPU memory to the disk; the CPU is used to call the DMA engine on the disk, so that the DMA engine directly writes the data to be calculated from the disk to the GPU memory and directly reads the computing result data from the GPU memory to the disk.
[0013] The beneficial technical effect of the present invention is that the present invention directly performs data interaction between the disk and the GPU through the DMA engine on the disk, so that the data exchange between the disk and the GPU no longer needs to be transferred through the system memory, which reduces the number of data movement times, improves the efficiency of data transmission, and thus improves computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A schematic diagram of the data migration route in the existing AI computing method;
[0015] Figure 2 A schematic diagram of the flow of the AI calculation method of the present invention;
[0016] Figure 3 Schematic diagram of the data migration route in the AI computing method of the present invention. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to more clearly understand the objectives, technical solutions and advantages of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0018] The present invention provides an AI computing method, which is applied to an AI computing system. The AI computing system includes an SSD, a CPU and a GPU. Data exchange between the disk and the GPU is directly performed through a DMA engine on the disk, so that the data exchange between the disk and the GPU is no longer transferred through a system memory (HostMemory), the number of data moves is reduced, the efficiency of data transmission is improved, and thus the computing efficiency is improved.
[0019] like Figure 2 , 3 As shown, in one embodiment of the present invention, the AI calculation method includes steps S10 to S30:
[0020] S10: The CPU calls the DMA engine on the disk to directly write the data to be calculated from the disk to the GPU memory.
[0021] S20: The GPU performs AI calculation according to the data to be calculated and stores the calculated calculation result data in the GPU memory.
[0022] S30: The CPU calls the DMA engine on the disk to directly read the calculation result data from the GPU memory to the disk and save it.
[0023] The disk in the present invention is a disk configured with a DMA engine. In order to adapt to AI application scenarios with high-speed computing, the disk in this embodiment adopts an SSD (Solid State Drive) with a high transmission rate.
[0024] The AI calculation method in this embodiment directly performs data exchange between the disk and the GPU through the DMA engine on the disk, so that the data exchange between the disk and the GPU no longer passes through the system memory (Host Memory), reducing the number of data movements, improving the efficiency of data transmission, and thus improving the computing efficiency.
[0025] In some preferred embodiments of the present invention, the step S10 further comprises:
[0026] S11: The CPU calls the disk driver to send a first data movement instruction to the disk and calls the DMA engine;
[0027] S12: The DMA engine in the disk executes the first data movement instruction to directly write the data to be calculated from the disk to the GPU memory.
[0028] In some preferred embodiments of the present invention, after executing step S12, that is, after the DMA engine writes the data to be calculated from the disk to the GPU memory, the following steps are also included: the DMA engine generates a first interrupt signal and sends the first interrupt signal to the CPU; the CPU notifies the GPU to start AI calculation after receiving the first interrupt signal. After receiving the first interrupt signal, the GPU performs AI calculation according to the data to be calculated.
[0029] In some preferred embodiments of the present invention, after executing step S20, that is, after the GPU stores the calculation result data in the GPU memory, the following steps are also included: the GPU generates a second interrupt signal and sends the second interrupt signal to the CPU; after receiving the second interrupt signal, the CPU calls the disk driver to send a second data movement instruction to the disk to call the DMA engine.
[0030] In some preferred embodiments of the present invention, the step S30 further comprises:
[0031] S31: The CPU calls the disk driver to send a second data moving instruction to the disk and calls the DMA engine;
[0032] S32: The DMA engine in the disk executes the second data moving instruction, directly reads the calculation result data from the GPU memory to the disk and saves it.
[0033] based on Figure 2 , 3 The AI computing method in the illustrated embodiment, the present invention also provides an AI computing system, which includes a disk, a CPU and a GPU; the GPU is used to perform AI calculation according to the data to be calculated to obtain calculation result data, and store the calculation result data in the GPU memory; the disk is used to cache the data to be calculated and save the calculation result data, and the disk is configured with a DMA engine, and the DMA engine is used to directly write the data to be calculated from the disk to the GPU memory and directly read the calculation result data from the GPU memory to the disk; the CPU is used to call the DMA engine on the disk so that the DMA engine directly writes the data to be calculated from the disk to the GPU memory and directly reads the calculation result data from the GPU memory to the disk.
[0034] The disk in the present invention is a disk configured with a DMA engine. In order to adapt to AI application scenarios with high-speed computing, the disk in this embodiment adopts an SSD (Solid State Drive) with a high transmission rate.
[0035] The AI computing system in this embodiment directly exchanges data between the disk and the GPU through the DMA engine on the disk, so that the data exchange between the disk and the GPU no longer needs to be transferred through the system memory (Host Memory), reducing the number of data movements, improving the efficiency of data transmission, and thus improving computing efficiency.
[0036] In some preferred embodiments of the present invention, the CPU sends a first data movement instruction to the disk by calling the disk driver to call the DMA engine; the DMA engine in the disk executes the first data movement instruction and directly writes the data to be calculated from the disk to the GPU memory.
[0037] In some preferred embodiments of the present invention, the DMA engine generates a first interrupt signal after writing the data to be calculated from the disk to the GPU memory, and sends the first interrupt signal to the CPU; after receiving the first interrupt signal, the CPU notifies the GPU to start AI calculation.
[0038] In some preferred embodiments of the present invention, the GPU generates a second interrupt signal after storing the calculation result data in the GPU memory, and sends the second interrupt signal to the CPU; after receiving the second interrupt signal, the CPU calls the disk driver to send a second data movement instruction to the disk to call the DMA engine.
[0039] In some preferred embodiments of the present invention, the CPU sends a second data movement instruction to the disk by calling the disk driver to call the DMA engine; the DMA engine in the disk executes the second data movement instruction, directly reads the calculation result data from the GPU memory to the disk and saves it.
[0040] In summary, the present invention directly performs data exchange between the disk and the GPU through the DMA engine on the disk, so that the data exchange between the disk and the GPU no longer needs to be transferred through the system memory (Host Memory), which reduces the number of data movements, improves the efficiency of data transmission, and thus improves computing efficiency.
[0041] The above description is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Those skilled in the art can make various equivalent changes and improvements based on the above embodiments, and all equivalent changes or modifications made within the scope of the claims should fall within the protection scope of the present invention.
Claims
1. An AI calculation method, characterized in that: include: The CPU calls the DMA engine on the disk to directly write the data to be calculated from the disk to the GPU memory; The GPU performs AI calculations based on the data to be calculated and stores the calculated calculation result data in the GPU memory; The CPU calls the DMA engine on the disk to directly read the calculation result data from the GPU memory to the disk and save it.
2. The AI calculation method according to claim 1, characterized in that: The CPU calls the DMA engine on the disk to directly write the data to be calculated from the disk to the GPU memory, including: The CPU calls the disk driver to send a first data movement instruction to the disk and calls the DMA engine; The DMA engine in the disk executes the first data movement instruction to directly write the data to be calculated from the disk to the GPU memory.
3. The AI calculation method according to claim 2, characterized in that: The CPU calls the DMA engine on the disk to directly read the calculation result data from the GPU memory to the disk and save it, including: The CPU calls the disk driver to send a second data movement instruction to the disk and calls the DMA engine; The DMA engine in the disk executes the second data movement instruction, directly reads the calculation result data from the GPU memory to the disk and saves it.
4. The AI calculation method according to claim 3, characterized in that: After the DMA engine writes the data to be calculated from disk to GPU memory, it also includes: The DMA engine generates a first interrupt signal, and sends the first interrupt signal to the CPU; After receiving the first interrupt signal, the CPU notifies the GPU to start AI calculation.
5. The AI calculation method according to claim 4, characterized in that: After the GPU stores the calculation result data into the GPU memory, the method further includes: The GPU generates a second interrupt signal, and sends the second interrupt signal to the CPU; After receiving the second interrupt signal, the CPU calls the disk driver to send a second data moving instruction to the disk to call the DMA engine.
6. An AI computing system, characterized in that: Including disk, CPU and GPU, The GPU is used to perform AI calculations based on the data to be calculated to obtain calculation result data, and store the calculation result data in the GPU memory; The disk is used to cache the data to be calculated and store the calculation result data. The disk is configured with a DMA engine, and the DMA engine is used to directly write the data to be calculated from the disk to the GPU memory and directly read the calculation result data from the GPU memory to the disk; The CPU is used to call the DMA engine on the disk, so that the DMA engine directly writes the data to be calculated from the disk to the GPU memory and directly reads the calculation result data from the GPU memory to the disk.
7. The AI computing system according to claim 6, wherein: The CPU sends a first data movement instruction to the disk by calling the disk driver to call the DMA engine; the DMA engine in the disk executes the first data movement instruction and directly writes the data to be calculated from the disk to the GPU memory.
8. The AI computing system according to claim 7, wherein: The CPU sends a second data moving instruction to the disk by calling the disk driver to call the DMA engine; the DMA engine in the disk executes the second data moving instruction, directly reads the calculation result data from the GPU memory to the disk and saves it.
9. The AI computing system of claim 8, wherein: The DMA engine generates a first interrupt signal after writing the data to be calculated from the disk to the GPU memory, and sends the first interrupt signal to the CPU; after receiving the first interrupt signal, the CPU notifies the GPU to start AI calculation.
10. The AI computing system according to claim 9, wherein: After storing the calculation result data in the GPU memory, the GPU generates a second interrupt signal and sends the second interrupt signal to the CPU; after receiving the second interrupt signal, the CPU calls the disk driver to send a second data movement instruction to the disk to call the DMA engine.