Method and apparatus for accelerating near-memory deep learning vector operation
The memory-near deep learning accelerator hardware structure and logic/method for vector reduction operations address the challenges of power consumption and memory delay in deep learning by efficiently processing large-scale matrix operations near memory, enhancing performance and energy efficiency.
Patent Information
- Application Number
- PCT/KR2023/019894
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
Existing deep learning architectures face challenges with increased power consumption and memory delay due to mass access to external memory during high-level deep learning operations, limiting efficient energy management and deep learning operation speed.
A memory-near deep learning accelerator hardware structure and logic/method for performing vector reduction operations inside the accelerator, considering complex operation conditions such as input data timing, ALU cycle, and number of instructions, to process large-scale matrix operations efficiently.
The proposed solution enables high-performance deep learning operations without memory delay, improving energy efficiency and operation speed by processing data closer to memory and adapting to complex operation conditions.
Smart Images

Figure KR2023019894_12062025_PF_FP_ABST
Abstract
Description
Method and device for accelerating memory-near deep learning vector operations
[0001] The present invention relates to artificial intelligence semiconductors and memory-near data processing technology, and more particularly, to logic and methods for performing vector reduction operations within an accelerator device near a memory for deep learning acceleration.
[0002] When processing high-level deep learning tasks in conventional computer architectures, such as large-scale matrix operations, power consumption increases due to the massive access to external memory, making efficient energy management difficult. Furthermore, limitations in external memory bandwidth lead to memory latency during data loading and storage, creating a data transfer bottleneck in the interaction between the processor and memory, limiting deep learning computational speed.
[0003] Therefore, data must be processed on an accelerator near memory, requiring computational logic capable of handling large matrices, such as BLAS (Basic Linear Algebra Subprogram), a fundamental deep learning operation. However, flexible logic that can be applied across a variety of computational situations, such as the timing of retrieving data from memory to the accelerator, the ALU (Arithmetic Logic Unit) computation cycle, and the number of instructions to be processed, is required, but such implementations are lacking.
[0004] The present invention has been made to solve the above problems, and the purpose of the present invention is to provide a memory-near deep learning accelerator hardware structure and a logic and method for accumulating data by considering complex operation conditions such as input data timing, ALU cycle, and number of instructions to be processed, in order to process high-performance deep learning without memory delay, and a vector reduction operation logic within the accelerator for large-capacity matrix operation in the structure.
[0005] In order to achieve the above object, a deep learning operation method according to one embodiment of the present invention includes the steps of: when a command is input, determining whether it is a new command for starting a new operation; if the determination result determines that it is a new command, storing the destination address of the command in a command slot; and loading vector data to be used as an operand from an external memory and a data memory and performing a first ALU operation using the operand.
[0006] The storage step may be initiated by setting the first flag in the instruction slot to the first value to indicate that it is the first operation of the instruction, and the first ALU operation execution step may be initiated by using the first value of the first flag as a control signal.
[0007] The deep learning operation method according to the present invention may further include, if it is determined as a result of the determination that the instruction is not new, a step of checking whether data having the same destination address is accumulated in the accumulation data slot; and, if it is determined as a result of the determination that data is accumulated, a step of loading the data stored in the external memory and the accumulated data in the accumulation data slot and performing a first ALU operation using them as operands.
[0008] The deep learning operation method according to the present invention may further include a step of performing a first ALU operation using data loaded from an external memory and 0 as an operand if the result of the verification is not accumulated.
[0009] The deep learning operation method according to the present invention may further include a step of determining whether the last operation of the instruction is completed when the first ALU operation is performed; and a step of storing data output as a result of the first ALU operation in an accumulation data slot together with a destination address if the last operation of the instruction is determined not to be completed as a result of the determination.
[0010] The deep learning operation method according to the present invention may further include a step of checking whether data having the same destination address is stored in the accumulation data slot; and, if the data is confirmed to be stored as a result of the check, a step of loading data having the same destination address from the accumulation data slot and performing a second ALU operation using the data stored in the result data slot as an operand.
[0011] The deep learning operation method according to the present invention may further include a step of storing the final data in a destination address if it is confirmed that the data is not stored.
[0012] The deep learning operation method according to the present invention may further include a step of performing a second ALU operation by loading data having the same destination address as the data output as the first ALU operation result and the data from the accumulated data slot, when it is determined that the last operation of the instruction has been completed as a result of the judgment.
[0013] The deep learning operation method according to the present invention may further include, when a second ALU operation is performed, a step of checking whether data having the same destination address is stored in the accumulation data slot; if it is confirmed as stored as a result of the check, a step of loading data having the same destination address from the accumulation data slot and performing the second ALU operation as an operand; if it is confirmed as not stored as a result of the check, a step of storing final data at the destination address.
[0014] According to another aspect of the present invention, a deep learning operation accelerator is provided, characterized by including: a command memory in which a command is stored; a data memory in which data is stored; and an operation logic that determines whether a command is a new command for starting a new operation when input from the command memory, and if the command is determined to be a new command as a result of the determination, stores a destination address of the command in an instruction slot, and loads vector data to be used as an operand from an external memory and a data memory and performs a first ALU operation using the operand.
[0015] According to another aspect of the present invention, a deep learning operation method is provided, characterized in that it includes the steps of: when a command is input, determining whether it is a new command for starting a new operation; if it is determined as a result of the determination that it is not a new command, determining whether data having the same destination address is accumulated in an accumulation data slot; and if it is determined as a result of the determination that data is accumulated, performing a first ALU operation by loading data stored in an external memory and the data accumulated in the accumulation data slot as operands.
[0016] According to another aspect of the present invention, a deep learning operation accelerator is provided, characterized by including: a command memory in which a command is stored; a data memory in which data is stored; and an operation logic which, when a command is input from the command memory, determines whether it is a new command for starting a new operation, and if it is determined as a result of the determination that it is not a new command, determines whether data having the same destination address is accumulated in an accumulation data slot, and if it is determined as a result of the determination that data is accumulated, loads data stored in an external memory and the data accumulated in the accumulation data slot to perform a first ALU operation using the data as an operand.
[0017] As described above, according to embodiments of the present invention, a vector reduction operation method used for large-scale matrix processing is proposed, which can be applied to the operation logic inside an accelerator near a memory, and in particular, the logic and method that take complex operation conditions into account can be applied to various operation situations occurring in the process of transferring and processing data between a memory and a deep learning accelerator.
[0018] Figure 1. Memory interface semiconductor in deep learning operations
[0019] Figure 2 is an internal structure of a memory module according to one embodiment of the present invention;
[0020] Figure 3 shows the internal structure of a deep learning operation accelerator.
[0021] Figure 4 shows the internal structure of the operation logic.
[0022] Figure 5 shows the number of slots according to the number of instructions and ALU cycles.
[0023] Figures 6 and 7 are drawings showing the process of performing a vector reduction operation when two instructions are input in a situation where the ALU cycle is 2.
[0024] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0025] Deep learning models require high-performance computing resources. Therefore, memory interface semiconductors require near-memory data processing technology to address the increased power consumption and data transfer latency issues associated with massive external memory access during deep learning operations, as shown in Figure 1.
[0026] One of these technologies is the development of a semiconductor structure that mounts a deep learning accelerator near the memory. However, in order to be adopted as hardware for processing high-performance deep learning models, logic for performing vector reduction operations within the accelerator is required to process large matrices such as BLAS (Basic Linear Algebra Subprogram) used in deep learning operations.
[0027] However, performing vector reduction operations using standard arithmetic logic is extremely difficult. This is because various operational situations arise during the transfer of instructions and data between external memory, the accelerator's internal memory, and the ALU unit. Complex conditions such as the timing of data input to the arithmetic logic, the ALU cycle, and the number of instructions to be processed must be considered. Depending on the situation, the arithmetic logic can be modified in a very complex manner.
[0028] Accordingly, in an embodiment of the present invention, logic and a method for performing vector reduction operations are presented by taking into account complex operation conditions occurring in a deep learning accelerator.
[0029] It is a vector reduction operation method and logic that can be applied inside an accelerator near memory. It is used for large-scale matrix processing in deep learning operations, and is a technology that can be flexibly applied to complex operation conditions such as data input time and ALU cycle.
[0030] FIG. 2 illustrates the internal structure of a DIMM (Dual In-line Memory Module) that can be applied as a memory module according to one embodiment of the present invention. As illustrated, the LRDIMM is composed of DRAMs for storing data, an RCD (Registered Clock Driver) and a DB (Data Buffer) for controlling the DRAMs, and is connected to a host core module.
[0031] A deep learning operation accelerator (100) is implemented in the DB. Fig. 3 is a diagram illustrating the internal structure of the deep learning operation accelerator (100). As illustrated, the deep learning operation accelerator (100) is configured to include an instruction memory (110), operation logic (120), and data memory (130).
[0032] The operation logic (120) is a logic for performing vector reduction operations within the deep learning operation accelerator (100). The internal structure of the operation logic (120) is illustrated in Fig. 4. As illustrated, the operation logic (120)
[0033] It is configured to include a command slot (121), an ALU (Arithmetic Logic Unit)-1 (122), an accumulation data slot (123), an ALU-2 (124), a result data slot (125), and switching elements (MUX / DEMUX) for logically connecting them.
[0034] The number of instruction slots (121), accumulated data slots (123), and result data slots (125) is variably adjusted according to the number of instructions input to the operation logic (120) and the cycle of the ALU (122, 124). Figure 5 shows the number of slots according to the number of instructions and the ALU cycle.
[0035] Figures 6 and 7 are diagrams illustrating a process of performing a vector reduction operation when two instructions are input in a situation where the ALU cycle is 2. Since the number of instructions is 2 and the ALU cycle is 2, as illustrated in Figure 5, there is 1 instruction slot (121), 6 accumulated data slots (123), and 1 result data slot (125).
[0036] In addition, the operation type of the command input to the operation logic (120) can be 'vector reduction sum' or 'vector reduction multiplication', and it is assumed that it includes 'opsize' indicating the operation size and a destination address where the final result data will be stored.
[0037] Vector data to be used for the operation is loaded from an external memory (DRAM) and an internal data memory (130) of a deep learning operation accelerator (100), and the operation logic (120) uses one of the elements of the vector as an operand. The timing for loading the vector into the operation logic (120) may vary during the data transfer process between the deep learning operation accelerator (100) and the external memory (DRAM), but the operation logic (120) according to an embodiment of the present invention can perform a vector reduction operation without affecting this timing.
[0038] The vector reduction operation process is as follows. When an instruction is input from the instruction memory (110), the operation logic (120) determines whether it is a new instruction for starting a new operation (1). If it is determined to be a new instruction, the operation logic (120) stores the destination address of the instruction in the instruction slot (121) and sets the first flag in the instruction slot (121) to 1 to indicate that it is the first operation of the instruction, and the first flag is used as a control signal to load vector data to be used as an operand from the external memory and data memory in the first operation and input it to the ALU-1 (122) (2).
[0039] Meanwhile, if the determination result is not a new instruction, that is, if it is determined to be a subsequent instruction rather than an instruction to start a new operation, the operation logic (120) checks the destination address stored in the instruction slot (121). If the same destination address is stored, the instruction has already been operated on before, so the first flag is set to 0. On the other hand, if the same destination address is not in the instruction slot (121), the instruction is a new instruction, so it is newly stored in the empty instruction slot (121) and the first flag is set to 1 in the same way. If the first flag is 0 (i.e., if it is not the start of a new operation), the operation logic (120) checks whether data with the same destination address is accumulated in the accumulation data slot (123). If the confirmation result is accumulated, the operation logic (120) loads the data stored in the external memory and the data accumulated in the accumulation data slot (123) and inputs them as operands to the ALU-1 (122). On the other hand, if the verification result is not accumulated, the operation logic (120) inputs the data loaded from the external memory and 0 as operands into ALU-1 (122) (3).
[0040] When data is output through an n-cycle operation of ALU-1 (122) (4), the operation logic (120) determines whether the last operation of the instruction is completed. If it is determined that the last operation of the instruction is not completed, the operation logic (120) stores the data output as a result of the operation of ALU-1 (122) together with the destination address in an empty accumulation data slot (123) (5).
[0041] If it is determined to be the last operation of the instruction (when opsize is 0), the operation logic (120) sets the end flag of the instruction slot (121) to 1. If the end flag is set to 1, the operation logic (120) loads the data output as the operation result of ALU-1 (122) and the data with the same destination address from the accumulation data slot (123) and inputs them as operands to ALU-2 (124) (6).
[0042] When data is output to the result data slot (125) over the n-cycle operation of ALU-2 (124) (7), the operation logic (120) checks whether data with the same destination address remains in the accumulation data slot (123). If it is confirmed that there is data left, the operation logic (120) loads the data stored in the result data slot (125) and the data with the same destination address from the accumulation data slot (123) and inputs them as operands into ALU-2 (124) (8).
[0043] On the other hand, if it is confirmed that there is no remaining, the operation logic (120) completes the operation and stores the output data of ALU-2 (124) in the destination address of the data memory (130) or external memory (9).
[0044] So far, a preferred embodiment of a method and device for accelerating memory-near deep learning vector operations has been described in detail.
[0045] In the above embodiment, in order to process high-performance deep learning without memory delay, a memory-near deep learning accelerator hardware structure and a logic and method for accumulating data by considering complex operation conditions such as input data timing, ALU cycle, and number of instructions to be processed, using vector reduction operation logic within the accelerator for large-scale matrix operations in this structure, are presented.
[0046] This presents a vector reduction operation method used in large-scale matrix processing, which can be applied to the operation logic inside the accelerator near the memory, and in particular, it can be applied to various operation situations that occur during data transfer and processing between the memory and the deep learning accelerator, with logic and methods that take complex operation conditions into account.
[0047] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. When a command is input, a step is taken to determine whether it is a new command to start a new operation; If the determination result is determined to be a new instruction, a step of storing the destination address of the instruction in the instruction slot; A deep learning operation method, characterized by including a step of loading vector data to be used as an operand from an external memory and a data memory and performing a first ALU operation with the operand.
2. In claim 1, The saving step is, To indicate that this is the first operation of an instruction, the first flag in the instruction slot is set to the first value, The first ALU operation execution stage is: A deep learning operation method characterized in that the first value of the first flag is used as a control signal and is initiated.
3. In claim 1, If the judgment result determines that it is not a new command, a step of checking whether data with the same destination address is accumulated in the accumulated data slot; A deep learning operation method, characterized in that it further includes a step of performing a first ALU operation by loading the data stored in the external memory and the accumulated data into the accumulated data slot as an operand if the verification result is accumulated.
4. In claim 3, A deep learning operation method, characterized in that it further includes a step of performing a first ALU operation using data loaded from an external memory and 0 as an operand if the verification result is not accumulated.
5. In claim 4, When the first ALU operation is performed, a step of determining whether the last operation of the instruction is completed; A deep learning operation method, characterized in that it further includes a step of storing data output as a result of a first ALU operation in an accumulation data slot together with a destination address if it is determined that the last operation of the judgment result instruction is not completed.
6. In claim 5, A step for checking whether data with the same destination address is stored in the accumulated data slot; A deep learning operation method, characterized in that it further includes a step of performing a second ALU operation by loading data having the same destination address from the accumulation data slot and the data stored in the result data slot if the verification result is confirmed to be stored.
7. In claim 6, A deep learning operation method, characterized in that it further includes a step of storing the final data in a destination address if the verification result is determined to be not stored.
8. In claim 5, A deep learning operation method, characterized in that it further includes a step of performing a second ALU operation by loading data having the same destination address from the accumulated data slot and data output as a result of the first ALU operation if it is determined that the last operation of the judgment result instruction has been completed; 9. In claim 1, When the second ALU operation is performed, a step of checking whether data with the same destination address is stored in the accumulation data slot; If the verification result is confirmed to be stored, a step of loading data having the same destination address from the accumulation data slot and the data stored in the result data slot to perform a second ALU operation as an operand; A deep learning operation method, characterized in that it further includes a step of storing the final data in a destination address if the verification result is determined to be not stored.
10. Command memory where commands are stored; Data memory where data is stored; A deep learning operation accelerator characterized by including an operation logic which determines whether a command is a new command for starting a new operation when a command is input from the command memory, and if the determination result determines that it is a new command, stores the destination address of the command in the command slot, and loads vector data to be used as an operand from an external memory and a data memory and performs a first ALU operation using the operand.
11. When a command is entered, a step is taken to determine whether it is a new command to start a new operation; If the judgment result determines that it is not a new command, a step of checking whether data with the same destination address is accumulated in the accumulated data slot; A deep learning operation method characterized by including a step of loading the data stored in the external memory and the accumulated data into the accumulated data slot and performing a first ALU operation as an operand if the verification result is accumulated.
12. Command memory where commands are stored; Data memory where data is stored; A deep learning operation accelerator characterized by including an operation logic which determines whether a command is a new command for starting a new operation when a command is input from the command memory, and if it is determined as not a new command as a result of the determination, checks whether data having the same destination address is accumulated in the accumulation data slot, and if the data is accumulated as a result of the determination, loads the data stored in the external memory and the data accumulated in the accumulation data slot to perform a first ALU operation using the operand.
Citation Information
Patent Citations
Semiconductor package
KR1020220058702A
Method For Automatic Control Of Incabin Environment For Passengers And The System For the Same
KR1020230064142A
Programmable accelerator for data-dependent, irregular operations
WO2023086353A1
KR20210154277A