Neural network parallel scheduling-oriented single-instruction multi-thread processor micro-architecture device

By adopting a single-instruction multi-threaded processor microarchitecture in the edge device, optimizing thread scheduling and resource sharing, the problems of high resource consumption and low computational efficiency are solved, and efficient and energy-saving neural network inference tasks are realized.

CN120832173APending Publication Date: 2025-10-24XI AN JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510742277.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

In resource-constrained edge devices, existing technologies struggle to balance computational efficiency and energy consumption for neural network inference tasks. Multi-core architectures lead to excessive resource consumption, and the fine-grained instruction set of GPGPUs causes efficiency issues.

Method used

It adopts a single-instruction multi-threaded processor microarchitecture for parallel scheduling of neural networks. By assigning a unique number to each thread and organizing them into thread groups, it uses a fair round-robin arbiter to select instructions and combines pipeline layout and resource sharing mechanism to optimize the parallelism of computing array and data transfer module, thereby reducing waiting time and resource requirements.

Benefits of technology

It improves the utilization of computing resources, reduces energy consumption, is suitable for resource-constrained edge devices, and enables efficient parallel processing, especially suitable for complex parallel tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832173A_ABST
    Figure CN120832173A_ABST
Patent Text Reader

Abstract

A single-instruction multi-thread processor micro-architecture device oriented to neural network parallel scheduling comprises a front-end instruction fetching module, an instruction cache module, a decoding module, an arithmetic logic operation unit, a multiplication and division module, a memory access unit and a data cache module, and the micro-architecture device allocates a unique thread number for each thread. Threads are organized into thread groups, and in each period, the fair alternate arbiter selects an instruction from an instruction buffer of the thread group and sends the instruction to a subsequent decoding stage. The micro-architecture not only solves challenges faced by end-side equipment when processing high-performance calculation requirements such as neural network reasoning, but also provides an effective solution capable of reducing energy consumption and improving calculation efficiency through an innovative architecture design. The method is of great significance in promoting development of end-side AI application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the field of artificial intelligence neural network technology, and in particular relates to a single instruction multi-threaded processor micro-architecture device for parallel scheduling of neural networks. Background Art

[0002] When deploying neural network inference acceleration tasks on edge devices, a multi-threaded concurrent preemptive model is often used to maximize compute array utilization and simplify the programming model. This model focuses on improving the average utilization of all threads rather than focusing on the performance of individual threads. While general-purpose graphics processing units (GPGPUs) offer a relatively effective solution, their fine instruction granularity makes their energy efficiency unsuitable for resource- and power-constrained edge devices.

[0003] The current mainstream approach to enhancing the CPU's ability to handle parallel tasks in edge devices is to adopt a multi-core architecture, allowing each core to independently execute different tasks. While this approach effectively achieves task-level parallelism, it also comes with high resource consumption. In resource-constrained applications, such as embedded systems, having too many cores can become a burden. Therefore, adopting a multi-threaded design is particularly important: on the one hand, it can utilize time slices when other threads encounter delays (such as cache misses or waiting for I / O operations to complete) to schedule other threads to run, thereby achieving implicit parallelism between tasks. On the other hand, multiple threads share the same computing units and instruction / data caches, which not only reduces hardware resource requirements but also makes them more suitable for resource-constrained applications.

[0004] In neural network inference tasks, operator operations based on the sliced ​​feature maps are used as the basic units of multi-threaded scheduling. By optimizing the scheduling strategy, the parallelism between the computing array and the data movement module is fully utilized, so that when some tasks are executed on the computing array, other necessary data can be loaded or moved out at the same time, thereby reducing waiting time and improving resource utilization. Since such tasks take a long time to execute, the real-time computing power requirements of the single-instruction multi-threaded SIMT processor are relatively low, effectively avoiding the efficiency problems caused by the fine granularity of instructions like GPGPU, and reducing the overhead caused by frequent context switching. Therefore, this method not only improves the overall execution efficiency and energy consumption ratio of the system, but is also particularly suitable for resource-constrained end-side devices, providing a new solution for achieving efficient and energy-saving neural network inference. Summary of the Invention

[0005] In view of this, the present disclosure provides a single instruction multi-threaded processor micro-architecture device for parallel scheduling of neural networks, including: a front-end instruction fetch module, an instruction cache module, a decoding module, an arithmetic logic unit, a multiplication and division module, a memory access unit and a data cache module.

[0006] The micro-architecture device allocates a unique thread number to each thread, and the threads are organized into thread groups. In each cycle, a fair round-robin arbiter selects instructions from the instruction buffers of the thread groups to send to the subsequent decode stage.

[0007] By the above technical solution, not only the needs of the end-side device for high energy efficiency are met, but also the utilization rate of the computing array is improved and the programming model is simplified. This architecture design is particularly optimized for the end-side environment, aiming to achieve efficient parallel processing capability with low energy consumption. An innovative coarse-grained single instruction multiple thread (SIMT) processor micro-architecture is proposed, which is a multi-threaded CPU micro-architecture design scheme based on the 32-bit RISC-V instruction set. Through the carefully designed thread switching mechanism and pipeline layout, efficient parallel execution between multiple threads is realized. Compared with the traditional single-core CPU, although the resource consumption of the scheme of the present application only increases by about 16%, it can support up to four threads running simultaneously, greatly improving the utilization efficiency of computing resources, and is particularly suitable for application scenarios that need to run complex parallel tasks under strict resource constraints. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 is the overall architecture schematic diagram of the multi-threaded processor provided in one embodiment of the present disclosure;

[0009] Figure 2 is the micro-architecture schematic diagram of the front-end instruction fetching module in one embodiment of the present disclosure;

[0010] Figure 3 is the instruction cache structure block diagram in one embodiment of the present disclosure;

[0011] Figure 4 is the decode stage schematic diagram in one embodiment of the present disclosure;

[0012] Figure 5 is the execution unit schematic diagram in one embodiment of the present disclosure;

[0013] Figure 6 is the internal structure diagram of the multiplication and division module in one embodiment of the present disclosure;

[0014] Figure 7 is the memory access unit pipeline structure schematic diagram in one embodiment of the present disclosure;

[0015] Figure 8 is the pipeline schematic diagram of three consecutive load operation instructions in one embodiment of the present disclosure;

[0016] Figure 9 is the schematic diagram of the case where consecutive load operation instructions encounter blocking in one embodiment of the present disclosure;

[0017] Figure 10 is a schematic diagram of the internal structure of a data cache in one embodiment of the present disclosure;

[0018] Figure 11 is a diagram of the content of a write buffer table entry in one embodiment of the present disclosure;

[0019] Figure 12 is a state diagram of a backfill unit in one embodiment of the present disclosure;

[0020] Figure 13 is a diagram of pipelining processing at the software level in one embodiment of the present disclosure;

[0021] Figure 14 is a diagram of the overall architecture of a neural network accelerator in one embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] To further describe the present application, the following will combine the accompanying drawings Figures 1 to 14 to make further descriptions.

[0023] In one embodiment, as Figure 1 shown, it discloses a single instruction multiple thread processor micro-architecture device for neural network parallel scheduling, including: a front-end instruction fetching module, an instruction cache module, a decoding module, an arithmetic logic unit, a multiplication division module, a memory access unit and a data cache module,

[0024] The micro-architecture device allocates a unique thread number to each thread, and the threads are organized into thread groups. In each cycle, a fair round-robin arbiter selects instructions from the instruction buffer of the thread group and sends them to the subsequent decoding stage.

[0025] In this embodiment, the present application not only solves the challenges faced by end-side devices in processing high-performance computing requirements such as neural network inference, but also provides an effective solution that can reduce energy consumption and improve computing efficiency through innovative architectural design. This has important significance for promoting the development of end-side AI applications.

[0026] To ensure that multiple threads can correctly and efficiently share the same pipeline, the system allocates a unique thread number to each thread. This number is used to index the corresponding register array and program counter (PC), ensuring the data independence and execution accuracy of each thread.

[0027] In the processor architecture, the front-end instruction fetching module is responsible for interacting with the instruction cache and mainly undertakes the task of fetching instructions. Specifically, the front-end will send a fetch request to the instruction cache according to the needs of each thread, and store the instructions obtained from the instruction cache into the instruction buffer of the corresponding thread. Figure 1The design example supports the parallel execution of 4 threads (i.e., contains 4 instruction buffers), but the architecture is not limited to supporting only 4 threads. By simply configuring the parameters, it can be flexibly adjusted to adapt to the needs of different numbers of threads, thereby realizing micro-architecture solutions with different numbers of threads. This design not only enhances the flexibility and scalability of the system, but also effectively improves resource utilization and overall processing efficiency.

[0028] In the 4-thread processor architecture, threads are organized into two thread groups, each containing 2 threads. In each cycle, the fair round robin arbiter (RR Arbiter for short) selects two instructions from the instruction buffers of the two thread groups to send to the subsequent decoding stage. This mechanism ensures fair scheduling and efficient execution among different threads.

[0029] The decoding stage of the pipeline is equipped with two decoding modules, which can simultaneously decode and process two instructions. According to the RISC-V instruction set definition, the instructions are parsed, and the corresponding source operands are read from the general register array according to the current thread number. In addition, the decoding stage also checks whether there is valid data available in the bypass network. If there is valid data in the bypass network, these data are used preferentially; otherwise, the required data are obtained from the register array.

[0030] After decoding and reading the source operands, the instructions are distributed to different execution units for calculation according to their operation codes. Given that the processor in the present application supports the RISC-V IM instruction set extension, it contains an arithmetic logic unit (ALU) dedicated to performing arithmetic and logical operations, as well as a multiply-divide unit (MDU) for multiplication and division operations. The ALU can complete an arithmetic or logical operation in each cycle. The multiplier in the MDU adopts a three-stage pipeline structure, and the entire multiplication operation takes three cycles to complete. The divider in the MDU can only process one instruction at a time, and it takes twenty clock cycles to complete a division operation. This design not only improves the execution efficiency of the processor in a multi-threaded environment, but also further improves the response speed and throughput of the overall system by optimizing the performance of key execution units, such as the multiplier using pipeline technology and the efficient bypass network.

[0031] Considering that the proportion of multiplication and division operations in actual applications is relatively low, in order to optimize the use of hardware resources and reduce unnecessary consumption, the present application specially designs a resource sharing mechanism. Specifically, we configure two arithmetic logic units (ALUs) for the two thread groups, ensuring that each thread group can efficiently perform arithmetic and logical operations. At the same time, only one multiply-divide unit (MDU) is set up for the two thread groups to share

[0032] Furthermore, to further optimize resource utilization and reduce consumption, the present application reuses the adder in the arithmetic logic unit (ALU) to not only perform the ADD operation defined in the RISC-V 32-bit I instruction set, but also to be used for address calculation of Load / Store instructions. The memory unit (LSU) is responsible for handling the specific execution flow of Load and Store instructions. After the ALU completes the address calculation, it writes the calculated address to a buffer in the LSU. Subsequently, the LSU takes out the memory request and its corresponding address from the buffer every cycle and sends it to the data cache (D-cache) to access the required data.

[0033] If the current memory request results in a D-cache miss, this situation will be reported to the LSU. In order to avoid the time delay caused by data cache missing affecting system performance, and to prevent blocking the execution of other threads, when data cache missing occurs, the LSU will notify the control module of the corresponding thread. Under the guidance of the control module, only the pipeline part related to the thread is cleared, and the thread is put into sleep state, suspending its instruction fetching process, until the required data is filled back to the data cache from the next level of memory, at which time the thread is awakened to continue execution.

[0034] In the data write-back phase, the multiplier module and the memory unit may attempt to write back data at the same time, causing competition for the write-back bus. To solve this problem, the present application sets the multiply-divide module to have a higher write-back priority. This means that when the multiply-divide module is performing data write-back, the write-back operation of the memory unit is temporarily prevented, causing the memory unit to stall for one clock cycle. This mechanism ensures that critical operation results can be written back first, while minimizing the efficiency loss due to resource contention, ensuring efficient operation and response speed of the system. Through these designs, the present application effectively improves resource utilization, reduces energy consumption, and enhances the concurrent processing capability of the overall system.

[0035] The number of threads involved in the present application can be configured by a global parameter, and can be flexibly configured for different application scenarios. The number of processor bits involved in the present application can also be flexibly configured, and a 64-bit processor can be realized by simple modification. The present application is applied to the control and scheduling module in the neural network accelerator, and can also be applied to the control and scheduling task of high-performance parallel computing tasks.

[0036] In another embodiment, the front-end instruction fetching module includes an instruction request generation module and a PC generation module, which is responsible for interacting with the instruction cache module and mainly undertakes the task of instruction fetching.

[0037] For this embodiment, the front-end fetch module is responsible for generating fetch requests for each thread according to its state, accessing the instruction cache, and writing the read data into the corresponding thread's instruction buffer. The micro-architecture diagram of the entire front-end fetch module is shown in FIG. 1. The front-end fetch module contains two key modules: the fetch request generation module and the PC generation module. The following will analyze these two modules in detail. Figure 2

[0038] a) Fetch request generation module

[0039] The fetch request generation module decides whether to issue a fetch request according to the current thread's state. The state of each thread consists of two parts: instruction cache miss state and data cache miss state. Only when both parts of the state are 0, the thread is allowed to issue a fetch request. Specifically, if an instruction cache miss occurs, the instruction cache miss state needs to be set to 1, indicating that the thread enters a sleep state due to the instruction cache miss, until the data is backfilled from the next level of memory into the instruction cache. Similarly, the processing logic of the data cache miss state also follows this principle to ensure that the thread state can be properly managed when the data cache miss occurs.

[0040] In order to optimize resource utilization and maximize clock frequency, this design only configures one access port for the instruction cache. At the same time, in order to avoid the situation that a thread is "starved" (i.e., its fetch request is always not processed) due to the inability to access the instruction cache, the fetch request generation module sends all threads' fetch requests to a fair round robin arbiter. The arbiter is responsible for deciding which thread can access the instruction cache in each cycle. When multiple threads issue fetch requests at the same time, the arbiter will select the most suitable thread for access based on historical records, which can not only ensure fairness, but also prevent a thread from monopolizing the instruction cache access channel for a long time, thus ensuring the overall efficiency of the system and balanced load among threads.

[0041] b) PC generation module

[0042] The PC generation module is responsible for specifying the fetch address for each thread. During program execution, the PC value will change according to various conditions. The following are several main scenarios that cause PC changes:

[0043] 1. Sequential increment: In the normal execution flow, the PC value will sequentially increment by the instruction length (4 bytes per instruction), i.e., PC = PC + 4.

[0044] 2. Instruction cache miss: When an instruction cache miss occurs, the PC needs to be reset to the fetch address at the time of the miss to ensure that the required instruction can be reacquired.​

[0045] 3. Branch misprediction: If the actual jump address or direction does not match the result of the ALU calculation, it indicates that a branch misprediction has occurred. At this time, the PC needs to be adjusted to the correct jump address.

[0046] 4. Data cache miss: To mask the delay of data cache miss and ensure that the thread can continue execution after data refill, when a data cache miss occurs, the thread will be put into a sleep state. After the thread is awakened, the PC needs to be restored to the address before the data cache miss occurs to continue execution.

[0047] 5. Interruption or exception: When an interruption or exception occurs, the program control flow needs to be immediately diverted to the corresponding interruption service program entry, at which time the PC should be set to the starting address of the interruption handling program.

[0048] Since the above five situations may occur independently or simultaneously affect the modification of the PC value, in order to avoid confusion and ensure the correct execution of the program, the priority must be determined according to the position of these situations in the pipeline. Table 1 shows the priority of the five situations and their relative position in the pipeline, where the larger the priority value, the higher the priority. Generally, the situation located at the rear of the pipeline has a higher priority, because the instructions at the rear of the pipeline should be executed in priority to the instructions at the front. This design ensures that even if multiple factors simultaneously affect the change of the PC value, they can be effectively managed through a clear priority order, thus maintaining the accuracy and efficiency of program execution. In this way, the present application not only improves the response speed and reliability of the system, but also optimizes the resource management and scheduling strategy in a multi-threaded environment.

[0049]

[0050] Table 1

[0051] Based on the above mechanism, each thread in each cycle will send a fetch request to the fair round-robin arbiter according to its current state. The arbiter is responsible for selecting the most appropriate thread request and according to the thread number index corresponding to the program counter (PC). Then, the arbiter will send the information including the fetch request, address and thread number to the instruction cache for instruction fetching. Once the fetch request of a thread is received by the instruction cache, the PC of the thread will automatically increment by 4, pointing to the location of the next instruction. If an instruction cache miss occurs, the PC needs to be reset to the fetch address that caused the miss, so that after data backfill, the instruction can be reattempted to be fetched. The processing of data cache miss is similar to this. When a data cache miss occurs, the related thread will be temporarily suspended, and when it resumes execution, the PC will be reset to the address at which the miss occurred, ensuring that the execution continues from the breakpoint. When an interrupt or exception occurs, the control module will intervene and reset the PC to the target entry address, making the program jump to the starting position of the corresponding interrupt service program. This mechanism ensures that even in a complex and variable running environment, the stability and response speed of the system can be maintained through precise control and efficient scheduling.

[0052] In this way, the present application not only optimizes the resource allocation and scheduling efficiency in a multi-threaded environment, but also ensures that the system can still maintain a high-efficiency and reliable running state when facing various abnormal situations. This design not only improves the overall performance, but also enhances the robustness and adaptability of the system, providing a solid foundation for efficient task processing.

[0053] In another embodiment, the fetch request sent by the front-end fetch module to the instruction cache module is first written to a fetch buffer, realizing the complete decoupling of the front-end fetch module and the instruction cache module.

[0054] For this embodiment, the main functions of the instruction cache include receiving the fetch request from the front-end module, querying the tag memory in the instruction cache, and comparing whether the tags match. If the tags match, a cache hit signal is generated; otherwise, if the tags do not match, a cache miss signal is generated. When a cache hit occurs, the instruction cache returns the required data to the front-end module to continue the instruction fetching process. When a cache miss occurs, the instruction cache not only needs to return the fetch address that caused the miss to reset the PC value of the corresponding thread, but also needs to read the data of the corresponding address from the next level of memory and backfill it into the instruction cache to ensure that subsequent access can hit.

[0055] The overall architecture of the instruction cache is as follows: Figure 3The architecture design aims to optimize the interaction efficiency between the front-end module and the instruction cache. By quickly responding to the fetch request, efficiently handling cache hit and miss cases, and quickly performing data backfill, the system ensures high-performance operation. In addition, this design also reduces the delay caused by cache miss, improves the overall resource utilization and execution efficiency in a multi-threaded environment. In this way, the invention ensures the stability and response speed of the system while effectively managing and scheduling key computing resources.

[0056] The fetch request sent by the front-end fetch module to the instruction cache is first written to the fetch buffer. This design takes into account the case where the instruction cache miss occurs for a certain thread, and the issued fetch request needs to be cleared. Each thread is equipped with an independent fetch buffer, so that when a cache miss occurs, only the fetch buffer of the corresponding thread needs to be cleared, simplifying the implementation of the flush operation.

[0057] To ensure that each thread has access to the instruction cache, the invention uses a fair round-robin arbiter that selects a fetch request from multiple fetch buffers in each cycle and forwards it to the control module of the instruction cache. The control module reads the data storage area and tag storage area in the instruction cache memory array according to the fetch address, and determines whether it is a cache hit or miss based on this. The criteria include whether the fetch address matches the corresponding field in the tag memory and whether the cache line is valid. If there is a cache hit, the required data is returned directly to the front-end module; if there is a cache miss, the thread number and the address causing the miss are reported to the front-end module, causing the thread to enter a sleep state, and a cache line backfill request is sent to the backfill unit.

[0058] The backfill unit is responsible for reading a complete cache line of data from the next level of memory (such as DDR or L2 cache) based on the address causing the cache miss, and writing it to the instruction cache memory. To handle cache line backfill requests from multiple threads, a buffer called the backfill buffer is set up between the instruction cache control module and the backfill unit to store these requests. Once the backfill unit successfully completes a data backfill, it sends a wake-up signal to the front-end module to wake up the relevant thread from the sleep state and allow it to continue execution.

[0059] Due to the existence of the instruction buffer, the front-end module and the instruction cache module are completely decoupled, which not only improves the flexibility and simplicity of the design, but also enhances the scalability and maintainability of the system. In addition, by using a fair round-robin arbiter to schedule the instruction fetch requests of each thread, the fairness between threads is guaranteed, and the problem of some threads being unable to be served for a long time is avoided. This architecture design significantly improves the resource utilization and overall performance in a multi-threaded environment.

[0060] In another embodiment, in order to avoid too many instructions accumulating in the instruction buffer without being processed, one instruction is taken out from the instruction buffer of each thread group for decoding every cycle, which means that two instructions are taken out every cycle in a 4-thread design.

[0061] Since the two instructions come from different threads, there is no data dependency between the instructions, which avoids the need to check the dependency, simplifies the circuit design, saves resources, and also improves the execution efficiency of the threads.

[0062] The decoding module parses the operation code and data addressing mode of the instruction according to the RISCV 32 standard. The design of the decoding module is general and will not be described again.

[0063] In the decoding stage, in addition to parsing the instructions, the operands also need to be read. The source of the operands is two: valid data on the bypass network and data in the general register array.

[0064] The specific structure of the decoding stage is shown in Figure 4 When the parsing of the instructions and the reading of the data are completed, the instructions will be dispatched to different execution units for calculation according to the type of the instructions. The present application supports the IM instruction set of RSICV 32, so the execution unit includes an arithmetic logic operation unit (ALU) and a multiplication division unit (MDU). The ALU and the MDU will be introduced below.

[0065] In another embodiment, the arithmetic logic operation unit is used for arithmetic and logical operations, which can be completed in one clock cycle. Since the decoding stage processes two instructions at a time, the subsequent calculation unit is also required to accept two instructions at a time.

[0066] In order to avoid multiple instructions competing for the calculation unit and causing pipeline blocking, the present application sets up two ALUs to ensure that the calculation of two instructions can be completed in one clock cycle. The specific architecture diagram of the execution unit is shown in Figure 5 .

[0067] Take 4 threads as an example, ALU0 is responsible for the arithmetic logic operation of thread 0 / 1, and ALU1 is responsible for the arithmetic logic operation of thread 2 / 3. When the ALU completes the calculation, the result needs to be placed on the bypass network for the decoding module to use.

[0068] In order to save resources, the ALU in the application is not only used for the arithmetic operation and logical operation operation defined in the instruction set, but also used as the address generation of the Load / Store instruction, so when the ALU completes the address calculation, the memory request needs to be sent to the memory unit to access the data cache.

[0069] In order to simplify the design of the write-back network, regardless of what calculation the ALU completes, the calculation result needs to be written into the buffer in the memory unit. If the calculation result of the ALU is not the calculation of the address, it also needs to pass through the pipeline of the memory unit before being written back to the structure. The difference is that although it passes through the pipeline of the memory unit, it does not initiate a request to the data cache.

[0070] In another embodiment, the multiply-divide module is responsible for performing multiplication and division operations, and its position in the overall architecture is as shown in Figure 5 .

[0071] Since the multiply-divide instruction is less used in actual programs and the resources required to implement the multiply-divide module are larger, in order to reduce the area of the processor, only one multiply-divide module is implemented.

[0072] The multiply-divide module is shared by all threads, and when a thread fails to obtain its use right due to competition for the multiply-divide unit, the pipeline needs to be paused.

[0073] The multiplication unit in the multiply-divide unit uses 3 pipeline stages to complete the operation of the multiplication instruction, which means that the multiply-divide unit can simultaneously process the multiplication instructions of multiple threads. The division unit in the multiply-divide unit can only accept one instruction at a time, and cannot accept other instructions until the division unit completes the calculation.

[0074] The internal structure of the multiply-divide unit is as shown in Figure 6 . Among them, the multiplication stages 0 / 1 / 2 represent the 3 pipeline stages of the multiplication unit.

[0075] At the entrance of the multiply-divide unit, the multiply-divide requests of each thread need to pass through the selection of the arbiter to be executed. The multiply-divide only contains one output port, when the multiplication unit and the division unit simultaneously produce valid results, the division unit has priority to write back, and the multiplication unit needs to pause for one clock cycle.

[0076] In another embodiment, the load store unit is responsible for sending the load request to the data cache module and waiting for the valid data signal returned by the data cache module; when a data cache miss occurs, the thread that generates the miss automatically enters a sleep state until the data cache completes the backfill operation of the required data and issues a wake-up signal before being reactivated.

[0077] For this embodiment, the load store unit (LSU) serves as the interface between the CPU pipeline and the data cache (D-cache), responsible for sending load requests to the data cache and waiting for the valid data signal returned by the data cache. To effectively avoid pipeline blocking caused by data cache misses, a mechanism is introduced: when a data cache miss occurs, the thread that generates the miss automatically enters a sleep state until the data cache completes the backfill operation of the required data and issues a wake-up signal before being reactivated.

[0078] This mechanism ensures that the system can maintain an efficient running state even in the face of cache misses, and will not be affected by the execution efficiency of the entire pipeline due to the delay of a single thread. Specifically, when a thread encounters a data cache miss, it is temporarily suspended, while other threads can continue normal execution, thereby minimizing resource idling and performance loss. Once the missing data is successfully read from the next level of memory (such as DDR or L2 cache) and filled into the data cache, the corresponding thread receives a wake-up signal and resumes its normal execution flow.

[0079] The load module is a design fully decoupled from the data cache, and the load module and the data cache communicate using specified messages. The data cache does not have to rely on the design of the load module, making the design of the data cache more flexible and having a larger design space. At the same time, the load module adopts a pipeline structure, which can continuously process load instructions, making it have high performance. In another embodiment, the load unit adopts a three-stage pipeline architecture, marked as load stage 0, load stage 1, and load stage 2.

[0080] For this embodiment, the overall structure of the load module is as shown in Figure 7 The design adopts a three-stage pipeline architecture, marked as load stage 0, load stage 1, and load stage 2. When the ALU completes address calculation, the result is written to the load module buffer (LSU Buffer). The buffer is equipped with two write ports, supporting the writing of two load instructions at a time. In each cycle, the load module reads a load instruction from the buffer into the subsequent pipeline processing stage.

[0081] To achieve decoupling between modules and increase design flexibility, the information exchange between the memory access module and the data cache uses a request / response (Req / Rsp) handshake protocol. Specifically, the memory access module sends a request message (Request) to the data cache, and the data cache must reply with a response message (Response) to confirm that the request has been successfully received. To avoid timing constraints, the present invention stipulates that the response message should be returned in the next cycle after the request message is sent. This one-clock cycle delay design provides greater design space for the data cache, avoiding increased design complexity or performance bottlenecks caused by strict timing requirements.

[0082] The specific process of the memory access module processing memory access requests is as follows:

[0083] 1. Memory Access Phase 0: In this phase, the memory access module sends the first memory access request in the buffer to the data cache and prepares to enter Memory Access Phase 1.

[0084] 2. Memory Access Phase 1: This phase waits for the data cache to return a response message. If the response value is 1 (Response=1), it indicates that the data cache has successfully received the request and has begun processing it, allowing the instruction to enter Memory Access Phase 2. If the response value is 0 (Response=0), it means that the data cache has rejected the memory access request. In this case, the instruction will be blocked in Memory Access Phase 1 and continue to send request messages to the data cache until it is successfully received.

[0085] 3. Memory Access Phase 2: All instructions eligible for Memory Access Phase 2 have received successful receipt confirmation from the data cache. In the cycle following the response (i.e., Memory Access Phase 2), the data cache returns a cache hit or cache miss message. For a cache hit, the instruction proceeds to the subsequent write-back phase; for a cache miss, the relevant thread is put to sleep until the missing data is filled from the next level of memory into the data cache.

[0086] Through the above design, the present invention not only optimizes the communication mechanism between the memory access unit and the data cache, improving the system's response speed and data processing efficiency, but also ensures that the system can maintain efficient operation even in abnormal situations such as cache misses, thereby improving overall reliability and performance. This structure makes resource utilization more rational in a multi-threaded environment, enhancing the system's stability and adaptability.

[0087] The flow of the access instruction in the access unit pipeline is summarized as follows: access stage 0 sends a request message to the data cache. Access stage 1 parses the response message returned by the data cache and decides whether to flow into access stage 2. Access stage 2 parses the cache hit / miss signal returned by the data cache and decides whether to enter the write-back stage.

[0088] When the access instruction is blocked in access stage 2 (the response message is 0), it is necessary to continue to send a request message to the data cache until the returned response signal is equal to 1. Since access stage 0 is in an idle state at this time, it can receive a new access instruction in the access unit buffer, and the situation that access stage 0 and access stage 1 send request messages to the data cache at the same time occurs in the pipeline. For this situation, it is stipulated that access stage 1 has greater priority.

[0089] The access unit in the present application supports simultaneous processing of multiple access instructions without causing blocking. For example, there are 3 load operation instructions in the access unit buffer at the same time, and it is assumed that the data cache can respond to the request in time (meaning that it will not be blocked in access stage 1), then the position of the 3 load operation instructions in the pipeline changes with time as shown in FIG. 8.

[0090] However, for some reason, the data cache may not be able to respond to the request in time, at which time the access instruction sending the request will be blocked in access stage 1, and will continue to send a request message to the data cache until the data cache replies the response message. For example, there are 3 consecutive load operation instructions in the access unit buffer, when the load operation 0 instruction is blocked in access stage 1, it will be back-pressured to access stage 0, and the position of the access instruction in the pipeline changes with time as shown in FIG. 9. Figure 9

[0091] In cycle 0, load operation 0 is in access stage 0, and the reply signal from the data cache is 0, which means that the data cache refuses to accept the access request of load operation 0, and load operation 0 instruction needs to be blocked in access stage 1, and this cycle will continue to send a request message to the data cache by load operation 0.

[0092] Enter cycle 2, the response message returned by the data cache is 1, which means that the data cache accepts the access request of load operation 0, and load operation 0 is allowed to enter access stage 2, and this cycle will send an access request to the data cache by load operation 1.

[0093] Enter cycle 3, load operation 0 is in access stage 2, and in this cycle the data cache will return the cache hit / miss message and the valid data when the cache hits. At this time, load operation 1 is in access stage 1 and the response message is 1, and this cycle will send an access request to the data cache by load operation 2 instruction.

[0094] ​In another embodiment, in access phase 0, the access unit sends the first access request in the buffer to the data cache module and prepares to enter access phase 1; in access phase 1, waits for the data cache module to return a response message; in access phase 2, for the case of cache hit, the instruction will continue to enter the subsequent write-back phase, and for the case of cache miss, the relevant thread is put into a sleep state until the missing data is filled into the data cache module from the next level of memory.

[0095] In another embodiment, the information interaction between the access unit and the data cache module adopts a request / response handshake protocol.

[0096] In another embodiment, the pipeline of the data cache module is divided into three phases: phase 0, phase 1 and phase 2.

[0097] For this embodiment, the data cache module (D-cache) and the upper access unit (LSU) interact through request / response (Request / Response) handshake signals. This design allows the data cache to only need to ensure the timing relationship between the request and the response (i.e. the response message returns in the next cycle after the request message is sent out), without relying on the specific implementation details of the access unit. In this way, the design space of the data cache is greatly expanded, enabling it to have higher flexibility and optimization potential while meeting performance requirements.

[0098] When the data cache receives an access request from the access unit, in the next cycle of returning the response message, the data cache needs to provide the access unit with the status information of cache hit or cache miss. This mechanism ensures that even in the case of cache miss, the system can quickly respond and take appropriate measures to handle the missing data, such as obtaining the required data from the next level of memory.

[0099] The structure of the entire data cache module is shown in Figure 10 The architecture not only simplifies the interface design between modules, but also improves the overall efficiency and reliability of the system. By setting the timing delay between the request and the response to one cycle, the data cache obtains greater design freedom, avoiding the design complexity and potential performance bottleneck caused by strict timing constraints.

[0100] Moreover, this design enables independent optimization and improvement of the data cache from the memory access unit, enhancing the scalability and maintainability of the system. Regardless of different application requirements or technology upgrades, such an architecture can effectively support the requirements of high-performance computing, ensuring the efficiency and stability of data access processes, while significantly enhancing resource utilization efficiency and system performance in a multi-threaded environment.

[0101] In another embodiment, stage 0 corresponds to the memory access stage 0 of the memory access unit pipeline, in which the memory access request issued by the memory access unit is received, and the tag memory and data memory in the data cache memory array are read according to the memory address; stage 1 is used to check whether a cache hit or miss occurs; and stage 2 is used to return valid data to the upper-level memory access unit.

[0102] For this embodiment, in order to improve the ability to process memory access requests, a pipeline structure is used inside the data cache to support continuous processing of requests issued by the memory access unit. The pipeline of the data cache is divided into three stages: stages 0 / 1 / 2.

[0103] Stage 0 corresponds to the memory access stage 0 of the memory access unit pipeline, in which the memory access request issued by the memory access unit is received, and the tag memory and data memory in the data cache array are read according to the memory address (the data in the memory is valid after one cycle), but multiple modules may access the memory array simultaneously, requiring an arbiter to select. If the memory access request obtains access to the memory array, the request will enter stage 1 and return a response message with a value of 1 to the upper-level memory access unit after entering stage 1, indicating that the data cache has successfully received the memory access request. If the memory access request does not obtain access to the memory array, it cannot enter stage 1 and needs to return a response message with a value of 0 to the upper-level memory access unit after one cycle, indicating that the data cache rejects the memory access request.

[0104] Stage 1 is used to check whether a cache hit / miss occurs. When the memory access request is a load operation instruction, it is necessary to check whether the cache memory array and the write buffer contain valid data for the address. As long as either the cache memory array or the write buffer contains data for the address, it is considered a cache hit, otherwise it is a cache miss. Whether there is a cache hit or miss, the next cycle will enter stage 2.

[0105] Stage 2 is used to return valid data to the upper-level memory access unit. When it is a load operation instruction and there is a cache hit, the data read from the write buffer or the data memory needs to be returned to the memory access unit. When it is a store operation instruction and there is a cache hit, the data and address are written to the write buffer.

[0106] Write buffer is used to store the data that needs to be written by the store operation instruction. Since the store operation instruction supports byte-level write, the content of each entry of the write buffer is as shown in Figure 11 The address represents 32-bit address, the data represents 32-bit (4 bytes) data to be written, and the enable signal represents which bytes of the 4-byte write data are valid.

[0107] When judging whether the load operation instruction is cache hit, not only the address of the load operation instruction needs to be judged to be consistent with the address of a certain entry in the write buffer, but also the data length read by the load operation instruction needs to be consistent with the enable signal. For example, the load operation instruction needs to read 2-byte (half word) data starting from address 0x40. At this time, the address of a certain entry in the write buffer is 0x40, but its enable signal is 1100, representing that only the high 2 bytes of the 4-byte data after the 0x40 address are valid. Although the address of the load operation instruction is consistent with the address in the write buffer, it cannot be counted as cache hit at this time.

[0108] When a cache miss occurs in stage 1, the address and thread number information of the miss are written into the backfill buffer, waiting for the backfill unit to process.

[0109] The backfill unit is responsible for loading data from the lower-level memory into the cache memory according to the address. The entire control logic of the backfill unit is composed of a state machine, and its state diagram is as shown in Figure 12

[0110] When the backfill unit receives the backfill request issued by the backfill buffer, it first enters the clear cache line state to set the replaced cache line to invalid state. Then it enters the check dirty state to judge whether the replaced cache line is dirty. If it is dirty, it enters the write back cache line state to write the dirty cache line to the lower-level memory. Otherwise, it enters the load cache line state to load the cache line to be read into the data cache, and finally enters the set cache line state to set the state of the cache line to valid state.

[0111] In another embodiment, the data cache module supports multi-threading.

[0112] For this embodiment, the data cache supports multi-threading, which means that the backfill unit and the read-write logic of the data cache can be executed in parallel and decoupled from each other. In order to avoid the data error caused by the write buffer of the data cache writing data to the same cache line when the backfill buffer writes back the dirty cache line, the cache line is set to invalid state before the backfill unit writes back the dirty data. In this way, when the write buffer accesses the cache line, cache line hit does not occur, and data cannot be written to cause errors.

[0113] ​In another embodiment, the application of multi-threaded processors in neural network accelerators is described. Neural network accelerators use a large number of compute arrays for performing efficient matrix-vector operations, while using DMA (Direct Memory Access) to load or store data to on-chip storage devices. The typical computation flow of a neural network can be simplified into the following three main steps:

[0114] 1. Data Load (LOAD): Load input data from external dynamic random access memory (DDR) to on-chip memory. This step ensures that all data required for subsequent computations can be accessed quickly, reducing latency and improving overall processing efficiency.

[0115] 2. Operation Execution (EXE): Start the compute array to process the loaded input data and temporarily store the operation results in on-chip memory. This phase uses efficient hardware resources to perform complex matrix and vector operations to accelerate the inference or training process of neural network models.

[0116] 3. Result Store (STORE): Write the computed output data from on-chip memory back to DDR. This step ensures that the processed data can be used by other system components or subsequent processing stages.

[0117] The above three steps can be implemented in a pipelined manner when writing software, as shown in Figure 13 In the figure, "Load Data" represents the process of loading data from DDR to on-chip memory; "Compute" represents the process of performing operations; and "Store Data" refers to the operation of writing the operation results from on-chip memory back to DDR. Through this pipelined processing method, tasks in different stages can be overlapped and executed, significantly improving the throughput and efficiency of the entire system.

[0118] For example, while starting the load operation in the first cycle, the previous batch of data may be in the computation stage, and the data processed earlier is being stored. Such a design not only maximizes the utilization of hardware resources, but also effectively shortens the overall processing time, enabling the neural network accelerator to process large-scale data sets while maintaining high performance. This optimization strategy is particularly crucial for improving the performance of AI applications on edge devices.

[0119] Through careful design at the software level, the pipeline processing of each part of the calculation process can be realized, thereby improving the utilization rate of the calculation unit. However, this programming method also brings challenges to programmers, increases the programming difficulty, and makes the code writing process complex and difficult to maintain, deviating from the conventional thinking mode. In order to simplify the programming difficulty and improve the utilization rate of the operation unit as much as possible, a multi-threaded processor can be used as a scheduling unit of the neural network accelerator. The multi-threaded processor supports multiple threads to occupy the same operation array at the hardware level, which not only improves the utilization rate of the operation unit, but also eliminates the need for programmers to write pipeline processing programs. In addition, only one operation unit can support the simultaneous operation of 4 threads, further reducing resource consumption.

[0120] The structure of the entire neural network accelerator is shown in Figure 14 The multi-threaded processor represents the multi-threaded processor in the present application; the matrix operation unit is used for matrix operation; the vector operation unit is used for vector operation; the DMA (Direct Memory Access) is responsible for data transmission between the DDR and the local memory (Local Memory); and the local memory is used for temporary storage of data.

[0121] In the neural network accelerator, each thread is equipped with a DMA module to manage data transmission between the DDR and the local memory, and each thread is allocated a piece of on-chip storage space. Each thread realizes data transmission between the local memory and the DDR by configuring its own DMA module. Since the process of DMA data transmission usually takes a long time, after starting the DMA, the thread can be put into a dormant state until the DMA operation is completed and it is awakened by an interrupt signal. This way provides more execution opportunities for other threads and improves the overall efficiency of the system.

[0122] When multiple threads run simultaneously and need to use matrix, vector operation unit and other computing resources, they first need to perform resource preemption. The prerequisite for preemption is that the required data has been fully prepared. Once a thread successfully preempts the computing resources, it will enter a dormant state until the calculation is completed. Threads that fail to preempt resources will continue to try to obtain computing resources. This mechanism ensures that computing resources can be efficiently utilized in a multi-threaded environment, while simplifying the programming model, reducing development difficulty and maintenance cost.

[0123] Using a multi-threaded processor as the scheduling center of the neural network accelerator can greatly improve the utilization of computing resources. Multiple threads use computing resources through preemption mechanism, ensuring that these resources are almost always in an efficient working state. This approach not only maximizes the efficiency of hardware use, but also significantly reduces the burden on programmers. Programmers no longer need to write complex code to manually optimize performance, as performance improvement is directly implemented at the hardware level, completely transparent to programmers. This design makes the development process more simple, while also reducing the possibility of human error. In addition, the multi-threaded processor supports multiple users to run their respective neural network tasks simultaneously without the need to allocate a complete set of hardware resources for each task, thereby reducing overall resource consumption. This means that under the same hardware conditions, the system can more effectively serve more users, improving the sharing and flexibility of resources.

[0124] In summary, the neural network accelerator equipped with a multi-threaded processor not only liberates programmers, allowing them to focus on algorithm and application logic development rather than underlying performance optimization, but also greatly improves the utilization of computing resources. At the same time, in the case of limited resources, this architecture can effectively support the concurrent needs of multiple users, improving the overall service capacity and response speed of the system. This design not only enhances the scalability and adaptability of the system, but also provides a solid foundation for building efficient and flexible neural network applications.

[0125] Although the embodiments of the present application are described above in combination with the drawings, the present application is not limited to the above specific embodiments and application fields, and the above specific embodiments are only illustrative and guiding, but not limiting. Those skilled in the art can make many forms under the inspiration of the present application and without departing from the scope protected by the claims of the present application, which all belong to the protection of the present application.

Claims

1. A single instruction multiple thread processor microarchitecture apparatus oriented towards neural network parallel scheduling, comprising: a front-end fetch module, an instruction cache module, a decode module, an arithmetic logic unit, a multiply-divide module, a memory access unit and a data cache module, The micro-architecture device assigns a unique thread number to each thread, and the threads are organized into thread groups. In each cycle, a fair round-robin arbiter selects an instruction from the instruction buffer of a thread group and sends it to the subsequent decode stage.

2. The micro-architecture device according to claim 1, preferably, the front-end fetch module comprises a fetch request generation module and a PC generation module, which is responsible for interacting with the instruction cache module and mainly undertakes the task of fetching instructions.

3. The micro-architecture device according to claim 1, the fetch request sent by the front-end fetch module to the instruction cache module is first written into a fetch buffer, realizing the complete decoupling of the front-end fetch module and the instruction cache module.

4. The micro-architecture apparatus of claim 1, the memory access unit is responsible for sending a memory request to the data cache module and waiting for a valid data signal returned from the data cache module; wherein, When a data cache miss occurs, the thread that generates the miss will automatically enter a sleep state until the data cache completes the backfill operation of the required data and sends a wake-up signal to reactivate.

5. The micro-architecture device according to claim 1, the memory access unit adopts a three-stage pipeline architecture, marked as memory stage 0, memory stage 1 and memory stage 2.

6. The microarchitectural apparatus of claim 5, wherein, In memory stage 0, the memory access unit sends the first memory request in the memory request buffer to the data cache module and prepares to enter memory stage 1; in memory stage 1, it waits for the data cache module to return a response message; in memory stage 2, for the case of cache hit, the instruction will continue to enter the subsequent write-back stage, and for the case of cache miss, the related thread will be put into a sleep state until the missing data is filled from the next level of memory to the data cache module.

7. The micro-architecture device according to claim 1, the information interaction between the memory access unit and the data cache module adopts a request / response handshake protocol.

8. The micro-architecture device according to claim 1, the pipeline of the data cache module is divided into three stages: stage 0, stage 1 and stage 2.

9. The microarchitectural apparatus of claim 8, wherein, Stage 0 corresponds to the memory stage 0 of the pipeline of the memory access unit, in which the memory request issued by the memory access unit is received, and the tag memory and data memory in the data cache array are read according to the memory address; Stage 1 is used to check whether a cache hit or miss occurs; stage 2 is used to return valid data to the upper-level memory access unit.

10. The micro-architecture device according to claim 1, the data cache module supports multi-threading.

Citation Information

Cited By

  • Multi-thread instruction fetch scheduling system

    CN121614182A

  • A multi-threaded instruction fetch scheduling system

    CN121614182B