Register renaming based on thread offset in multithreaded processing systems
The hardware-based register renaming circuit using thread offsets efficiently resolves register conflicts in multithreaded systems, enhancing performance and simplifying design complexity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2026-04-01
Smart Images

Figure 2026056594000001_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates to computer architecture, and more particularly to register renaming based on thread offsets in a multithreaded processing system. [Background technology]
[0002] Advances in data science, artificial intelligence (AI), and machine learning (ML) are transforming technology across various industries. To cope with these changes, semiconductor devices and systems have also been developed with new technologies, including computer architecture, processor and memory design, network security, and communication interfaces. Among these developments, the importance of processor architecture is increasing, especially in applications that require high throughput, low power consumption, and small physical space, such as mobile devices.
[0003] Among advanced processor architecture designs, instruction pipeline structures are widely used in many processing applications, including multithreading and parallel processing. As the demand for high-performance computing increases, designing efficient processor architectures faces an ever-growing number of challenges. Architectural register dependencies, out-of-order and in-order execution, circuit complexity, inefficient compiler techniques, and the complexity of communication and interfaces in multiprocessor environments all contribute to numerous problems in instruction pipeline design. These problems become even more pronounced in multithreaded processing systems. Compilers tend to generate inefficient code when resolving register race conditions, failing to utilize the internal structure of the microarchitecture. Hardware solutions tend to be overly complex, resulting in large silicon area and making them unsuitable for high-performance computing. This is the challenge. [Overview of the project] [Problems that the invention aims to solve]
[0004] The present invention has been made in view of the problems in the conventional multithreaded processing system described above, and the object of the present invention is to provide a system and method that can provide an efficient structure for resolving conflicts in register usage with respect to register renaming technology in a microarchitecture. This technology includes a hardware implementation of a circuit that performs register renaming at runtime as instructions pass through the instruction pipeline of a processing element (PE). [Means for solving the problem]
[0005] To achieve the above objective, the present invention provides an apparatus (register renaming circuit) comprising: an offset acquisition unit configured to acquire a first offset and a second offset, respectively, based on a thread identifier that identifies at least one of a first thread or a second thread executed on a processing element (PE), in accordance with the use of a first register and a second register; and an address pointer configured to generate at least one of a first register address or a second register address, respectively, based on at least one of the first offset or the second offset, wherein the first register address and the second register address correspond to a first operand and a second operand stored in a register file, respectively, and the first thread and the second thread include a first decoded instruction and a second decoded instruction that operate on the first operand and the second operand, respectively. [Effects of the Invention]
[0006] According to the apparatus, method, and system of the present invention, a hardware implementation of a circuit that performs register renaming at runtime as an instruction passes through the instruction pipeline of a processing element (PE) is included, thereby providing an efficient structure for resolving conflicts in register usage through register renaming. [Brief explanation of the drawing]
[0007] [Figure 1] This block diagram shows a schematic configuration of a system according to an embodiment of the present invention. [Figure 2] This is a block diagram showing a schematic configuration of the execution circuit in a PE according to an embodiment of the present invention. [Figure 3] This figure shows a multithreaded program according to an embodiment of the present invention. [Figure 4] This figure illustrates the simultaneous execution of a multithreaded program according to an embodiment of the present invention. [Figure 5] This is a diagram illustrating a register renaming circuit according to an embodiment of the present invention. [Figure 6] This is a flowchart illustrating the register renaming process using offset according to an embodiment of the present invention. [Modes for carrying out the invention]
[0008] Next, specific examples of embodiments for implementing the apparatus, method, and system according to the present invention relating to register renaming based on thread offset in a multithreaded processing system will be described with reference to the drawings.
[0009] The following detailed description includes many specific details to provide a complete understanding of this disclosure. However, those skilled in the art will understand that the disclosed embodiments can be implemented even without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the subject matter disclosed herein. Throughout this specification, the phrase "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment can be included in at least one embodiment disclosed herein. Thus, the phrases "in one embodiment", "in an embodiment", "according to one embodiment", or other expressions of similar import appear in various places throughout this specification, but they do not necessarily all refer to the same embodiment.
[0010] Furthermore, particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In this regard, the term "exemplary" as used herein means "serving as an example, instance, or illustration". Embodiments described herein as "exemplary" are not necessarily to be construed as preferred or advantageous over other embodiments. Furthermore, particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Also, depending on the context of the discussion herein, singular terms may include the corresponding plural forms, and plural terms may include the corresponding singular forms. Similarly, terms that are hyphenated (e.g., "two-dimensional", "pre-determined", "pixel-specific", etc.) may be used interchangeably with corresponding terms that are not hyphenated (e.g., "two dimensional", "predetermined", "pixel specific", etc.) in some cases, and descriptions using capital letters (e.g., "Counter Clock", "Row Select", "PIXOUT", etc.) may be used interchangeably with corresponding descriptions using non-capital letters (e.g., "counter clock", "row select", "pixout", etc.). Such occasionally occurring alternative usages are not considered to be conflicting with each other.
[0011] Also, depending on the context of the discussion in this specification, singular terms may include the corresponding plurals, and plural terms may include the corresponding singulars. Furthermore, it should be noted that the various figures (including component diagrams) shown and discussed in this specification are for illustrative purposes only and are not drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Additionally, where appropriate, reference numbers are repeated between figures to indicate corresponding elements and / or similar elements. The terms used in this specification are for the purpose of describing only some exemplary embodiments and are not intended to limit the claimed subject matter. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that, when used herein, the terms “comprises” and / or “comprising” identify the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0012] When an element or layer is said to be on another element or layer, or to be "connected to" or "coupled to" another element or layer, it will be understood that the element or layer may be directly on another element or layer, or to be connected to or coupled to another element or layer, or there may be an intervening element or layer. In contrast, when an element is described as being "directly on," "directly connected to," or "directly coupled to" another element or layer, there is no intervening element or layer. The same number refers to the same element throughout. As used herein, the term "and / or" includes any combination of one or more of the related enumerated items. As used herein, terms such as “first,” “second,” etc., are used as labels for the nouns that precede them and do not imply any kind of order (e.g., spatial, temporal, logical, etc.) unless explicitly defined as such.
[0013] Furthermore, the same reference number may be used across two or more figures to refer to parts, components, blocks, circuits, units, or modules that have the same or similar functions. However, such usage is solely for the sake of simplifying explanation and ease of discussion, and does not imply that the structural or architectural details of such components or units are the same across all embodiments, or that such commonly referred parts / modules are the only way to carry out some of the exemplary embodiments disclosed herein. Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as those generally understood by those skilled in the art in which this subject matter pertains. Furthermore, terms as defined in commonly used dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant technology, and should not be interpreted in an idealized or overly formal sense unless explicitly defined otherwise herein.
[0014] As used herein, the term “module” refers to any combination of software, firmware, and / or hardware configured to provide the functions described herein in relation to the module. For example, software may be embodied as a software package, code and / or instruction set, or instructions, and the term “hardware” as used in any implementation described herein may include, for example, assemblies, hardwired circuits, programmable circuits, state machine circuits, and / or firmware that stores instructions executed by programmable circuits, either alone or in any combination. Modules can be embodied collectively or individually as circuits that form part of a larger system, such as, but are not limited to, integrated circuits (ICs), systems-on-a-chip (SoCs), or assemblies.
[0015] As used herein, the term “solid-state” in the context of storage refers to storage technology that uses integrated circuits to store data instead of moving parts (e.g., rotating disks, platters, read / write heads, etc.). "Flash memory" refers to a type of non-volatile memory that retains data even when the power supply is cut off. This is commonly used in solid-state drives (SSDs). There are two types of flash memory. NAND flash memory and NOR flash memory. NAND flash memory is suitable for SSDs and mobile applications because it has high storage density and low cost per bit. NOR flash memory is optimized for random access and is often used in applications that require high-speed code execution.
[0016] As used herein, the term “buffer” in the context of storage refers to a memory device that temporarily stores data or information as part of an operation to move data from one location to another. Buffers are typically implemented using static random-access memory (SRAM) for high-speed access. The buffer can be configured as standard SRAM or in a first-in, first-out (FIFO) configuration.
[0017] In embodiments of the present invention, a register renaming technique is disclosed. This technology provides efficient instruction pipeline processing in the microarchitecture of multithreaded processing elements (PEs) in systems that use multiple processing elements (PEs). This technology offers several advantages, including fast resolution of register contention, a simple design, and efficient use of registers in the register file. In one embodiment of the present invention, the compiler compiles a multithreaded program and provides a thread identifier and a decoded instruction. The register renaming circuit includes a table, an offset acquisition unit, and an address pointer.
[0018] A table is created containing the offset for each thread, based on the number of threads being executed and the size of the register file. The decoded instruction indicates register usage, which includes references to registers. The offset acquisition unit is configured to acquire the first offset and the second offset, respectively, according to the use of the first register and the use of the second register, based on a thread identifier that identifies at least one of the first thread or the second thread. The first and second threads are executed on the Processing Element (PE). The address pointer is configured to generate at least one of the first register address or the second register address based on at least one of the first offset or the second offset, respectively. The first register address and the second register address correspond to the first and second operands stored in the register file, respectively. The first thread and the second thread each include a first decoded instruction and a second decoded instruction that operate on the first operand and the second operand, respectively.
[0019] Figure 1 is a block diagram showing a schematic configuration of system 100 according to an embodiment of the present invention. System 100 can be implemented as one or more system-on-chip (SoC) packages that include high-density devices such as three-dimensional (3D) packages. System 100 includes a host processor 110, an input / output (I / O) controller 140, a network interface card (NIC) 146, a graphics display controller (GDC) 150, a bus 155, a memory controller 160, and multiple processing elements 170. k Includes (k=1...N).
[0020] These components may interface with, or include, other components described further below. System 100 may include the above components in an increased or decreased manner. Furthermore, one component may be integrated with another. For example, the input / output controller 140, GDC 150, and memory controller 160 can be integrated into a single module. Integration may be carried out partially or overlappingly. For example, the GDC150 may be integrated into the processor 110, and the input / output controller 140 and the memory controller 160 may be integrated into a single controller.
[0021] System 100 serves as an example of the role of high-bandwidth memory (HBM) circuits in a high-performance computing (HC) platform. Many HC platforms may utilize multiple HBM circuits, such as stacked dynamic random access memory (DRAM) that operate in conjunction with processing units and I / O circuits. In many cases, the application environment imposes additional criteria such as low power consumption, reliable signal integrity, fault tolerance, and reliable operation under extreme conditions, including high temperatures and confined spaces. Other applications that benefit from highly integrated HBM designs include mobile communications (smartphones, base stations, user equipment, etc.), cameras, automobiles, entertainment (games, multimedia, music, movies, etc.), technical design (animation, graphics, etc.), medical (visualization, medical image processing, etc.), robotics, drones, automated testing equipment, speech processing, speech synthesis, video and image analysis, visual processing, automated facial recognition, artificial intelligence (AI) applications, and data centers.
[0022] The host processor 110 is a programmable device capable of executing a program or set of instructions to perform a task. The processor 110 may be a general-purpose processor, a digital signal processor, a microcontroller, a neural processing unit (NPU), or a specially designed processor such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The processor 110 may include a single core or multiple cores. Each core has multi-way multithreading. Processor 110 has a simultaneous multithreading function to further utilize the parallelism provided by multiple threads spanning multiple cores. The host processor 110 includes a memory management circuit (MMC) 120 and a processing management circuit (PMC) 130. The host processor 110 may include additional or subtractive of the above components. Bus 155 connects processor 110 to multiple PE170 k This could be any suitable bus connecting to other devices, including (k=1…N). Bus 155 may also be a Direct Media Interface (DMI).
[0023] The input / output controller 140 controls the mass storage device 142 and the input / output device 144. The mass storage device 142 may include a CD-ROM, a hard disk, and a solid-state drive (SSD). Input devices may include styluses, keyboards, mice, microphones, and image sensors. Output devices may include audio equipment, speakers, scanners, and printers. The network interface card (NIC) 146 provides an interface to the network 148, for example, via a wireless medium.
[0024] The memory controller 160 may be an extension of the MMC120. The memory controller 160 controls memory devices such as memory 162, HBM 164, and non-volatile memory (NVM) 166. Memory 162 may include static random access memory (SRAM) and dynamic random access memory (DRAM). The HBM164 includes a three-dimensional stack of memory dies to provide high bandwidth, low latency, low power consumption, and large storage capacity. The HBM164 features processing-in-memory (PIM) functionality. The NVM166 may include read-only memory (ROM), flash memory, Wide-IO NAND, MRAM, and / or other types of memory. Memory 162 stores instructions or programs loaded from NVM 166 or mass storage device 142, and these are then processed by processor 110 or either PE 170 k When executed by, processor 110 or either PE170 k This causes the system to perform the operations described in various embodiments. It also stores data used in its operation. The NVM166 contains instructions, programs, constants, or data that are retained regardless of whether power is on or off. The instruction or program corresponds to the functions described below.
[0025] In one embodiment, the program is one of the PE170 k Includes a compiler that compiles programs to be executed. This compiler is executed by either the host processor 100 or a processor connected to bus 155. The GDC150 controls the display device 152 and provides graphical processing. The GDC150 may be integrated inside the processor 110. The GDC150 typically features a graphical user interface (GUI) that allows for user interaction, enabling commands to be sent and functions to be activated. Additional devices or bus interfaces may be used for interconnection and / or expansion. Examples include the Peripheral Component Interconnect Express (PCIe) bus and the Universal Serial Bus (USB).
[0026] Host processor 110 and multiple processing elements 170 k It maintains a tight communication interface, at least via bus 155 and other independent lines. Multiple PE170 k The cluster consisting of these components operates under the control and management of the host processor 110. Once activated and started, each PE170 executes its own program and accesses data in the private instruction memory 192 and data memory 194. The host processor 110 provides a layer of abstraction for the entire architecture. In short, the host processor 110 hides the complexity of program execution from the user or high-level applications. The application program may specify what it should do, and the host processor 120 handles the details of how to execute it by assigning the task to individual PEs.
[0027] The host processor 110 includes a memory management circuit (MMC) 120 and a process management circuit (PMC) 130. Each PE 170 k (k = 1…N) includes an L1 cache 172 k , a configuration (CFG) circuit 174 k , an execution circuit 180 k , an interrupt circuit 182 k , an instruction memory 192 k , a data memory 194 k , a calculation circuit 196 k , and a communication interface 198. k The host processor 110 and the PE 170k may include more or fewer of the above components. Hereinafter, for clarity, the index k attached to the PE 170 k and its related elements may be omitted.
[0028] Based on the memory access by at least one of the PMC 130 and the PE 170, the MMC 120 is configured to operate using the memory controller 160 or without using the memory controller 160 to manage memory operations for at least one of the L2 cache 125, the main memory 134, the L1 cache 172 in the PE 170, the memory 162, the HBM 164, or the NVM 166. The memory operations include at least one of read access, write access, page table update, translation lookaside buffer (TLB) update, cache response, and access violation response.
[0029] The L2 cache 125 is configured to function as a translation lookaside buffer (TLB) for converting virtual memory to physical memory. The L2 cache 125 is typically implemented by a high-speed memory such as high-speed SRAM, enabling the MMC 120 to quickly obtain page mapping from virtual memory to physical memory without accessing a slow page table. The L2 cache 125 can be used as cache storage to achieve fast response times to memory accesses. The MMC120 updates the page table in memory 162 or the TLB in L2 cache 125 if there are new entries in the table. The MMC120 responds to access violations such as non-existent memory addresses, buffer overflows, and null pointers. The MMC120 reports any violations to a test controller (not shown) for debugging or testing purposes.
[0030] The PMC130 includes a main execution circuit 132, a main memory 134, and an interrupt controller 136. It is configured to manage at least one processor operation performed by at least one PE170. Processor operations include at least one of the following: program initiation, program execution, and interrupt delivery. The main execution circuit 132 is a processing unit or circuit capable of executing a program or instruction stored in the main memory 134. The program is any suitable program. In one embodiment, the program is a compiler that compiles programs to be executed on multiple PE170s.
[0031] Main memory 134 is private memory for the PMC130. Main memory 134 can be any suitable type of memory, such as DRAM, SRAM, MRAM (Magnetoresistive Random-Access Memory), flash, or a combination thereof. Main memory 134 contains a page table for converting virtual pages into physical pages as part of the memory management tasks performed by MMC120. The main execution circuit 132 accesses memory 162, HBM 164, and NVM 166 via the memory controller 160. The interrupt controller 136 controls and manages interrupt requests from PE170 and interrupt services to PE170. This control and management includes prioritizing interrupt requests and sending commands or messages to PE170.
[0032] Each PE170 is configured to operate independently or in cooperation with other PE170s and the host processor 110. The PE170 and the host processor 110 together form a multiprocessor system and operate in coordination to operate in parallel or sequentially based on the overall system objectives. In each PE170, the execution circuit 180 is configured as a circuit capable of executing a program, instruction, or command stored in the instruction memory 192. The execution circuit 180 is configured to interface with or communicate with the instruction memory 192, data memory 194, calculation circuit 196, communication interface 198, and memory controller 160 via the bus 190.
[0033] The execution circuit 180 accesses memory 162, HBM 164, and NVM 166 via the memory controller 160. In some embodiments, the execution circuit 180 includes an instruction pipeline that processes instructions from the instruction memory 192. The instruction pipeline in the execution circuit 180 is explained in Figure 2. The execution circuit 180 accesses the data stored in the data memory 194. Data memory 194 is used to store temporary data for program execution and data structures such as the stack or heap. The instruction memory 192 and data memory 194 are private or local to the associated PE and are implemented by any suitable memory, including DRAM, SRAM, MRAM, flash, or a combination thereof. The calculation circuit 196 is configured to perform logical operations and / or calculation processes. The computational circuit 196 may include multiple functional units, tensor units, mathematical units, and buffers and interconnects. These computing units are scheduled by a PE scheduler (not shown). The PE scheduler is configured by the host processor 110.
[0034] The communication interface 198 provides an interface for communication between PEs and between the associated PEs and the host processor 110. The L1 cache 172 provides high-speed cache memory to the execution circuit 180. L1 cache 172 can be used to implement the TLB for address translation. The L1 cache 172 is connected to the L2 cache 125 of the host processor 110 for additional caching operations. Each PE's L1 cache 172 is enabled to communicate with the L2 cache 125, allowing information to be shared between PEs.
[0035] The interrupt circuit 182 provides services related to interrupt requests and responses between PEs for inter-processor interrupts (IPI), as well as interrupt requests and responses between PEs and the host processor 110. The interrupt circuit 182 generates an IPI to another PE and receives an IPI response from the other PE. Before requesting an interrupt, the PE preloads data and status into memory 162, so that other PEs can retrieve that data when handling interrupts. The interrupt circuit 182 generates an interrupt in the main execution circuit 132 via the interrupt controller 136 when the PE requests a service or reports a status. For example, when the PE completes the task currently assigned to it, it sends an interrupt to the main execution circuit 132. Before sending an interrupt, the PE sends a message, result, data, status, or state to memory 162 or HBM164 so that the main execution circuit 132 can check the message when responding to the interrupt. This enables an efficient communication protocol between the PE and the host processor 110. The CFG circuit 174 includes CFG data that configures the PE170 to perform operations or calculations as needed. Furthermore, the CFG circuit 174 enables or disables the PE under the control of the host processor 110.
[0036] Figure 2 shows a schematic configuration of the execution circuit 180 shown in Figure 1 within the PE170 according to an embodiment of the present invention. The execution circuit 180 includes an instruction buffer 210, an instruction pipeline 220, a register buffer 260, and a register file 270. The execution circuit 180 may include more or fewer components than those described above.
[0037] The instruction buffer 210 stores instructions from the instruction memory 192 and queues them for supply to the instruction pipeline 220. The instruction buffer 210 includes small buffers that are inserted between stages of the pipeline 220 to keep the instruction flow moving. The instruction buffer 210 buffers in-order instructions, out-of-order instructions, or predictive branch instructions by corresponding circuits (not shown). The instruction pipeline 220 includes multiple stages for preparing instructions to be executed in the program flow. The stages of the instruction pipeline 220 include a fetch stage 222, a decode stage 226, a register renaming stage 230, an issue stage 234, an execution stage 238, a memory stage 242, a write-back stage 246, and a retirement stage 250. Pipeline 220 may include more or fewer stages than those described above. Furthermore, not all stages are active for all instructions.
[0038] The fetch stage 222 retrieves instructions from the instruction buffer 210 or directly from the instruction memory 192. The program counter (not shown) tracks the address of an instruction. Fetched instructions are typically held in the instruction register, ready to be decoded. The fetch stage includes branch prediction, out-of-order execution, reorder buffers, false branch handling, and other instruction fetch mechanisms to provide a smooth flow of instructions through the pipeline. Decode stage 226 includes decoding, interpreting, and translating the binary representation of the instruction into an understandable and executable portion. This involves separating instructions into opcodes and operands and determining the actions to be performed. The result of the decode stage 226 includes the decoded instruction, which is further examined to determine whether there is any contention in register usage during the execution of that instruction. Decode stage 226 includes branching to the microcode corresponding to the instruction.
[0039] The register renaming stage 230 resolves register usage conflicts in the decoded instruction. The register renaming stage 230 eliminates false data dependencies, particularly write-after-write (WAW) and write-after-read (WAR) hazards. The register renaming stage 230 maps logical register names or architecture register names to physical register names in the processor's internal register file. In this process, instructions are dynamically updated as they pass through the pipeline, ensuring that operands are correctly read and written without conflicts. The register renaming stage 230 includes a register renaming circuit 235, which will be further described in Figure 3.
[0040] Issue stage (or issuing stage) 234 prepares for the execution of the instruction. The issuing stage 234 includes selecting and scheduling instructions, taking into account factors such as program order and dependency analysis. The issue stage 234 includes allocating resources (e.g., functional units, memory access) and obtaining operands. Issuance methods include in-order and out-of-order methods. The execution stage 238 executes the instructions that were issued and prepared by the issue stage 234. The execution stage 238 uses the calculation circuit 196 (Figure 1) to perform executions including arithmetic operations, logical operations, and other operations provided by the functional unit. Execution stage 238 includes retrieving operands from register file 270, calculating memory addresses, and evaluating branch predictions. Execution stage 238 includes retrieving operands from memory stage 242, which accesses L1 cache 172 and bus 190.
[0041] The memory stage 242 retrieves operands from memory or writes data to memory. Examples of instructions that are allowed to access memory include load and store instructions. If the instruction does not access memory, memory stage 242 is bypassed. The write-back stage 246 writes the results of the execution stage 238 back to the register file 270. The write backstage 246 writes data to the register buffer 260, and the register buffer 260 sends the data to the register file 270 when it is ready. In response to an instruction, the write-back stage 246 may obtain the data to be written back directly from the calculation circuit 196, or it may obtain it from the memory stage 242. Retirement Stage 250 confirms the execution of the said order. Retirement Stage 250 can be integrated with Lightback Stage 246. Retirement Stage 250 includes handling exceptions, freeing up resources, and other housekeeping functions.
[0042] Figure 3 shows a multithreaded program 300 according to an embodiment of the present invention. The multithreaded program 300 includes a main program 310 and K threads 320j (j==1...K).
[0043] The main program 310 is a loop that repeats the loop body N times. Each iteration involves the following operation, where "%r" refers to an architecture register or a logical register. 20:load%r1,a[i] means r1←a[i] 30:load%r2,b[i] means r2←b[i] 40:mult%r3,%r1,%r2 means r3 ← r1 × r2 50:add%r3,%r4 means r3←r3+r4
[0044] As shown above, each iteration is independent of the others. Furthermore, within each iteration, there are no dependencies between elements. All N iterations are independent of each other. In other words, iteration j does not require the result of any other iteration k such that "j≠k". Since there are no dependencies, these iterations can be executed in parallel. In a truly parallel environment, these iterations are assigned to multiple physical processors and executed in parallel. In a multithreaded programming environment, this means that program 310 may be broken down into multiple threads. Now, let's assume there are K threads. The main program 310 is divided into N / K threads. N / K threads are allowed to run simultaneously. Each thread performs N / K iterations, changing the index value each time.
[0045] Program 3201 is assigned to thread 1, which executes a loop from i=1 to N / K. Program 320p is assigned to thread p, which executes a loop from i=(p-1)N / K+1 to pN / K. Program 320 K This is assigned to thread K, which executes a loop from i=(K-1)N / K up to N. These are independent, but each iteration involves access to the registers (r1, r2, r3, r4). If the iterations are executed sequentially within a single thread, no register usage contention occurs. However, if these operations are performed simultaneously, a register usage conflict occurs, and therefore, register renaming is used to resolve this conflict.
[0046] Figure 4 is a diagram illustrating the simultaneous execution 400 of a multithreaded program according to an embodiment of the present invention. Concurrent execution 400 executes two threads, 410 and 420, simultaneously. Concurrent execution of threads within a single processor element means that the processor switches execution resources such as memory, registers, and compute units between threads. Concurrent execution may appear to be parallel execution, but in reality, only one thread is running at any given time. Thread switching is extremely fast, making concurrent execution appear parallel. The two threads, 410 and 420, are executed on the timeline.
[0047] Thread 1 (410) has the following instructions: 21:load%r1,a[i];r1←a[i] 31:load%r2,b[i]:r2←b[i]) 41:mult,%3,%1,%2:r3←r1×r2)
[0048] The runtime index i in thread 1 is i=263. The instructions (21, 31, 41) are executed in their respective execution windows (412, 414, 416). For simplicity, let's assume that each of these execution windows completes the execution of the instruction in question.
[0049] Thread 2 (420) has the following instructions: 22:load%r1,a[i];r1←a[i] 32:load%r2,b[i]:r2←b[i] 42:mult,%3,%1,%2:r3←r1×r2
[0050] The runtime index i in thread 2 is i=471. The commands (22, 32, 42) are executed in their respective execution windows (422, 424, 428). For simplicity, let's assume that each of these execution windows completes the execution of the instruction in question.
[0051] In concurrent execution, only one execution is active at any given time. Between t0 and t1, execution window 412 is active. The contents of a
[0263] are loaded into register r1. Between t1 and t2, execution window 422 is active. The contents of a
[0471] are loaded into register r1. Therefore, the previous contents of a
[0263] in register r1 are overwritten and destroyed.
[0052] Between t2 and t3, execution window 414 is active. The contents of b
[0263] are loaded into register r2. Between t3 and t4, execution window 424 is active. The contents of b
[0471] are loaded into register r2. Therefore, the previous contents of b
[0263] in register r2 are overwritten and destroyed. Between t4 and t5, execution window 416 is active. Register r3 is currently loaded with the product of registers r1 and r2, which include a
[0471] and b
[0471] . Therefore, register r3 is loaded with a
[0471] × b
[0471] . This is incorrect because it should actually be the product of a
[0263] and b
[0263] . Between t5 and t6, execution window 426 is active. Currently, register r3 is loaded with the product of registers r1 and r2, which include a
[0471] and b
[0471] . Therefore, register r3 is loaded with a
[0471] × b
[0471] . This will yield the correct result, however, if a subsequent instruction in thread 1 uses the result of the product a
[0263] × b
[0263] , it will yield an incorrect result.
[0053] The above example demonstrates that simultaneous execution of multiple threads accessing the same register can lead to unstable, inaccurate, or unpredictable results. One solution to this problem is to allow each thread to access its own register file. In other words, the physical register file is statically partitioned, and partitions (divided areas) are allocated to multiple threads.
[0054] Figure 5 is a diagram illustrating the register renaming circuit 235 shown in Figure 2 according to an embodiment of the present invention. The register renaming circuit 235 includes an offset table 510, an offset acquisition unit 520, and an address pointer 530. The register renaming circuit 235 may include more or fewer components than those described above. The register renaming circuit 235 operates with the assistance of a compiler 501 that compiles instructions and generates multithreaded decoded instructions.
[0055] Compilers are designed for multithreaded programming. The compiler is configured to allow the user to specify the number of threads. During execution, the compiler generates a thread identifier 515 for the currently running thread and a decoded instruction 525 for that active thread. The register renaming circuit 235 uses this information to divide the register file 270 and assign the divided register files to each thread according to the thread identifier. This is achieved by using an offset assigned to each thread. Each thread is assigned a partitioned register file using an offset that points to a register. The offset is determined before the execution of a multithreaded program and is maintained throughout the execution of the entire program.
[0056] For example, suppose register file 270 has N=512 registers, and these are assigned names from r1 to r512. Assume the number of threads is K=4. To maximize partition size, the entire register file 270 is divided into four partitions, each assigned to a thread (1, 2, 3, and 4). Therefore, each partition has a total of N / K registers, i.e., 512 / 4 = 128 registers. Thread 1 (541) is assigned to the partition containing registers 1 through 128. Thread 2 (542) is allocated to the partition containing registers 129 through 256. Thread 3 (543) is assigned to the partition containing registers 257 through 384. Thread 4 (544) is allocated to the partition containing registers 385 through 512. Each thread has its own small register file, so register contention between multiple threads does not occur. This partitioning provides the basis for the offsets stored in the offset table, and these offsets are equal to the first register in the partition.
[0057] The offset is stored in table 510. This table can be implemented as high-speed SRAM, similar to a cache. This table is used as a lookup table to generate offsets that are indexed by thread identifiers. For example, if N=512 and K=4, the table will store offset values (0, 128, 256, 384) for each thread ID (1, 2, 3, 4). The offset is the starting register number or address of the partition. By allocating a separate partition to each thread, register contention does not occur. Architectural registers are converted to physical registers by adding an offset to the register address.
[0058] Decoded instructions have register usage that references registers. Since the register needs to be renamed, the offset is retrieved from table 510. For N threads, N offsets are obtained. Assume that two threads are running on the PE, and that these two threads have two decoded instructions. Decoded instructions manipulate operands stored in a register file. Each decoded instruction has register usage. Table 510 is configured to store the first offset and the second offset based on the usage of the first register and the second register, respectively.
[0059] The offset acquisition unit 520 is configured to acquire the first offset and the second offset according to the use of the first register and the use of the second register, respectively, based on a thread identifier that identifies at least one of the first thread or the second thread. Once the offset is obtained, it is used to rename the register or to calculate the register's address. Each address pointer 530 is configured to generate at least one of a first register address or a second register address based on at least one of a first offset or a second offset. The first register address and the second register address correspond to the first operand and the second operand stored in the register file 270, respectively.
[0060] The register address or renamed register is determined by adding an offset to the register name in the decoded instruction, as follows: "Renaming register = Decoded instruction register + Thread offset" ... (1)
[0061] For example, let's assume that table 510 is configured for a 4-thread program. Two instructions are executed on two separate threads. Instruction 1 is "loadr5,a" in thread 1, and instruction 2 is "loadr73,c" in thread 2. When instruction 1 is decoded, its thread identifier (ID=1) is determined. The offset acquisition unit 520 acquires the offset from the offset table 810. This offset is 0. Therefore, an address pointer adds an offset to the register address or register number. The renamed register is "r(5+0)=r5". When instruction 2 is decoded, its thread identifier (ID=2) is determined. The offset acquisition unit 520 acquires the offset from the offset table 510. This offset is 128. Therefore, the address pointer adds an offset to the register address or register number in the decoded instruction. The renamed register is "r(73+128)=r191".
[0062] Block 550 shows an example of a running 4-thread program. The offset for thread 1 is 0. Thread 1 executes instructions in three execution windows (561, 569, 577) that reference the architecture registers (0, 1, 3), respectively. The offset for thread 2 is 128. Thread 2 executes instructions in three execution windows (567, 573, 581) that reference the architecture registers (2, 1, 0), respectively. The offset for thread 3 is 256. Thread 3 executes instructions in two execution windows (563, 575) that reference architecture registers (0, 18), respectively. The offset for thread 4 is 384. Thread 4 executes instructions in three execution windows (565, 571, 579) that reference the architecture registers (0, 4, 1), respectively.
[0063] Subsequently, the registers are renamed, that is, the register addresses for the physical registers in register file 270 are generated by address pointer 530 as follows: Thread 1: r0→r0, r1→r1, r3→r3 Thread 2: r2 → r130 (= 2 + 128), r1 → r129 (= 1 + 128), r0 → r128 (= 0 + 128) Thread 3: r0 → r256 (= 0 + 256), r18 → r274 (= 18 + 256) Thread 4: r0→r384 (=0+384), r4→r388 (=4+384), r1→r385 (=1+384)
[0064] In the example above, the four threads may refer to the same architectural register, but the renamed register will point to a different physical register. Therefore, register contention does not occur.
[0065] Figure 6 is a flowchart illustrating the register renaming process 600 using an offset according to an embodiment of the present invention. When "started," the register renaming process 600 stores the first offset and the second offset in the table, respectively, based on the use of the first register and the use of the second register (step S610). The first or second offset is determined based on the size of the register file and the number of threads running in the processing element (PE). Next, the register renaming process 600 obtains the first offset and the second offset according to the first register usage and the second register usage, respectively, based on a thread identifier that identifies at least one of the first thread or the second thread. (Step S620) The first and second threads are executed on the Processing Element (PE).
[0066] Next, the register renaming process 600 generates at least one of the first register address or the second register address based on at least one of the first offset or the second offset (step S630). The first and second register addresses are renamed registers and correspond to physical registers in the register file. The first register address and the second register address correspond to the first and second operands stored in the register file, respectively. The first thread and the second thread each include a first decoded instruction and a second decoded instruction that operate on the first and second operands, respectively. After that, the register renaming process 600 is completed.
[0067] All or part of the embodiments of the present invention can be implemented by various means according to the application, depending on the specific features and functions. These means may include hardware, software, firmware, or any combination thereof. Hardware, software, or firmware elements may have multiple modules that are coupled together. Hardware modules are coupled to other modules by mechanical, electrical, optical, electromagnetic, or other physical connections. Software modules are linked to other modules through functions, procedures, methods, subprograms, subroutine calls, jumps, links, parameter, variable, argument passing, function returns, and so on. Software modules are linked to other modules to receive variables, parameters, arguments, pointers, etc., and / or to generate or pass resulting, updated variables, pointers, etc. The firmware module is coupled to other modules by any combination of the hardware coupling method and software coupling method described above. A hardware module, software module, or firmware module may be coupled to another hardware module, software module, or firmware module. A module can also be a software driver or interface for interacting with the operating system running on the platform. A module can also be a hardware driver for configuring, setting up, and initializing hardware devices, and for sending and receiving data to and from those hardware devices. The device may include any combination of hardware modules, software modules, and firmware modules.
[0068] The embodiments and operations of the present invention described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed herein and their structural equivalents), or in one or more combinations thereof. Embodiments of the present invention described herein can be implemented as one or more computer programs, that is, as one or more modules of computer program instructions encoded on a computer storage medium for executing or controlling the operation of a data processing device. Furthermore, or alternatively, program instructions can be encoded into artificially generated propagation signals, such as mechanically generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiving device for execution by a data processing device.
[0069] Computer storage media may be computer-readable storage devices, computer-readable storage boards, memory arrays or devices for random access or serial access, or combinations thereof, or may be included in such devices. Furthermore, although computer storage media are not propagating signals themselves, they can be sources or destinations for computer program instructions encoded with artificially generated propagating signals. Computer storage media can also be, or be comprised of, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices). Furthermore, the operations described herein can be performed as operations executed by a data processing device on data stored in one or more computer-readable storage devices, or on data received from other sources.
[0070] This specification contains many specific implementation details, but these implementation details should not be interpreted as limitations on the claimed scope of the invention, but rather as descriptions of features specific to particular embodiments. Certain features described herein in the context of a separate embodiment may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any appropriate subordinate combination. Furthermore, while features are described above as acting in certain combinations, and may even be initially claimed in this manner, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be related to a lower-level combination or a variation of a lower-level combination.
[0071] Similarly, although operations are depicted in a specific order in the drawings, this should not be understood as requiring that such operations be performed in a specific illustrated order or sequential order, or that all illustrated operations be performed, in order to achieve a preferred result. In some situations, multitasking or parallel processing can be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as necessary in all embodiments, and the described program components and systems can generally be integrated in a single software product or packaged in multiple software products.
[0072] Thus, specific embodiments of the present invention have been described herein. Other embodiments are within the scope of the following claims. In some cases, favorable results can still be obtained by performing the operations described in the claims in a different order. Furthermore, the steps shown in the attached diagram do not necessarily require the specific order or sequence indicated to obtain favorable results. In certain implementations, multitasking and parallel processing can be advantageous. As will be apparent to those skilled in the art, the innovative concepts described herein are modifiable and adaptable to a wide range of applications. Therefore, the scope of the claimed subject matter should not be limited to the specific exemplary teachings described above, but rather defined by the following claims. [Explanation of Symbols]
[0073] 100 Systems 120 Memory Management Circuit (MMC) 125 L2 Cache 130 Processing management circuit (PMC) 132 Main execution circuit 134 Main Memory 136 Interrupt Controller 140 Input / Output Controllers 142 Mass storage 144 Input / Output Devices 146 NIC (Network Interface Card) 148 Networks 150 GDC (Graphics Display Controller) 152 Display device 155 Bus 160 memory controllers 162 memory 164 HBM (High Bandwidth Memory) 166 NVM (Non-Volatile Memory) 1701…170 N PE (Processing element) 172, 1721L1 cache 1741CFG (Configuration Circuit) 180 Execution Circuit 182 Interrupt Circuits 190 bus 192 instruction memory 194 data memory 196 Calculation circuit 198 Communication Interface 210 Instruction Buffer 220 instruction pipeline 222 Fetch 226 Decode 230 Register Renaming 235 Register renaming circuit Issued 234 238 Execution 242 memory 246 Lightback 250 Retirement 260 register buffer 270 Register File
Claims
1. It is a device, An offset acquisition unit is configured to acquire a first offset and a second offset, respectively, based on a thread identifier that identifies at least one of a first thread or a second thread executed on a processing element (PE), in accordance with the use of a first register and a second register, respectively. The system includes an address pointer configured to generate at least one of a first register address or a second register address based on at least one of the first offset or the second offset, The first register address and the second register address correspond to the first operand and the second operand stored in the register file, respectively. The apparatus is characterized in that the first thread and the second thread each include a first decoded instruction and a second decoded instruction that operate on the first operand and the second operand, respectively.
2. The apparatus according to claim 1, characterized in that at least one of the first offset or the second offset is determined based on the size of the register file and the number of threads executed in the processing element (PE).
3. The apparatus according to claim 1, further comprising tables configured to store the first offset and the second offset, respectively, based on the use of the first register and the use of the second register.
4. The apparatus according to claim 3, characterized in that at least one of the use of the first register or the use of the second register is provided by the compiler.
5. The apparatus according to claim 1, characterized in that the first thread and the second thread are executed simultaneously within the processing element (PE).
6. The apparatus according to claim 1, characterized in that the use of the first register and the use of the second register result in a register contention.
7. The apparatus according to claim 6, characterized in that at least one of the first register use or the second register use maps an architecture register to a physical register in the register file.
8. The apparatus according to claim 7, characterized in that at least one of the first register use or the second register use maps an architecture register to a physical register based on the resolution of the register contention.
9. The apparatus according to claim 4, characterized in that the register file is statically divided based on the number of concurrent threads executed on the processing element (PE).
10. The apparatus according to claim 1, characterized in that the processing element (PE) is part of a PE cluster in a high-bandwidth memory (HBM) processing system.
11. It is a method, The steps include obtaining a first offset and a second offset according to the use of a first register and a second register, respectively, based on a thread identifier that identifies at least one of a first thread or a second thread running on a processing element (PE), The process includes the step of generating at least one of a first register address or a second register address based on at least one of the first offset or the second offset, The first register address and the second register address correspond to the first operand and the second operand stored in the register file, respectively. A method characterized in that the first thread and the second thread each include a first decoded instruction and a second decoded instruction that operate on the first operand and the second operand, respectively.
12. The method according to 11, characterized in that at least one of the first offset or the second offset is determined based on the size of the register file and the number of threads executed in the processing element (PE).
13. The method according to 11, further comprising the step of storing the first offset and the second offset in a table, respectively, based on the use of the first register and the use of the second register.
14. The method according to 13, characterized in that at least one of the use of the first register or the use of the second register is provided by the compiler.
15. The method according to 11, characterized in that the first thread and the second thread are executed simultaneously within the processing element (PE).
16. The method according to 11, characterized in that the use of the first register and the use of the second register result in a register contention.
17. The method according to 16, characterized in that at least one of the first register use or the second register use maps an architecture register to a physical register in the register file.
18. The method according to 17, characterized in that at least one of the first register use or the second register use maps an architecture register to a physical register based on the resolution of the register contention.
19. The method according to 14, characterized in that the register file is statically divided based on the number of concurrent threads executed on the processing element (PE).
20. It is a system, A management processor configured to manage processor and memory operations, The system comprises a processing element (PE) in a processing element (PE) cluster configured to be managed by the management processor, The processing element (PE) includes a register renaming circuit, The aforementioned register renaming circuit is A conflict detection circuit configured to detect register conflicts between a first decoded instruction and a second decoded instruction, Here, the register contention is associated with a first architecture register and a first physical register corresponding to the first architecture register. The system includes a mapping circuit configured to map the first architecture register to a second available physical register different from the first physical register, A system characterized in that the first decoded instruction and the second decoded instruction are decoded from a single thread.