Apparatus including a register renaming circuit for executing mixed instruction words with register collisions, and a register renaming method and system thereof.
The register renaming circuit addresses register collisions in processor architectures by dynamically mapping architecture registers to physical registers, enhancing performance and reducing hardware complexity for high-performance computing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-04-01
Smart Images

Figure 2026056600000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a computer architecture, and more particularly to an apparatus including a register renaming circuit for executing mixed instruction words having register collisions, a register renaming method, and a system thereof. [Background technology]
[0002] The technical section providing the background to this invention is for context only, and the disclosure of a concept in this section does not constitute prior art.
[0003] Advances in data science, artificial intelligence (AI), and machine learning (ML) are driving revolutionary technological changes across diverse industrial sectors. To accommodate these changes, semiconductor devices and systems incorporating new technologies in areas such as computing architecture, processors, memory design, network security, and communication interfaces are being developed. Among these advancements, processor architecture is becoming increasingly important, particularly in applications requiring high processing power, low power consumption, and small physical space, such as mobile devices.
[0004] Among the most advanced processor architecture designs, the instruction pipeline structure is widely used in many processing applications, including multithreading and parallel computing. As the demand for high-performance computing increases, designing efficient processor architectures faces an increasing number of challenges. Many problems have arisen in instruction pipeline design due to issues such as architectural register dependency, out-of-order and in-order execution, circuit complexity, inefficient compiler techniques, communication in multiprocessor environments, and interface complexity. Compilers tend to generate inefficient code when resolving register collisions and do not leverage the internal structure of the microarchitecture. Hardware solutions tend to be overly complex, resulting in large silicon area and making them unsuitable for high-performance computing.
[0005] The information disclosed in the technical section forming the background of this invention is intended to enhance understanding of the background of this invention and therefore includes information that does not constitute prior art. [Overview of the project] [Problems that the invention aims to solve]
[0006] The present invention has been made in view of the above-mentioned conventional problems, and the object of the present invention is to provide an apparatus including a register renaming circuit for executing a mixed instruction word that has register collisions, a register renaming method and a system thereof. [Means for solving the problem]
[0007] To overcome these problems, a system and method for register renaming in microarchitectures are described here. This technique aims to provide an efficient structure for resolving register usage conflicts. This technique includes a hardware implementation of a circuit that renames registers at runtime as the instruction word passes through the instruction pipeline of the processing element (PE).
[0008] An apparatus including a register renaming circuit according to one aspect of the present invention, made to achieve the above objective, comprises: a collision sensing circuit configured to detect a register collision between a first decoded instruction and a second decoded instruction; and a mapping circuit configured to change the first architecture register to a second architecture register and map the second architecture register to a second physical register different from the first physical register, wherein the register collision relates to the first architecture register and the first physical register corresponding to the first architecture register, and the first decoded instruction and the second decoded instruction are decoded in a single thread of a processing element (PE).
[0009] A register renaming method for a device including a register renaming circuit according to one aspect of the present invention, made to achieve the above objective, comprises the steps of: detecting a register collision between a first decoded instruction and a second decoded instruction; changing a first architecture register to a second architecture register; and mapping the second architecture register to a second physical register that is available differently from the first physical register, wherein the register collision relates to the first architecture register and the first physical register corresponding to the first architecture register, and the first decoded instruction and the second decoded instruction are decoded in a single thread of a processing element (PE).
[0010] A system including a register renaming circuit according to one aspect of the present invention, made to achieve the above objective, comprises a host processor configured to manage processor operations and memory operations, and a processing element of a processing element (PE) cluster configured to be managed by a management processor, wherein the processing element includes a collision sensing circuit configured to detect register collisions between a first decoded instruction and a second decoded instruction, and a mapping circuit configured to change a first architecture register to a second architecture register and map the second architecture register to a second physical register different from the first physical register, wherein the register collision relates to the first architecture register and the first physical register corresponding to the first architecture register, and the first decoded instruction and the second decoded instruction are decoded in a single thread of the processing element (PE). [Effects of the Invention]
[0011] According to the present invention, it is possible to provide an apparatus including a register renaming circuit for executing a mixed instruction word with register collisions having improved performance, as well as a register renaming method and system thereof. [Brief explanation of the drawing]
[0012] [Figure 1] Block diagram showing a system with a schematic representation. [Figure 2] This is a block diagram showing the execution circuit in PE shown in Figure 1, according to one embodiment of the present invention. [Figure 3] This is a diagram showing a register renaming circuit used in the register renaming stage of Figure 2 according to one embodiment of the present invention. [Figure 4] Figure 3 is a diagram showing the mapping circuit configuration. [Figure 5] This is a flowchart showing the process for register renaming according to one embodiment of the present invention. [Figure 6] A flowchart showing a process of mapping to the physical register shown in FIG. 5 according to an embodiment of the present invention.
Embodiments of the Invention
[0013] Hereinafter, specific examples of embodiments for carrying out the present invention will be described in detail with reference to the drawings.
[0014] In the following detailed description, numerous specific details are presented in order to provide a thorough understanding of the present invention. However, a person of ordinary skill in the art should understand that the disclosed aspects can be implemented without such specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the subject matter of the present invention.
[0015] Throughout this specification, the terms “a particular embodiment” or “one embodiment” mean that certain features, structures, or characteristics described in relation to the embodiment may be included in at least one embodiment of the present invention. Therefore, even if “a particular embodiment,” “one embodiment,” or “by one embodiment” (or other similar subsections) appear in various places throughout this specification, they do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, the word “exemplary” as used herein means “serving as an example, case, or illustration.” Embodiments described as “exemplary” in this specification should not be construed as necessarily preferable or advantageous to other embodiments. Also, certain features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Depending on the context discussed herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. Similarly, hyphenated terms (e.g., "two-dimensional," "pre-determined," "pixel-specific") can be used interchangeably with their non-hyphenated versions (e.g., "two-dimensional," "predetermined," "pixel-specific"), and capitalized terms (e.g., "Counter Clock," "Row Select," "PIXOUT") can be used interchangeably with their non-capitalized versions (e.g., "counter clock," "row select," "pixout"). Such interchangeable uses should not be considered contradictory.
[0016] Also, according to the context discussed in this specification, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. Also, it should be noted that the various drawings (including the drawings of components) illustrated and discussed in this specification are for illustrative purposes only and are not drawn to actual scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Also, reference numerals are repeated in the drawings to indicate corresponding and / or similar elements where appropriate.
[0017] The terms used in this specification are for the purpose of describing some exemplary embodiments only and are not intended to limit the claimed subject matter. The singular forms used in this specification are intended to include the plural forms as well, unless the context clearly dictates otherwise. The terms "comprising" and / or "comprises" used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0018] When an element or layer is referred to as being "on," "connected to," or "coupled to" another element or layer, it should be understood that this may mean directly on, connected to, or coupled to the other element or layer, or intervening elements or layers may be present. In contrast, when an element is referred to as being "immediately on," "directly connected to," or "directly coupled to" another element or layer, there are no other elements or layers therebetween. Like numbers refer to like elements throughout. The term "and / or" used in this specification includes any and all combinations of one or more of the associated listed items.
[0019] As used herein, terms such as “first,” “second,” etc., are used as labels for preceding nouns and do not imply any order of a type (e.g., spatial, temporal, logical, etc.) unless explicitly defined. Furthermore, the same reference numerals are used across two or more drawings to refer to parts, components, blocks, circuits, units, or modules having the same or similar function. However, such specifications are for the sake of simplification of explanation and for the convenience of discussion, and do not imply that the structure or structural details of such components or units are the same in all embodiments, or that a commonly referenced part / module is the only way to embody some of the exemplary embodiments of the present invention.
[0020] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as those generally understood by an ordinary technician in the technical field to which this subject pertains. Terms that are the same as those defined in commonly used dictionaries should be interpreted to have the same meaning in the context of the relevant technology and should not be interpreted in an idealized or overly formal sense unless expressly defined herein.
[0021] As used herein, the term “module” refers to any combination of software, firmware, and / or hardware configured to provide the functions described herein in relation to the module. For example, software may be embodied as a software package, code, and / or instruction set, or instruction; and as used in the embodiments described herein, the term “hardware” may include, for example, assemblies, wired circuits, programmable circuits, state-machine circuits, and / or firmware that stores instructions executed by programmable circuits, either alone or in any combination. Modules may be embodied in circuits that constitute part of a larger system, collectively or individually, including, but not limited to, integrated circuits (ICs), systems-on-a-chip (SoCs), assemblies, and the like.
[0022] As used herein, the term “solid-state” refers to storage technology that uses integrated circuits instead of moving parts (e.g., rotating disks, platters, read / write heads) to store data in relation to storage. The term “flash memory” refers to a type of non-volatile memory in which data is retained even when the power supply is removed. Flash memory is commonly used in solid-state drives (SSDs). There are two types of flash memory: NAND flash and NOR flash. NAND flash memory has high storage density and low cost per bit, making it suitable for SSDs and mobile applications. NOR flash is optimized for random access and is often used in applications that require fast code execution.
[0023] In the context of storage as used herein, the term "buffer" refers to a memory device that temporarily stores data or information as part of the process of moving data from one location to another. Buffers are generally embodied in static random access memory for fast access. Buffers consist of standard SRAM or a first-in-first-out (FIFO) organization.
[0024] This embodiment discloses a register renaming technique. This technique provides a method for efficiently processing the instruction pipeline in the microarchitecture of a PE in a system using a large number of processing elements (PEs). This technique offers various advantages, such as faster processing of stages in the instruction pipeline, dedicated circuits for specific tasks, and independent operation of the compiler. In the instruction pipeline, an instruction goes through various steps, including decoding, register renaming, issuing, executing, and discarding. Register renaming is the step that follows the instruction decoding step. This step is performed by a register renaming circuit that includes a collision detection circuit and a mapping circuit. The collision detection circuit is configured to detect a register collision between a first decoded instruction and a second decoded instruction. The register collision relates to a first architecture register and a first physical register corresponding to the first architecture register. The mapping circuit is configured to change the first architecture register to a second architecture register and map the second architecture register to a second physical register different from the first physical register. The first decoded instruction and the second decoded instruction are decoded in a single thread of the processing element (PE). The mapping circuit includes a renaming circuit and a mapping table. The renaming circuit is configured to change the first architecture register to the second architecture register. The mapping table is configured to map the second architecture register to the second physical register. The mapping table stores an architecture identifier that identifies either the first or second architecture register, and a physical identifier that identifies the second physical register.
[0025] Figure 1 is a block diagram showing system 100 according to one embodiment of the present invention.
[0026] System 100 is embodied in one or more System-on-a-Chip (SoC) packages, including high-density devices such as three-dimensional (3D) packages. System 100 comprises a host processor 110, an Input / Output (I / O) controller 140, a Network Interface Card (NIC) 146, a Graphic Display Controller (GDC) 150, a bus 155, a memory controller 160, and multiple processing elements (PEs) 170. kThe system includes (k=1, ..., N). These components interface with each other or include other components described later. System 100 may include more or fewer components than those described above. Also, one component may be integrated with other components. For example, the I / O controller 140, GDC 150, and memory controller 160 may be integrated into a single module. Such integration may be partial or superimposed on each other. For example, the GDC 150 may be integrated into the processor 110, and the I / O controller 140 and memory controller 160 may be integrated into a single controller. System 100 exemplifies the role that high-bandwidth memory (HBM) circuits play in high-performance computing (HC) platforms. Many HC platforms use multiple HBM circuits, which include stacked dynamic random access memory (DRAM) that operates with processing units or I / O circuits. In many cases, application environments require low power consumption, reliable signal integrity, fault tolerance, and reliable operation under extreme conditions such as high temperatures and confined spaces. Other examples of applications that can benefit from highly integrated HBM designs include mobile communications (e.g., smartphones, base stations, user terminals), cameras, vehicles, entertainment (e.g., games, multimedia, music, movies), technical design (e.g., animation, graphics), medical (e.g., visualization, medical imaging), robotics, drones, automated test equipment, audio processing, speech synthesizers, video and image analysis, vision, automated facial recognition, artificial intelligence (AI), applications, and data centers.
[0027] The host processor 110 is a programmable device that executes a program or set of instructions to perform tasks. The host processor 110 is a specially designed processor such as a general-purpose processor, a digital signal processor (DSP), a microcontroller, a neural processing unit (NPU), a field programmable gate array (FPGA), or an applications-specific integrated circuit (ASIC). The host processor 110 may contain one core or multiple cores. Each core has multi-way multi-threading. The processor 110 has simultaneous multi-threading capabilities to further utilize the parallelism generated from multiple threads across multiple cores. The host processor 110 includes a memory management circuit (MMC) 120 and a processing management circuit (PMC) 130. The host processor 110 may include more or fewer components than those described above. The bus 155 connects the processor 110 to multiple devices, for example, multiple PE 170 k Any bus suitable for connecting to (k=1, ..., N). Bus 155 is a Direct Media Interface (DMI).
[0028] The I / O controller 140 controls the large-capacity storage 142 and the input / output devices 144. The large-capacity storage 142 includes CD-ROMs, hard disks, and solid-state drives (SSDs). The input devices include styluses, keyboards, mice, microphones, and image sensors. The output devices include audio devices, speakers, scanners, and printers. The network interface card (NIC) 146 provides an interface to the network, for example, via a wireless medium 148.
[0029] The memory controller 160 is an extension of the memory management circuit (MMC) 120. The memory controller 160 controls memory devices such as memory 162, high-bandwidth memory (HBM) 164, and non-volatile memory (NVM) 166. Memory 162 includes static random access memory (SRAM) and dynamic random access memory (DRAM). HBM 164 includes a three-dimensional stacked structure of memory dies to provide high bandwidth, low latency, low power consumption, and high storage capacity. It may also have processing-in-memory (PIM) functionality. NVM 166 includes ROM (Read-Only Memory), flash memory, wide-IO NAND, MRAM, and / or other forms of memory. Memory 162 controls the processor 110 or multiple PEs 170 kWhen executed by any one of these, the instruction words or program that cause the calculation described in various embodiments are stored, and this instruction words or program are loaded from NVM 166 or mass storage 142. Memory 162 also stores the data used for the calculation. NVM 166 contains instruction words, programs, constants, or data that are maintained regardless of whether power is supplied or not. The instruction words or program correspond to the functions described later. In one embodiment, the program is a plurality of PE 170 k It includes a compiler that compiles a program to be executed on one of the following. The compiler is executed on either the host processor 110 or one of the processors coupled to the bus 155.
[0030] The GDC 150 controls the display device 152 and provides graphics processing. The GDC 150 may be integrated inside the processor 110 and generally includes a graphical user interface (GUI) that enables interaction with the user, allowing the user to transmit commands or activate functions.
[0031] Additional devices and bus interfaces are provided for interconnection or expansion. Examples include the Peripheral Component Interconnect Express (PCIe) bus and the Universal Serial Bus (USB).
[0032] Host processor 110 and multiple PE 170 k This maintains a close communication interface via minimum bus 155 and other separate lines. Multiple PE 170 kThe cluster consisting of operates under the control and management of the host processor 110. Once activated and started, each PE 170 executes its own program and accesses its own instruction word memory 192 and data cache 194. The host processor 110 provides an abstraction layer for the overall architecture and essentially hides the complexity of program execution from the user or higher-level applications. The application program specifies the work to be done, and the host processor 110 processes the execution method by assigning or distributing the work to individual PEs.
[0033] The host processor 110 includes the MMC 120 and the PMC 130. Each PE 170 k (k = 1,..., N) has an L1 cache 172 k , a configuration circuit (CFG) 174 k , an execution circuit 180 k , an interrupt circuit 182 k , an instruction word memory 192 k , a data cache 194 k , an arithmetic circuit 196 k , and a communication interface 198 k including. The host processor 110 and the PE 170 k may include a greater or lesser number of components than the above components. For clarity of explanation below, the index k for the PE 170 k and related components is omitted.
[0034] The MMC 120 is configured to operate with or without the memory controller 160 and manages memory operations to at least one of the following based on memory access to the PMC 130 or at least one of the multiple PEs 170: the L2 cache 125, main memory 134, L1 cache 172 within the PE 170, memory 162, HBM 164, or NVM 166. Memory operations include at least one of read access, write access, page table update, Translation Lookaside Buffer (TLB) update, cache response, and access violation response. The L2 cache 125 is configured to act as a TLB, translating virtual memory to physical memory. The L2 cache 125 is typically implemented with fast memory, such as fast SRAM, so that the MMC 120 can quickly look up virtual-physical page mappings without accessing slow page tables. It is also used as a cache storage location to provide fast responses to memory accesses. The MMC 120 updates the page table in memory 162 or the TLB in the L2 cache 125 if there is a new end value in the page table. The MMC 120 handles access violations such as non-existent memory addresses, buffer overflows, and null pointers. The MMC 120 also reports violations to a test controller (not shown) for debugging or testing purposes.
[0035] The PMC 130 includes a Main Executing Circuit 132, a main memory 134, and an interrupt controller 136. The PMC 130 is configured to manage at least one processor operation performed by at least one of a plurality of PEs 170. The processor operation includes at least one of program initiation, program execution, and interrupt transmission. The Main Executing Circuit 132 is a processing unit or circuit that executes a program or instruction stored in the main memory 134. The program is any suitable program. In one embodiment, the program is a compiler that compiles the program to be executed on the PE 170. The main memory 134 is private to the PMC 130 and is a suitable type of memory such as DRAM, SRAM, magnetoresistive Random-Access Memory (MRAM), flash memory, or a combination thereof. The main memory 134 includes a page table for converting virtual pages into physical pages as part of the memory management work performed by the MMC 120. The main execution circuit 132 may access memory 162, HBM 164, and NVM 166 via the memory controller 160. The interrupt controller 136 controls and manages interrupt requests and services from or to PE 170. Management includes assigning priorities to interrupt requests and transmitting commands or messages to PE 170.
[0036] Each PE 170 is configured to operate independently or in conjunction with other PE 170s and the host processor 110. Together they form a multiprocessor system and operate in parallel or sequentially depending on the purpose of the overall system. Within each PE 170, the execution circuit 180 consists of circuits that execute programs, instructions, or commands stored in the instruction memory 192. The execution circuit 180 interfaces with or communicates with the memory controller 160 via the instruction memory 192, data memory 194, arithmetic circuit 196, communication interface 198, and bus 190. Through the memory controller 160, the execution circuit 180 accesses memory 162, HBM 164, and NVM 166. In one embodiment, the execution circuit 180 includes an instruction pipeline that processes instructions from the instruction memory 192. The instruction pipeline of the execution circuit 180 is illustrated in Figure 2. The execution circuit 180 accesses data stored in the data memory 194. The data cache 194 is used to store temporary data for executing the program and data structures such as the stack or heap. The instruction word memory 192 and the data cache 194 are private or local to the associated PE and are embodied in suitable memory such as DRAM, SRAM, MRAM, flash memory, or a combination thereof. The arithmetic circuit 196 is configured to perform logic and / or arithmetic operations and includes a number of functional units, tensor units, mathematical units, buffers, and interconnects. The arithmetic units are scheduled by a PE scheduler (not shown). The PE scheduler is configured by the host processor 110. The communication interface 198 provides an interface for communication between PEs and between associated PEs and the host processor 110. The L1 cache 172 provides fast cache memory to the execution circuit 180 and is used to implement the TLB for address translation. For additional caching operations, it is coupled to the L2 cache 125 of the host processor 110. Each PE shares information with the others by communicating between its L1 cache 172 and its L2 cache 125.The interrupt circuit 182 provides inter-processor interrupt (IPI) request and response services between PEs, as well as interrupt request and response services between PEs and the host processor 110. The interrupt circuit 182 generates IPIs to other PEs and receives IPI responses from other PEs. Each PE preloads data and status into memory 162 before requesting an interrupt, brings the relevant data when other PEs process the interrupt, and generates an interrupt to the main execution circuit 132 via the interrupt controller 136 when a PE requests a service or reports a status. For example, a PE sends an interrupt to the main execution circuit 132 when it has completed the task currently assigned to it. Before sending an interrupt, the PE transmits a message, result, data, status, or condition to memory 162 or HBM 164 so that the main execution circuit 132 can see it when responding to the interrupt. This enables an efficient communication protocol between each PE and the host processor 110. The configuration circuit (CFG) 174 includes CFG data that configures each PE to perform the necessary operations or calculations. The CFG circuit 174 activates or deactivates each PE under the control of the host processor 110.
[0037] Figure 2 is a block diagram showing the execution circuit 180 shown in Figure 1 within PE 170 according to one embodiment of the present invention.
[0038] The execution circuit 180 includes an instruction word buffer 210, an instruction word pipeline 220, a register buffer 260, and a register file 270. The execution circuit 180 may include more or fewer components than those described above.
[0039] The instruction buffer 210 stores instruction words from the instruction memory 192 and queues the instruction words to be supplied to the instruction pipeline 220. Small buffers inserted between stages of the pipeline 220 are included to maintain the flow of instruction words. The buffers buffer correct, incorrect, or predicted branch instruction words by corresponding circuits (not shown).
[0040] The instruction pipeline 220 includes numerous stages for preparing instructions to be executed in the program flow. The stages of the instruction pipeline 220 include a fetch stage 222, a decode stage 226, a register renaming stage 230, an issue stage 234, an execution stage 238, a memory stage 242, a writeback stage 246, and a retirement stage 250. The pipeline 220 may contain more or fewer stages than those listed above. Furthermore, not all stages are activated for all instructions.
[0041] The fetch stage 222 retrieves instruction words from the instruction word buffer 210 or directly from the instruction word memory 192. A program counter (not shown) tracks the address of the instruction word. The fetched instruction words are generally stored in an instruction word register ready to be decoded. The fetch stage 222 includes branch prediction, non-sequential execution, a reorder buffer, handling of mispredicted branches, and other instruction word fetching mechanisms to provide a smooth flow of instruction words through pipeline.
[0042] Decode stage 226 includes decoding, interpreting, and translating the binary representation of the instruction word into parts that can be understood and executed. Decode stage 226 includes determining the actions to be performed by separating the instruction word into opcode and operand. The result of decode stage 226 includes the decoded instruction word, which is further checked for register usage conflicts in the instruction word execution. Decode stage 226 includes branching to the microcode corresponding to the instruction word.
[0043] The register renaming stage 230 resolves register usage conflicts within the decoded instruction. In particular, it eliminates false data dependencies, such as the dangers of write-after-write (WAW) and write-after-read (WAR). The register renaming stage 230 maps logical or architectural register names to physical register names in the processor's internal register file. This process ensures that operands are read and written correctly without conflicts as the instruction passes through the pipeline and is dynamically updated. The register renaming stage 230 includes a register renaming circuit 235, which is further described in Figure 3.
[0044] Issue stage 234 prepares instruction words for execution. Issue stage 234 includes selecting and scheduling instruction words, taking into account factors such as program order and dependency analysis. Issue stage 234 includes resource allocation (e.g., function units, memory access) and operand lookup. Issues can be in both sequential and non-sequential forms.
[0045] The execution stage 238 executes the instructions provided by the issue stage 234. It uses the arithmetic circuit 196 in Figure 1 to perform arithmetic and logical operations and other operations provided by the functional unit. The execution stage 238 takes operands from the register file 270, calculates memory access, and evaluates branch prediction. It also obtains operands from the memory stage 242, which accesses the L1 cache 172 and the bus 190.
[0046] Memory stage 242 performs operations that either retrieve operands from memory or write data to memory. Examples of instruction words that access memory include load and store. If an instruction word does not access memory, memory stage 242 is bypassed.
[0047] The write-back stage 246 overwrites the results of the execution stage 238 in the register file 270. The data is recorded in the register buffer 260 and, when ready, is transmitted to the register file 270. Depending on the instruction, the write-back stage 246 obtains the data to be written either directly from the arithmetic circuit 196 or from the memory stage 242.
[0048] The retirement stage 250 finally completes the execution of the instruction word and is integrated into the write-back stage 246. The retirement stage 250 includes exception handling, resource release, and other cleanup operations.
[0049] Figure 3 is a diagram showing a register renaming circuit 235 used in the register renaming stage 230 of Figure 2 according to one embodiment of the present invention.
[0050] Figure 3 also illustrates the register renaming process with examples shown in blocks (310, 315, 320, 356, and 370).
[0051] The register renaming circuit 235 includes a collision detection circuit 340 and a mapping circuit 350. The register renaming circuit 235 may include more or fewer components than those described above. For example, the collision detection circuit 340 and the mapping circuit 350 may be integrated into a single unit or circuit. In one embodiment, the collision detection circuit 340 and the mapping circuit 350 are used with the compiler to perform register renaming during compile time. In other embodiments, the collision detection circuit 340 and the mapping circuit 350 are used for register renaming during runtime.
[0052] The collision detection circuit 340 is configured to detect a register collision between a first decoded instruction 332 and a second decoded instruction 334. The register collision involves a first architecture register and a first physical register corresponding to the first architecture register. The first and second decoded instruction words (332 and 334) are part of a microcode sequence in the PE 170 microarchitecture.
[0053] Block 310 shows three instruction words (A, B, C) in sequence. Block 315 shows the mnemonics for instruction words (A, B, C). Instruction word A is load%r1,a[i], where %r1 is the architecture register or logical register r1 and a[i] is the i-th element of the array a[]. The mnemonic for instruction word A is r1←a[i], where the arrow ← indicates a load or move operation. Similarly, instruction word B is load%r2,b[i], meaning r2←b[i], and instruction word C is mul_add c,d, meaning c←c*d+c, where * is the multiplication operator and + is the addition operator. Instruction words (1, 2) are considered simple instructions because they refer to only a single destination register. Simple instructions are easy to decode and the registers are explicitly defined. On the other hand, instruction 3 is classified as a compound instruction because it involves multiple registers in both the source and destination. Compound instructions can cause register collisions because they may not explicitly define registers.
[0054] After the decode stage 226, the instruction words (1, 2, 3) in block 310 are converted to decoded instruction words (1, 2, 4, 5, 6, 7, 8). The compiler included in the host processor 110 in Figure 1 compiles the instruction words (1, 2, 3) in block 310. Block 320 contains the decoded instruction words (1, 2, 4, 5, 6, 7, 8). Instruction words (4, 5, 6, 7, 8) are instructions compiled from instruction word 3 and then decoded. It can be seen that register collisions exist between instruction word 1 and instruction word 6, and between instruction word 2 and instruction word 7. The cause of the collisions is that in the case of instruction word 1 and instruction word 6, destination register r1 is used by both instruction words, and in the case of instruction word 2 and instruction word 7, destination register r2 is used by both instruction words. As a result, the contents of registers r1 and r2 that were previously saved are overwritten and lost.
[0055] In the current embodiment, instruction 1 and instruction 6 are considered to be the first decoded instruction 332 and the second decoded instruction 334, respectively. The collision detection circuit 340 receives the first and second decoded instruction words (332, 334) and determines whether a register collision exists. This is done by comparing the destination registers of the two instruction words. If the two registers are the same and the contents of the first register are not stored, a collision is considered to exist. On the other hand, if the registers are different, or if they are the same but the contents of the first register are stored, there is no collision, and the program can proceed to the next instruction word pair or the next stage. The program searches for collision possibility for all possible instruction word pairs within the group. If a collision exists, a register renaming operation is performed to resolve the collision. The register renaming operation is performed by the mapping circuit 350. The registers that appear in the decoded instruction words are called architecture registers or logical registers. These do not directly represent the physical registers in the register file 270 that actually store the data. The first and second decoded instruction words (332, 334) are decoded from a single thread of PE 170. In one embodiment, the actual physical registers are two sets, grouped into a primary set 353 and a reserve set 354. The primary set 353 includes the registers used as primary registers among the physical registers used for mapping. The reserve set 354 acts as a backup set and includes reserve physical registers that are used when the primary registers are in use and therefore unavailable for mapping during the renaming operation. Initially, all primary registers are available. As the instruction word progresses, these registers are used, and the number of available primary registers gradually decreases. If there are no available primary registers at all, the reserve registers are used and are used until the primary registers become available again. The reserve set 354 is hidden from the user to simplify the translation or decoding of the instruction word.
[0056] When a collision occurs, register renaming renames one of the architecture registers so that it references another physical register. In the example shown in Figure 3, a collision involving architecture register r1 exists between instruction words (1, 6). Initially, architecture register r1 is mapped from the main set to the first physical register pr1. When the collision is detected in instruction word 6, architecture register r1 is changed or converted to the second architecture register r5, and the second architecture register r5 is mapped to the second physical register pr5 to prevent it from overwriting the first physical register pr1. In other words, the mapping circuit 350 is configured to change the first architecture register r1 in instruction word 6 to the second architecture register r5, which is available in the main set 353 and maps to the second physical register pr5, which is different from the first physical register pr1. In this embodiment, it is assumed that four architecture registers are mapped to four physical registers before the collision occurs, and the use of five physical registers is merely illustrative for illustrative purposes. When a collision occurs, assuming that the next physical register pr5 is available in the main set, the next physical register pr5 is used for mapping to the architecture register r5. The mapping circuit 350 is further described in Figure 4. By selecting different physical registers, the register contents are preserved even for previously decoded instruction words. For this purpose, the mapping circuit 350 includes a mapping table 352 that maps architecture registers to physical registers. The mapping table 352 shows the following mapping:
[0057] [Table 1]
[0058] The next collision occurs in instruction(2, 7) because both use architecture register r2 as the destination register. In instruction(2), architecture register r2 is mapped to physical register pr2. When the collision is detected in instruction(7), architecture register r2 is mapped to architecture register r6, so that it can be mapped to another available primary physical register (e.g., pr6), as in the case of instruction(6) discussed above. However, if all primary physical registers are unavailable because they are in use by other instruction(s), architecture register r6 is mapped to a backup or duplicate physical register, which is selected from spare set 354. Since primary physical registers are likely to be exhausted, especially when many instruction(s) are waiting, the physical registers in spare set 354 are used as backups. Spare set 354 is hidden from the user, and if a primary physical register is unavailable for renaming within register file 370, a spare register is used. Alternatively, if there are no available physical registers at all, other techniques are employed, such as temporarily saving the contents of a physical register to memory to free it up. If the spare set 354 is also exhausted, another technique is used. In this embodiment, assuming that the spare physical register rpr2 is available, the architecture register r6 is mapped to the spare physical register rpr2 to illustrate this concept.
[0059] Block 356 illustrates the register renaming process. For instruction words 1 and 6 in Block 320, architecture register r1 (which is mapped to physical register pr1 in mapping table 352) is renamed to architecture register r5, which is mapped to physical register pr5 in the case of instruction word 6, resulting in instruction word 9. Physical register pr5, unlike physical register pr1, is used in the main set 353. In relation to instruction words 2 and 6, for instruction word 2, architecture register r2 is mapped to physical register pr2. For instruction word 7, architecture registers r2 and r1 are renamed to architecture register r6 and r5, respectively. For illustrative purposes, we assume that there are no physical registers used in the main set 353 after register renaming in instruction word 6. Therefore, architecture register r6 is mapped to a register in the spare set 354. We assume that this spare register is spare physical register rpr2 in spare set 354, and that it is available differently from physical register pr2 by mapping table 352. As a result, the conflict between architecture register r1 and architecture register r2 is resolved.
[0060] Block 370 contains the sequence of the final instruction words after register renaming. Instruction words (6, 7, 8) are converted to instruction words (9, 10, 11) respectively, and all register collisions are resolved. The instruction words are then forwarded to issue stage 234.
[0061] Figure 4 is a diagram showing the mapping circuit 350 shown in Figure 3.
[0062] The mapping circuit 350 includes a register name changer 420 and a mapping table 430. The number of components of the mapping circuit 350 may be greater or less than the above-mentioned components.
[0063] The register renamer 420 receives the decoded instruction 410 from the collision detection circuit 340. The decoded instruction 410 includes the architecture registers shown in block 320 of Figure 3. The register renamer 420 is a circuit configured to change the first architecture register to the second architecture register with the decoded instruction 410. This is shown in block 356 of Figure 3, where architecture register r1 is changed to architecture register r5 and architecture register r2 is changed to architecture register r6. The register renamer 420 generates an architecture identifier 425 that identifies the register, such as a register number.
[0064] The mapping table 430 is configured to map the second architecture registers to the second physical registers. The mapping table 430 is a high-speed memory (e.g., SRAM) containing identifiers of the physical registers corresponding to the architecture registers, including a spare set 354, as shown in the mapping table 352 in Figure 3. Since the registers used in the instruction temporarily store the table, the only criterion for selecting a register is whether the register is available. Therefore, the mapping table 430 selects any available register to store the data. As shown in Figure 3, either the primary set 353 or the spare set 354 is used depending on availability. In one embodiment, the mapping table 430 is embodied in a lookup table or a hash function. The mapping table 430 generates a physical identifier 435 corresponding to the architecture identifier 425. The physical identifier 435 is used to access the register file 270 when necessary, and then forwarded to the issue stage 234.
[0065] Figure 5 is a flowchart showing process 500 for register renaming according to one embodiment of the present invention.
[0066] At the start, process 500 receives decoded instruction words from the decoder (step 510). Decoded instruction words may contain instruction words with register collisions. Next, process 500 detects register collisions between the first decoded instruction word and the second decoded instruction word (step 520). The first and second decoded instruction words are decoded in a single thread of the processing element (PE). Register collisions relate to the first architecture register and the first physical register corresponding to the first architecture register. This is done by scanning the register fields of the instruction word to see if there are any matching items. This operation also includes checking the state of the register to see if the register is saved or not. Next, process 500 determines whether a register collision exists or not (step 530). If no register collision exists (no in step 530), process 500 determines whether all instruction words in the step have been processed or not (step 535). If all instruction words have not been processed, process 500 returns to step 520 to continue checking for other instruction words. Otherwise, process 500 terminates after all renaming instructions have been processed.
[0067] If a register collision exists (yes in step 530), process 500 changes the first architecture register to the second architecture register and maps the second architecture register to a second physical register different from the first physical register (step 540). Next, process 500 communicates to the executor to execute the first decoded instruction associated with the first physical register and the second decoded instruction associated with the second physical register (step 550). This is to advance the instruction to the issue stage. Then process 500 terminates.
[0068] Figure 6 is a flowchart showing a process 540 that maps to the physical register shown in Figure 5 according to one embodiment of the present invention.
[0069] At the start, process 540 changes the first architecture register to the second architecture register (step 610). This is illustrated in step 356 of Figure 3, where architecture register r1 and architecture register r2 are changed to architecture register r5 and architecture register r6, respectively. Next, process 540 maps the second architecture register to the second physical register (step 620). This mapping is illustrated in mapping table 352 of Figure 3, where r5 is mapped to pr5 and r6 is mapped to rpr2. Mapping table 352 stores an architecture identifier that identifies either the first or second architecture register, and a physical identifier that identifies the second physical register. Next, process 540 terminates.
[0070] The whole or a part of this embodiment can be embodied by a variety of means depending on the specific features, functions, and applications. Such means include hardware, software, firmware, or a combination thereof. Hardware, software, or firmware elements have various modules that are coupled to one another. Hardware modules are coupled to other modules via mechanical, electrical, optical, electromagnetic, or other physical connections. Software modules are coupled to other modules in ways such as functions, procedures, methods, subprograms, or subroutine calls, jumps, links, parameter, variable, factor transfer, and function returns. Software modules are coupled to receive variables, parameters, factors, pointers, etc. from other modules, or to generate or transmit results, updated variables, pointers, etc. Firmware modules are coupled to other modules via any combination of hardware and software coupling schemes. Hardware, software, or firmware modules are each coupled to other hardware, software, or firmware modules. A module is a software driver or interface for interacting with an operating system (OS) running on a platform. A module may also be a hardware driver for configuring, setting up, and initializing hardware devices and for sending and receiving data on such devices. The device includes any combination of hardware, software, and firmware modules.
[0071] The themes and operations described herein are embodied in computer software, firmware, hardware, or combinations thereof, including digital electronic circuits, or structures and their structural transparencys disclosed herein. Embodiments of the themes described herein are embodied in one or more computer programs, i.e., one or more modular computer program instructions encoded on a computer recording medium, which are used to execute or control the operation of a data processing device. Alternatively, the program instructions are encoded in artificially generated radio signals, such as mechanically generated electrical signals, optical signals, or electromagnetic signals generated for the transmission of information, which are executed by a suitable receiving device. Computer recording media include, and are contained within, computer-readable storage devices, computer-readable memory boards, random or sequential access memory arrays or devices, or combinations thereof. Computer recording media can be the source or destination of computer program instructions encoded in artificially generated radio signals, although these are not radio signals. Computer recording media may also consist of one or more individual physical components or media (e.g., multiple CDs, disks, or other storage devices). Furthermore, the operations described herein are embodied in operations performed by a data processing device on data stored in one or more computer-readable storage devices or on data received from other references.
[0072] This specification contains many specific details of implementation, but these details should be interpreted as describing features specific to particular embodiments, rather than limiting the scope of the claimed theme. In this specification, certain features described in separate embodiments may be embodied by combining them in a single embodiment. Conversely, diverse features described in a single embodiment may be embodied by separating them into various embodiments or by appropriate partial combinations. Furthermore, one or more features described or claimed as operating in a particular combination may, in some cases, be removed from the combination, and the claimed combination may be limited to partial combinations or variations thereof.
[0073] Similarly, even if the diagrams show operations in a specific order, it should not be understood that such operations must necessarily be performed in the specific order or sequential order shown, or that all operations should be performed to obtain a desirable result. In certain situations, multitasking and parallel processing methods are advantageous. Furthermore, even if the system components are described separately in the above embodiments, it should not be understood that such separation is necessary in all embodiments, and the described program components and systems are generally integrated into a single software product or packaged into multiple software products.
[0074] Therefore, specific embodiments have been described herein. Other embodiments are also included within the scope of the claims. In some cases, preferred results can be obtained even if the operations expressed in the claims are performed in a different order. Furthermore, the processes shown in the drawings do not necessarily have to be performed in the illustrated order or sequentially in order to obtain preferred results. In certain embodiments, multitasking and parallel processing methods are advantageous.
[0075] Although embodiments of the present invention have been described in detail above with reference to the drawings, the present invention is not limited to the embodiments described above, and can be modified and implemented in various ways without departing from the technical spirit of the present invention. [Explanation of Symbols]
[0076] 100 Systems 110 host processors 120 Memory Management Circuit (MMC) 125 L2 Cache 130 Processing Management Circuit (PMC) 132 Main execution circuit 134 Main Memory 136 Interrupt Controllers 140 I / O controllers 142 Large Capacity Storage 144 Input / Output (I / O) Devices 146 Network Interface Card (NIC) 148 Wireless Medium 150 Graphics Display Controller (GDC) 152 Display devices Buses 155 and 190 160 memory controllers 162 memory 164 High-bandwidth memory (HBM) 166 Non-volatile memory (NVM) 1701-170 N Processing elements (PE) 172, 1721L1 cache 1741 configuration circuit (CFG) 1801 execution circuit 1821 Interrupt Circuit 192, 1921 instruction word memory 1941 data cache 196, 1961 arithmetic circuit 1981 Communication Interface 210 Instruction word buffer 220 command word pipeline 222 Fetch Stage 226 Decode Stage 230 Register Renaming Stage 234 Issue Stage 235 Register renaming circuit 238 Execution Stages 242 memory stages 246 Light Backstage 250 Retirement Stages 260 register buffer 270 Register File 332, 334 First and second decoded command words 340 Collision Detection Circuit 350 Mapping Circuit 352 Mapping Tables 353 Main Set 354 Spare Set 410 Decoded command words 420 Register renamer 425 Architecture Identifier 430 Mapping Tables 435 Physical Identifier
Claims
1. A device including a register renaming circuit, A collision detection circuit configured to detect register collisions between a first decoded instruction and a second decoded instruction, The system includes a mapping circuit configured to change a first architecture register to a second architecture register and to map the second architecture register to a second physical register different from the first physical register, The register collision relates to the first architecture register and the first physical register corresponding to the first architecture register, The apparatus is characterized in that the first decoded instruction and the second decoded instruction are decoded by a single thread of a processing element (PE).
2. The apparatus according to claim 1, characterized in that the first architecture register and the first physical register are destinations for the first decoded instruction and the second decoded instruction.
3. The first decoded instruction is a simple instruction that references one register, which is the first architecture register. The apparatus according to claim 1, characterized in that the second decoded instruction is a compound instruction that references at least a source register and a destination register which is the second architecture register.
4. The apparatus according to claim 1, characterized in that the second physical register is selected from a primary set and a backup set based on availability.
5. The apparatus according to claim 1, characterized in that the first decoded instruction word associated with the first physical register and the second decoded instruction word associated with the second physical register are issued to an execution circuit by an instruction word issuer for execution.
6. The mapping circuit is A register renamer configured to change the first architecture register to the second architecture register, The system includes a mapping table configured to map the second architecture register to a second physical register, The apparatus according to claim 1, characterized in that the mapping table stores an architecture identifier that identifies the first architecture register or the second architecture register and a physical identifier that identifies the second physical register.
7. The apparatus according to claim 1, characterized in that the first decoded instruction and the second decoded instruction are part of a microcode sequence in the microarchitecture of a processing element (PE).
8. The apparatus according to claim 1, characterized in that the collision detection circuit and the mapping circuit are used for register renaming at compile time.
9. The apparatus according to claim 1, characterized in that the collision detection circuit and the mapping circuit are used for register renaming at runtime.
10. The apparatus according to claim 1, characterized in that the processing element is part of a processing element cluster of a high-bandwidth memory (HBM) processing system.
11. A method for register renaming a device including a register renaming circuit, A step of detecting a register collision between a first decoded instruction and a second decoded instruction, The first architecture register is changed to the second architecture register, The method includes the step of mapping the second architecture register to a second physical register that is available differently from the first physical register, The register collision relates to the first architecture register and the first physical register corresponding to the first architecture register, A method characterized in that the first decoded instruction and the second decoded instruction are decoded in a single thread of a processing element (PE).
12. The method according to 11, characterized in that the first architecture register and the first physical register are destinations for the first decoded instruction and the second decoded instruction.
13. The first decoded instruction is a simple instruction that references one register, which is the first architecture register. The method according to 11, characterized in that the second decoded instruction is a compound instruction that references at least a source register and a destination register which is the second architecture register.
14. The method according to 11, characterized in that the second physical register is selected from a primary set and a backup set based on availability.
15. The method according to 11, further comprising the step of issuing an execution circuit to execute the first decoded instruction associated with the first physical register and the second decoded instruction associated with the second physical register.
16. Includes a mapping table for mapping the second architecture register to the second physical register using a mapping circuit, The method according to 11, characterized in that the mapping table stores an architecture identifier that identifies the first architecture register or the second architecture register and a physical identifier that identifies the second physical register.
17. The method according to 11, characterized in that the first decoded instruction and the second decoded instruction are part of a microcode sequence in the microarchitecture of a processing element (PE).
18. The method according to 11, characterized in that the detection of the register collision and the mapping are performed by the compiler at compile time.
19. The method according to 11, characterized in that the detection of the register collision and the mapping are performed at runtime.
20. A system including a register renaming circuit, A host processor configured to manage processor and memory operations, A processing element of a processing element (PE) cluster configured to be managed by a management processor, comprising: The aforementioned processing element is A collision detection circuit configured to detect register collisions between a first decoded instruction and a second decoded instruction, The system includes a mapping circuit configured to change a first architecture register to a second architecture register and to map the second architecture register to a second physical register different from the first physical register, The register collision relates to the first architecture register and the first physical register corresponding to the first architecture register, A system characterized in that the first decoded instruction and the second decoded instruction are decoded by a single thread of a processing element (PE).