Apparatus, method and system for register renaming

By employing a register renaming technique based on thread offset in a multi-threaded processing system, register addresses are dynamically managed, resolving register conflict issues and improving processor performance and resource utilization efficiency.

CN121704901APending Publication Date: 2026-03-20SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511295828.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-07-15
Filing Date
2025-09-11
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In multi-threaded processing systems, existing technologies struggle to effectively resolve register conflicts, leading to inefficient processor architecture design and complex hardware solutions that impact high-performance computing.

Method used

A thread offset-based register renaming technique is adopted to manage register usage in multi-threaded processing elements through an offset getter and an address pointer generator, including conflict detection and mapping circuitry, and dynamically renaming register addresses to avoid conflicts.

Benefits of technology

It enables efficient register usage in multi-threaded processing systems, quickly resolves register conflicts, simplifies design, and improves processor performance and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121704901A_ABST
    Figure CN121704901A_ABST
Patent Text Reader

Abstract

Apparatuses, methods, and systems for register renaming are disclosed. The offset fetcher is configured to obtain a first offset and a second offset in accordance with a first register usage and a second register usage, respectively, based on thread identifiers identifying at least one of the first thread and the second thread, respectively. The first thread and the second thread execute on a processing element (PE). The address pointer generator is configured to generate at least one of the first register address and the second register address based on at least one of the first offset and the second offset, respectively. The first register address and the second register address correspond to a first operand and a second operand stored in the register file, respectively. The first thread and the second thread respectively include a first decode instruction and a second decode instruction that respectively operate on the first operand and the second operand.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 696,797, filed September 19, 2024, and U.S. Patent Application No. 19 / 270,462, filed July 15, 2025, which are incorporated herein by reference in their entirety as fully set forth herein. Technical Field

[0002] The disclosure generally relates to computer architecture. More specifically, the subject matter disclosed herein relates to thread offset-based register renaming in multithreaded processing systems. Background Technology

[0003] This background section is intended to provide context only, and the disclosure of any concept in this section does not constitute an admission that the concept is prior art.

[0004] Advances in data science, artificial intelligence (AI), and machine learning (ML) have driven transformative changes in technology across various industries. To accommodate these changes, new technologies have been developed for semiconductor devices and systems, including computing architectures, processor and memory designs, cybersecurity, and communication interfaces. Among these developments, processor architecture has become increasingly important, particularly in applications requiring high throughput, low power consumption, and limited physical space, such as mobile devices.

[0005] In advanced processor architecture design, instruction pipelining has become popular for many processing applications involving multithreading and parallel operations. With the increasing demand for high-performance computing, the design of efficient processor architectures has faced numerous challenges. Issues such as architectural register dependencies, out-of-order execution versus ordered execution, circuit complexity, inefficient compiler techniques, and the complexity of communication and interfaces in multiprocessor environments have created many problems in instruction pipelining design. These problems are even more prevalent in multithreaded processing systems. Compilers tend to generate inefficient code when resolving register conflicts and do not utilize the internal structure of the microarchitecture. Hardware solutions tend to be overly complex, resulting in large silicon areas and being unsuitable for high-performance computing.

[0006] The information disclosed in this background section is only intended to enhance the understanding of the background information disclosed, and therefore may contain information that does not constitute prior art. Summary of the Invention

[0007] To overcome these problems, a system and method for register renaming techniques in microarchitecture are described herein. The techniques aim to provide an efficient structure for resolving conflicts in register usage. The techniques include a hardware implementation of circuitry for renaming registers at runtime as instructions flow through the instruction pipeline in a processing element (PE).

[0008] In one embodiment, a register renaming circuit includes an offset fetcher and an address pointer. The offset fetcher is configured to obtain a first offset and a second offset based on thread identifiers that identify at least one of a first thread or a second thread, respectively, according to first register usage and second register usage. The first thread and the second thread execute on a processing element (PE). The address pointer is configured to generate at least one of a first register address and a second register address based on at least one of the first offset and the second offset, respectively. The first register address and the second register address correspond to a first operand and a second operand stored in a register file, respectively. The first thread and the second thread each include a first decoding instruction and a second decoding instruction that operate on the first operand and the second operand, respectively.

[0009] In one embodiment, a method includes: obtaining a first offset and a second offset based on thread identifiers that identify at least one of a first thread or a second thread, respectively, according to first register usage and second register usage, wherein the first thread and the second thread execute on a processing element; generating at least one of a first register address or a second register address based on at least one of the first offset or the second offset, wherein the first register address and the second register address correspond to a first operand and a second operand stored in a register file, respectively, and wherein the first thread and the second thread each include a first decoding instruction and a second decoding instruction that operate on the first operand and the second operand, respectively.

[0010] In one embodiment, a system includes: a management processor configured to manage processor operations and memory operations; and processing elements, in a cluster of processing elements, configured to be managed by the management processor, the processing elements including register renaming circuitry, the register renaming circuitry including: a conflict detector circuit configured to: detect a register conflict between a first decoding instruction and a second decoding instruction, the register conflict being associated with a first architecture register and a first physical register corresponding to the first architecture register; and a mapping circuit configured to: map the first architecture register to a second physical register that is available and different from the first physical register, wherein the first decoding instruction and the second decoding instruction are decoded from a single thread. Attached Figure Description

[0011] In the following sections, aspects of the subject matter disclosed herein will be described with reference to exemplary embodiments shown in the accompanying drawings.

[0012] Figure 1 This is a block diagram illustrating a system according to an embodiment.

[0013] Figure 2 This is a diagram illustrating the execution circuitry in the PE according to an embodiment.

[0014] Figure 3 This is a diagram illustrating a multithreaded program according to an embodiment.

[0015] Figure 4 This is a diagram illustrating the concurrent execution of a multithreaded program according to an embodiment.

[0016] Figure 5 This is a diagram illustrating a register renaming circuit according to an embodiment.

[0017] Figure 6 This is a flowchart illustrating the process of renaming a register using an offset according to an embodiment. Detailed Implementation

[0018] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosure. However, those skilled in the art will understand that the aspects disclosed may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the subject matter disclosed herein.

[0019] Throughout this specification, references to “an embodiment” or “an embodiment” mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment disclosed herein. Therefore, the appearance of the phrases “in an embodiment”, “in an embodiment”, or “according to an embodiment” (or other phrases with similar meanings) throughout this specification does not necessarily indicate the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, as used herein, the word “exemplary” means “serving as an example, instance, or illustration.” Any embodiment described herein as “exemplary” should not be construed as necessarily preferred or advantageous over other embodiments. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Additionally, depending on the context discussed herein, singular terms may include their corresponding plural forms, and plural terms may include their corresponding singular forms. Similarly, hyphenated terms (e.g., “two-dimensional,” “pre-defined,” “pixel-specific,” etc.) may occasionally be used interchangeably with their corresponding non-hyphenated versions (e.g., “two-dimensional,” “pre-defined,” “pixel-specific,” etc.), and uppercase entries may be used interchangeably with their corresponding non-uppercase versions. Such occasional interchangeable use should not be considered as inconsistency between them.

[0020] Furthermore, depending on the context of this discussion, singular terms may include their corresponding plural forms, and plural terms may include their corresponding singular forms. It should also be noted that the various figures shown and discussed herein (including component diagrams) are for illustrative purposes only and are not drawn to scale. For example, the dimensions of some elements may be exaggerated relative to others for clarity. Additionally, reference numerals have been repeated in the figures where appropriate to indicate corresponding and / or similar elements.

[0021] The terminology used herein is for the purpose of describing some exemplary embodiments only and is not intended to limit the claimed subject matter. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. It will also be understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof.

[0022] It will be understood that when an element or layer is referred to as being on, "connected to," or "bonded to" another element or layer, it may be directly on, directly connected to, or bonded to the other element or layer, or there may be intermediate elements or layers present. Conversely, when an element is referred to as being "directly on," "directly connected to," or "directly bonded to" another element or layer, there are no intermediate elements or layers present. The same reference numerals always denote the same element. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0023] Unless explicitly defined herein, terms such as “first,” “second,” etc., as used herein, serve as labels for nouns following them and do not imply any kind of ordering (e.g., spatial, temporal, logical, etc.). Furthermore, the same reference numerals may be used across two or more figures to denote parts, components, blocks, circuits, units, or modules having the same or similar functions. However, such usage is merely for simplicity of description and ease of discussion; it does not imply that the construction or architectural details of such components or units are the same across all embodiments, or that such commonly referenced parts / modules are the only way to implement some of the exemplary embodiments disclosed herein.

[0024] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject pertains. It will also be understood that, unless expressly defined herein, terms (such as those defined in general dictionaries) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and shall not be interpreted in an idealized or overly formalized sense.

[0025] As used herein, the term "module" means any combination of software, firmware, and / or hardware configured to provide the functionality described herein in conjunction with modules. For example, software may be implemented as a software package, code, and / or instruction set or instructions, and the term "hardware" as used in any implementation described herein may, for example, individually or in any combination, include components, hardwired circuitry, programmable circuitry, state machine circuitry, and / or firmware storing instructions executed by programmable circuitry. Modules may be implemented collectively or individually as circuitry forming part of a larger system, such as, but not limited to, integrated circuits (ICs), systems-on-chips (SoCs), components, etc.

[0026] As used herein, in the context of storage devices, the term "solid-state" refers to a storage technology that uses integrated circuits rather than moving parts (e.g., spinning disks, platters, read / write heads) to store data. The term "flash memory" refers to a non-volatile memory that retains data even when power is removed. It is commonly used in solid-state drives (SSDs). There are two types of flash memory: NAND flash and NOR flash. NAND flash has high storage density and low cost per bit and is suitable for SSDs and mobile applications. NOR flash is optimized for random access and is typically used in applications that require fast code execution.

[0027] As used herein, the term "buffer" in the context of storage devices refers to a memory device that temporarily stores data or information as part of an operation involving moving data from one location to another. Buffers are typically implemented using static random access memory (RAM) for fast access. Buffers can be organized as standard SRAM or first-in-first-out (FIFO) organization.

[0028] In one embodiment, a technique for register renaming is disclosed. This technique provides high efficiency in instruction pipeline processing within a microarchitecture of multi-threaded PEs in a system using multiple processing elements (PEs). The technique offers several advantages, including rapid resolution of register conflicts, simple design, and efficient register use of register files.

[0029] In one embodiment, the compiler compiles a multithreaded program and provides thread identifiers and decoding instructions. The register renaming circuitry includes a table, an offset fetcher, and an address pointer generator. A table with offsets for the threads is created based on the number of threads to be executed and the size of the register file. The decoding instructions provide register usage including references to registers. The offset fetcher is configured to obtain a first offset and a second offset based on thread identifiers that identify at least one of the first and second threads, respectively, according to the first register usage and the second register usage. The first and second threads execute on a processing element (PE). The address pointer generator is configured to generate at least one of a first register address and a second register address based on at least one of the first and second offsets, respectively. The first and second register addresses correspond to a first operand and a second operand stored in the register file, respectively. The first and second threads each include a first decoding instruction and a second decoding instruction that operate on the first operand and the second operand, respectively.

[0030] Figure 1 This is a block diagram illustrating a system 100 according to an embodiment. The system (e.g., a system for register renaming) 100 may be implemented as one or more system-on-a-chip (SoC) packages including high-density devices such as three-dimensional (3D) packages. System 100 includes a host processor 110 (or processor 110), an input / output (I / O) controller 140, a network interface card (NIC) 146, a graphics display controller (GDC) 150, a bus 155, a memory controller 160, and multiple processing elements 170. k(k = 1, ..., N, where N is a positive integer greater than 1). These components may interface to each other or may include other components further described below. System 100 may include more or fewer components than those described above. Furthermore, components may be integrated into other components. For example, I / O controller 140, GDC 150, and memory controller 160 may be integrated into a module. Integration may be partial and / or overlapping. For example, GDC 150 may be integrated into processor 110, and I / O controller 140 and memory controller 160 may be integrated into a single controller, etc. System 100 may be an example illustrating the role of high-bandwidth memory (HBM) circuitry in high-computing (HC) platforms. Many HC platforms may use several HBM circuits comprising stacked dynamic random access memory (DRAM) that operates in conjunction with processing units or I / O circuitry. In many cases, the application environment adds additional criteria including low power consumption, reliable signal integrity, fault tolerance, and reliable operation under extreme conditions including heat and confined spaces. Examples of other applications that will benefit from highly integrated HBM designs include mobile communications (e.g., smartphones, base stations, user devices), cameras, vehicles, entertainment (e.g., games, multimedia, music, movies), technical design (e.g., animation, graphics), medical (e.g., visualization, medical imaging), robotics, drones, automated test equipment, audio processing, speech synthesizers, video and image analytics, vision, automated facial recognition, artificial intelligence (AI) applications, and data centers.

[0031] The host processor 110 is a programmable device that executes a set of programs or instructions to perform a task. It may be a general-purpose processor, a digital signal processor, a microcontroller, a neural processor (NPU), or a specially designed processor (such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)). It may include a single core or multiple cores. Each core may have multi-way multithreading. The processor 110 may have synchronous multithreading features to further utilize the parallelism of multiple threads across multiple cores. The host processor 110 may include memory management circuitry (MMC) 120 and processing management circuitry (PMC) 130. The host processor 110 may include more or fewer components than those described above. Bus 155 may be any suitable bus for connecting the processor 110 to other devices, including multiple physical devices (PEs) 170. k (k=1, ..., N). Bus 155 can be a Direct Media Interface (DMI).

[0032] I / O controller 140 controls mass storage device 142 and input / output device 144. Mass storage device 142 may include CD-ROM, hard disk, and solid-state drive (SSD). Input devices may include stylus, keyboard, mouse, microphone, and image sensor. Output devices may include audio devices, speakers, scanners, and printers. Network interface card (NIC) 146 provides an interface to a network, for example, via wireless medium 148.

[0033] Memory controller 160 may be an extension of MMC 120. It controls memory devices such as memory 162, HBM 164, and non-volatile memory (NVM) 166. Memory 162 may include static random access memory (SRAM) and dynamic random access memory (DRAM). HBM 164 may include a 3D stack of memory dies to provide high bandwidth, low latency, low power consumption, and high storage capacity. It may also have in-memory processing (PIM) capabilities. NVM 166 may include read-only memory (ROM), flash memory, wide-IO NAND, MRAM, and / or other types of memory. Memory 162 may store instructions or programs loaded from NVM 166 or mass storage device 142, which are then processed by processor 110 or PE 170. K When either of them executes, it causes processor 110 or PE 170 to... K Any of them performs the operations described in the various embodiments. It may also store data used in the operation. The NVM 166 may include instructions, programs, constants, or data that are maintained regardless of whether it is powered on. Instructions or programs may correspond to the functions described below. In one embodiment, a program may include data compiled to be used in the PE170. k The compiler of the program executed in one of the processors. The compiler can be executed by either the host processor 100 or any of the processors connected to the bus 155.

[0034] The GDC 150 controls the display device 152 and provides graphical operation. It can be integrated into the processor 110. It typically has a graphical user interface (GUI) to allow interaction with a user who can send commands or activate functions.

[0035] Additional devices or bus interfaces can be used for interconnection and / or expansion. Some examples may include the PCIe (High-Speed ​​Peripheral Component Interconnect) bus, Universal Serial Bus (USB), etc.

[0036] Main processor 110 and multiple processing elements 170 k A tight communication interface is maintained at least via bus 155 and other separate lines. Multiple PE 170 kThe cluster operates under the control and management of the host processor 110. Once enabled and started, each of the PEs 170 can execute its own program and access data in its private instruction memory 192 and data memory 194. The host processor 110 provides a layer of abstraction for the overall architecture. Essentially, it hides the complexity of program execution from users or high-level applications. The application can specify what needs to be done, and the host processor 110 will handle the details of how to execute by assigning or delegating tasks to individual PEs.

[0037] The host processor 110 includes a memory management circuit (MMC) 120 and a processing management circuit (PMC) 130. PE 170 K Each of (k=1, ..., N) includes an L1 cache of 172. k Configuration (CFG) circuit 174 k (or CFG 174), Execution circuit 180 k Interrupt circuit 182 k Instruction memory 192 k Data storage 194 k Calculation circuit 196 k and communication interface 198 k Host processor 110 and PE 170 k It may include more or fewer components than those described above. For clarity, PE 170 may be discarded in the following text. k The index k of its associated components.

[0038] MMC 120 is configured to operate with or without memory controller 160 to manage memory operations on at least one of L2 cache 125, main memory 134, L1 cache 172 in PE 170, memory 162, HBM 164, and NVM 166 based on memory accesses via at least one of PE 170 and at least one of PMC 130. Memory operations may include at least one of read access, write access, page table update, translation lookup buffer (TLB) update, cache response, and access violation response. L2 cache 125 may be configured to act as a translation lookup buffer (TLB) to translate virtual memory into physical memory. L2 cache 125 is typically implemented using fast memory (such as fast SRAM) to allow MMC 120 to quickly fetch virtual-to-physical page mappings without accessing slower page tables. It may also be used as a cache storage device to provide fast response to memory accesses. When a new entry exists in the table, the MMC 120 can update the page table in memory 162 or the TLB in L2 cache 125. The MMC 120 can respond to any access violations (such as non-existent memory addresses, buffer overflows, null pointers, etc.). It can report any violations to a test controller (not shown) for debugging or testing purposes.

[0039] PMC 130 includes main execution circuitry 132, main memory 134, and interrupt controller 136. It is configured to manage at least one processor operation executed by at least one of PE 170. Processor operations may include at least one of program initiation, program execution, and interrupt delivery. Main execution circuitry 132 may be a processing unit or circuit capable of executing a program or instructions stored in main memory 134. The program may be any suitable program. In one embodiment, the program is a compiler that compiles a program for execution in PE 170. Main memory 134 is dedicated to PMC 130. It may be any suitable type of memory (such as DRAM, SRAM, magnetoresistive random access memory (MRAM), flash memory, or any combination thereof). Main memory 134 may include page tables to translate virtual pages into physical pages as part of memory management tasks performed by MMC 120. Main execution circuitry 132 may also access memory 162, HBM 164, and NVM 166 via memory controller 160. Interrupt controller 136 controls and manages interrupt requests and interrupt services from or to PE 170. This may include prioritizing interrupt requests and sending commands or messages to PE 170.

[0040] Each of the PEs 170 is configured to operate independently or collaboratively with other PEs 170 and the host processor 110. Together, they form a multiprocessor system capable of cooperating to work in parallel or sequentially based on overall system objectives. In each of the PEs 170, execution circuitry 180 is configured to execute circuitry of programs, instructions, or commands stored in the instruction memory 192. Execution circuitry 180 is connected to or communicates with the instruction memory 192, data memory 194, computing circuitry 196, communication interface 198, and memory controller 160 via a bus interface 190. Execution circuitry 180 accesses memory 162, HBM 164, and NVM 166 through memory controller 160. In some embodiments, execution circuitry 180 includes an instruction pipeline for processing instructions from instruction memory 192. Figure 2The instruction pipeline in execution circuitry 180 is described. Execution circuitry 180 can access data stored in data memory 194. Data memory 194 can be used to store temporary data and data structures (such as stacks or heaps) for program execution. Instruction memory and data memory 192 and 194 are private or local to the associated PE and can be implemented by any suitable memory including DRAM, SRAM, MRAM, flash memory, or any combination thereof. Computation circuitry 196 is configured to perform logical and / or computational operations. It may include multiple functional units, tensor units, mathematical units, buffers, and interconnects. These computational units can be scheduled by a PE scheduler (not shown). The PE scheduler can be configured by host processor 110. Communication interface 198 provides an interface for communication between PEs and between an associated PE and host processor 110. L1 cache 172 provides a fast cache to execution circuitry 180. It can be used to implement a TLB for address translation. It can be connected to L2 cache 125 in host processor 110 for additional cache operations. By allowing L1 cache 172 and L2 cache 125 to communicate in each PE, PEs can share information among themselves. Interrupt circuit 182 servicing interrupt requests and responses within PEs for inter-processor interrupts (IPIs) and between PEs and the host processor 110. Interrupt circuit 182 generates IPIs to other PEs and receives IPI responses from other PEs. PEs can preload data or status in memory 162 before requesting an interrupt, allowing another PE to retrieve the data when servicing an interrupt. Interrupt circuit 182 can also generate interrupts to the main execution circuit 132 via interrupt controller 136 when a PE requests service or reports status. For example, a PE can send an interrupt to the main execution circuit 132 when it completes its currently assigned task. Before sending the interrupt, it can send messages, results, data, status, or conditions to memory 162 or HBM 164 to allow the main execution circuit 132 to examine the messages when responding to the interrupt. This allows for an efficient communication protocol between PEs and the host processor 110. CFG circuit 174 includes CFG data that configures PE 170 to perform operations or calculations as needed. CFG circuit 174 can also enable or disable PE under the control of host processor 110.

[0041] Figure 2 This illustrates the PE 170 according to an embodiment. Figure 1 The diagram shows an execution circuit 180. The execution circuit 180 includes an instruction buffer 210, an instruction pipeline 220 (or pipeline 220), a register buffer 260, and a register file (or register stack) 270. The execution circuit 180 may include more or fewer components than those described above.

[0042] Instruction buffer 210 stores instructions from instruction memory 192 and queues them for feeding into instruction pipeline 220. Instruction buffer 210 may include small buffers between stages in pipeline 220 to maintain the flow of instructions. Instruction buffer 210 may buffer ordered instructions, out-of-order instructions, or predicted branch instructions through corresponding circuitry (not shown).

[0043] Instruction pipeline 220 includes multiple stages to prepare instructions for execution in program flow. The stages in instruction pipeline 220 include fetch stage 222 (or fetch 222), decode stage 226 (or decode 226), register renaming stage 230 (or register renaming 230), initiation stage 234 (or initiation 234), execution stage 238 (or execution 238), memory stage 242 (or memory 242), write-back stage 246 (or write-back 246), and retirement stage 250 (or retirement 250). Pipeline 220 may include more or fewer stages than those described above. Furthermore, not all stages are active for all instructions.

[0044] The fetch stage 222 fetches (or retrieves) instructions from the instruction buffer 210 or directly from the instruction memory 192. A program counter (not shown) tracks the address of the instruction. The fetched instruction is typically stored in the instruction register for decoding. The fetch stage may include branch prediction, out-of-order execution, reorder buffering, mispredicted branch handling, and other instruction fetching mechanisms to provide a smooth flow of instructions through the pipeline.

[0045] Decoding level 226 involves decoding, interpreting, and translating the binary representation of an instruction into understandable and executable parts. This includes separating the instruction into opcodes and operands and determining what action to perform. The result of decoding level 226 includes the decoded instruction, which is further examined to determine if there are any conflicts in register usage during instruction execution. Decoding level 226 may include branching to the microcode corresponding to the instruction.

[0046] Register renaming level 230 is used to resolve register conflicts in decoded instructions. It eliminates spurious data dependencies, especially write-after-write (WAW) and read-after-write (WAR) hazards. Register renaming level 230 maps logical register names or architectural register names to physical register names in the processor's internal register file. This process is dynamically updated as instructions flow through the pipeline, ensuring that conflicts do not occur and operands are read and written correctly. Register renaming level 230 includes... Figure 5 The register renaming circuit 235 is further described in the text.

[0047] Initiator (or issuer) level 234 prepares instructions for execution. It may include selecting and scheduling instructions considering factors such as program order and dependency analysis. Initiator level 234 may include allocating resources (e.g., functional units, memory accesses) and fetching operands. Initiation may include ordered and out-of-order schemes.

[0048] Execution stage 238 executes instructions initiated and prepared by initiating stage 234. It uses computing circuitry 196 (in... Figure 1 The function (in the middle) can perform arithmetic and logical operations, as well as other operations, provided by the functional unit. It can obtain operands from register file 270, calculate memory addresses, and evaluate branch predictions. It can obtain operands from memory level 242, which accesses L1 cache 172 and bus 190.

[0049] Memory level 242 retrieves operands from memory or writes data to memory. Examples of instructions that can access memory include load and store. If an instruction does not access memory, memory level 242 can be bypassed.

[0050] Write-back stage 246 writes the results of execution stage 238 back to register file 270. It writes data to register buffer 260, which sends data to register file 270 when ready. Write-back stage 246 can obtain the data to be written back directly from computing circuitry 196 or from memory stage 242 according to instructions.

[0051] The retirement level 250 terminates the execution of instructions. It can be combined with the write-back level 246. It may include exception handling, resource release, and other housekeeping functions.

[0052] Figure 3 This is a diagram illustrating a multithreaded program 300 according to an embodiment. The multithreaded program 300 includes a main program 310 and K threads (or programs) 320. j (j=1, ..., K, where K is a positive integer greater than 1).

[0053] The main program 310 is a loop that iterates the main body N times. Each iteration (or loop) includes the following operations (where %r represents the architecture register or logic register): 20: Load %r1, a[i] represents r1 a[i] 30: Load %r2, b[i] represents r2 b[i] 40: Multiplication %r3, %r1, %r2 represent r3 r1×r2 50: Addition %r3, %r4 means r3 r3+r4 As can be seen from the above, each iteration is independent of the others. Furthermore, within each iteration, there are no dependencies between elements. All N iterations are independent of each other. In other words, iteration j does not require the results of any other iteration k, where j ≠ k. Because there are no dependencies, these iterations can be performed in parallel. In a truly parallel environment, they can be allocated to multiple physical processors, and they can be executed in parallel. In a multithreaded programming environment, this means that program 310 can be decomposed into several threads. Assume there are K threads. The main program 310 can be divided into K threads, each of which includes N / K iterations. The K threads will be allowed to execute concurrently. Each thread will execute N / K iterations with different index values.

[0054] Program 3201 will be assigned to thread 1, and thread 1 will execute a loop from i=1 to N / K. Program 320 p It will be assigned to thread p, which will execute a loop from i = (p-1)N / K+1 to pN / K. (Program 320) K The loop will be assigned to thread K, which will execute a loop from i = (K-1)N / K+1 to N. Although they are independent, each iteration involves accessing registers r1, r2, r3, and r4. If the iterations are executed sequentially within a single thread, there will be no register conflicts. However, if they are executed concurrently, register conflicts will occur, and therefore register renaming is used to resolve these conflicts.

[0055] Figure 4 This is a diagram illustrating concurrent execution 400 of a multithreaded coding program according to an embodiment. Concurrent execution 400 executes two threads 410 and 420 concurrently.

[0056] Concurrent execution of threads within a single processing element means that the processor switches execution resources, including memory, registers, and computation units, between threads. Concurrent execution gives the appearance of parallel execution, but in reality, only one thread is running at any given time. Thread switching can happen very quickly, so concurrent execution appears to be parallel.

[0057] Two threads, 410 and 420, execute on the timeline. Thread 1, 410, has the following instructions: 21: Load %r1, a[i]: r1 a[i] 31: Load %r2, b[i]: r2 b[i] 41: Multiplication %r3, %r1, %r2: r3 r1×r2 At execution, the index i of thread 1 is i=263. Instructions 21, 31, and 41 are executed in execution windows 412, 414, and 416, respectively. For simplicity, it is assumed that each of these execution windows has completed the execution of its instructions.

[0058] Thread 2 420 has the following instructions: 22: Load %r1, a[i]: r1 a[i] 32: Load %r2, b[i]: r2 b[i] 42: Multiplication %r3, %r1, %r2: r3 r1×r2 At execution, the index i of thread 2 is i=471. Instructions 22, 32, and 42 are executed at execution windows 422, 424, and 426, respectively. For simplicity, it is assumed that each of these execution windows has completed the execution of its instructions.

[0059] For concurrent execution, only one execution (or execution window) is active at any given time. From t0 to t1, execution window 412 is active. Register r1 is loaded with the contents of a

[263] . From t1 to t2, execution window 422 is active. Register r1 is loaded with the contents of a

[471] . Therefore, the previous contents of a

[263] in r1 are overwritten and corrupted. From t2 to t3, execution window 414 is active. Register r2 is loaded with the contents of b

[263] . From t3 to t4, execution window 424 is active. Register r2 is loaded with the contents of b

[471] . Therefore, the previous contents of b

[263] in r2 are overwritten and corrupted.

[0060] From t4 to t5, execution window 416 is active. Register r3 is loaded with the product of r1 and r2, which now contains a

[471] and b

[471] . Therefore, register r3 is loaded with a

[471] × b

[471] . This is incorrect because it should be the product of a

[263] and b

[263] . From t5 to t6, execution window 426 is active. Register r3 is loaded with the product of r1 and r2, which now contains a

[471] and b

[471] . Therefore, register r3 is loaded with a

[471] × b

[471] . This proves to be correct, but if a subsequent instruction in thread 1 uses the result of the product a

[263] × b

[263] , it will get an incorrect result.

[0061] The examples above illustrate that concurrent execution of multiple threads accessing the same registers can lead to unstable, incorrect, or unpredictable results. One solution to this problem is to allow each thread to access its own register file. In other words, the physical register file is statically partitioned, and these partitions are allocated to multiple threads.

[0062] Figure 5 This illustrates an embodiment. Figure 2 A diagram of register renaming circuit 235 is shown. Register renaming circuit 235 includes offset table 510, offset fetcher 520, and address pointer generator 530. Register renaming circuit 235 may include more or fewer components than those described above. In one example, register renaming circuit 235 also includes conflict detector circuitry (not shown) and mapping circuitry (not shown), the conflict detector circuitry being configured to detect register conflicts between two decoded instructions, and the mapping circuitry being configured to map schema registers to other available physical registers different from the physical registers associated with schema registers. For example, the two decoded instructions are decoded from a single thread. Register renaming circuit 235 operates with the assistance of compiler 501, which compiles the instructions and generates instructions for multi-threaded decoding.

[0063] The compiler is designed for multithreaded programming. It is configured to allow the user to specify the number of threads. During execution, it generates thread identifiers (IDs) 515 for the active running threads and decoded instructions 525 for the active threads. Register renaming circuitry 235 uses this information to partition the register file 270 and allocate the partitioned register file to the threads based on the thread identifiers. This is achieved by using offsets allocated to each thread. Each thread uses an offset to point to the registers to allocate the partitioned register file. The offsets are determined before the multithreaded program is executed and remain unchanged throughout the entire execution of the program.

[0064] For example, suppose register file 270 has N=512 registers, named r1 to r512. Suppose the number of threads is K=4. Maximizing the partition size, the entire register file 270 is partitioned into four partitions allocated to the four threads: thread 1, thread 2, thread 3, and thread 4. Therefore, each partition has a total of N / K registers (or 512 / 4=128 registers). Thread 1541 is assigned to the partition containing registers 1 through 128. Thread 2542 is assigned to the partition containing registers 129 through 256. Thread 3543 is assigned to the partition containing registers 257 through 384. Thread 4544 is assigned to the partition containing registers 385 through 512. Since each thread has its own register file, there are no register conflicts due to multiple threads, although the partition is smaller. This partitioning is based on offsets stored in an offset table, which correspond to (e.g., equal to) the first register in the partition (e.g., the register number or address of the first register).

[0065] Offsets are stored in table 510. This table can be implemented as fast SRAM (such as a cache). It will be used as a lookup table to generate offsets indexed by thread identifiers. For example, for N=512 and K=4, the table stores offset values ​​of 0, 128, 256, and 384 for thread ID 1, thread ID 2, thread ID 3, and thread ID 4, respectively. The offset is the starting register number or address of the partition. By allocating separate partitions for threads, there will be no register conflicts. Architectural registers will be translated into physical registers by adding the register address to the offset.

[0066] The decoded instructions have register usages that reference registers. Since the registers need to be renamed, offsets are obtained from Table 510. For N threads, N offsets are obtained. Assume there are two threads executing on the PE, and two decoded instructions from those two threads. The decoded instructions operate on operands stored in the register file. Each decoded instruction has register usages. Table 510 is configured to store the first offset and the second offset based on the first register usage and the second register usage, respectively.

[0067] Offset acquirer 520 is configured to obtain a first offset and a second offset based on thread identifiers that identify at least one of the first and second threads, respectively, according to the usage of the first and second registers. After the offsets are obtained, they are used to rename registers or calculate the addresses of registers.

[0068] Address pointer generator 530 can be configured to generate at least one of a first register address and a second register address based on at least one of a first offset and a second offset, respectively. The first register address and the second register address correspond to a first operand and a second operand stored in register file 270, respectively. The register address or renamed register is determined by adding the offset to the register name in the decoding instruction as follows (so that the architectural register can be mapped to the physical register): The renamed register = the register in the decoded instruction + the thread offset (1) For example, suppose Table 510 is constructed for a 4-thread program. There are two instructions executed on two threads. Instruction 1 is "load r5, a" in thread 1, and instruction 2 is "load r73, c" in thread 2. When instruction 1 is decoded, its thread identifier (ID=1) is determined. Offset fetcher 520 fetches the offset from offset table 810. This offset is 0. Therefore, address pointer generator adds the offset to the register address or register number. The renamed register is r(5+0) = r5. When instruction 2 is decoded, its thread identifier (ID=2) is determined. Offset fetcher 520 fetches the offset from offset table 510. This offset is 128. Therefore, address pointer generator adds the offset to the register address or register number in the decoded instruction. The renamed register is r(73+128) = r191.

[0069] Box 550 shows an example of a 4-thread program in execution. Thread 1 has offset = 0. It executes instructions at three execution windows 561, 569, and 577, respectively, referencing architecture registers 0, 1, and 3. Thread 2 has offset = 128. It executes instructions at three execution windows 567, 573, and 581, respectively, referencing architecture registers 2, 1, and 0. Thread 3 has offset = 256. It executes instructions at two execution windows 563 and 575, respectively, referencing architecture registers 0 and 18. Thread 4 has offset = 384. It executes instructions at three execution windows 565, 571, and 579, respectively, referencing architecture registers 0, 4, and 1. The registers are then renamed, or the register addresses of the physical registers in register file 270 are generated by address pointer generator 530 as follows: Thread 1: r0→r0, r1→r1, r3→r3 Thread 2: r2→r130 (=2+128), r1→r129 (=1+128), r0→r128 (=0+128) Thread 3: r0→r256 (=0+256), r18→r274 (=18+256) Thread 4: r0→r384 (=0+384), r4→r388 (=4+384), r1→r385 (=1+384) In the example above, four threads can reference the same architectural register, but the renamed register points to different physical registers. Therefore, register conflicts can never occur.

[0070] Figure 6 This is a flowchart illustrating a process 600 of renaming a register using an offset according to an embodiment.

[0071] At the outset, process 600 stores a first offset and a second offset in a table (box 610) based on first register usage and second register usage, respectively. The first offset or the second offset is determined based on the size of the register file and the number of threads executing in the processing element (PE). Next, process 600 obtains the first offset and the second offset, respectively, based on thread identifiers that identify at least one of the first thread and the second thread, according to the first register usage and the second register usage, respectively (box 620). The first thread and the second thread execute on the PE.

[0072] Then, process 600 generates at least one of a first register address and a second register address based on at least one of a first offset and a second offset (box 630). The first register address and the second register address are addresses of renamed registers and correspond to physical registers in the register file. The first register address and the second register address correspond to a first operand and a second operand stored in the register file, respectively. The first thread and the second thread each include a first decoding instruction and a second decoding instruction that operate on the first operand and the second operand, respectively. Then, process 600 terminates.

[0073] All or part of the embodiments may be implemented based on the application through various means according to specific features and functions. These means may include hardware, software, or firmware, or any combination thereof. Hardware, software, or firmware elements may have several modules combined with each other. Hardware modules are combined with other modules through mechanical, electrical, optical, electromagnetic, or any physical connection. Software modules are combined with other modules through function, process, method, subroutine or subroutine calls, jumps, links, parameter, variable and argument passing, function returns, etc. Software modules are combined with other modules to receive variables, parameters, arguments, pointers, etc., and / or generate or pass results, update variables, pointers, etc. Firmware modules are combined with other modules through any combination of the above hardware and software combination methods. Hardware, software, or firmware modules may be combined with any of other hardware, software, or firmware modules. Modules may also be software drivers or interfaces for interacting with an operating system running on the platform. Modules may also be hardware drivers for configuring data, setting data, initializing data, sending data to hardware devices, and receiving data from hardware devices. Devices may include any combination of hardware, software, and firmware modules.

[0074] Embodiments of the subject matter and operations described herein may be implemented in digital electronic circuits, or in computer software, firmware, or hardware including the structures disclosed herein and their equivalents, or in a combination of one or more of these. Embodiments of the subject matter described herein may be implemented as one or more computer programs (i.e., one or more modules of computer program instructions) encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Optionally or additionally, the program instructions may be encoded on artificially generated propagation signals (e.g., machine-generated electrical, optical, or electromagnetic signals), which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or may be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagation signal, it may be a source or destination of computer program instructions encoded in artificially generated propagation signals. Computer storage media may also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or may be included within one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices). Furthermore, the operations described herein can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.

[0075] While this specification may contain numerous specific details of implementation, these details should not be construed as limiting the scope of any claimed subject matter, but rather as descriptions of features specific to particular embodiments. Specific features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in a particular combination and even initially claimed in this way, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may refer to a sub-combination or a variation of a sub-combination.

[0076] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in sequential order, or all of the shown operations to be performed, in order to achieve the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0077] Therefore, specific embodiments of the subject matter have been described herein. Other embodiments are within the scope of the appended claims. In some cases, the actions set forth in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result. In certain embodiments, multitasking and parallel processing can be advantageous.

[0078] As those skilled in the art will recognize, the innovative concepts described herein can be modified and varied across a wide range of applications. Therefore, the scope of the claimed subject matter should not be limited to any particular exemplary teachings discussed above, but rather is defined instead by the appended claims.

Claims

1. A device for renaming registers, comprising: An offset acquirer is configured to: obtain a first offset and a second offset based on a thread identifier that identifies at least one of a first thread and a second thread, respectively, according to the use of a first register and the use of a second register, wherein the first thread and the second thread execute on a processing element; as well as An address pointer generator is configured to generate at least one of a first register address and a second register address based on at least one of a first offset and a second offset, respectively. Wherein, the first register address and the second register address correspond to the first operand and the second operand stored in the register file, respectively, and The first thread and the second thread each include a first decoding instruction and a second decoding instruction that perform operations on the first operand and the second operand, respectively.

2. The device as claimed in claim 1, wherein, At least one of the first offset and the second offset is determined based on the size of the register file and the number of threads executing in the processing element.

3. The device as described in claim 1, further comprising: The table is configured to store the first offset and the second offset based on the use of the first register and the use of the second register, respectively.

4. The device as described in claim 3, wherein, At least one of the first register usage and the second register usage is provided by the compiler.

5. The device as claimed in claim 1, wherein, The first and second threads are executed concurrently in the processing element.

6. The device as claimed in claim 1, wherein, The use of the first register and the use of the second register result in a register conflict.

7. The device as claimed in claim 6, wherein, At least one of the first register use and the second register use maps the architecture register to the physical register in the register file.

8. The device as claimed in claim 7, wherein, At least one of the first register use and the second register use maps the architectural register to the physical register in order to resolve register conflicts.

9. The device as claimed in claim 4, wherein, The register file is statically partitioned based on the number of concurrent threads executing on the processing element.

10. The device as claimed in claim 1, wherein, A processing element is part of a cluster of processing elements in a high-bandwidth memory processing system.

11. A method for renaming a register, comprising: Based on thread identifiers that identify at least one of the first thread and the second thread respectively, a first offset and a second offset are obtained according to the use of the first register and the use of the second register respectively, and the first thread and the second thread execute on the processing element; as well as At least one of the first register address and the second register address is generated based on at least one of the first offset and the second offset, respectively. Wherein, the first register address and the second register address correspond to the first operand and the second operand stored in the register file, respectively, and The first thread and the second thread each include a first decoding instruction and a second decoding instruction that perform operations on the first operand and the second operand, respectively.

12. The method of claim 11, wherein, At least one of the first offset and the second offset is determined based on the size of the register file and the number of threads executing in the processing element.

13. The method of claim 11, further comprising: The first offset and the second offset are stored in the table based on the use of the first register and the use of the second register, respectively.

14. The method of claim 13, wherein, At least one of the first register usage and the second register usage is provided by the compiler.

15. The method of claim 11, wherein, The first and second threads are executed concurrently in the processing element.

16. The method of claim 11, wherein, The use of the first register and the use of the second register result in a register conflict.

17. The method of claim 16, wherein, At least one of the first register use and the second register use maps the architecture register to the physical register in the register file.

18. The method of claim 17, wherein, At least one of the first register use and the second register use maps the architectural register to the physical register in order to resolve register conflicts.

19. The method of claim 14, wherein, The register file is statically partitioned based on the number of concurrent threads executing on the processing element.

20. A system for renaming registers, comprising: The management processor is configured to manage processor operations and memory operations; as well as Processing elements, within a cluster of processing elements and configured to be managed by a management processor, include register renaming circuitry. The register renaming circuit includes: A collision detector circuit is configured to: detect register collisions between a first decoding instruction and a second decoding instruction, the register collisions being associated with a first architecture register and a first physical register corresponding to the first architecture register; and The mapping circuit is configured to map a first architecture register to a second physical register that is available and different from the first physical register. The first and second decoding instructions are decoded from a single thread.