Register renaming after non-selectable scheduler queue
By buffering load operations in a non-selectable scheduler queue until load data is returned, the FPU architecture optimizes the allocation of physical register numbers, enhancing performance by maintaining a larger free list size.
Patent Information
- Application Number
- JP2022521686
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-22
- Filing Date
- 2020-10-22
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2040-10-22
AI Technical Summary
In existing floating-point unit (FPU) architectures, load operations consume physical register numbers from the free list even before the load data is returned, reducing the available size of the free list and impacting FPU performance.
The proposed solution involves buffering load operations in a non-selectable scheduler queue (NSQ) until the load data is returned from the load-store unit, then assigning physical register numbers from either the free list or a physical register number buffer, depending on the timing of these events.
This approach effectively increases the size of the free list by delaying the allocation of physical register numbers until the load data is available, thereby improving FPU performance by optimizing the use of physical register numbers.
Smart Images

Figure 0007681580000001 
Figure 0007681580000002 
Figure 0007681580000003
Abstract
Description
[Background technology]
[0001] A processing system often includes a coprocessor, such as a floating-point unit (FPU), to complement the functionality of a primary processor, such as a central processing unit (CPU). For example, an FPU performs mathematical operations, such as addition, subtraction, multiplication, division, and other floating-point instructions, including transcendental and bitwise operations. The FPU receives instructions for execution, decodes the instructions, and performs address translations required for the operations contained in the instructions. The FPU also performs register renaming by assigning one or more physical register numbers to one or more architectural registers associated with an operation. The physical register numbers indicate entries in a physical register file that store the operands or results of the operation. The FPU also includes a scheduler for scheduling operations that have entries assigned to them in the physical register file. In some cases, the FPU scheduler is a distributed scheduler that uses at least two levels of scheduler queues: (1) a first level having a non-pickable scheduler queue; and (2) a second level having two or more pickable scheduler queues. The selectable scheduler queue stores instruction operations for a corresponding subset of the multiple execution pipes, and the non-selectable scheduler queue serves to temporarily buffer instruction operations from the instruction pipeline front-end before the instruction operations are assigned to the selectable scheduler queues.
[0002] The present disclosure can be better understood, and its numerous features and advantages made apparent to those skilled in the art by reference to the following drawings, in which the use of the same reference numbers in different drawings indicates similar or identical items. [Brief description of the drawings]
[0003] [Figure 1]FIG. 1 is a block diagram of a processing system that performs register renaming after buffering instructions or operations in a non-selectable scheduler queue (NSQ) in accordance with some embodiments. [Diagram 2] FIG. 2 is a block diagram of a processor implementing a multi-modal distributed scheduler for queue and execution pipe balancing according to some embodiments. [Diagram 3] FIG. 2 is a block diagram of a processor including an FPU that implements NSQ to buffer operations before renaming, according to some embodiments. [Figure 4] FIG. 2 is a flow diagram of a method for selectively allocating physical register numbers from a free list or a physical register number buffer according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0004] A load operation of an instruction executed in a floating-point unit (FPU) is provided to the load-store unit at the same time that the FPU allocates an entry in the physical register file to hold the load data of the load operation. First, a conventional FPU performs renaming of the load operation before buffering the load operation in a non-selectable scheduler queue. The load operation is not scheduled from the non-selectable scheduler queue (to one of the selectable scheduler queues) until the load-store unit returns the load data loaded by the load operation. Since it usually takes several cycles to obtain the load data from memory or a cache, the load operation remains in the non-selectable scheduler queue for at least this time interval. However, as described above, a physical register number is assigned to the load operation before adding it to the non-selectable scheduler queue. Thus, the load operation consumes a physical register number from the free list (free list) for at least the time interval required for the load-store unit to return the load data of the load operation, which effectively reduces the size of the free list (e.g., the number of free physical register numbers) available for other operations of the FPU.
[0005] 1-4 disclose an embodiment of an architecture that improves the performance of an FPU by storing instructions in a non-selectable scheduler queue prior to renaming of architectural registers used by operations in the instruction and assignment of physical register numbers to the operations. In response to renaming of architectural registers associated with an operation, such as a load operation, the operation is added to one set of selectable scheduler queues that store operations for a corresponding subset of execution pipes. The non-selectable scheduler queue buffers the load operation at the same time that the load store unit obtains load data for the operands loaded by the load operation. The FPU also includes a load mapper that assigns physical register numbers to the load operation in response to the load data being returned from the load store unit. In some embodiments, the physical register numbers are assigned from a physical register number (PRN) buffer that holds a subset of available physical register numbers provided by a free list. The physical register numbers assigned to the load operation are mapped to a retirement identifier in a corresponding mapping structure. If the load operation is popped from the non-selectable scheduler queue before the load-store unit returns the load data, a physical register number is assigned to the load operation from the free list. If the load-store unit returns the load data before the load operation is popped from the non-selectable scheduler queue, a physical register number is assigned to the load operation from the PRN buffer. In either case, the mapping of physical register numbers to retired identifiers is stored in a mapping structure to prevent different physical register numbers from being assigned to the same load operation. Thus, the effective size of the free list in the FPU is increased, improving FPU performance, because a physical register number is not assigned to the load operation from the free list until either the load operation is popped from the non-selectable scheduler queue or the load-store unit returns the load data, whichever comes first.
[0006] 1 is a block diagram of a processing system 100 that performs register renaming after buffering instructions or operations in a non-selectable scheduler queue, according to some embodiments. The processing system 100 includes or has access to a memory 105 (e.g., system memory) or other storage components implemented using a non-transitory computer-readable medium, such as dynamic random access memory (DRAM). However, some embodiments of the memory 105 are implemented using other types of memory, including static random access memory (SRAM), non-volatile RAM, etc. The processing system 100 also includes a bus 110 to support communication between entities implemented in the processing system 100, such as the memory 105. Some embodiments of the processing system 100 include other buses, bridges, switches, routers, etc. that are not shown in FIG. 1 for clarity.
[0007] The processing system 100 includes a central processing unit (CPU) 115. Some embodiments of the CPU 115 include multiple processing elements (not shown in FIG. 1 for clarity) that execute instructions simultaneously or in parallel. The processing elements may be referred to as processor cores, computing units, or other terms. The CPU 115 is connected to a bus 110 and communicates with a memory 105 via the bus 110. The CPU 115 executes instructions, such as program code 120, stored in the memory 105, and the CPU 115 stores information, such as results of executed instructions, in the memory 105. The CPU 115 may also initiate graphics operations by issuing draw calls.
[0008] An input / output (I / O) engine 125 processes input or output operations related to a display 130 and other elements of the processing system 100, such as a keyboard, a mouse, a printer, an external disk, etc. The I / O engine 125 is coupled to the bus 110 such that the I / O engine 125 communicates with the memory 105, the CPU 115, or other entities coupled to the bus 110. In the illustrated embodiment, the I / O engine 125 reads information stored in an external storage component 135, which is implemented using a non-transitory computer-readable medium, such as a compact disc (CD), a digital video disc (DVD), etc. The I / O engine 125 also writes information, such as results of processing by the CPU 115, to the external storage component 135.
[0009] The processing system 100 includes a graphics processing unit (GPU) 140 that renders images that are displayed on the display 130. For example, the GPU 140 renders objects to generate picture element (pixel) values that are provided to the display 130, which uses the pixel values to display an image representing the rendered objects. Some embodiments of the GPU 140 are used for general-purpose computing, performing reduction and scan operations on ordered sets of elements, among other operations. In the illustrated embodiment, the GPU 140 communicates with the memory 105 (and other entities connected to the bus 110) via the bus 110. However, some embodiments of the GPU 140 communicate with the memory 105 via a direct connection or via other buses, bridges, switches, routers, and the like. The GPU 140 executes instructions stored in the memory 105, and the GPU 140 stores information in the memory 105, such as results of executed instructions. For example, the memory 105 stores a copy 145 of instructions representing program code to be executed by the GPU 140 .
[0010] Floating-point unit (FPU) 150 complements the functionality of CPU 115 and GPU 140. FPU 150 executes mathematical operations such as addition, subtraction, multiplication, division, and other floating-point instructions including transcendental operations, bit operations, and the like. Although not shown in FIG. 1 for clarity, FPU 150 includes a non-selectable scheduler queue (NSQ), a renamer, and a set of selectable scheduler queues associated with corresponding execution pipes. The NSQ buffers load operations at the same time that the load-store unit obtains the load data for the operands loaded by the load operation. In response to receiving a load operation from the NSQ, the renamer renames the architectural register used by the load operation and assigns a physical register number to the load operation. The set of selectable scheduler queues receives the load operations from the renamer and stores the load operations before execution. FPU 150 also implements (or has access to) a physical register file and a free list that stores the physical register numbers of allocatable entries in the physical register file.
[0011] 2 is a block diagram of a processor 200 implementing a multi-modal distributed scheduler 202 for queue and execution pipe balancing according to some embodiments. The processor 200 includes an instruction front end 204, one or more instruction execution units, such as an integer unit 206 and a floating point / single instruction multiple data (SIMD) unit 208, and a cache / memory interface subsystem 210. The instruction front end 204 operates to fetch instructions as part of an instruction stream, decode the instructions into one or more instruction operations (e.g., micro-ops or uops), and then dispatch each instruction operation to one of the execution units 206, 208 for execution. In executing instruction operations, the execution units frequently make use of one or more caches implemented in cache / memory interface subsystem 210, or provide data to them for accessing or storing data in external memory (e.g., random access memory or RAM) or external input / output (I / O) devices via a load / store unit (LSU), memory controller, or I / O controller (not shown) of cache / memory interface subsystem 210.
[0012] Processor 200 implements a multi-modal distributed scheduler 202 in each of one or more of the execution units of processor 200. For purposes of explanation, an embodiment is described herein in which multi-modal distributed scheduler 202 is implemented in floating point / SIMD unit 208. However, in other embodiments, integer unit 206 or other execution units of processor 200 implement a multi-modal distributed scheduler in addition to or instead of that implemented by floating point / SIMD unit 208, using the guidelines provided herein.
[0013] The multimodal distributed scheduler 202 implements a two-level queuing process whereby a first scheduler queue 212 temporarily buffers instruction operations, which are allocated among a plurality of second scheduler queues 214 via a multiplexer (mux) network 216. A picker 218 of each of the second scheduler queues 214 selects instruction operations buffered in the corresponding second scheduler queue 214 for allocation to a subset of execution pipes associated with the corresponding second scheduler queue 214 or for other allocation. Because instruction operations cannot be selected for execution directly from the first scheduler queues 212, the first scheduler queues 212 are referred to herein as “non-selectable scheduler queues 212” or “NSQs 212.” Conversely, because instruction operations are selectable from the second scheduler queues 214 for execution, each of the second scheduler queues 214 is referred to as a "selectable scheduler queue 214" or "SQ 214."
[0014] The floating point / SIMD unit 208 includes a rename module 220, a physical register file 222, a plurality of execution pipes 224, such as six execution pipes 224-1 through 224-6 in the illustrated embodiment, and a multi-modal distributed scheduler 202. The rename module 220 is disposed between the NSQ 212 and the selectable scheduler queues 214. As described in more detail with respect to FIG. 3, the rename module 220 performs rename operations on instruction operations received from the NSQ 212, including renaming architectural registers to physical registers in the physical register file 222, and outputs the renamed instruction operations to the multiplexer network 216 for distribution to a plurality of second scheduler queues, e.g., by distributing the operations to the two illustrated second scheduler queues 214-1, 214-2 using corresponding selectors 218-1, 218-2 and queue controllers 226. Each of the second scheduler queues 214 is responsible for buffering instruction operations for a corresponding respective subset of the plurality of execution pipes 224. For example, in the illustrated embodiment, second scheduler queue 214-1 buffers instruction operations for a subset consisting of execution pipes 224-1, 224-2, and 224-3, and second scheduler queue 214-2 buffers instruction operations for another subset consisting of execution pipes 224-4, 224-5, and 224-6.
[0015] 3 is a block diagram of a processor 300 including an FPU 305 that implements NSQ to buffer operations before renaming, according to some embodiments. The processor 300 may be used to implement some embodiments of the processing system 100 shown in FIG. 1 and the processor 200 shown in FIG. 2. An instruction front end (DE) 310 provides instructions or operations to the FPU 305 and other entities within the processor 300, including a load store unit 315. Instructions provided to the FPU 305 often include load operations that are used to load data stored in a memory, such as the memory 105 shown in FIG. 1. In the illustrated embodiment, the DE 310 provides load operations to the FPU 305 and the load store unit 315 simultaneously.
[0016] Operations received from DE 310 are first buffered in NSQ 320. Buffering load operations before further processing in FPU 305 allows additional time for load store unit 315 to read the load data from memory before allocating a physical register file to the load operation, thereby increasing the number of available physical register files. After buffering the load operation, the operation is popped from NSQ 320 and provided to decode / translate module 325, which decodes the instruction into one or more instruction operations (e.g., micro-ops or uops) and translates the virtual address contained in the instruction or operation. The decoded load operation is provided to renamer 330, which renames architectural registers and allocates physical registers from physical register file 335. Depending on the renaming and allocation of physical registers, the load operation is provided to selectable scheduler queue 340.
[0017] The FPU 305 performs various processes of architectural register renaming and physical register allocation depending on the relative age of the buffering in the NSQ 320 and the load data from the load store unit 315. The FPU 305 includes a free list 345 that indicates the physical register numbers of the available physical registers in the physical register file 335. The free list 345 is implemented using storage components such as memory, registers, buffers, etc. The FPU also includes a physical register number buffer 350 that stores the physical register numbers of a subset of the available physical registers. A physical register number (and corresponding physical register) is assigned to the load operation from either the free list 345 or the physical register number buffer 350 depending on whether the load operation is popped from the NSQ 320 before or after the load store unit 315 returns the load data. In some embodiments, a physical register number is assigned to the load operation from the free list 345 in response to the load operation being popped from the NSQ 320 before the load store unit 315 returns the load data. Before the load operation is popped from NSQ 320, a physical register number is assigned to the load operation from physical register number buffer 350 in response to load data being returned by load store unit 315. Load mapper 355 assigns a physical register number to the load operation in response to load data being returned from load store unit 315.
[0018] Additional circuitry is incorporated into FPU 305 to coordinate the allocation of physical register numbers to load operations by free list 345 and physical register number buffer 350. Some embodiments of FPU 305 include a mapping structure 360 that maps retirement identifiers to physical register numbers that are assigned to load operations. Information stored in mapping structure 360 is used to track assigned physical register numbers until the corresponding load operation retires. Thus, a physical register number assigned to an in-flight load operation is not assigned to other operations. Information stored in mapping structure 360 is used to determine whether a physical register number has already been assigned to a load operation. For example, an entry in mapping structure 360 is checked before assigning a physical register number to a load operation that is popped off NSQ 320 to ensure that a different physical register number has not already been assigned to the load operation in response to load store unit 315 returning the load data for the load operation. Thus, mapping structure 360 prevents different physical register numbers from being assigned to the same load operation. Multiplexers 365 , 370 are used to coordinate the distribution of information from the free list 345 and physical register number buffer 350 to the mapping structure 360 and renamer 330 .
[0019] FPU 305 provides the load operation to data path 375 responsive to load store unit 315 successfully loading the load data and the necessary physical registers being allocated for the load operation.
[0020] 4 is a flow diagram of a method 400 for selectively allocating physical register numbers from a free list or physical register number buffer according to some embodiments. Method 400 is implemented in an FPU, such as FPU 150 shown in FIG. 1, floating point / SIMD unit 208 shown in FIG. 2, and some embodiments of FPU 305 shown in FIG. 3.
[0021] In block 405, the FPU buffers the load operation in an NSQ (such as NSQ 212 shown in FIG. 2 or NSQ 320 shown in FIG. 3), and in block 410, the load operation is dispatched to a load store unit, such as load store unit 315 shown in FIG. 3. The buffering of the load operation in block 405 is performed simultaneously with the dispatching of the load operation in block 410.
[0022] At decision block 415, the FPU determines whether the load operation has been popped from the NSQ. If so, the method 400 proceeds to block 420 where a physical register number is assigned to the load operation from a free list, such as free list 345 shown in Figure 3. If the load operation has not yet been popped from the NSQ, the method 400 proceeds to decision block 425.
[0023] At decision block 425, the FPU determines whether the load store unit has returned the load data requested by the load operation. If so, method 400 proceeds to block 430 where a physical register number is assigned to the load operation in a physical register number buffer, such as physical register number buffer 350 shown in Figure 3. If the load store unit has not yet returned the load data, method 400 proceeds to decision block 415 where another iteration of method 400 is performed. Some embodiments of method 400 are repeated once every clock cycle, while other iteration time intervals are used in other embodiments.
[0024] In some embodiments, the above apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also called integrated circuit packages or microchips), such as the FPUs described above with reference to FIGS. 1-4. Electronic design automation (EDA) and computer-aided design (CAD) software tools are used to design and manufacture these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs include code executable by a computer system that operates the computer system to operate on the code representing the circuits of the one or more IC devices to perform at least a portion of the processing to design or adapt a manufacturing system to manufacture the circuits. This code may include instructions, data, or a combination of instructions and data. The software instructions representing the design or manufacturing tools are typically stored in a computer readable storage medium accessible by the computing system. Similarly, the code representing one or more phases of the design or manufacture of the IC devices may be stored in or accessed from the same or different computer readable storage medium.
[0025] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tape, magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or micro-electromechanical systems (MEMS) based storage media. The computer-readable storage medium (e.g., system RAM or ROM) may be internal to the computing system, the computer-readable storage medium (e.g., a magnetic hard drive) may be permanently attached to the computing system, the computer-readable storage medium (e.g., an optical disk or Universal Serial Bus (USB)-based flash memory) may be removably attached to the computing system, or the computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to the computer system via a wired or wireless network.
[0026] In some embodiments, some aspects of the above techniques may be implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored in or tangibly embodied on a non-transitory computer-readable storage medium. The software may include instructions and specific data that, when executed by one or more processors, operate the one or more processors to perform one or more aspects of the above techniques. The non-transitory computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, a cache, a random access memory (RAM), or one or more other non-volatile memory devices, etc. The executable instructions stored on the non-transitory computer-readable storage medium may be source code, assembly language code, object code, or other instruction formats that can be interpreted or executed by one or more processors.
[0027] In addition to the above, it should be noted that not all activities or elements described in the summary description are required, some of the specific activities or devices may not be required, one or more additional activities may be performed, and one or more additional elements may be included. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will appreciate that various changes and modifications can be made without departing from the scope of the invention as set forth in the claims. Thus, the specification and drawings should be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the invention.
[0028] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, the benefits, advantages, solutions to problems, and features by which any benefit, advantage, or solution may occur or be manifested are not to be construed as critical, essential, or essential features of any or all claims. Moreover, the specific embodiments described above are illustrative only, as the disclosed invention may be modified and practiced in different but similar manners apparent to those skilled in the art having the benefit of the teachings herein. No limitations are intended to the details of construction or design herein shown, other than as set forth in the appended claims. It is therefore apparent that the specific embodiments described above may be altered or modified, and all such variations are considered to be within the scope of the disclosed invention. Accordingly, the protection sought herein is set forth in the appended claims.
Claims
1. a first scheduler queue configured to temporarily buffer a load operation simultaneously with the load store unit obtaining load data for an operand to be loaded by the load operation; a renamer configured to, in response to receiving the load operation from the first scheduler queue, rename an architectural register used by the load operation and assign a physical register number to the load operation; a second set of scheduler queues configured to receive the load operations from the renamer and to store the load operations prior to execution; Device.
2. A physical register file; a storage component configured to store a free list indicating physical register numbers of allocatable entries in the physical register file.
2. The apparatus of claim 1.
3. a physical register number buffer configured to store a subset of the physical register numbers of the allocatable entries in the physical register file; 3. The apparatus of claim 2.
4. a physical register number is assigned to the load operation from the physical register number buffer in response to the load store unit returning the load data before the load operation is popped from the first scheduler queue; 4. The apparatus of claim 3.
5. a load mapper that assigns a physical register number to the load operation in response to the load data being returned by the load store unit; 4. The apparatus of claim 3.
6. a mapping structure for mapping a retirement identifier to a physical register number assigned to a load operation to inhibit assignment of another physical register number to the same load operation; 6. The apparatus of claim 5.
7. the second set of scheduler queues configured to store operations for a plurality of execution pipes; each second scheduler queue in the set stores operations for a different subset of the plurality of execution pipes; 2. The apparatus of claim 1.
8. A method for scheduling a load operation in a first scheduler queue, the method comprising: buffering the load operation in a first scheduler queue at the same time that a load store unit obtains load data for an operand to be loaded by the load operation; at a renamer, in response to receiving the load operation from the first scheduler queue, renaming an architectural register used by the load operation and assigning a physical register number to the load operation; providing the load operation from the renamer to any of a set of second scheduler queues configured to store the load operation prior to execution; method.
9. storing physical register numbers of allocatable entries in the physical register file in a free list.
9. The method of claim 8.
10. storing a subset of the physical register numbers of the allocatable entries in the physical register file in a physical register number buffer; Further comprising:
10. The method of claim 9.
11. and assigning a physical register number from the physical register number buffer to the load operation in response to the load store unit returning the load data before the load operation is popped from the first scheduler queue. The method of claim 10.
12. assigning the physical register number includes assigning the physical register number to the load operation using a load mapper in response to the load data being returned by the load store unit.
12. The method of claim 11.
13. and mapping a retirement identifier to a physical register number assigned to the load operation to inhibit another physical register number from being assigned to the same load operation.
13. The method of claim 12.
Citation Information
Patent Citations
System and method for merging partially-writhing result in retirement phase
JP2017228267A
Renaming with generation numbers
US20150309796A1