Matrix scheduler and matrix scheduling method for aggregating dependency information into column

The matrix scheduler addresses the issue of increased wiring delay in miniaturized semiconductor technology by aggregating dependency information into columns, reducing circuit complexity and delay, and enabling more scheduler entries for improved processor performance.

JP2025094604APending Publication Date: 2025-06-25FUJITSU LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023210272
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

The miniaturization of semiconductor technology has led to an increase in wiring delay, which hinders the performance improvement of processors due to the larger proportion of circuit delay caused by wiring and gate delays, making it difficult to increase the number of scheduler entries.

Method used

A matrix scheduler that aggregates dependency information into columns, utilizing a matrix table with cells that store a dep signal indicating dependency relationships, and includes a grant signal to manage instruction execution, reducing the need for pend signals and minimizing circuit increase and delay.

Benefits of technology

This approach allows for minimizing circuit increase and delay while increasing the number of scheduler entries, simplifying control logic and reducing circuit volume by half, thus enhancing processor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094604000001_ABST
    Figure 2025094604000001_ABST
Patent Text Reader

Abstract

To provide a matrix scheduler that minimizes a circuit increase and a circuit delay when the number of entries in the scheduler is increased.SOLUTION: A matrix scheduler includes a matrix table 11 having at most N lines and M columns of cells (N and M are natural numbers of 2 or greater) and a 1 line and M columns of cells for storing a grant signal, and includes a processing unit to, in each cell of the matrix table 11, store a dep signal 15 indicating a dependency with a producer of the cell's entry, and when a command is issued from the scheduler, set a 1 to a bit of the grant signal corresponding to an issued scheduler entry, and when a product of the dep signal 15 of a cell in a line direction of the matrix table and an inverted signal of the grant signal becomes all 0 bit, execute a command.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a matrix scheduler that aggregates dependency information into columns and a matrix scheduling method.

Background Art

[0002] In the core part of a processor, there is an optimization method called instruction scheduling that arranges instruction sequences to be executed in the shortest possible time. In recent years, the number of scheduler entries has been increasing to meet the high performance requirements of processors.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Due to the miniaturization of semiconductor technology, the proportion of wiring delay in circuit delay has become larger. The increase in the circuit amount leads not only to gate delay due to the number of transistor stages but also to the deterioration of wiring delay due to the increase in circuit area, which becomes an obstacle to increasing the number of scheduler entries and hinders performance improvement.

[0005] On one side, it aims to minimize circuit increase and circuit delay when increasing the number of scheduler entries.

Means for Solving the Problems

[0006] On one side, a matrix scheduler that aggregates dependency information into columns has a matrix table with cells of at most N rows and M columns (N and M are natural numbers greater than or equal to 2) and a cell of 1 row and M columns for storing the grant signal. Each cell of the matrix table stores a dep signal indicating that it is dependent on the producer of the entry of the cell. When an instruction is issued from the scheduler, 1 is set in the bit of the grant signal corresponding to the issued scheduler entry. When the product of the dep signal of the cells in the row direction of the matrix table and the inverted signal of the grant signal all become 0 bits, the instruction is executed, and it includes a processing unit.

Advantages of the Invention

[0007] On one side, it is possible to minimize circuit increase and circuit delay when increasing the number of entries in the scheduler.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

DETAILED DESCRIPTION OF THE INVENTION

[0009] 〔A〕Embodiment Hereinafter, an embodiment will be described with reference to the drawings. However, the embodiments shown below are merely examples, and there is no intention of excluding various modifications and applications of technologies not explicitly shown in the embodiments. That is, the present embodiment can be variously modified and implemented without departing from its gist. Also, each figure is not intended to include only the components shown in the figure, but can include other components and the like.

[0010] 〔A-1〕Related Example FIG. 1 is a block diagram schematically showing a configuration example of the processor core 1.

[0011] The processor core 1 includes an instruction cache 61, an instruction buffer 62, an instruction decoder 63, Scheduler-A 64a, Scheduler-E 64b, an arithmetic execution unit 7, and a load / store unit 8.

[0012] The arithmetic execution unit 7 includes a physical GPR (General Purpose Register) 71, a fixed-point arithmetic unit 72, and an address generation arithmetic unit 73.

[0013] The load / store unit 8 includes an LDSTQ (Load Store Queue) 81 and a data cache 82.

[0014] Instructions for instructing the operation of the processor core 1 are stored in the instruction cache 61. The instruction codes read from the instruction cache 61 are accumulated in the instruction buffer 62 and sequentially sent to the instruction decoder 63.

[0015] The instruction decoder 63 interprets the instructions and registers information such as the instruction codes with the scheduler.

[0016] The scheduler accumulates instructions and speculatively issues instructions to the arithmetic unit and the cache memory from the instructions that are ready.

[0017] In FIG. 1, it includes a Scheduler-E64b that accumulates arithmetic instructions and a Scheduler-A64a that accumulates memory access instructions such as load and store.

[0018] Since arithmetic instructions are accumulated in the Scheduler-E64b, its output is connected to the physical GPR 71 and the fixed-point arithmetic unit 72 of the arithmetic execution unit 7. Since memory access instructions are accumulated in the Scheduler-A64a, its output is connected to the physical GPR 71 and the address generation arithmetic unit 73 of the arithmetic execution unit 7, and the load / store unit 8 is connected thereto.

[0019] The memory access instructions output from the Scheduler-A64a refer to the physical GPR 71 to calculate the access address, and perform processing such as addition by the address generation arithmetic unit 73 based on the read data.

[0020] The obtained address is sent to the load store unit 8, stored in the queue named LDSTQ81, and sequentially accesses the data cache. When the instruction is a load, data is output from the cache, and in the case of a fixed-point load, the data is written to the physical GPR71.

[0021] The arithmetic instruction output from the Scheduler-E64b refers to the physical GPR71, executes a fixed-point operation based on the read data, and writes the result to the physical GPR71. Although omitted in Figure 1, the Scheduler-E64b can also store floating-point operation instructions. In that case, the floating-point operation instruction issued from the Scheduler-E64b executes a floating-point operation by referring to the physical GPR71 in the arithmetic execution unit 7.

[0022] The scheduler controls the issue order between dependent instructions and has a function to perform adjustment control for out-of-order issue from issuable instructions.

[0023] Figure 2 is a diagram illustrating an instruction sequence in a related example.

[0024] For example, in the instruction sequence as shown in Figure 2, (1) sub subtracts by referring to the fixed-point registers x1 and x2, and x3 is updated with the result. (2) mul multiplies x1 and x2, and x4 is updated. (3) add adds x3 and x4, and x5 is updated.

[0025] (3) add uses the result of (1) sub for x3 and the result of (2) mul for x4. Therefore, it can be said that (1) sub and (2) mul have a dependency relationship with (3) add.

[0026] The scheduler can issue instructions regardless of the original instruction order. However, in such a case with a dependency relationship, the issue of (3) add must wait for the execution of (1) sub and (2) mul.

[0027] In the case of instructions with such dependencies, an instruction that updates a register, such as (1) sub or (2) mul, is called a producer, and an instruction that uses the register updated by the producer is called a consumer.

[0028] The scheduler manages this producer - consumer relationship with a Dependency Matrix Table corresponding to cancellation after instruction issuance.

[0029] Figure 3 is a diagram for explaining the first state of the Dependency Matrix Table in the related example.

[0030] In Figure 1, a configuration using two schedulers was introduced. For simplicity of explanation, it will be described using a single scheduler composed of, for example, 8 entries. When the instruction sequence in Figure 2 is executed in this example, if (1) sub is registered in entry 3 of the scheduler, (2) mul is registered in entry 5, and (3) add is registered in entry 1 respectively, it will be as shown in Figure 3.

[0031] In the example shown in Figure 3, the rows of the matrix table 91 correspond to the consumer entry numbers, and the columns correspond to the producer entry numbers.

[0032] Each entry of the scheduler has 7 cells in the row direction, and its position indicates the producer entry number.

[0033] Each cell of the matrix holds 2 - bit information indicating the dependency relationships of pend96 and dep95. The dep signal indicates that there is a dependency with the producer of the corresponding entry, and pend indicates the execution status of that producer. When pend is 1, it indicates that the producer has not been executed yet. When it is 0 as shown by the reference numeral 92, it indicates that there is no dependency, or although there is a dependency, the producer has already been issued an instruction and the operation has been executed.

[0034] For example, in Figure 3, since the (3)add stored in Entry 1 depends on Entries 3 and 5 of (1)sub and (2)mul respectively, bit3 and bit5 indicate 1, and the other bits indicate 0. The rdy signal (see reference numeral 93) indicating that it is issuable from the scheduler becomes 1 only when all the bits of that entry are not 0. In the state of Figure 3, Entry 1 has rdy = 0 and cannot be issued from Select94 of the scheduler.

[0035] Figure 4 is a diagram for explaining the second state of the dependence resolution matrix table in the related example.

[0036] When sub is issued from the scheduler, as shown in Figure 4, since (1)sub is in Entry 3, signals are asserted in the column direction of the matrix table 91, and all the pends that are 1 are dropped to 0.

[0037] After the instruction is issued from Entry 5, all the pends in the row direction of Entry 1 (see reference numeral 92) become 0, entry1 has rdy = 1 (see reference numeral 93), and it becomes issuable from Select94.

[0038] In this way, the scheduler grasps the dependence relationship between instructions and performs issue control according to the order by updating the issue status of the producer.

[0039] An instruction issued from the scheduler may be canceled for some reason and returned to the scheduler. In that case, the cancel signal is asserted in the column direction, the signal is set back to 1, and the state of Figure 3 is restored. At this time, since only the pend signal that originally had a dependence can be set to 1, only the pend of the cell where the dep signal is 1 is set to 1.

[0040] Figure 5 is a circuit diagram for explaining the dep signal 95 and the pend signal 96 in the related example.

[0041] The dep signal 95 is set to 1 when the instruction is registered with the scheduler (allocate). Therefore, the result of ORing the loopback from the output of the FlipFlop for value retention and the OR gate 951 is set to the FlipFlop.

[0042] Valid is the valid of the scheduler entry and is a signal of the consumer. When the instruction is issued from the scheduler and the processing is completed, the valid of its own entry is cleared (in other words, released). Next, when another instruction is registered, the dep signal 95 needs to be 0. Therefore, the valid is ANDed by the AND gate 952 for setting the dep signal, and the dep signal 95 is reset in synchronization with valid = 0.

[0043] Also, the dep signal 95 is reset by the rst_dep signal separately. The rst_dep signal is a signal notified when the dependent producer is released from the scheduler and is a signal notified in the column direction of the table.

[0044] Since the scheduler entry from which the instruction has been released will have another instruction registered later, the dep signal is set to drop to prevent the set / reset operation of pend by another instruction. Therefore, the dep signal 95 is ANDed by the AND gate 961 for the set condition of the pend signal 96, and the case of dep = 0 is such that pend becomes 0.

[0045] The set_pend and rst_pend connected to the pend signal 96 are signals that match the set and rst shown in FIGS. 3 and 4.

[0046] The rst_pend signal becomes a signal of 1 by ANDing by the AND gate 962 when an instruction is issued from the scheduler.

[0047] The output from the AND gates 961 and 962 and the allocate are ORed by the OR gate 963 to become the pend signal 96.

[0048] Instructions issued by the scheduler may fail to execute due to cache misses or other reasons. In such a case, the instructions are reissued back to the scheduler. Since returning to the scheduler means returning to the unissued state, the pend signal of the consumer must be set again. Therefore, when an instruction returns to the scheduler, the set_pend signal is notified in the column direction.

[0049] The notification of the signal indicating that an instruction has been issued is only for one cycle, so the conventional method is called the pulse method here. In the pulse method, if the consumer is not present in the scheduler at the timing when the producer issues an instruction, the bit in the dependency_table cannot be set to 0.

[0050] In case the timing of the consumer being registered in the scheduler is later than the issue notification by the producer, it is necessary to separately have a management table or the like to grasp the issue status of the producer and refer to it immediately before scheduler registration, which raises concerns about an increase in circuit volume and circuit complexity.

[0051] 〔A-2〕Configuration Example FIG. 6 is a diagram for explaining the first state of the dependency resolution matrix table in the embodiment. FIG. 7 is a diagram for explaining the second state of the dependency resolution matrix table in the embodiment.

[0052] In this embodiment, a level-type Dependency Matrix Table that minimizes circuit increase is proposed by sharing the pend signal and using the resources that were in the matrix as column resources.

[0053] The dependency resolution matrix table in FIG. 6 is obtained by adding a grant signal to the upper part of the matrix table 11 as compared with the dependency resolution matrix table in the related example shown in FIG. 3.

[0054] The grant signal is a signal that is held and updated by a flip-flop (FF) or latch composed of a number of bits corresponding to the number of scheduler entries (in other words, the number of columns in the dependency table). By adding the grant signal, the pend signal of each cell disappears, and the number of bits of memory elements such as FFs is halved in each cell.

[0055] Also, in FIGS. 3 and 4, the values recorded in the matrix table 91 were the values of the pend signal, but in FIG. 5, the values recorded in the matrix table 11 are the values of the dep signal.

[0056] The grant signal is a signal with the same number of bits as the number of scheduler entries and is a signal indicating the issuance status of the producer. Since it is a signal indicating the issuance status of the producer, when an instruction is issued from Select14 of the scheduler as shown in FIG. 7, a 1 is set in the bit of the grant signal corresponding to the issued scheduler entry. In the case of cancellation, it is only necessary to set the grant signal to 0 again.

[0057] The grant signal 17 is notified to all entries of the scheduler.

[0058] Each cell of the matrix table 11 has only a dep signal indicating which scheduler entry it depends on, and when dep & ~grant shown in reference numeral 12 becomes 0 in all bits in the row direction, rdy = 1 shown in reference numeral 13 can be set.

[0059] By providing the grant signal in this way, the pend signal that each cell conventionally had becomes unnecessary. The dependency resolution matrix table has a table size determined by O(n^2) from its structure, and an increase in the number of scheduler entries greatly affects the circuit amount. In this embodiment, the circuit amount of memory elements such as FFs and latches including peripheral circuits can be halved.

[0060] The grant signal remains at 1 until it is canceled when 1 is set, continuously notifying the consumer of the issuance status at all times. The signal's operation is based on a level system.

[0061] Figure 8(a) is a circuit diagram for explaining the dep signal in the embodiment, and Figure 8(b) is a circuit diagram for explaining the grant signal in the embodiment.

[0062] The circuit related to the dep signal 15 and its set (in other words, the OR gate 151 and AND gate 152) shown in Figure 8(a) is the same as that in Figure 5.

[0063] The output of the dep signal 15 is ANDed with the grant signal by the AND gate 16 and output as shown in Figure 7.

[0064] The dep signal 15 is held by each entry with the same number of bits as the number of entries in the scheduler, and in the case of dependencies from multiple instructions, it can hold multiple bits set in one row in a decoded format.

[0065] In Figure 8(b), the set_grant signal is similar to the rst_pend signal in Figure 5, and the rst_grant signal and the set_pend signal in Figure 6 are signals where the set / reset is inverted. That is, when the producer is issued, the set_grant signal goes high and sets 1 in the grant. When the producer is canceled and returns to the scheduler, the rst_grant signal goes high and the grant is reset to 0. Therefore, although there is a difference in set / rst, it has the same movement as the pend signal 96 in Figure 5.

[0066] In Figure 5, the dep signal 95 was ANDed with the set of pend, but instead, in Figure 8(b), the valid is ANDed by the AND gate 171. This is the valid of the producer's scheduler entry and is a signal expected to have the same movement as the rst_dep in Figure 5 (that is, when the producer is released, the grant is set to 0).

[0067] The output from the AND gate 171 and the set_grant signal are input to the OR gate 172, and the output of the OR gate 172 becomes the grant signal 17.

[0068] The grant signal 17 has a function of setting or resetting a signal according to the issuance or cancellation status of the corresponding scheduler entry.

[0069] Comparing Fig. 5 and Fig. 8(b), although the circuit operations and expected movements are the same, by using the pend signal 96 as the grant signal 17, the circuits for the number of scheduler entries, which used to be separate, can be made into one common circuit in the column direction.

[0070] In particular, it is generally known that a FF (or a storage element such as a latch) has a larger number of transistors than logic gates such as AND and OR, and they can be replaced by one AND gate on the matrix table 11, and a significant reduction in the circuit amount can be expected.

[0071] The processor core in the embodiment has the same configuration as the processor core 1 shown in Fig. 1.

[0072] Instructions for instructing the operation of the processor core 1 are stored in the instruction cache 61, the instruction codes read from the instruction cache 61 are accumulated in the instruction buffer 62, and are sequentially sent to the instruction decoder 63.

[0073] The instruction decoder 63 interprets instructions and registers information such as instruction codes in the scheduler.

[0074] The scheduler accumulates instructions and speculatively issues instructions to the arithmetic unit or the cache memory from the instructions that are ready.

[0075] In FIG. 1, there is a Scheduler-E64b that stores arithmetic instructions and a Scheduler-A64a that stores memory access instructions such as load and store. Since arithmetic instructions are stored in Scheduler-E64b, its output is connected to the physical GPR71 of the arithmetic execution unit 7 and the fixed-point arithmetic unit 72.

[0076] Since memory access instructions are stored in Scheduler-A64a, its output is connected to the physical GPR71 of the arithmetic execution unit 7 and the address generation arithmetic unit 73, and a load / store unit 8 is connected thereto.

[0077] The memory access instruction output from Scheduler-A64a refers to the physical GPR71 to calculate its access address, and processes such as addition are performed by the address generation arithmetic unit 73 based on the read data. The obtained address is sent to the load / store unit 8, stored in a queue called LDSTQ81, and sequentially accesses the data cache 82.

[0078] When the instruction is a load, data is output from the cache, and when it is a fixed-point load, the data is written to the physical GPR71.

[0079] The arithmetic instruction output from Scheduler-E64b refers to the physical GPR71, executes fixed-point arithmetic based on the read data, and the result is written to the physical GPR71.

[0080] Although omitted in FIG. 1, Scheduler-E64b can also store floating-point arithmetic instructions. In that case, the floating-point arithmetic instruction issued from Scheduler-E64b refers to the physical GPR71 in the arithmetic execution unit 7 and the floating-point arithmetic is executed.

[0081] The scheduler has a function of controlling the issue order between dependent instructions and performing adjustment control for out-of-order issue from issuable instructions.

[0082] The operation execution unit 7 has a matrix table 11 having at most N rows and M columns (N and M are natural numbers of 2 or more) of cells and one row and M columns of cells for storing the grant signal. The operation execution unit 7 stores a dep signal 15 indicating that it is in a dependency relationship with the producer of the entry of the cell in each cell of the matrix table 11. When an instruction is issued from the scheduler, the operation execution unit 7 sets 1 in the bit of the grant signal 17 corresponding to the issued scheduler entry, and when the product of the dep signal 15 of the cells in the row direction of the matrix table 11 and the inverted signal of the grant signal 17 all becomes 0 bits, the instruction is executed.

[0083] FIG. 9 is a diagram illustrating an instruction sequence in the embodiment.

[0084] For example, in an instruction sequence as shown in FIG. 9, (1) ldr calculates an access address to memory with reference to x1 and x2 of the fixed-point register, and x3 is updated with the data read from the L1 cache using the result.

[0085] (2) mul multiplies x1 and x3, and x4 is updated. Since x3 uses the result updated by (1) ldr, it can be said that (2) mul has a dependency relationship with (1) ldr.

[0086] (3) add adds x1 and x4, and x5 is updated. Since x4 used by (3) add uses the result of (2) mul, it can be said that (2) mul has a dependency relationship with (3) add.

[0087] Although the scheduler can issue instructions regardless of the original instruction order, in such a case of having a dependency relationship, the issuance of (2) mul has to wait for the execution of (1) ldr, and the issuance of (3) add has to wait for the execution of (2) mul.

[0088] In the case of instructions with such dependencies, the (1)ldr for (2)mul and the (2)mul for (3)add are called producers, and the (2)mul for (1)ldr and the (3)add for (2)mul are called consumers.

[0089] The scheduler manages this producer - consumer relationship in a dependency resolution matrix table. There are two schedulers, Scheduler - A64a and Scheduler - E64b, and each can be both a producer and a consumer for the other.

[0090] 〔A - 3〕First Variation Figure 10 is a diagram for explaining the first state of the dependency resolution matrix table in the first variation.

[0091] In the first variation, it is explained when each scheduler is composed of 4 entries.

[0092] In this example, when the instruction sequence in Figure 2 is executed, if (1)ldr is registered in entry1 of Scheduler - A64a, (2)mul is registered in entry2 of Scheduler - E64b, and (3)add is registered in entry 3 of Scheduler - E64b, it will be as shown in Figure 10.

[0093] Since Scheuler - E64b and Scheduler - A64a can be both producers and consumers for each other, the matrix table 11a is composed of an 8 - row and 8 - column matrix. The row direction corresponds to the consumer, and the column direction corresponds to the producer.

[0094] Each consumer holds a 7 - bit bit - vector excluding itself in Scheduler - E64b and Scheduler - A64a in the row direction. These bit - vectors are also called dep signals, and each bit of the bit - vector indicates the position of the scheduler entry of the producer on which the consumer depends.

[0095] In Figure 10, since the left 4 bits correspond to Scheduler-E64b and the right 4 bits correspond to Scheduler-A64a, for example, (2)mul depends on (1)ldr registered in entry1 of Scheduler-A64a, so the corresponding bit is 1. Similarly, for (3)add, the bit corresponding to entry2 of Scheduler-E64b where (2)mul is registered is 1.

[0096] Grant_e[0:3] and grant_a[0:3] at the top of Figure 10 indicate the instruction execution status of the corresponding entries of Scheduler-E64b and Scheduler-A64a. When the bit here becomes 1, the instruction of the corresponding entry is issued from the scheduler and the operation or load is being executed, notifying the subsequent consumer that the dependency has been resolved. For example, when (1)ldr is issued, 1 is set in grant_a[1].

[0097] The circuit configuration of each cell in the matrix table 11a of Figure 10 is the same as that in (a) of Figure 8.

[0098] The dep signal 15 is made of memory elements such as FFs and latches, and is designed to hold the value set by the corresponding allocate signal when an instruction is registered in the scheduler.

[0099] When the corresponding producer is issued from the scheduler and the operation or load process is completed, it is released from the scheduler. At that time, the producer issues an instruction to rst the dep signal 15 to all consumers. This is the rst_dep signal, which resets the dep signal 15 to 0.

[0100] Also, the AND gate 152 that sets the dep signal 15 has the valid of the consumer's scheduler entry ANDed. By ANDing the valid, in case the rst_dep does not come and the valid of the scheduler becomes 0 due to a branch prediction miss or cancellation by an asynchronous interrupt, the dep_val also becomes 0.

[0101] The dep signal 15 is ANDed with the ~grant signal by the AND gate 16 (see reference numeral 12 in FIG. 10) and output as dep_not_grant. That is, the dep_not_grant is a signal indicating that there is a dependency on the corresponding producer but the grant has already come and the dependency has been resolved, or there is no dependency on the corresponding producer in the first place.

[0102] The consumer checks all the dep_not_grant signals in the row direction of the dependency table. If all of them are 0, it determines that the instruction can be issued and sets 1 to the rdy signal shown by reference numeral 13 in FIG. 10.

[0103] For example, (1) For a load, since all the dep signals in the row direction in FIG. 10 are 0, the dep_not_grant is of course all 0 and the rdy is set to 1. (2) and (3) have a place where the dep signal 15 has a 1, and since the corresponding grant signal is 0, there is a place where the dep_not_grant = 1, so the rdy = 0 and the instruction cannot be issued.

[0104] FIG. 11 is a diagram for explaining the second state of the dependency resolution matrix table in the first modification example. FIG. 12 is a diagram for explaining the third state of the dependency resolution matrix table in the first modification example. FIG. 13 is a diagram for explaining the fourth state of the dependency resolution matrix table in the first modification example.

[0105] As shown in Fig. 11, when (1) a load is issued, 1 is set in the corresponding grant_a[1]. Therefore, since all the dep_not_grant in the Scheduler-E entry2 of (2) mul become 0, rdy = 1.

[0106] Next, as shown in Fig. 12, when (2) a mul is issued, 1 is set in grant_e[2], the rdy of the Scheduler-E entry2 of (3) add becomes 1, and all instructions can be issued.

[0107] As shown in Fig. 13, when (1) a load is issued and after passing through an appropriate plurality of cycles 18, when the load process is completed, (1) the load is unregistered from Scheduler-A64a. At that time, the valid of Scheruler-A entry1 is reset (omitted in Fig. 13). On the other hand, rst_dep is notified to all the dep signals 15 in the column of the depnency_table, and the places where 1 is set are dropped to 0.

[0108] The dep signal 15 is reset when the source instruction of the dependency is released from the scheduler. The dep signal 15 is reset when the corresponding scheduler entry is released from the scheduler.

[0109] 〔A-4〕Second Modified Example The matrix table 11a described with reference to Figs. 10 to 13 etc. had each consumer hold the number of bits of the corresponding producer using a memory element such as an FF. In the second modified example, a method of encoding and holding these will be described. The first modified example may be called a decoding method, and the second modified example may be called an encoding method.

[0110] Fig. 14 is a diagram for explaining the dependency resolution matrix table in the second modified example.

[0111] In the decoding method described in the first modification example, the matrix table 11a was composed of memory elements such as FF, but in the second modification example shown in FIG. 14, it is replaced by a combination circuit (the internal circuit of the cell will be described later with reference to FIGS. 15 and 16).

[0112] Therefore, each consumer holds dep_val and dep_id[2:0] instead of the dep signal (see reference numerals 191a and 191b). Since there are as many dep_val and dep_id as the number of operands of the instruction, for example, an instruction that operates on two data holds two dep_val / ids.

[0113] Decode dep_val and id (see reference numerals 192a and 192b), and when all operands are ORed by the OR gate 193, it becomes equivalent to the dep signal in the decoding method.

[0114] The dep signal 15 is reset when the dependent source instruction is released from the scheduler. The dep signal 15 is reset when the corresponding scheduler entry is released from the scheduler.

[0115] FIG. 15 is a circuit diagram of each cell in the encoding method in the second modification example.

[0116] In the circuit diagram shown in FIG. 15, the inverted signal of grant and the dep signal are ANDed by the AND gate 20 and output as dep_not_grant. In the encoding method, compared with the decoding method, the FF and its peripheral circuits are eliminated, so the circuit is simpler.

[0117] FIG. 16 is a circuit diagram for explaining the dep_val signal and the dep_id signal in the second modification example.

[0118] dep_val31 becomes 1 when the producer exists somewhere in the scheduler, and dep_id32 is registered in the encoded form of the entry number of the producer.

[0119] It is set (allocated) when the consumer is registered in the reservation station. The allocate and dep_val31 are ORed by the OR gate 311.

[0120] When the consumer is released, the valid of the scheduler is ANDed by the AND gate 312 to set dep_val31 to 0. In the AND gate 312, the inverted signal of rst_dep, valid, and the output of the OR gate 311 are ANDed and output as dep_val31.

[0121] The dep_id32 has a selector 321 at the input to continuously hold the set value while valid is 1.

[0122] dep_val31 is reset upon the release of the producer in the same way as the dep signal of the decoding method (rst_dep). Although omitted in FIG. 16, rst_dep can be created by comparing dep_id32 and the id of the producer.

[0123] In this way, the level-based dependency resolution matrix table can hold its dependency information in an encoded form. Since the amount of circuitry for the encoding method and the decoding method can vary depending on the number of operands, the number of entries, etc., the method can be selected according to the embodiment.

[0124] 〔B〕Effect According to the matrix scheduler and the matrix scheduling method that aggregate the dependency information in columns in the above-described embodiment, for example, the following operational effects can be achieved.

[0125] The operation execution unit 7 has a matrix table 11 having cells of at most N rows and M columns (N and M are natural numbers of 2 or more) and cells of 1 row and M columns for storing a grant signal. The operation execution unit 7 stores a dep signal 15 indicating that it is in a dependency relationship with the producer of the entry of the cell in each cell of the matrix table 11. When an instruction is issued from the scheduler, the operation execution unit 7 sets a bit of the grant signal 17 corresponding to the issued scheduler entry to 1, and when the product of the dep signal 15 of the cells in the row direction of the matrix table 11 and the inverted signal of the grant signal 17 all becomes 0 bits, the instruction is executed.

[0126] This can minimize the circuit increase and circuit delay when increasing the number of scheduler entries. Also, it is no longer necessary to store the pend signal in each conventional cell, and the storage capacity of the storage unit can be halved.

[0127] Specifically, in the pulse method of the related example, the consumer referred to a table for managing the issuance status of the producer immediately before registering with the scheduler. On the other hand, in the level method in the embodiment, even when the consumer is registered with the scheduler later, the issuance status of the producer is always notified as the grant signal 17, so a table like the conventional issuance status management table is unnecessary, and a significant reduction in circuit volume and simplification of control can be expected. Also, in this embodiment, it is widely applicable regardless of the configuration of the scheduler or the dependency resolution matrix table.

[0128] The grant signal 17 has a function of setting or resetting the signal according to the issuance or cancellation status of the corresponding scheduler entry. The grant signal 17 is notified to all entries of the scheduler.

[0129] This can appropriately control the grant signal 17.

[0130] The dep signal 15 is held in a decoded format where each entry is held with the same number of bits as the number of entries in the scheduler, and when there are dependencies from multiple instructions, multiple bits can be set in one row. The dep signal 15 is reset when the source instruction of the dependency is released from the scheduler. The dep signal 15 is reset when the corresponding scheduler entry is released from the scheduler.

[0131] In this way, the dep signal 15 can be appropriately controlled.

[0132] The dep signal 15 is held in an encoding method where each entry encodes the number of the source scheduler for each operand of the instruction.

[0133] In this way, the circuit configuration in each cell can be made simpler compared to the decoding method.

[0134] [C] Others The disclosed technology is not limited to the above-described embodiments, and can be implemented with various modifications without departing from the spirit of the present embodiment. Each configuration and each process of the present embodiment can be selectively adopted as necessary, or can be appropriately combined.

[0135] [D] Supplementary Note Regarding the above embodiments, the following supplementary note is further disclosed.

[0136] (Supplementary Note 1) It has a matrix table having at most N rows and M columns (N and M are natural numbers of 2 or more) of cells and 1 row and M columns of cells for storing the grant signal. In each cell of the matrix table, a dep signal indicating that the entry of the cell is in a dependency relationship with the producer is stored. When an instruction is issued from the scheduler, 1 is set in the bit of the grant signal corresponding to the issued scheduler entry, and when the product of the dep signal of the cells in the row direction of the matrix table and the inverted signal of the grant signal all becomes 0 bits, the instruction is executed. A matrix scheduler that aggregates dependency information into columns and includes a processing unit.

[0137] (Appendix 2) The grant signal has a function of setting or resetting the signal according to the issuance or cancellation status of the corresponding scheduler entry. The matrix scheduler according to Appendix 1, which aggregates dependency information into columns.

[0138] (Appendix 3) The grant signal is notified to all entries of the scheduler. The matrix scheduler according to Appendix 1 or 2, which aggregates dependency information into columns.

[0139] (Appendix 4) The dep signal is held by each entry in a decoding format with the same number of bits as the number of scheduler entries, and multiple bits can be set in one row when there are dependencies from multiple instructions. The matrix scheduler according to Appendix 1 or 2, which aggregates dependency information into columns.

[0140] (Appendix 5) The dep signal is held by each entry in an encoding method that encodes the number of the source scheduler for each operand of the instruction. The matrix scheduler according to Appendix 1 or 2, which aggregates dependency information into columns.

[0141] (Appendix 6) The dep signal is reset when the source instruction is released from the scheduler. The matrix scheduler according to Appendix 1 or 2, which aggregates dependency information into columns.

[0142] (Appendix 7) The dep signal is reset when the corresponding scheduler entry is released from the scheduler. The matrix scheduler according to Appendix 1 or 2, which aggregates dependency information into columns.

[0143] (Appendix 8) It has a matrix table having at most N rows and M columns (N and M are natural numbers of 2 or more) of cells and one row and M columns of cells for storing a grant signal, In each cell of the matrix table, a dep signal indicating that the entry of the cell is in a dependency relationship with the producer is stored, When an instruction is issued from the scheduler, a bit of the grant signal corresponding to the issued scheduler entry is set to 1, and when the product of the dep signal of the cells in the row direction of the matrix table and the inverted signal of the grant signal all becomes 0 bits, the instruction is executed. A matrix scheduling method in which a computer executes a process and aggregates dependency information into columns.

[0144] (Appendix 9) The grant signal has a function of setting or resetting a signal according to the issuance or cancellation status of the corresponding scheduler entry. The matrix scheduling method according to Appendix 8, which aggregates dependency information into columns.

[0145] (Appendix 10) The grant signal is notified to all entries of the scheduler. The matrix scheduling method according to Appendix 8 or 9, which aggregates dependency information into columns.

[0146] (Appendix 11) The dep signal is held by each entry in a decoding format having the same number of bits as the number of entries of the scheduler, and a plurality of bits can be set in one row when there are dependencies from a plurality of instructions. The matrix scheduling method according to Appendix 8 or 9, which aggregates dependency information into columns.

[0147] (Appendix 12) The dep signal is held by each entry in an encoding method in which the number of the scheduler of the dependency source is encoded for each operand of the instruction. The matrix scheduling method that aggregates the dependency information described in Appendix 8 or 9 into columns.

[0148] (Appendix 13) The dep signal is reset when the source instruction of the dependency is released from the scheduler. The matrix scheduling method that aggregates the dependency information described in Appendix 8 or 9 into columns.

[0149] (Appendix 14) The dep signal is reset when the corresponding scheduler entry is released from the scheduler. The matrix scheduling method that aggregates the dependency information described in Appendix 8 or 9 into columns.

Explanation of symbols

[0150] 1: Processor core 11, 11a, 91: Matrix table 14, 94: Select 15, 95: Dep signal 151, 172, 193, 311, 951, 963: OR gate 152, 16, 171, 20, 312, 952, 961, 962: AND gate 16: AND gate 17: Grant signal 18: Cycle 321: Selector 61: Instruction cache 62: Instruction buffer 63: Instruction decoder 64a: Scheduler-A 64b: Scheduler-E 7: Arithmetic execution unit 71: Physical GPR 72: Fixed-point arithmetic unit 73: Address generation arithmetic unit 8: Load / store unit 81: LDSTQ 82: Data cache 96: Pend signal

Claims

1. It has a matrix table having at most N rows and M columns (N and M are natural numbers of 2 or more) of cells and a 1-row and M-column cell for storing a grant signal, In each cell of the matrix table, a dep signal indicating that it is in a dependency relationship with the producer of the entry of the cell is stored, When an instruction is issued from the scheduler, a bit of the grant signal corresponding to the issued scheduler entry is set to 1, and when the product of the dep signal of the cells in the row direction of the matrix table and the inverted signal of the grant signal all becomes 0 bits, the instruction is executed, A matrix scheduler that aggregates dependency information into columns and includes a processing unit.

2. The grant signal has a function of setting or resetting the signal according to the issuance or cancellation status of the corresponding scheduler entry, The matrix scheduler according to claim 1, which aggregates dependency information into columns.

3. The grant signal is notified to all entries of the scheduler, The matrix scheduler according to claim 1 or 2, which aggregates dependency information into columns.

4. The dep signal is held in a decoding format in which each entry holds a number of bits equal to the number of entries of the scheduler, and when there are dependencies from a plurality of instructions, a plurality of bits can be set in one row, The matrix scheduler according to claim 1 or 2, which aggregates dependency information into columns.

5. The dep signal is held in an encoding method in which each entry encodes the number of the scheduler of the dependency source for each operand of the instruction, The matrix scheduler according to claim 1 or 2, which aggregates dependency information into columns.

6. The dep signal is reset when the dependency source instruction is released from the scheduler, The matrix scheduler according to claim 1 or 2, which aggregates dependency information into columns.

7. The dep signal is reset when the corresponding scheduler entry is released from the scheduler, The matrix scheduler according to claim 1 or 2, which aggregates dependency information into columns.

8. It has a matrix table having at most N rows and M columns (N and M are natural numbers of 2 or more) of cells and a 1-row and M-column cell for storing a grant signal, In each cell of the matrix table, a dep signal indicating that it is in a dependency relationship with the producer of the entry of the cell is stored, When an instruction is issued from the scheduler, a bit of the grant signal corresponding to the issued scheduler entry is set to 1, and when the product of the dep signal and the inverted grant signal of the cells in the row direction of the matrix table all become 0 bits, the instruction is executed. A matrix scheduling method in which a computer executes a process and aggregates dependency information into columns.

Citation Information

Patent Citations

  • Parallel computer and compiler

    JP1994028324A