Configurable Inter-Processor Synchronization System

The direct register-level communication system addresses synchronization delays in multi-core systems by using point-to-point synchronization chains, enabling efficient parallel processing and reduced latency in processor synchronization.

CN111382112BActive Publication Date: 2025-07-15KALRAY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201911353795.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-27
Filing Date
2019-12-25
Publication Date
2025-07-15
Estimated Expiration
2039-12-25

AI Technical Summary

Technical Problem

In the prior art, the synchronization scheme between processors has a delay problem in multi-core systems, especially the synchronization of processors outside the group depends on a shared communication channel, resulting in an increase in delay.

Method used

Using a configurable point-to-point synchronization link, low-latency inter-processor synchronization is achieved by introducing synchronization registers and gate structures into each processor, and the notification bit state is propagated on the link using waiting and notification machine instructions to form forward and backward synchronization units to optimize the synchronization process.

Benefits of technology

It realizes low-latency inter-processor synchronization, supports efficient communication between parallel processors, especially iterative non-independent loops and parallel processing of vector loops, reducing synchronization delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111382112B_ABST
    Figure CN111382112B_ABST
Patent Text Reader

Abstract

The present invention relates to an inter-processor synchronization system, comprising a plurality of processors (PEs); a plurality of unidirectional notification lines (FN, BN) connecting the processors in a chain; in each processor (PEi): a synchronization register (FE) having bits respectively associated with the notification lines, the synchronization register (FE) being connected to record the respective states of the upstream notification lines (FNin) propagated by the upstream processor (PEi-1), and a gate (12) controlled by a configuration register (FM) to propagate the states of the upstream notification lines (FNin) on the downstream notification lines (FNout) to the downstream processor (PEi+1).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the synchronization of multiple processors that run threads of the same program sharing the same resources in parallel, and more particularly, to an inter-processor point-to-point communication system that allows processors to communicate directly with other processors through a register file. Background Art

[0002] Figure 1 An example of a 4x4 processor PE array that can be integrated on the same chip is shown. The processors are then also referred to as "cores" and form a multi-core system. The cores or processor PEs can be connected to a common bus, which is also connected to a shared memory MEM. The memory MEM can contain program code run by the processors and working data of the programs.

[0003] Patent application US2015-0339256 discloses a point-to-point inter-processor communication technology, in which processors can be divided into groups of four adjacent processors as shown in the figure. Each processor in the group is connected to the other three processors in the group through wired point-to-point links. The point-to-point links are designed to allow each processor to directly write a notification into a register of a specified register file of any other processor in the group.

[0004] Therefore, the point-to-point links mentioned here are physical links that directly transfer bit states between processor registers. Do not confuse these physical links with general communication channels between processors (e.g., buses or on-chip networks), which also allow data transfer between registers, but in a software-based manner, by running general instructions on the processors, and the processors participate in data exchange through the shared memory.

[0005] More specifically, the processor has a group synchronization instruction in its instruction set, which simultaneously runs a wait command with a first parameter and a notification command with a second parameter. The wait command causes the processor to pause and wait for a bit pattern corresponding to the bit pattern transmitted in its parameter to appear in an event register. The notification command causes the valid bits of its parameter to be directly written into the event registers of other processors in the group. This writing only occurs when the processor exits its waiting state. Summary of the Invention

[0006] Generally, an inter-processor synchronization system is provided, including multiple processors; multiple one-way notification lines that connect processors in a chain; in each processor: a synchronization register having bits respectively associated with the notification lines, the synchronization register being connected to record the corresponding states of the upstream notification lines propagated by upstream processors; and a gate controlled by a configuration register to propagate the states of the upstream notification lines on the downstream notification lines to downstream processors.

[0007] Each processor may be configured to selectively activate a downstream notification line based on parameters of a notification machine instruction run by the processor.

[0008] Each processor may be configured to pause the running of a corresponding program based on parameters of a wait machine instruction run by the processor, and the pause is triggered when a synchronization register contains a valid bit pattern corresponding to the parameters of the wait instruction.

[0009] Each processor may be configured to reset the synchronization register when the pause is triggered.

[0010] The wait instruction and the notification instruction may form part of a single machine instruction executable by the processor.

[0011] The configuration register may include bits respectively associated with upstream notification lines, and the gates are configured to selectively propagate the state of the upstream notification lines based on the corresponding states of the bits in the configuration register.

[0012] An inter-processor synchronization method is also provided, including the following steps: through a plurality of processors in a row connection chain, the rows are configured to transmit corresponding notification bits in the same direction; in the first processor of the chain, a notification bit is sent to a second processor after the first processor in the chain; and in the second processor, depending on the state of local configuration bits, the notification bit is propagated to a third processor after the second processor in the chain.

[0013] The second processor may perform the following steps: save the notification bit in the synchronization register; run a wait machine instruction with parameters to stop the processor; and release the processor from the stopped state when the synchronization register contains a bit pattern corresponding to the parameters of the wait instruction.

[0014] The second processor may perform the following steps: run a notification machine instruction with parameters; and configure the notification bit to be sent to the third processor according to the notification instruction parameters.

[0015] The second processor may reset the synchronization register when exiting the stopped state. Description of the Drawings

[0016] In combination with the accompanying drawings, embodiments will be described in the following non-limiting description, where:

[0017] Previously described Figure 1 is a block diagram of a processor array divided into several groups, where the processors may communicate through point-to-point links;

[0018] Figure 2 is a block diagram of an embodiment of a processor array, where all processors are connected in a chain of notification links;

[0019] Figure 3 is a block diagram of an embodiment of a forward synchronization unit of a processor;

[0020] Figure 4 is a block diagram of an embodiment of a backward synchronization unit of a processor; and

[0021] Figure 5 shows an example configuration of a processor chain. Detailed Description

[0022] Figure 1 The known structure in each group achieves near-zero latency synchronization among the four processors in each group. Processors outside the group can also participate in the synchronization scheme, but synchronization between groups is software-based, using a communication channel shared by all processors, which significantly increases latency.

[0023] Figure 2 is a block diagram of an embodiment of a processor array organized according to a synchronization structure that allows any number of processors in the array to participate in a synchronization scheme with low latency. More specifically, the processors are organized in a chain of configurable point-to-point synchronization links. By "chain", it can be understood that each processor is physically connected to only one other processor or upstream processor in front of it, and physically connected to only one other processor or downstream processor behind it. As shown by the arrows, each link between two processors can be bidirectional, and the chain can be closed. Preferably, the chain is configured such that each processor is connected to two physically adjacent processors. As shown, the chain can pass through the processors row by row and reverse direction from one row to the next.

[0024] Each processor includes a pair of synchronization units, a first unit FSYNC that manages the link in the so-called "forward" direction, and another unit BSYNC that manages the link in the opposite "backward" direction. Each forward FN or backward FN link between two processors can include several physical rows, each physical row being designed to convey the state of bits representing events or notifications. Thus, such a multi-row link can convey several notifications distinguished by the ranking of the rows.

[0025] FSYNC and BSYNC units extend the execution units traditionally provided in a processor and can be configured to respond to two dedicated machine instructions, namely a wait instruction and a notify instruction. As described in the above US2015-0339256 patent application, the wait instruction and the notify instruction can be part of a single machine instruction called SYNCGROUP, which has two parameters, one of which identifies a wait channel and the other identifies a notify channel. If the wait channel parameter is zero, the instruction behaves like a simple notify instruction. Conversely, if the notify channel parameter is zero, the instruction behaves like a simple wait instruction. When the SYNCGROUP instruction identifies at least one wait channel, its execution has the particularity of putting the processor in a waiting state before any notification is issued. Once the processor is released from its waiting state, a notification is issued.

[0026] In the present disclosure, the SYNCGROUP instruction can have a 64-bit composite parameter, which is divided into 16 forward notification bits notifyF, 16 backward notification bits notifyB, 16 forward wait channel bits waitclrF, and 16 backward wait channel bits waitclrB.

[0027] Figure 3 is a block diagram of an embodiment of the forward synchronization unit FSYNC of processor PEi, where i is the rank of the processor in the chain. The processor includes execution units configured to respond to machine instructions (including SYNCGROUP instructions) running on the processor. The processor also includes a register file, which includes registers dedicated to synchronization, called Inter-Processor Event (IPE) registers. For example, a 64-bit IPE register can have a 16-bit forward event field FE, a 16-bit backward event field BE, a 16-bit forward mode field FM, and a final 16-bit backward mode field BM.

[0028] Figure 3 The FSYNC unit shown for managing the forward notification link FN uses the FE and FM fields and the waitclrF and notifyF parameters of the SYNCGROUP instruction. The BSYNC unit for managing the backward notification link BN uses the BE and BM fields of the IPE register and the waitclrB and notifyB parameters, which will be described later.

[0029] The FE field of the IPE register is configured to record notifications generated by the previous processor PEi-1 that arrive on 16 incoming forward notification lines FNin. For example, when the corresponding notification line goes to '1', each bit initially '0' in the FE field goes to '1', and this bit remains '1' even if the line subsequently goes back to '0'.

[0030] The 16 bits of the mode field FM of the IPE register form the first input of a bitwise AND gate 12, with the other inputs receiving 16 incoming notification lines Fnin. The 16-bit output of gate 12 contributes to the status of 16 outgoing forward notification lines FNout, thus leading to the next processor PEi+1.

[0031] Thus, depending on the bits contained in the mode field FM, the status of the incoming lines Fnin can be blocked or propagated individually on the outgoing notification lines FNout. In other words, notifications from the previous processor PEi-1 can be selectively propagated to the next processor PEi+1.

[0032] The current processor PEi can also send forward notifications to the next processor. To this end, the bits of the notifyF parameter of the synchronization instruction run by the processor Pei are combined by a bitwise OR gate 14 to the output of the AND gate 12. Thus, regardless of the status of the corresponding bits of the output of the AND gate 12, any bit at "1" in the notifyF parameter is transmitted on the corresponding outgoing notification line.

[0033] When a processor Pei with a non-zero waitclrF parameter runs a synchronization instruction, the processor is placed in a waiting state, and the waitclrF parameter determines the conditions required to release the processor from its waiting state. More specifically, the waitclrF parameter identifies, by corresponding bit positions, the notifications expected by the processor Pei. All incoming notifications are recorded in the FE register field, including those that pass through the processor but are not used by the processor. Thus, at 16, the content of the FE field is compared with the waitclrF parameter such that when the bits at "1" in the FE register field include the bits at "1" in the waitclrF parameter, this comparison yields a "true" result. The execution unit 10 of the processor considers this "true" result to release the processor from its waiting state, so that the processor resumes the execution of its program. In addition, the "true" result of this comparison resets the bits in the FE register field corresponding to the bits at "1" in the waitclrF parameter, so that a new wave of notifications can be considered.

[0034] Figure 4 is a block diagram of an embodiment of the backward synchronization unit BSYNC of the processor PEi. Such a unit uses the BE and BM fields of the IPE register and the waitclrB and notifyB parameters of the synchronization instruction to manage the backward notification link BN in a manner similar to the unit in Figure 3 The structure is symmetric for handling the incoming notification line Bnin from the processor PEi+1 and the outgoing notification line Bnout to the processor PEi-1.

[0035] The FM and BM mode fields of the IPE register of the processor can be initialized by the operating system at system startup to configure a processor group within which the processors are synchronized with each other. Preferably, to reduce latency, each group consists of consecutive processors in the chain.

[0036] By default, when the BM and FN fields of all processors are set to "0", notifications are not propagated, but direct notifications are still possible. This creates a group of three processors, where each processor PEi can send 16 different notifications to each of its two adjacent processors PEi-1, PEi+1, and receive 16 different notifications from each of its two adjacent processors.

[0037] Setting a bit of the mode field BM or FM to "1" in processor PEi enables the corresponding notification to be propagated between processors PEi-1 and PEi+1 on either side of processor PEi.

[0038] When the same bit is set to "1" in all mode fields, in the link direction, all downstream processors receive the corresponding notification from any processor. If the chain is configured as a ring, as Figure 2 shown, an infinite loop of propagation is prevented by setting the corresponding bit in one of the mode fields to "0".

[0039] By using bit values in the mode fields, a large number of combinations of processor groups can be configured between these two extremes, and these groups can differ between different notification lines.

[0040] Furthermore, regardless of how the grouping is chosen, each processor at the end of a group is at the intersection of two adjacent groups. In fact, even if such a processor does not propagate notifications from one group to another, it can itself send notifications to and receive notifications from each of the two groups.

[0041] The following is an example of the application of such a structure in barrier-synchronization. In such a synchronization, all involved processors are expected to reach the same point or barrier during operation in order to continue their processing.

[0042] In the structure of patent application US2015 - 0339256, a barrier involving at most four processors in a group is materialized in each processor by a register that has one bit for each processor in the group. Once a processor reaches the barrier, it notifies the other processors by setting the bit associated with it in the registers of the other processors and then stops to wait for the other processors to reach the barrier. Once all the bits corresponding to the other processors are set in its register, the stopped processor resumes its activity and starts by resetting the bit in its register.

[0043] Figure 5 shows Figure 2 the configuration of the processors in, for example, performing barrier synchronization for a group of eight consecutive processors 0 to 7 in a chain. For this purpose, only one forward notification line and one backward notification line are needed between each pair of adjacent processors in the group. Although the notification lines may vary between different pairs of processors, for clarity, it is assumed here that all the notification lines used have the same rank k, where k is between 0 and 15. Thus, a forward notification line designated as FNk and a backward notification line designated as BNk are used between pairs of processors.

[0044] Processors 0 to 7 are all configured not to propagate the forward notification FNk, i.e., the bit of rank k in their FE register fields is "0". Thus, the processors in the group can only receive the notification FNk from its immediate previous one, as shown by the arrows turning right inside the processors.

[0045] In addition, processors 1 to 6 are configured to propagate the backward notification BNk, i.e., the bit of rank k in their BE register fields is "1". This state is shown by the arrows pointing horizontally to the left.

[0046] Of course, each processor can still send any notification in both directions, and its FE and BE register fields record the notifications passing through the processor. Thus, processor 7 is shown by an arrow turning left, indicating that it can send the backward notification BNk.

[0047] When processors 1 to 6 reach the barrier, they are programmed to run continuously:

[0048] - Wait for an instruction for the forward notification FNk,

[0049] - Issue an instruction for the forward notification FNk, and

[0050] - Wait for an instruction for the backward notification BNk.

[0051] The first two instructions can be implemented by a single SYNCGROUP instruction whose waitclrF and notifyF parameters each identify only rank k. The third instruction can be a SYNCGROUP instruction with all parameters empty except for the waitclrB parameter which identifies only rank k.

[0052] When processor 0 reaches the barrier, it is programmed to run continuously:

[0053] - an instruction to issue a forward notification FNk, and

[0054] - an instruction to wait for a backward notification BNk.

[0055] These two instructions can be implemented by two consecutive SYNCGROUP instructions with all parameters zero except for the notifyF and waitclrB parameters which each identify only rank k.

[0056] Finally, when processor 7 reaches the barrier, it is programmed to run continuously:

[0057] - an instruction to wait for a forward notification FNk, and

[0058] - an instruction to issue a backward notification BFk.

[0059] Both of these instructions can be implemented by a single SYNCGROUP instruction whose waitclrF and notifyB parameters each identify only rank k.

[0060] Thus, processors 0 through 7 enter a waiting state when they reach the barrier. Processors 1 through 7 start by running a wait instruction and do not issue a notification until the barrier is removed. Only processor 0 starts the notification instruction FNk as soon as it reaches the barrier, then pauses and waits.

[0061] The notification FNk issued by processor 0 is recorded by processor 1, and processor 1 exits its waiting state by issuing the notification FNk to the next processor 2. Processor 1 enters the waiting state again, this time waiting for the backward notification BNk. In fact, although processor 1 has reached the barrier, it does not know whether the other downstream processors have reached the barrier.

[0062] These events propagate from one processor to another until processor 7. Once the notification FNk is received, processor 7 issues a backward notification BNk to processor 6 and resumes running its program. Since the last processor 7 only receives the notification when all the previous processors have issued a notification when they reached the barrier, all processors have reached the barrier.

[0063] Since the processor 6 and the processors 1 to 5 are configured to propagate the backward notification BNk, the notification BNk arrives at all the processors 0 to 6 almost simultaneously. Each of these processors exits its waiting state and resumes the execution of its program.

[0064] This synchronization structure also opens up new possibilities for parallel processing, especially for loops whose iterations are not independent. Iterations are independent in a loop when accesses to shared data (e.g., arrays) are performed on different elements.

[0065] Traditional multi-core architectures allow multiple processors to be assigned to run multiple iterations of a loop in parallel. For example, each iteration of the following type of loop is independent and can be assigned to a different core:

[0066] for(i = 0; i < n; i++){

[0067] a[i] = a[i] + b[i];

[0068] }

[0069] In fact, it is well known that the variable a[i] is defined as an operand and is updated. This type of loop is called a parallel loop. In fact, such a loop is transformed into NB_PE parallel sub-loops in the following way, where NB_PE is the number of processors assigned to run and pid is the number of the processor running the sub-loop:

[0070]

[0071] The "barrier" instruction refers to a function typically available in the operating environment of a multi-core processor, e.g., pthread_barrier_wait.

[0072] There are so-called vector loops, where the value of the operand depends on the order in which the iterations are run, e.g., the following type of loop:

[0073] for(i = 0; i < n; i++){

[0074] a[i] = a[i + 1] + b[i];

[0075] }

[0076] In this case, running two iterations of the loop in parallel is incorrect, e.g.:

[0077] a[1] = a[2] + b[1], and

[0078] a[2] = a[3] + b[2]

[0079] In fact, if the second iteration is completed before the first iteration, the variable a[2] will contain the new value, while the first iteration requires the old value.

[0080] To avoid such pitfalls in a conventional manner, the loop is run by a single core, or the loop is decomposed into two parallel loops via a temporary array temp[], in the following form:

[0081]

[0082] In the structure of the present disclosure, by rewriting the loop as follows, vector loop iterations can be processed in parallel on a processor chain:

[0083]

[0084] Note that the variables ii and t1 have local scopes limited to their respective loop bodies, while the arrays a[] and b[] are global scope variables shared by the processors.

[0085] Thus, for 8 processors, processor 0 runs:

[0086]

[0087] While processor 1 runs in parallel:

[0088]

[0089] And so on, processor 7 runs:

[0090]

[0091] As in Figure 5 In the example, processors 0 to 7 are all configured not to propagate forward notifications. In addition, processors 1 to 6 are configured to propagate backward notifications. The waitclrF and notifyF parameters identify the forward notification lines used, which have rank k in the example considered. Similarly, the waitclrB and notifyB parameters identify the backward notification lines used, which have rank k in the example considered.

[0092] In the first iteration, processor 0 runs:

[0093] 1: t1 = a[1] + b[0];

[0094] 2: syncgroup(notifyF);

[0095] 3: syncgroup(waitclrB);

[0096] 4: a[0] = t1;

[0097] In the first iteration, processor 1 runs in parallel:

[0098] 5: t1 = a[2] + b[1];

[0099] 6: syncgroup(notifyF);

[0100] 7: syncgroup(waitclrF);

[0101] 8: a[1] = t1;

[0102] Thus, in line 1, processor 0 reads the old value of variable a[1]. Importantly, this operation occurs before variable a[1] receives the new value updated by processor 1 in line 8. Thus, the system is configured and programmed such that line 8 always runs after line 1.

[0103] In line 2, processor 0 notifies processor 1 that it has read variable a[1]. (In line 3, processor 0 normally waits for a backward notification to continue: this notification is issued by the syncgroup(notifyB) instruction run by processor 7 before entering the sub-loop.)

[0104] In parallel, in line 7, after saving the new value of a[1] in variable t1, processor 1 enters a waiting state and only runs line 8 after receiving a notification from processor 0.

[0105] Step by step, each processor releases the next processor after reading the old variable value, so that the next processor can update the variable. The last processor in the chain sends a backward notification (syncgroup(notifyB)) during the iteration, which releases the first processor to start a new iteration.

[0106] Finally, when each processor exits its sub-loop, it runs the last notification, which releases the processors still waiting, and then runs barrier-type synchronization as needed in the case of a parallel loop.

[0107] Note that in the loop body, each processor runs two SYNCGROUP instructions. The first performs a notification-type operation, and the second performs a "waitclr"-type operation. Given the chained synchronization scheme for a processor, in principle, the order of these two instructions can be reversed, which allows them to be combined into one instruction. However, in cases where calculations not involving global variables can be inserted between these two instructions, it may be more efficient to keep the two instructions separate. On the other hand, the reversal of the two SYNCGROUP instructions and their combination into a single instruction allows synchronization of more general vector loops, where the gap in the number of iterations between reading a variable and writing it back is not exact. For example, this is the case for the following loop, where j > 0 and the variables are:

[0108] for(i = 0; i < n; i++){

[0109] a[i] = a[i + j] + b[i];

[0110] }

Claims

1. An inter-processor synchronization system, comprising: At least three processors, each processor being configured to release the processor from a waiting state in operation in response to a notification; A plurality of one-way point-to-point notification lines connecting the processors in a chain; And In each processor: i) A synchronization register having bits respectively associated with the notification lines, the synchronization register being connected to record the corresponding states of the upstream notification lines propagated by the upstream processors, and ii) A gate controlled by a configuration register to selectively propagate the state of the upstream notification line on the downstream notification line to the downstream processors.

2. The system according to claim 1, wherein Each processor is configured to selectively activate the downstream notification line according to the parameters of the notification machine instruction run by the processor.

3. The system according to claim 2, wherein Each processor is configured to pause the running of the corresponding program according to the parameters of the wait machine instruction run by the processor, and the pause is triggered when the synchronization register contains a valid bit pattern corresponding to the parameters of the wait machine instruction.

4. The system according to claim 3, wherein, Each processor is configured to reset the synchronization register when the pause is triggered.

5. The system according to claim 4, wherein, The wait machine instruction and the notification machine instruction form part of a single machine instruction executable by the processor.

6. The system according to claim 1, wherein The configuration register includes bits respectively associated with the upstream notification lines, and the gate is configured to selectively propagate the state of the upstream notification line according to the corresponding states of the bits in the configuration register.

7. An inter-processor synchronization method, comprising the following steps: Connecting at least three processors in a chain by point-to-point lines, the point-to-point lines being configured to transmit corresponding notification bits in the same direction; Configuring each processor to release the processor from a waiting state in response to a notification; In the first processor of the chain, sending the notification bits to a second processor after the first processor in the chain; And In the second processor, selectively propagating the notification bits to a third processor after the second processor in the chain depending on the state of local configuration bits.

8. The method according to claim 7, wherein, The second processor performs the following steps: Saving the notification bits in a synchronization register; Running a wait machine instruction with parameters to stop the processor; and Releasing the processor from the stopped state when the synchronization register contains a bit pattern corresponding to the parameters of the wait machine instruction.

9. The method according to claim 8, wherein The second processor performs the following steps: Running a notification machine instruction with parameters; and Configuring the notification bits to be sent to the third processor according to the parameters of the notification machine instruction.

10. The method according to claim 9, wherein, The second processor resets the synchronization register when exiting the stopped state.

Citation Information

Patent Citations

  • Inter-processor synchronization system

    US20150339256A1

  • Method of purging erroneous signals from closed ring data communication networks capable of repeatedly circulating such signals

    US4468734A

  • Protocol for read write transfers via switching logic by transmitting and retransmitting an address

    US5163138A

  • Protocol independent performance monitor with selectable FEC encoding and decoding

    US6775799B1