Transfer control element and networking system

The transfer control element with simplified logic circuits and timing acquisition elements addresses the miniaturization and timing verification challenges of self-synchronous pipelines, enhancing data processing efficiency and throughput.

JP7789375B2Active Publication Date: 2025-12-22UNIV OF TSUKUBA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022575129
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-30
Filing Date
2021-12-07
Publication Date
2025-12-22
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

Conventional transfer control elements for self-synchronous pipelines, such as Muller's C elements, are difficult to miniaturize and do not effectively verify timing, leading to reduced throughput and increased circuit size.

Method used

A transfer control element comprising a first and second logic circuit, a logic gate, and timing information acquisition elements, which utilize look-up tables and simplified wiring to control data transfer between pipeline stages, allowing for miniaturization and improved timing verification.

Benefits of technology

The solution enhances data processing capacity per unit time while reducing circuit size and improving throughput by ensuring efficient timing verification and placement of pipeline stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007789375000001
    Figure 0007789375000001
  • Figure 0007789375000002
    Figure 0007789375000002
  • Figure 0007789375000003
    Figure 0007789375000003
Patent Text Reader

Abstract

To increase the data processing quantity per unit time while reducing the size of circuit structure for transfer control compared with the conventional structure of a combination of logic gates, a first logic circuit (61) of a transfer control element (54) outputs a first output signal (q1) on the basis of an upstream transfer request signal (s1), a reset signal (r1), a first feedback signal (pq1), and a first lookup table, a logic gate (62) outputs an output signal (s2) obtained by applying the logical NOR operation to the upstream transfer request signal (s1), the inverted signal of the first output signal, a second feedback signal (pq2), and a downstream transfer enabling signal (acki), and a second logic circuit outputs a second output signal (q2) on the basis of the output signal (s2) obtained by the logical NOR operation, a downstream transfer enabling signal (r2), the second feedback signal (pq2), and a second lookup table.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a transfer control element that controls data transfer in a self-synchronous pipeline, and to a networking system that allows data to be transmitted and received between nodes that include the transfer control element. [Background technology]

[0002] BACKGROUND ART With regard to a networking system that allows data transmission and reception between a plurality of nodes, the technology described in Patent Document 1 below is known. Patent document 1 (International Publication No. 2018 / 097242) describes a configuration in which, when nodes having data-driven processors in an autonomous distributed communication network (ad hoc network) perform processing in a self-synchronous pipeline at each node, a transfer control element (C element) controls the transfer of data between pipeline stages. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2018 / 097242 (WO-A1-2018 / 097242) Summary of the Invention [Problem to be solved by the invention]

[0004] (Problems with the prior art) The technology described in Patent Document 1 uses a circuit (hereinafter simply referred to as a "C element" or "conventional transfer control element") that is an extension of Muller's C element, which is commonly used in asynchronous circuits. The C element in Patent Document 1 is configured to enable handshakes using only standard logic gates, but does not disclose a method for verifying the timing of transfer control in a self-synchronous pipeline, nor does it describe the placement and wiring of each element. In other words, with conventional C elements that combine standard logic gates, when implemented in an FPGA, it is difficult to miniaturize the circuit configuration of the C element. This makes it difficult to improve the amount of data processed per unit time (throughput) by reducing the time required for transfer control by the C element, and there is a risk of a decrease in the throughput of the entire circuit.

[0005] The present invention has as its technical object to improve the amount of data processed per unit time while miniaturizing the circuit configuration of transfer control in a transfer control element that controls data transfer between self-synchronous pipeline stages, compared to conventional configurations that combine logic gates. [Means for solving the problem]

[0006] In order to solve the above technical problem, the transfer control element of the invention as set forth in claim 1 comprises: A transfer control element used in a circuit consisting of a self-synchronous pipeline in which a plurality of pipeline stages are connected, and which controls data transfer between an upstream pipeline stage and a downstream pipeline stage in data processing, comprising: a first logic circuit, a logic gate, and a second logic circuit; the first logic circuit, based on an upstream transfer request signal input from an upstream-stage transfer control element, a reset signal obtained by feeding back an output signal of the logic gate, a first feedback signal obtained by feeding back a first output signal of the first logic circuit, and a predetermined first look-up table, the upstream transfer request signal or outputting a first output signal so that the first feedback signal is maintained; the logic gate outputs an output signal that is a negative OR of the upstream transfer request signal, an inverted signal of the first output signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a downstream transfer permission signal input from a downstream-stage transfer control element; the second logic circuit outputs a second output signal so that the output signal of the negative OR or the second feedback signal is maintained based on the output signal of the negative OR, the downstream transfer enable signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a predetermined second look-up table; the transfer control element outputs the first output signal as a transfer permission signal to an upstream transfer control element, and outputs the second output signal as a transfer request signal to a downstream transfer control element; It is characterized by:

[0007] The invention as recited in claim 2 is the transfer control element as recited in claim 1, a first timing information acquisition element disposed on a wiring between the first logic circuit and the logic gate, for acquiring a transition timing of a first output signal; a second timing information acquisition element disposed on a wiring between the logic gate and the second logic circuit, for acquiring transition timing of an output signal of the negative OR; a third timing information acquisition element disposed on a wiring on the output side of the second logic circuit, for acquiring a transition timing of a second output signal; The present invention is characterized by the following.

[0008] The invention as set forth in claim 3 is the transfer control element as set forth in claim 1 or 2, a data latch to which the output signal of the logic gate and a function control signal are input; Equipped with The result of a logical operation between the output signal from the data latch and the second output signal is output as a transfer request signal to the downstream stage. It is characterized by:

[0009] In order to solve the above technical problem, the networking system of the invention described in claim 4 comprises: A networking system in which a plurality of nodes are connected by a communication network, Each node has a data-driven processor consisting of a self-synchronous pipeline, The self-synchronous pipeline has a plurality of pipeline stages connected in a plurality of stages, A transfer control element controls the transfer of data between an upstream pipeline stage and a downstream pipeline stage of the data processing; the transfer control element includes a first logic circuit, a logic gate, and a second logic circuit; the first logic circuit, based on an upstream transfer request signal input from an upstream-stage transfer control element, a reset signal obtained by feeding back an output signal of the logic gate, a first feedback signal obtained by feeding back a first output signal of the first logic circuit, and a predetermined first look-up table, the upstream transfer request signal or outputting a first output signal so that the first feedback signal is maintained; the logic gate outputs an output signal that is a negative OR of the upstream transfer request signal, an inverted signal of the first output signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a downstream transfer permission signal input from a downstream-stage transfer control element; the second logic circuit outputs a second output signal so that the output signal of the negative OR or the second feedback signal is maintained based on the output signal of the negative OR, the downstream transfer enable signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a predetermined second look-up table; the transfer control element outputs the first output signal as a transfer permission signal to an upstream transfer control element, and outputs the second output signal as a transfer request signal to a downstream transfer control element; It is characterized by: [Effects of the Invention]

[0010] According to the inventions recited in claims 1 and 4, in a transfer control element that controls data transfer between self-synchronous pipeline stages, it is possible to improve the amount of data processed per unit time while miniaturizing the circuit configuration of the transfer control compared to conventional configurations that combine logic gates. According to the second aspect of the present invention, in the transfer control element of the present invention, timing can be verified by using a timing information acquisition element. According to the invention as recited in claim 3, the maximum throughput of the processor can be improved compared to when the output from the second logic circuit is input to the data latch. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is an explanatory diagram of the entire networking system equipped with a transfer control element according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a functional block diagram showing the functions of the data-driven processor according to the first embodiment. [Figure 3] FIG. 3 is a block diagram of a circuit in which the data-driven processor of the first embodiment is implemented, and is an explanatory diagram of the basic structure of a self-synchronous pipeline. [Figure 4] FIG. 4 is a detailed explanatory diagram of the transfer control element of the first embodiment. [Figure 5] FIG. 5 is an explanatory diagram of the logic circuit of the transfer control element of the first embodiment, where FIG. 5A is a circuit configuration diagram and FIG. 5B is a truth table of the look-up table. [Figure 6] FIG. 6 is an explanatory diagram of the configuration of a data-driven processor according to a first embodiment of the present invention, which is an application example of a self-synchronous pipeline. [Figure 7] FIG. 7 is an explanatory diagram of transfer control in the diverter. [Figure 8] FIG. 8 is an explanatory diagram of a transfer control element for a diversion section. [Figure 9] FIG. 9 is a circuit diagram illustrating a conventional transfer control element. [Figure 10] FIG. 10 is a diagram illustrating the experimental results of the number of resources used in the transfer control circuit. [Figure 11] FIG. 11 is a diagram illustrating the experimental results of throughput, where FIG. 11A is a graph of the experimental results for Cyclone IV, and FIG. 11B is a graph of the experimental results for Zynq-7000. [Figure 12] FIG. 12 is a diagram illustrating the experimental results of throughput in a circular pipeline. FIG. 12A is a graph of the experimental results for Cyclone IV, and FIG. 12B is a graph of the experimental results for Zynq-7000. [Figure 13] FIG. 13 is an explanatory diagram of a transfer control element according to the second embodiment, and corresponds to FIG. 4 of the first embodiment. [Figure 14] FIG. 14 is an explanatory diagram of a transfer control element according to the third embodiment, and corresponds to FIGS. 4 and 8 of the first embodiment. [Figure 15] 15A is an explanatory diagram of the experimental results when the conventional configuration shown in FIG. 9 is used, FIG. 15B is an explanatory diagram of the experimental results when the configuration of Example 1 is used, and FIG. 15C is an explanatory diagram of the experimental results when the configuration of Example 3 is used in the functional transfer control circuit together with the configuration of Example 1. DETAILED DESCRIPTION OF THE INVENTION

[0012] Next, examples that are specific examples of the embodiment of the present invention will be described with reference to the drawings, but the present invention is not limited to the following examples. In the following description using the drawings, illustrations of components other than those necessary for the description are omitted as appropriate to facilitate understanding. [Example]

[0013] FIG. 1 is an explanatory diagram of the entire networking system equipped with a transfer control element according to a first embodiment of the present invention. 1, the networking system S of the first embodiment of the present invention has a plurality of sensor nodes (nodes) N1 as an example of a data-driven processing device. In addition, the networking system S of the first embodiment can be applied to, as examples of the sensor nodes (nodes) N1, mechanical security devices, monitoring devices for structures such as water pipes and gas pipes, monitoring devices for viaducts and places prone to landslides, and monitoring devices for air pollutants and radiation. Furthermore, the networking system S of the first embodiment is also provided with a node (center node) N2 as an example of a data-driven processing device installed at the center of a security company, management company, or monitoring agency. In addition, in the first embodiment, each of the nodes N1 and N2 is assigned unique identification information (node ​​ID).

[0014] FIG. 2 is a functional block diagram showing the functions of the data-driven processor according to the first embodiment. 1 and 2, data driven processor 1 is electrically connected to wireless communication module 2 as an example of a communication unit, sensor 3 as an example of a monitoring means, and display 4 as an example of a display member. The wireless communication module 2 is configured to be able to transmit and receive data to and from another node N1 via wireless communication. The wireless communication method of the first embodiment can employ any conventionally known configuration, such as a wireless LAN, a mobile phone network, or near-field communication (NFC). The first embodiment can employ a sensor network method or an ad hoc method as an example of an autonomous distributed communication network. That is, the first embodiment employs a method in which direct communication is performed between nodes, rather than communication between an access point that manages the network and each node, as in the infrastructure method. The communication network is not limited to the autonomous distributed communication network exemplified in the embodiment, and can also be a wired communication network or a communication network in which communication paths are managed (routed), as in the infrastructure method.

[0015] Furthermore, the sensor 3 of the first embodiment is configured by, for example, an acceleration sensor, and senses and monitors vibrations of the node N1. Therefore, it is possible to detect vibrations when the node N1 is removed from its installation location due to theft or the like, or when it is carried away. Note that, although an acceleration sensor is exemplified as the sensor 3, the present invention is not limited to this. For example, the sensor 3 can be applied to any sensor, such as a security sensor that detects people or animals as an example of a monitored object using electromagnetic waves such as infrared rays or visible light, or a fire alarm sensor that detects smoke or heat from a fire as an example of a monitored object. Furthermore, any component that receives an input signal, such as an image signal from a security camera, can be used. Other applications include sensors that monitor public and private infrastructure such as flammable gas leaks and water leaks, smart meters that remotely detect power consumption in homes and offices, monitoring sensors to keep an eye on people living alone, M2M (Machine-to-Machine) that detects sold-out or malfunctioning vending machines, and any other sensor (sensing network) including those related to IoT (Internet of Things) technology, where devices with communication capabilities are connected via the Internet to perform automatic recognition, automatic control, remote measurement, etc.

[0016] In the first embodiment, data transmitted and received between the nodes N1 and N2 includes identification information (node ​​ID) of the source node N1, identification information (data ID) of the data, and the data itself such as the detection result of the sensor 3. Furthermore, data-driven processor 1 of Example 1 is configured as a so-called multi-core processor, having a configuration including multiple cores 1a, 1b, 1c, ... In Example 1, processing is assigned to each of cores 1a, 1b, 1c, ... For example, first core 1a is configured to process the detection results of sensor 3 of node N1 itself and to distribute received data. Second core 1b is configured to process received data from adjacent first nodes N1 and N2, and third core 1c, fourth core 1d, ... are configured to process received data from adjacent second nodes N1 and N2 and third nodes N1 and N2, ...

[0017] FIG. 3 is a block diagram of a circuit in which the data-driven processor of the first embodiment is implemented, and is an explanatory diagram of the basic structure of a self-synchronous pipeline. 3, data-driven processor 1 of the first embodiment has a pipeline structure using a self-synchronous pipeline. Data-driven processor 1 of the first embodiment has multiple pipeline stages (pipeline stages) 51 corresponding to the pipeline stages on the functional blocks (architecture).

[0018] 3, each pipeline stage 51 has a data latch (DL: Data Latch) 52 that holds a packet transmitted from an upstream pipeline stage 51 along the flow of packets (blocks of data), a processing circuit (FL: Function Logic) 53 that executes the processing of each pipeline stage 51 based on the packet held in the data latch 52, and a C' element 54 as an example of a transfer control element that supplies a signal (clock signal, trigger signal) to the data latch 52 to permit / prohibit data transfer. The C' element of the first embodiment is configured by a self-timed transfer control mechanism (STCM: Self-timed Transfer Control Mechanism).

[0019] Valid data is transferred between adjacent pipeline stages 51 by a type of negotiation called four-way handshake using a transfer request (send) signal and a transfer acknowledge (ack) signal by the C' element 54. This local data transfer ensures that only valid data is transferred. In other words, signal gating at the pipeline stage 51 level is naturally realized, and circuits that consume power due to switching are truly limited to the pipeline stage 51 that is processing data, leading to power savings at the system level. The four-way handshake is achieved by the following steps: (1) When the input signal (transfer request signal sendi) from the C' element 54 (Ci) of the i+1th pipeline stage 51 is asserted (the signal becomes valid), the C' element 54 (Ci+1) of the i+1th pipeline stage 51 asserts the transfer acknowledge signal acki+1 (output signal). (2) When the acki+1 output signal of Ci+1 is asserted, the preceding Ci negates its sendi output signal (invalidates the signal). (3) When the sendi input signal of Ci+1 is negated, Ci+1 negates the acki+1 output signal and simultaneously asserts the cpi+1 signal to the i+1th stage data latch 52. This causes the data output from the previous stage to be latched into the data latch 52 of the i+1th stage pipeline stage 51 to which Ci+1 belongs. At the same time, Ci+1 asserts the sendi+1 output signal. (The cpi+1 signal is negated when the acki+2 input signal is asserted.) Similarly, the C' element 54 in the subsequent stage repeats steps (1) to (3).

[0020] (Detailed configuration of C' element) FIG. 4 is a detailed explanatory diagram of the transfer control element of the first embodiment. FIG. 5 is an explanatory diagram of the logic circuit of the transfer control element of the first embodiment, where FIG. 5A is a circuit configuration diagram and FIG. 5B is a truth table of the look-up table. In FIG. 4, the C′ element 54 (54i+1) of the first embodiment has a first logic circuit 61 (QSR1), a negative OR gate (NOR gate) 62 as an example of a third logic circuit, and a second logic circuit 63 (QSR2). The first logic circuit 61 outputs a first output signal (q1=acko) so that the upstream transfer signal (s1) or the first feedback signal (pq1) is maintained based on an upstream transfer request signal (sendi=send in, s1) input from the upstream C′ element 54 (54i), a reset signal (r2) obtained by feeding back the output signal of the NOR gate 62, a first feedback signal (pq1) obtained by feeding back the first output signal (q1) of the first logic circuit 61, a master reset signal (mr) for initializing the outputs of all transfer control elements to achieve synchronization, and a first look-up table (LUT, see Figure 5B).

[0021] As shown in Figure 5B, when the master reset signal (mr) and the reset signal (r1) are both "0", if either the upstream transfer signal (s1) or the first feedback signal (pq1) is "1", the first output signal (q1) becomes "1", and if both the upstream transfer signal (s1) and the first feedback signal (pq1) are "0", the first output signal (q1) becomes "0". Furthermore, when the master reset signal (mr) or the reset signal (r1) is "1", the first output signal (q1) is "0", with one exception. The exception is that even if the reset signal (r1) is "1", if the upstream transfer signal (s1) and the first feedback signal (pq1) are "1", the first output signal (q1) is "1".

[0022] Therefore, the first logic circuit 61 of the first embodiment realizes a pseudo SR-FF (QSR) having a function similar to that of a NAND gate latch composed of two cross-connected NAND gates in a conventional C element described later. The QSR of the first embodiment realizes the logic of a so-called SR-FF (Set Reset Flip Flop) by using one LUT (LUTqsr) and feedback. That is, the first logic circuit 61 of the first embodiment is a circuit that can specify the value to be held internally and output by using a set (s) that sets the output to 1 and a reset (r) that sets the output to 0, and can also hold and output the value (pq) that was previously specified when both s and r are asserted (i.e., set to 1).

[0023] The NOR gate 62, which is an example of a logic gate, outputs an output signal (s2) of the NOR of the upstream transfer request signal (sendi), the inverted signal (not q1) of the first output signal (q1), a second feedback signal (pq2) obtained by feeding back the second output signal (q2) of the second logic circuit 63, a downstream transfer permission signal (acki = ack in) input from the downstream C' element 54 (54i+2), and a g signal for function expansion. The g signal is used to realize transfer control having functions such as merging, and is not used in FIG. 5 (it is always input with "0"). In the embodiment, a NOR gate (NOR gate) is used as an example of a logic gate, but this is not limiting, and an equivalent circuit using a logical product gate (AND gate), a logical non-product gate (NAND gate), or the like may also be used.

[0024] The second logic circuit 63 outputs a second output signal (q2) so that the negative OR output signal (s2) or the second feedback signal (pq2) is maintained based on the negative OR output signal (s2), the downstream transfer acknowledge signal (r2=acki), a second feedback signal (pq2) obtained by feeding back the second output signal (q2) of the second logic circuit 63, a master reset signal (mr), and a second look-up table. The second lookup table is similar to the first lookup table, and therefore a detailed description thereof will be omitted.

[0025] The C' element 54 (54i+1) outputs the first output signal (q1) to the upstream C' element (54i) as a transfer acknowledge signal (acko = ack out) via the delay element Da. The C' element 54 (54i+1) also outputs the second output signal (q2) to the downstream C' element 54 (54i+2) as a transfer request signal (sendo = send out) via the delay element Ds. The C' element 54 (54i+1) also outputs the second output signal (q2) to the data latch 52 as a transfer control signal cp. The delay elements Da and Ds are elements that delay the timing of data transfer in order to adjust the timing of data transfer. The amount of delay in the timing of data transfer is set according to the design, specifications, etc.

[0026] When observing each output signal (q1, s2, q2), latches 66-68, which are an example of timing information acquisition elements, can be virtually installed to observe transition timing (the timing at which a signal that changes depending on the length of the circuit wiring, etc., transitions (changes) from "0" to "1" and from "1" to "0"). Specifically, by placing a first latch (first timing information acquisition element) 66 on the wiring between the first logic circuit 61 and the NOR gate 62, the first output signal (q1) can be observed. Furthermore, by placing a second latch (second timing information acquisition element) 67 on the wiring between the NOR gate 62 and the second logic circuit 63, the NOR output signal (s2) can be observed. Furthermore, by placing third latches (third timing information acquisition elements) 68a, 68b on the wiring on the output side of the second logic circuit 63, it is possible to observe the second output signal (q2) going to the downstream C' element 54 and the second output signal (q2) going to the data latch 52.

[0027] FIG. 6 is an explanatory diagram of the configuration of a data-driven processor according to a first embodiment of the present invention, which is an application example of a self-synchronous pipeline. In Figure 6, data-driven processor 1 of embodiment 1 has a matching memory unit (MM) 72 that detects data pairs for operations input from I / O interface 71, which inputs and outputs data, a program storage unit (PS) 73 that realizes the fetching of operations, a function processing unit (FP) 74 that executes operations, and a memory access unit (MA) 75. Between I / O interface 71 and matching memory unit 72 and between matching memory unit 72 and program storage unit 73, merge units (M) 76a, 76b that realize pipeline merging are arranged, and between memory access unit 75 and each merge unit 76a, 76b, branch units (B) 77a, 77b that realize pipeline branching are arranged. In the case where the processing of input data is a unary operation that does not require detection of data pairs, data-driven processor 1 of FIG. 6 can reduce power consumption and processing time by providing a path (Bypass in FIG. 8) that can be executed by bypassing MM72. Note that the self-synchronous pipeline structure is also described in, for example, International Publication No. 2013 / 011653, and various configurations can be adopted.

[0028] FIG. 7 is an explanatory diagram of transfer control in the diverter. FIG. 8 is an explanatory diagram of a transfer control element for a diversion section. 6, the C' elements of MM72, PS73, FP74, and MA75 can use the configurations shown in FIGS. 4 and 5 as they are, but the branch sections (diverging sections) 77a and 77b and merge sections (merging sections) 76a and 76b require modifications. In the branch sections 77a and 77b shown in FIGS. 7 and 8, a signal br indicating divergence and an output signal cp from the second logic circuit 63 are input to a D latch 81. The D latch 81 is a latch (element) that outputs data in response to a clock signal; data is output while the clock signal is HIGH and the output does not change while the clock signal is LOW. 4 and the output from the D latch 81, a logical AND is taken and a transfer request signal (sendo-a, sendo-b) is sent to send data exclusively to one of the downstream pipeline stages 51, and a logical AND is taken with the transfer acknowledge signals (acki-a, acki-b) from the C' element and input to the C' element 54. The merge units 76a, 76b, unlike those in Figures 7 and 8, simply combine the transfer request signals (sendi-a, sendi-b) and transfer acknowledge signals (acko-a, acko-b) for the upstream stage and input / output them exclusively through a logic circuit, and the signals for the downstream stage are as in Figure 4, so a detailed description will be omitted.

[0029] (Explanation of the Function of Example 1) The C' element 54 of Example 1 having the above configuration consists of a first logic circuit 61 (QSR1), a negative OR gate (NOR gate) 62, and a second logic circuit 63 (QSR2), and is different from a configuration that combines multiple logic gates as in the conventional configuration. The configuration of a conventional C element will be described below.

[0030] (Explanation of a conventional transfer control element) FIG. 9 is a circuit diagram illustrating a conventional transfer control element. As shown in FIG. 9, the conventional transfer control element 01 extends the circuit used for data transfer control in an asynchronous circuit so that the above-mentioned handshake can be realized using only standard logic gates. Specifically, the conventional transfer control element 01 has two NAND gates (negative AND gates) 02 and 03 cross-connected together, a third NAND gate 04, and two NAND gates 06 and 07 cross-connected together.

[0031] The two upstream NAND gates 02 and 03 output signals based on the transfer request signal sendi (send in) from the upstream stage, the signal of the other NAND gate 02 or 03 cross-connected with it, the output signal of the third NAND gate 04, and a master reset signal mr that initializes all transfer control elements 01. The output signals are inverted and output to the upstream stage via a delay element 08a as a transfer acknowledge signal acko (ack out). The third NAND gate 04 outputs a negative logical product based on the transfer request signal sendi from the upstream stage, the outputs from the upstream NAND gates 02 and 03, the transfer acknowledge signal acki (ack in) from the downstream stage, and the outputs from the downstream NAND gates 06 and 07. The two downstream NAND gates 06 and 07 output signals based on the signal from the third NAND gate 04, the signal from the cross-connected NAND gates 06 and 07, the transfer acknowledge signal acki (ack in) from the downstream stage, and the master reset signal mr. The output signals are output as a signal cp to the data latch 52 or as a transfer request signal sendo (send out) to the downstream stage via a delay element 08b.

[0032] As described above, the conventional transfer control element 01 has multiple logic gates 02-07 and requires cross-connected wiring, which requires space for installing the logic gates 02-07 and space for routing the cross-connected wiring. This tends to increase the overall circuit size, and the longer wiring reduces throughput. In contrast, the C' element 54 of Example 1 has fewer components and simpler wiring than the conventional C element. This allows for a smaller overall circuit configuration and improves throughput.

[0033] Several studies have been reported on implementing asynchronous circuits on FPGAs (Field Programmable Gate Arrays). Non-Patent Document 1 (M. Tranchero and LM Reyneri, "Implementation of Self-Timed Circuits onto FPGAs Using Commercial Tools," Proc. DSD, pp. 373-380, Jan. 2008) describes how asynchronous circuits can be implemented by adding a type of directive to a circuit description that can restrict logic synthesis and optimization, thereby enabling the implementation of Muller's C-elements on an FPGA. However, Non-Patent Document 1 does not disclose placement and routing and timing verification techniques, nor does it disclose any circuit configuration methods other than Muller's C-elements. Furthermore, in Non-Patent Document 2 (J. Furushima, M. Nakajima, and H. Saito, "Design of an Asynchronous Processor with Bundled-data Implementation on a Commercial Field Programmable Gate Array," Informatica, Vol. 40, pp. 399-408, Nov. 2016.), they propose a design support tool for Intel FPGAs that automates design constraint generation, timing verification, and delay adjustment for asynchronous circuits, and propose a method for building a bundled-data processor on an FPGA that achieves a handshake similar to that of a self-synchronous pipeline. However, this method assumes a control circuit different from C elements for data transfer, so the proposed method cannot be applied to a self-synchronous pipeline.

[0034] In contrast to these, the C' element 54 of the first embodiment uses a look-up table, and is simple in configuration and also supports a self-synchronous pipeline. Generally, in a QSR, if a change in the input causes a change in the output, and the input signal changes again before that change is fed back to the input, oscillation can occur. However, in the C' element 54 of the first embodiment, such an input change does not occur during a signal transition during a handshake, so oscillation does not occur in the circuit configuration of the C' element of the first embodiment.

[0035] Currently, design support tools (CAD tools) are distributed for commercially available FPGAs, and the general procedure for implementing a typical synchronous circuit on an FPGA is to perform logic synthesis and placement and routing of the described (designed) circuit, and then verify (simulate) the timing. However, when attempting to implement a self-synchronous pipeline using asynchronous C elements, the following difficulties arise. (i) When the design support tool optimizes and automatically corrects the wiring, some circuits may be degenerated, resulting in a circuit that cannot realize four-way handshake. Therefore, a circuit configuration suitable for FPGA implementation is required. (ii) CAD tools cannot detect the critical path (the path along which the signal propagation is slowest) of the self-synchronous transfer control circuit. Therefore, it is necessary to clarify the critical path and its timing constraints. (iii) Placement and routing that ignores the high modularity of each pipeline stage in a self-synchronous pipeline can significantly increase processing delays. To address this issue, it is necessary to specify the placement area of ​​the self-synchronous transfer control circuit (C element). (iv) Inefficient allocation of global wiring can increase processing delays. To address this issue, it is necessary to control the range of global wiring usage. In addition, conventional architectures use the cp signal to latch function control signals for functions such as duplication, deletion, and branching. This architecture needs to be changed because it increases the circuit size and the time required for data transfer. (v) Timing verification by CAD tools is not supported. Instead, timing constraints must be verified independently.

[0036] (Regarding (i)) In the conventional transfer control element 01, the cp signal is implemented as positive logic and the send and ack signals as negative logic using NAND gate latches 02 and 03. With such NAND gate latches 02 and 03, the circuit sometimes degenerates during logic synthesis when using CAD tools, preventing the intended operation from being achieved, although the cause is unknown. In contrast, in the first embodiment, the circuit configuration is simplified, making it less susceptible to degeneration, and the send signal and ack signal are implemented using positive logic. Therefore, compared to when the conventional transfer control element 01 is used, it is possible to simultaneously avoid circuit degeneration and achieve a smaller scale that leads to higher throughput.

[0037] (Regarding (ii)) In pipeline stage 51, as in synchronous circuits, data arriving at data latch 52 must be latched after the setup time of data latch 52 has elapsed. The signal propagation time on the critical path between DL52 and FL53 in the i-th pipeline stage 51i and the time required to ensure the setup time of DL52 in the (i+1)th stage are called Tfi. After DL52 latches the data, the data must not be changed until the hold time has elapsed. The time required to ensure the hold time of DL52 in pipeline stage i+1 is called Tri. The sum of Tfi and Tri (Tfi + Tri) determines the performance characteristics of the pipeline and also represents the minimum time required for data transfer between pipeline stages 51, i.e., the pipeline tact. The signal propagation times on the critical path that determine Tfi and Tri are clarified using points P1, P2, and P3 shown in Figure 4 as starting points.

[0038] The handshake with the subsequent stage starts with the rising edge of the output (point P2) of the five-input NOR NOR gate 62. In other words, the rising edge of point P2 indicates that the handshake with the subsequent stage has started. The signal transitions that realize the handshake in the adjacent C' element and the time required for them are as follows. Note that it is assumed that after sendo and acko are negated by asserting the master reset signal mr, mr is also negated. When point P2 of the C' element 54 in the (i-1)th stage rises (changes from "0" to "1"), point P3 rises. The time required for this is defined as T1i. At this time, cp also rises. When point P3 of the C' element 54 in the (i-1)th stage rises, sendo rises, sendi of the C' element 54 in the i-th stage rises, and point P1 rises. The time required for this is T2i. When point P1 of the i-th C' element 54 rises, acko also rises, acki of the (i-1)th C' element 54 rises, and point P3 of the (i-1)th C' element 54 falls (changing from "1" to "0"). The time required for this is T3i. When P3 of the C' element 54 in the (i-1)th stage falls, sendo falls, sendi of the C' element 54 in the i-th stage falls, and point P2 rises. The time required for this is T4i. At this time, cp also falls. When point P2 of the C' element 54 in the (i-1)th stage rises, point P1 falls. The time required for this is T5i. At this time, a handshake with the subsequent stage begins. When point P1 of C' element 54 in the ith stage falls, acko also falls, acki of C' element 54 in the (i-1)th stage falls, and point P2 also falls. The time required for this is T6i.

[0039] From the above, the path consisting of T1i to T6i is the critical path of the handshake. Here, Tfi is the time T1i+T2i+T3i+T4i from when the handshake is started at the C' element 54 in the (i-1)th stage until when the handshake is started at the C' element 54 in the i-th stage, and Tri is the time T5i+T6i from when the handshake is started at the C' element 54 in the i-1th stage until the handshake can be started at the C' element 54 in the (i-1)th stage.

[0040] The constraint for guaranteeing the signal propagation time on the critical path of DL and FL in pipeline stage i and the setup time of DL 52 in pipeline stage i+1 by Tfi is given by the following equation (1). Tfi+Tcpi+1>TDLFLiMax+TSi+1…Formula (1) Here, Tcpi+1 is the time from when the handshake starts in the (i+1)th pipeline stage until cpi+1 reaches DL52, TDLFLiMax is the maximum time from when the handshake starts in the i-th pipeline stage 51 until data reaches DL52 of the next pipeline stage, and TSi+1 is the setup time of DL52 of the (i+1)th pipeline stage 51. Furthermore, the constraint for guaranteeing the hold time of DL52 in pipeline stage i+1 by Tri is given by the following equation (2). Tri+TDLFLiMin>Tcpi+1+THi+1 …Equation (2) Here, THi+1 is the hold time of DL 52 of the (i+1)th pipeline stage 51. TDLFLiMin is the minimum time from when handshake starts in the i-th pipeline stage 51 until data arrives at DL 52 of the next pipeline stage.

[0041] In addition, the C' element 54 for erasure, duplication, branching, and merging has a timing constraint shown in the following equation (3) between the function control signal that controls erasure, duplication, branching, and merging and the handshake signal that conveys send and ack. Ttc≧Tfc…Formula (3) Here, Ttc is the time required for the handshake signal to propagate, and is the constraint in this equation. Tfc is the time required for the function control signal to propagate. For example, in Cb (Figure 8), the control signal br, which indicates the shunt direction, is latched by the cp signal and must reach the subsequent two-input NAND gate A before the handshake signal reaches the other input of the two-input NAND gate A. Tfc is the time from the start of the handshake until the br signal propagates to the two-input NAND gate A, and Ttc is the time from the start of the handshake until the handshake signal propagates to the two-input NAND gate A. Similar constraints exist for the other two-input NAND gate B, which must be satisfied to guarantee data transfer. Similarly, there are constraints on the gates to which the function control signal and handshake signal are input for the erase, copy, and merge C' elements.

[0042] (Regarding (iii)) FPGA CAD tools typically flatten the hierarchy of the circuit description during logic synthesis before optimizing the logic. This does not take into account which pipeline stage each placement target belongs to. Furthermore, CAD tools cannot detect that the critical path of the self-synchronous transfer control circuit is the path that constitutes T1i to T6i, which can result in the circuit elements of T1i to T6i being placed in physically distant locations. To avoid this, LogicLock (Intel FPGA software) or Pblock (Xilinx FPGA software), which specify the module placement range, is used to specify placement areas according to the size of each circuit, so that the self-synchronous transfer control circuit of each pipeline stage is placed as close as possible to other circuits without interfering with them. This prevents unnecessary expansion of T1i to T6i. This previously required users to manually fine-tune the placement areas through trial and error, which was time-consuming. In contrast to this, in the first embodiment, the placement and routing is simplified, the effort required to specify the placement area is reduced, and the processing delay is also reduced.

[0043] (Regarding (iv)) Because the gate signals controlling the opening and closing of data latches are generally assumed to be clock signals, global routing is generally used to conserve wiring resources and reduce latency. Global routing can be explicitly specified using the global primitive (in Intel FPGA software) or the BUFG primitive (in Xilinx FPGA software). However, routing gate signals to data latches, such as D latches, in data transfer control circuits using global routing does not necessarily result in reduced latency. For example, a data latch for a function control signal is physically close to the logic gate that generates that gate signal, and it is expected that the signal will arrive faster if it does not go through global routing. In such cases, local routing is used. Furthermore, adjusting the timing of the erase, copy, and shunt function control signals to satisfy Equation (3) requires increasing the corresponding delay in Ds, which increases Tfi. To address this dilemma, we introduce a data latch (D latch) 81 to latch the function control signal, and use the P2 signal instead of the cp signal as its gate signal. To reduce wiring delay, this gate signal uses local wiring instead of global wiring. With this configuration, the P2 signal rises before the cp signal, allowing the function control signal to be latched earlier, reducing Tfc in equation (3). Therefore, the increase in Tfi can be suppressed.

[0044] (Regarding (v)) Timing verification using commercially available CAD tools does not verify the constraints shown in Equations (1), (2), and (3), so these must be verified independently. To achieve this, it is necessary to extract timing information for the paths that are subject to the constraints. CAD tools can extract timing information between adjacent registers. That is, even in a self-synchronous pipeline, timing information for the data path consisting of DL52, which is considered a register, and FL53 can be extracted in the same way as for a synchronous circuit. On the other hand, there are no registers in the path through which the send and ack signals that realize the handshake propagate. In contrast, although a pseudo-SR-FF like that in Example 1 is included, it is not recognized as a register by the CAD tool, so timing information related to the handshake cannot be extracted as is.

[0045] Therefore, in the first embodiment, for timing verification using a CAD tool, latches 66 to 68 that can be logically ignored and recognized as registers are inserted into the circuit description, and paths that constitute the critical path are sandwiched between the latches 66 to 68, thereby making it possible to extract timing information. Note that the elements are not limited to latches, and any elements recognized as registers can be used, but logically ignorable latches are preferable. The latch primitive (Intel FPGA software) or the LDCE primitive (Xilinx FPGA software) is used to implement the latches 66 to 68. The latches 66 to 68 in Figure 4 are the result of applying this method. By inputting an open signal from the outside as the clock signal for the latches 66 to 68 and constantly asserting it, it is possible to avoid degeneration during logic synthesis. For simplicity, this open signal has been omitted from the figure.

[0046] As a timing verification procedure for a self-synchronous pipeline, first, the timing information required for timing verification is extracted using a CAD tool according to the following procedure. (1) The entire pipeline circuit is logically synthesized. (2) To identify critical paths during timing verification, the hierarchy of all pipeline stages and self-synchronous transfer control circuits is preserved. This preserves the hierarchy information and wiring names even after placement and routing. To preserve the hierarchy, use the Design partition (Intel FPGA) or flatten hierarchy option. (3) The entire circuit is logically synthesized and then placed and routed. (4) Use the report timing command to obtain timing information for all paths to be verified. This operation is not performed automatically by the CAD tool, so provide it as a timing information extraction script and have it run automatically.

[0047] Next, it is verified that each pipeline stage satisfies the timing constraints. To perform this verification automatically, a unique tool was created that performs verification, i.e., compares the magnitude relationships based on the acquired timing information and the timing constraints shown in equations (1) to (3), and outputs the results, thereby achieving semi-automation of timing verification. If the verification results show that a pipeline stage 51 does not satisfy the timing constraints, the delay amount of the delay modules (Da, Ds) of that self-synchronous transfer control circuit (C' element 54) is adjusted, and steps (3) and (4) above are performed again. This process is repeated until all constraints are satisfied. Therefore, even in the configuration of the first embodiment, it is possible to verify the timing using a commercially available CAD tool, and it is also possible to adjust the amount of delay.

[0048] (Experimental example) Next, an experiment was conducted to confirm the effects of the present invention. The circuit configuration method of the embodiment is evaluated using commercial FPGAs, namely Xilinx Zynq-7000 FPGA and Intel Cyclone IV FPGA. The CAD tools used are Quartus Prime Standard 18.1 and Vivado 2019.2. To quantitatively evaluate the reduction in circuit size and the resulting improvement in throughput achieved by the circuit of Example 1, a 10-stage linear self-synchronous pipeline was implemented using a conventional transfer control element 01 and the C' element 54 of Example 1, based on the design procedure of Example 1. The pipeline tact, which determines the throughput, varies depending on the latency associated with DL and FL on the data path. To evaluate the throughput when the pipeline can be ideally divided so that these latencies fit within the time required for handshake, that is, when the shortest pipeline tact is achieved, FL, Da, and Ds were removed.

[0049] FIG. 10 is a diagram illustrating the experimental results of the number of resources used in the transfer control circuit. We calculated the resource usage of the conventional transfer control element 01 and the C' element 54 in the pipeline. The unit of each resource is the Logic Cell in Cyclone IV, and the LUT and FF in Zynq-7000. The experimental results are shown in Figure 10. 10, the conventional transfer control element 01 could not be implemented in the Zynq-7000 FPGA. Specifically, placement and routing using a CAD tool could not be performed. In contrast, the C′ element 54 of the first embodiment could be implemented. In Cyclone IV, the number of logic cells in the transfer control circuit was reduced by 50% by using the C' element 54 of the first embodiment.

[0050] FIG. 11 is a diagram illustrating the experimental results of throughput, where FIG. 11A is a graph of the experimental results for Cyclone IV, and FIG. 11B is a graph of the experimental results for Zynq-7000. Next, to evaluate the throughput, the pipeline tact time (Tfi+Tri) was calculated by static timing analysis for each pipeline stage 51. The results are shown in FIG. The pipeline tact can increase or decrease depending on placement and routing using a CAD tool. In FIG. 11, the maximum pipeline tact of the C′ element 54 of Example 1 in each FPGA is normalized to 1. The maximum pipeline throughput is the reciprocal of the longest pipeline tact. As shown in FIG. 11A, in Cyclone IV, the longest pipeline tact of the C′ element 54 of Example 1 was reduced by approximately 68% compared to the conventional transfer control element 01. Therefore, it was confirmed that the throughput (= 1 / longest pipeline tact) could ideally be improved by approximately 3.2 times at most.

[0051] Next, the throughput in the processor configuration was evaluated. In order to evaluate the throughput when C' element 54 of Example 1 is applied to data-driven processor 1, a circular pipeline was implemented in which FL was removed from the processor configuration of Figure 6, and the pipeline tact (Tfi + Tri) was found in each pipeline stage 51 by static timing analysis. As explained in (ii) above, the pipeline stage 51, which has the functions of data deletion, duplication, and pipeline branching and merging, has the timing constraints shown in equation (3). While the constraints can be satisfied by adding delays to the Da and Ds corresponding to each constraint, in this experimental example, we evaluated the pipeline tact under ideal timing constraint satisfaction. Specifically, we calculated the pipeline tact by setting the propagation time of the handshake signal, shown on the left side of equation (3), to the propagation time of the function control signal, shown on the right side of equation (3), which is its lower limit. Regarding the latency associated with each DL and FL, we assumed that the pipeline could be ideally divided so that these latencies fit within the time required for the handshake, and removed Da and Ds. The MM, PS, FP, and MA of the data-driven processor 1 shown in Figure 6 each consist of two or more pipeline stages. The results are shown in Figure 12.

[0052] FIG. 12 is a diagram illustrating the experimental results of throughput in a circular pipeline. FIG. 12A is a graph of the experimental results for Cyclone IV, and FIG. 12B is a graph of the experimental results for Zynq-7000. In Fig. 12, for example, MM0 represents the first stage of MM. Also, Cex represents the erase C element and C' element used in the pipeline stage that erases data, Cce represents the erase / copy C element and C' element used in the pipeline stage that erases and copies data, and Cm represents the merge C element and C' element used in the pipeline stage that realizes the merge of pipelines. 12, in Cyclone IV, the longest pipeline tact time when the C' element 54 of the first embodiment is used is reduced by approximately 37% compared to when the conventional transfer control element 01 is used. Therefore, it was confirmed that the throughput can ideally be improved by a maximum of approximately 1.6 times. [Example]

[0053] FIG. 13 is an explanatory diagram of a transfer control element according to the second embodiment, and corresponds to FIG. 4 of the first embodiment. 13, first logic circuit 61′ of C′ element 54′ in Example 2 differs from first logic circuit 61 in Example 1 in which a reset signal (r) and a master reset signal (mr) are input, in that first logic circuit 61′ is configured as a latch to which a signal resulting from ORing the reset signal and the master reset signal in OR gate 101 is input. Note that in first logic circuit 61′ in Example 2, one of the three inputs always receives a signal “1” (“1′b1” = “1” in binary). In the second embodiment, the first logic circuit 61' is configured to use a latch instead of an LUT, which allows it to be used not only to acquire timing information but also to hold the state, thereby reducing the circuit scale on the critical path and improving throughput.

[0054] Furthermore, a latch 102 as an example of a state-holding element is disposed downstream of the NOR gate 62. Therefore, in the second embodiment, the signal is held by the latch 102, and therefore a signal line for a feedback signal (pq signal) such as that in the first logic circuit 61 of the first embodiment is not required. Furthermore, instead of the second logic circuit 63 of the first embodiment, the second logic circuit 63' of the second embodiment has a configuration in which a second logic circuit 63a' for a transfer control signal (cp) and a second logic circuit 63b' for a transfer request signal (sendo) are connected in parallel. As with the first logic circuit 61', the second logic circuit 63' also receives a signal resulting from ORing the transfer acknowledge signal (acki) and the master reset signal (mr) in the OR gate 103. Therefore, the circuit configuration shown in FIG. 13 can also achieve the same operations and functions as those in the first embodiment, and can also achieve further miniaturization of the circuit and improvement of throughput. [Example]

[0055] FIG. 14 is an explanatory diagram of a transfer control element according to the third embodiment, and corresponds to FIGS. 4 and 8 of the first embodiment. In FIG. 14, when the transfer control element (C′ element) 54″ of the third embodiment is used in the branch sections (diverging sections) 77a, 77b (when it has a diverting function), in the first embodiment, the output signal from the second logic circuit 63 is input to a diverting D latch 81, which is an example of a data latch. However, in the C′ element 54″ of the third embodiment, the output signal of the NOR gate 62 is input to a diverting D latch 201, which is an example of a data latch. This is different from the first embodiment in that Furthermore, a latch 202 can be provided on the output side of the shunt D latch 201 for timing verification using a CAD tool.

[0056] (Function of Example 3) In the C' element 54" of the third embodiment having the above configuration, when the cp signal changes (rises), the output signal of the NOR gate 62 before the second logic circuit 63 is input to the shunt D latch 201. As a result, after a new handshake is started (P2 rises), P2 (the output signal of the NOR gate 62) rises before cp (the output signal of the second logic circuit 63) rises. In other words, after P2 rises, the function control signal is latched in the branch data latch before cp (the output signal of the second logic circuit 63) rises.

[0057] As mentioned above, to construct a functional nonlinear pipeline such as the circular pipeline of a self-synchronous data-driven processor, the self-synchronous transfer control circuit has the functions of pipeline merging and branching, as well as data copying and erasure. The self-synchronous transfer control circuit for branching, copying, and erasure has the time constraint formulated in the above equation (3). If equation (3) is not satisfied during circuit design, the number of corresponding delay modules must be increased to adjust the signal propagation delay. However, this adjustment results in an undesirably large increase in Tfi, because increasing Ds to satisfy the constraint in equation (3) leads to a large increase in Tfi, since Ds delays signal propagation twice during one handshake cycle. To overcome this dilemma, in the third embodiment, a shunt D latch 201 is introduced, which latches the function control signal by using the output signal of the NOR gate 62 (P2 of C') as the gate signal of the data latch instead of the cp signal. The gate signal of the shunt D latch 201 is connected using local wiring instead of global wiring to reduce wiring delay.

[0058] 14, in the third embodiment, as described above, after a new handshake is started, P2 (the output signal of NOR gate 62) rises before cp (the output signal of second logic circuit 63) rises. Therefore, in C′ element 54″ of the third embodiment, signal br instructing diversion, which is an example of a function control signal, can be latched in advance by diversion D latch 201, thereby reducing Tfc in equation (3). In FIG. 14, a transfer control circuit (C′ element 54″) having a shunting function has been described. Similarly, a self-synchronous transfer control circuit having copy and erase functions can be provided with a D latch to which a function control signal such as a copy signal or an erase signal and P2 are input, instead of the shunting D latch 201.

[0059] 15A is an explanatory diagram of the experimental results when the conventional configuration shown in FIG. 9 is used, FIG. 15B is an explanatory diagram of the experimental results when the configuration of Example 1 is used, and FIG. 15C is an explanatory diagram of the experimental results when the configuration of Example 3 is used in the functional transfer control circuit together with the configuration of Example 1. In the experiment, a simulation was carried out to confirm the effect of the present invention. Based on the simulation results, Tfi, Tdpfi, and Tri were calculated for each of the three cases. It should be noted that "Tdpfi" is the lower limit of Tfi obtained from equation (1), that is, "TDLFLiMax+TSi+1-Tcpi+1". The horizontal axis in Figure 15 indicates each pipeline stage using the name of the data-driven processor module to which that pipeline stage corresponds. Furthermore, if the corresponding module consists of multiple pipeline stages, it is numbered to indicate the position, counting from 0. For example, the first pipeline stage of MM is designated MM0.

[0060] Ideally, the circuit should be designed so that Tfi = Tdpfi, i.e., Tdpfi, which indicates the minimum time required for data processing and transfer, is the same as Tfi, which indicates the time actually required for transfer control. If Tfi > Tdpfi, the actual transfer control time includes wasted time, and the larger Tfi - Tdpfi is, the worse the throughput will be. In addition, since the maximum throughput is the reciprocal of (Tfi+Tri), which indicates the pipeline tact time, it is necessary to make (Tfi+Tri) as small as possible in order to improve the maximum throughput.

[0061] Figure 15A shows the results when the conventional transfer control circuit is used, and shows that Tfi is significantly greater than Tdpfi in most pipeline stages. 15B shows the results when the transfer control circuit of the first embodiment is used. According to FIG. 15B, the configuration of the first embodiment is improved compared to the conventional configuration of FIG. 15A. However, in some pipeline stages (e.g., BB-MB) where the transfer control circuit with the function is used, Tfi is significantly greater than Tdpfi. This is due to the effect of increasing Ds to satisfy the constraint of equation (3).

[0062] Fig. 15C shows the results when using the functional transfer control circuit of Example 3. According to Fig. 15C, Tfi is smaller than in the cases of Fig. 15A and Fig. 15B, and high throughput is achieved. Therefore, in the configuration of Example 3, as shown in FIG. 15C, it is possible to reduce the adverse effects caused by increasing Ds in the configuration of Example 1, and to eliminate the problem that Tfi greatly exceeds Tdpfi. As a result, the maximum pipeline tact time of the configuration of Example 3 was reduced to about 60% of that of the configuration of Example 1. In other words, the maximum throughput increased by about 66% (=1 / 0.60-1).

[0063] This result shows that the configuration illustrated in Example 3 makes it possible to bring Tfi close to its lower limit, Tdpfi. The effectiveness of the proposed improvement was demonstrated by implementing it on a state-of-the-art self-synchronous data-driven processor equipped with a circular pipeline and capable of all functions of merging, branching, copying, and erasing. As a result, it was confirmed that the maximum throughput of the processor can be increased by approximately 66% by using the circuit configuration of Example 3.

[0064] (Example of change) Although the embodiments of the present invention have been described in detail above, the present invention is not limited to the above embodiments, and various modifications can be made within the scope of the gist of the present invention as set forth in the claims. For example, in the above embodiment, a configuration including a display 4 is exemplified, but the present invention is not limited to this and a configuration without a display is also possible. Furthermore, the node N1 may be configured to include any other components, such as an actuator, a lamp, or a buzzer, in addition to the display 4 and the wireless communication module 2. In the above embodiment, each sensor node N1 has a configuration including a sensor 3 as an example of a monitoring means, but this is not limiting. For example, it is also possible to have a node that does not have a sensor 3 in the network, but receives data from other nodes and transmits (transfers, rebroadcasts) the data to yet another node, i.e., a so-called relay node. Also, in the embodiment, a configuration including both receiving means and transmitting means has been exemplified, but this is not limiting. For example, a node such as a monitoring center can be configured as a node that does not have broadcasting means (transmitting means) and only receives data transmitted from other sensor nodes N1. [Explanation of symbols]

[0065] 1...Data-driven processor, 51...pipeline stage, 54...transfer control element, 61...first logic circuit, 62...Logic gates, 63...second logic circuit, 66...first timing information acquisition element, 67...second timing information acquisition element, 68...Third timing information acquisition element, 201...D latch, acki: downstream transfer acknowledge signal, acko...transfer acknowledge signal, N1, N2...nodes, pq1...first feedback signal, pq2...second feedback signal, q1...first output signal, q2...second output signal, r1...reset signal, S...Networking system, s1, sendi...upstream transfer request signal, s2: Negative OR output signal, sendo...transfer request signal.

Claims

1. A transfer control element used in a circuit consisting of a self-synchronous pipeline in which a plurality of pipeline stages are connected, and which controls data transfer between an upstream pipeline stage and a downstream pipeline stage in data processing, comprising: a first logic circuit, a logic gate, and a second logic circuit; the first logic circuit outputs a first output signal so as to hold the upstream transfer request signal or the first feedback signal, based on an upstream transfer request signal input from a transfer control element in an upstream stage, a reset signal obtained by feeding back an output signal of the logic gate, a first feedback signal obtained by feeding back a first output signal of the first logic circuit, and a predetermined first look-up table; the logic gate outputs an output signal that is a negative OR of the upstream transfer request signal, an inverted signal of the first output signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a downstream transfer permission signal input from a transfer control element at a downstream stage; the second logic circuit outputs a second output signal so that the output signal of the negative OR or the second feedback signal is maintained based on the output signal of the negative OR, the downstream transfer permission signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a predetermined second look-up table; the transfer control element outputs the first output signal as a transfer permission signal to an upstream transfer control element, and outputs the second output signal as a transfer request signal to a downstream transfer control element; A transfer control element characterized by:

2. a first timing information acquisition element disposed on a wiring between the first logic circuit and the logic gate, for acquiring transition timing of a first output signal; a second timing information acquisition element disposed on a wiring between the logic gate and the second logic circuit, for acquiring transition timing of an output signal of the negative OR; a third timing information acquisition element disposed on a wiring on the output side of the second logic circuit, for acquiring transition timing of a second output signal; 2. The transfer control element according to claim 1, comprising:

3. a data latch to which the output signal of the logic gate and a function control signal are input; Equipped with The result of a logical operation between the output signal from the data latch and the second output signal is output as a transfer request signal to the downstream stage.

3. The transfer control element according to claim 1, wherein the first and second transfer control elements are connected to each other.

4. A networking system in which a plurality of nodes are connected by a communication network, Each node has a data-driven processor consisting of a self-synchronous pipeline, The self-synchronous pipeline has a plurality of pipeline stages connected in a plurality of stages, A transfer control element controls the transfer of data between an upstream pipeline stage and a downstream pipeline stage of the data processing; the transfer control element includes a first logic circuit, a logic gate, and a second logic circuit; the first logic circuit outputs a first output signal so as to hold the upstream transfer request signal or the first feedback signal, based on an upstream transfer request signal input from a transfer control element in an upstream stage, a reset signal obtained by feeding back an output signal of the logic gate, a first feedback signal obtained by feeding back a first output signal of the first logic circuit, and a predetermined first look-up table; the logic gate outputs an output signal that is a negative OR of the upstream transfer request signal, an inverted signal of the first output signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a downstream transfer permission signal input from a transfer control element at a downstream stage; the second logic circuit outputs a second output signal so that the output signal of the negative OR or the second feedback signal is maintained based on the output signal of the negative OR, the downstream transfer permission signal, a second feedback signal obtained by feeding back the second output signal of the second logic circuit, and a predetermined second look-up table; the transfer control element outputs the first output signal as a transfer permission signal to an upstream transfer control element, and outputs the second output signal as a transfer request signal to a downstream transfer control element; A networking system comprising:

Citation Information

Patent Citations

  • Operation processor

    JP1993233853A

  • Networking system

    WO2018097242A1