Interprocessor interconnection architecture for hardware simulation acceleration system

By using a hierarchical interconnect network architecture, large-size cross switches are decomposed into multiple small-size multiplexer arrays. Combined with the interconnection between processors with different topologies, the problems of resource consumption and long latency in hardware simulation acceleration systems are solved, and efficient data communication and simulation acceleration are achieved.

CN121833598APending Publication Date: 2026-04-10XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing processor interconnect architectures suffer from problems such as huge resource consumption, long logic latency, poor wiring friendliness, and severe communication resource contention and conflict when facing hardware simulation acceleration systems with a large number of processors, making it difficult to meet the needs of efficient simulation and verification.

Method used

A hierarchical interconnection network architecture is adopted, which decomposes the large-size cross switch into multiple small-size multiplexer arrays. Combined with the fully connected Crossbar structure inside the Boolean processor cluster, the CMESH topology between Boolean processor clusters, and the dedicated interconnection network of the TDM interface cluster, a three-layer interconnection network is formed. The communication path is optimized by combining the Boolean processor cluster, the TDM interface cluster, and the router.

Benefits of technology

It significantly reduces chip area and wiring complexity, lowers data communication latency, improves simulation efficiency and debuggability, simplifies compiler task allocation, and the topology facilitates compiler optimization and routing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833598A_ABST
    Figure CN121833598A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of simulation acceleration. The invention provides an inter-processor interconnection architecture for a hardware simulation acceleration system. According to the layered interconnection network architecture provided by the embodiment of the invention, a large-size crossbar switch is decomposed into a combination of a plurality of small-size multiplexer arrays, and the combination is divided into an intra-cluster network, a Boolean processor inter-cluster network and a TDM interface cluster interconnection network for connection, so that the chip area and the wiring complexity are remarkably reduced. Meanwhile, the performance evaluation method of the interconnection architecture between the processors is matched, and high-precision and high-efficiency simulation can be carried out on the interconnection architecture. The performance of different interconnection structures in the aspects of communication cycle number, area, utilization rate and the like can be accurately evaluated, and a key basis is provided for interconnection network optimization. Interconnection resource consumption and communication delay of the simulation acceleration chip are effectively reduced, a regular topological structure of the simulation acceleration chip facilitates task scheduling and optimization of a compiler, and the overall simulation efficiency and debuggeability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the field of hardware simulation acceleration, and in particular to an inter-processor interconnect architecture for a hardware simulation acceleration system. BACKGROUND

[0002] Currently, the logic design scale of various processor chips continues to expand, while the time window for market access is continuously narrowing. Especially for super large scale integrated circuits, especially for AI chips, supercomputer CPUs, server CPUs and other high-end chips, the demand for simulation during the design and debugging stage is increasing. Software simulation is an important means of chip logic function verification, and its advantage is to provide comprehensive signal observability, but the low efficiency of simulation has become a key bottleneck restricting the verification speed. With the explosive growth of chip design complexity, the contradiction between verification efficiency requirements and simulation performance has become increasingly prominent, greatly increasing the difficulty and time cost of functional verification. Another method for verifying and debugging chip logic design is to use hardware logic simulation for acceleration. Hardware logic simulation can greatly speed up the simulation speed and improve the simulation efficiency.

[0003] There are two basic types of design verification systems: hardware-driven systems that implement logic designs in programmable logic devices, and software-driven systems that simulate designs in one or more simulation processors. The first is a hardware-based design verification system that uses a large number of interconnected field programmable gate arrays (such as Synopsys' ZeBu and Mentor's Veloce), and the other is a hardware-based functional verification system that uses a large number of processor modules. Each processor module is provided with a dedicated computing unit, and in this processor-based system, the functions of the DUT are executed in the simulation processor, and the simulation processor calculates the output of the design. (such as Cadence's Palladium, which is composed of a large number of high-speed interconnected Boolean processors; the compiler divides the design between the processors to generate a schedule of Boolean operations in the correct temporal order).

[0004] The existing two hardware simulation acceleration systems have advantages and disadvantages. In terms of speed, the FPGA-based simulation acceleration system is faster than the processor-based simulation acceleration system. In terms of compilation time, the processor-based simulation acceleration system is less than the FPGA simulation acceleration system. The former only needs to ensure the correct order of the Boolean logic, and the latter is time-consuming in layout and wiring. In terms of debuggability, FPGA needs to add probes ila in advance to track signals, and the visibility is poor. The Boolean processor uses a special data memory to store intermediate results, and the signal visibility is 100%. In order to make full use of the advantages of FPGA and Boolean processor platform, the Sphinx simulation acceleration system has been proposed, which is composed of a layered Boolean processor cluster and an FPGA interconnection, connected through a dynamically configured interface circuit. By placing the modified part of the DUT on the Boolean processor and the rest on the FPGA, the respective advantages are used for functional verification.

[0005] The existing processor interconnection architecture faces great challenges in performance, power consumption and area when facing the continuous expansion of the number of cores, the increasing competition and conflict of communication resources. The interconnection structure in the prior art can be roughly divided into two directions: 1) Generally, for the interconnection between a small number of processors, the traditional Crossbar structure can exhibit good characteristics, such as no blocking, high bandwidth, low delay, but for hardware simulation acceleration systems, the number of processors increases sharply, at this time the resource consumption of the full connection Crossbar structure is huge, the logic delay is longer, and the routability is not friendly. The performance of the interconnection network is poor.

[0006] 2) With the development of NOC technology, many topological structures can also be applied to processor interconnection. Common on-chip interconnection schemes include Mesh structure and optimized 2DMesh, 3DMesh structure, star structure, ring structure, but when facing a large number of simulation processors, the routing hop count in these topological structures will be uncertain, causing unpredictable congestion and conflict.

[0007] According to the above, the basic feature of the interconnection network architecture is the modular design of the on-chip interconnection structure, and the modules are communicated in parallel through the micro network on the chip to improve the bandwidth to reduce the delay, while improving the scalability. However, the design of the interconnection network is closely related to the power consumption, area limitation and complex wiring problem, so how to build the topology of the interconnection network is crucial.

[0008] Therefore, it is necessary to improve one or more problems in the above-mentioned related technical solutions.

[0009] It should be noted that this section is intended to provide background information concerning the technical field of the disclosure and should not be considered as admitting prior art. SUMMARY

[0010] Embodiments of the present disclosure aim to provide a processor interconnection architecture for hardware simulation acceleration system, thereby at least partially overcoming one or more problems caused by limitations and defects of the related art.

[0011] According to a first aspect of embodiments of the present disclosure, a processor interconnection architecture for hardware simulation acceleration system is provided, comprising: a plurality of Boolean processor clusters interconnected by a first interconnection network, each Boolean processor cluster comprising a plurality of Boolean processors, and each Boolean processor being interconnected by a second interconnection network within the Boolean processor cluster; a plurality of interconnected TDM interface clusters connected to the plurality of Boolean processor clusters by a third interconnection network between the TDM interface clusters and the Boolean processor clusters.

[0012] Further, the first interconnection network comprises: a plurality of routers, each router comprising a first multiplexer array; wherein the first multiplexer array comprises: a first type of multiplexer array for supporting communication between the Boolean processor clusters connected by the same router, the number and width of the multiplexers in the first type of multiplexer array being related to the number of connected clusters; a second type of multiplexer array for supporting communication between the Boolean processor clusters connected by different routers.

[0013] Further, the second interconnection network comprises: a plurality of second multiplexer arrays, each second multiplexer array comprising a plurality of second multiplexers, the output of each second multiplexer being connected to an input port of a Boolean processor; wherein, the number of second multiplexers in each Boolean processor cluster is the same as the total number of input ports of the Boolean processors included in the Boolean processor cluster, the width of each second multiplexer is twice the number of Boolean processors in the cluster, half of the width of the second multiplexer is used for broadcasting of intra-cluster data, and the other half is used for accepting signals from outside the cluster, each second multiplexer outputs to a port of a Boolean processor, and each Boolean processor outputs to a second multiplexer.

[0014] Further, the third interconnection network comprises: A plurality of third multiplexer arrays, each third multiplexer array comprising a plurality of third multiplexers, the number and width of the third multiplexers being determined by the connection mode of the router.

[0015] Further, the first interconnection network between the clusters of Boolean processors adopts a CMesh topology, and the second interconnection network within the clusters of Boolean processors adopts a fully connected Crossbar structure.

[0016] According to a second aspect of the embodiments of the present disclosure, a performance evaluation method of an inter-processor interconnection architecture for a hardware simulation acceleration system is provided, comprising: generating periodic non-conflict transaction stimuli according to the netlist of the circuit to be verified; wherein the transaction stimuli simulates the task scheduling process of the compiler in the hardware simulation acceleration system; generating the inter-processor interconnection architecture corresponding to the real circuit structure composed of the first multiplexer array, the second multiplexer array and the third multiplexer array according to the interconnection architecture parameters defined by the user; loading the transaction stimuli into the inter-processor interconnection architecture for simulation, performing signal routing and recording simulation process data; generating a simulation report reflecting the performance indicators of the inter-processor interconnection architecture based on the simulation process data.

[0017] Further, in the step of generating periodic non-conflict transaction stimuli according to the netlist of the circuit to be verified, comprising: generating the netlist file of the benchmark circuit through a synthesis tool; constructing the mapping relationship of the lookup table to obtain the directed acyclic graph of the netlist file; topologically sorting the directed acyclic graph to obtain the critical path; calculating the task partitioning and mapping to each Boolean processor according to the constraint conditions and the critical path, and generating the non-conflict transaction stimuli distributed by period.

[0018] Further, the interconnection network structure is defined in a parameterized manner, and the parameters of the interconnection network structure at least include the total number of Boolean processors, the number of Boolean processors in each cluster of Boolean processors, the total number of clusters of Boolean processors, the number of TDM interfaces and the number of clusters of TDM interfaces.

[0019] Further, the performance indicators of the inter-processor interconnection architecture include: the total number of cycles required for processing transactions, the interconnection network area estimated based on the size of the multiplexer, and the utilization rate of the multiplexer.

[0020] The technical solutions provided by the embodiments of the present disclosure can include the following beneficial effects: In the embodiments of the present disclosure, through the above inter-processor interconnection architecture, on the one hand, the present application proposes a hierarchical interconnection network architecture, which divides a large-size crossbar into a combination of multiple small-size multiplexer arrays, and divides into an intra-cluster network, a Boolean processor inter-cluster network, and a TDM interface cluster interconnection network connected by a method, which significantly reduces the chip area and wiring complexity. At the same time, the present application provides a performance evaluation method of the inter-processor interconnection architecture, which can simulate the above interconnection architecture with high precision and high efficiency. The performance of different interconnection structures in terms of communication cycle number, area, utilization rate, etc. can be accurately evaluated, which provides a key basis for interconnection network optimization. The present application effectively reduces the interconnection resource consumption and communication delay of the simulation acceleration chip, and the regular topology structure facilitates the task scheduling and optimization of the compiler, which improves the overall simulation efficiency and debuggability. On the other hand, the hierarchical CMesh interconnection network structure is adopted to realize the data communication between the multi-core processors, which greatly saves the resources by reducing the multiplexer width, significantly reduces the data communication delay, ensures the frequency of processor data communication, and cooperates with the compiler to ensure the non-blocking and conflict-free data transmission. The number of hops of data communication between any two nodes of the interconnection network of the present application is determined, which simplifies the difficulty of task allocation of the compiler, and the coupling between the three networks of the present application adopts a regular topology, which is very friendly to the later wiring. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate one embodiment consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0022] Figure 1 A schematic diagram of an inter-processor interconnection architecture for a hardware simulation acceleration system in an exemplary embodiment of the present disclosure is shown; Figure 2 A schematic diagram of a hardware simulation acceleration system based on a processor in an exemplary embodiment of the present disclosure is shown; Figure 3 A schematic diagram of a multiplexer array within a Boolean processor cluster in an exemplary embodiment of the present disclosure is shown; Figure 4 A detailed diagram of a router between clusters in an exemplary embodiment of the present disclosure is shown; Figure 5 A data flow diagram of communication between processors within a cluster in an exemplary embodiment of the present disclosure is shown; Figure 6 A data flow diagram of communication between processors between clusters in the same BLOCK in an exemplary embodiment of the present disclosure is shown; Figure 7 A data flow diagram showing the communication between inter-cluster processors in adjacent BLOCKs in an example embodiment of the present disclosure; Figure 8 A data flow diagram showing the communication between inter-cluster processors in diagonal BLOCKs in an example embodiment of the present disclosure; Figure 9 A flow chart showing the steps of a method for performance evaluation of an inter-processor interconnect architecture for a hardware emulation acceleration system in an example embodiment of the present disclosure; Figure 10 A flow chart showing the simulation procedure of a dedicated interconnect network emulator in an example embodiment of the present disclosure; Figure 11 A diagram showing the stimulation, topology, and report file in an example embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. Features described in one implementation can be combined with features described in a different implementation. The described features, structures, or characteristics can be combined in one or more implementations.

[0024] In addition, the drawings are to be considered in all respects as illustrative and not restrictive; identical reference numerals have been used, where possible, to denote identical or similar features, and thus repetition of the description thereof will be omitted. Some of the blocks in the drawings can be functional blocks that do not necessarily have a corresponding physical or logical structure.

[0025] An inter-processor interconnect architecture for a hardware emulation acceleration system is provided in the present example implementation. Referring to FIG. 1, the inter-processor interconnect architecture for a hardware emulation acceleration system can include: Figure 1 a plurality of Boolean processor clusters interconnected by a first interconnect network, each Boolean processor cluster including a plurality of Boolean processors, and each Boolean processor being interconnected by a second interconnect network within the Boolean processor cluster; a plurality of interconnected TDM interface clusters connected to the plurality of Boolean processor clusters by a third interconnect network between the TDM interface clusters and the Boolean processor clusters.

[0026] ​By the above-mentioned inter-processor interconnection architecture for hardware simulation acceleration system, on the one hand, the application proposes a layered interconnection network architecture, which divides a large-size crossbar into a combination of multiple small-size multiplexer arrays, divides into an intra-cluster network, a Boolean processor inter-cluster network, and a TDM interface cluster interconnection network connected by a method, which significantly reduces the chip area and wiring complexity. On the other hand, the layered CMesh interconnection network structure is adopted to realize the data communication between the multi-core processors, which greatly saves the resources by reducing the multiplexer width, significantly reduces the delay of data communication, guarantees the frequency of processor data communication, and cooperates with the compiler to guarantee the non-blocking and non-conflict of data transmission. The number of hops of data communication between any two nodes of the interconnection network of the application is determined, which simplifies the difficulty of task allocation of the compiler, and the three networks and the coupling between them of the application all adopt regular topologies, which are very friendly to the later wiring.

[0027] In the following, the above-mentioned inter-processor interconnection architecture for hardware simulation acceleration system in the present example embodiment will be described in more detail. Figures 1 to 8 The above-mentioned inter-processor interconnection architecture for hardware simulation acceleration system in the present example embodiment will be described in more detail.

[0028] In one embodiment, the application proposes an inter-processor interconnection architecture for a hardware simulation acceleration system. It includes a plurality of Boolean processors (i.e. simulation processors) and a plurality of TDM interfaces, a TDM interface for data communication with an FPGA, a plurality of processor clusters, a cluster containing a plurality of Boolean processors, a plurality of multiplexer arrays in the cluster supporting communication between processors, a plurality of routers supporting communication between clusters, a plurality of TDM dedicated multiplexer arrays supporting communication between TDM and Boolean processors, and a lightweight dedicated simulation platform.

[0029] The inter-processor interconnection architecture proposed by the application is based on the idea of layered interconnection. The complete system interconnection architecture is divided into three parts, the first part is the interconnection network within the Boolean processor cluster (i.e. the second interconnection network), the second part is the interconnection network between the Boolean processor clusters (i.e. the first interconnection network), and the third part is the interconnection network between the TDM cluster and the Boolean processor cluster (i.e. the third interconnection network).

[0030] For the proposed interconnection network inside the Boolean processor cluster, the overall full connection Crossbar architecture is adopted, and multiple Boolean processors are aggregated into a cluster, where each hardware emulation acceleration chip includes multiple processor clusters. In the hardware emulation acceleration system of the present application, the number of multiplexers in each cluster is the same as the total number of input ports of the Boolean processors included in the cluster. Half of the width of the multiplexer of the cluster is used for broadcasting of data inside the cluster, and the other half is used to accept signals outside the cluster. Therefore, the width of each multiplexer is twice the number of Boolean processors in the cluster. Overall, each multiplexer outputs to a port of a Boolean processor, and each Boolean processor outputs to a multiplexer.

[0031] For the proposed interconnection network between Boolean processor clusters, the CMesh topology in NOC is adopted, which evolves from the 2D MESH topology. The key component in the MESH structure is the Router, which is used to support routing. The interconnection network between Boolean processor clusters is composed of multiple Routers. The cluster connected by a Router is called a Block. The Router is responsible for data communication between Blocks and clusters. Each Router is composed of a multiplexer array. The multiplexers in the Router are divided into two categories. One category is responsible for communication between Blocks, i.e., communication between clusters inside the Block. The number of multiplexers is the number of clusters plus 3, and the width of the multiplexer is the number of clusters plus 2. The other category is responsible for communication between Blocks. The number of multiplexers is 3, and the width of the multiplexer is 2.

[0032] For the proposed interconnection network between TDM clusters and Boolean processor clusters, the network topology is relatively simple compared to the previous two network topologies. Because according to the actual demand, there is no need for communication between TDM clusters, the interconnection network design is greatly simplified. Referring to the construction method of the Boolean processor cluster, multiple TDM interfaces are aggregated into a cluster, where each hardware emulation acceleration chip includes multiple TDM clusters. According to the actual demand, the TDM data communication frequency is much smaller than the Boolean processor communication frequency. A dedicated TDM cluster interconnection network, i.e., a dedicated multiplexer array, is customized. The number and width of the multiplexers are strongly related to the connection method of the Router.

[0033] In one specific embodiment, Figure 2A processor-based hardware emulation acceleration system overview of an embodiment is shown, the system includes a host or computer workstation, a Boolean processor coupled with FPGA emulation acceleration chip. The host workstation acts as a software driver for the emulation acceleration chip. The initial step of the design compilation converts the HDL design into a netlist description. The netlist is a description of the design components and electrical interconnections. The netlist includes all the circuit elements required to implement the design, including: combinational logic (i.e. gates), sequential logic (i.e. flip-flops and latches), and memory (i.e. SRAM, DRAM, etc.). The stable circuit portions that do not require frequent modification and have been verified are then generated into a bitstream file loaded into the programmable logic device FPGA, the netlist that requires frequent modification is converted into a series of statements that will be executed by the emulation processor, usually in the form of Boolean equations. These statements (also known as steps) are loaded into the emulation processor, which executes the statements in order, one step at a time. The processor calculates the output and saves it in a data storage array, which can then be transferred back to the host workstation or used for future processing steps. The Boolean processor coupled with FPGA emulation acceleration chip is composed of a large number of special-purpose Boolean processors coupled with FPGAs, because the communication between the Boolean processor and the FPGA is not tight, only the key signals at the partition of the netlist are needed, so the Boolean processor and the FPGA are connected through time division multiplexing. The Boolean processor schedules the required instructions through the program counter, and sends the required data to the special-purpose Boolean operation calculation unit, the output of the calculation unit is output through the multiplexer array, the input of the calculation unit can come directly from the output data of the multiplexer array, or from the previously cached data in the data storage unit.

[0034] In one specific embodiment, Figure 1 A specific hardware emulation acceleration system interconnection structure diagram of the present patent is shown, the system is composed of 1024 Boolean processors, 1280 TDM interfaces, 16 Boolean processor emulation clusters, 20 TDM interface clusters, and 4 special-purpose Routers, each Boolean processor cluster is composed of 64 Boolean processors and a multiplexer array. As Figure 3The figure shows a schematic diagram of a Boolean processor cluster, which internally contains 64 Boolean processors and 256 multiplexer groups. The interface of the Boolean processor is 4-input 1-output. Each multiplexer outputs to the port of the Boolean processor. Each port has a dedicated multiplexer and does not appear in the form of multiplexing. The width of each multiplexer is 128, of which 64 signals come from the output of the 64 Boolean processor cluster inside the cluster to support the broadcast communication between the processors inside the cluster, and the other 64 signals come from the output signals of other clusters outside the cluster. Unlike the connection method in patent US 7555.423 B2, the output of the Boolean processor in this application is directly output through the multiplexer array and is directly connected to the input port of the Boolean processor. In patent US 7555.423 B2, the output of the multiplexer array is output to the shared data storage array, and then the processor selects the signal. The connection method of this application is based on the more advanced nature of the dedicated cache unit inside the Boolean processor and software scheduling, which ensures lower data communication delay between processors and higher communication efficiency.

[0035] In a specific embodiment, Figure 4 The detailed diagram of the Router is shown, Figure 1 The hardware simulation acceleration system shown is composed of 4 Routers, each of which is connected with 4 Boolean processor clusters and 5 TDM interface clusters, as well as two adjacent Routers in the horizontal and vertical directions, to realize data communication between any Boolean processor of any two clusters. The Router structure in this application cancels the path cache method and adopts combination logic plus timing logic to optimize data transmission delay and achieve stronger performance. The parameters of the 4 Routers are completely consistent, and each is a Router with a radix of 7, realizing 7-to-7 switching. Each Router integrates 4 Boolean processor clusters, 1 TDM cluster (5 TDM clusters are first gated outside the Router to transmit only 1 TDM cluster connected to the Router), and two adjacent Routers. The 7 ports (3 types of ports) of the Router can perform multicast or unicast, but cannot perform the process of sending data to itself, so each port can only propagate in 6 directions. The Router is internally composed of two multiplexer arrays. The first multiplexer array is responsible for supporting communication between adjacent Blocks and the same Block, which is composed of 6 multiplexers, each with a width of 5. The second multiplexer array supports communication between diagonal Blocks, which selects the input through 1 multiplexer and realizes the bypass output function through 2 multiplexers. The multiplexer arrays work together to realize communication between any processors.

[0036] In one specific embodiment, Figures 5 to 8 The dashed lines in FIG. 1 represent the data flow direction between the processors, and the specific hop count between the processors in different locations can be derived from the data flow. Figure 5 The data flow path 1 between the Boolean processors in the same cluster is shown, and the communication delay is 1 delay of a 128-wide multiplexer, denoted as t1, due to the intra-cluster broadcast. Figure 6 The data flow path 2 between the clusters of the same router is shown, and the communication delay is 1 delay of a 128-wide multiplexer plus 1 delay of a 5-wide multiplexer, denoted as t1+t2. Figure 7 The data flow path 3 between the clusters of adjacent routers is shown, and the communication delay is 1 delay of a 128-wide multiplexer plus 2 delays of a 5-wide multiplexer plus 2 delays of a 2-wide multiplexer, denoted as t1+2t2+2t3. Figure 8 The data flow path 4 between the clusters of diagonal routers is shown, and the communication delay is 1 delay of a 128-wide multiplexer plus 2 delays of a 5-wide multiplexer plus 3 delays of a 2-wide multiplexer, denoted as t1+2t2+3t3. The hop count between any two processors is within the above four paths, so the hop count is relatively certain, and the advantage of this design is that the compiler for software is friendly to task allocation, and the compiler understands the interconnection structure.

[0037] Further, the example embodiment also provides a performance evaluation method for an inter-processor interconnection architecture of a hardware simulation acceleration system. Referring to FIG. 2, the inter-processor interconnection architecture for the hardware simulation acceleration system can include: Figure 9 According to a user-defined interconnection architecture parameter, a first multiplexer array, a second multiplexer array, and a third multiplexer array are generated to form an inter-processor interconnection architecture corresponding to a real circuit structure. According to a user-defined interconnection architecture parameter, a first multiplexer array, a second multiplexer array, and a third multiplexer array are generated to form an inter-processor interconnection architecture corresponding to a real circuit structure. According to a user-defined interconnection architecture parameter, a first multiplexer array, a second multiplexer array, and a third multiplexer array are generated to form an inter-processor interconnection architecture corresponding to a real circuit structure. The transaction stimulus is loaded into the inter-processor interconnection architecture for simulation, signal routing is performed, and simulation process data is recorded. Based on the simulation process data, a simulation report reflecting the performance indicators of the inter-processor interconnection architecture is generated.

[0038] In one embodiment, the performance evaluation method for the inter-processor interconnection architecture of the hardware simulation acceleration system corresponds to a simulation platform as shown in FIG. 3. Figure 10The simulation platform can include: an excitation generation module, a network topology module, a routing function module, and a statistics module. The excitation generation module is configured to generate periodic non-conflict transaction excitation according to a circuit netlist to be verified, wherein the transaction excitation simulates a task scheduling process of a compiler in a hardware simulation acceleration system. The network topology module is configured to generate an inter-processor interconnection architecture corresponding to a real circuit structure, which is composed of a first multiplexer array, a second multiplexer array, and a third multiplexer array, according to user-defined interconnection architecture parameters. The routing function module is configured to load the transaction excitation into the inter-processor interconnection architecture for simulation, perform signal routing, and record simulation process data. The statistics module is configured to generate a simulation report reflecting performance indicators of the inter-processor interconnection architecture based on the simulation process data.

[0039] In one embodiment, Python, as a modern high-level language, has a rich ecosystem and concise syntax structure, and exhibits many advantages in building an interconnection network simulator. Compared with SystemC based on C++ library, Python has significant advantages in rapid prototyping, code readability, third-party library integration, and cross-platform deployment. The network simulator implemented using Python can flexibly load multiple information streams for parallel processing, and Python can effectively simulate the time allocation of instruction execution, greatly improving development efficiency and model scalability while ensuring sufficient accuracy.

[0040] Since the Boolean processors are homogeneous units, the communication between many Boolean processors has symmetry, that is, at a certain moment, multiple processors perform the same operation, causing serious blocking of the multiplexer array, while at a certain moment, a large number of processors are idle, so the simulated model is not accurate and does not conform to the real scene in the hardware simulation acceleration system. The compiler in the hardware simulation acceleration system learns the interconnection network structure and schedules non-conflict transactions at each cycle, and delays conflict transactions for processing. Therefore, to simulate the real compiler scheduling process, the transaction generation of the simulation platform also follows this principle.

[0041] Since the interconnection structure mentioned in the present application is optimized between multiplexer arrays, and the router is also built by different selector arrays, many traditional network simulators fail to reflect this important physical property, so a plurality of width multiplexer networks are constructed in the present application to reflect the real circuit structure. At the same time, since the area of the interconnection network can be reflected by the width and size of the multiplexer, an area evaluation model can also be established.

[0042] The transaction excitation based on the simulator (i.e., the simulation platform) is distributed by cycle from the beginning, so the simulator provides a time-parallel simulation method, which greatly improves the simulation efficiency.

[0043] The simulation platform evaluation standard proposed in the application is that, firstly, internal transactions per cycle are conflict-free, a large number of transactions are distributed to multiple cycles, and the total number of cycles required for processing transactions in different topological structures is compared, so that the performance of the interconnection structure can be judged. In order to be closer to the real situation, the excitation of the simulation platform is evolved from the real circuit netlist, the network topology is built by a large number of multiplexers, and the evaluation of the cycle is obtained by the labeling of the multiplexer.

[0044] In one specific embodiment, Figure 10 The simulation process of the special interconnection network simulator is shown, and it has been mentioned in the foregoing that the network simulator in the application includes an excitation generator, a network topology generator, a routing function module, a statistical module, and the corresponding simulation process is divided into three steps. The first step is that the excitation generator generates the excitation corresponding to the real circuit, the second step is that the network topology generator generates the corresponding network topology structure according to the user-defined parameters, and the third step is that the excitation generated in the first two steps is loaded into the routing function module for simulation, and the relevant results are recorded during the simulation process.

[0045] The use steps and functional characteristics of the network simulator of the application will be specifically described below through an example. It is assumed that a hierarchical Mesh interconnection network mentioned in the foregoing is to be generated, that is, a main topology structure including 1024 nodes and 4 routers is to be generated, and the processor nodes are connected through the inter-cluster multiplexer array and the router, as shown in Figure 10 The key parameters such as size, bpus_num, ports_num, bpus_per_cluster, and cluster_num are defined in the network topology generator to construct the multiplexer network, and the required network topology is generated.

[0046] The main steps of the excitation generator mentioned in the application include: the first step is to generate a plurality of benchmark circuit netlist files based on 4-input LUT through the open source synthesis kit YOSYS; the second step is to construct the LUT mapping and the mapping relationship from the port to the LUT, that is, the DAG graph of the real netlist is obtained; the third step is to design a topology sorting algorithm with the number of nodes as the core to obtain the ASAP and ALAP graphs of the DAG graph, and the critical path is obtained; the fourth step is to divide the graph into each Boolean processor according to the out-degree number and the critical path constraints, and generate the excitation of each cycle.

[0047] Figure 11The incentive, topology, report file schematic diagram is shown, the incentive file can be generated as a txt or other format file, after the generator generates various files, the user calls the simulation platform to read the file, and the new project loads these files to simulate and generate the report file, which can reflect multiple indicators of the interconnected circuit, such as cycle number, area amount, utilization rate and the like.

[0048] The inherent benefits of the embodiments described in the application also include fewer selection signals. Compared with the traditional Crossbar interconnection structure, taking 1024 cores as an example, 4096 multiplexers are needed, and the width of each multiplexer is 1024. In the application, the large-width multiplexer is simplified to a 128-width multiplexer matched with some small-size selectors to realize the same function. The reduction of the size of the multiplexer will greatly reduce the number of selection signals, further reduce the interconnection scale on the chip, and thus be beneficial to the wiring work and reduce the area required for realizing the hardware simulation acceleration chip.

[0049] The application describes an interconnection circuit structure for connecting a large number of Boolean processors and TDM interfaces. The method and system have the stability and low delay characteristics of connection, greatly reduce the chip amount required by the interconnection structure, and significantly reduce the power consumption of the integrated circuit, and the regular topology structure is also beneficial to the design of the compiler.

[0050] The beneficial effects of the application are: The layered CMesh interconnection network structure is adopted to realize the data communication between the multi-core processors, the width of the multiplexer is greatly reduced to save resources, the delay of data communication is significantly reduced, the frequency of processor data communication is ensured, and the non-blocking and non-conflict of data transmission are ensured in cooperation with the compiler. The number of hops of data communication between any two nodes of the interconnection network of the application is determined, the difficulty of task allocation of the compiler is simplified, and the coupling between the three networks of the application adopts a regular topology, which is very friendly to the later wiring.

[0051] A lightweight simulation platform is built by using python language, which is dedicated to the interconnection network scene of the application, supports the design space exploration of different bases and different topologies, and the topology structure is built by multiplexers, which is close to the real circuit and increases the accurate modeling of the physical layer constraints. Parallel simulation transactions improve the simulation efficiency.

[0052] It should be understood that the orientation or positional relationship indicated by terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise" and the like in the above description is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the embodiments of the present disclosure and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the embodiments of the present disclosure.

[0053] In addition, the terms "first", "second", "third", "fourth", "fifth", "sixth", "seventh" and "eighth" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implying the number of the technical features indicated. Therefore, the features defined as "first", "second", "third", "fourth", "fifth", "sixth", "seventh" and "eighth" can explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, the meaning of "plurality" is two or more, unless otherwise explicitly specified and limited.

[0054] In the embodiments of the present disclosure, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting", "fixing" and the like should be understood in a broad sense, for example, can be fixedly connected, or can be detachably connected, or can be integrated; can be mechanically connected, or can be electrically connected; can be directly connected, or can be indirectly connected through an intermediate medium; can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present disclosure can be understood according to the specific circumstances.

[0055] In the embodiments of the present disclosure, unless otherwise explicitly specified and limited, the first feature "on" or "under" the second feature can include that the first and second features are in direct contact, or can include that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, the first feature "on", "above" and "on" the second feature includes that the first feature is directly above and obliquely above the second feature, or only indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature "under", "below" and "under" the second feature includes that the first feature is directly below and obliquely below the second feature, or only indicates that the horizontal height of the first feature is less than that of the second feature.

[0056] In the description of the disclosure, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the disclosure. In the description of the disclosure, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in the specification.

[0057] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the disclosure that follow the general principles thereof and include the general principles thereof and include other known or customary features not specifically mentioned herein. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the disclosure are indicated by the appended claims.

Claims

1. A processor interconnect architecture for a hardware emulation acceleration system, characterized in that, include: Multiple Boolean processor clusters are interconnected by a first interconnection network between Boolean processor clusters, each Boolean processor cluster includes multiple Boolean processors, and each Boolean processor is interconnected by a second interconnection network within the Boolean processor cluster; Multiple interconnected TDM interface clusters are connected to multiple Boolean processor clusters via a third interconnection network between the TDM interface clusters and the Boolean processor clusters.

2. The processor interconnect architecture for hardware simulation acceleration systems according to claim 1, characterized in that, The first interconnection network includes: Several routers, each router including a first multiplexer array; wherein the first multiplexer array includes: The first type of multiplexer array is used to support communication between Boolean processor clusters connected to the same router. The number and width of multiplexers in the first type of multiplexer array are related to the number of clusters connected. The second type of multiplexer array is used to support communication between Boolean processor clusters connected to different routers.

3. The processor interconnect architecture for hardware simulation acceleration systems according to claim 2, characterized in that, The second interconnection network includes: Multiple arrays of second multiplexers, each array comprising multiple second multiplexers, the output of each second multiplexer connected to the input port of a Boolean processor; wherein, The number of second multiplexers in each Boolean processor cluster is the same as the total number of input ports of the Boolean processors contained in the Boolean processor cluster. The width of each second multiplexer is twice the number of Boolean processors in the cluster. Half of the width of the second multiplexer is used for broadcasting data within the cluster, and the other half is used for receiving signals from outside the cluster. Each second multiplexer outputs to a port of a Boolean processor, and each Boolean processor outputs to a second multiplexer.

4. The processor interconnect architecture for hardware simulation acceleration systems according to claim 3, characterized in that, The third interconnection network includes: A dedicated array of multiple third multiplexers, each array comprising multiple third multiplexers, the number and width of which are determined by the router’s connection method.

5. The processor interconnect architecture for hardware simulation acceleration systems according to claim 4, characterized in that, The first interconnection network between Boolean processor clusters adopts a CMESH topology, while the second interconnection network within the Boolean processor cluster adopts a fully connected Crossbar structure.

6. A performance evaluation method for inter-processor interconnect architecture in a hardware simulation acceleration system, characterized in that, include: Periodic conflict-free transaction stimuli are generated based on the netlist of the circuit to be verified; wherein, the transaction stimuli simulate the task scheduling process of the compiler in the hardware simulation acceleration system; Based on user-defined interconnect architecture parameters, an inter-processor interconnect architecture corresponding to the actual circuit structure is generated, consisting of a first multiplexer array, a second multiplexer array, and a third multiplexer array. The transaction stimulus is loaded into the inter-processor interconnect architecture for simulation, signal routing is executed, and simulation process data is recorded. Based on the simulation process data, a simulation report reflecting the performance indicators of the inter-processor interconnect architecture is generated.

7. The performance evaluation method for inter-processor interconnect architecture of a hardware simulation acceleration system according to claim 6, characterized in that, The steps of generating periodic conflict-free transaction stimuli based on the netlist of the circuit to be verified include: Generate the netlist file for the reference circuit using synthesis tools; Construct lookup table mapping relationships to obtain a directed acyclic graph of the netlist file; Perform a topological sort on the directed acyclic graph to obtain the critical path; Based on constraints and critical path calculations, task partitioning is performed and mapped to various Boolean processors, generating conflict-free transaction incentives allocated periodically.

8. The performance evaluation method for inter-processor interconnect architecture of a hardware simulation acceleration system according to claim 7, characterized in that, The interconnection network structure is defined using a parameterized approach. The interconnection network structure parameters include at least the total number of Boolean processors, the number of Boolean processors in each Boolean processor cluster, the total number of Boolean processor clusters, the number of TDM interfaces, and the number of TDM interface clusters.

9. The performance evaluation method for inter-processor interconnect architecture of a hardware simulation acceleration system according to claim 8, characterized in that, Performance metrics reflecting the inter-processor interconnect architecture include: The total number of cycles required to process transactions, the interconnection area estimated based on multiplexer size, and the utilization rate of the multiplexer.