Hardware accelerator, die, multi-die co-system, method
By using hardware accelerators in a multi-die collaborative system to achieve synchronized state transitions between multiple dies, the limitations of single-chip NPU performance improvement are overcome, providing greater computing power and flexible scalability, making it suitable for high-end application scenarios.
Patent Information
- Application Number
- CN202511241116.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-02
AI Technical Summary
The performance improvement of single-chip NPU is limited by factors such as chip area and manufacturing process, making it difficult to meet the needs of high-end application scenarios such as artificial intelligence computing and ultra-high performance computing in large data centers. Moreover, the single-chip design has fixed functions, which are difficult to expand or upgrade, and cannot be applied to complex application scenarios with multiple functions working together.
A multi-die collaborative system is adopted, which realizes the synchronization state transition between multiple dies through the synchronization control module and synchronization transceiver module in the hardware accelerator, generates and processes synchronization signal output data, and realizes collaborative computing between multiple dies.
Without altering the chip architecture, multi-die technology enables greater computing power and superior computing capabilities, supporting flexible expansion and collaborative computing while reducing design costs and complexity.
Smart Images

Figure CN120806008B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of multi-die cooperation, and in particular to a hardware accelerator, a die, a multi-die cooperation system and a method. BACKGROUND
[0002] Generally, the performance improvement of a single-chip NPU (Neural Processing Unit) is limited by factors such as chip area and process technology. With the increasing demand for computing power, a single chip is difficult to meet the needs of some high-end application scenarios, such as artificial intelligence computing in large data centers and ultra-high performance computing, which require extreme computing power. Moreover, once the design of a single-chip NPU is completed, its functions are relatively fixed, and it is difficult to perform large-scale functional expansion or upgrade. It is also difficult to be well adapted to complex application scenarios that require multiple functions to work together and have high requirements for computing and storage capabilities, such as high-end servers and professional graphics processing workstations. SUMMARY
[0003] In view of this, embodiments of the present disclosure provide a hardware accelerator, a die, a multi-die cooperation system and a method to at least solve or alleviate the above problems.
[0004] According to a first aspect of embodiments of the present disclosure, a hardware accelerator is provided, which is applied to a first die in a multi-die cooperation system. The hardware accelerator comprises: a synchronization control module, configured to generate and process synchronization signal output data, the synchronization signal being used to control the synchronization state transition between multiple dies; and a synchronization transceiver module, configured to receive the synchronization signal output data sent by the synchronization control module, and send synchronization signal input data sent by other dies to the synchronization control module.
[0005] According to a second aspect of embodiments of the present disclosure, a die is provided, which applies the hardware accelerator of the first aspect.
[0006] According to a third aspect of embodiments of the present disclosure, a multi-die cooperation system is provided, which comprises a first die and a second die, and the hardware accelerator of the first aspect is run on the first die and the second die.
[0007] According to a fourth aspect of embodiments of the present disclosure, a hardware acceleration method is provided, which is applied to a first die in a multi-die cooperation system. The method comprises: generating and processing synchronization signal output data, the synchronization signal being used to control the synchronization state transition between multiple dies; and receiving the synchronization signal output data, and sending synchronization signal input data sent by other dies to the synchronization control module.
[0008] According to a fifth aspect of the embodiments of the present disclosure, a hardware acceleration method is provided, applied to a multi-die cooperative system including a first die and a second die, and the method includes: the first die generating synchronization signal output data, entering a pause state and recording a second die identifier, and sending the synchronization signal output data to the second die; after the second die receives the synchronization signal output data, the second die returns response synchronization signal input data; and the first die jumps to a working state according to the response synchronization signal input data.
[0009] According to the hardware accelerator scheme provided by the embodiments of the present disclosure, the synchronization control module of the hardware accelerator generates and processes synchronization signal output data, and the synchronization signal is used to control the synchronization state jump between the multiple dies. The synchronization control module sends the synchronization signal output data to the synchronization transceiver module, and the synchronization transceiver module sends the received synchronization signal input data sent by other dies to the synchronization control module. The embodiments of the present disclosure realize a neural network computing platform with greater computing power and excellent cooperative computing capability through the multi-die (DIE) technology without changing the chip architecture. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art based on these drawings.
[0011] Figure 1 is a structural schematic diagram of a hardware accelerator according to an embodiment of the present disclosure;
[0012] Figure 2 is a structural schematic diagram of a hardware accelerator according to another embodiment of the present disclosure;
[0013] Figure 3 is an example diagram of synchronization signal input data of a hardware accelerator according to still another embodiment of the present disclosure;
[0014] Figure 4 is a structural schematic diagram of a synchronization control module of a hardware accelerator according to still another embodiment of the present disclosure;
[0015] Figure 5 is a structural schematic diagram of a front-end unit of a hardware accelerator according to still another embodiment of the present disclosure;
[0016] Figure 6 is a schematic diagram of a finite state machine of a hardware accelerator according to still another embodiment of the present disclosure;
[0017] Figure 7is a structural schematic diagram of a global synchronization unit of a hardware accelerator according to another embodiment of the present disclosure;
[0018] Figure 8 is a structural schematic diagram of a synchronization bus unit of a hardware accelerator according to another embodiment of the present disclosure;
[0019] Figure 9 is a structural schematic diagram of a synchronization transceiver module of a hardware accelerator according to another embodiment of the present disclosure;
[0020] Figure 10 is a schematic diagram of a state machine of a protocol conversion unit of a hardware accelerator according to another embodiment of the present disclosure;
[0021] Figure 11 is a schematic diagram of an application scenario of a hardware accelerator according to another embodiment of the present disclosure;
[0022] Figure 12 is a schematic diagram of a die according to another embodiment of the present disclosure;
[0023] Figure 13 is a schematic diagram of a multi-die cooperative system according to another embodiment of the present disclosure;
[0024] Figure 14 is a flowchart of a hardware acceleration method according to another embodiment of the present disclosure;
[0025] Figure 15 is a flowchart of a hardware acceleration method according to another embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0027] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or in a different section / subsection in any manner.
[0028] In the description of the embodiments of the disclosure, the term "comprising" and similar terms thereof are to be understood as open-ended, i.e., "including but not limited to". The term "based on" is to be understood as "based at least in part on". The term "one embodiment" or "the embodiment" is to be understood as "at least one embodiment". The term "some embodiments" is to be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or same objects. Other explicit and implicit definitions can also be included below.
[0029] In some embodiments of the disclosure, data of users, acquisition and / or use of data, etc. can be involved. These aspects all comply with corresponding laws and regulations and relevant provisions. In some embodiments of the disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are performed on the premise that the user is aware of and confirms. Accordingly, when implementing each embodiment of the disclosure, the type of data or information that can be involved, the use range, the use scenario, etc. should be notified to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations. The specific notification and / or authorization manner can vary according to the actual situation and application scenario, and the scope of the disclosure is not limited in this aspect.
[0030] In the specification and embodiments of the disclosure, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. The user refuses to process personal information other than the necessary information required for the basic function, which does not affect the user's use of the basic function.
[0031] First, some nouns or terms appearing in the process of describing the embodiments of the disclosure are applicable to the following explanations.
[0032] Die: refers to a chip that has not been packaged after completing the wafer manufacturing process.
[0033] Die2Die: a high-speed communication interface technology used to connect multiple dies (chips) within the same package.
[0034] AXI: an extensible interface protocol used to realize high-speed data transmission between modules in a system on chip (SoC).
[0035] The encoding scheme of the transmission identifier provided by the embodiments of the disclosure is described in detail below with reference to the accompanying drawings.
[0036] Hardware accelerator
[0037] See Figure 1The embodiment of the present disclosure provides a hardware accelerator applied to a first die in a multi-die cooperative system, and the hardware accelerator comprises:
[0038] A synchronization control module 11 is configured to generate and process synchronization signal output data, and the synchronization signal is used for controlling synchronization state transition between the plurality of dies.
[0039] A synchronization transceiver module 12 is configured to receive the synchronization signal output data sent by the synchronization control module 11, and send the synchronization signal input data sent by other dies to the synchronization control module 11.
[0040] The embodiment of the present disclosure realizes the cooperation between the plurality of dies through the synchronization control module 11 and the synchronization transceiver module 12, the plurality of dies are integrated together, and each die can independently perform calculation, so that large-scale parallel calculation is realized. The NPU of the multi-die can flexibly expand the calculation capability according to actual needs, and different computing resources are called according to different tasks, so that the flexibility and expansibility are higher. The embodiment of the present disclosure reduces the design cost and complexity, and improves the chip yield.
[0041] In addition, the multi-die NPU can make each die be the same design, reduce the chip area and design difficulty, and then reduce the cost and implementation difficulty.
[0042] Referring to Figure 2 In some specific implementations of the embodiment of the present disclosure, the hardware accelerator further comprises:
[0043] An on-chip network module 13 is configured to connect the synchronization control module 11 and the synchronization transceiver module 12 through an on-chip network protocol.
[0044] A neural network processor 14 is configured to interact with the synchronization control module 11, and send a control instruction to the synchronization control module 11 to generate synchronization signal output data.
[0045] A configuration bus module 15 is configured to set a control instruction buffer of the synchronization control module 11 and configure a transmission address of the synchronization transceiver module 12.
[0046] In some specific implementations of the embodiment of the present disclosure, the synchronization control module sends the synchronization signal output data to the synchronization transceiver module through a synchronization output bus, the synchronization transceiver module receives the synchronization signal input data sent by other dies through a synchronization input bus, and the synchronization signal output data and the synchronization signal input data both comprise a source identification (Source ID), a target identification (Destination ID) and a transmission type, and handshake is performed through a valid signal (vld) and a response signal (ack).
[0047] vld (Valid): Valid signal, indicating that the data sent by the sending end (such as the synchronization signal input data sent by the synchronization input bus) is valid and ready to be received.
[0048] ack (Acknowledge): Acknowledgment signal, indicating that the receiving end has successfully received the data.
[0049] handshake (handshake): A communication protocol that ensures data between the sending end and the receiving end.
[0050] Source ID: Indicates the identification of the die sending the synchronization signal (such as VIP0, VIP1).
[0051] Destination ID: Indicates the identification of the target die of the synchronization signal (such as VIP0, VIP1).
[0052] Transmission type: Indicates the type or purpose of the synchronization signal, which can be used to distinguish different synchronization operations (such as triggering pause release, state jump, etc.).
[0053] Specifically, the source identification is used to indicate the sending die of the data, the destination identification is used to indicate the receiving die of the data (i.e. whether the die is the target die corresponding to the data), and the transmission type is used to indicate the specific type of the data.
[0054] Referring to Figure 3 , the embodiment is an example of synchronization signal input data, which also includes the category of identification and supports 16 categories, and the embodiment of the disclosure only takes one type as an example, which is set to 0. The embodiment of the disclosure will not be described again.
[0055] In some specific implementations of the embodiment of the disclosure, referring to Figure 4 , the synchronization control module 11 includes:
[0056] The front-end unit 111 (FE) is configured to generate synchronization signal output data and control the synchronization state jump of the first die according to the control instruction corresponding to the received synchronization signal input data.
[0057] The global synchronization unit 112 is configured to synchronize and clock the synchronization signal input data and the synchronization signal output data.
[0058] The synchronization bus unit 113 is configured to process the synchronization signal output data and determine whether the synchronization signal input data corresponds to the first die.
[0059] The embodiments of the present disclosure realize generation and processing of the synchronization signal output data by the front-end unit 111, the synchronous bus unit 113 and the global synchronization unit 112, and control the synchronization state jump of the first die according to the control signal corresponding to the synchronization input data.
[0060] In some specific implementations of the embodiments of the present disclosure, referring to Figure 5 , the front-end unit 111 comprises:
[0061] The finite state machine 1111 comprises a read state (FE_OUT_CMD_READ_ST), a stall state (FE_OUT_CMD_STALL_CHIP_ST) and a pre-stall state (FE_OUT_PRE_STALL_ST), and is configured to generate the pop-up signal according to the control instruction corresponding to the received synchronization signal input data.
[0062] Specifically, the information carried by the synchronization signal tells the target die (the second die) the operation to be performed and the state to be entered.
[0063] The synchronization signal comprises:
[0064] The read synchronization signal is a control instruction for entering the read state from the current state.
[0065] The stall synchronization signal is a control instruction for entering the stall state from the current state. The read state is for generating the pop-up signal and entering the pre-stall state after receiving a valid pre-stall instruction; the stall state is for returning the first die to the idle state after waiting for the response of the synchronization signal; and the pre-stall state is a transition state and will jump to the stall state.
[0066] The first-in-first-out buffer 1112 is configured to store the pop-up signal received from the finite state machine 1111 and provide a valid pre-stall instruction for the finite state machine 1111.
[0067] The control instruction buffer 1113 is configured to store and output the synchronization signal output data and output the control instruction to the first-in-first-out buffer 1112.
[0068] Referring to Figure 6When the control instruction received by the front-end unit 111 is a pre-stall (stall_fe) instruction, the finite state machine 1111 will enter the pre-stall state (FE_OUT_PRE_STALL_ST) from the read state (FE_OUT_CMD_READ_ST), and when the received control instruction is a stall (stall_chip_fe) instruction, it will enter the stall state (FE_OUT_CMD_STALL_CHIP_ST) from the pre-stall state (FE_OUT_PRE_STALL_ST), and record which chip needs to be used to release the stall state of the first chip. After the first chip receives the synchronization signal response (fe_out_stall) of the corresponding chip, it will enter the read state (FE_OUT_CMD_READ_ST) and exit the stall state (FE_OUT_CMD_STALL_CHIP_ST).
[0069] Specifically, referring to Figure 7 The global synchronization unit 112 includes:
[0070] The receive input register (rcv input reg) 1123 is used to receive the synchronization signal input data sent by other chips sent by the synchronization bus unit 113.
[0071] The receive buffer (rcv buffer) 1122 is used to buffer the synchronization signal input data.
[0072] The receive first-in-first-out queue (rcv afifo) 1121 is used to receive the synchronization signal input data and complete the asynchronous-to-synchronous conversion, so that the synchronization input data can be used within the first chip.
[0073] The asynchronous-to-synchronous conversion and the use of the synchronization input data within the chip are processed by using a standard circuit structure.
[0074] The receive first-in-first-out queue 1121 sends the synchronization input data that can be used within the first chip to the finite state machine 1111 of the front-end unit 111 to generate a control instruction corresponding to the synchronization signal, and control the synchronization state jump between multiple chips.
[0075] The send first-in-first-out queue (fe fifo) 1124 is used to receive the synchronization signal output data sent by the finite state machine 1111 of the front-end unit 111 and complete the synchronous-to-asynchronous conversion, so that the synchronization output data can be sent to other chips.
[0076] The synchronous-to-asynchronous conversion and the sending of the synchronization output data to other chips are processed by using a standard circuit structure.
[0077] A sending buffer (fe buffer) 1125 is configured to buffer the synchronization signal output data.
[0078] A sending output register (fe output reg) 1126 is configured to send the synchronization signal output data to the synchronization bus unit 113.
[0079] In some implementations of the embodiments of the present disclosure, the synchronization bus unit 113 is specifically configured to:
[0080] process the synchronization signal output data and transmit the synchronization signal output data to the synchronization transceiver module;
[0081] receive the synchronization signal input data, and if it is determined through comparison of the target identifier that the synchronization signal input data does not correspond to the first die, forward the synchronization signal input data to other dies, and if it is determined that the synchronization signal input data corresponds to the first die, send a control instruction to the front-end module to control the synchronization state of the die to jump.
[0082] Specifically, the target identifier is used to determine whether the first die corresponds to the received synchronization signal input data.
[0083] Referring to Figure 8 , the synchronization bus unit 113 includes:
[0084] A first data synchronization or buffering flip-flop 1131 is configured to receive the synchronization signal output data through the synchronization output bus.
[0085] An arbiter (ARB) 1133 is configured to receive the synchronization signal output data sent by the first data synchronization or buffering flip-flop 1131, send the synchronization signal output data to the synchronization transceiver module 12 through an output buffer (Send Buffer) 1134, and instruct the synchronization transceiver module 12 to send the synchronization signal output data to other dies.
[0086] The arbiter (ARB) 1133 is further configured to determine, according to a target identifier of the received synchronization signal input data, whether the synchronization signal input data corresponds to the first die, and if so, send the synchronization signal input data to the buffer area 1132 for storage.
[0087] A second data synchronization or buffering flip-flop 1135 is configured to send the synchronization signal input data buffered and stored in the buffer area (RVC Buffer) 1132 to the global synchronization unit 112.
[0088] Specifically, the front-end unit 111 of the first die VIP0 generates synchronization signal output data (such as sema0to1), the synchronization bus unit 113 transmits the synchronization signal output data to the synchronization transceiver module 12 through the arbiter 1133 and the output buffer 1134, and finally sends the synchronization signal output data to the second die VIP1.
[0089] The synchronization bus unit 113 of the second die VIP1 receives the synchronization signal input data (sema0to1) of the first die VIP0, confirms that the target identification is the second die, and sends the synchronization signal input data to the front-end unit 111 through the global synchronization unit 112 to trigger state jumping (such as from the pause state to the working state). If the target identification does not match, the second die VIP1 forwards the synchronization signal input data to the third die VIP2.
[0090] The synchronization bus unit 113 of the embodiment of the present disclosure avoids multi-synchronization signal conflicts through the arbitrator 1133, and ensures correct transmission of data according to priority or target.
[0091] The output buffer 1134 serves as a sending buffer and supports temporary storage and forwarding of synchronization signals, especially when cross-die transmission or target die mismatch occurs.
[0092] When the target identification of the synchronization signal input data corresponds to the current die, the control instruction of the synchronization signal input data triggers the state machine of the front-end unit 111 to jump (such as releasing the pause state, see Figure 6 The pause state of the front-end unit 111 jumps to the pre-pause state).
[0093] The synchronization signal output data and the synchronization signal input data both include source identification, target identification, and transmission type, and are handshake through valid signals (vld) and response signals (ack).
[0094] In some specific implementations of the embodiment of the present disclosure, referring to Figure 9 The synchronization transceiver module 12 includes:
[0095] The configuration interface unit 121 is configured to receive synchronization signal input data sent by other dies and configure transmission addresses of the other dies.
[0096] The protocol conversion unit 122 is configured to receive the transmission addresses of the other dies and generate bus signals of an advanced extensible interface bus protocol (AXI) according to the transmission addresses.
[0097] Specifically, the configuration interface unit 121 sends the synchronization signal input data sent by the other dies to the arbitrator 1133 of the synchronization bus unit 113, and the arbitrator 1133 determines whether the synchronization signal input data corresponds to the first die according to the target identification of the received synchronization signal input data.
[0098] The protocol conversion unit 122 sends the bus signals of the advanced extensible interface bus protocol to the other dies.
[0099] The embodiment of the present disclosure realizes data encapsulation and transmission by configuring the interface unit 121 and the protocol conversion unit 122, and guarantees the efficiency and accuracy of data transmission.
[0100] Specifically, referring to Figure 10 , the state machine of the protocol conversion unit 122 includes:
[0101] An idle state (IDLE) is used for suspending the processing operation until a valid signal is output.
[0102] A write driving state (W_DRIVEW) is used for driving the write address and control signal.
[0103] After receiving the valid signal (sema_out_vld), the state machine jumps to the write driving state. In this state, the protocol conversion unit 122 sends the address pre-configured in the register to the advanced extensible interface bus protocol bus.
[0104] A write handshake state (W_HANDS) is used for sending the write data and waiting for the ready signal.
[0105] After the protocol conversion unit 122 detects that the preliminary write waiting (awready) signal is pulled high (i.e., sema_axi_awready), the state machine jumps to the write handshake state. In this state, the protocol conversion unit 122 sends the data to the advanced extensible interface bus protocol bus and waits for the write waiting (wready) signal to be pulled high (i.e., wready_over=wready).
[0106] A write waiting state (W_WAIT) is used for waiting for the write response channel to be prepared.
[0107] After waiting for the write waiting (wready) signal to be pulled high (i.e., wready_over=wready), the state machine enters the write waiting state. This state is used for waiting for the write response (B channel).
[0108] A write end state (W_END) is used for waiting for the write response valid signal or the bypass signal to complete the transmission.
[0109] The write waiting state directly jumps to the write end state. In the write end state, the protocol conversion unit 122 waits for the write operation completion (bvalid) signal of the B channel to be pulled high, or the register configuration signal (reg_axi_bchannel_bypass) to be pulled high. This indicates that the protocol conversion unit 122 is waiting for the completion confirmation of the write operation, or skipping the response waiting process through a certain mechanism.
[0110] AWREADY: This is a signal in the AXI bus protocol, indicating that the master can send a write address.
[0111] WREADY: This is a signal in the AXI bus protocol, indicating that the master can send a write data.
[0112] BVALID: This is a signal in the AXI bus protocol, indicating that the write operation has been completed, and the slave is ready to send a write response.
[0113] REG_AXI_BCHANNEL_BYPASS: This is a register signal for controlling whether to skip waiting for a B-channel response.
[0114] Figure 11 For a complete multi-die synchronization process, the multi-dies are respectively a first die VIP0, a second die VIP1, a third die VIP(n-1), and a fourth die VIP(n).
[0115] When the first die VIP0, the second die VIP1, the third die VIP(n-1), and the fourth die VIP(n) obtain a control instruction, the front-end unit 111 of the first die VIP0 receives the control instruction fe_pe_sema, and the first die VIP0 enters a stall state and sends a first synchronization signal to the dies that need to be synchronized, while recording the transmission target identifier. When the first die VIP0 receives the synchronization signal sema1to0, the first die VIP0 enters a work state.
[0116] For the second die VIP1, according to the instruction, the second die VIP1 can enter a stall state, but will not output synchronization signal output data. When receiving the synchronization signal sema0to1 sent by the first die VIP0, the state machine of the second die VIP1 re-flushes to the next stall state, and sends a synchronization signal sema1to2 to the third die VIP(n-1) in this stage. When receiving the synchronization signal sema2to1 sent by the third die VIP(n-1), the state jumps to a work state, and sends a synchronization signal sema1to0 to the first die VIP0.
[0117] Referring to Figure 12 The embodiments of the present disclosure also provide a die, and the first die applies any one of the above hardware accelerators.
[0118] Referring to Figure 13The embodiment of the present disclosure also provides a multi-die cooperative system, comprising a first die and a second die, and the first die and the second die both run the hardware accelerator of any one of the above.
[0119] Referring to Figure 14 The embodiment of the present disclosure also provides a hardware acceleration method applied to a first die in a multi-die cooperative system, which comprises the following steps:
[0120] Step S1: generating and processing synchronization signal output data, the synchronization signal being used for controlling the synchronization state transition between the dies.
[0121] Step S2: receiving the synchronization signal output data and sending the received synchronization signal input data sent by other dies to the synchronization control module.
[0122] According to the hardware accelerator scheme provided by the embodiment of the present disclosure, the hardware accelerator is applied to a first die in a multi-die cooperative system, and the synchronization control module of the hardware accelerator generates and processes synchronization signal output data, the synchronization signal being used for controlling the synchronization state transition between the dies. The synchronization control module sends the synchronization signal output data to the synchronization transceiver module, and the synchronization transceiver module sends the received synchronization signal input data sent by other dies to the synchronization control module. The embodiment of the present disclosure realizes a neural network computing platform with greater computing power and excellent cooperative computing capability through the multi-die (DIE) technology without changing the chip architecture.
[0123] Referring to Figure 15 The embodiment of the present disclosure also provides a hardware acceleration method applied to a multi-die cooperative system comprising a first die and a second die, which comprises the following steps:
[0124] Step T1: the first die generates synchronization signal output data, enters a pause state, records the second die identifier, and sends the synchronization signal output data to the second die.
[0125] Step T2: the second die returns response synchronization signal input data after receiving the synchronization signal output data.
[0126] Step T3: the first die jumps to a working state according to the response synchronization signal input data.
[0127] According to the hardware accelerator scheme provided by the embodiment of the present disclosure, the synchronization control module of the hardware accelerator generates and processes synchronization signal output data, and the synchronization signal is used to control the synchronization state transition between the multiple dies. The synchronization control module sends the synchronization signal output data to the synchronization transceiver module, and the synchronization transceiver module sends the received synchronization signal input data sent by other dies to the synchronization control module. The embodiment of the present disclosure realizes a neural network computing platform with greater computing power and excellent collaborative computing capability through the multi-die (DIE) technology without changing the chip architecture.
[0128] It should be understood that each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment mainly describes the difference from other embodiments. Especially, for the method embodiments, since the method described is basically similar to the method described in the device and system embodiments, the description is relatively simple, and the relevant parts can refer to the part of the description of other embodiments.
[0129] It should be understood that the above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in which the embodiments are described, and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.
[0130] It should be understood that the elements described herein in singular form or shown in the figures only one do not represent the number of the elements limited to one. In addition, the modules or elements described or shown as separate in this paper can be combined into a single module or element, and the modules or elements described or shown as single in this paper can be split into multiple modules or elements.
[0131] It should also be understood that the terms and expressions used herein are only used for description, and one or more embodiments of the present disclosure should not be limited to these terms and expressions. The use of these terms and expressions does not mean that any illustration and description (or part thereof) is excluded, and it should be recognized that various modifications that can exist should be included in the scope of the claims. Other modifications, changes and replacements can also exist. Accordingly, the claims should be considered to cover all these equivalents.
Claims
1. A hardware accelerator, wherein, The hardware accelerator is applied to a first die in a multi-die cooperative system, and the hardware accelerator comprises: a synchronization control module, configured to generate and process synchronization signal output data, the synchronization signal being used to control synchronization state transition between the dies; a synchronization transceiver module, configured to receive the synchronization signal output data sent by the synchronization control module, and send synchronization signal input data sent by other dies to the synchronization control module; the synchronization control module comprises: a synchronization bus unit, configured to process the synchronization signal output data, and receive the synchronization signal input data, and forward the synchronization signal input data to other dies if it is determined that the synchronization signal input data does not correspond to the first die by comparing a target identifier.
2. The hardware accelerator of claim 1, wherein, The synchronization control module sends the synchronization signal output data to the synchronization transceiver module through a synchronization output bus, and the synchronization transceiver module receives synchronization signal input data sent by other dies through a synchronization input bus, the synchronization signal output data and the synchronization signal input data both comprise a source identifier, a target identifier and a transmission type, and handshake is performed through an effective signal and a response signal.
3. The hardware accelerator of claim 2, wherein, The synchronization control module further comprises: a front-end unit, configured to generate the synchronization signal output data, and control synchronization state transition of the first die according to a control instruction corresponding to the synchronization signal input data received; a global synchronization unit, configured to synchronize and clock the synchronization signal input data and the synchronization signal output data.
4. The hardware accelerator of claim 3, wherein, The front-end unit comprises: a finite state machine, comprising a reading state, a pause state and a pre-pause state, configured to generate a pop-up signal according to a control instruction corresponding to the synchronization signal input data received; a first-in-first-out buffer, configured to store the pop-up signal received from the finite state machine, and provide an effective pre-pause instruction for the finite state machine; a control instruction buffer, configured to store and output the synchronization signal output data, and output a control instruction to the first-in-first-out buffer; wherein the reading state is entered into the pre-pause state after the pop-up signal is generated and an effective pre-pause instruction is received; the pause state is used to return the first die to the reading state after the synchronization signal response is waited for; and the pre-pause state is a transition state, and is jumped to the pause state.
5. The hardware accelerator of claim 4, wherein, The synchronization bus unit is further configured to: process the synchronization signal output data, and transmit the synchronization signal output data to the synchronization transceiver module; receive the synchronization signal input data, and send a control instruction to the front-end unit to control synchronization state transition of the die if it is determined that the synchronization signal input data corresponds to the first die by comparing a target identifier.
6. The hardware accelerator of claim 5, wherein, The synchronization transceiver module comprises: a configuration interface unit, configured to receive the synchronization signal input data sent by the other dies, and configure a transmission address of the other dies; a protocol conversion unit, configured to receive the transmission address of the other dies, and generate a bus signal of an advanced extensible interface bus protocol according to the transmission address.
7. The hardware accelerator of claim 6, wherein, The state machine of the protocol conversion unit comprises: a write driving state, used to drive a write address and a control signal; Write handshake state, for sending write data and waiting for ready signal; Write waiting state, for waiting for write response channel preparation; Write end state, for waiting for write response valid signal or bypass signal to complete transmission.
8. A die, wherein, The die applies the hardware accelerator of any one of claims 1-7.
9. A multi-die co-system, wherein, The multi-die cooperative system includes a first die and a second die, and the hardware accelerator of any one of claims 1-7 runs on the first die and the second die.
10. A hardware acceleration method applied to a first die in a multi-die co-system, wherein, The method comprises: Generating and processing synchronization signal output data, the synchronization signal being used to control synchronization state jump between multiple dies; Receiving the synchronization signal output data and sending the received synchronization signal input data sent by other dies to a synchronization control module; The generating and processing synchronization signal output data, the synchronization signal being used to control synchronization state jump between multiple dies, comprises: Processing the synchronization signal output data and receiving the synchronization signal input data, and if it is determined through comparing target identification that the first die does not correspond to the synchronization signal input data, forwarding the synchronization signal input data to other dies.
11. A hardware acceleration method, wherein, The method is applied to a multi-die cooperative system including a first die and a second die, and the method comprises: The first die generates synchronization signal output data, enters a pause state and records a second die identification, and sends the synchronization signal output data to the second die; After the second die receives the synchronization signal output data, the second die returns response synchronization signal input data; The first die jumps to a working state according to the response synchronization signal input data.
Citation Information
Patent Citations
Multi-core synchronous management system applied to neural network accelerator
CN117892781A