System and chip for supporting near memory operation based on Internet on chip
By designing a system based on the on-chip Internet network that supports near-memory computing in the scenario of multiple execution units and multiple on-chip memory, the network blockage problem caused by many data requests is solved, and the performance of the neural network processing unit is optimized.
Patent Information
- Application Number
- CN202510038151.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
AI Technical Summary
In the scenario of multiple execution units and multiple on-chip memory, the on-chip Internet network is prone to network blockage due to the large number of data requests, which in turn leads to degradation in the performance of the neural network processing unit.
A system that supports near-memory computing based on the on-chip Internet network is designed, including a write receiving module, a first read command generation module, a write information cache module and an operation module. By performing operations in the near-memory, the occupation of on-chip Internet network is reduced.
By performing operations in near memory, the number of accesses to the on-chip Internet network is reduced, the risk of network blockage is reduced, and the performance of neural network processing units is optimized.
Smart Images

Figure CN119961207A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of chip technology, and in particular to a system and chip supporting near-memory computing based on an on-chip interconnect network. Background Art
[0002] The system design of neural-network processing unit (NPU) generally adopts network-on-chip (NoC) to realize the communication between execution unit (EU) and on-chip memory (OCM).
[0003] The execution unit can read data from the on-chip memory through the on-chip interconnect network, perform the target operation inside the execution unit, and write the target operation result to the on-chip memory through the on-chip interconnect network after obtaining the target operation result. However, this method of completing the target operation inside the execution unit requires the execution unit to access the on-chip memory through the on-chip interconnect network many times during the operation process, and the data path is long, resulting in a decrease in the performance of the neural network processing unit.
[0004] In addition, in the scenario where there are multiple execution units and multiple on-chip memories in the on-chip interconnect network system, the on-chip interconnect network is more likely to be congested due to multiple data requests issued by the execution units, thereby causing the performance of the neural network processing unit to decline. Summary of the invention
[0005] The present application provides a system and chip that supports near-memory computing based on an on-chip interconnect network to solve the problem that in a scenario with multiple execution units and multiple on-chip memories, the on-chip interconnect network is prone to network congestion due to the large number of data requests in the on-chip interconnect network, thereby causing performance degradation.
[0006] In a first aspect, the present application provides a system for supporting near memory operations based on an on-chip interconnect network, comprising: a write receiving module, a first read command generating module, a write information caching module, and an operation module;
[0007] The write receiving module is configured to: in response to a received first near memory write command, obtain the storage status of the first read command generation module, the write information cache module and the write feedback module; the first near memory write command is a command sent by the master device and the slave device when performing a near memory operation based on the AIX bus; the near memory write command includes a near memory write address and original data;
[0008] If the storage status of the first read command generation module and the write information cache module is not full, the near memory write address and the original data are sent to the write information cache module; and the near memory write address is sent to the first read command generation module;
[0009] The first read command generation module is configured to: generate a first read command in response to a received near memory write address;
[0010] Sending the first read command to the storage module, so that the storage module sends the near memory read data corresponding to the first read command in the target cache to the operation module; and sending the first read command to the write information cache module, so that the write information cache module outputs the original data and the near memory write address to the operation module;
[0011] The operation module is configured to: in response to the received near memory read data and the original data, perform a target operation on the near memory read data and the original data to obtain target data;
[0012] The target data is written into the storage module according to the near memory write address.
[0013] In some feasible embodiments, the near memory write command further includes a write operation identifier; and the write receiving module is further configured to: send the write operation identifier to the write information cache module.
[0014] In some feasible embodiments, the near memory computing system further includes a first flow distribution module;
[0015] The write receiving module executes, in response to the received near memory write command, acquiring the storage status of the first read command generating module, the write information caching module and the write feedback module, and is further configured to: send the first near memory write command to the first diversion module;
[0016] The first shunting module is configured to: extract the near memory write address and original data from the first near memory write command;
[0017] The near memory write address is stored in the first read command generation module, and the near memory write address and the original data are stored in the write information cache module.
[0018] In some feasible embodiments, the near memory computing system further includes a write arbitration module;
[0019] The operation module executes writing the target data into the target cache, and is specifically configured to: send the second near memory write command to the write arbitration module; the second near memory write command includes a near memory write identifier, the near memory write address, the target data and a near memory write ID;
[0020] The write arbitration module is configured to, when detecting a near memory write identifier in the second near memory write command, preferentially write the target data into the storage module based on the near memory write address.
[0021] In some feasible embodiments, the near memory computing system further includes a write feedback module;
[0022] The write arbitration module is further configured to: send the near memory write id to the write feedback module;
[0023] The write feedback module is configured to: feed back the near memory write ID to the master device.
[0024] In some feasible embodiments, the near memory computing system further includes a first beat module;
[0025] The first read command generation module executes sending the first read command to the write information cache module, and is further configured to: send the first read command to the write information cache module through the first beat module;
[0026] The first beat module is configured to: in response to the received first read command, beat according to a first preset beat number and then forward the first read command to the write information cache module;
[0027] The write information cache module is configured to send the original data to the operation module in response to the first read command.
[0028] In some feasible embodiments, the near memory computing system further includes a second flow distribution module;
[0029] The storage module is configured to: when sending the near memory read data to the operation module, send the near memory read data to the second diversion module;
[0030] The first beat module is further configured to: forward the first read command to the second shunt module;
[0031] The second traffic distribution module is configured to send the received near memory read data to the operation module in response to the first read command.
[0032] In some feasible embodiments, the write receiving module is further configured to: receive a non-near memory write command, and obtain a write operation identifier of the second write command;
[0033] If the write operation identifier is used to indicate that the current write command is a non-near memory write command, the non-near memory write command is sent to the first diversion module; the non-near memory write command includes a non-near memory write address, non-near memory write data, and a non-near memory write ID;
[0034] The first traffic distribution module is configured to: forward the non-near memory write command to the write arbitration module;
[0035] The write arbitration module is configured to: in response to the received non-near memory write command, write the non-near memory write data to the storage module according to the non-near memory write address; and send the non-near memory write ID to the write feedback module so that the write feedback module returns the non-near memory write ID to the master device.
[0036] In some feasible embodiments, the near memory computing system further includes: a read receiving module, a second read command generating module, a read ID cache module, and a read feedback module;
[0037] The read receiving module is configured to: receive a first non-near memory read command, wherein the first non-near memory read command includes a non-near memory read address and a non-near memory read id;
[0038] Obtaining the storage status of the second read command generation module, the read ID cache module, and the read feedback module;
[0039] If the storage status of the second read command generation module, the read ID cache module and the read feedback module is not full, the non-near memory read address is sent to the second read command generation module, and the non-near memory read ID is sent to the read cache module;
[0040] The second read command generation module is configured to: generate a second non-near memory read command in response to the received non-near memory read address;
[0041] Send the second non-near memory read command to the storage module so that the storage module sends the non-near memory read data corresponding to the non-near memory read address to the read feedback module; and send the second non-near memory read command to the read ID cache module so that the read ID cache module sends the non-near memory read ID to the read feedback module.
[0042] In some feasible embodiments, the near memory computing system further includes a second beat module;
[0043] The second read command generation module executes sending the second non-near memory read command to the read ID cache module, and is further configured to: send the second non-near memory read command to the second beat module;
[0044] The second beat module is configured to: in response to the received second non-near memory read command, send the second non-near memory read command to the read ID cache module after beating according to a second preset beat number;
[0045] The read id cache module is configured to: in response to receiving the second non-near memory read command, send the non-near memory read id to the read feedback module.
[0046] In some feasible embodiments, the second beat module is further configured to: in response to the received second non-near memory read command, after beating according to a second preset beat number, send the second non-near memory read command to the second diversion module;
[0047] The second traffic distribution module is configured to: receive the non-near memory read data fed back by the storage module;
[0048] When the second non-near memory read command is received, the non-near memory read data is sent to the read feedback module.
[0049] In some feasible embodiments, the second beat module is further configured to: send a second non-near memory read command to the read feedback module;
[0050] The read feedback module is configured to: generate a non-near memory read feedback instruction based on the non-near memory read data and the non-near memory read id;
[0051] The non-near memory read feedback instruction is sent to the master device.
[0052] In some feasible embodiments, the near memory computing system further includes a read arbitration module;
[0053] The read arbitration module is configured to, upon receiving at least two read instructions, forward the read instructions to the storage module in turn, so that the preceding read instruction and the succeeding read instruction received by the storage module correspond to different data transmission transactions.
[0054] In a second aspect, the present application provides a chip, comprising a master device, a slave device, and the near memory computing system described in the first aspect. The master device is configured to: send near memory data processing instructions and / or non-near memory data processing instructions to the system supporting near memory computing based on an on-chip interconnect network, so as to perform a data interaction process with the slave device based on the system supporting near memory computing based on the on-chip interconnect network.
[0055] It can be seen from the above technical content that the present application provides a system and chip that supports near memory operations based on an on-chip interconnect network. The system includes a write receiving module, a first read command generating module, a write information cache module and an operation module. The first near memory write command is received by the write receiving module, the near memory write address is stored by the first read command generating module, and the near memory write address and the original data are stored by the write information cache module. The original data sent by the write information cache module is received by the operation module, as well as the near memory read data fed back by the storage module, and the target operation is performed. The target data obtained by the target operation is then written into the storage module. The system can perform the target operation on the side adjacent to the storage module to reduce the occupancy of the on-chip interconnect network during the operation process and optimize the performance of the neural network processing unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0057] Figure 1 A schematic diagram of the system architecture of a neural network processing unit provided in an embodiment of the present application;
[0058] Figure 2 A schematic diagram of modules of a system supporting near-memory computing based on an on-chip interconnect network provided in an embodiment of the present application;
[0059] Figure 3 An operating timing diagram of the write receiving module provided in an embodiment of the present application when receiving a first near memory write command;
[0060] Figure 4 A timing diagram of the operation of the first shunt module provided in an embodiment of the present application;
[0061] Figure 5 The operation sequence diagram of the write arbitration module provided in the embodiment of the present application;
[0062] Figure 6 A timing diagram of the operation of the first beat module provided in an embodiment of the present application;
[0063] Figure 7 A timing diagram of the operation of the second shunt module provided in an embodiment of the present application;
[0064] Figure 8 A timing diagram of the operation of the write receiving module provided in an embodiment of the present application when receiving a non-near memory write command;
[0065] Fig. 9 A timing diagram of the operation of the read receiving module provided in an embodiment of the present application when receiving a non-near memory read command;
[0066] Fig.10 A timing diagram of the operation of the second beat module provided in an embodiment of the present application;
[0067] Fig.11 This is a timing diagram of the operation of the polling arbitration module provided in an embodiment of the present application. DETAILED DESCRIPTION
[0068] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.
[0069] The neural network processing unit includes an execution unit, an on-chip memory, and an on-chip interconnect network. The execution unit can communicate with the on-chip memory through the on-chip interconnect network. The execution unit can access data from the on-chip memory through the on-chip interconnect network, and can also write data to the on-chip memory through the on-chip interconnect network.
[0070] For example, when the execution unit performs an accumulation operation, it needs to read data from the on-chip memory through the on-chip interconnect network and perform the accumulation operation inside the execution unit. After obtaining the target data through the accumulation operation, the execution unit writes the target data into the on-chip memory through the on-chip interconnect network. It can be understood that after the execution unit reads the data from the on-chip memory, the space in the on-chip memory used to store the data can be released, and then the execution unit can write the target data into the released space. In this way, the data read by the execution unit can be the same as the storage address of the target data in the on-chip memory.
[0071] However, in the process of the execution unit performing the accumulation operation, the execution unit needs to frequently access the on-chip interconnect network. When the data path of the on-chip interconnect network is long, the system delay of the neural network processing unit will increase. In addition, the frequent access of the execution unit to the on-chip interconnect network will cause congestion of the on-chip interconnect network and occupy the bandwidth resources of the on-chip interconnect network, resulting in low bandwidth utilization of the on-chip interconnect network, thereby reducing the performance of the neural network processing unit.
[0072] In order to solve the above problems, Figure 1As shown, some embodiments of the present application provide a system architecture of a neural network processing unit. Among them, the on-chip memory is divided into equal parts, so that each execution unit can have a corresponding on-chip memory, so that the data traffic can be transmitted in a dispersed manner, which is conducive to reducing the congestion risk of a single path or area in the on-chip interconnect network. In addition, different execution units can access the on-chip memory through different access paths in the on-chip interconnect network, which can also alleviate the problems of data conflict and access delay.
[0073] exist Figure 1 In the system architecture of the neural network processing unit shown in FIG. 1 , the execution unit can perform general read and write access operations to the on-chip memory through the on-chip interconnect network. Figure 2 The near memory computing system shown in the figure moves the computing logic inside the execution unit to the near end of the on-chip memory to implement near memory computing (NMC). This is to alleviate the problem that when the neural network processing unit performs the target operation, the execution unit frequently accesses the on-chip memory, resulting in the performance degradation of the neural network processing unit. Figure 2 The figure is only used to illustrate the internal structure of a near memory computing system. Figure 1 In the on-chip interconnect network shown, a near memory operation system can be provided between the slave device side of the on-chip interconnect network and the on-chip memory, so that the on-chip interconnect network can receive a near memory / non-near memory command issued by the execution unit based on Figure 2 The system shown supports near memory operations and non-near memory reads and writes to perform target tasks.
[0074] like Figure 2 and Figure 3 As shown, some embodiments of the present application provide a system supporting near memory operations based on an on-chip interconnect network, including a write receiving module (wemd_pipe_fifo), a first read command generation module (nmc_rd_fifo), a write information cache module (nmc_wr_fifo); a write feedback module (bch_fifo) and an operation module.
[0075] The write receiving module is configured to: in response to a received first near memory write command, obtain the storage status of the first read command generation module, the write information cache module and the write feedback module. The first near memory write command is a command sent by the master device and the slave device when performing a near memory operation based on the AXI bus; the near memory write command includes a near memory write address and original data;
[0076] If the storage status of the first read command generation module and the write information cache module is not full, the near memory write address and the original data are sent to the write information cache module; and the near memory write address is sent to the first read command generation module.
[0077] It can be understood that the first near memory write command is a command sent by the master device and the slave device when performing near memory operations based on the AXI bus, and the first near memory write command may include a near memory write address and original data. Among them, based on the near memory write address in the first near memory write command, the slave device side of the on-chip interconnection network system can read the near memory read data from the on-chip memory, and perform the target operation with the original data. After obtaining the target data, the target data is written to the storage module according to the near memory write address to realize the near memory operation. Among them, the original data refers to the data sent by the master device to the slave device through the on-chip interconnection network for the near memory write operation. This part of the data arrives at the near memory operation system with the on-chip interconnection network to participate in the near memory operation.
[0078] The storage module may be an on-chip memory group, and Figure 1 As shown, the number of on-chip memory groups can be the same as the number of execution units, and the execution units can access the on-chip memory groups based on the on-chip interconnect network. The on-chip memory groups can be SRAMs (Static Random-Access Memory) pre-divided into multiple equal parts. And, as Figure 2 As shown, a system supporting near-memory operations may be provided between the slave device side of the on-chip interconnect network and the on-chip memory, so that the execution unit can access the on-chip memory through the on-chip interconnect network.
[0079] In some embodiments, the master device may refer to the slave device side of the on-chip interconnect network. After receiving the near memory operation instruction initiated by the execution unit on the other side of the on-chip interconnect network, the slave device side may issue a request to the on-chip memory according to the near memory operation instruction. Furthermore, the slave device side of the on-chip interconnect network is equivalent to the master device in the near memory operation process, and the on-chip memory being accessed is equivalent to the slave device.
[0080] It should be noted that the first near memory write command may be a write command sent by the master device to the write receiving module. The first near memory write command includes a write operation identifier, a near memory write address, a near memory write ID, a write strobe signal and original data.
[0081] In addition, in the embodiment of the present application, it also includes the necessary identifiers for executing data interaction based on the axi protocol, such as last (the last data beat identifier in the burst transmission) and rresp (read response identifier). This part of the identifier can be used to maintain the read and write access method based on the axi protocol between the master device and the slave device. However, the focus of the embodiment of the present application is to improve the system that supports near-memory operations to alleviate the problem that the on-chip interconnect network is easily congested due to many data requests, and does not affect the normal use of commonly used identifiers such as last and rresp, so the uses of commonly used identifiers are not explained one by one.
[0082] In some embodiments, the near memory write command also includes a write operation identifier. The write receiving module is further configured to: send the write operation identifier to the write cache module. In this way, the write cache module can store a write operation identifier for indicating the target operation type, and when receiving the first read command sent by the first read command generation module, the write operation identifier is sent to the operation module to instruct the operation module to perform the target operation according to the write operation identifier.
[0083] In some embodiments, the write operation identifier can be recorded as awuser (custom field), and the write operation identifier is used to indicate the data operation type used when executing the target operation. For example, awuser=1 indicates that a signed 32-bit integer operation (s32bit) is used; awuser=2 indicates that a signed 16-bit integer operation (s16bit) is used; awuser=3 indicates that an unsigned 16-bit integer (u16) is used; awuser=4 indicates that a 16-bit floating point number (fp16) is used; user=5 indicates that a 32-bit floating point number (fp32) is used; awuser=6 indicates that a 16-bit floating point number (bf16) is used.
[0084] In addition, the near memory operation method provided in some embodiments is not limited to the several operations mentioned above, and different write operation identifiers can be set according to actual operation requirements to control the operation module to perform the operation specified by the write operation identifier through the write operation identifier.
[0085] The near memory write ID is used to distinguish multiple write commands issued by the execution unit. The write select signal is used to indicate the valid part of the data being transmitted, so that the near memory write command can be converted into a signal that the SRAM can recognize according to the write select signal.
[0086] In some embodiments, the master device sends a first near memory write command to a system that supports near memory operations. When the write receiving module receives the first near memory write command, it obtains the storage status of the first read command generation module and the write information cache module. When the storage status of these modules is not full, it means that there is enough space to receive data and perform near memory operations. And the write receiving module can send the near memory write address in the first near memory write command to the write information cache module, and send the near memory write address to the first read command generation module.
[0087] It should be noted that when the first near memory write command is sent, a write operation identifier is also sent simultaneously with the first near memory write command, and in the near memory scenario, the write operation identifier awuser≠0. In some embodiments, the near memory operation can continue to be executed only when awuser≠0, and the address that is the same as the near memory write address of the current first near memory write command does not appear in the current on-chip interconnection network system, and the storage status of the first read command generation module and the write information cache module is not full.
[0088] In some embodiments, the first read command generation module is used to generate a first read command to read near memory read data from the SRAM. The first read command generation module is configured to: generate the first read command in response to a received near memory write address.
[0089] The first read command is sent to the storage module so that the storage module sends the near memory read data corresponding to the first read command in the target cache to the operation module; and the first read command is sent to the write information cache module so that the write information cache module outputs the original data and the near memory write address to the operation module.
[0090] In some embodiments, when receiving the near memory write address, the first read command generation module can generate a first read command according to the near memory write address and the near memory write ID corresponding to the near memory write address, and send the first read command to the SRAM to read the near memory read data in the SRAM, and the SRAM sends the near memory read data to the operation module. And send the first read command to the write information cache module, so that the write information cache module sends the original data to the operation module.
[0091] It is understandable that the write information cache module can send the original data to the operation module when receiving the first read command. In this way, the operation module can perform operations on the original data and the near memory read data, without having to return the near memory read data to the main device before performing the operation, thereby saving data transmission distance and improving data transmission efficiency.
[0092] The operation module is configured to: in response to the received near memory read data and original data, perform a target operation on the near memory read data and original data to obtain target data, and write the target data into the storage module according to the near memory write address.
[0093] In some embodiments, when the operation module receives the near memory read data and the original data, it can perform the target operation on the near memory read data and the original data according to the operation mode corresponding to the write operation identifier. Taking the target operation as an accumulation operation as an example, there is no need to perform the accumulation operation inside the main device and then write the accumulation result to the SRAM. The target data can be obtained on the side close to the SRAM, thereby reducing the number of times the main device accesses the on-chip interconnect network, thereby optimizing the performance of the main device that processes data through the on-chip interconnect network during operation, such as a neural network processing unit.
[0094] like Figure 2 and Figure 4 As shown, the system supporting near memory operation also includes a first diversion module. The write receiving module executes in response to the received near memory write command, obtains the storage status of the first read command generation module, the write information cache module and the write feedback module, and is also configured to: send the first near memory write command to the first diversion module;
[0095] The first diversion module is configured to extract the near memory write address and original data from the first near memory write command.
[0096] The near memory write address is stored in the first read command generation module, and the near memory write address and the original data are stored in the write information cache module.
[0097] In some embodiments, the first shunt module is used to extract various parts of information in the first near memory write command. After receiving the first near memory write command, the write receiving module can send the first near memory write command to the first shunt module, and the first shunt module extracts the near memory write address, original data, near memory write ID and other information from the first near memory write command. Furthermore, the first shunt module can send the near memory write address to the first read command generation module for storage, and send the original data and the near memory write address to the write information cache module for storage.
[0098] In this way, after receiving the write address, the storage state of the first read command generation module is a non-empty state, and the first read command generation module can generate a first read command according to the near memory write address. It can be understood that when the storage module is SRAM, the first read command generation module can convert the first read command in combination with the write selection signal in the first near memory write command to convert the axi access signal into an SRAM access signal, thereby realizing interface adaptation.
[0099] like Figure 2 and Figure 5As shown, the system supporting near memory operation also includes a write arbitration module. The operation module executes writing the target data into the storage module, and is specifically configured to: send the second near memory write command to the write arbitration module; the second near memory write command includes a near memory write identifier, the near memory write address, the target data, and a near memory write ID.
[0100] The write arbitration module is configured to, when detecting a near memory write identifier in the second near memory write command, preferentially write the target data into the storage module based on the near memory write address.
[0101] In some embodiments, after the operation module performs the target operation based on the original data and the near memory read data to obtain the target data, the target data can be sent to the write arbitration module. In this way, the write arbitration module can receive the target data and write the target data into the storage module according to the near memory write address.
[0102] It is understandable that after obtaining the target data, the operation module can generate a second near memory write command according to the target data, the near memory write identifier, the near memory write address, and the near memory write ID, and send the second near memory write command to the write arbitration module.
[0103] In this way, when the write arbitration module detects the near memory identifier in the second near memory write command, it can set a write priority for the target data so that the write priority of the target data obtained based on the near memory operation system is the highest, and then the target data is written to the target storage module first.
[0104] It should be noted that when the write data module writes the target data into the destination storage module, it can also send a write id to the write feedback module so that the write feedback module saves the near memory write id (awid) of the first near memory write command and feeds back the write feedback id (bid) to the master device so that the master device knows that the first near memory write command has been executed.
[0105] In some embodiments, the on-chip interconnect network system also includes a write feedback module, which can receive the near memory write ID sent by the write arbitration module and feedback the near memory write ID to the master device so that the master device knows that the first near memory write command has been executed.
[0106] like Figure 2 and Figure 6 As shown, the system supporting near memory operation also includes a first beat module. The first read command generation module executes sending the first read command to the write information cache module, and is further configured to: send the first read command to the write information cache module through the first beat module.
[0107] The first beat module is configured to: in response to the received first read command, beat according to a first preset beat number and then forward the first read command to the write information cache module.
[0108] The write information cache module is configured to send the original data to the operation module in response to the first read command.
[0109] In some embodiments, the first beat module can be used to adjust the timing of the first read command, so that each module in the near memory computing system has enough time to process the current data, thereby avoiding conflicts between the modules when processing data. For example, the first preset beat number can be set to 4 beats, so that the SRAM has enough time to find the near memory read data according to the near memory write address and send the near memory read data to the computing module.
[0110] It should be noted that the structure of SRAM in actual application scenarios is complex and can be a multi-level, multi-part structure. Therefore, SRAM needs response time when receiving the first read command, and the first beat module provides response time for SRAM through beats, thereby improving the overall timing of the near memory computing system. Among them, the first preset beat number can be adjusted according to actual needs.
[0111] In some embodiments, the first read command generation module may send the first read command to the first beat module, so that the first beat module beats the first read command. After beating, the first beat module may send the first read command to the write information cache module, so that the write information cache module responds to the first read command and sends the original data to the operation module.
[0112] In this way, based on the adjustment of the first beat module, the response time can be provided for the SRAM, thereby improving the overall timing of the near memory operation system when executing the near memory write command.
[0113] like Figure 2 and Figure 7 As shown, the system supporting near memory operation also includes a second shunt module. The storage module is further configured to: when sending near memory read data to the operation module, send the near memory read data to the second shunt module. The first beat module is further configured to: send the first read command to the second shunt module. The second shunt module is configured to: in response to the first read command, send the received near memory read data to the operation module.
[0114] In some embodiments, after receiving the first read command, the storage module sends the near memory read data to the second shunt module. In this way, when the first beat module sends the first beat read command to the second shunt module, the second shunt module can send the near memory read data to the operation module, so that the operation module performs the target operation. In addition, the second shunt module can send the first read command to the write information cache module, so that the write information cache module sends the original data to the operation module, so that the operation module performs the target operation.
[0115] By setting up the second diversion module, the execution timing of each module of the near memory computing system can be adjusted, and the data fed back by the storage module can be sent to the appropriate module in different scenarios to cooperate with the overall operation of the near memory computing system.
[0116] like Figure 2 and Figure 8 As shown, the system supporting near memory operations provided by some embodiments of the present application can also be used to process non-near memory write commands issued by the execution unit. For example, the master device can write the data in the non-near memory write command to the storage module by issuing the non-near memory write command. The write receiving module is further configured to: receive the non-near memory write command, and obtain the write operation identifier of the second write command.
[0117] If the write operation identifier is used to characterize that the current write command is a non-near memory write command, the non-near memory write command is sent to the first diversion module; the non-near memory write command includes a non-near memory write address, non-near memory write data and a non-near memory write ID.
[0118] The first diversion module is configured to: forward the non-near memory write command to the write arbitration module;
[0119] The write arbitration module is configured to: in response to the received non-near memory write command, write the non-near memory write data to the storage module according to the non-near memory write address; and send the non-near memory write ID to the write feedback module so that the write feedback module returns the non-near memory write ID to the master device.
[0120] It is understandable that the non-near memory write command is different from the near memory write command. The non-near memory write command refers to the host device writing the write data directly to the storage module. If the specified location in the storage module has data pre-stored, after executing the non-near memory write command, the pre-stored data will be overwritten by the write data without involving the calculation process.
[0121] It is understandable that when the write command sent by the master device is a non-near memory write command, the write operation identifier awuser in the write channel is 0. In this way, the write receiving module can distinguish the type of the write command through the write operation identifier, and then when the write command is a non-near memory write command, forward the non-near memory write command to the first shunt module, and the first shunt module parses and forwards the non-near memory write command.
[0122] For example, the first traffic distribution module may further forward the non-near memory write command to the write arbitration module. In this way, the write arbitration module may write the non-near memory write data into the storage module according to the non-near memory write command.
[0123] It should be noted that when writing non-near memory write data to a storage module, signal conversion needs to be performed according to the type of the storage module. For example, when the storage module is an SRAM, the axi signal needs to be converted into an sram signal to adapt to the interface of the SRAM.
[0124] It can be understood that when the write arbitration module writes non-near memory write data to the storage module, it can also send a non-near memory write ID to the write feedback module, so that the write feedback module can feedback the non-near memory write ID to the master device, so that the master device can obtain the non-near memory write command execution completion.
[0125] like Figure 2 and Fig. 9 As shown, the system supporting near memory operation can support general read access of the master device. The near memory operation system also includes a read receiving module (arcmd_pipe_fifo), a second read command generating module (normal_rd_fifo), a read id cache module (arid_fifo) and a read feedback module (rch_fifo).
[0126] The read receiving module is configured to: receive a first non-near memory read command, wherein the first non-near memory read command includes a non-near memory read address and a non-near memory read ID.
[0127] Obtain the storage status of the second read command generation module, the read ID cache module and the read feedback module.
[0128] If the storage status of the second read command generation module, the read ID cache module and the read feedback module is not full, the non-near memory read address is sent to the second read command generation module, and the non-near memory read ID is sent to the read cache module.
[0129] The second read command generation module is configured to generate a second non-near memory read command in response to the received non-near memory read address.
[0130] Send the second non-near memory read command to the storage module so that the storage module sends the non-near memory read data corresponding to the non-near memory read address to the read feedback module; and send the second non-near memory read command to the read ID cache module so that the read ID cache module sends the non-near memory read ID to the read feedback module.
[0131] In some embodiments, the first non-near memory read command issued by the master device can be received by the read receiving module, so that the read receiving module can extract the non-near memory read address and the non-near memory read ID from the first non-near memory read command. Furthermore, the read receiving module can detect the storage status of the second read command generation module, the read cache module and the read feedback module, and when the storage status of the second read command generation module, the read cache module and the read feedback module are all in a non-full state, the non-near memory read address is sent to the second read command generation module for storage, and the non-near memory read ID is sent to the read ID cache module for storage.
[0132] It can be understood that if the storage status of the second read command generation module, the read cache module and the read feedback module is not full, it means that there is no data processing process in the current near memory computing system that conflicts with the current read command. Therefore, the read receiving module can respond to the read command and execute the operation of accessing the read data from the target storage module.
[0133] The second read command generation module can generate a second non-near memory read command for accessing read data according to the read command. The second non-near memory read command is distinguished from the first read command in the above embodiment. The second non-near memory read command is a read command for reading non-near memory read data from the storage module further generated by the second read command generation module when the master device issues the first non-near memory read command. That is, the master device only reads data from the storage module without involving computing operations.
[0134] Among them, when the second read command generation module receives the non-near memory read address, it can generate and output the second non-near memory read command. It can be understood that when generating the second non-near memory read command, it is also necessary to convert the signal type of the second non-near memory read command according to the signal type sent by the master device and the signal type adapted by the storage module. For example, convert the axi access signal into an SRAM access signal.
[0135] In this way, the second read command generation module can send the second non-near memory read command to the storage module to read the non-near memory read data from the storage module. It can be understood that when the storage module receives the second non-near memory read command, it can feed back the non-near memory read data to the second shunt module, so that the second shunt module feeds back the non-near memory read data to the read feedback module, and then the read feedback module feeds back the non-near memory read data and the read ID to the execution unit.
[0136] It is understandable that, because the data that the read feedback module needs to feed back to the master device include the read ID and the non-near memory read data, when the read receiving module detects / obtains the storage state of the read feedback module, it is necessary not only to determine whether the storage state of the read feedback module is not full, but also to determine whether the remaining storage space of the read feedback module is sufficient to accommodate the non-near memory read data.
[0137] like Figure 2 and Fig.10 As shown, the system supporting near memory operation also includes a second beat module. The second read command generation module executes sending the second non-near memory read command to the read ID cache module, and is further configured to: send the second non-near memory read command to the second beat module;
[0138] The second beat module is configured to: in response to the received second non-near memory read command, send the second non-near memory read command to the read ID cache module after beating according to a second preset beat number;
[0139] The read id cache module is configured to: in response to receiving the second non-near memory read command, send the non-near memory read id to the read feedback module.
[0140] In some embodiments, the second read command generation module sends a second non-near memory read command to the read id cache module through the second beat module, so that the read id cache module sends the stored non-near memory read id to the read feedback module. Among them, the second beat module can beat according to the second preset beat number to provide SRAM with sufficient time to search and feedback non-near memory read data. The second beat number can be set according to actual application requirements. For example, the second beat number can be set to 4 beats. When the read id cache module receives the second non-near memory read command again, it can send the non-near memory read id to the read feedback module so that the read feedback module can feedback the read data to the master device.
[0141] In some embodiments, the second beat module is further configured to: in response to the received second non-near memory read command, beat according to a second preset beat number and then send the second non-near memory read command to the second diversion module.
[0142] The second traffic distribution module is configured to receive the non-near memory read data fed back by the storage module.
[0143] When the second non-near memory read command is received, the non-near memory read data is sent to the write feedback module.
[0144] like Fig.10As shown, when the second beat module receives the second non-near memory read command, it also sends the second non-near memory read command to the second shunt module, so that the second shunt module sends the non-near memory read data fed back by the SRAM to the read feedback module. In this way, the read feedback module can receive the non-near memory read data and feed back the non-near memory read data to the master device, so that the master device can obtain the non-near memory read data from the SRAM by sending the first non-near memory read command.
[0145] It is understandable that the second diversion module may have received the non-near memory read data before receiving the second non-near memory read command, but it still needs to wait until the second non-near memory read command is received before sending the non-near memory read data to the read feedback module to cooperate with the overall operating timing of the near memory operating system.
[0146] like Fig.10 As shown, the second beat module is further configured to: send a second non-near memory read command to the read feedback module.
[0147] The read feedback module is configured to: generate a non-near memory read feedback instruction based on the non-near memory read data and the non-near memory read id;
[0148] The non-near memory read feedback instruction is sent to the master device.
[0149] In some embodiments, the second beat module can also send a second non-near memory read command to the read feedback module. In this way, when the feedback module receives the second non-near memory read command, it can generate a non-near memory read feedback instruction based on the non-near memory read data received from the second diversion module and the non-near memory read ID received from the read ID cache module, and then send the non-near memory read feedback instruction to the master device. Then, when the master device receives the non-near memory read feedback instruction, it can obtain the non-near memory read data based on the non-near memory read feedback instruction, and can determine that the first non-near memory read command has been executed based on the non-near memory read ID.
[0150] It is understandable that after receiving the second non-near memory read command, the second beat module can forward the second non-near memory read command to the read feedback module, the read ID cache module and the second diversion module at the same time. In this way, the second beat module can provide SRAM with sufficient time to search and send non-near memory read data through beating, and can also forward the second non-near memory read command at the same time, so that the read feedback module can obtain non-near memory read data in time and feedback non-near memory read data to the main device, thereby cooperating with the overall operation of the near memory operation system.
[0151] like Fig.11As shown, the system supporting near memory operations also includes a read arbitration module. The read arbitration module is configured to: when receiving at least two read instructions, forward the read instructions to the storage module in turn, so that the preceding read instructions and the succeeding read instructions received by the storage module correspond to different data transmission transactions.
[0152] In some embodiments, the first read command generated by the first read command generation module and the second non-near memory read command generated by the second read command generation module are both polled and arbitrated by the read arbitration module before being sent to the storage module. The read arbitration module can make the first read command and the second non-near memory read command be sent to the SRAM in turn, so that the near memory command and the non-near memory command are carried out in turn in the near memory computing system to cooperate with the overall operation of the near memory computing system. In this way, the near memory computing system can meet the host device's demand for non-near memory data and the host device's demand for near memory data.
[0153] Some embodiments of the present application also provide a chip, including a master device, a slave device, and the system supporting near memory operations described in the above embodiments. The master device may be a neural network execution unit, and the target storage module may be an SRAM. The execution units each have a corresponding SRAM, and the execution units can perform general read and write accesses to the SRAM through the system supporting near memory operations, as well as read and write accesses including accumulation operations.
[0154] The master device is configured to send near memory data processing instructions and / or non-near memory data processing instructions to the system supporting near memory operations, so as to perform a data interaction process with the slave device based on the system supporting near memory operations.
[0155] It can be seen from the above technical content that the present application provides a system and chip that supports near memory operations based on an on-chip interconnect network. The system includes a write receiving module, a first read command generating module, a write information cache module and an operation module. The first near memory write command is received by the write receiving module, the near memory write address is stored by the first read command generating module, and the near memory write address and the original data are stored by the write information cache module. The original data sent by the write information cache module is received by the operation module, as well as the near memory read data fed back by the storage module, and the target operation is performed. The target data obtained by the target operation is then written into the storage module. The system can perform the target operation on the side adjacent to the storage module to reduce the occupancy of the on-chip interconnect network during the operation process and optimize the performance of the neural network processing unit.
[0156] Similar parts between the embodiments provided in this application can be referenced to each other. The specific implementation methods provided above are only a few examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other implementation methods expanded based on the scheme of this application without creative work belong to the protection scope of this application.
Claims
1. A system supporting near-memory computing based on an on-chip interconnect network, characterized in that: include: A write receiving module, a first read command generating module, a write information caching module and a calculation module; The write receiving module is configured to: in response to a received first near memory write command, obtain the storage status of the first read command generating module, the write information caching module and the write feedback module; The first near memory write command is a command sent by the master device and the slave device when performing near memory operation based on the AIX bus; The near memory write command includes a near memory write address and original data; If the storage status of the first read command generating module and the write information cache module is not full, the near memory write address and the original data are sent to the write information cache module; and, sending the near memory write address to the first read command generating module; The first read command generation module is configured to: generate a first read command in response to a received near memory write address; Sending the first read command to a storage module, so that the storage module sends the near memory read data corresponding to the first read command in the target cache to the computing module; and sending the first read command to the write information cache module, so that the write information cache module outputs the original data and the near memory write address to the operation module; The operation module is configured to: in response to the received near memory read data and the original data, perform a target operation on the near memory read data and the original data to obtain target data; The target data is written into the storage module according to the near memory write address.
2. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 1, characterized in that: The near memory write command also includes a write operation identifier; the write receiving module is further configured to: send the write operation identifier to the write information cache module.
3. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 1, characterized in that: The near memory computing system further includes a first flow diversion module; The write receiving module executes, in response to the received near memory write command, acquiring the storage status of the first read command generating module, the write information caching module and the write feedback module, and is further configured to: send the first near memory write command to the first diversion module; The first shunting module is configured to: extract the near memory write address and original data from the first near memory write command; The near memory write address is stored in the first read command generation module, and the near memory write address and the original data are stored in the write information cache module.
4. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 3, characterized in that: The near memory computing system also includes a write arbitration module; The operation module executes writing the target data into the target cache, and is specifically configured to: send a second near memory write command to the write arbitration module; The second near memory write command includes a near memory write identifier, the near memory write address, the target data and a near memory write ID; The write arbitration module is configured to, when detecting a near memory write identifier in the second near memory write command, preferentially write the target data into the storage module based on the near memory write address.
5. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 4, characterized in that: The near memory computing system further includes a write feedback module; The write arbitration module is further configured to: send the near memory write id to the write feedback module; The write feedback module is configured to: feed back the near memory write ID to the master device.
6. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 1, characterized in that: The near memory computing system further includes a first beat module; The first read command generation module executes sending the first read command to the write information cache module, and is further configured to: send the first read command to the write information cache module through the first beat module; The first beat module is configured to: in response to the received first read command, beat according to a first preset beat number and then forward the first read command to the write information cache module; The write information cache module is configured to send the original data to the operation module in response to the first read command.
7. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 6, characterized in that: The near memory computing system further includes a second flow diversion module; The storage module is configured to: when sending the near memory read data to the operation module, send the near memory read data to the second diversion module; The first beat module is further configured to: forward the first read command to the second shunt module; The second traffic distribution module is configured to send the received near memory read data to the operation module in response to the first read command.
8. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 5, characterized in that: The write receiving module is further configured to: receive a non-near memory write command, and obtain a write operation identifier of the non-near memory write command; If the write operation identifier is used to indicate that the current write command is a non-near memory write command, the non-near memory write command is sent to the first diversion module; the non-near memory write command includes a non-near memory write address, non-near memory write data, and a non-near memory write ID; The first shunt module is configured to: forward the non-near memory write command to the write arbitration module; The write arbitration module is configured to: in response to the received non-near memory write command, write the non-near memory write data into the storage module according to the non-near memory write address; And sending the non-near memory write ID to the write feedback module, so that the write feedback module returns the non-near memory write ID to the master device.
9. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 1, characterized in that: The near memory computing system further includes: a read receiving module, a second read command generating module, a read ID cache module, and a read feedback module; The read receiving module is configured to: receive a first non-near memory read command, wherein the first non-near memory read command includes a non-near memory read address and a non-near memory read id; Obtaining the storage status of the second read command generation module, the read ID cache module, and the read feedback module; If the storage status of the second read command generation module, the read ID cache module and the read feedback module is not full, the non-near memory read address is sent to the second read command generation module, and the non-near memory read ID is sent to the read ID cache module; The second read command generation module is configured to: generate a second non-near memory read command in response to the received non-near memory read address; Send the second non-near memory read command to the storage module so that the storage module sends the non-near memory read data corresponding to the non-near memory read address to the read feedback module; and send the second non-near memory read command to the read ID cache module so that the read ID cache module sends the non-near memory read ID to the read feedback module.
10. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 9, characterized in that: The near memory computing system further includes a second beat module; The second read command generation module executes sending the second non-near memory read command to the read ID cache module, and is further configured to: send the second non-near memory read command to the second beat module; The second beat module is configured to: in response to the received second non-near memory read command, send the second non-near memory read command to the read ID cache module after beating according to a second preset beat number; The read id cache module is configured to: in response to receiving the second non-near memory read command, send the non-near memory read id to the read feedback module.
11. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 10, characterized in that: The second beat module is further configured to: in response to the received second non-near memory read command, after beating according to a second preset beat number, send the second non-near memory read command to the second diversion module; The second traffic distribution module is configured to: receive the non-near memory read data fed back by the storage module; When the second non-near memory read command is received, the non-near memory read data is sent to the read feedback module.
12. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 10, characterized in that: The second beat module is further configured to: send a second non-near memory read command to the read feedback module; The read feedback module is configured to: generate a non-near memory read feedback instruction based on the non-near memory read data and the non-near memory read id; The non-near memory read feedback instruction is sent to the master device.
13. The system for supporting near-memory computing based on an on-chip interconnect network according to claim 10, characterized in that: The near memory computing system also includes a read arbitration module; The read arbitration module is configured to, upon receiving at least two read instructions, forward the read instructions to the storage module in turn, so that the preceding read instruction and the succeeding read instruction received by the storage module correspond to different data transmission transactions.
14. A chip, characterized in that: include: A master device, a slave device, and a system supporting near memory computing based on an on-chip interconnect network as described in any one of claims 1 to 13; The master device is configured to send near memory data processing instructions and / or non-near memory data processing instructions to the system supporting near memory operations based on the on-chip interconnect network, so as to perform a data interaction process with the slave device based on the system supporting near memory operations based on the on-chip interconnect network.