Accelerator, multi-accelerator system and data transfer method
By generating communication commands in the accelerator, locating and storing data in the buffer partition, integrating and packaging the data before sending it, the problem of low bandwidth utilization in multi-accelerator systems is solved, achieving efficient data transmission and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
In the communication process of a multi-accelerator system, the bandwidth utilization is low because the amount of data transmitted each time is small and a single communication still requires a fixed bandwidth, resulting in resource waste and reduced communication efficiency.
By introducing a computing unit into the accelerator to generate communication commands, the communication module locates the buffer partition of the target accelerator and obtains the base address, stores the communication data in the buffer partition, integrates and sends it to the packet module for packaging when the transmission conditions are met, and the packet module sends it to the target accelerator based on the base address, thus achieving efficient transmission of communication data.
It effectively reduces the waste of bandwidth resources, improves communication efficiency, and enhances the data transmission performance of multi-accelerator systems.
Smart Images

Figure CN121441868B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, specifically to an accelerator, a multi-accelerator system, and a data transmission method. Background Technology
[0002] As the scale and complexity of computing tasks continue to increase, multi-accelerator systems have become a key architecture for meeting high-performance requirements. During task execution, frequent data exchange and communication are required between multiple accelerators. In related technologies, communication between multiple accelerators typically relies on shared system memory or direct point-to-point transmission via the system bus.
[0003] In the process of realizing the concept of this invention, it was found that at least the following problems exist in the related technology: In the communication process of multi-accelerator systems, since the amount of data transmitted each time is often small, and a single communication still requires a fixed bandwidth, the bandwidth utilization rate is low, resulting in a waste of resources and thus reducing communication efficiency. Summary of the Invention
[0004] In view of the above problems, the present invention provides an accelerator and a data transmission method.
[0005] According to a first aspect of the present invention, an accelerator is provided, comprising: a computing unit configured to generate communication instructions, the communication instructions including a communication address and communication data; a communication module configured to receive the communication instructions from the computing unit, determine a buffer partition corresponding to the target accelerator based on a target accelerator pointed to by the communication address in the communication instructions, and obtain a base address for the buffer partition; if it is determined that the communication address is within an address window determined by the base address, store the communication data in the communication instructions in the buffer partition; in response to triggering a transmission condition, send the base address and a plurality of communication data stored in the buffer partition to a packet module, wherein the plurality of communication data includes the communication data in the communication instructions and historical communication data in historical communication instructions whose communication addresses point to the target accelerator; the packet module is configured to package the received plurality of communication data based on the base address and send them to the target accelerator.
[0006] A second aspect of the present invention provides a multi-accelerator system, comprising: a plurality of the above-described accelerators; and an interconnection network connecting the plurality of accelerators, wherein the interconnection network is adapted to a preset communication protocol.
[0007] A third aspect of the present invention provides a data transmission method applied to an accelerator, the accelerator including a computing unit, a communication module, and a packet module; the method includes: the computing unit generating a communication instruction, the communication instruction including a communication address and communication data; the communication module receiving the communication instruction from the computing unit, determining a buffer partition corresponding to the target accelerator based on the target accelerator pointed to by the communication address in the communication instruction, and obtaining a base address for the buffer partition; if it is determined that the communication address is within an address window determined by the base address, storing the communication data in the communication instruction into the buffer partition; in response to triggering a transmission condition, sending the base address and a plurality of communication data stored in the buffer partition to the packet module, wherein the plurality of communication data includes the communication data in the communication instruction and historical communication data in a historical communication instruction pointing to the target accelerator; the packet module, based on the base address, packaging the received plurality of communication data and sending it to the target accelerator.
[0008] According to an embodiment of the present invention, a communication instruction including a communication address and communication data is generated by a computing unit. After receiving the communication instruction, the communication module locates the buffer partition corresponding to the target accelerator and obtains the base address. Based on the address window, it stores the communication data in the buffer partition, completing the targeted storage and orderly management of the communication data. When the transmission conditions are met, current and historical communication data are integrated and sent to the packet module. The packet module packages multiple data packets based on the base address and sends them to the target accelerator, achieving efficient transmission of communication data, effectively reducing bandwidth waste, and improving communication efficiency. Attached Figure Description
[0009] The above-mentioned contents, as well as other objects, features and advantages of the present invention, will become clearer from the following description of embodiments of the present invention with reference to the accompanying drawings.
[0010] Figure 1 A schematic diagram of an accelerator according to an embodiment of the present invention is shown.
[0011] Figure 2 A schematic diagram of the operation of a communication module according to an embodiment of the present invention is shown.
[0012] Figure 3A A schematic diagram of an accelerator according to another embodiment of the present invention is shown.
[0013] Figure 3B A schematic diagram of the communication architecture of an accelerator according to an embodiment of the present invention is shown.
[0014] Figure 4 A schematic diagram of the packaging of the grouping module according to an embodiment of the present invention is shown.
[0015] Figure 5 A schematic diagram of a preset communication protocol according to an embodiment of the present invention is shown.
[0016] Figure 6 A schematic diagram of the unpacking module according to an embodiment of the present invention is shown.
[0017] Figure 7 A schematic diagram of a multi-accelerator system according to an embodiment of the present invention is shown.
[0018] Figure 8 A flowchart of a data transmission method according to an embodiment of the present invention is shown. Detailed Implementation
[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0022] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0023] In the technical solution of this invention, the data involved (including but not limited to data used for analysis, data stored, data displayed, etc.) are all information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of related data all comply with relevant laws, regulations and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse.
[0024] An embodiment of the present invention provides an accelerator, comprising: a computing unit configured to generate communication instructions, the communication instructions including a communication address and communication data; a communication module configured to receive the communication instructions from the computing unit, determine a buffer partition corresponding to the target accelerator based on the target accelerator pointed to by the communication address in the communication instructions, and obtain a base address for the buffer partition; if the communication address is determined to be within the address window determined by the base address, store the communication data in the communication instructions in the buffer partition; in response to triggering a transmission condition, send the base address and multiple communication data stored in the buffer partition to a packet module, wherein the multiple communication data includes the communication data in the communication instructions and historical communication data in historical communication instructions pointing to the target accelerator; and a packet module configured to package the received multiple communication data based on the base address and send them to the target accelerator.
[0025] Figure 1 A schematic diagram of an accelerator according to an embodiment of the present invention is shown.
[0026] like Figure 1 As shown, the accelerator 100 includes: a computing unit 101 configured to generate communication instructions, the communication instructions including a communication address and communication data; a communication module 102 configured to receive the communication instructions from the computing unit 101, determine a buffer partition corresponding to the target accelerator based on the target accelerator pointed to by the communication address in the communication instructions, and obtain a base address for the buffer partition; if the communication address is determined to be within the address window determined by the base address, store the communication data in the communication instructions in the buffer partition; in response to triggering a transmission condition, send the base address and multiple communication data stored in the buffer partition to a packet module 103, wherein the multiple communication data includes the communication data in the communication instructions and the historical communication data in the historical communication instructions whose communication addresses point to the target accelerator; the packet module 103 is configured to package the received multiple communication data based on the base address and send them to the target accelerator.
[0027] According to an embodiment of the present invention, the computing unit 101 of the accelerator 100 can generate communication instructions through built-in instruction generation logic. This logic can parse the data flow requirements of the computing task, determine the communication address in combination with the topological location of the target accelerator, extract the intermediate results or configuration parameters to be transmitted during the computing process as communication data, and finally encapsulate them into a communication instruction including the communication address and communication data according to a preset instruction format.
[0028] The communication module 102 pre-stores a mapping table between each target accelerator and a buffer partition. This mapping table records the correspondence between the identifier of the target accelerator and the buffer partition. After receiving a communication command, the communication module 102 first extracts the identifier representing the target accelerator from the communication command, determines the corresponding buffer partition by looking up the table, and then reads the base address of the buffer partition from its own configuration register.
[0029] The communication module 102 calculates the address window (i.e., base address + preset address offset) based on the base address and the preset address offset. The communication address is compared with the boundary value of the address window. If the communication address is within the address window, the communication data is written to the data storage area corresponding to the buffer partition through the data bus. During storage, a precise mapping method based on the address offset is used, and the storage index of historical communication data is retained to achieve data accumulation.
[0030] When a transmission condition is triggered, the communication module 102 sends the base address and multiple communication data stored in the buffer partition to the packet module 103. Transmission conditions may include the amount of data stored in the buffer partition reaching a preset threshold (e.g., 90% of the partition capacity) or the communication address being outside the address window. The communication module 102 collects these trigger signals in real time. When either condition is met, it immediately integrates the current communication data and historical communication data within the buffer partition to form a data set, synchronously extracts the base address of the buffer partition, and transmits the base address and data set together to the packet module 103.
[0031] The packet module 103 encapsulates multiple communication data into transaction layer data packets according to the data packet format of a preset communication protocol. After packaging, the transaction layer data packets are sent to the target accelerator through the interconnection channel with the target accelerator.
[0032] The computing unit generates communication commands including communication addresses and data. Upon receiving these commands, the communication module locates the corresponding buffer partition of the target accelerator, obtains the base address, and stores the communication data in the buffer partition according to the address window, thus completing the targeted storage and orderly management of the communication data. When the transmission conditions are met, current and historical communication data are integrated and sent to the packet module. The packet module packages multiple data packets based on the base address and sends them to the target accelerator, achieving efficient transmission of communication data, effectively reducing bandwidth waste, and improving communication efficiency.
[0033] According to an embodiment of the present invention, the buffer partition includes multiple buffer units; the communication module is further configured to: match the address label of the communication address with the address label of each buffer unit to obtain a matching result; if the matching result indicates that there is an existing unit in the multiple buffer units that corresponds to the address label of the communication address, determine the existing unit as the target buffer unit so as to store the communication data in the target buffer unit; if the matching result indicates that there is no existing unit in the multiple buffer units that corresponds to the address label of the communication address, determine the target buffer unit from the free units of the multiple buffer units so as to store the communication data in the target buffer unit.
[0034] The buffer partition comprises multiple buffer units, each including an independent address tag storage area, a data storage area, and status flags (used to indicate whether the buffer unit is currently idle or occupied). Each buffer unit is connected to the control logic of the communication module via an internal bus, supporting parallel read and write operations to improve data processing efficiency.
[0035] When the communication module performs address tag matching, it extracts the communication address from the received communication command to determine the address tag based on the communication address. Specifically, the address tag can be obtained through a preset address segmentation rule, such as truncating the high 16 bits of the communication address as the address tag to distinguish different data transmission contexts.
[0036] Subsequently, the extracted address tags are compared with the address tags stored in the address tag storage area of all buffer units. The matching process can be implemented using a parallel comparator at the hardware level to reduce matching latency, and finally generate a matching result including whether the match was successful and the index of the matching unit.
[0037] When the matching result indicates that an existing unit corresponding to the address label of the communication address exists among multiple buffer units, the communication module will first locate the existing unit using the unit index in the matching result. Simultaneously, it will verify whether its status flag is in an "occupied" state (to reduce data overwriting issues caused by abnormal unit status). If confirmed, the existing unit will be designated as the target buffer unit. Subsequently, the communication data in the communication command will be transmitted to the data storage area of the target buffer unit. If the target buffer unit already stores historical communication data, an overwrite or append write operation can be performed as needed (the specific method is determined by the system's preset storage strategy).
[0038] If the matching results show that no existing unit corresponds to the address tag of the current communication address among multiple buffer units, the communication module will immediately query the status flag bits of each buffer unit and filter out all buffer units with a status of "idle". To ensure balanced utilization of buffer resources, a round-robin selection strategy can be used to determine the target buffer unit from multiple idle units, or the optimal idle unit can be dynamically selected based on the storage performance of the buffer unit (such as read / write speed, access priority). After determining the target buffer unit from multiple idle units, the communication module will first write the address tag of the current communication address into the address tag storage area of the target buffer unit and update its status flag bit to "occupied". After completing the unit initialization configuration, the communication data will be stored in the data storage area of the target buffer unit.
[0039] By matching the address tags of the communication address with the address tags of the buffer unit, the target buffer unit can be accurately located, enabling efficient storage of communication data. When a matching existing unit exists, it is used directly to store the communication data, reducing redundant resource allocation. When no matching unit exists, the target buffer unit is selected from the available units, making full use of the available resources of the buffer partition, improving the utilization rate of the buffer partition and the flexibility of data storage, thereby optimizing the storage management efficiency of the communication module.
[0040] According to an embodiment of the present invention, the communication module is further configured to: determine the starting position of the communication data in the target buffer unit based on the address offset of the communication address relative to the base address; write the communication data into the data storage area of the target buffer unit based on the starting position; determine byte masks corresponding to multiple byte positions in the target buffer unit according to the starting position and the data length of the communication data; and mark the corresponding byte enable bits as valid by performing a bitwise OR operation between the byte mask and the byte enable bits corresponding to each byte position in the target buffer unit.
[0041] The communication module has a built-in address arithmetic unit for calculating the address offset of the communication address relative to the base address. Specifically, the address offset, expressed in bytes, is obtained by subtracting the base address from the communication address. During the calculation, the consistency of the address bit width must be checked simultaneously to ensure that both are 32-bit or 64-bit physical addresses, thereby reducing calculation errors caused by bit width mismatch.
[0042] After calculating the address offset, the starting position of the communication data in the target buffer unit is determined. The determination of the starting position is based on the internal storage architecture of the target buffer unit, which is divided into contiguous byte-level storage units. The communication module directly maps the calculated address offset to the starting byte address within this data storage area. Simultaneously, it verifies whether the starting position is within the storage address range of the target buffer unit (i.e., starting position < total number of bytes in the target buffer unit) using address boundary detection logic. If it exceeds the range, an address error correction mechanism is triggered to ensure the legality of the communication data write.
[0043] Based on the starting position, the communication module writes the communication data into the data storage area of the target buffer unit. Specifically, the communication module generates corresponding read / write timing signals according to the bit width of the communication data (e.g., 32-bit, 64-bit), and transmits the communication data in parallel to the storage unit corresponding to the starting byte address of the target buffer unit through the data bus. Subsequent bytes of data are then written sequentially at consecutive addresses.
[0044] During the writing process, the communication module determines the byte mask corresponding to multiple byte positions in the target buffer unit based on the starting position and the data length of the communication data. Subsequently, the byte mask is bitwise ORed with the byte enable bits corresponding to each byte position in the target buffer unit, marking the corresponding byte enable bits as valid.
[0045] When generating the byte mask, the communication module first obtains the byte index corresponding to the starting position (e.g., byte index 5 when the starting position is 5 bytes) and the data length of the communication data (specified by the length field in the communication instruction, in bytes). Next, it generates an initial mask (initially all 0s) through a shift register, and sets N consecutive bits (N being the data length) starting from the corresponding bit of the initial mask to 1 according to the byte index, forming the byte mask corresponding to the communication data (e.g., byte mask 0b00001110 when the starting position is 2 and the data length is 3).
[0046] Each byte position in the target buffer unit corresponds to an independent byte enable bit, which is stored in a dedicated enable register (1 indicates valid, 0 indicates invalid). The communication module inputs the generated byte mask and the current value in the enable register into the bitwise OR operation unit. The bitwise OR operation preserves the marked valid byte bits and marks the newly added valid byte positions as 1. The operation result is written back to the enable register in real time, completing the update of the byte enable bits. This process ensures that subsequent data readings only extract the marked valid byte data, reducing the overhead of transmitting invalid data.
[0047] By determining the starting position of communication data in the target buffer unit based on the address offset of the communication address, the communication data can be accurately written into the specified data storage area. At the same time, a byte mask is generated based on the starting position and data length, and the enable bit of the byte is marked as valid by bitwise OR operation, thereby achieving precise control and efficient management of data storage, and improving the accuracy and flexibility of data storage.
[0048] According to an embodiment of the present invention, the communication module is further configured to: generate a sub-header for communication data based on the address offset, the address tag of the target buffer unit, and the payload length of the communication data identified by the byte enable bit in the target buffer unit; bind the sub-header and the communication data, and store them in the target buffer unit.
[0049] When the communication module generates the sub-header, it first processes the three core parameters—address offset, target buffer unit address tag, and payload length—using built-in field extraction and calculation logic. For the address offset field, the communication module directly calls the previously calculated byte-level address offset of the communication address relative to the base address. Considering the bit width adaptation requirements for storage and transmission, this address offset is converted into binary data of a preset bit width (e.g., 16 bits). If the address offset value is less than the preset bit width, high-order bits are padded with 0s to ensure the uniformity of the field format.
[0050] The address tag of the target buffer unit is obtained through the unit indexing logic of the communication module. This logic reads the corresponding address tag from the tag storage area of the buffer unit based on the determined index of the target buffer unit. Similarly, the format is standardized according to the preset tag field width (e.g., 24 bits) in the sub-header to reduce subsequent parsing anomalies caused by inconsistent address tag widths.
[0051] The payload length is obtained based on the state of the byte enable bits in the target buffer unit. The communication module scans and counts the byte enable bits corresponding to all byte positions in the target buffer unit bit by bit, and uses the number of valid byte enable bits as the payload length of the communication data (in bytes). Subsequently, the payload length is converted into field data of a preset bit width (e.g., 8 bits).
[0052] After processing the three core parameters, the communication module uses sub-header encapsulation logic to concatenate the fields in a preset field order (e.g., "address label - address offset - payload length - reserved field") to form a fixed-length (e.g., 64-bit) sub-header. The reserved field is used for subsequent function expansion and is filled with 0 by default.
[0053] The generated sub-header is bound to the communication data. Specifically, the generated sub-header is used as a prefix and concatenated with the communication data byte-by-byte to form a combined data structure of "sub-header + communication data". After binding, the communication module writes the sub-header and communication data completely into the designated storage area of the target buffer unit through the write enable signal of the target buffer unit. At the same time, the status flag bit of the target buffer unit is updated to mark that the unit has stored "communication data with sub-header", so that the complete bound data can be quickly identified and retrieved when the transmission condition is triggered subsequently.
[0054] The communication module generates a sub-header for the communication data based on the address offset, the address tag of the target buffer unit, and the payload length. This sub-header is then bound to the communication data and stored in the target buffer unit. This process provides more detailed metadata for the communication data, facilitating subsequent rapid location and efficient management.
[0055] According to an embodiment of the present invention, the communication module further includes an available payload length register; wherein the available payload length register is initialized to a preset maximum payload length; the available payload length register is configured to: for a communication instruction, when the communication data has been stored in the buffer partition, determine the remaining payload length based on the data length of the communication data, the field length of the sub-header, and the current payload length; and in response to receiving a next communication instruction, trigger a transmission condition if it is determined that the remaining payload length is less than the sum of the data length of the communication data in the next communication instruction and the field length of the corresponding sub-header.
[0056] The communication module also includes an available payload length register, which adopts a hardware-level register design. The bit width is configured according to the preset maximum payload length (e.g., 32 bits, which can cover the payload length range from 0 to 4GB), and has high-speed read and write characteristics to adapt to the real-time processing rhythm of communication data.
[0057] The initialization operation of the available payload length register is automatically performed during the power-on reset phase of the communication module, reading the preset maximum payload length from the preset configuration register. The preset maximum payload length is pre-calculated and determined based on the total storage capacity of the buffer partition, the maximum field length of the sub-header, and data transmission efficiency requirements. For example, when the total buffer partition capacity is 1MB and the maximum field length of the sub-header is 8 bytes, the preset maximum payload length can be configured to 1048568 bytes, reserving some redundant space to reduce overflow. The preset maximum payload length is then written to the available payload length register to complete the initialization configuration.
[0058] When processing the current communication command, if the communication data has already been stored in the buffer partition, the communication module will calculate the remaining payload length. During the calculation, it first reads the current payload length from the available payload length register, and then extracts the data length of the current communication data and the field length of the sub-header. The field length of the sub-header depends on the number of bits in the address offset and the maximum payload length of the communication data. The remaining payload length is obtained through the operation logic of "current payload length - (data length of communication data + field length of sub-header)". After the calculation is completed, the result is immediately written back to the available payload length register, achieving real-time updates.
[0059] When the communication module receives the next communication command, it first parses the data length of the communication data in the next communication command, and then calculates the total length required to store the communication data and corresponding sub-header in the next communication command, based on the length of the corresponding sub-header field. Subsequently, the communication module compares the remaining payload length in the available payload length register with the required total length in real time.
[0060] If the comparison result shows that the remaining payload length is less than the required total length (i.e., the remaining space in the buffer partition is insufficient to accommodate the communication data and corresponding sub-header in the next communication command), a high-level trigger signal is output. This signal is directly connected to the transmission control logic of the communication module, thereby triggering the transmission conditions and initiating the process of transmitting the communication data stored in the buffer partition.
[0061] In addition, the communication module also includes a base address register. The base address register is used to store the base address for the buffer partition corresponding to the target accelerator, providing support for subsequent operations such as address offset calculation and address window determination.
[0062] By setting the available payload length register, the remaining space in the buffer partition can be tracked in real time, and the sending condition can be triggered when the space is insufficient, reducing data overflow and ensuring the integrity and reliability of data transmission.
[0063] According to an embodiment of the present invention, the communication module is further configured to: trigger a transmission condition when it is determined that the communication address is outside the address window; initialize a buffer partition to redetermine the base address and the corresponding address window based on the communication address in the next received communication instruction.
[0064] The communication module compares the communication address with the boundary value of the address window. When the comparison result indicates that the communication address is outside the address window, it outputs a high-level trigger signal, thereby triggering the transmission condition.
[0065] To maintain system-wide memory consistency, the communication module also integrates a counter to track the number of communication instructions being processed in each buffer partition in real time. When the communication module receives a release operation from the system level (e.g., a memory barrier operation or a synchronization signal indicating the end of a compute kernel's execution), it must flush the entire relevant buffer partition. In the accelerator memory consistency model, "system-wide" refers to synchronization operations that act on all devices in the system, while "release operations" refer to synchronization operations performed by data producers (such as processors or other accelerators) before making their data visible to other consumers.
[0066] The accelerator's release operation requires the hardware implementation to flush all in-transit communication commands to a persistent point, ensuring these operations are visible to all threads within the specified synchronization scope. Therefore, unconditionally flushing the communication module upon receiving a system-wide release operation is crucial for ensuring compatibility with the strict memory consistency model of heterogeneous accelerators.
[0067] After determining the trigger condition for transmission (whether due to address out-of-bounds access or system release operation), the communication module will send the communication data already stored in the buffer partition, bound to the corresponding sub-header, and its base address to the downstream packet module, and then initialize the buffer partition. The initialization process is not a simple sequential reset, but rather a predictive resource mapping operation performed by the reset control logic. Specifically, while clearing the buffer unit, based on the size and address pattern of the just-sent data block, a lightweight predictor dynamically pre-calculates and adjusts the address alignment parameters that may be used in the next cycle (such as temporarily fine-tuning the preset address offset bits), and temporarily stores these predicted parameters along with the invalid base address. When the next communication instruction arrives, its address resolution process will prioritize using these predicted parameters for alignment calculation. If a match is successful, some hardware states can be directly reused, significantly reducing window reconstruction latency and achieving overlapping optimization between the initialization and next data transmission preparation phases.
[0068] To ensure the order of loading communications at the same address, any load operation on remote memory that matches the address of a write operation temporarily stored in the grouping or communication module must refresh all communication operations currently queued in the remote write queue that match the address. These matching operations can be refreshed individually, or the load hit event itself can trigger a refresh of the entire remote write queue, just like a synchronization operation, thus strictly guaranteeing the read-write order.
[0069] The initialization of the buffer partition is performed by the reset control logic built into the communication module. After initialization is triggered, this logic first sends a clear signal to all buffer units within the partition, clearing the address tags in the address tag storage area and all communication data in the data storage area of each unit. At the same time, it resets the byte enable bit corresponding to each byte position to an invalid state (logic '0') and updates the status flag bit of each buffer unit to "idle".
[0070] Synchronously, the reset control logic sends an initialization signal to the available payload length register associated with the partition, restoring it to the preset maximum payload length value; its base address register is then temporarily stored as an invalid address identifier (e.g., an address with all '1's).
[0071] When the communication module receives the next communication command, it first extracts the target accelerator identifier from the command and determines the buffer partition by querying the internal mapping table. Next, it parses the communication address carried in the command and determines the new base address by performing alignment calculations. For example, assuming the preset address offset is 12 bits and the received communication address is 0x4078, the new base address obtained after clearing the lower 12 bits is 0x4000.
[0072] Subsequently, the communication module writes the calculated new base address into the base address register corresponding to the buffer partition. Based on this base address and the preset address offset bits, a new address window can be defined. Continuing the previous example, the lower bound of the address window is the base address 0x4000, and the upper bound is base address + 2^12 = 0x4000 + 0x1000 = 0x5000 (the actual covered address range is [0x4000, 0x4FFF]). Therefore, the communication address 0x4078 is exactly within this newly defined address window, and its address offset relative to the base address is 0x78.
[0073] By triggering the transmission and initialization process when the communication address exceeds the current address window, and by strictly responding to system consistency events and sequence constraints, this design ensures efficient turnover and reuse of buffer partitions. This reduces address space conflicts and data error overwriting, while also guaranteeing consistency with the upper-level memory model, thereby significantly improving the flexibility, reliability, and overall communication efficiency of the communication module in a multi-accelerator system.
[0074] Figure 2 A schematic diagram of the operation of a communication module according to an embodiment of the present invention is shown.
[0075] like Figure 2As shown, after receiving the communication instruction, which includes the communication address and communication data, from the computing unit, the communication module calls the base address stored in the base address register for buffer partitioning and performs a matching check between the communication address and the address window corresponding to the base address, effectively reducing the problem of buffer partition data corruption caused by address out-of-bounds errors.
[0076] After the address matching is successful, the communication module calls the current value of the available payload length register to compare in real time whether the total length of the communication data and the corresponding sub-header is within the remaining payload range. This real-time register comparison method replaces the traditional software polling, which can reduce the waiting delay before data storage and improve the utilization efficiency of buffer resources.
[0077] After length verification, the communication module performs parallel matching of the address tag of the communication address with the address tags of each buffer unit within the buffer partition. If a matching existing unit exists, the communication data is directly written to that existing unit. If no match exists, the target buffer unit is selected from the idle units to complete the writing of the communication data. This parallel tag matching mechanism allows for simultaneous traversal of multiple buffer units, effectively accelerating the data writing location speed.
[0078] If the address match fails or the length check is not met, the communication module will immediately trigger the transmission condition, sending the base address in the current buffer partition and multiple stored communication data to the packet module, while simultaneously initializing the buffer partition. This condition-triggered batch transmission replaces fixed-period transmission, which can promptly release space when buffer resources are insufficient, reduce the data residence time in the buffer partition, and thus reduce the end-to-end latency of cross-accelerator data transmission.
[0079] Figure 3A A schematic diagram of an accelerator according to another embodiment of the present invention is shown.
[0080] like Figure 3A As shown, the accelerator adopts a modular integrated architecture. Internally, the unpacking module, computing unit, communication module and grouping module are connected through a high-speed internal bus. The functional boundaries of each module are clear and the interaction path is fixed, which can reduce communication redundancy between modules and improve the efficiency of data flow within the accelerator.
[0081] The communication module, serving as the core of data temporary storage, has n independent buffer partitions (e.g., buffer partition 1 to partition n), each employing a physically isolated storage space design. This multi-partition structure can classify and store communication data from different sources or senders, reducing mutual overwriting or interference between data streams. It also supports parallel reading and writing of multiple sets of data, adapting to the batch data temporary storage needs in high-throughput scenarios and significantly improving the flexibility of buffer resource utilization.
[0082] The computing unit transmits the generated communication instructions to the communication module via the internal bus. The communication module, based on the identifiers in the communication instructions, distributes the corresponding communication data to the matching buffer partitions (such as buffer partitions 2 and 3 marked with specific data writes). The packet module then extracts the communication data that meets the sending conditions from each buffer partition, encapsulates it into target transaction layer data packets, and sends them externally.
[0083] After receiving the target transaction layer data packet from the external interconnect channel, the unpacking module parses the communication data and transmits it to the computing unit for further processing. This inter-module collaborative process effectively reduces data scheduling latency, while the dynamic reuse of multiple buffer partitions reduces transmission congestion caused by a single partition being fully loaded, further improving the accelerator's communication reliability.
[0084] Figure 3B A schematic diagram of the communication architecture of an accelerator according to an embodiment of the present invention is shown.
[0085] like Figure 3B As shown, in this multi-accelerator system, each accelerator (e.g., accelerator 0, accelerator n) is equipped with multiple computing units with independent L1 caches. The L1 cache temporarily stores intermediate data from the local computation of the computing unit, and a subsequent translation buffer enables fast translation from virtual to physical addresses. This configuration reduces the number of memory accesses during address translation, improving the data read / write response speed of the computing unit.
[0086] Data flows from multiple computing units are scheduled via a crossbar switch. This crossbar switch supports non-blocking multi-port data forwarding, reducing data flow congestion caused by a single computing unit monopolizing the link. Simultaneously, it aggregates the processed data from each computing unit into the final-level cache. This final-level cache serves as a shared storage layer within the accelerator, enabling efficient data sharing between computing units and significantly reducing latency in cross-unit data interaction.
[0087] The communication module within the accelerator is configured with multiple independent buffer partitions to isolate and temporarily store communication data to be sent to different target accelerators. This isolated storage method reduces data confusion between different targets and improves the orderliness of data management. When the data in the buffer partition meets the transmission conditions, the packet module encapsulates the corresponding sub-headers of each buffer partition with the communication data into a data packet with a transaction layer header. The unified transaction layer encapsulation format can adapt to the intermediate accelerator consistency interconnection protocol, ensuring format compatibility of cross-accelerator data transmission.
[0088] Accelerator 0 and accelerator n exchange data through an accelerator consistency interconnect. This interconnect link supports high-speed transmission of transaction layer data packets. Its consistency design can ensure the state consistency of shared data between different accelerators and reduce the problem of dirty reads when multiple accelerators cooperate.
[0089] The receiving end's unpacking module parses the incoming data packets, separates the sub-header and communication data, and distributes them to the corresponding computing unit's conversion backup buffer through a cross switch. This enables precise data transfer to the computing unit, ensuring the effective landing of cross-accelerator data and supporting the collaborative efficiency of multiple accelerators in large-scale parallel computing scenarios.
[0090] According to an embodiment of the present invention, the packet module is further configured to: in response to multiple received communication data, encapsulate the communication data and corresponding sub-headers in each buffer unit of the buffer partition into a transaction layer data packet; fill the base address into the address field of the transaction layer data packet, and fill the total length of the payload of the multiple communication data in the transaction layer data packet into the length field of the transaction layer data packet to obtain a target transaction layer data packet; and send the target transaction layer data packet to the target accelerator through an interconnection channel corresponding to a preset communication protocol.
[0091] After receiving multiple communication data and their corresponding base addresses from the communication module, the packet module sequentially locates the data stored in each buffer unit using the unit index, and extracts the communication data and corresponding sub-headers from each buffer unit. Information such as the address offset and payload length in the sub-headers is used as the core metadata of the transaction layer data packet, while the communication data serves as the payload body of the transaction layer data packet.
[0092] Subsequently, the grouping module uses its built-in transaction layer encapsulation engine to encapsulate data according to the specifications of a preset communication protocol, such as Compute Express Link (CXL). The sub-headers and communication data in each buffer unit are encapsulated into independent transaction layer data packets. During the encapsulation process, necessary control information such as transaction type identifiers and priority fields are automatically added to ensure that the transaction layer data packets meet the transmission requirements of the preset communication protocol.
[0093] Specifically, during the encapsulation process, the grouping module can encapsulate data in ascending order of the address offsets corresponding to the buffer units, ultimately obtaining the transaction layer data packet. During this process, if the valid bytes within a single buffer unit are not contiguous, the data in that buffer unit can be split into multiple contiguous data blocks, and corresponding sub-header information can be generated for each data block.
[0094] After completing the basic encapsulation of each buffer unit, the packet module initiates the construction process for the target transaction layer data packet. Regarding base address padding, the packet module uses address field mapping logic to write the base address of the received buffer partition into the reserved address field of the transaction layer data packet.
[0095] To calculate the total payload length, the grouping module iterates through all encapsulated transaction layer data packets corresponding to each buffer unit, extracts the payload length of the communication data in each transaction layer data packet (derived from the payload length field in the sub-header), and performs an accumulation operation to obtain the total payload length of multiple communication data packets. Then, the total payload length is written into the length field of the transaction layer data packet, completing the filling of the core fields.
[0096] After the target transaction layer data packet is constructed, the packet module identifies the preset communication protocol type through protocol adaptation logic and converts the target transaction layer data packet into a frame format that conforms to the link layer specification of the preset communication protocol.
[0097] Subsequently, the packet module sends the target transaction layer data packet through an interconnection channel (such as a CXL link) corresponding to a preset communication protocol. Before sending, it confirms the connection status of the interconnection channel and the target accelerator's receive-ready signal to ensure a smooth transmission link. During transmission, if the target transaction layer data packet is lost or erroneous, the packet module will retransmit until the target accelerator successfully receives it and sends back an acknowledgment signal.
[0098] By encapsulating communication data and sub-headers into transaction layer packets, filling in the base address and total payload length, a complete target transaction layer packet is formed and sent to the target accelerator through the interconnection channel. This optimizes the data transmission format, effectively improves transmission efficiency and accuracy, and ensures that the data can be correctly received and processed by the target accelerator.
[0099] Figure 4 A schematic diagram of the packaging of the grouping module according to an embodiment of the present invention is shown.
[0100] Figure 5 A schematic diagram of a preset communication protocol according to an embodiment of the present invention is shown.
[0101] like Figure 4 As shown, in the batch data encapsulation process of the grouping module, the communication module first generates a sub-header for each independent communication data (the address offset is one of the core fields of the sub-header, which is integrated with metadata such as address tag and payload length). Then, the initial splicing of a single group of data is completed in a fixed order of "communication data - sub-header containing address offset" to form a basic data group.
[0102] For multiple pieces of communication data, the above basic data group splicing process will be repeated. Specifically, each line corresponds to an independent combination of "communication data - sub-header containing address offset". The multi-line reuse structure allows the grouping module to schedule batch data through a unified loop processing logic, reducing the redundant operation of encapsulating individual data, effectively improving the encapsulation efficiency of batch data, and adapting to the high-throughput data transmission requirements in multi-accelerator collaborative scenarios.
[0103] Finally, the grouping module performs a sequential overlay operation on these multi-row basic data groups, integrating them into a continuous data stream. A transaction layer header is added to the beginning of the integrated data stream. This sequential overlay method strictly aligns the generation order of the basic data groups with the content order of the encapsulated data packet, reducing data out-of-order issues during cross-accelerator transmission. The unified addition of the transaction layer header integrates the necessary global control and verification dimensions for the entire data packet, improving the reliability of data transmission across the interconnected network.
[0104] like Figure 5 As shown, the main structure of the target transaction layer data packet is further improved by the grouping module according to a fixed framework. First, a sequence number field is added. This field is automatically assigned a continuous identifier by the grouping module according to the generation order of the target transaction layer data packet, so that the receiving end unpacking module can quickly calibrate the data packet transmission order in multi-packet concurrent scenarios and effectively avoid the risk of data out-of-order transmission.
[0105] Subsequently, the transaction layer header is constructed. The grouping module extracts the base address of the buffer partition from the base address register of the communication module, and simultaneously obtains the enable bit information corresponding to each communication data, integrating these contents into the transaction layer header. This centralized encapsulation of core control information allows the receiving end to obtain the key parameters for data location by parsing the transaction layer header only once during unpacking, significantly reducing the time overhead caused by scattered field queries.
[0106] The data filling in the transaction layer relies on the address offset within the sub-header. The grouping module reads each group of data sequentially from the communication module's buffer partition, combining the address offset, data length, and data bits (the actual communication data content) in the sub-header to complete the filling using structured, repetitive grouping logic. This arrangement based on an integrated sub-header not only adapts to the encapsulation requirements of batch data but also allows the grouping module to efficiently process multiple groups of data through standardized loop parsing logic. Furthermore, it facilitates data splitting by group during unpacking at the receiving end, reducing the complexity of the unpacking process.
[0107] Finally, the grouping module generates a checksum based on the entire transaction layer packet content and fills it into the error check segment. This checksum can quickly verify the integrity of the data packet during the unpacking stage at the receiving end, promptly identify data corruption issues during transmission, reduce invalid data flowing into subsequent processing flows, and further improve the reliability of data transmission.
[0108] According to an embodiment of the present invention, the accelerator further includes an unpacking module; wherein the unpacking module is configured to receive a target transaction layer data packet via an interconnection channel; parse the address field of the target transaction layer data packet to obtain a base address; parse the length field of the target transaction layer data packet to obtain the total payload length of multiple communication data in the target transaction layer data packet, and extract each sub-header and the corresponding communication data in sequence; for each sub-header, decode to obtain the address offset of the corresponding communication address relative to the base address, and the payload length of the corresponding communication data; determine the starting position of the communication data in multiple buffer units in the buffer partition determined by the base address according to the address offset, and write the communication data into the corresponding buffer unit according to the payload length; and send the communication data written into each buffer unit to the computing unit for processing.
[0109] When the unpacking module receives a target transaction layer data packet, it first parses the address field of the packet to obtain the base address corresponding to the buffer partition. Then, it writes this base address into the unpacking module's base address temporary register, providing a reference for subsequent address calculations. Next, it parses the length field of the target transaction layer data packet and reads the total payload length recorded within it. This total payload length is used to verify whether the subsequently extracted sub-header matches the total length of the communication data.
[0110] Based on the encapsulation structure of the target transaction layer data packet ("base address + length field + multiple sub-headers - communication data combination"), each sub-header and its corresponding communication data are extracted sequentially using the address offset positioning method. During extraction, the payload part of the target transaction layer data packet is split in turn according to the field length of the sub-header, forming an independent data group of a single sub-header and its corresponding communication data.
[0111] For each extracted sub-header, the unpacking module splits the fields according to a preset field format (such as "address identifier-address offset-payload length"), and decodes them to obtain the address offset of the communication address relative to the base address and the payload length of the corresponding communication data. Subsequently, the unpacking module initiates the address mapping logic, accumulates the base address and the decoded address offset to obtain the starting address of the communication data in the buffer partition, and then combines it with the address range division rules of each buffer unit in the buffer partition (such as each buffer unit occupying a fixed N bytes of address space) to locate the buffer unit to which the starting address belongs, and determines the relative starting position of the communication data within that buffer unit.
[0112] Based on the payload length, the unpacking module writes the communication data into the corresponding data storage area in the buffer unit in byte order. During the writing process, the byte enable bit (marking the valid data range) and status flag bit (set to "ready") of the buffer unit are updated synchronously.
[0113] After each buffer unit completes writing the communication data, the unpacking module sends a data ready signal to the computing unit via the internal data bus. Once the computing unit responds with a receive ready signal, the data transmission logic is initiated. Specifically, the communication data in each buffer unit is read sequentially according to their processing priority. During transmission, metadata such as the identifier of the buffer unit containing the communication data and the effective payload length are included to facilitate rapid identification and processing by the computing unit.
[0114] If the computing unit needs to process data in batches, the unpacking module can temporarily store the communication data of each buffer unit, integrate it according to the needs of the computing unit, and then transmit it, ensuring the efficiency and orderliness of data processing.
[0115] The unpacking module receives and parses target transaction layer data packets, extracting the base address, total payload length, and sub-header information. It then decodes the data to obtain the address offset of the communication address and the payload length of the communication data. Based on this, the communication data is accurately written into the buffer unit and finally sent to the computing unit for processing. This process achieves efficient data reception, parsing, and storage, ensuring data integrity and accuracy, and improving the accelerator's communication efficiency and overall performance.
[0116] Figure 6 A schematic diagram of the unpacking module according to an embodiment of the present invention is shown.
[0117] like Figure 6 As shown, the unpacking module first receives the target transaction layer data packet from the interconnection channel. It first parses the transaction layer packet header of the target transaction layer data packet and extracts global control information such as the buffer partition base address from it. This enables the unpacking module to quickly obtain the reference parameters for data location, saving the overhead of multiple subsequent queries and improving the startup efficiency of parsing.
[0118] Subsequently, the unpacking module, based on the fixed field format of the sub-header, splits the main body of the transaction layer data packet, obtaining multiple encapsulation units of "communication data - sub-header containing address offset". The address offset is embedded as a core field in the sub-header structure. This streamlined encapsulation unit form allows the unpacking module to directly complete the splitting through standardized field length matching logic, without needing to separately identify the offset field. This reduces redundant operations of byte-by-byte parsing and significantly speeds up the splitting of batch data.
[0119] For each encapsulation unit, when the unpacking module parses the sub-header, it simultaneously extracts metadata such as address offset, address tag, and payload length. This metadata is then combined with the already obtained base address to perform address calculations and reconstruct the original communication address. Since the address offset and other core metadata are integrated into the same sub-header, all address-related information can be obtained in a single parsing step. This reduces parsing steps and makes the association between metadata and offsets closer, enabling more accurate matching of the original address corresponding to the data and reducing data address misalignment issues after cross-accelerator transmission.
[0120] Subsequently, the unpacking module extracts communication data from the encapsulation unit. At the same time, it combines the effective payload length in the sub-header to identify the enable bits (marking the valid state of the data) and invalid bits (removing placeholder content) of the corresponding data. It then reassembles the data into a single data record according to the structure of "communication address-communication data-enable bit". This reassembly logic is repeated for multiple rows to form a complete data list. This not only effectively filters out valid data but also restores the original organization of the data, making it easier to accurately distribute the data to the computing unit and improving the efficiency of data flow.
[0121] Figure 7 A schematic diagram of a multi-accelerator system according to an embodiment of the present invention is shown.
[0122] The present invention also provides a multi-accelerator system, such as Figure 7 As shown, the system includes: multiple accelerators; an interconnection network connecting the multiple accelerators, wherein the interconnection network is adapted to a preset communication protocol.
[0123] According to embodiments of the present invention, the multiple accelerators of the multi-accelerator system adopt a standardized hardware architecture design, supporting unified functional module expansion and interface adaptation. In actual deployment, the number of accelerators can be configured in different ways, such as board-level integration or rack-level cluster deployment, depending on the system's computing power requirements. Each accelerator contains a complete computing unit, communication module, packetization module, and unpacking module, and each accelerator is equipped with a unique device identifier (such as physical address code or logical number). This identifier is pre-stored in the configuration register of each accelerator and is used for addressing and communication routing in the interconnection network.
[0124] The interconnect network is the core of data interaction between multiple accelerators, employing a high-bandwidth, low-latency architecture. Its hardware components include interconnect switches, high-speed link interfaces, routing control units, and protocol adaptation circuits. The specific topology can be selected based on the number of accelerators and communication requirements, and is adapted to a pre-defined communication protocol (such as CXL).
[0125] In addition, the interconnection network supports a dynamic bandwidth allocation mechanism, which can flexibly adjust the proportion of link bandwidth resources according to the communication load requirements of each accelerator, reducing the increase in communication latency of other accelerators caused by a single accelerator occupying too many resources.
[0126] By connecting multiple accelerators through an interconnected network and adapting to a pre-defined communication protocol, the multi-accelerator system enables efficient communication and collaborative work between accelerators, significantly improving the overall performance and scalability of the accelerator system.
[0127] Figure 8 A flowchart of a data transmission method according to an embodiment of the present invention is shown.
[0128] The present invention also provides a data transmission method, such as... Figure 8 As shown, the data transmission method includes operations S810 to S850.
[0129] When operating the S810, the computing unit generates communication instructions, which include communication address and communication data.
[0130] When operating the S820, the communication module receives communication instructions from the computing unit, determines the buffer partition corresponding to the target accelerator based on the target accelerator pointed to by the communication address in the communication instructions, and obtains the base address for the buffer partition.
[0131] When operating the S830, if it is determined that the communication address is within the address window determined by the base address, the communication data in the communication instruction is stored in the buffer partition.
[0132] In operation S840, in response to the triggering of the transmission condition, the base address and multiple communication data stored in the buffer partition are sent to the packet module. The multiple communication data include communication data in the communication command and historical communication data in the historical communication command pointing to the target accelerator.
[0133] When operating the S850, the packet module packages multiple received communication data based on the base address and sends them to the target accelerator.
[0134] According to an embodiment of the present invention, the computing unit, as the core of communication instruction generation, first parses the data flow requirements of the current computing task to determine the communication address, and extracts intermediate results and configuration parameters generated during the computing process as communication data. Subsequently, the computing unit encapsulates the communication address and communication data into a complete communication instruction and sends it to the communication module through the internal high-speed data bus.
[0135] After receiving a communication command through the bus interface unit, the communication module extracts the target accelerator identifier contained in the communication address of the command, queries the mapping table pre-stored in the local configuration register, and determines the buffer partition that matches the target accelerator. Subsequently, the communication module reads the base address of the buffer partition and obtains the address window corresponding to the base address.
[0136] The communication module compares the communication address with the upper and lower boundary values of the address window. If the communication address is determined to be within the address window, the write logic of the buffer partition is initiated. The buffer partition consists of multiple independent buffer units. The communication module matches the target buffer unit based on the address label of the communication address. After determining the target buffer unit, the starting position is located by combining the address offset of the communication address relative to the base address, so as to store the communication data in the data storage area of the target buffer unit. At the same time, the byte enable bit of the buffer unit is updated to mark the valid data range.
[0137] The sending conditions can be triggered in multiple ways, including when the total length of the communication data stored in the buffer partition reaches a preset threshold, when the remaining space recorded in the available payload length register is insufficient to accommodate the communication data in the next communication instruction, when the communication address exceeds the current address window, or when the timer times out.
[0138] When any transmission condition is met, the communication module traverses all buffer units in the buffer partition, extracts the communication data in each buffer unit, and synchronously reads the base address of the buffer partition in the base address register, so as to transmit the base address and multiple communication data together to the packet module.
[0139] After receiving data, the packet module parses the sub-headers (including address offset, payload length, etc.) of each communication data point and encapsulates them into a unified transaction layer data packet according to the transaction layer format of a preset communication protocol. During encapsulation, the packet module fills the address field of the transaction layer data packet with the received base address and adds the payload length of all communication data to the length field of the transaction layer data packet using a length statistics unit, forming a complete target transaction layer data packet. Finally, the packet module sends the target transaction layer data packet to the target accelerator through an interconnection channel adapted to the preset communication protocol.
[0140] The computing unit generates communication commands, the communication module receives and processes these commands, accurately locates the buffer partition and stores the communication data, and the grouping module packages and sends the data. This achieves efficient data transmission and management, effectively optimizes the communication process between accelerators, improves communication efficiency and reliability, and enhances the overall performance of the system.
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0142] Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0143] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
Claims
1. An accelerator, characterized in that, The accelerator includes: A computing unit is configured to generate communication instructions, the communication instructions including a communication address and communication data; The communication module is configured to receive communication instructions from the computing unit, determine a physically isolated buffer partition corresponding to the target accelerator based on the target accelerator pointed to by the communication address in the communication instructions, and obtain a base address for the buffer partition. The base address is obtained by aligning the received communication address after the buffer partition is initialized. The buffer partition includes multiple buffer units. If it is determined that the communication address is within the address window determined by the base address and the preset address offset, the communication data in the communication instruction is stored in the buffer partition; In response to the triggering of the transmission condition, the base address and multiple communication data and corresponding sub-headers stored in the buffer partition are sent to the packet module, wherein the multiple communication data include communication data in the communication instruction and historical communication data in the historical communication instruction pointing to the target accelerator; The communication module is further configured as follows: Based on the address offset of the communication address relative to the base address, the starting position of the communication data in the target buffer unit determined from the plurality of buffer units is determined; based on the starting position, the communication data is written into the data storage area of the target buffer unit; Based on the starting position and the data length of the communication data, determine the byte mask corresponding to the multiple byte positions in the target buffer unit; By performing a bitwise OR operation between the byte mask and the byte enable bits corresponding to each byte position in the target buffer unit, the corresponding byte enable bits are marked as valid; Based on the address offset, the address tag of the target buffer unit, and the payload length of the communication data in the target buffer unit (marked by the byte enable bit), a sub-header for the communication data is generated; the sub-header and the communication data are bound together and stored in the target buffer unit. The packet module is configured to package multiple received communication data based on the base address, generate a transaction layer data packet including the base address, the total effective payload length of the multiple communication data, and the sub-header of each communication data, and send it to the target accelerator through an interconnection channel corresponding to a preset communication protocol; If the valid bytes within a single buffer unit are not contiguous, the data in the buffer unit is split into multiple contiguous data blocks, and a corresponding sub-header is generated for each data block.
2. The accelerator according to claim 1, characterized in that, The communication module is also configured to: The address label of the communication address is matched with the address label of each of the buffer units to obtain the matching result; If the matching result indicates that there is an existing unit among the plurality of buffer units that corresponds to the address label of the communication address, the existing unit is identified as the target buffer unit, and the communication data is stored in the target buffer unit. If the matching result indicates that there is no existing unit corresponding to the address label of the communication address among the plurality of buffer units, a target buffer unit is determined from the free units of the plurality of buffer units to store the communication data in the target buffer unit.
3. The accelerator according to claim 1, characterized in that, The communication module further includes an available payload length register; wherein the available payload length register is initialized to a preset maximum payload length; The available payload length register is configured to: for the communication instruction, when the communication data has been stored in the buffer partition, determine the remaining payload length based on the data length of the communication data, the field length of the sub-header, and the current payload length; In response to receiving a next communication instruction, if it is determined that the remaining payload length is less than the sum of the data length of the communication data in the next communication instruction and the field length of the corresponding sub-header, the sending condition is triggered.
4. The accelerator according to claim 1 or 3, characterized in that, The communication module is also configured to: If it is determined that the communication address is outside the address window, the sending condition is triggered; The buffer partition is initialized to redetermine the base address and the corresponding address window based on the communication address in the next received communication instruction.
5. The accelerator according to claim 1, characterized in that, The grouping module is also configured to: In response to the received multiple communication data, for each buffer unit in the buffer partition, the communication data in each buffer unit and the corresponding sub-header are encapsulated into a transaction layer data packet; The base address is filled into the address field of the transaction layer data packet, and the total length of the payload of multiple communication data in the transaction layer data packet is filled into the length field of the transaction layer data packet to obtain the target transaction layer data packet; The target transaction layer data packet is sent to the target accelerator through an interconnection channel corresponding to a preset communication protocol.
6. The accelerator according to claim 5, characterized in that, The accelerator also includes an unpacking module; wherein... The unpacking module is configured to receive target transaction layer data packets via the interconnection channel; Parse the address field of the target transaction layer data packet to obtain the base address; Parse the length field of the target transaction layer data packet to obtain the total payload length of multiple communication data in the target transaction layer data packet, and extract each sub-header and corresponding communication data in sequence; For each of the sub-headers, the address offset of the corresponding communication address relative to the base address and the effective payload length of the corresponding communication data are decoded. Based on the address offset, the starting position of the communication data in multiple buffer units of the buffer partition determined by the base address is determined, and the communication data is written into the corresponding buffer unit according to the payload length. The communication data written into each of the buffer units is sent to the computing unit for processing.
7. A multi-accelerator system, characterized in that, The system includes: Multiple accelerators as described in any one of claims 1-6; An interconnection network connects multiple accelerators, wherein the interconnection network is adapted to a preset communication protocol.
8. A data transmission method, characterized in that, Applied to an accelerator, the accelerator including a computing unit, a communication module, and a packet module; the method includes: The computing unit generates communication instructions, which include communication addresses and communication data. The communication module receives a communication instruction from the computing unit, determines a physically isolated buffer partition corresponding to the target accelerator based on the target accelerator pointed to by the communication address in the communication instruction, and obtains a base address for the buffer partition. The base address is obtained by aligning the received communication address after the buffer partition is initialized. The buffer partition includes multiple buffer units. If it is determined that the communication address is within the address window determined by the base address and the preset address offset, the communication data in the communication instruction is stored in the buffer partition; In response to the triggering of the transmission condition, the base address and multiple communication data and corresponding sub-headers stored in the buffer partition are sent to the packet module, wherein the multiple communication data include communication data in the communication instruction and historical communication data in the historical communication instruction pointing to the target accelerator; The method further includes: the communication module determining the starting position of the communication data in a target buffer unit determined from the plurality of buffer units based on the address offset of the communication address relative to the base address; and writing the communication data into the data storage area of the target buffer unit based on the starting position. Based on the starting position and the data length of the communication data, determine the byte mask corresponding to the multiple byte positions in the target buffer unit; By performing a bitwise OR operation between the byte mask and the byte enable bits corresponding to each byte position in the target buffer unit, the corresponding byte enable bits are marked as valid; The communication module generates a sub-header for the communication data based on the address offset, the address tag of the target buffer unit, and the payload length of the communication data in the target buffer unit, which is indicated by the byte enable bit; the sub-header and the communication data are bound together and stored in the target buffer unit. The packet module packages the received multiple communication data based on the base address to generate a transaction layer data packet including the base address, the total effective payload length of the multiple communication data, and the sub-header of each communication data, and sends it to the target accelerator through an interconnection channel corresponding to a preset communication protocol; If the valid bytes within a single buffer unit are not contiguous, the data in the buffer unit is split into multiple contiguous data blocks, and a corresponding sub-header is generated for each data block.
Citation Information
Patent Citations
Unified address space for multiple hardware accelerators using dedicated low latency links
CN112543925A
Data migration method, data migration processing device, chip and electronic equipment
CN120596406A
Data packet receiving method, data processing unit, host and network card
CN120710964A