Inter-chip interconnection device and electronic equipment
By setting up data reuse and request merging modules in the inter-chip interconnection device, the problem of limited data bandwidth utilization in inter-chip communication is solved, and more efficient data transmission and communication performance improvement are achieved.
Patent Information
- Application Number
- CN202510535244.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The performance of inter-chip communication is limited by data bandwidth utilization. Existing technologies have limited improvements by expanding port width and frequency, resulting in increased communication overhead and reduced efficiency.
A data reuse module and a request merging module are set up in the inter-chip interconnection device. The data reuse module determines the destination address of the request to be issued and compares it with the address of the historical request to determine the request to be merged. The request merging module merges the received requests to be merged to reduce unnecessary repeated requests and communication overhead.
Optimize communication efficiency, reduce communication overhead of inter-chip interconnection, and improve data transmission efficiency and bandwidth utilization.
Smart Images

Figure CN120067035B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to, but is not limited to, the field of circuit technology, and in particular to inter-chip interconnection devices and electronic devices. Background Art
[0002] As chip sizes continue to increase and emerging applications demand ever-increasing chip computing power, the need for inter-chip interconnect communication has arisen. Inter-chip communication performance is typically achieved through wider data bit widths, higher clock frequencies, and longer data packets specified in the consistency protocol. However, this information cannot be increased indefinitely, limiting data bandwidth utilization and thus hindering inter-chip communication performance. Summary of the Invention
[0003] In view of this, embodiments of the present application at least provide an inter-chip interconnection device and an electronic device, which can improve inter-chip communication performance.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] On the one hand, an embodiment of the present application provides an inter-chip interconnection device, which is arranged between a first processor and an inter-chip interconnection logic. The inter-chip interconnection device includes a data reuse module and a request merging module, wherein the data reuse module is used to receive a request to be issued and determine a request to be merged based on the destination address of the request to be issued; the data reuse module is also used to send the request to be merged to the request merging module; the request merging module is used to merge the received requests to be merged to obtain a merged request, and send the merged request to the inter-chip interconnection logic; the inter-chip interconnection logic is used to send the merged request to the second processor.
[0006] On the other hand, an embodiment of the present application provides an electronic device, including multiple processors and an inter-chip interconnection device provided in the above embodiment corresponding one-to-one to the multiple processors.
[0007] In an embodiment of the present application, a data reuse module is set in the inter-chip interconnection device to receive requests to be issued, and the destination address of the requests to be issued is compared with the address of the requests issued in the past to determine which requests to be issued need to be passed to the request merging module as requests to be merged and which requests to be issued need to be ignored, thereby reducing unnecessary repeated requests and optimizing communication efficiency; at the same time, the requests to be merged are sent to the request merging module through the data reuse module, so that the request merging module can merge the received requests to be merged to obtain merged requests, thereby reducing the number of requests that the inter-chip interconnection logic needs to process and reducing communication overhead; based on the embodiments provided in the present application, the communication overhead of the inter-chip interconnection can be reduced to a certain extent, and the data transmission efficiency of the inter-chip interconnection can be improved.
[0008] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.
[0010] Figure 1 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 1 ;
[0011] Figure 2 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 2 ;
[0012] Figure 3 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 3 ;
[0013] Figure 4 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 4 ;
[0014] Figure 5 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 5 ;
[0015] Figure 6 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 6 ;
[0016] Figure 7 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 7 ;
[0017] Figure 8 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 8 ;
[0018] Figure 9 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 9 ;
[0019] Figure 10 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 10 ;
[0020] Figure 11A schematic diagram of an interconnection structure in an inter-GPU interconnection scenario provided in an embodiment of the present application;
[0021] Figure 12 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 10 one;
[0022] Figure 13 A hardware entity diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0024] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or sequence of "first / second / third" may be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.
[0026] To facilitate understanding of this solution, before describing the embodiments of the present application, the application background of the embodiments of the present application will be described.
[0027] Inter-chip communication includes communication between multiple dies (die-to-die, D2D) packaged on the same chip and communication between multiple boards (chip-to-chip C2C). Because the distance between dies is greater than the physical connection within a die, communication performance is affected when accessing one die from another via D2D. Similarly, because the distance between chips is greater than the physical connection within a chip, communication performance is also affected when accessing one chip from another via C2C. To improve inter-chip communication performance, related technologies typically expand port width and frequency to provide greater inter-chip interconnect bandwidth. They also support coherence protocols through more complex logic to support longer data packets, improve transmission efficiency, and reduce the ratio of instruction information to actual data in a data packet. However, if the amount of data transmitted between chips does not decrease, the limited inter-chip interconnect bandwidth may require more complex instruction information to maintain consistent communication, resulting in additional communication overhead. In addition, the length of the data packets supported by the consistency protocol cannot be increased indefinitely because it is limited by the actual transmission capability of the graphical processing unit (GPU) and the data interleaving granularity when the GPU transmits across multiple ports. This limits the overall data bandwidth utilization.
[0028] To address the aforementioned issues, embodiments of the present application provide an inter-chip interconnect device, which is disposed between a processor and inter-chip interconnect logic. When a processor transmits data, the data must be transmitted via the inter-chip interconnect device to the inter-chip interconnect logic, and then transmitted to other processors via the inter-chip interconnect logic. Because the inter-chip interconnect device consolidates requests, the number of transmissions is reduced, thereby improving data bandwidth utilization and, in turn, enhancing inter-chip communication performance.
[0029] Figure 1 A schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application Figure 1 ,like Figure 1 As shown, the inter-chip interconnect device is arranged between the first processor 10 and the inter-chip interconnect logic 30. The inter-chip interconnect device 20 includes a data reuse module 21 and a request merging module 22, wherein the data reuse module 21 is used to receive a request to be issued and determine a request to be merged based on the destination address of the request to be issued; the data reuse module 21 is also used to send the request to be merged to the request merging module 22; the request merging module 22 is used to merge the received requests to be merged to obtain a merged request, and send the merged request to the inter-chip interconnect logic 30; the inter-chip interconnect logic 30 is used to send the merged request to the second processor 40.
[0030] Among them, the above-mentioned inter-chip interconnect device is a dedicated hardware module arranged between the processor and the inter-chip interconnect logic, which is used to optimize the data transmission process between processors, by merging requests and reducing the number of transmissions, improving data bandwidth utilization, and thus improving inter-chip communication performance; among them, the inter-chip interconnect logic is the hardware (or software component) responsible for transmitting data between processors. The inter-chip interconnect logic is used to receive merged requests from the inter-chip interconnect device and forward these merged requests to the target processor (such as the second processor) to realize data communication between processors.
[0031] In some embodiments, the above-mentioned data reuse module is used to receive pending requests issued by the first processor 10, and compare and judge based on the destination address of the pending requests and the addresses of the requests issued historically, to determine which pending requests need to be passed to the request merging module as pending requests, and which pending requests need to be ignored (that is, they have already been requested and do not need to be sent to the next unit), thereby reducing unnecessary repeated requests and optimizing communication efficiency.
[0032] In some embodiments, the request merging module is configured to receive pending merge requests from the data reuse module and merge these pending merge requests to generate merged requests. The purpose of merging requests is to reduce the number of requests, reduce the burden on inter-chip interconnect logic, and improve overall data transmission efficiency.
[0033] In some possible implementations, the data reuse module receives pending requests from the first processor and records the pending requests; the data reuse module compares the destination addresses of the pending requests with the recorded addresses to determine which requests need to be passed to the request merging module as pending requests and which requests need to be ignored (i.e., the corresponding destination addresses have been requested and do not need to be sent to the request merging module). Then, the data reuse module sends the pending requests to the request merging module. The request merging module merges the received pending requests to generate a merged request; finally, the request merging module sends the merged request to the inter-chip interconnect logic, which forwards the request to the second processor.
[0034] Based on the above-mentioned embodiments disclosed in the present application, a data reuse module is set in the inter-chip interconnection device to receive requests to be issued, and the destination address of the requests to be issued is compared with the address of the requests issued in the past to determine which requests to be issued need to be passed to the request merging module as requests to be merged, and which requests to be issued need to be ignored, thereby reducing unnecessary repeated requests and optimizing communication efficiency; at the same time, the requests to be merged are sent to the request merging module through the data reuse module, so that the request merging module can merge the received requests to be merged to obtain merged requests, thereby reducing the number of requests that the inter-chip interconnection logic needs to process and reducing communication overhead; based on the embodiments provided in the present application, the communication overhead of the inter-chip interconnection can be reduced to a certain extent, and the data transmission efficiency of the inter-chip interconnection can be improved.
[0035] Figure 2 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 2 .based on Figure 1 The data reuse module 21 includes a waiting request recording unit 211, which is used to store historical requests in the inter-chip communication process and the historical addresses of the historical requests; wherein, the data reuse module 21 is also used to determine the requests to be merged based on the destination address of the request to be issued and the historical address of the historical request; the data reuse module 21 is also used to update the waiting request recording unit 211 based on the request to be issued.
[0036] The inter-chip communication process refers to all processes from when a request is received from the inter-chip interconnection device until the inter-chip interconnection device returns a response corresponding to the request.
[0037] The waiting request recording unit 211 is used to store historical requests and their corresponding historical addresses in the inter-chip communication process.
[0038] In some embodiments, the destination address of the request to be merged does not overlap with the address range of the historical address of the previous request. In other words, the request to be merged is a request that, after comparing the destination address of the request to be sent with the historical address of the previous request, is determined not to cause duplication and does not overlap with the address range. The request to be merged is passed to the request merging module to reduce communication times and optimize communication efficiency.
[0039] In some possible implementations, when the inter-chip interconnect device receives a pending request, the data reuse module 21 compares its destination address with the historical addresses in the pending request recording unit 211. Based on the comparison result, it is determined whether the request will cause duplicate transmission, that is, whether the address ranges do not overlap. If it will not cause duplication, it is determined as a pending request to be merged and is prepared for transmission to the request merging module. After completing the above determination, the pending request recording unit 211 records the pending request and its corresponding destination address as a new historical request and corresponding historical address, thereby updating the pending request recording unit 211.
[0040] For example, assume that there are three pending requests: request 1, request 2, and request 3. The destination address range of request 1 overlaps with a previous historical request, so request 1 is determined to be a duplicate request and is ignored. The destination address ranges of request 2 and request 3 do not overlap with historical requests, so request 2 and request 3 are determined to be requests to be merged and are passed to the request merging module for merging processing. After completing the judgment, the waiting request recording unit 211 records request 1, request 2, and request 3 and their corresponding destination addresses as new historical requests and corresponding historical addresses to update the waiting request recording unit 211.
[0041] Based on the above-mentioned embodiments disclosed in the present application, by determining the requests to be merged based on the destination address of the request to be issued and the historical address of the historical request, it is possible to filter requests with overlapping addresses, reduce unnecessary communication overhead, and improve communication efficiency; in addition, by updating the waiting request recording unit 211 based on the request to be issued, the real-time and accuracy of the waiting request recording unit 211 can be maintained, and the latest data support can be provided for subsequent request processing; based on the embodiments provided in the present application, the efficiency of inter-chip communication can be effectively improved and the communication overhead can be reduced.
[0042] In some embodiments, the data reuse module 21 is also used to determine the request to be issued as the request to be merged when there is no target historical request whose address range overlaps with the request to be issued; the data reuse module 21 is also used to write the request to be issued as a historical request into the waiting request recording unit 211.
[0043] In some possible implementations, when the data reuse module 21 receives a pending request, it reads all historical requests and their corresponding historical addresses from the pending request recording unit 211. It compares the destination address of the pending request with the historical address of each historical request to determine whether the address ranges overlap. If the address range of the pending request does not overlap with the address range of any historical request, the pending request is determined to be a pending request to be merged. Accordingly, the pending request recording unit 211 treats the pending request as a new historical request and its destination address as the corresponding historical address. The new historical request and historical address are then written into the pending request recording unit 211.
[0044] For example, assume that there are three pending requests: request 1, request 2, and request 3. Among them, the destination address range of request 1 is 0x1000-0x1FFF, and there is a historical request in the waiting request recording unit 211, whose address range is 0x2000-0x2FFF. Since the address range of request 1 does not overlap with the address range of the historical requests, request 1 is determined to be a request to be merged; the destination address range of request 2 is 0x2000-0x2FFF. Since the address range of request 2 overlaps with the address range of the historical requests, request 2 is not determined to be a request to be merged. The destination address range of request 3 is 0x3000-0x3FFF. Since the address range of request 3 does not overlap with the address range of any historical requests, request 3 is determined to be a request to be merged. Accordingly, after the above judgment is completed, the waiting request recording unit 211 is updated, and requests 1 to 3 are treated as new historical requests, and the corresponding destination addresses are written into the waiting request recording unit 211 as new historical addresses.
[0045] Based on the above-mentioned embodiments disclosed in this application, by comparing the destination address of the request to be issued with the historical address of the historical request, the request to be merged is determined, which can effectively reduce the sending of duplicate requests and reduce communication overhead. At the same time, writing the request to be issued as a historical request into the waiting request recording unit 211 can maintain the real-time and accuracy of the recording unit and provide the latest data support for subsequent request processing. Based on the embodiments provided in this application, the efficiency of inter-chip communication can be significantly improved, the network load can be reduced, and the overall performance can be improved.
[0046] Figure 3 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 3 .based on Figure 2, the data reuse module 21 includes a write data cache 212; the data reuse module 21 is also used to, when the request type of the request to be issued is a write request, allocate a new write data address for the request to be issued in the write data cache 212; based on the new write data address, write the write data corresponding to the request to be issued into the write data cache 212; and write the new write data address into the waiting request recording unit 211.
[0047] In some possible implementations, after receiving the request to be issued, the data reuse module 21 can parse the request type information in the request to determine whether the request to be issued is a read request or a write request; if the request to be issued is a write request, the data reuse module 21 searches for free storage space in the write data cache 212 and allocates a new write data address for the request; the write data corresponding to the request to be issued is written into the write data cache 212 according to the allocated write data address; accordingly, the allocated new write data address is written into the waiting request recording unit 211 and stored together with other information of the write request (such as request type, destination address, etc.) for subsequent processing and tracking of the write request.
[0048] For example, assume that there is a pending request with a write request type, a destination address of 0x2000, and a write data length of 128 bytes. The data reuse module 21 parses the request information of the pending request and determines that the pending request has a write request type. The module searches for free space in the write data cache 212 and allocates a new write data address, such as 0x3000, for the write request. The module writes the 128-byte write data to the write data cache 212 at the address 0x3000. The module writes the write data address 0x3000 to the waiting request recording unit 211 and stores the address together with the write request's destination address 0x2000, the request type, and other information.
[0049] Based on the above-mentioned embodiments disclosed in the present application, by setting a write data cache 212 in the data reuse module 21, a dedicated storage space can be provided for write data, so that the write data has a reasonable storage location before transmission, and centralized storage of write data can be achieved, which facilitates subsequent management and operation of write data. For example, these write data can be more conveniently found and processed when needed, thereby improving data processing efficiency; and the new write data address is written into the waiting request recording unit 211, which provides key information for subsequent tracking and processing of write requests, and can clearly understand the write data storage location corresponding to each write request, which helps to perform corresponding processing when the write data transmission is completed or an exception occurs.
[0050] In some embodiments, the data reuse module 21 is also used to ignore the request to be issued when there is a target historical request whose address range overlaps with the request to be issued; the data reuse module 21 is also used to write the request to be issued as a historical request into the waiting request recording unit 211, and establish an association relationship between the historical request corresponding to the request to be issued and the target historical request.
[0051] In the above embodiment, if the address ranges of multiple requests overlap, processing these requests simultaneously may lead to data conflicts, resource waste, and other issues. Therefore, upon receiving a pending request, the data reuse module 21 checks the pending request record unit 211 to see if there is a target historical request with an address range that overlaps with the pending request. If so, the pending request is ignored to avoid duplicate processing.
[0052] For example, there is a request to be issued, whose request type is a write request, the destination address range is 0x2000-0x2FFF, and the write data length is 4KB. A historical request has been recorded in the waiting request recording unit 211, and its address range is 0x2500-0x3000. The data reuse module 21 receives the write request and parses the address range as 0x2000-0x2FFF; compares the address range with the historical request address range in the waiting request recording unit 211; finds that there is a historical request with an address range of 0x2500-0x3000 that overlaps with the address range of the request to be issued 0x2000-0x2FFF; the data reuse module 21 ignores the write request and does not pass it to the next processing unit.
[0053] In some embodiments, by ignoring the request to be issued when there is a target historical request whose address range overlaps with the request to be issued, repeated processing of data in the same address range can be avoided, and unnecessary resource consumption can be reduced. For example, the processor's parsing and execution time for repeated requests can be reduced, and the memory bandwidth occupancy can be reduced. At the same time, by writing the request to be issued as a historical request into the waiting request recording unit 211, and establishing an association relationship between the historical request corresponding to the request to be issued and the target historical request, the subsequent management and tracking of the request can be facilitated, and in subsequent processing, related requests can be uniformly processed according to the association relationship, thereby improving the efficiency of request processing. Based on the embodiments provided in the present application, the request processing flow in the inter-chip communication process can be optimized, the complexity of request processing can be reduced, and the efficiency and overall performance of inter-chip communication can be improved.
[0054] In some embodiments, the data reuse module 21 is also used to obtain a target write data address when the request type of the request to be issued is a write request; based on the target write data address, write the write data corresponding to the request to be issued into the write data cache 212; the target write data address is the write data address of the earliest historical request whose address range overlaps with the request to be issued.
[0055] Among them, when there is a pending request whose request type is a write request, and there is a target historical request whose address range overlaps with the pending request, it means that before the current pending request, there is at least one historical request whose historical address overlaps with the destination address of the pending request. Therefore, when writing data to the same destination address, the latest write data needs to be written to the target address, that is, the write data corresponding to the current pending request needs to be written to the target address.
[0056] In some embodiments, there may be at least one historical request whose address range overlaps with the destination address of the request to be issued. In this case, it is necessary to find the earliest historical request among the at least one historical request and obtain the write data address corresponding to the historical request as the target write data address. It is understandable that when there are at least two historical requests whose address range overlaps with the destination address of the request to be issued, the write data stored in the target write data address is no longer the write data corresponding to the earliest historical request, but the write data corresponding to the latest historical request among the at least two historical requests. Of course, after the write data of the current request to be issued is written to the target write data address, the write data stored in the target write data address is also the write data of the latest request among these historical requests.
[0057] In the above embodiment, in order to solve the problem of how to ensure that the data is correctly written to the target address in the write request and keep the data up to date and consistent. When there is a write request to be issued and the address range of the request overlaps with the previous historical request, it is necessary to find the earliest one of these historical requests and obtain its write data address as the target write data address. In this way, when writing the data of the current request to be issued, it can be ensured that the data is overwritten to the correct location and is the latest write data. In this way, after the write request is issued, the latest write data can be found based on the target write data address, thereby avoiding data conflicts and inconsistencies and improving the efficiency and accuracy of data processing.
[0058] In some possible implementations, when the data reuse module receives a write request to be issued, it checks whether the address range of the request overlaps with previous historical requests. If there is an overlap, the data reuse module will further search for the earliest of these historical requests and obtain its write data address as the target write data address. Then, the data reuse module will write the write data corresponding to the request to be issued into the write data cache. Finally, the data reuse module will formally write the write data in the cache to the target write data address, thereby completing the data write operation. The entire process needs to ensure the accuracy and consistency of the data to avoid data loss or errors.
[0059] For example, it is assumed that three historical requests are recorded in the waiting request recording unit 211, namely historical request A, historical request B and historical request C. Among them: Historical request A: the request type is a write request, the destination address range is 0x1000-0x2000, the write data length is 4KB, and it is the first write request to arrive. From the perspective of logical processing, its corresponding initial target write data address (used for subsequent overwrite logic tracking) is set to Addr_A (actually corresponding to a physical or logical address space that can store 4KB of data). Historical request B: the request type is a write request, the destination address range is 0x1000-0x2000, the write data length is 4KB, and it arrives after historical request A. Historical request C: the request type is a write request, the destination address range is 0x1000-0x2000, the write data length is 4KB, and it arrives after historical request B.
[0060] At this point, the first processor issues a new pending request D, whose request type is a write request, whose destination address range is 0x1000 - 0x2000, whose write data length is 4KB, and whose write data content is different from (or may be the same as) the historical requests A, B, and C. After receiving the pending request D, the data reuse module 21 first determines that its request type is a write request. It then compares the address range of the pending request D with the address ranges of all historical requests in the pending request record unit 211. This comparison reveals that the address ranges of the historical requests A, B, and C all overlap with the address range of the pending request D.
[0061] The update process of the target write data address for the above historical requests A to C can be understood as follows: according to the order of arrival time of the requests, the data reuse module 21 will process these write requests in sequence. Since historical request A is the first to arrive, its write data will first be written to the storage location corresponding to the target write data address Addr_A. Then, the write data of historical request B will overwrite the data of historical request A in the address range of Addr_A (corresponding to the overlapping part of 0x1000 - 0x2000). At this time, although it is still logically tracked with Addr_A, the actual storage content of the target write data address has been updated to the data of historical request B. Finally, the write data of historical request C will overwrite the data of historical request B in the address range of Addr_A.
[0062] When a pending request D arrives, the data reuse module 21 will also write its write data to the storage location corresponding to the target write data address Addr_A in chronological order, overwriting the data of the previous historical request C within the address range of Addr_A. Therefore, the storage location corresponding to Addr_A ultimately stores the data of the pending request D.
[0063] Based on the above-mentioned embodiment disclosed in the present application, based on the obtained target write data address, the write data corresponding to the request to be issued is written into the write data cache 212, and the cache is used to temporarily store the write data. On the one hand, it can alleviate the pressure that may be caused by directly writing data to the target storage location, and on the other hand, it also provides convenience for subsequent unified data processing. In addition, the target write data address is set to the write data address of the earliest historical request whose address range overlaps with the request to be issued. When processing the write request, the overwriting and writing of data can be considered in a relatively orderly manner, logically laying the foundation for subsequent data overwriting operations. Based on the embodiment provided by the present application, in complex scenarios where there are multiple write request address ranges overlapping, the writing and overwriting problems of write data can be handled relatively reasonably, the processing flow of write data can be optimized to a certain extent, and the efficiency of data processing of the inter-chip interconnect device can be improved.
[0064] In some embodiments, the data reuse module 21 is further configured to sort the unissued requests to be merged to obtain an issuance order; and send the unissued requests to be merged to the request merging module 22 in sequence according to the issuance order.
[0065] The aforementioned dispatch order is the order in which requests are sent, determined by the data reuse module after sorting the requests to be merged. In some embodiments, this dispatch order can be determined based on a variety of factors, such as request priority, overlap in address ranges, and the time sequence of request arrival. By properly arranging the dispatch order, conflicts between requests can be avoided, improving the efficiency of request processing.
[0066] In some possible implementations, after obtaining the unissued requests to be merged, the data reuse module can store these requests to be merged in an internal request queue; analyze the requests to be merged in the request queue, extract the request attributes of the requests to be merged, sort the requests to be merged in the request queue according to the request attributes of the requests to be merged, and obtain the issuance order; take out the requests to be merged from the request queue in sequence according to the issuance order, and send them to the request merging module.
[0067] Based on the above embodiments disclosed in the present application, unissued requests to be merged are sent to the request merging module 22 in the order of issuance, which can avoid conflicts between requests, reduce waiting and confusion during request processing, and improve the processing efficiency of the request merging module for requests.
[0068] In some embodiments, the issuance order is determined based on the request attributes of the unissued request to be merged; the request attributes include at least one of the following: the time of receiving the unissued request to be merged, the address interval between the destination address of the unissued request to be merged and the last issued request to be merged, and the number of historical requests associated with the unissued request to be merged.
[0069] The time of receiving the unsent request to be merged refers to the specific moment when the data reuse module receives a certain unsent request to be merged. Based on this time, the order of issuing requests can be reasonably arranged to avoid long waiting times for requests.
[0070] For example, there are three unsent requests A, B, and C to be merged. Request A is received by the data reuse module at 10:00:00 AM, request B is received at 10:00:01 AM, and request C is received at 10:00:02 AM. When determining the order in which requests are sent, if the time of receipt is prioritized, request A will be sent to the request merging module first.
[0071] Among them, each unissued request to be merged has a corresponding destination address, which is used to specify the location of the data or resources to be read / written by the request to be merged. The above-mentioned address interval refers to the difference between the destination address of the currently unissued request to be merged and the destination address of the last request to be merged that has been issued. If the interval between the destination addresses of two requests to be merged is small, it means that the data or resource locations accessed by the two requests to be merged are closer, then the unissued request to be merged with the smaller interval will be sent to the request merging module address first, which can facilitate the subsequent request merging module to merge the requests to be merged.
[0072] For example, assume the data reuse module has already sent a pending request to the request merging module, with a destination address of 0x1000. There are now three pending requests, A, B, and C, with destination addresses of 0x1050, 0x2000, and 0x2050, respectively. The address interval between request A's destination address 0x1050 and the previously issued request's destination address 0x1000 is 50; the address interval between request B's destination address 0x2000 and the previously issued request's destination address 0x1000 is 1000; and the address interval between request C's destination address 0x2050 and the previously issued request's destination address 0x1000 is 1050. When determining the order of request issuance, if address intervals are prioritized, request C will be processed first because it has the smallest address interval and is closest to the destination address of the previous request. This makes it easier for the subsequent request merging module to merge request C with the previous request (or subsequent adjacent requests), reducing overhead. Request A is then processed, followed by request B.
[0073] The number of historical requests associated with the unissued pending request refers to the number of historical requests directly associated with the pending request, as determined by a mechanism (e.g., a data reuse module ignoring the pending request and writing it as a historical request into the pending request record unit while simultaneously establishing an association with the target historical request) when a target historical request with an address range that overlaps with the pending request exists. When the unissued pending request has associations with multiple historical requests, such as overlapping address ranges, the number of associated historical requests is the number of these overlapping historical requests. Considering the number of associated historical requests can help understand the complexity of the request's association within the set of historical requests, facilitating the proper ordering of request processing. Generally speaking, a large number of associated historical requests may indicate that the destination address of the pending request has been read / written multiple times. Accordingly, the method further includes, if the pending request is a write request, prioritizing the issuance of the pending request with a smaller number of requests; and if the pending request is a read request, prioritizing the issuance of the pending request with a larger number of requests.
[0074] For example, there are three unissued requests to be merged, A, B, and C. Among them, request A to be merged has an address range overlapping association with a historical request, and the number of associated historical requests is 1; request B to be merged has an address range overlapping association with two historical requests, and the number of associated historical requests is 2; request C to be merged has an address range overlapping association with three historical requests, and the number of associated historical requests is 3. When the request to be merged is a write request, the request to be merged with a smaller number is issued first. Therefore, write request A will be processed first because it has the smallest number of associated historical requests, which can reduce frequent modifications and potential conflicts to the data and ensure data consistency and integrity. When the request to be merged is a read request, a larger number of requests to be merged will be issued first. Therefore, read request C will be processed first because it has the largest number of associated historical requests, which can read more relevant data in one operation and improve the efficiency of data reading.
[0075] Of course, it is also possible to combine two or more of the above request attributes to make a comprehensive judgment on each unissued request to be merged, and determine the issuance order of each unissued request to be merged.
[0076] Based on the above-mentioned embodiments disclosed in the present application, by determining the issuance order based on the request attribute of the time when the unissued request to be merged is received, the order of requests can be taken into account, so that requests received earlier have a greater probability of being processed first, which helps to operate in the logical order of request arrival and avoid long-term backlog of requests; at the same time, by determining the issuance order based on the request attribute of the address interval between the destination address of the unissued request to be merged and the last issued request to be merged, the distribution of requests in the address space can be taken into account. When the address interval is small, it means that the two requests may be in a similar address area. Giving priority to such requests may help improve the efficiency of subsequent request merging modules in merging requests; in addition, by determining the issuance order based on the request attribute of the number of historical requests associated with the unissued request to be merged, the frequency of reading / writing the destination address corresponding to the request to be merged can be understood, and then the order of priority can be determined based on the request type, which can resolve potential conflicts or dependency problems between requests; based on the embodiments provided in the present application, multiple request attributes can be comprehensively considered to determine the issuance order of requests to be merged, so as to arrange request processing more flexibly and reasonably, and improve overall performance and resource utilization.
[0077] Figure 4 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 4 .based on Figure 1 , the request merging module 22 includes a merging waiting unit 221 and a merging processing unit 222;
[0078] The merging waiting unit 221 is configured to receive and store the request to be merged;
[0079] The merging processing unit 222 is configured to merge the requests to be merged according to the destination addresses of the requests to be merged to obtain the merged requests.
[0080] The merged request is a result obtained after being processed by the merge processing unit, and is a new request formed by merging multiple requests to be merged.
[0081] In the above embodiment, a request merging module is provided, comprising a merging waiting unit and a merging processing unit, to manage received requests to be merged. The merging waiting unit is responsible for receiving and storing requests to be merged, providing a buffer area for the requests; the merging processing unit, based on the destination addresses of the requests to be merged, merges requests with the same or similar destination addresses into a single merged request, thereby reducing resource usage and improving request processing efficiency.
[0082] In some possible implementations, when a request to be merged is received, the request to be merged is first sent to a merge waiting unit. The merge waiting unit assigns a unique identifier to each request to be merged and stores it in an internal storage structure, such as a queue or a list; the merge waiting unit sorts the requests to be merged according to certain rules (such as the time order of request arrival) for subsequent merge processing; the merge processing unit obtains the requests to be merged from the merge waiting unit regularly or according to certain trigger conditions (such as the number of requests to be merged in the merge waiting unit reaches a certain threshold, or every preset time interval); then, the merge processing unit analyzes and compares the destination addresses of these requests to be merged, and merges the requests to be merged with consecutive destination addresses into one merged request. In some embodiments, during the merging process, the merge processing unit records the relevant information of each request to be merged and integrates it into the merged request.
[0083] For example, assume that there are three requests to be merged: request A, request B, and request C. Request A: destination address is 0x1000, data length is 128 bytes. Request B: destination address is 0x1080, data length is 64 bytes. Request C: destination address is 0x1100, data length is 32 bytes. The merge waiting unit receives these three requests and stores them in the internal request queue. When the merge processing unit starts processing, it finds that the destination addresses of request A and request B are adjacent (and the data lengths are compatible), so request A and request B are merged into a new merged request. The destination address of the new merged request is 0x1000, and the data length is 192 bytes (128 bytes + 64 bytes). Request C remains as a separate request because its destination address is not adjacent to the merged request.
[0084] Based on the above-described embodiments disclosed in this application, by receiving and storing pending requests through the merge waiting unit 221, a certain number of requests can be accumulated, providing a sufficient data source for subsequent merge processing, thereby improving the flexibility and efficiency of the merge processing. At the same time, the merge processing unit 222 performs merge processing based on the destination addresses of the pending requests, and can merge requests with consecutive destination addresses into a single merged request, thereby effectively reducing the number of requests in inter-chip interconnect communication, reducing communication overhead, and improving data transmission efficiency and bandwidth utilization.
[0085] Figure 5 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 5 .based on Figure 4 The merge processing unit 222 includes an address comparison unit 2221 and an address calculation unit 2222, wherein the address comparison unit 2221 is configured to compare destination addresses of the requests to be merged stored in the merge waiting unit 221, divide the requests to be merged stored in the merge waiting unit 221 into at least one request group to be merged, and send the request group to the address calculation unit 2222;
[0086] The address calculation unit 2222 is configured to determine request information of the merged request based on the requests to be merged in the group of requests to be merged, and generate the merged request based on the request information.
[0087] In some possible implementations, the address comparison unit compares the destination address of each pending request stored in the merge waiting unit, groups pending requests with consecutive destination addresses into the same pending request group; and sends the divided pending request group to the address calculation unit for subsequent processing. The address calculation unit receives the pending request group from the address comparison unit. Based on the pending requests in the group, the request information of the merged request, including the destination address and data length, is determined; based on this request information, a merged request is generated, and the merged request is sent to the inter-chip interconnect logic for transmission.
[0088] In some other possible implementations, the address comparison unit receives pending requests from the merge waiting unit, compares the destination address of each pending request, and groups requests that can be merged into a first request group, while grouping requests that cannot be merged into a second request group. The address calculation unit receives the first and second request groups from the address comparison unit, and for the first request group, determines request information of the merged request, including the destination address and data length, based on the pending requests within the group. For the second request group, it directly treats one of the pending requests as the merged request.
[0089] In some embodiments, when the packet length is not restricted, all requests to be merged with continuous destination addresses are divided into the same request group to be merged to form the longest continuous request; when the packet length is restricted, according to the packet length restriction, the requests to be merged that can be merged into the maximum packet length are divided into the same request group to be merged to obtain the most requests with the maximum packet length.
[0090] For example, assume that in the inter-chip interconnection communication, there are the following five requests to be merged: Request A: destination address is 0x1000, data length is 256 bytes; Request B: destination address is 0x1100, data length is 256 bytes; Request C: destination address is 0x1400, data length is 64 bytes; Request D: destination address is 0x1200, data length is 256 bytes; Request E: destination address is 0x1300, data length is 64 bytes.
[0091] In some embodiments, the request information includes a request address and a data length; the request address is the first logically preceding destination address among the destination addresses of the requests to be merged in the request group to be merged; the data length of the merged request is the sum of the data lengths of the requests to be merged in the request group to be merged.
[0092] Without a packet length limit, the address comparison unit groups requests A, B, D, and E into the same request group to be merged (the first request group), and requests C into a request group to be merged (the second request group). Accordingly, the destination address of the first request group to be merged is 0x1000, and the data length is 832 bytes; the destination address of the second request group to be merged is 0x1400, and the data length is 64 bytes.
[0093] When the packet length is limited to 512 bytes, the address comparison unit groups requests A and B into the same request group to be merged (the first request group), requests D and E into the same request group to be merged (the first request group), and request C into a request group to be merged (the second request group). Accordingly, the destination address of the first request group to be merged is 0x1000, and the data length is 512 bytes; the destination address of the second request group to be merged is 0x1400, and the data length is 64 bytes; and the destination address of the third request group to be merged is 0x1300, and the data length is 320 bytes.
[0094] Based on the above-mentioned embodiment disclosed in the present application, the address comparison unit 2221 compares the destination addresses of the requests to be merged stored in the merge waiting unit 221, and divides the requests to be merged into at least one request group to be merged. In this way, requests with continuous destination addresses can be effectively grouped together, providing a basis for subsequent merge processing and helping to improve the efficiency of merge processing. At the same time, the address calculation unit 2222 determines the request information of the merged request based on the requests to be merged in the request group to be merged, and generates a merged request, which can further reduce the number of requests that need to be transmitted and reduce communication overhead.
[0095] Figure 6 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 6 .based on Figure 5 The merging processing unit 222 further includes a writing data processing unit 2223, wherein:
[0096] The write data processing unit 2223 is configured to sequentially retrieve the write data corresponding to each of the requests to be merged in the group of requests to be merged from the data reuse module 21 based on the address sequence and write data addresses of each of the requests to be merged in the group of requests to be merged;
[0097] The address calculation unit 2222 is further configured to construct the merged request based on the request information and the write data corresponding to each of the requests to be merged in the group of requests to be merged.
[0098] Among them, the above-mentioned write data processing unit can sequentially retrieve the corresponding write data from the data reuse module according to the address sequence and write data address of each request to be merged in the request group to be merged, and provide necessary data support for the subsequent construction of the merged request.
[0099] In some possible implementations, the write data processing unit receives a group of pending requests from the address comparison unit, sequentially retrieves the corresponding write data from the write data cache 212 in the data reuse module based on the address order and write data address of each pending request in the pending request group, and passes the retrieved write data to the address calculation unit to construct a merged request. The address calculation unit receives information about the pending request group from the address comparison unit and the write data corresponding to each pending request from the write data processing unit. Based on information about the pending request group (such as the destination address, data length, etc.) and the write data, it constructs request information for the merged request. The request information is combined with the write data to generate a complete merged request.
[0100] For example, assume that in inter-chip interconnect communication, there are the following five requests to be merged: Request A: The destination address is 0x1000, the data length is 256 bytes, and the write data is stored in the write data cache 212 in the data reuse module at a starting address of 0x0000; Request B: The destination address is 0x1100, the data length is 256 bytes, and the write data is stored in the write data cache 212 at a starting address of 0x0100. Request C: The destination address is 0x1400, the data length is 64 bytes, and the write data is stored in the write data cache 212 at a starting address of 0x0200. Request D: The destination address is 0x1200, the data length is 256 bytes, and the write data is stored in the write data cache 212 at a starting address of 0x0240. Request E: The destination address is 0x1300, the data length is 64 bytes, and the write data is stored in the write data cache 212 at a starting address of 0x0340.
[0101] Without limiting the packet length, the write data processing unit sequentially retrieves the write data corresponding to requests A, B, D, and E. Starting at address 0x0000, it continuously reads 832 bytes of data (256 bytes + 256 bytes + 256 bytes + 64 bytes). The address calculation unit constructs merged request 1 based on the information from requests A, B, D, and E (destination address 0x1000, data length 832 bytes) and the write data. Simultaneously, based on the information from request C (destination address 0x1400, data length 64 bytes) and the corresponding write data, it constructs merged request 2.
[0102] When the packet length is limited, the write data processing unit: extracts the write data corresponding to requests A and B, that is, starting from address 0x0000, continuously reads 512 bytes of data (256 bytes + 256 bytes) to prepare data for the first request group to be merged; extracts the write data corresponding to requests D and E, that is, starting from address 0x0240, continuously reads 320 bytes of data (256 bytes + 64 bytes) to prepare data for the second request group to be merged; the address calculation unit constructs merged request 1 based on the information of requests A and B (destination address is 0x1000, data length is 512 bytes) and the write data; constructs merged request 2 based on the information of request C (destination address is 0x1400, data length is 64 bytes) and the corresponding write data; constructs merged request 3 based on the information of requests D and E (destination address is 0x1300, data length is 320 bytes) and the corresponding write data.
[0103] Based on the above embodiments disclosed in the present application, the write data processing unit 2223 sequentially takes out the corresponding write data in the data reuse module 21 based on the address sequence and write data address of each request to be merged in the request group to be merged, thereby effectively integrating the relevant data of the request to be merged.
[0104] In some embodiments, the data reuse module 21 is further configured to retrieve the write data corresponding to the request to be merged in response to the request merging module 22, and to treat the historical request corresponding to the request to be merged in the waiting request recording unit 211 as the issued request.
[0105] It is understandable that if, in response to the request merging module 22 retrieving the write data corresponding to the request to be merged, the historical request corresponding to the request to be merged in the waiting request recording unit 211 is not recorded as an issued request, the historical request corresponding to the request to be merged in the waiting request recording unit 211 will still be a historical request. Therefore, when a new request to be issued arrives, if the address range overlaps with that of the historical request corresponding to the request to be merged, it will be ignored, thereby causing the request to be missed.
[0106] Based on the above-mentioned embodiments disclosed in the present application, by marking the historical requests corresponding to the requests to be merged in the waiting request recording unit as issued requests, the processing status of the requests can be correctly tracked and managed, thereby avoiding the situation where, when a new request to be issued arrives, the historical requests that are not marked as issued requests are ignored because the address range overlaps with the new request, thereby causing the request to be missed.
[0107] The above embodiment describes the process of receiving a request to be issued and generating a corresponding merged request. The second processor can generate and send a corresponding response message in response to the merged request. The following embodiment will illustrate the processing process of the response message.
[0108] In some embodiments, the request merging module 22 is further configured to, upon receiving a response message corresponding to a merged request, split the response message to obtain sub-response messages corresponding to each of the requests to be merged corresponding to the merged request, and send the sub-response messages corresponding to each of the requests to be merged to the data reuse module 21;
[0109] The data reuse module 21 is further configured to determine a target response message of the request to be issued corresponding to the request to be merged based on the sub-response message corresponding to the request to be merged, and feed the target response message back to the request object that issued the request to be issued.
[0110] Among them, the response message is the result returned by the second processor after processing the merged request, and may include information obtained by processing the merged request, such as data reading results, operation status, etc.; the sub-response message is the various parts obtained after the response message is split, each part corresponds to a request to be merged in the merged request, and accordingly, the target response message corresponds to each request to be issued; the request object is the object that issues the request to be issued.
[0111] It is understood that the content of the response message will vary depending on the type of request being issued (read or write). If it is a read request, the response message contains the read data; if it is a write request, the response message is a write response message indicating whether the write is completed.
[0112] In some possible implementations, the request merging module receives a response message corresponding to a merged request; based on the request information corresponding to the merged request (such as the destination address and data length of the corresponding request to be merged), it splits the response message to obtain sub-response messages corresponding to each request to be merged, and sends the sub-response messages to the data reuse module. The data reuse module receives the sub-response message from the request merging module; based on the sub-response message and information in the pending request recording unit (such as the correspondence between the request to be merged and the request to be issued), it determines the target response message for the pending request corresponding to the pending request to be merged, and feeds the target response message back to the requesting party that issued the pending request.
[0113] Based on the above-mentioned embodiment disclosed in the present application, when the request merging module receives the response message corresponding to the merged request, the response message is split to obtain the sub-response messages corresponding to each request to be merged, and these sub-response messages are sent to the data reuse module. The data reuse module determines the target response message of the request to be issued corresponding to the request to be merged based on the sub-response message, and feeds it back to the request object that issued the request to be issued. In this way, each request to be issued can receive its corresponding response message, even if the request to be issued may be ignored by the data reuse module or merged and processed by the request merging module.
[0114] Figure 7 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 7 .based on Figure 4 The request merging module 22 further includes a request merging recording unit 223 and a response data processing unit 224; wherein,
[0115] The address comparison unit 2221 is further configured to store the request group to be merged in the request merging recording unit 223;
[0116] The response data processing unit 224 is used to obtain the request group to be merged corresponding to the merged request in the request merging recording unit 223; based on the request group to be merged, split the response message to obtain sub-response messages corresponding to each request to be merged in the request group to be merged.
[0117] In some embodiments, during the issuance process before the merged request is issued, the address comparison unit 2221 needs to store the request group to be merged in the request merging recording unit 223 after completing the grouping.
[0118] In some possible implementations, when the response data processing unit receives a response message corresponding to a merged request, it obtains information of a request group to be merged corresponding to the merged request from the request merging record unit; the response data processing unit splits the response message according to information of the request group to be merged (such as the destination address and data length of the request to be merged) to obtain sub-response messages corresponding to each request to be merged.
[0119] For example, assume that in inter-chip interconnect communication, there are two pending requests (both read requests): pending request A, with a destination address of 0x1000 and a data length of 128 bytes, issued by requestee A; pending request B, with a destination address of 0x1080 (consecutive to request A's destination address), and a data length of 128 bytes, issued by requestee B. The address comparison unit groups these two pending requests into a pending request group and stores the pending request group and its related information in the request merging recording unit. These two requests are then merged into a single merged request and sent to the second processor. The second processor returns a response message, which includes the read data results of the merged requests (i.e., requests A and B); after receiving the response message, the response data processing unit obtains the information of the request group to be merged from the request merging record unit; the response data processing unit splits the response message into two sub-response messages based on this information, where sub-response message 1 (corresponding to request A): 128 bytes of data read starting from address 0x1000; sub-response message 2 (corresponding to request B): 128 bytes of data read starting from address 0x1080.
[0120] As another example, assume that in inter-chip interconnect communication, there are the following two pending requests (both write requests): pending request C, with a destination address of 0x1100 and a data length of 64 bytes, issued by request object C; pending request D, with a destination address of 0x1140 and a data length of 128 bytes, issued by request object D. The address comparison unit divides these two requests into a pending request group and sends it to the second processor for data writing. After the second processor completes the write operation, it returns a response message, namely a write response message, indicating whether the write operation was successful (for example, a status code or confirmation information, but not the actual written data). After receiving the write response message, the response data processing unit obtains information about the pending request group from the request merging record unit. Based on this information, the response data processing unit copies the write response message into two sub-response messages (because the write response message generally does not contain specific data, the copy operation is actually a copy of the write response status).
[0121] Based on the above-mentioned embodiments disclosed in the present application, the requests to be merged are divided into groups and stored in the request merging record unit through the address comparison unit, and the response message is split (or copied) according to the group information of the requests to be merged and the content of the response message through the response data processing unit, so that each request to be merged can receive a response result that matches its request type.
[0122] In some embodiments, the response data processing unit 224 is further configured to delete the to-be-merged request group corresponding to the merged request in the request merging recording unit 223 in response to completing the splitting of the response message.
[0123] It is understood that the response data processing unit reduces storage burden and improves resource utilization efficiency by deleting the pending request group information corresponding to the merged requests in the request merging record unit. This step is intended to keep the request merging record unit clean and efficient, avoid the accumulation of invalid data, and thus improve overall performance.
[0124] In some possible implementations, after receiving a response message corresponding to a merged request, the response data processing unit first splits the response message according to the pending request group information in the request merge record unit to obtain sub-response messages corresponding to each pending request. The response data processing unit then checks whether all sub-response messages have been successfully split. Once it is confirmed that all sub-response messages have been processed, the response data processing unit searches for the pending request group information corresponding to the merged request in the request merge record unit and deletes it.
[0125] Based on the above-mentioned embodiments disclosed in the present application, after the response data processing unit 224 completes the splitting of the response message, the group of requests to be merged corresponding to the merged request is deleted in the request merging record unit 223, which can reduce the storage burden of the request merging record unit and improve resource utilization efficiency; in addition, since the deletion operation is performed after all sub-response messages have been processed, no important information will be lost, thereby improving stability and reliability.
[0126] In some embodiments, when the request type of the merged request is a read request, the response message includes original read data; wherein, the response data processing unit 224 is further used to determine the destination address corresponding to each of the requests to be merged in the request group to be merged based on the request group to be merged; based on the target address corresponding to each of the requests to be merged, the original read data is split to obtain the split read data corresponding to each of the sub-response messages.
[0127] The original read data is the data read from the target address contained in the response message and is the object of the subsequent splitting by the response data processing unit. In the current embodiment, the sub-response message contains the split read data, that is, the split read data.
[0128] In some possible implementations, after the response data processing unit receives the response message (including the original read data) corresponding to the merged request, it obtains the information of the request group to be merged from the request merging record unit, including the destination address and data length of each request to be merged; based on this information, the response data processing unit extracts the read data segment corresponding to each request to be merged from the original read data in the order of the destination addresses of the requests to be merged, as split read data; and then the split read data is encapsulated into a sub-response message for subsequent processing or feedback.
[0129] For example, assume that in inter-chip interconnect communication, there are two pending requests (both read requests): pending request A: destination address 0x1000, data length 128 bytes; pending request B: destination address 0x1080 (contiguous with request A's destination address), data length 128 bytes. The second processor returns a response message containing 256 bytes of original read data starting at address 0x1000 (i.e., the sum of the data from requests A and B). After receiving the response message, the response data processing unit obtains information about the pending request group from the request merging record unit. Based on this information, it determines the starting position and length of the split read data: the read data for request A starts at address 0x1000 and is 128 bytes long; the read data for request B starts at address 0x1080 and is 128 bytes long. The response data processing unit extracts these two data segments from the original read data as split read data and encapsulates them into two sub-response messages.
[0130] Based on the above-mentioned embodiments disclosed in the present application, the response data processing unit splits the original read data according to the information of the request group to be merged, and the split read data corresponding to each sub-response message can be obtained; the read request processing flow of the inter-chip interconnection communication is optimized, and the communication efficiency is improved.
[0131] In some embodiments, the data reuse module 21 is further used to obtain at least one request to be issued that is associated with the request to be merged; and generate a target response message corresponding to each request to be issued based on the sub-response message corresponding to the request to be merged.
[0132] The target response message is generated by the data reuse module based on the sub-response message, corresponding to the processing result of each pending request. The target response message is the response message corresponding to the pending request that is ultimately fed back to the request object.
[0133] In some possible implementations, after receiving the request to be issued, the data reuse module checks whether there is a target historical request in the waiting request record unit whose address range overlaps with the request to be issued. If there is an overlap, the data reuse module will write the request to be issued as a historical request into the waiting request record unit according to the implementation provided in the above embodiment, and establish an association relationship between the historical request corresponding to the request to be issued and the target historical request. In this way, when a sub-response message corresponding to the request to be merged is received, the data reuse module will split the sub-response message into the corresponding request to be issued according to the previously stored association relationship, and generate a target response message for the request to be issued.
[0134] For example, assume that in inter-chip interconnect communication, there are the following two pending requests (both read requests): pending request A: destination address 0x1000, data length 128 bytes; pending request B: destination address 0x1000, data length 128 bytes. During the issuance process, if pending request A is sent to the request merging module as a pending request to be merged and stored as historical request A, upon receiving pending request B, the data reuse module can determine that the address range of pending request B overlaps with that of historical request A. Therefore, pending request B will be ignored, stored as historical request B, and an association relationship will be established between pending request B and historical request A. During the reception process, after obtaining the sub-response message corresponding to the pending request to be merged, all historical requests corresponding to the pending request to be merged, namely historical request A and historical request B (pending request A and pending request B), can be found based on the aforementioned association relationship. The sub-response message can then be split into pending request A and pending request B, respectively, to generate a target response message for the pending request.
[0135] Based on the above-described embodiments disclosed in this application, when the data reuse module receives a sub-response message corresponding to a pending request, it can split the sub-response message into the corresponding pending request based on the previously stored association relationship, generating a target response message for each pending request. This step ensures that each pending request receives its corresponding processing result even if it is ignored (merged).
[0136] Figure 8 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 8 The data reuse module 21 is configured with a read data cache 214 ; the data reuse module 21 is further configured to allocate a read data address for the request to be merged in the read data cache 214 when the request type of the request to be issued is a read request, and write the read data address into the waiting request recording unit 211 .
[0137] In some possible implementations, after receiving a pending request, the data reuse module may check the request type. If the request type is a read request, the data reuse module may allocate a read data address in the read data cache for the pending request and write the read data address into the pending request recording unit. In this way, in subsequent processing, the split read data corresponding to each pending request may be stored in the read data cache based on the read data address, or the corresponding read data may be read from the read data cache.
[0138] In some embodiments, the data reuse module 21 is also used to receive the split read data corresponding to the sub-response message corresponding to each of the requests to be merged sent by the request merging module; obtain the read data address corresponding to each of the requests to be merged in the waiting request recording unit 211; and based on the read data address corresponding to each of the requests to be merged, store the split read data corresponding to each of the requests to be merged in the read data cache 214.
[0139] In some possible implementations, after the data reuse module receives the sub-response message corresponding to each request to be merged sent by the request merging module, it obtains the split read data in the sub-response message; the data reuse module accesses the waiting request recording unit to obtain the read data address corresponding to each request to be merged; based on the read data address, the data reuse module stores the split read data corresponding to each request to be merged in the read data cache.
[0140] For example, it is assumed that in the inter-chip interconnection communication, there is a pending request A: the destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. A pending request B: the destination address is 0x1080, the data length is 128 bytes, and the request type is a read request. These two requests are merged into one pending request and sent to the second processor for processing. The data reuse module allocates a read data address in the read data cache for the pending request, such as 0x10000, and writes the address to the waiting request record unit. The second processor reads 256 bytes of data from the target address (covering the destination address range of requests A and B) and returns a response message; based on the foregoing embodiment, the data reuse module can receive the split read data corresponding to the sub-response message corresponding to the pending request sent by the request merging module, and store the split read data in the read data cache based on the read data address 0x10000 corresponding to the pending request.
[0141] Figure 9 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 9The data reuse module further includes a read data distribution unit 213, which is configured to, when the request type of the request to be issued is a read request, read the split read data corresponding to the request to be merged in the read data cache 214 based on the read data address corresponding to the request to be merged; and generate a target response message corresponding to the request to be issued based on the split read data and read response identifier corresponding to the request to be merged for each request to be issued.
[0142] In some possible implementations, the read data distribution unit can obtain the read data address corresponding to the pending request from the waiting request recording unit and read the split read data from the read data cache using the read data address corresponding to the pending request. The read data distribution unit then generates a target response message corresponding to each pending request based on the read response identifier and the split read data. The target response message is then fed back to the corresponding requesting party.
[0143] It is understood that for each pending request, when the address ranges read by each pending request are the same, complete and identical split read data are used in the process of generating the target response message for each pending request. Of course, when the address ranges read by each pending request are different, partial read data can be obtained from the split read data based on the address ranges read by the pending request.
[0144] In some embodiments, based on the address range of the read request to be issued and the read data address corresponding to the request to be merged, the address range of the partial read data corresponding to the request to be issued in the read data cache can be determined, and then based on the address range of the partial read data in the read data cache, the partial read data can be obtained from the split read data stored in the read data cache.
[0145] For example, assume that there are two pending requests in the inter-chip interconnection communication: pending request A: the destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. pending request B: the destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. These two requests are merged into one pending request and sent to the second processor for processing. The second processor returns a response message containing the complete split read data (128 bytes, covering the destination address range of requests A and B); after the read data distribution unit receives the response message, since the address range read by pending requests A and B is the same (both 0x1000-0x107F), the target response message can be generated directly for the two requests based on the complete split read data.
[0146] As another example, assume that in the inter-chip interconnect communication, there are the following three pending requests: pending request C: the destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. pending request D: the destination address is 0x1000, the data length is 64 bytes, and the request type is a read request. pending request E: the destination address is 0x1040, the data length is 64 bytes, and the request type is a read request. These three requests are merged into one pending request with a destination address of 0x1000 and a data length of 128 bytes, and sent to the second processor for processing. The second processor reads data starting from the target address 0x1000, covering the destination address range of requests C, D, and E, and returns a response message. The response message contains the complete split read data (a total of 128 bytes, covering the range of 0x1000-0x10FF), and stores it in the pre-assigned read data address corresponding to the pending request (taking 0x10000 as an example). For request C to be issued, since its data length is 128 bytes, the read data distribution unit obtains the 128-byte data starting at 0x10000 from the read data cache as part of the target response message, that is, the complete split read data. For request D to be issued, since its data length is 64 bytes and is the same as the destination address of request C, the read data distribution unit obtains the 64-byte data starting at 0x10000 from the read data cache as part of the target response message, that is, the first half of the split read data. For request E to be issued, since its destination address is 0x1040, the read data distribution unit obtains the 64-byte data starting at 0x10040 from the read data cache as part of the target response message, that is, the second half of the split read data.
[0147] Based on the above-described embodiments disclosed in this application, the read data distribution unit can flexibly retrieve partial or full data from the split read data to generate corresponding target response messages, based on the address range and data length of different pending read requests. This not only improves the processing efficiency of read requests and reduces unnecessary data transmission and processing, but also ensures that each pending request receives its corresponding target response message.
[0148] Figure 10 This is a schematic diagram of the structure of an inter-chip interconnection device provided in an embodiment of the present application. Figure 10 The device further comprises:
[0149] The low-latency bypass unit 50 is configured to, upon receiving a request from the first processor and the request having a low-latency identifier, directly send the request with the low-latency identifier to the inter-chip interconnection logic; or, upon receiving a request from the first processor and the request not having the low-latency identifier, determine the request without the low-latency identifier as the request to be issued.
[0150] The low-latency bypass unit is used to identify whether the request from the first processor carries a low-latency flag and, based on the flag, decide whether to send the request directly to the inter-chip interconnect logic for processing or to identify it as a pending request for subsequent processing. This can reduce processing delays for high-priority requests and improve response speed.
[0151] For example, assume that in inter-chip interconnection communication, the first processor generates three requests: request 1 carries a low-latency identifier, request 2 does not carry a low-latency identifier, and request 3 does not carry a low-latency identifier. After receiving these three requests, the low-latency bypass unit checks whether each request carries a low-latency identifier. For request 1, since it carries a low-latency identifier, the low-latency bypass unit will immediately send it directly to the inter-chip interconnection logic for processing. For request 2 and request 3, since they do not carry a low-latency identifier, the low-latency bypass unit determines them as requests to be issued and passes them to the data reuse module for subsequent processing.
[0152] Based on the above-disclosed embodiments of this application, the low-latency bypass unit identifies and processes requests with low-latency identifiers, effectively reducing processing delays for high-priority requests and improving response speed. This mechanism enables this application to prioritize requests with low-latency requirements, ensuring the timely completion of critical tasks.
[0153] Based on the foregoing embodiments, an embodiment of the present application provides an inter-chip interconnection device, which includes the units included and the modules included in each unit, and can be implemented by a processor in an electronic device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0154] An embodiment of the present application provides an electronic device, including multiple processors and an inter-chip interconnection device as described in the above embodiment corresponding one-to-one to the multiple processors.
[0155] In some implementation scenarios, the first processor is the current GPU and the second processor is the peer GPU. Figure 11, which shows a schematic diagram of the interconnection structure in the inter-GPU interconnection scenario provided by this application. Among them, the current GPU 1110 may include multiple GPU cores, and the inter-chip interconnection device in the above embodiment may include a low-latency bypass unit 1130, a data reuse module 1140, and a request merging module 1150; accordingly, the current GPU 1110 may be connected to the low-latency bypass unit 1130 in the inter-chip interconnection device via an on-chip interconnection bus 1120; the low-latency bypass unit 1130, the data reuse module 1140, and the request merging module 1150 are connected in sequence; and the inter-chip interconnection device is connected to the opposite-end GPU 1170 via the inter-chip interconnection logic 1160.
[0156] Continue reading Figure 11 , the request from the current GPU 1110 passes through the data reuse module 1140, and determines whether the data can be shared with the issued request based on the operation type of the request. If the read request address has been accessed by the issued request, then after the data of the issued request is returned, the part of the data can be directly retrieved and returned; or if the write request address has a corresponding request that has not yet been issued and is waiting to be sent, then the data of the write request waiting to be sent can be directly overwritten to reduce the actual cross-card communication generated. After passing through the data reuse module, the request enters the request merging module 1150. The request merging module 1150 will combine the requests based on all the request information collected within a period of time, calculate the longest possible packet length, and communicate through the inter-chip interconnection logic 1160.
[0157] Reverse access, such as the return of read data and the receipt of write data, follows the reverse order. For the received write data, the request merging module 1150 splits the write data into several independent requests according to the bus protocol, and sends them to the data bus through the data reuse module 1140 and the low-latency bypass unit 1130. For the returned read data, it is sent to the data reuse module 1140 based on the data information recorded when the request is integrated. When the data reuse module 1140 receives the returned data, it will split the data and send it to each request object based on the previously recorded request reuse relationship (i.e., the association relationship in the above embodiment). The low-latency bypass unit 1130 receives a low-latency request with a special identifier (i.e., the low-latency identifier in the above embodiment), and skips the aforementioned data reuse module 1140 and request merging module 1150, and sends it directly to the inter-chip interconnection logic 1160, so as to minimize the access delay as much as possible.
[0158] See also Figure 12The data reuse module 1210 includes: a waiting request recording unit 1213, a write data cache 1214, a read data cache 1211, and a read data distribution unit 1212. The write data merging logic behavior is as follows: after the write request enters the data reuse module 1210, the waiting request recording unit 1213 is queried to check whether there is a write request still in the write data cache 1214: if there is, the write data address corresponding to this write data is taken out, and the write data is overwritten to the buffer corresponding to the write data address of the unfinished write request recorded in the waiting request recording unit 1213, and the current write request is merged with the waiting request; if there is no write request, the write request is recorded separately, and a new write data address is allocated, and the write data is written to the write data cache 1214. The write data cache 1214 will select the write data according to specific rules and send it to the next-level request merging module 1220 for request merging. Possible rules include the first request to arrive first; the request with the smallest difference in address from the previous request to arrive first; the request with the most merges to arrive first, and so on.
[0159] After the read request enters the data reuse module 1210, it queries the waiting request recording unit 1213. If the waiting request recording unit 1213 already covers the data required for the request, the read request is only recorded in the waiting request recording unit 1213 without actually sending the request to the request merging module 1220, and the previous request number that the request is waiting for is marked.
[0160] Read data dispatch unit 1212 operates as follows: When a read request returns, it reports the request number to pending request record unit 1213. Other pending requests (i.e., ignored historical requests) use this number to determine whether they have already retrieved their desired data. They then use this number to look up the original pending request in pending request record unit 1213, and then determine the location of their desired data in read data cache 1211. They then read the data from read data cache 1211, append the corresponding read identification information, and return it to the corresponding requesting party. Once all possible read data dispatches for a given address have been completed, the corresponding read data cache space is released.
[0161] See also Figure 12 , request to merge modules such as Figure 12 As shown, it includes: a request merge recording unit 1223, a merge waiting unit 1221, an address calculation unit 1226, an address comparison unit 1222, a write data processing unit 1225 and a response data processing unit 1224.
[0162] After a request (i.e., a pending request) from the data reuse module 1210 enters the request merging module 1220, it is written to the merging waiting unit 1221. The address comparison unit 1222 compares the addresses of all similar requests in the merging waiting unit 1221, selects a request sequence that can generate the longest continuous request, and sends it to the address calculation unit 1226 to calculate information such as the packet length and starting address. The information about this request sequence is also recorded in the request merging recording unit 1223. After the read data is returned from the peer end, the request merging recording unit 1223 is queried to obtain the pending request corresponding to the merged request (corresponding to the merged request in the above embodiment). The pending request is then split according to the pending request and sent to the data reuse module 1210. For the merged write request, the write data processing unit 1225 retrieves the write data from the data reuse module 1210 based on the address of the sorted write data in the write data cache 1214 of the data reuse module, packages it, and sends it to the inter-chip interconnect logic.
[0163] The aforementioned data reuse module 1210 and request merging module 1220 occupy a relatively large number of logical levels, resulting in additional communication delays. Therefore, for data requiring low latency in cross-card communications, the low-latency bypass unit will be passed directly to the inter-chip interconnect logic for low-latency transmission. The low-latency bypass unit will serve as the front-end input of the entire system. Based on information such as the priority signal on the bus, it will determine whether the request qualifies as a low-latency request and choose to bypass it directly to the inter-chip interconnect logic or to the data reuse module 1210.
[0164] For example, a current GPU has four computing cores, each of which needs to read 128 bytes of data and write 32 bytes of data from the video memory of another GPU. The four computing cores need to issue read and write requests as follows: Core 0: Read request 1, read address: 0x0; Read request 2, read address: 0x80; Write request, write address: 0x1000; Core 1: Read request 1, read address: 0x80; Read request 2, read address: 0x100; Write request, write address: 0x1020; Core 2: Read request 1, read address: 0x100; Read request 2, read address: 0x180; Write request, write address: 0x1040; Core 3: Read request 1, read address: 0x180; Read request 2, read address: 0x200; Write request, write address: 0x1060. The read address of the read request is the destination address of the read request, and the write address of the write request is the destination address of the write request.
[0165] The data reuse module receives the first read request from each core, records it in the pending request record unit, and sends it to the request merging module. Next, it receives the second read request from each core. After comparing it with the read address recorded in the pending request record unit, it determines that the second requests from cores 0, 1, and 2 (0x80, 0x100, and 0x180) have already been requested. Therefore, the request is not repeated. Instead, it is recorded in the pending request record unit and marked with the number of the pending request. However, core 3's second request (0x200) has not yet been requested and is therefore sent to the request merging module. The pending request record unit ignores read request 1 from cores 1-3 and adds the request identifier of core 1's read request 1 to core 0's read request 2, the request identifier of core 2's read request 1 to core 1's read request 2, and the request identifier of core 3's read request 1 to core 2's read request 2. Furthermore, the duplicate count is recorded as 1 in read request 2 from cores 0-2. The read request 2 of core 3 is treated as a new request to be merged. Thus, a total of 8 requests are recorded in the waiting request recording unit, including 5 requests to be merged and 3 requests that were ignored.
[0166] The request merging module receives a total of five requests to be merged and enters them into the merge waiting unit: 0x0, 0x80, 0x100, 0x180, and 0x200. The address comparison unit determines that the data required by these five requests is adjacent, so the information about these five requests is sent to the address calculation unit and recorded in the request merging recording unit. The address calculation unit calculates the starting address of the final packet to be sent as 0x0 and the packet length as 0x280 (640) bytes, and then sends the packet to the inter-chip interconnect logic.
[0167] After a read request is returned from the inter-chip interconnect logic, the request merging module uses the returned read request number to locate the original five pending requests in the request merging record unit. It then splits the original read data based on the address information of the pending requests and fills the corresponding read data buffer. After the split is complete, the resources in the request merging record unit are released.
[0168] After receiving the returned read data, the data reuse module searches the waiting request record unit based on the identification information carried in the read data and broadcasts the entry number to all waiting request objects. Taking the data corresponding to the return address 0x100 as an example, the corresponding number (02) broadcast will inform Core 1's second read request that the required data is also ready. Based on this, the read data distribution unit will search the corresponding read data cache address in item 02 in the waiting request record unit and retrieve the data and return it to Core 1 as the return data for its second read request.
[0169] For the four core write requests, after entering the waiting request recording unit, the four requests found no address overlap and were subsequently sent to the request merging module. After the merging waiting unit detected that the four data were consecutive addresses, the request merging module merged the four 32-byte write requests into a single 128-byte write request. The write data was retrieved from the write data cache in address order, formed into a packet, and sent to the inter-chip interconnect logic.
[0170] After Core 3 has received all of its read data, it sends a signal to the other GPU. This signal is intended to have the lowest possible latency. Core 3 marks this request with a priority signal. This request is detected by the low-latency bypass unit and sent directly to the inter-chip interconnect logic, bypassing the data reuse module and request merging module to achieve the lowest possible latency.
[0171] As a result, the data reuse module reduces the total amount of data requiring inter-chip communication. The request merging module increases the packet length of inter-chip communication, thereby improving the actual data bandwidth utilization efficiency of inter-chip communication. The low-latency bypass unit eliminates the additional delay introduced by the data reuse module and the request merging module, ensuring that low-latency access requirements still achieve the expected low latency.
[0172] Figure 13 A hardware entity diagram of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 13 As shown, the hardware entity of the electronic device 1300 includes: a first processor 1310, an inter-chip interconnection device 1320, an inter-chip interconnection logic 1330, and a second processor 1340. The inter-chip interconnection device 1320 can be understood with reference to the description of the device embodiment of the present application.
[0173] The first processor and the second processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the above-mentioned processors may also be other electronic devices, and the embodiments of the present application are not specifically limited thereto.
[0174] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0175] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0176] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0177] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0178] In addition, the functional units in the various embodiments of the present application can all be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units. It can be understood by those skilled in the art that all or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memories (ROMs), magnetic disks, or optical disks.
[0179] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0180] The above is only an implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. An inter-chip interconnection device, characterized in that: The inter-chip interconnection device is provided between the first processor and the inter-chip interconnection logic, and the inter-chip interconnection device includes a data reuse module and a request merging module. The data reuse module includes a waiting request recording unit, and the waiting request recording unit is used to store historical requests in the inter-chip communication process and historical addresses of the historical requests; wherein, The data reuse module is configured to receive a request to be issued, and determine the request to be issued as a request to be merged if there is no target historical request in the historical requests whose address range overlaps with the request to be issued; The data reuse module is further configured to sort the unissued requests to be merged to obtain an issuance order; and sequentially send the unissued requests to be merged to the request merging module according to the issuance order; the issuance order is determined based on request attributes of the unissued requests to be merged; the request attributes include at least one of the following: the time when the unissued request to be merged is received, the address interval between the destination address of the unissued request to be merged and the last issued request to be merged, and the number of historical requests associated with the unissued request to be merged; Wherein, when the request attribute includes the number of historical requests associated with the unissued request to be merged, if the request type of the request to be merged is a write request, the request to be merged with a smaller number is issued first; if the request type of the request to be merged is a read request, the request to be merged with a larger number is issued first; The request merging module is configured to merge the received requests to be merged to obtain a merged request, and send the merged request to the inter-chip interconnect logic; the inter-chip interconnect logic is configured to send the merged request to the second processor.
2. The inter-chip interconnection device according to claim 1, characterized in that: in, The data reuse module is further configured to update the waiting request recording unit based on the request to be issued.
3. The inter-chip interconnection device according to claim 2, characterized in that: The data reuse module is further configured to write the request to be issued as a historical request into the waiting request recording unit.
4. The inter-chip interconnection device according to claim 3, characterized in that: The data reuse module includes a write data cache; The data reuse module is further configured to, when the request type of the request to be issued is a write request, allocate a new write data address for the request to be issued in the write data cache; and write the write data corresponding to the request to be issued into the write data cache based on the new write data address; And write the new write data address into the waiting request recording unit.
5. The inter-chip interconnection device according to claim 2, characterized in that: The data reuse module is further configured to ignore the request to be issued if there is a target historical request whose address range overlaps with the request to be issued; The data reuse module is further configured to write the request to be issued as a historical request into the waiting request recording unit, and establish an association relationship between the historical request corresponding to the request to be issued and the target historical request.
6. The inter-chip interconnection device according to claim 5, characterized in that: The data reuse module is further configured to, when the request type of the request to be issued is a write request, obtain a target write data address; and write the write data corresponding to the request to be issued into the write data cache based on the target write data address; The target write data address is the write data address of the earliest historical request whose address range overlaps with the request to be issued.
7. The inter-chip interconnection device according to any one of claims 1 to 6, characterized in that: The request merging module includes a merging waiting unit and a merging processing unit; The merging waiting unit is configured to receive and store the request to be merged; The merging processing unit is used to merge the requests to be merged according to the destination addresses of the requests to be merged to obtain the merged requests.
8. The inter-chip interconnection device according to claim 7, characterized in that: The merging processing unit includes an address comparing unit and an address calculating unit, wherein: the address comparison unit is configured to compare destination addresses of the requests to be merged stored in the merging waiting unit, divide the requests to be merged stored in the merging waiting unit into at least one request group to be merged, and send the request group to the address calculation unit; The address calculation unit is configured to determine request information of the merged request based on the requests to be merged in the group of requests to be merged, and to generate the merged request based on the request information.
9. The inter-chip interconnection device according to claim 8, characterized in that: The request information includes a request address and a data length; the request address is the first logically preceding destination address among the destination addresses of the requests to be merged in the request group to be merged; the data length of the merged request is the sum of the data lengths of the requests to be merged in the request group to be merged.
10. The inter-chip interconnection device according to claim 8, characterized in that: In the case where the request type of the request to be merged is a write request, the merge processing unit further includes a write data processing unit, wherein: The write data processing unit is configured to sequentially retrieve, from the data reuse module, the write data corresponding to each of the requests to be merged in the group of requests to be merged based on the address sequence and write data addresses of each of the requests to be merged in the group of requests to be merged; The address calculation unit is further configured to construct the merged request based on the request information and the write data corresponding to each of the requests to be merged in the group of requests to be merged.
11. The inter-chip interconnection device according to claim 7, wherein: The data reuse module is further configured to, in response to the request merging module, retrieve the write data corresponding to the request to be merged, and use the historical request corresponding to the request to be merged in the waiting request recording unit as the issued request.
12. The inter-chip interconnection device according to any one of claims 8 to 10, characterized in that: The request merging module is further configured to, upon receiving a response message corresponding to a merged request, split the response message to obtain sub-response messages corresponding to each of the requests to be merged corresponding to the merged request, and send the sub-response messages corresponding to each of the requests to be merged to the data reuse module; The data reuse module is further configured to determine a target response message of the request to be issued corresponding to the request to be merged based on the sub-response message corresponding to the request to be merged, and feed the target response message back to the request object that issued the request to be issued.
13. The inter-chip interconnection device according to claim 12, wherein: The request merging module includes an address comparison unit; the request merging module also includes a request merging recording unit and a response data processing unit; wherein, The address comparison unit is further configured to store the request group to be merged in the request merging recording unit; The response data processing unit is used to obtain the group of requests to be merged corresponding to the merged request in the request merging recording unit; based on the group of requests to be merged, split the response message to obtain sub-response messages corresponding to each of the requests to be merged in the group of requests to be merged.
14. The inter-chip interconnection device according to claim 13, wherein: The response data processing unit is further configured to delete the to-be-merged request group corresponding to the merged request in the request merging record unit in response to completing the splitting of the response message.
15. The inter-chip interconnection device according to claim 13, wherein: In the case where the request type of the merged request is a read request, the response message includes original read data; wherein, The response data processing unit is also used to determine the destination address corresponding to each of the requests to be merged in the request group to be merged based on the request group to be merged; based on the target address corresponding to each of the requests to be merged, the original read data is split to obtain the split read data corresponding to each of the sub-response messages.
16. The inter-chip interconnection device according to claim 12, wherein: The data reuse module is further configured to obtain at least one request to be issued that is associated with the request to be merged; and generate a target response message corresponding to each request to be issued based on the sub-response message corresponding to the request to be merged.
17. The inter-chip interconnection device according to claim 16, characterized in that: The data reuse module is configured with a read data cache; The data reuse module is further configured to allocate a read data address for the request to be merged in the read data cache and write the read data address into the waiting request recording unit when the request type of the request to be issued is a read request.
18. The inter-chip interconnection device according to claim 17, characterized in that: The data reuse module is also used to receive the split read data corresponding to the sub-response message corresponding to each of the requests to be merged sent by the request merging module; obtain the read data address corresponding to each of the requests to be merged in the waiting request recording unit; and based on the read data address corresponding to each of the requests to be merged, store the split read data corresponding to each of the requests to be merged in the read data cache.
19. The inter-chip interconnection device according to claim 18, further comprising a read data distribution unit, characterized in that: The read data distribution unit is configured to, when the request type of the request to be issued is a read request, read the split read data corresponding to the request to be merged in the read data cache based on the read data address corresponding to the request to be merged; For each of the requests to be issued, a target response message corresponding to the request to be issued is generated based on the split read data and the read response identifier corresponding to the request to be merged.
20. The inter-chip interconnection device according to any one of claims 1 to 6, characterized in that: The inter-chip interconnection device further includes: A low-latency bypass unit is configured to, upon receiving a request from the first processor and the request having a low-latency identifier, directly send the request with the low-latency identifier to the inter-chip interconnection logic; or, upon receiving a request from the first processor and the request not having the low-latency identifier, determine the request without the low-latency identifier as the request to be issued.
21. An electronic device, characterized in that: include: A plurality of processors and an inter-chip interconnection device according to any one of claims 1 to 20 corresponding one-to-one to the plurality of processors.
Citation Information
Patent Citations
Memory access control structure and method, memory system, processor and electronic equipment
CN116680089A
Packet Combiner for a Packetized Bus with Dynamic Holdoff time
US20070079044A1