Inter-chip interconnection device and electronic equipment

By setting up a data reuse module and a request merging module in the inter-chip interconnection device, the problem of low data bandwidth utilization in the prior art is solved, and more efficient inter-chip interconnection communication performance is achieved.

CN120067035AActive Publication Date: 2025-05-30MOORE THREADS TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510535244.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

When the prior art improves the performance of inter-chip interconnected communication, the utilization of data bandwidth is limited, resulting in limited improvement in communication performance.

Method used

By setting up a data reuse module and a request merging module in the inter-chip interconnection device, the data reuse module receives the request to be issued and determines its merge request. The request merging module processes the merge request, reducing unnecessary duplicate requests and communication overhead.

Benefits of technology

Optimize communication efficiency, reduce communication overhead for inter-chip interconnection, and improve data transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067035A_ABST
    Figure CN120067035A_ABST
Patent Text Reader

Abstract

The invention discloses an inter-chip interconnection device, the inter-chip interconnection device is arranged between a first processor and an inter-chip interconnection logic, the inter-chip interconnection device comprises a data reuse module and a request merging module, the data reuse module is used for receiving a request to be sent, and the request merging module is used for merging the request to be sent; determining a to-be-merged request based on the destination address of the to-be-sent request; the data reusing module is also used for sending the to-be-merged request to the request merging module; the request merging module is used for performing merging processing on the received to-be-merged requests to obtain merged requests and sending the merged requests to the inter-chip interconnection logic; the inter-chip interconnection logic is to send the merged request to a second processor. According to the invention, the inter-chip communication performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of circuit technologies, and particularly relates to an inter-chip interconnect device and an electronic device. Background Art

[0002] With the increasing scale of chips and the growing demand for chip computing power in emerging applications, the demand for inter-chip interconnect communication has emerged. The communication performance of inter-chip interconnect communication is usually achieved through a wider data bit width, a higher clock frequency, and longer data packets specified in the coherence protocol; however, these information cannot be increased without limit, resulting in limited data bandwidth utilization, thereby affecting the improvement of inter-chip communication performance. Summary of the Invention

[0003] In view of this, embodiments of this application at least provide an inter-chip interconnect device and an electronic device, which can improve inter-chip communication performance.

[0004] The technical solution of the embodiments of this application is implemented as follows: On the one hand, an embodiment of this application provides an inter-chip interconnect device. The inter-chip interconnect device is disposed between a first processor and inter-chip interconnect logic. The inter-chip interconnect device includes a data reuse module and a request merging module. Among them, the data reuse module is configured to receive a request to be sent, and determine a request to be merged based on the destination address of the request to be sent; the data reuse module is further configured to send the request to be merged to the request merging module; the request merging module is configured to perform a merging process on the received request to be merged to obtain a merged request, and send the merged request to the inter-chip interconnect logic; the inter-chip interconnect logic is configured to send the merged request to a second processor.

[0005] On the other hand, an embodiment of this application provides an electronic device, including a plurality of processors and the inter-chip interconnect devices provided in the above embodiments corresponding to the plurality of processors one by one.

[0006] In the embodiments of this application, by setting a data reuse module in the inter-chip interconnect device to receive a request to be sent, and comparing and judging the destination address of the request to be sent with the addresses of the previously sent requests, it is determined which requests to be sent need to be passed as requests to be merged to the request merging module and which requests to be sent need to be ignored, thereby reducing unnecessary repeated requests and optimizing communication efficiency; at the same time, by sending the request to be merged from the data reuse module to the request merging module, the request merging module can perform a merging process on the received request to be merged to obtain a merged request, reducing the number of requests that the inter-chip interconnect logic needs to process and reducing communication overhead; based on the embodiments provided in this application, the communication overhead of inter-chip interconnect can be reduced to a certain extent, and the data transmission efficiency of inter-chip interconnect can be improved.

[0007] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, rather than limiting the technical solutions of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with this application and, together with the specification, are used to explain the technical solutions of this application.

[0009] Figure 1 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 1 ; Figure 2 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 2 ; Figure 3 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 3 ; Figure 4 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 4 ; Figure 5 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 5 ; Figure 6 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 6 ; Figure 7 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 7 ; Figure 8 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 8 ; Figure 9 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 9 ; Figure 10 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 10 ; Figure 11 Schematic diagram of the interconnect structure in the inter-GPU interconnect scenario provided by an embodiment of this application; Figure 12 Schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of this application Figure 10 One; Figure 13 Schematic diagram of the hardware entity of an electronic device provided by an embodiment of this application. Detailed implementation manners

[0010] To make the objectives, technical solutions and advantages of this application clearer, the technical solutions of this application will be further elaborated in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0011] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.

[0013] To facilitate the understanding of this solution, before describing the embodiments of this application, the application background in the embodiments of this application will be described.

[0014] Inter-die communication includes communication between several dies encapsulated on the same chip (die to die, D2D), and also includes communication between multiple boards (chip to chip C2C). Since the physical connection distance between dies is longer than that inside the die, when accessing one die from another through D2D, the communication performance will be affected. Similarly, the physical connection distance between chips is longer than that inside the chip, and when accessing one chip from another through C2C, the communication performance will also be affected. To improve the inter-die communication performance, in related technologies, the port width and frequency are usually expanded to provide a larger inter-die interconnect bandwidth, and a more complex logic is used to support the coherence protocol to support longer packet lengths of packets, improve the transmission efficiency, and reduce the ratio of instruction information to actual data in the packet. However, when the amount of data transmitted between dies does not decrease, due to the need to maintain coherent communication with the limited inter-die interconnect bandwidth, the packet may require more complex instruction information, thus bringing an additional communication burden. Moreover, the length of the packet supported by the coherence protocol is also limited by the actual emission ability of the graphics processing unit (GPU) and the data interleaving granularity during multi-port emission of the GPU, and cannot be increased indefinitely, resulting in limited total data bandwidth utilization.

[0015] To solve the above problems, an embodiment of the present application provides an inter-die interconnect device, which is arranged between a processor and inter-die interconnect logic. When the processor transmits data, the data needs to be transmitted to the inter-die interconnect logic through the inter-die interconnect device, and then transmitted to other processors through the inter-die interconnect logic. Since the inter-die interconnect device merges requests and reduces the number of transmissions, the data bandwidth utilization is improved, and thus the inter-die communication performance is enhanced.

[0016] Figure 1 Schematic diagram of the device structure of an inter-die interconnect device provided by an embodiment of the present application Figure 1 , as Figure 1 shown, the inter-die interconnect device is arranged between the first processor 10 and the inter-die interconnect logic 30. The inter-die interconnect device 20 includes a data reuse module 21 and a request merge module 22. Among them, the data reuse module 21 is configured to receive a request to be sent, and determine a request to be merged based on the destination address of the request to be sent; the data reuse module 21 is further configured to send the request to be merged to the request merge module 22; the request merge module 22 is configured to perform a merge process on the received request to be merged to obtain a merged request, and send the merged request to the inter-die interconnect logic 30; the inter-die interconnect logic 30 is configured to send the merged request to the second processor 40.

[0017] Among them, the above-mentioned inter-chip interconnect device is a dedicated hardware module disposed between the processor and the inter-chip interconnect logic, which is used to optimize the data transmission process between processors. By merging requests and reducing the number of transmissions, the data bandwidth utilization rate is improved, thereby enhancing the inter-chip communication performance. Among them, the inter-chip interconnect logic is the hardware (or software component) responsible for transmitting data between processors. The inter-chip interconnect logic is used to receive the merged requests from the inter-chip interconnect device and forward these merged requests to the target processor (such as the second processor) to achieve data communication between processors.

[0018] In some embodiments, the above-mentioned data reuse module is used to receive the requests to be sent issued by the first processor 10, compare and judge according to the destination address of the requests to be sent and the addresses of the requests issued historically, and determine which requests to be sent need to be passed as requests to be merged to the request merging module, and which requests to be sent need to be ignored (that is, they have been requested and do not need to be sent to the next unit), so as to reduce unnecessary repeated requests and optimize the communication efficiency.

[0019] In some embodiments, the above-mentioned request merging module is used to receive the requests to be merged from the data reuse module and perform merging processing on these requests to be merged to generate merged requests. The purpose of merging requests is to reduce the number of requests, reduce the burden on the inter-chip interconnect logic, and improve the overall data transmission efficiency.

[0020] In some possible implementation manners, the data reuse module receives the requests to be sent from the first processor and records the requests to be sent. The data reuse module compares and judges according to the destination address of the requests to be sent and the recorded addresses, and determines which requests need to be passed as requests to be merged to the request merging module, and which requests need to be ignored (that is, the corresponding destination addresses have been requested and do not need to be sent to the request merging module). Then, the data reuse module sends the requests to be merged to the request merging module. The request merging module performs merging processing on the received requests to be merged to generate merged requests. Finally, the request merging module sends the merged requests to the inter-chip interconnect logic, and the inter-chip interconnect logic forwards the requests to the second processor.

[0021] Based on the above embodiments disclosed in the present application, by setting up a data reuse module in the inter-chip interconnect device to receive requests to be sent, and comparing and judging the destination address of the requests to be sent with the addresses of the requests sent historically, it is determined which requests to be sent need to be passed to the request merging module as requests to be merged, and which requests to be sent need to be ignored, thereby reducing unnecessary duplicate requests and optimizing communication efficiency. At the same time, by sending requests to be merged from the data reuse module to the request merging module, the request merging module can perform merging processing on the received requests to be merged to obtain merged requests, reducing the number of requests that the inter-chip interconnect logic needs to process and reducing communication overhead. Based on the embodiments provided in the present application, the communication overhead of inter-chip interconnect can be reduced to a certain extent, and the data transmission efficiency of inter-chip interconnect can be improved.

[0022] Figure 2 is a schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of the present application Figure 2 Based on Figure 1 the data reuse module 21 includes a waiting request recording unit 211, and the waiting request recording unit 211 is used to store historical requests and historical addresses of the historical requests during the inter-chip communication process. Wherein, the data reuse module 21 is further configured to determine requests to be merged based on the destination address of the requests to be sent and the historical addresses of the historical requests. The data reuse module 21 is further configured to update the waiting request recording unit 211 based on the requests to be sent.

[0023] Among them, the inter-chip communication process refers to all processes from when an inter-chip interconnect device receives a request until the inter-chip interconnect device returns a response corresponding to the request.

[0024] Among them, the waiting request recording unit 211 is used to store historical requests and their corresponding historical addresses during the inter-chip communication process.

[0025] In some embodiments, the address ranges between the destination addresses of the requests to be merged and the historical addresses of the historical requests do not overlap. That is to say, the requests to be merged are: requests that are determined not to cause duplicate sending and whose address ranges do not overlap by comparing the destination address of the requests to be sent with the historical addresses of the historical requests. The requests to be merged will be passed to the request merging module to reduce the number of communication times and optimize communication efficiency.

[0026] In some possible implementation manners, when the inter-chip interconnect device receives a request to be sent out, the data reuse module 21 compares its destination address with the historical addresses in the waiting request recording unit 211. According to the comparison result, it is determined whether the request will cause duplicate sending, that is, whether the address ranges do not overlap. If it will not cause duplication, it is determined as a request to be merged and prepared to be passed to the request merging module; after the above determination is completed, the waiting request recording unit 211 records the request to be sent out and its corresponding destination address as a new historical request and the corresponding historical address to update the waiting request recording unit 211.

[0027] Exemplarily, assume that there are three requests to be sent out: Request 1, Request 2, and Request 3. The destination address range of Request 1 overlaps with a previous historical request, so Request 1 is determined as a duplicate request and is ignored. The destination address ranges of Request 2 and Request 3 do not overlap with the historical requests, so Request 2 and Request 3 are determined as requests to be merged and are passed to the request merging module for merging processing. After the determination is completed, the waiting request recording unit 211 records Request 1, Request 2, and Request 3 and their corresponding destination addresses as new historical requests and the corresponding historical addresses to update the waiting request recording unit 211.

[0028] Based on the above embodiments disclosed in the present application, by determining the request to be merged based on the destination address of the request to be sent out and the historical address of the historical request, it is possible to filter requests with overlapping addresses, reduce unnecessary communication overhead, and improve communication efficiency; in addition, by updating the waiting request recording unit 211 based on the request to be sent out, the real-time performance and accuracy of the waiting request recording unit 211 can be maintained, providing the latest data support for subsequent request processing; based on the embodiments provided in the present application, the efficiency of inter-chip communication can be effectively improved and the communication overhead can be reduced.

[0029] In some embodiments, the data reuse module 21 is further configured to determine the request to be sent out as the request to be merged when there is no target historical request whose address range overlaps with the request to be sent out; the data reuse module 21 is further configured to write the request to be sent out as a historical request into the waiting request recording unit 211.

[0030] In some possible implementation manners, when the data reuse module 21 receives a request to be sent, it reads all historical requests and their corresponding historical addresses from the waiting request recording unit 211; compares the destination address of the request to be sent with the historical addresses of each historical request to determine whether the address ranges overlap; if the address range of the request to be sent does not overlap with the address range of any historical request, the request to be sent is determined as a request to be merged. Correspondingly, the waiting request recording unit 211 takes the request to be sent as a new historical request and its destination address as the corresponding historical address; writes the new historical request and the historical address into the waiting request recording unit 211.

[0031] Exemplarily, assume that there are three requests to be sent: Request 1, Request 2, and Request 3. Among them, the destination address range of Request 1 is 0x1000 - 0x1FFF, and there is a historical request in the waiting request recording unit 211 with an address range of 0x2000 - 0x2FFF. Since the address range of Request 1 does not overlap with the address range of the historical request, Request 1 is determined as a request to be merged; the destination address range of Request 2 is 0x2000 - 0x2FFF. Since the address range of Request 2 overlaps with the address range of the historical request, Request 2 is not determined as a request to be merged. The destination address range of Request 3 is 0x3000 - 0x3FFF. Since the address range of Request 3 does not overlap with the address range of any historical request, Request 3 is determined as a request to be merged. Correspondingly, after the above judgment is completed, the waiting request recording unit 211 takes Requests 1 to 3 as new historical requests and writes the corresponding destination addresses as new historical addresses into the waiting request recording unit 211.

[0032] Based on the above embodiments disclosed in the present application, by comparing the destination address of the request to be sent with the historical address of the historical request to determine the request to be merged, the sending of duplicate requests can be effectively reduced, and the communication overhead can be reduced. At the same time, taking the request to be sent as a historical request and writing it into the waiting request recording unit 211 can maintain the timeliness and accuracy of the recording unit and provide the latest data support for subsequent request processing. Based on the embodiments provided in the present application, the efficiency of inter-chip communication can be significantly improved, the network load can be reduced, and the overall performance can be enhanced.

[0033] Figure 3 is a schematic structural diagram of an inter-chip interconnection device provided by an embodiment of the present application Figure 3 Based on Figure 2, the data reuse module 21 includes a write data cache 212; the data reuse module 21 is further configured to, when the request type of the to-be-issued request is a write request, allocate a new write data address for the to-be-issued request in the write data cache 212; write the write data corresponding to the to-be-issued request into the write data cache 212 based on the new write data address; and write the new write data address into the waiting request recording unit 211.

[0034] In some possible implementation manners, after receiving a to-be-issued request, the data reuse module 21 may parse the request type information in the request to determine whether the to-be-issued request is a read request or a write request; if the to-be-issued request is a write request, the data reuse module 21 searches for free storage space in the write data cache 212 to allocate a new write data address for the request; writes the write data corresponding to the to-be-issued request into the write data cache 212 according to the allocated write data address; correspondingly, writes the allocated new write data address into the waiting request recording unit 211 and stores it together with other information of the write request (such as the request type, destination address, etc.) for subsequent processing and tracking of the write request.

[0035] Exemplarily, assume that there is a to-be-issued request with a request type of write request, a destination address of 0x2000, and a write data length of 128 bytes. The data reuse module 21 parses the request information of the to-be-issued request to determine that the request type of the to-be-issued request is a write request; searches for free space in the write data cache 212 to allocate a new write data address for the write request, such as 0x3000; writes the 128-byte write data into the write data cache 212 according to the address 0x3000; writes the write data address 0x3000 into the waiting request recording unit 211 and stores it together with information such as the destination address 0x2000 and the request type of the write request.

[0036] Based on the above embodiments disclosed in the present application, by setting the write data cache 212 in the data reuse module 21, a dedicated storage space can be provided for the write data, enabling the write data to have a reasonable storage location before transmission, realizing the centralized storage of the write data, facilitating subsequent management and operation of the write data. For example, these write data can be more conveniently searched and processed when needed, improving the efficiency of data processing; moreover, writing the new write data address into the waiting request recording unit 211 provides key information for subsequent tracking and processing of the write request, enabling a clear understanding of the storage location of the write data corresponding to each write request, which helps to perform corresponding processing when the write data transmission is completed or an exception occurs.

[0037] In some embodiments, the data reuse module 21 is further configured to ignore the to-be-issued request when there is a target historical request whose address range overlaps with the to-be-issued request; the data reuse module 21 is further configured to write the to-be-issued request as a historical request into the waiting request recording unit 211 and establish an association relationship between the historical request corresponding to the to-be-issued request and the target historical request.

[0038] In the above embodiments, if the address ranges of multiple requests overlap, processing these requests simultaneously may cause problems such as data conflicts and resource waste. Therefore, after receiving the to-be-issued request, the data reuse module 21 checks whether there is a target historical request in the waiting request recording unit 211 whose address range overlaps with the to-be-issued request. If so, the to-be-issued request is ignored to avoid duplicate processing.

[0039] Exemplarily, there is a to-be-issued request with a request type of write request, a destination address range of 0x2000 - 0x2FFF, and a write data length of 4KB. The waiting request recording unit 211 has already recorded a historical request with an address range of 0x2500 - 0x3000. The data reuse module 21 receives this write request, parses out the address range of 0x2000 - 0x2FFF; compares this address range with the address range of the historical request in the waiting request recording unit 211; finds that there is an overlap between the address range 0x2500 - 0x3000 of a historical request and the address range 0x2000 - 0x2FFF of the to-be-issued request; the data reuse module 21 ignores this write request and does not pass it to the next processing unit.

[0040] In some embodiments, by ignoring the to-be-issued request when there is a target historical request whose address range overlaps with the to-be-issued request, duplicate processing of data in the same address range can be avoided, unnecessary resource consumption can be reduced, such as reducing the parsing and execution time of the processor for duplicate requests and reducing the occupation of memory bandwidth; at the same time, by writing the to-be-issued request as a historical request into the waiting request recording unit 211 and establishing an association relationship between the historical request corresponding to the to-be-issued request and the target historical request, subsequent management and tracking of requests can be facilitated, and relevant requests can be uniformly processed based on this association relationship in subsequent processing, improving the efficiency of request processing; based on the embodiments provided in this application, the request processing flow in the inter-chip communication process can be optimized, the complexity of request processing can be reduced, and the efficiency and overall performance of inter-chip communication can be improved.

[0041] In some embodiments, the data reuse module 21 is further configured to, when the request type of the to-be-issued request is a write request, obtain a target write data address; based on the target write data address, write the write data corresponding to the to-be-issued request into the write data cache 212; the target write data address is the write data address of the earliest historical request whose address range overlaps with the to-be-issued request.

[0042] Among them, when there is a to-be-issued request with a request type of write request and there is a target historical request whose address range overlaps with the to-be-issued request, it means that before the current to-be-issued request, there is at least one historical request whose historical address overlaps with the destination address of this to-be-issued request. Therefore, when writing data to the same destination address, the latest write data needs to be written to this target address, that is, the write data corresponding to the current to-be-issued request needs to be written to this target address.

[0043] In some embodiments, there may be at least one historical request whose address range overlaps with the destination address of the to-be-issued request. At this time, it is necessary to find the earliest historical request among the at least one historical request and obtain the write data address corresponding to this historical request as the target write data address. It can be understood that when there are at least two historical requests whose address ranges overlap with the destination address of the to-be-issued request, the write data stored in the target write data address is no longer the write data corresponding to the earliest historical request, but the write data corresponding to the latest historical request among the at least two historical requests. Of course, after the write data of the current to-be-issued request is written to the target write data address, the write data stored in the target write data address is also the write data of the latest request among these historical requests.

[0044] In the above embodiments, in order to solve the problem of how to ensure that data is correctly written to the target address and maintain the latestness and consistency of the data in a write request. When there is a to-be-issued write request and the address range of this request overlaps with a previous historical request, it is necessary to find the earliest one among these historical requests and obtain its write data address as the target write data address. In this way, when writing the data of the current to-be-issued request, it can be ensured that the data is overwritten to the correct position and is the latest write data. In this way, after the write request is issued, the latest write data can be found based on the target write data address, thereby avoiding data conflicts and inconsistencies and improving the efficiency and accuracy of data processing.

[0045] In some possible implementations, when the data reuse module receives a write request to be sent out, it checks whether the address range of the request overlaps with previous historical requests. If there is an overlap, the data reuse module further looks for the earliest one among these historical requests and obtains its write data address as the target write data address. Then, the data reuse module writes the write data corresponding to the request to be sent out into the write data cache. Finally, the data reuse module officially writes the write data in the cache to the target write data address, thus completing the data writing operation. The entire process needs to ensure the accuracy and consistency of the data to avoid data loss or errors.

[0046] Exemplarily, assume that there are three historical requests recorded in the waiting request recording unit 211, namely historical request A, historical request B, and historical request C. Among them: Historical request A: The request type is a write request, the destination address range is 0x1000 - 0x2000, the write data length is 4KB, and it is the first write request to arrive. From a logical processing perspective, set its corresponding initial target write data address (for subsequent overwrite logic tracking) as Addr_A (which actually corresponds to a physical or logical address space that can store 4KB of data). Historical request B: The request type is a write request, the destination address range is 0x1000 - 0x2000, the write data length is 4KB, and it arrives after historical request A. Historical request C: The request type is a write request, the destination address range is 0x1000 - 0x2000, the write data length is 4KB, and it arrives after historical request B.

[0047] At this time, the first processor issues a new request D to be sent out. The request type is a write request, the destination address range is 0x1000 - 0x2000, the write data length is 4KB, and the write data content is different from that of historical requests A, B, and C (it can also be the same). After receiving the request D to be sent out, the data reuse module 21 first determines that its request type is a write request. Then, it compares the address range of the request D to be sent out with the address ranges of all historical requests in the waiting request recording unit 211; through comparison, it is found that the address ranges of historical requests A, B, and C all overlap with the address range of the request D to be sent out.

[0048] The update process of the above historical requests A to C for the target write data address can be understood as follows: According to the arrival time order of the requests, the data reuse module 21 will process these write requests in sequence. Since historical request A is the first to arrive, its write data will be first written to the storage location corresponding to the target write data address Addr_A. Then, the write data of historical request B will overwrite the data of historical request A within the address range of Addr_A (corresponding to the overlapping part of 0x1000 - 0x2000). At this time, although logically still tracked by Addr_A, the content stored at this target write data address has actually been updated to the data of historical request B. Finally, the write data of historical request C will overwrite the data of historical request B within the address range of Addr_A.

[0049] When the pending request D arrives, the data reuse module 21 will also write its write data to the storage location corresponding to the target write data address Addr_A in chronological order, overwriting the data of the previous historical request C within the address range of Addr_A. Therefore, the data stored at the storage location corresponding to Addr_A finally is the data of the pending request D.

[0050] Based on the above embodiments disclosed in the present application, based on the obtained target write data address, writing the write data corresponding to the pending request into the write data cache 212 and using the cache to temporarily store the write data can, on the one hand, relieve the possible pressure caused by directly writing data to the target storage location, and on the other hand, also facilitate the subsequent unified data processing; in addition, setting the target write data address as the write data address of the earliest historical request whose address range overlaps with the pending request can, when processing write requests, consider data overwriting and writing in a relatively orderly manner, laying a logical foundation for subsequent data overwriting operations. Based on the embodiments provided in the present application, in a complex scenario where there are multiple overlapping write request address ranges, the writing and overwriting problems of write data can be processed relatively reasonably, optimizing the write data processing flow to a certain extent and improving the data processing efficiency of the inter-chip connection device.

[0051] In some embodiments, the data reuse module 21 is further configured to sort the unissued pending merge requests to obtain an issuance order; and sequentially send the unissued pending merge requests to the request merge module 22 according to the issuance order.

[0052] Among them, the above issuance order is the request sending order determined after the data reuse module sorts the pending merge requests. In some embodiments, this issuance order can be determined based on multiple factors, such as the priority of the requests, the overlapping situation of the address ranges, the arrival time order of the requests, etc. By arranging a reasonable issuance order, conflicts between requests can be avoided and the request processing efficiency can be improved.

[0053] In some possible implementation manners, after the data reuse module obtains the unissued merge requests to be merged, it may store these merge requests to be merged in an internal request queue; analyze the merge requests to be merged in the request queue, extract the request attributes of the merge requests to be merged, sort the merge requests to be merged in the request queue according to the request attributes of the merge requests to be merged to obtain the sending order; and sequentially take out the merge requests to be merged from the request queue according to the sending order and send them to the request merging module.

[0054] Based on the above embodiments disclosed in this application, sending the unissued merge requests to be merged to the request merging module 22 sequentially according to the sending order can avoid conflicts between requests, reduce waiting and chaos during the request processing, and improve the processing efficiency of the request merging module for requests.

[0055] In some embodiments, the sending order is determined based on the request attributes of the unissued merge requests to be merged; the request attributes include at least one of the following: the time when the unissued merge request is received, the address interval between the destination address of the unissued merge request and the destination address of the previous issued merge request, and the number of historical requests associated with the unissued merge request.

[0056] Among them, the time when the unissued merge request is received refers to the specific moment when the data reuse module receives a certain unissued merge request. Determining the sending order based on this time can reasonably arrange the processing order of requests and avoid long waiting of requests.

[0057] Exemplarily, there are three unissued merge requests A, B, and C. Request A is received by the data reuse module at 10:00:00 am, request B is received at 10:00:01 am, and request C is received at 10:00:02 am. Then when determining the request sending order, if the receiving time is given priority, request A will be sent to the request merging module first.

[0058] Among them, each unissued merge request has a corresponding destination address for specifying the location where the data or resource to be read / written by the merge request is located. The above address interval refers to the gap between the destination address of the current unissued merge request and the destination address of the previous issued merge request. If the destination address intervals of two merge requests to be merged are small, it means that the locations of the data or resources accessed by these two merge requests are closer. Then, giving priority to sending the unissued merge request with a smaller address interval to the request merging module can facilitate the subsequent request merging module to merge the merge requests to be merged.

[0059] Exemplarily, assume that the data reuse module has sent a merge - pending request to the request merging module, and its destination address is 0x1000. Now there are three un - sent merge - pending requests A, B, and C, with destination addresses 0x1050, 0x2000, and 0x2050 respectively. Among them, the address interval between the destination address 0x1050 of request A and the destination address 0x1000 of the previously sent request is 50; the address interval between the destination address 0x2000 of request B and the destination address 0x1000 of the previously sent request is 1000; the address interval between the destination address 0x2050 of request C and the destination address 0x1000 of the previously sent request is 1050. When determining the request sending order, if the address interval is given priority, then request C will be processed first because the address interval of request C is the smallest and it is the closest to the destination address of the previous request. This facilitates the subsequent request merging module to merge request C with the previous request (or subsequent adjacent requests), reducing overhead. Then process request A, and finally process request B.

[0060] Among them, the number of historical requests associated with the un - sent merge - pending request refers to the number of historical requests that are directly associated with the to - be - sent request in the association relationship established through a certain mechanism (such as the data reuse module ignoring the to - be - sent request, writing it as a historical request into the waiting request record unit, and establishing an association relationship with the target historical request) when there is a target historical request whose address range overlaps with the to - be - sent request. When the un - sent merge - pending request has an association situation such as an address range overlap with multiple historical requests, the number of its associated historical requests is the number of these overlapping associated historical requests. By considering the number of associated historical requests, the association complexity of the request in the historical request set can be understood, which helps to reasonably arrange the processing order of requests. Generally speaking, if the number of associated historical requests is large, it may mean that the destination address of the merge - pending request has been read / written multiple times. Correspondingly, the method further includes, when the request type of the merge - pending request is a write request, giving priority to sending the merge - pending requests with a smaller number; when the request type of the merge - pending request is a read request, giving priority to sending the merge - pending requests with a larger number.

[0061] Exemplarily, there are three unissued merge requests A, B, and C to be merged. Among them, merge request A has an overlapping address range association with one historical request, and the number of associated historical requests is 1; merge request B has an overlapping address range association with two historical requests, and the number of associated historical requests is 2; merge request C has an overlapping address range association with three historical requests, and the number of associated historical requests is 3. When the merge request is a write request, the merge request with a smaller quantity is preferentially issued. Therefore, write request A will be processed first because the number of its associated historical requests is the smallest, which can reduce frequent modification and potential conflicts of data and ensure data consistency and integrity. When the merge request is a read request, the merge request with a larger quantity is preferentially issued. Therefore, read request C will be processed first because the number of its associated historical requests is the largest, which can read more relevant data in one operation and improve the efficiency of data reading.

[0062] Of course, two or more of the above request attributes can also be combined to comprehensively judge each unissued merge request and determine the issuance order of each unissued merge request.

[0063] Based on the above embodiments disclosed in the present application, by determining the issuance order based on the request attribute of the time when the unissued merge request is received, the order of requests can be considered, so that the earlier received requests have a higher probability of being processed first, which helps to operate in the logical order of request arrival and avoid long-term backlog of requests; at the same time, by determining the issuance order based on the request attribute of the address interval between the destination address of the unissued merge request and the previously issued merge request, the distribution of requests in the address space can be considered. When the address interval is small, it indicates that the two requests may be in adjacent address areas, and preferentially processing such requests may help improve the merging efficiency of the subsequent request merging module for the merge requests; in addition, by determining the issuance order based on the number of historical requests associated with the unissued merge request, the frequency of reading / writing of the destination address corresponding to the merge request can be understood, and then the order can be determined based on the request type, which can solve potential conflicts or dependency problems between requests; based on the embodiments provided in the present application, multiple request attributes can be comprehensively considered to determine the issuance order of the merge requests, so as to arrange request processing more flexibly and reasonably, and improve the overall performance and resource utilization rate.

[0064] Figure 4 It is a schematic diagram of the device structure of an inter-chip interconnection device provided by an embodiment of the present application Figure 4 Based on Figure 1 , the request merging module 22 includes a merge waiting unit 221 and a merge processing unit 222; The merge waiting unit 221 is configured to receive and store the merge requests; The merging processing unit 222 is configured to perform merging processing on the to-be-merged request according to the destination address of the to-be-merged request, so as to obtain the merged request.

[0065] Wherein, the merged request is the result obtained after being processed by the merging processing unit, and is a new request formed by merging multiple to-be-merged requests.

[0066] In the above embodiment, by setting up a request merging module including a merging waiting unit and a merging processing unit, the received to-be-merged requests are managed. Among them, the merging waiting unit is responsible for receiving and storing the to-be-merged requests, providing a buffer area for the requests; the merging processing unit, according to the destination address of the to-be-merged requests, merges the requests with the same or similar destination addresses into a merged request, thereby reducing resource occupation and improving the efficiency of request processing.

[0067] In some possible implementation manners, when a to-be-merged request is received, the to-be-merged request is first sent to the merging waiting unit. The merging waiting unit assigns a unique identifier to each to-be-merged request and stores it in an internal storage structure, such as a queue or a list; the merging waiting unit sorts the to-be-merged requests according to certain rules (such as the arrival time order of the requests) for subsequent merging processing; the merging processing unit periodically or according to certain triggering conditions (such as the number of to-be-merged requests in the merging waiting unit reaches a certain threshold, or at every preset time interval) retrieves the to-be-merged requests from the merging waiting unit; then, the merging processing unit analyzes and compares the destination addresses of these to-be-merged requests, and merges the to-be-merged requests with consecutive destination addresses into a merged request. In some embodiments, during the merging process, the merging processing unit records the relevant information of each to-be-merged request and integrates it into the merged request.

[0068] Exemplarily, assume there are three to-be-merged requests: Request A, Request B, and Request C. Request A: The destination address is 0x1000, and the data length is 128 bytes. Request B: The destination address is 0x1080, and the data length is 64 bytes. Request C: The destination address is 0x1100, and the data length is 32 bytes. The merging waiting unit receives these three requests and stores them in the internal request queue. When the merging processing unit starts to process, it is found that the destination addresses of Request A and Request B are adjacent (and the data lengths are compatible), so Request A and Request B are merged into a new merged request. The destination address of the new merged request is 0x1000, and the data length is 192 bytes (128 bytes + 64 bytes). Request C remains a separate request because its destination address is not adjacent to the merged request.

[0069] Based on the above embodiments disclosed in the present application, the merge waiting unit 221 receives and stores the merge requests to be merged, which can accumulate a certain number of requests, provide a sufficient data source for subsequent merge processing, and improve the flexibility and efficiency of the merge processing. At the same time, the merge processing unit 222 performs merge processing according to the destination addresses of the merge requests to be merged, and can merge the requests with consecutive destination addresses into a merged request, thereby effectively reducing the number of requests in the inter-chip interconnection communication, reducing the communication overhead, and improving the efficiency and bandwidth utilization of data transmission.

[0070] Figure 5 is a schematic structural diagram of an inter-chip interconnection device provided by an embodiment of the present application Figure 5 Based on Figure 4 , the merge processing unit 222 includes an address comparison unit 2221 and an address calculation unit 2222. Among them, the address comparison unit 2221 is configured to compare the destination addresses of the merge requests to be merged stored in the merge waiting unit 221, divide the merge requests to be merged stored in the merge waiting unit 221 into at least one merge request group, and send the merge request group to the address calculation unit 2222; The address calculation unit 2222 is configured to determine the request information of the merged request based on the merge requests to be merged within the merge request group, and generate the merged request based on the request information.

[0071] In some possible implementation manners, the address comparison unit compares the destination addresses of each merge request to be merged stored in the merge waiting unit, divides the merge requests to be merged with consecutive destination addresses into the same merge request group; and sends the divided merge request group to the address calculation unit for subsequent processing. The address calculation unit receives the merge request group from the address comparison unit. Determines the request information of the merged request according to the merge requests to be merged within the group, including the destination address, data length, etc.; generates the merged request based on this request information, and sends the merged request to the inter-chip interconnection logic for transmission.

[0072] In some other possible implementation manners, the address comparison unit receives the merge requests to be merged from the merge waiting unit; compares the destination addresses of each merge request to be merged, divides the merge requests to be merged that can be merged into a first request group, and divides the requests that cannot be merged into a second request group separately. The address calculation unit receives the first request group and the second request group from the address comparison unit; for the first request group, determines the request information of the merged request according to the merge requests to be merged within the group, including the destination address, data length, etc.; for the second request group, directly uses one of the merge requests to be merged as the merged request.

[0073] In some embodiments, without limiting the packet length, all the merge requests with consecutive destination addresses are divided into the same merge request group to form the longest consecutive requests; in the case of limited packet length, according to the packet length limit, the merge requests that can be merged into the maximum packet length are divided into the same merge request group to obtain the most requests with the maximum packet length.

[0074] Exemplarily, assume that in the inter-chip interconnection communication, there are the following five merge requests: Request A: the destination address is 0x1000 and the data length is 256 bytes; Request B: the destination address is 0x1100 and the data length is 256 bytes; Request C: the destination address is 0x1400 and the data length is 64 bytes; Request D: the destination address is 0x1200 and the data length is 256 bytes; Request E: the destination address is 0x1300 and the data length is 64 bytes.

[0075] In some embodiments, the request information includes a request address and a data length; the request address is the first destination address that is logically earlier among the destination addresses of the merge requests in the merge request group; the data length of the merged request is the sum of the data lengths of the merge requests in the merge request group.

[0076] In the case of not limiting the packet length: the address comparison unit divides Request A, B, D, and E into the same merge request group (the first request group), and takes Request C as a merge request group (the second request group). Accordingly, the destination address of the first merge request group is 0x1000 and the data length is 832 bytes; the destination address of the second merge request group is 0x1400 and the data length is 64 bytes.

[0077] In the case of limiting the packet length to 512 bytes, the address comparison unit divides Request A and B into the same merge request group (the first request group), divides Request D and E into the same merge request group (the first request group), and takes Request C as a merge request group (the second request group). Accordingly, the destination address of the first merge request group is 0x1000 and the data length is 512 bytes; the destination address of the second merge request group is 0x1400 and the data length is 64 bytes; the destination address of the third merge request group is 0x1300 and the data length is 320 bytes.

[0078] Based on the above embodiments disclosed in the present application, the destination addresses of the merge-waiting requests stored in the merge-waiting unit 221 are compared by the address comparison unit 2221, and the merge-waiting requests are divided into at least one merge-waiting request group. In this way, requests with consecutive destination addresses can be effectively aggregated together, providing a basis for subsequent merge processing and helping to improve the efficiency of merge processing. At the same time, the address calculation unit 2222 determines the request information of the merged request based on the merge-waiting requests within the merge-waiting request group and generates the merged request, which can further reduce the number of requests to be transmitted and reduce the communication overhead.

[0079] Figure 6 is a schematic diagram of the device structure of an inter-chip interconnection device provided by an embodiment of the present application Figure 6 Based on Figure 5 , the merge processing unit 222 further includes a write data processing unit 2223, where the write data processing unit 2223 is configured to sequentially fetch the write data corresponding to each of the merge-waiting requests in the merge-waiting request group from the data reuse module 21 based on the address order of each of the merge-waiting requests in the merge-waiting request group and the write data address; the address calculation unit 2222 is further configured to construct the merged request based on the request information and the write data corresponding to each of the merge-waiting requests in the merge-waiting request group.

[0080] Among them, the above write data processing unit can sequentially fetch the corresponding write data from the data reuse module according to the address order of each of the merge-waiting requests in the merge-waiting request group and the write data address, providing necessary data support for the subsequent construction of the merged request.

[0081] In some possible implementation manners, the write data processing unit receives the merge-waiting request group from the address comparison unit, and sequentially fetches the corresponding write data from the write data cache 212 in the data reuse module according to the address order of each of the merge-waiting requests in the merge-waiting request group and the write data address; and transfers the fetched write data to the address calculation unit for constructing the merged request. The address calculation unit receives the merge-waiting request group information from the address comparison unit and simultaneously receives the write data corresponding to each of the merge-waiting requests from the write data processing unit; constructs the request information of the merged request based on the information of the merge-waiting request group (such as destination address, data length, etc.) and the write data; and combines the request information with the write data to generate a complete merged request.

[0082] Exemplarily, assume that in the inter-chip interconnect communication, there are the following five requests to be merged: Request A: The destination address is 0x1000, the data length is 256 bytes, and the starting address of the write data stored in the write data cache 212 of the data reuse module is 0x0000; Request B: The destination address is 0x1100, the data length is 256 bytes, and the starting address of the write data stored in the write data cache 212 is 0x0100. Request C: The destination address is 0x1400, the data length is 64 bytes, and the starting address of the write data stored in the write data cache 212 is 0x0200. Request D: The destination address is 0x1200, the data length is 256 bytes, and the starting address of the write data stored in the write data cache 212 is 0x0240. Request E: The destination address is 0x1300, the data length is 64 bytes, and the starting address of the write data stored in the write data cache 212 is 0x0340.

[0083] Without limiting the packet length, the write data processing unit sequentially fetches the write data corresponding to Requests A, B, D, and E, that is, starting from address 0x0000, continuously reads 832 bytes of data (256 bytes + 256 bytes + 256 bytes + 64 bytes); the address calculation unit constructs the merged request 1 based on the information of Requests A, B, D, and E (destination address is 0x1000, data length is 832 bytes) and the write data. At the same time, based on the information of Request C (destination address is 0x1400, data length is 64 bytes) and the corresponding write data, constructs the merged request 2.

[0084] When the packet length is limited, the write data processing unit: fetches the write data corresponding to Requests A and B, that is, starting from address 0x0000, continuously reads 512 bytes of data (256 bytes + 256 bytes), to prepare data for the first group of requests to be merged; fetches the write data corresponding to Requests D and E, that is, starting from address 0x0240, continuously reads 320 bytes of data (256 bytes + 64 bytes), to prepare data for the second group of requests to be merged; the address calculation unit constructs the merged request 1 based on the information of Requests A and B (destination address is 0x1000, data length is 512 bytes) and the write data; constructs the merged request 2 based on the information of Request C (destination address is 0x1400, data length is 64 bytes) and the corresponding write data; constructs the merged request 3 based on the information of Requests D and E (destination address is 0x1300, data length is 320 bytes) and the corresponding write data.

[0085] Based on the above embodiments disclosed in the present application, by the write data processing unit 2223 fetching the corresponding write data in sequence in the data reuse module 21 based on the address order of each request to be merged in the group of requests to be merged and the write data address, the relevant data of the requests to be merged can be effectively integrated.

[0086] In some embodiments, the data reuse module 21 is further configured to, in response to the request merging module 22 fetching the write data corresponding to the to-be-merged request, regard the historical request corresponding to the to-be-merged request in the waiting request recording unit 211 as an issued request.

[0087] It can be understood that if, in response to the request merging module 22 fetching the write data corresponding to the to-be-merged request, the historical request corresponding to the to-be-merged request in the waiting request recording unit 211 is not regarded as an issued request. Then, in the waiting request recording unit 211, the historical request corresponding to the to-be-merged request remains a historical request. So, when a new to-be-issued request arrives, if there is an overlap in the address range with the historical request corresponding to the to-be-merged request, it will be ignored, thereby resulting in missed requests.

[0088] Based on the above embodiments disclosed in the present application, by marking the historical request corresponding to the to-be-merged request in the waiting request recording unit as an issued request, the processing status of requests can be correctly tracked and managed, avoiding being ignored due to the overlap in the address range between the historical request not marked as an issued request and the new request when a new to-be-issued request arrives, thereby resulting in missed requests.

[0089] The above embodiments describe the process of receiving to-be-issued requests and generating corresponding merged requests. The second processor can generate and send a corresponding response message in response to the merged request. The following embodiments will illustrate the processing process of the response message.

[0090] In some embodiments, the request merging module 22 is further configured to, when receiving a response message corresponding to the merged request, split the response message to obtain sub-response messages corresponding to the respective to-be-merged requests corresponding to the merged request, and send the sub-response messages corresponding to the respective to-be-merged requests to the data reuse module 21; The data reuse module 21 is further configured to determine a target response message of the to-be-issued request corresponding to the to-be-merged request based on the sub-response message corresponding to the to-be-merged request, and feedback it to the request object that issues the to-be-issued request.

[0091] Among them, the response message is the result returned by the second processor after processing the merged request, and may include information obtained by processing the merged request, such as data reading results, operation status, etc.; the sub-response message is each part obtained after splitting the response message, and each part corresponds to one of the to-be-merged requests in the merged request. Correspondingly, the target response message corresponds to each to-be-issued request; the request object is the object that issues the to-be-issued request.

[0092] It can be understood that, according to the request type (read or write) of the request to be sent, the content of the response message will be different. If it is a read request, the response message is the read data read; if it is a write request, the response message is a write response message indicating whether the write is completed.

[0093] In some possible implementation manners, the request merging module receives the response message corresponding to the merged request; according to the request information corresponding to the merged request (such as the destination address, data length, etc. of the corresponding requests to be merged), the response message is split to obtain sub-response messages corresponding to each request to be merged, and the sub-response messages are sent to the data reuse module. The data reuse module receives the sub-response messages from the request merging module; according to the sub-response messages and the information in the waiting request record unit (such as the correspondence between the requests to be merged and the requests to be sent), the target response message of the request to be sent corresponding to the request to be merged is determined, and the target response message is fed back to the request object that sends the request to be sent.

[0094] Based on the above embodiments disclosed in the present application, when the request merging module receives the response message corresponding to the merged request, the response message is split to obtain sub-response messages corresponding to each request to be merged, and these sub-response messages are sent to the data reuse module. The data reuse module determines the target response message of the request to be sent corresponding to the request to be merged according to the sub-response messages, and feeds it back to the request object that sends the request to be sent. In this way, each request to be sent can receive its corresponding response message, even if the request to be sent may be ignored by the data reuse module or merged and processed by the request merging module.

[0095] Figure 7 is a schematic diagram of the device structure of an inter-chip interconnection device provided by an embodiment of the present application Figure 7 Based on Figure 4 , the request merging module 22 further includes a request merging record unit 223 and a response data processing unit 224; wherein, The address comparison unit 2221 is further configured to store the group of requests to be merged in the request merging record unit 223; The response data processing unit 224 is configured to obtain the group of requests to be merged corresponding to the merged request in the request merging record unit 223; based on the group of requests to be merged, the response message is split to obtain sub-response messages corresponding to each of the requests to be merged in the group of requests to be merged.

[0096] In some embodiments, during the sending process before the merged request is sent, after the above address comparison unit 2221 completes grouping, it is also necessary to store the group of requests to be merged in the request merging record unit 223.

[0097] In some possible implementation manners, when the response data processing unit receives the response message corresponding to the merged request, it obtains the information of the group of requests to be merged corresponding to the merged request from the request merging record unit; the response data processing unit splits the response message according to the information of the group of requests to be merged (such as the destination address, data length, etc. of the requests to be merged), and obtains the sub-response messages corresponding to each request to be merged.

[0098] Exemplarily, it is assumed that in the inter-chip interconnect communication, there are the following two requests to be sent (both are read requests): Request to be sent A: The destination address is 0x1000, the data length is 128 bytes, and it is sent by request object A; Request to be sent B: The destination address is 0x1080 (continuous with the destination address of request A), the data length is 128 bytes, and it is sent by request object B. The address comparison unit divides these two requests to be sent into a group of requests to be merged, and stores the group of requests to be merged and its related information in the request merging record unit. Then, these two requests are merged into a merged request and sent to the second processor. The second processor returns a response message, which contains the read data result of the merged request (i.e., requests A and B); after receiving the response message, the response data processing unit obtains the information of the group of requests to be merged from the request merging record unit; the response data processing unit splits the response message into two sub-response messages according to this information, where sub-response message 1 (corresponding to request A): 128 bytes of data read starting from address 0x1000; sub-response message 2 (corresponding to request B): 128 bytes of data read starting from address 0x1080.

[0099] Another exemplarily, it is assumed that in the inter-chip interconnect communication, there are the following two requests to be sent (both are write requests): Request to be sent C: The destination address is 0x1100, the data length is 64 bytes, and it is sent by request object C; Request to be sent D: The destination address is 0x1140, the data length is 128 bytes, and it is sent by request object D. The address comparison unit divides these two requests into a group of requests to be merged and sends them to the second processor for data writing. After the second processor completes the write operation, it returns a response message, that is, a write response message, which indicates whether the write operation is successful (for example, a status code or confirmation information, and does not contain the actually written data); after receiving the write response message, the response data processing unit obtains the information of the group of requests to be merged from the request merging record unit, and according to this information, the response data processing unit copies the write response message into two sub-response messages (since the write response message usually does not contain specific data, the copying operation is actually a copy of the write response status).

[0100] Based on the above embodiments disclosed in the present application, by dividing the requests to be merged into groups through the address comparison unit and storing them in the request merging record unit, and by splitting (or copying) the response message by the response data processing unit according to the information of the requests to be merged group and the content of the response message, each request to be merged can receive a response result matching its request type.

[0101] In some embodiments, the response data processing unit 224 is further configured to, in response to completing the splitting of the response message, delete the group of requests to be merged corresponding to the merged request in the request merging record unit 223.

[0102] It can be understood that the response data processing unit reduces the storage burden and improves the resource utilization efficiency by deleting the information of the group of requests to be merged corresponding to the merged request in the request merging record unit. This step aims to keep the request merging record unit clean and efficient, avoid the accumulation of invalid data, and thus improve the overall performance.

[0103] In some possible implementation manners, after receiving the response message corresponding to the merged request, the response data processing unit first splits the response message according to the information of the group of requests to be merged in the request merging record unit to obtain sub-response messages corresponding to each request to be merged. Then, the response data processing unit checks whether all the sub-response messages have been successfully split. Once it is confirmed that all the sub-response messages have been processed, the response data processing unit searches for the information of the group of requests to be merged corresponding to the merged request in the request merging record unit and deletes it.

[0104] Based on the above embodiments disclosed in the present application, by the response data processing unit 224 deleting the group of requests to be merged corresponding to the merged request in the request merging record unit 223 after completing the splitting of the response message, the storage burden of the request merging record unit can be reduced and the resource utilization efficiency can be improved; in addition, since the deletion operation is performed after all the sub-response messages have been processed, no important information will be lost, enhancing the stability and reliability.

[0105] In some embodiments, when the request type of the merged request is a read request, the response message includes the original read data; wherein, the response data processing unit 224 is further configured to, based on the group of requests to be merged, determine the destination addresses corresponding to each of the requests to be merged in the group of requests to be merged; and based on the destination addresses corresponding to each of the requests to be merged, split the original read data to obtain split read data corresponding to each of the sub-response messages.

[0106] Among them, the original read data is the data read from the target address included in the response message, and is the object to be split by the subsequent response data processing unit. In the current embodiment, the sub-response message contains the split read data, that is, the split read data.

[0107] In some possible implementation manners, after receiving the response message corresponding to the merged request (including the original read data), the response data processing unit obtains the information of the pending merge request group from the request merge record unit, including the destination address and data length of each pending merge request; according to this information, the response data processing unit extracts the read data segment corresponding to each pending merge request from the original read data in the order of the destination addresses of the pending merge requests, as the split read data; and then encapsulates the split read data into sub-response messages for subsequent processing or feedback.

[0108] Exemplarily, it is assumed that in the inter-chip interconnect communication, there are the following two pending requests (both are read requests): Pending request A: the destination address is 0x1000, and the data length is 128 bytes; Pending request B: the destination address is 0x1080 (continuous with the destination address of request A), and the data length is 128 bytes. The second processor returns a response message, which contains 256 bytes of original read data starting from address 0x1000 (that is, the sum of the data of requests A and B). After receiving the response message, the response data processing unit obtains the information of the pending merge request group from the request merge record unit. According to this information, the starting position and length of the split read data are determined: the read data of request A starts from address 0x1000 and has a length of 128 bytes; the read data of request B starts from address 0x1080 and has a length of 128 bytes; the response data processing unit extracts these two data segments from the original read data as the split read data and encapsulates them into two sub-response messages.

[0109] Based on the above embodiments disclosed in the present application, by splitting the original read data by the response data processing unit according to the information of the pending merge request group, the split read data corresponding to each sub-response message can be obtained; the read request processing flow of the inter-chip interconnect communication is optimized, and the communication efficiency is improved.

[0110] In some embodiments, the data reuse module 21 is further configured to obtain at least one pending request having an association relationship with the pending merge request; and generate a target response message corresponding to each pending request based on the sub-response message corresponding to the pending merge request.

[0111] Among them, the target response message is the processing result corresponding to each pending request generated by the data reuse module based on the sub-response message. The target response message is the response message corresponding to the pending request that is finally fed back to the request object.

[0112] In some possible implementation manners, after the data reuse module receives a request to be sent, it checks whether there is a target historical request in the waiting request record unit whose address range overlaps with the request to be sent. If there is an overlap, the data reuse module writes the request to be sent as a historical request into the waiting request record unit according to the implementation manner provided in the foregoing embodiments, and establishes an association relationship between the historical request corresponding to the request to be sent and the target historical request. In this way, when receiving a sub-response message corresponding to a request to be merged, the data reuse module splits the sub-response message onto the corresponding requests to be sent according to the previously stored association relationship, and generates a target response message for the requests to be sent.

[0113] Exemplarily, assume that in the inter-chip interconnect communication, there are the following two requests to be sent (both are read requests): Request to be sent A: destination address is 0x1000, data length is 128 bytes; Request to be sent B: destination address is 0x1000, data length is 128 bytes. During the sending process, if the request to be sent A is sent to the request merging module as a request to be merged and stored as historical request A; after receiving the request to be sent B, the data reuse module can determine that the address range of the request to be sent B overlaps with that of historical request A. Therefore, the request to be sent B will be ignored, stored as historical request B, and an association relationship with historical request A is established. During the receiving process, after obtaining the sub-response message corresponding to the request to be merged, based on the foregoing association relationship, all historical requests corresponding to the request to be merged can be found, that is, historical request A and historical request B (request to be sent A and request to be sent B), and then the sub-response message can be split onto the request to be sent A and the request to be sent B respectively to generate a target response message for the requests to be sent.

[0114] Based on the foregoing embodiments disclosed in the present application, when the data reuse module receives a sub-response message corresponding to a request to be merged, it can split the sub-response message onto the corresponding requests to be sent according to the previously stored association relationship, and generate a target response message for each request to be sent. This step ensures that each request to be sent can receive its corresponding processing result even if it is ignored (merged).

[0115] Figure 8 is a schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of the present application Figure 8 The data reuse module 21 is configured with a read data cache 214; the data reuse module 21 is further configured to, when the request type of the request to be sent is a read request, allocate a read data address for the request to be merged in the read data cache 214, and write the read data address into the waiting request record unit 211.

[0116] In some possible implementations, after receiving a request to be sent, the data reuse module may check its request type. If the request type is a read request, the data reuse module will allocate a read data address in the read data cache for the request to be merged and write the read data address to the waiting request record unit. In this way, during subsequent processing, the split read data corresponding to each of the requests to be merged can be stored in the read data cache according to this read data address, or the corresponding read data can be read from the read data cache.

[0117] In some embodiments, the data reuse module 21 is further configured to receive the split read data corresponding to each of the sub-response messages corresponding to the requests to be merged sent by the request merging module; obtain the read data addresses corresponding to each of the requests to be merged in the waiting request record unit 211; and store the split read data corresponding to each of the requests to be merged in the read data cache 214 based on the read data addresses corresponding to each of the requests to be merged.

[0118] In some possible implementations, after receiving the sub-response messages corresponding to each of the requests to be merged sent by the request merging module, the data reuse module obtains the split read data in the sub-response messages; the data reuse module accesses the waiting request record unit to obtain the read data addresses corresponding to each of the requests to be merged; and based on the read data addresses, the data reuse module stores the split read data corresponding to each of the requests to be merged in the read data cache.

[0119] Exemplarily, assume that in the inter-chip interconnection communication, there is a request to be sent A: the destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. The request to be sent B: the destination address is 0x1080, the data length is 128 bytes, and the request type is a read request. These two requests are merged into one request to be merged and sent to the second processor for processing. The data reuse module allocates a read data address, such as 0x10000, in the read data cache for the request to be merged and writes the address to the waiting request record unit. The second processor reads 256 bytes of data from the target address (covering the destination address ranges of requests A and B) and returns a response message; based on the foregoing embodiments, the data reuse module can receive the split read data corresponding to the sub-response message corresponding to the request to be merged sent by the request merging module and store the split read data in the read data cache based on the read data address 0x10000 corresponding to the request to be merged.

[0120] Figure 9 It is a schematic diagram of the device structure of an inter-chip interconnection device provided by an embodiment of the present application Figure 9The data reuse module further includes a read data distribution unit 213. When the request type of the request to be sent is a read request, the read data distribution unit 213 is configured to read the split read data corresponding to the request to be merged from the read data cache 214 based on the read data address corresponding to the request to be merged; for each request to be sent, a target response message corresponding to the request to be sent is generated based on the split read data corresponding to the request to be merged and the read response identifier.

[0121] In some possible implementation manners, the read data distribution unit may obtain the read data address corresponding to the request to be merged from the waiting request record unit; and read the split read data from the read data cache based on the read data address corresponding to the request to be merged. Then, the read data distribution unit generates a target response message corresponding to each request to be sent according to the read response identifier and the split read data, and the target response message is fed back to the corresponding request object.

[0122] It can be understood that for each request to be sent, when the address ranges read by the requests to be sent are the same, the complete and identical split read data is used in the process of generating the target response messages for the requests to be sent. Of course, when the address ranges read by the requests to be sent are different, partial read data can be obtained from the split read data based on the address range read by the request to be sent.

[0123] In some embodiments, the address range of the partial read data corresponding to the request to be sent in the read data cache may be determined based on the address range read by the request to be sent and the read data address corresponding to the request to be merged, and then the partial read data is obtained from the split read data stored in the read data cache based on the address range of the partial read data in the read data cache.

[0124] Exemplarily, it is assumed that in the inter-chip interconnection communication, there are two requests to be sent: Request to be sent A: The destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. Request to be sent B: The destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. These two requests are merged into one request to be merged and sent to the second processor for processing. The second processor returns a response message containing the complete split read data (128 bytes, covering the destination address ranges of requests A and B); after receiving the response message, since the address ranges read by requests to be sent A and B are the same (both are 0x1000 - 0x107F), the target response messages for these two requests can be directly generated based on the complete split read data.

[0125] Exemplarily, assume that in the inter-chip interconnect communication, there are the following three requests to be sent: Request C to be sent: the destination address is 0x1000, the data length is 128 bytes, and the request type is a read request. Request D to be sent: the destination address is 0x1000, the data length is 64 bytes, and the request type is a read request. Request E to be sent: the destination address is 0x1040, the data length is 64 bytes, and the request type is a read request. These three requests are merged into a request to be merged, with the destination address being 0x1000 and the data length being 128 bytes, and then sent to the second processor for processing. The second processor reads data starting from the target address 0x1000, covering the destination address ranges of requests C, D, and E, and returns a response message. The response message contains the complete split read data (a total of 128 bytes, covering the range of 0x1000 - 0x10FF), and stores it in the read data address corresponding to the pre-allocated request to be merged (taking 0x10000 as an example). For request C to be sent, since its data length is 128 bytes, the read data distribution unit obtains 128 bytes of data starting from 0x10000 from the read data cache as part of the target response message, that is, the complete split read data. For request D to be sent, since its data length is 64 bytes and the destination address is the same as that of request C, the read data distribution unit obtains 64 bytes of data starting from 0x10000 from the read data cache as part of the target response message, that is, the first half of the split read data. For request E to be sent, since its destination address is 0x1040, the read data distribution unit obtains 64 bytes of data starting from 0x10040 from the read data cache as part of the target response message, that is, the second half of the split read data.

[0126] Based on the above embodiments disclosed in the present application, the read data distribution unit can flexibly obtain some or all of the data from the split read data according to the address range and data length read by different requests to be sent, so as to generate the corresponding target response message. This not only improves the processing efficiency of read requests, reduces unnecessary data transmission and processing, but also ensures that each request to be sent can receive its corresponding target response message.

[0127] Figure 10 is a schematic diagram of the device structure of an inter-chip interconnect device provided by an embodiment of the present application Figure 10 . The device further includes: A low-latency bypass unit 50, configured to directly send a request with a low-latency identifier to the inter-chip interconnect logic when receiving a request from the first processor and the request has a low-latency identifier; or, when receiving a request from the first processor and the request does not have the low-latency identifier, determine the request without the low-latency identifier as the request to be sent.

[0128] Among them, the low-latency bypass unit is used to identify whether a low-latency identifier is carried in a request from the first processor, and determine whether to directly send the request to the inter-chip interconnection logic for processing or determine it as a request to be sent for subsequent processing according to the identifier. In this way, the processing delay of high-priority requests can be reduced and the response speed can be improved.

[0129] Exemplarily, assume that in inter-chip interconnection communication, the first processor generates three requests: Request 1 carries a low-latency identifier, Request 2 does not carry a low-latency identifier, and Request 3 does not carry a low-latency identifier. After receiving these three requests, the low-latency bypass unit checks whether each request carries a low-latency identifier. For Request 1, since it carries a low-latency identifier, the low-latency bypass unit will immediately send it directly to the inter-chip interconnection logic for processing. For Request 2 and Request 3, since they do not carry a low-latency identifier, the low-latency bypass unit determines them as requests to be sent and passes them to the data reuse module for subsequent processing.

[0130] Based on the above embodiments disclosed in this application, by using the low-latency bypass unit to identify and process requests with low-latency identifiers, the processing delay of high-priority requests can be effectively reduced and the response speed can be improved. This mechanism enables this application to give priority to processing requests with low-latency requirements and ensure the timely completion of key tasks.

[0131] Based on the foregoing embodiments, an embodiment of this application provides an inter-chip interconnection device. The device includes each unit included and each module included in each unit, and can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0132] An embodiment of this application provides an electronic device, including a plurality of processors and an inter-chip interconnection device corresponding to each of the plurality of processors as described in the above embodiments.

[0133] In some implementation scenarios, the first processor is the current GPU and the second processor is the peer GPU. Please refer to Figure 11, which shows a schematic diagram of the interconnection structure in the inter-GPU interconnection scenario provided by this application. Among them, the current GPU 1110 may include multiple GPU cores. The inter-chip interconnection device in the above embodiment may include a low-latency bypass unit 1130, a data reuse module 1140, and a request merging module 1150. Correspondingly, the current GPU 1110 may be connected to the low-latency bypass unit 1130 in the inter-chip interconnection device through the on-chip interconnection bus 1120; the low-latency bypass unit 1130, the data reuse module 1140, and the request merging module 1150 are connected in sequence; the inter-chip interconnection device is connected to the peer GPU 1170 through the inter-chip interconnection logic 1160.

[0134] Continue to refer to Figure 11 , requests from the current GPU 1110 pass through the data reuse module 1140, and it is judged whether the data can be shared with the requests that have been sent according to the operation type of the request. If the read request address has been accessed by the requests that have been sent, then after the data of the requests that have been sent is returned, this part of the data can be directly retrieved and returned; or if there are corresponding requests that have not been sent waiting to be sent for the write request address, then the data of the write request waiting to be sent can be directly overwritten to reduce the actual cross-card communication. After the request passes through the data reuse module, it enters the request merging module 1150. The request merging module 1150 will combine the requests according to all the request information collected within a period of time, calculate the longest possible packet length, and communicate through the inter-chip interconnection logic 1160.

[0135] For reverse access, such as the return of read data and the reception of write data, the reverse order is followed. For the received write data, the request merging module 1150 splits the write data into several independent requests according to the bus protocol, and sends them to the data bus through the data reuse module 1140 and the low-latency bypass unit 1130. For the returned read data, it is sent to the data reuse module 1140 according to the data information recorded during request integration. When the data reuse module 1140 receives the returned data, it will split and send the data to each request object according to the previously recorded request reuse relationship (i.e., the association relationship in the above embodiment). The low-latency bypass unit 1130 receives low-latency requests with special identifiers (i.e., the low-latency identifiers in the above embodiment), skips the aforementioned data reuse module 1140 and request merging module 1150, and directly sends them to the inter-chip interconnection logic 1160 to minimize the access latency.

[0136] Please refer to Figure 12, the data reuse module 1210 includes: a wait request record unit 1213, a write data cache 1214, a read data cache 1211, and a read data distribution unit 1212. Among them, the write data merging logic is as follows: after a write request enters the data reuse module 1210, it queries the wait request record unit 1213 to check whether there is a write request that is still in the write data cache 1214. If there is, it takes out the write data address corresponding to this write data, overwrites the write data to the buffer corresponding to the write data address of the unfinished write request recorded in the wait request record unit 1213, and merges the current write request with the wait request. If not, it separately records this write request, allocates a new write data address, and writes the write data into the write data cache 1214. The write data cache 1214 will select write data according to specific rules and send it to the next-level request merging module 1220 for request merging. Possible rules include the request that arrives first has the highest priority; the request with the smallest difference in address from the previous request sent has the highest priority; the request with the most merge occurrences has the highest priority, and so on.

[0137] After a read request enters the data reuse module 1210, it queries the wait request record unit 1213. If the wait request record unit 1213 already contains the data required by this request, it only records the read request in the wait request record unit 1213, instead of actually sending this request to the request merging module 1220, and marks the previous request number that this request is waiting for.

[0138] The behavior of the read data distribution unit 1212 is as follows: when a read request returns, it reports the number of the returned request to the wait request record unit 1213. The remaining wait requests (i.e., the ignored historical requests) will identify whether they have obtained the data they need based on this number, and find the original waiting request in the wait request record unit 1213 based on this number. Then, they can obtain the position of the data they need in the read data cache 1211, read the read data from the read data cache 1211, add the read response identification information corresponding to the request, and return it to the corresponding request object. When all possible read data distributions for a certain address are completed, the corresponding read data cache space is released.

[0139] Please refer to Figure 12 , the request merging module is as Figure 12 shown, and includes: a request merging record unit 1223, a merge waiting unit 1221, an address calculation unit 1226, an address comparison unit 1222, a write data processing unit 1225, and a response data processing unit 1224.

[0140] After the request from the data reuse module 1210 (i.e., the request to be merged) enters the request merging module 1220, it is written into the merge waiting unit 1221. The address comparison unit 1222 will compare the addresses of all requests of the same type in the merge waiting unit 1221, select a request sequence that can generate the longest continuous requests, send it to the address calculation unit 1226 to calculate information such as the packet length and the starting address, and at the same time record the information of this group of requests into the request merge record unit 1223. After the read data returns from the peer, the request merge record unit 1223 will be queried to obtain the requests to be merged corresponding to the merged requests (corresponding to the merged requests in the above embodiments), and split according to the requests to be merged and sent to the data reuse module 1210. For the merged write requests, the write data processing unit 1225 will, according to the address of the sorted write data in the write data cache 1214 of the data reuse module, fetch the write data from the data reuse module 1210, pack it and send it to the inter-chip interconnection logic.

[0141] The above data reuse module 1210 and the request merging module 1220 will occupy more logic levels, resulting in additional communication delays. Therefore, for the data with low latency requirements in cross-card communication, it will be handed over to the low latency bypass unit and directly sent to the inter-chip interconnection logic to achieve low latency transmission. The low latency bypass unit will be the front-end input of the entire system, and judge whether the request belongs to the request with low latency requirements according to information such as the priority signal on the bus, and select to directly bypass and send it to the inter-chip interconnection logic or send it to the data reuse module 1210.

[0142] Exemplarily, the current GPU has 4 computing cores, and each core needs to read 128 bytes of data from the video memory of another GPU and write 32 bytes of data. The 4 computing cores need to issue the following read requests and write requests: Core 0: Read address of read request 1: 0x0; Read address of read request 2: 0x80; Write address of the write request is 0x1000; Core 1: Read address of read request 1: 0x80; Read address of read request 2: 0x100; Write address of the write request is 0x1020; Core 2: Read address of read request 1: 0x100; Read address of read request 2: 0x180; Write address of the write request is 0x1040; Core 3: Read address of read request 1: 0x180; Read address of read request 2: 0x200; Write address of the write request is 0x1060; where the read address of the read request is the destination address of the read request; the write address of the write request is the destination address of the write request.

[0143] The data reuse module successively receives the first read requests of each core, records them in the waiting request record unit, and sends them to the request merging module. Next, it receives the second read requests of each core. After comparing and judging with the read addresses recorded in the waiting request record unit, it is found that the second requests of cores 0, 1, and 2 (0x80, 0x100, 0x180) have already been requested, so the requests are not sent repeatedly, but are recorded in the waiting request record unit, and the numbers of the waiting requests are marked. However, the second request of core 3 (0x200) has not been requested yet, so it will be sent to the request merging module. The waiting request record unit can ignore the read request 1 of cores 1-3, add the request identifier of the read request 1 of core 1 to the read request 2 of core 0, add the request identifier of the read request 1 of core 2 to the read request 2 of core 1, add the request identifier of the read request 1 of core 3 to the read request 2 of core 2; and, record the repetition number as 1 on the read requests 2 of cores 0-2. The read request 2 of core 3 is a new request to be merged; thus, a total of 8 requests are recorded in the waiting request record unit, including 5 requests to be merged and 3 requests that are ignored and sent.

[0144] The request merging module will receive a total of 5 requests to be merged and fill them into the merge waiting unit, which are 0x0, 0x80, 0x100, 0x180, and 0x200 respectively. After being judged by the address comparison unit, the data required by these 5 requests is adjacent before and after, so the information of these 5 requests is sent to the address calculation unit, and the information of these 5 requests is recorded in the request merging record unit. After calculation by the address calculation unit, it is determined that the starting address of the finally sent packet is 0x0, and the packet length is 0x280 (640) bytes, and the packet is sent to the inter-chip interconnection logic.

[0145] After the read requests are returned from the inter-chip interconnection logic, the request merging module searches for the original 5 requests to be merged in the request merging record unit according to the numbers of the returned read requests, and splits the original read data obtained according to the address information of the requests to be merged, and fills it into the corresponding read data cache. After the splitting is completed, the resources in the request merging record unit are released.

[0146] After the data reuse module receives the returned read data, it searches for the waiting request record unit according to the identification information carried by the read data, and broadcasts the table item number to all waiting request objects. Taking the data corresponding to the return address 0x100 as an example, the broadcast of its corresponding number (02) will inform the second read request of core 1 that the data it needs is also ready. The read data distribution unit will accordingly search for the address of the corresponding read data cache in item 02 of the waiting request record unit and retrieve the data to return to core 1 as the return data of its second read request.

[0147] For four core write requests, after the four requests enter the waiting request record unit and find that there is no address overlap with each other, they are successively sent to the request merging module. After detecting that these four pieces of data are consecutive addresses in the merging waiting unit, the request merging module merges these four write requests with a length of 32 bytes into one write request with a length of 128 bytes, and fetches the write data from the write data cache in the address order to form a packet and send it to the inter-chip interconnect logic.

[0148] After core 3 receives all its read data, core 3 will send a signal to inform the peer GPU. This signal is expected to have as low a latency as possible. Core 3 marks this request with a priority signal. After this request is detected by the low-latency bypass unit, it directly bypasses the data reuse module and the request merging module and is sent to the inter-chip interconnect logic to obtain as low a latency as possible.

[0149] Thus, the data reuse module reduces the total amount of data that needs to be communicated between chips. The request merging module brings a longer packet length for inter-chip communication, thereby improving the actual data bandwidth utilization efficiency of inter-chip communication; the low-latency bypass unit solves the additional latency introduced by the data reuse module and the request merging module, so that the low-latency access requirements can still obtain the expected low latency.

[0150] Figure 13 The schematic diagram of the hardware entity of an electronic device provided by an embodiment of the present application is shown in Figure 13 As shown, the hardware entity of the electronic device 1300 includes: a first processor 1310, an inter-chip interconnect device 1320, an inter-chip interconnect logic 1330, and a second processor 1340. Among them, the inter-chip interconnect device 1320 can be understood with reference to the description of the device embodiment of the present application.

[0151] The above-mentioned first processor and second processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device implementing the above processor functions can also be other, and the embodiments of the present application do not make specific limitations.

[0152] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above steps / processes do not mean the sequence of execution, and the execution sequence of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0153] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0154] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0155] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0156] In addition, each functional unit in the embodiments of the present application may be entirely integrated into one processing unit, or each unit may be separately regarded as one unit, or two or more units may be integrated into one unit; the above-mentioned integrated unit may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program may be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical discs, etc., which can store program codes of various types.

[0157] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical discs, etc., which can store program codes of various types.

[0158] The above are only the implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.

Claims

1. An inter-chip interconnection device, characterized in that: The inter-chip interconnection device is arranged between the first processor and the inter-chip interconnection logic, and the inter-chip interconnection device includes a data reuse module and a request merging module, wherein: The data reuse module is used to receive the request to be sent, and determine the request to be merged based on the destination address of the request to be sent; The data reuse module is further used to send the request to be merged to the request merging module; The request merging module is used to merge the received requests to be merged to obtain a merged request, and send the merged request to the inter-chip interconnection logic; the inter-chip interconnection logic is used to send the merged request to the second processor.

2. The chip interconnection device according to claim 1, characterized in that: The data reuse module includes a waiting request recording unit, which is used to store historical requests in the inter-chip communication process and historical addresses of the historical requests; wherein, The data reuse module is further used to determine the request to be merged based on the destination address of the request to be sent and the historical address of the historical request; The data reuse module is further used to update the waiting request recording unit based on the request to be issued.

3. The inter-chip interconnection device according to claim 2, characterized in that: The data reuse module is further configured to determine the request to be issued as the request to be merged when there is no target historical request with an address range overlapping with the request to be issued; The data reuse module is further configured to write the request to be issued as a historical request into the waiting request recording unit.

4. The chip-to-chip interconnection device according to claim 3, characterized in that: The data reuse module includes a write data cache; The data reuse module is further configured to allocate a new write data address for the request to be issued in the write data cache when the request type of the request to be issued is a write request; and write the write data corresponding to the request to be issued into the write data cache based on the new write data address; And write the new write data address into the waiting request recording unit.

5. The chip-to-chip interconnection device according to claim 2, characterized in that: The data reuse module is further configured to ignore the request to be issued if there is a target historical request whose address range overlaps with the request to be issued; The data reuse module is further configured to write the request to be issued as a historical request into the waiting request recording unit, and establish an association relationship between the historical request corresponding to the request to be issued and the target historical request.

6. The inter-chip interconnection device according to claim 5, characterized in that: The data reuse module is further configured to obtain a target write data address when the request type of the request to be issued is a write request; based on the target write data address, write the write data corresponding to the request to be issued into the write data cache; The target write data address is the write data address of the earliest historical request whose address range overlaps with the request to be issued.

7. The inter-chip interconnection device according to claim 2, characterized in that: The data reuse module is further used to sort the unissued requests to be merged to obtain an issuance order; and send the unissued requests to be merged to the request merging module in sequence according to the issuance order.

8. The inter-chip interconnection device according to claim 7, characterized in that: The issuance order is determined based on the request attributes of the unissued request to be merged; the request attributes include at least one of the following: the time of receiving the unissued request to be merged, the address interval between the destination address of the unissued request to be merged and the last issued request to be merged, and the number of historical requests associated with the unissued request to be merged.

9. The inter-chip interconnection device according to any one of claims 1 to 8, characterized in that: The request merging module includes a merging waiting unit and a merging processing unit; The merging waiting unit is used to receive and store the request to be merged; The merging processing unit is used to merge the requests to be merged according to the destination addresses of the requests to be merged to obtain the merged requests.

10. The inter-chip interconnection device according to claim 9, characterized in that: The merging processing unit includes an address comparing unit and an address calculating unit, wherein: The address comparison unit is used to compare the destination addresses of the requests to be merged stored in the merge waiting unit, divide the requests to be merged stored in the merge waiting unit into at least one request group to be merged, and send the request group to be merged to the address calculation unit; The address calculation unit is used to determine the request information of the merged request based on the requests to be merged in the request group to be merged, and generate the merged request based on the request information.

11. The chip-to-chip interconnection device according to claim 10, characterized in that: The request information includes a request address and a data length; the request address is the first logically preceding destination address among the destination addresses of the requests to be merged in the request group to be merged; the data length of the merged request is the sum of the data lengths of the requests to be merged in the request group to be merged.

12. The chip-to-chip interconnection device according to claim 10, characterized in that: In the case where the request type of the request to be merged is a write request, the merge processing unit further includes a write data processing unit, wherein: The write data processing unit is used to sequentially retrieve the write data corresponding to each of the requests to be merged in the request group to be merged and the write data address based on the address sequence of each of the requests to be merged in the request group to be merged and the write data address; The address calculation unit is further used to construct the merged request based on the request information and the write data corresponding to each of the requests to be merged in the group of requests to be merged.

13. The chip-to-chip interconnection device according to claim 9, characterized in that: The data reuse module is further configured to, in response to the request merging module taking out the write data corresponding to the request to be merged, take the historical request corresponding to the request to be merged in the waiting request recording unit as the issued request.

14. The inter-chip interconnection device according to any one of claims 1 to 8, characterized in that: The request merging module is further configured to, upon receiving a response message corresponding to a merged request, split the response message to obtain sub-response messages corresponding to each of the requests to be merged corresponding to the merged request, and send the sub-response messages corresponding to each of the requests to be merged to the data reuse module; The data reuse module is further used to determine the target response message of the request to be issued corresponding to the request to be merged based on the sub-response message corresponding to the request to be merged, and feed it back to the request object that issued the request to be issued.

15. The chip-to-chip interconnection device according to claim 14, characterized in that: The request merging module includes an address comparison unit; the request merging module also includes a request merging recording unit and a response data processing unit; wherein, The address comparison unit is further used to store the request group to be merged in the request merging recording unit; The response data processing unit is used to obtain the group of requests to be merged corresponding to the merged request in the request merging record unit; based on the group of requests to be merged, split the response message to obtain sub-response messages corresponding to each of the requests to be merged in the group of requests to be merged.

16. The chip-to-chip interconnection device according to claim 15, characterized in that: The response data processing unit is further configured to delete the to-be-merged request group corresponding to the merged request in the request merging record unit in response to completing the splitting of the response message.

17. The chip-to-chip interconnection device according to claim 15, characterized in that: In the case where the request type of the merged request is a read request, the response message includes original read data; wherein, The response data processing unit is also used to determine the destination address corresponding to each of the requests to be merged in the request group to be merged based on the request group to be merged; based on the target address corresponding to each of the requests to be merged, the original read data is split to obtain the split read data corresponding to each of the sub-response messages.

18. The chip-to-chip interconnection device according to claim 14, characterized in that: The data reuse module is further used to obtain at least one request to be issued that is associated with the request to be merged; and generate a target response message corresponding to each request to be issued based on the sub-response message corresponding to the request to be merged.

19. The chip-to-chip interconnection device according to claim 18, characterized in that: The data reuse module is configured with a read data cache; The data reuse module is further configured to allocate a read data address for the request to be merged in the read data cache and write the read data address into the waiting request recording unit when the request type of the request to be issued is a read request.

20. The chip-to-chip interconnection device according to claim 19, characterized in that: The data reuse module is also used to receive the split read data corresponding to the sub-response message corresponding to each of the requests to be merged sent by the request merging module; obtain the read data address corresponding to each of the requests to be merged in the waiting request recording unit; and based on the read data address corresponding to each of the requests to be merged, store the split read data corresponding to each of the requests to be merged in the read data cache.

21. The inter-chip interconnection device according to claim 20, further comprising a read data distribution unit, characterized in that: The read data distribution unit is configured to read the split read data corresponding to the request to be merged in the read data cache based on the read data address corresponding to the request to be merged when the request type of the request to be issued is a read request; For each of the requests to be issued, a target response message corresponding to the request to be issued is generated based on the split read data and the read response identifier corresponding to the request to be merged.

22. The inter-chip interconnection device according to any one of claims 1 to 8, characterized in that: The inter-chip interconnection device also includes: A low-latency bypass unit is used to, when receiving a request from the first processor and the request has a low-latency flag, directly send the request with a low-latency flag to the inter-chip interconnection logic; or, when receiving a request from the first processor and the request does not have the low-latency flag, determine the request without a low-latency flag as the request to be issued.

23. An electronic device, characterized in that: include: A plurality of processors and an inter-chip interconnection device as claimed in any one of claims 1 to 22 corresponding one to each of the plurality of processors.

Citation Information

Patent Citations

  • Cache management request fusing

    CN105793832A

  • Method and device for improving performance of distributed file system

    CN111737212A

  • Memory access control structure and method, memory system, processor and electronic equipment

    CN116680089A

  • Data reading method for chip, chip, computer equipment and storage medium

    CN117785043A

  • Packet Combiner for a Packetized Bus with Dynamic Holdoff time

    US20070079044A1