Artificial intelligence chip and operation method thereof, multi-chip system, and electronic device

CN122817153APending Publication Date: 2026-09-25SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611318901.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]然而,在采用亲和性多路径桥接的场景中,难以确保指向同一地址粒度区间的读写请求之间的RAW(Read After Write,读操作在写操作完成之后)一致性

Benefits of technology

[0038]基于本申请的实施例,AI芯片的桥接模块可以利用在写请求等待列表中对应任意写请求的记录项,对与该写请求通过同一交换路径指向同一地址粒度区间的读请求实施使能传输限制,以在该地址粒度区间对应的交换路径上维持读写请求之间的RAW一致性。而且,AI芯片的桥接模块还可以在任意地址粒度区间对应的交换路径发生故障时,立即触发指向该地址粒度区间的写请求和读请求向替代交换路径迁移,同时,即便写请求等待列表中的记录项由于迁移的写请求的交换路径改变而失效,桥接模块也可以通过在迁移窗口列表中添加与该地址粒度区间对应的窗口项,以继续维持迁移到替代交换路径上的读写请求之间的RAW一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817153A_ABST
    Figure CN122817153A_ABST
Patent Text Reader

Abstract

The application relates to an artificial intelligence chip, an operation method thereof, a multi-chip system and an electronic device. Based on the application, a bridge module of the artificial intelligence chip can utilize a record item corresponding to an arbitrary write request in a write request waiting list to implement transmission enabling restriction on a read request directed to a same address granularity interval through a same exchange path as the write request, so as to maintain RAW consistency between read and write requests on an exchange path corresponding to the address granularity interval; the bridge module of the artificial intelligence chip can also immediately trigger migration of a write request and a read request directed to an address granularity interval to an alternative exchange path when a fault occurs in an exchange path corresponding to the address granularity interval, and simultaneously, the bridge module can also add a window item corresponding to the address granularity interval in a migration window list, so as to continue to maintain RAW consistency between read and write requests migrated to the alternative exchange path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of AI (Artificial Intelligence) chips, and in particular to an AI chip, an operating method of an AI chip, a multi-chip system, and an electronic device. Background Technology

[0002] To meet the computational demands of tasks such as training large models, multiple AI chips can be interconnected via a switching network, enabling these tasks to be completed collaboratively by the chips. In this scenario, collaboration between multiple AI chips requires cross-chip read / write operations.

[0003] Specifically, the on-chip bus of each AI chip can be bridged with multiple switching paths in an off-chip switching network through the AI ​​chip's bridging module, allowing write and read requests transmitted from one AI chip to other AI chips to be distributed across multiple switching paths. In this case, write and read requests pointing to the same address granularity range (i.e., the same address granularity range of the on-chip memory of another AI chip) can be assigned to the same specified switching path to achieve affinity multipath bridging based on address granularity range.

[0004] However, in scenarios employing affinity multipath bridging, it is difficult to ensure RAW (Read After Write) consistency between read and write requests pointing to the same address granularity range. RAW consistency between read and write requests pointing to the same address granularity range means that a read request generated after a write request should be completed (i.e., arrive at the destination) after the write request.

[0005] Therefore, how to maintain RAW consistency between read and write requests based on affinity multipath bridging has become a technical problem to be solved.

[0006] It is understood that the content of the "Background Art" section is intended to aid in understanding this disclosure. Some (or all) of the content disclosed in the "Background Art" section may not be known to those skilled in the art. The content disclosed in the "Background Art" section does not imply that such content was known to those skilled in the art prior to this disclosure. Summary of the Invention

[0007] Embodiments of this application provide an AI chip, an operating method for an AI chip, a multi-chip system, and an electronic device that help maintain RAW consistency between read and write requests based on affinity multipath bridging.

[0008] In one embodiment of this application, an AI chip is provided, comprising:

[0009] Processing module;

[0010] The bridging module is configured to interact with the processing module via an on-chip bus, and to evenly distribute write requests and read requests generated based on the on-chip interaction to multiple switching paths in an off-chip switching network for transmission. The even distribution mechanism includes: write requests and read requests pointing to the same address granularity range are assigned to the same specified switching path.

[0011] The bridging module is further configured as follows:

[0012] During the period of dynamically maintaining the write request wait list, the path status of multiple swap paths is checked;

[0013] In response to the detection that any switching path is in a fault state, a mechanism is triggered to migrate pending write and read requests that are directed to the same address granularity range via that switching path to an alternative switching path selected in the switching network. Furthermore, based on the write request wait list, a window entry corresponding to the address granularity range to which the migrated write and read requests are directed is added to the migration window list. This window entry is deleted only after all write requests to that address granularity range have been successfully retransmitted, thereby restricting the migration of read requests to that address granularity range to be enabled for transmission only after the migration of write requests to that address granularity range has been successfully retransmitted via the alternative switching path.

[0014] In some examples, optionally, the bridging module is specifically configured to perform the operation of adding window items to the migration window list by: searching the write request wait list for all records corresponding to write requests with any exchange path currently detected as faulty as the specified exchange path; using the records found in the write request wait list, adding window items to the migration window list corresponding to the address granularity range pointed to by the migrated write and read requests; wherein, the record item in the write request wait list corresponding to any write request is used to: restrict read requests generated after the write request and pointing to the same address granularity range via the same specified exchange path as the write request to be enabled for transmission after the write request is successfully transmitted.

[0015] In some examples, optionally, the record entry corresponding to any write request in the write request wait list includes: the path identifier of the specified exchange path to which the write request was assigned when it was generated; the bridging module is specifically configured to perform the operation of querying record entries in the write request wait list in the following manner: using the path identifier of any exchange path currently detected as faulty as an index, search the write request wait list for all record entries corresponding to write requests with that exchange path as the specified exchange path.

[0016] In some examples, optionally, the record entry corresponding to any write request in the write request wait list includes: a range identifier of the address granularity range pointed to by the write request; the bridging module is specifically configured to perform the operation of adding a window entry to the migration window list using the searched record entries in the following manner: identifying the range identifier in the record entries searched in the write request wait list; and creating a window entry corresponding to an address granularity range represented by the range identifier based on all record entries in the write request wait list that include the same range identifier.

[0017] In some examples, optionally, the window entry in the migration window list corresponding to any address granularity interval includes: an interval identifier for the address granularity interval, a path identifier for an alternative switching path pointing to the address granularity interval for write and read requests, and a number of retransmissions pending for write requests pointing to the address granularity interval; wherein, the initial value of the number of retransmissions pending in the window entry corresponding to any address granularity interval is determined by the number of record entries on which the window entry is based; and, the current value of the number of retransmissions pending in the window entry corresponding to any address granularity interval decreases in response to an increase in the number of successful retransmissions of write requests pointing to any address granularity interval via alternative switching paths.

[0018] In some examples, optionally, the window entry in the migration window list corresponding to any address granularity interval includes: an interval identifier for the address granularity interval, a path identifier for the alternative switching path pointing to the write request and read request of the address granularity interval, and a number of retransmissions pending for the write request pointing to the address granularity interval; wherein, the initial value of the number of retransmissions pending in the window entry corresponding to any address granularity interval is used to characterize the total number of migrations of write requests pointing to the address granularity interval; the bridging module is specifically configured to perform the operation of restricting the migration of read requests in the following manner: in response to the successful retransmission of any write request pointing to any address granularity interval through the alternative switching path, decrementing the number of retransmissions pending in the window entry in the migration window list corresponding to the address granularity interval by one; and determining the release timing of enabling transmission of the migrated read requests based on the current value of the number of retransmissions pending in the window entry corresponding to any address granularity interval in the migration window list.

[0019] In some examples, the dynamic maintenance of the write request wait list may optionally include: adding an entry corresponding to the write request in response to the generation of any write request; deleting an entry corresponding to the write request in response to the successful transmission of any write request; and deleting an entry added to the write request wait list in response to the successful retransmission of any write request pointing to any address granularity range via an alternative switching path.

[0020] In some examples, the bridging module is optionally configured to perform the operation of releasing the enable transmission of the migration read requests by determining the timing of the migration in such a way that, in response to the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity interval in the migration window list becoming 0, the enable transmission of all read requests of the migration is released in batches.

[0021] In some examples, the bridging module is optionally configured to perform the operation of determining the release timing of the enabled transmission of the migrated read requests by releasing the enabled transmission of each migrated read request one by one according to the sorting relationship between the migrated read requests and write requests during the change of the current value of the number of retransmissions in the window item corresponding to any address granularity interval in the migration window list.

[0022] In some examples, the address granularity range may optionally be a memory page, the size of which is 4KB.

[0023] In some examples, optionally, any switching path in the switching network is used to implement switching forwarding with flow control based on PDC, and the bridging module is specifically configured to: determine whether the write request has been successfully transmitted or successfully retransmitted in response to receiving a PDS ACK carrying a PSN carrying an arbitrary write request, and confirm that the path state of the specified switching path carrying the write request is in a fault state in response to a timeout in receiving a PDS ACK carrying a PSN carrying an arbitrary write request or in response to receiving a PDS NACK carrying a PSN carrying an arbitrary write request.

[0024] In another embodiment of this application, an operation method for an AI chip is provided. The AI ​​chip includes a processing module and a bridging module. The bridging module interacts with the processing module via an on-chip bus. The bridging module evenly distributes write requests and read requests generated based on the on-chip interaction to multiple switching paths in an off-chip switching network for transmission. The mechanism for even distribution includes: write requests and read requests pointing to the same address granularity range are assigned to the same specified switching path.

[0025] The operation method includes the following steps performed by the bridging module:

[0026] During the period of dynamically maintaining the write request wait list, the path status of multiple swap paths is checked;

[0027] In response to the detection that any switching path is in a fault state, a mechanism is triggered to migrate pending write and read requests that are directed to the same address granularity range via that switching path to an alternative switching path selected in the switching network. Furthermore, based on the write request wait list, a window entry corresponding to the address granularity range to which the migrated write and read requests are directed is added to the migration window list. This window entry is deleted only after all write requests to that address granularity range have been successfully retransmitted, thereby restricting the migration of read requests to that address granularity range to be enabled for transmission only after the migration of write requests to that address granularity range has been successfully retransmitted via the alternative switching path.

[0028] Optionally, in some examples, adding window entries to the migration window list corresponding to the address granularity ranges pointed to by the migrated write and read requests based on the write request wait list includes: searching the write request wait list for all records corresponding to write requests with any exchange path currently detected as faulty as the specified exchange path; using the records found in the write request wait list, adding window entries to the migration window list corresponding to the address granularity ranges pointed to by the migrated write and read requests; wherein, the record entries in the write request wait list corresponding to any write request are used to: restrict read requests generated after the write request and pointing to the same address granularity range via the same specified exchange path as the write request to be enabled for transmission after the write request is successfully transmitted.

[0029] In some examples, optionally, the record entry corresponding to any write request in the write request wait list includes: the path identifier of the specified switch path to which the write request was assigned when it was generated; the step of searching the write request wait list for all records corresponding to write requests with any switch path currently detected as faulty as the specified switch path includes: using the path identifier of any switch path currently detected as faulty as an index, searching the write request wait list for all records corresponding to write requests with that switch path as the specified switch path.

[0030] In some examples, optionally, the record entry corresponding to any write request in the write request wait list includes: a range identifier of the address granularity range pointed to by the write request; the step of adding a window entry corresponding to the address granularity range pointed to by the migrated write request and read request in the migration window list using the record entries searched in the write request wait list includes: identifying the range identifier in the record entries searched in the write request wait list; and creating a window entry corresponding to an address granularity range represented by the range identifier based on all record entries in the write request wait list that include the same range identifier.

[0031] In some examples, optionally, the window entry in the migration window list corresponding to any address granularity interval includes: an interval identifier for the address granularity interval, a path identifier for an alternative switching path pointing to the address granularity interval for write and read requests, and a number of retransmissions pending for write requests pointing to the address granularity interval; wherein, the initial value of the number of retransmissions pending in the window entry corresponding to any address granularity interval is determined by the number of record entries on which the window entry is based; and, the current value of the number of retransmissions pending in the window entry corresponding to any address granularity interval decreases in response to an increase in the number of successful retransmissions of write requests pointing to any address granularity interval via alternative switching paths.

[0032] In some examples, optionally, the window entry in the migration window list corresponding to any address granularity interval includes: an interval identifier for the address granularity interval, a path identifier for an alternative switching path pointing to the write request and read request of the address granularity interval, and a number of retransmissions pending for the write request pointing to the address granularity interval; wherein, the initial value of the number of retransmissions pending in the window entry corresponding to any address granularity interval is used to characterize the total number of migrations of write requests pointing to the address granularity interval; the step of using the window entry to restrict the migration of read requests to be enabled for transmission after the migration of write requests is successfully retransmitted through the alternative switching path includes: in response to the successful retransmission of any write request pointing to any address granularity interval through the alternative switching path, decrementing the number of retransmissions pending in the window entry in the migration window list corresponding to the address granularity interval by one; and determining the release timing for enabling transmission of the migration of read requests based on the current value of the number of retransmissions pending in the window entry corresponding to any address granularity interval in the migration window list.

[0033] In some examples, the dynamic maintenance of the write request wait list may optionally include: adding an entry corresponding to the write request in response to the generation of any write request; deleting an entry corresponding to the write request in response to the successful transmission of any write request; and deleting an entry added to the write request wait list in response to the successful retransmission of any write request pointing to any address granularity range via an alternative switching path.

[0034] In some examples, optionally, determining the release timing of enabling transmission for the migration read requests based on the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity interval in the migration window list includes: releasing the enabling transmission of all migration read requests in batches in response to the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity interval in the migration window list becoming 0.

[0035] In some examples, optionally, determining the release timing of enabling transmission for the migration read request based on the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity interval in the migration window list includes: releasing the enabling transmission of each migration read request one by one according to the sorting relationship between the migration read requests and write requests during the change of the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity interval in the migration window list.

[0036] In another embodiment of this application, a multi-chip system is provided, including at least two artificial intelligence chips interconnected via a switching network, wherein at least one of the at least two artificial intelligence chips is the AI ​​chip in the foregoing embodiment.

[0037] In another embodiment of this application, an electronic device is provided, including the multi-chip system described in the foregoing embodiments.

[0038] Based on embodiments of this application, the bridging module of the AI ​​chip can utilize the record entry corresponding to any write request in the write request wait list to enable transmission restrictions on read requests that point to the same address granularity range via the same switching path as the write request, thereby maintaining RAW consistency between read and write requests on the switching path corresponding to that address granularity range. Furthermore, when a switching path corresponding to any address granularity range fails, the bridging module of the AI ​​chip can immediately trigger the migration of write and read requests pointing to that address granularity range to an alternative switching path. Simultaneously, even if the record entry in the write request wait list becomes invalid due to the change in the switching path of the migrated write request, the bridging module can continue to maintain RAW consistency between read and write requests migrated to the alternative switching path by adding a window entry corresponding to that address granularity range to the migration window list.

[0039] In other words, the embodiments of this application utilize a migration window list to track rewrite requests pointing to a specific address granularity range at a dual granularity of the failed switching path and the address granularity range. This accurately identifies address granularity ranges at risk of migration, rather than migrating or globally pausing all address granularity ranges. As a result, RAW consistency during migration is maintained without affecting the normal routing of unmigrated address granularity ranges. Attached Figure Description

[0040] The following figures are for illustrative purposes only and do not limit the scope of this application:

[0041] Figure 1 This is an exemplary structural diagram of a multi-chip system in an embodiment of this application;

[0042] Figure 2 This is a schematic diagram illustrating the principle of the RAW consistency maintenance mechanism at the write request granularity in a multi-chip system according to an embodiment of this application.

[0043] Figure 3 This is a schematic diagram illustrating the principle of the RAW consistency maintenance mechanism at the address granularity level during fault migration in a multi-chip system according to an embodiment of this application.

[0044] Figure 4 This is a first example schematic diagram of a RAW consistency maintenance mechanism at the address granularity interval during fault migration in a multi-chip system according to an embodiment of this application.

[0045] Figure 5 This is a second example schematic diagram of a RAW consistency maintenance mechanism with address granularity intervals during fault migration in a multi-chip system according to an embodiment of this application.

[0046] Figure 6 This is an exemplary flowchart illustrating the operation method of the AI ​​chip in the embodiments of this application. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided with reference to the accompanying drawings and embodiments.

[0048] Figure 1 This is a schematic diagram illustrating an exemplary structure of a multi-chip system in an embodiment of this application. Please refer to [link / reference]. Figure 1 In embodiments of this application, the multi-chip system may include at least two AI chips interconnected via a switching network, and, Figure 1 Taking at least two AI chips, including a first AI chip U1 and a second AI chip U2, as an example.

[0049] For example, in the embodiments of this application, each AI chip in the multi-chip system (i.e., each of the first AI chip U1 and the second AI chip U2) can be any of the integrated circuit chips suitable for AI, such as GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural Network Processing Unit), DPU (Deep Learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose computing on Graphics Processing Units).

[0050] For example, in embodiments of this application, each AI chip in a multi-chip system (i.e., each of the first AI chip U1 and the second AI chip U2) may include a processing module, a memory controller, on-chip memory, and a bridging module.

[0051] For example, in the embodiments of this application, the memory controller and bridge module of each AI chip can be interconnected with the processing module and memory controller of the AI ​​chip through the on-chip bus of the AI ​​chip, and the bridge module of each AI chip can also be interconnected with at least one other AI chip through an off-chip switching network. For example, the bridge module of the first AI chip U1 can also be interconnected with the bridge module of the second AI chip U2 through an off-chip switching network.

[0052] For example, in the embodiments of this application, the on-chip bus of each AI chip in the multi-chip system (i.e., each of the first AI chip U1 and the second AI chip U2) can be an AMBA (Advanced Microcontroller Bus Architecture) protocol bus. Specifically, the on-chip bus of each AI chip in the multi-chip system can be any one of the following AMBA protocol buses: AXI (Advanced eXtensible Interface) protocol, AHB (Advanced High-performance Bus) protocol, and APB (Advanced Peripheral Bus) protocol.

[0053] For example, in the embodiments of this application, the processing module of each AI chip in the multi-chip system (i.e., each of the first AI chip U1 and the second AI chip U2) can initiate read and write operations on the on-chip memory of its own AI chip to the memory controller of its own AI chip via the on-chip bus (i.e., on-chip read and write operations), and initiate read and write operations on the on-chip memory of other AI chips to the bridging module of its own AI chip via the on-chip bus (i.e., cross-chip read and write operations). For example, the processing module of each AI chip in the multi-chip system may include a DMA (Direct Memory Access) module or an AI core not shown in the figure. The embodiments of this application do not limit the processing module.

[0054] Exemplarily, in the embodiments of this application, the focus is mainly on the case where the processing module of any AI chip (e.g., the first AI chip U1) in a multi-chip system initiates cross-chip read / write operations on other AI chips (e.g., the second AI chip U2). The bridging module of any AI chip whose processing module initiates the cross-chip read / write operation can generate read / write requests that are forwarded to other AI chips via a switching network. Furthermore, the cross-chip read / write operations implemented based on the forwarding of read / write requests via the switching network have RAW operation timing requirements between write and read operations at the same memory address within the same AI chip. For ease of description, the first AI chip U1 represents any AI chip in the multi-chip system that acts as the source AI chip for the read / write request, and the second AI chip U2 represents any other AI chip in the multi-chip system that acts as the destination AI chip for the read / write request. Additionally, it is understood that any AI chip in the multi-chip system may act only as a source AI chip, only as a destination AI chip, or both. In the case where one AI chip in a multi-chip system serves as both a source AI chip and a destination AI chip, the descriptions of the first AI chip U1 and the second AI chip U2 in the following text can be applied to the same AI chip.

[0055] For example, in an embodiment of this application, if the on-chip bus of each AI chip in a multi-chip system is an AMBA protocol bus, then the bridging module of each AI chip for implementing cross-chip read / write operations can interconnect and communicate with the processing module and memory controller of the AI ​​chip through the bus logic channels of the AMBA protocol bus. The bus logic channels of the AMBA protocol bus include an AW channel for transmitting write addresses, a W channel for transmitting write data, an AR channel for transmitting read addresses, an R channel for transmitting read data, and a B channel for transmitting write responses.

[0056] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can: generate a write request carrying the write address and write data in response to a write address received from the AW channel and write data received from the W channel; generate a read request carrying the read address in response to a read address received from the AR channel; transmit the read data in the read response to the read response received from the switching network to the R channel in response to the read response received from the switching network; and transmit a write response to the B channel in response to the write address received from the AW channel and the write data received from the W channel.

[0057] For example, in an embodiment of this application, the bridging module of the destination AI chip (e.g., the second AI chip U2) can: transmit the write address in the write request to the AW channel and the write data in the write request to the W channel in response to a write request received from the switching network; transmit the read address in the read request to the AR channel in response to a read request received from the switching network; generate a read response carrying the read data in response to the read data received from the R channel; and receive a write response from the B channel.

[0058] For example, in embodiments of this application, the switching network can provide multiple switching paths between any pair of source AI chips (e.g., the first AI chip U1) and destination AI chips (e.g., the second AI chip U2), and each switching path can be an Ethernet link. Accordingly, write requests, read requests, and read responses can all be Ethernet messages.

[0059] For example, in embodiments of this application, for scenarios involving complex topologies in multi-chip systems, the switching network may include at least two cascaded switches, such that at least one of the multiple switching paths between any pair of source AI chips (e.g., the first AI chip U1) and destination AI chips (e.g., the second AI chip U2) is a multi-level switching path via multi-level switching forwarding through at least two levels of switches.

[0060] For example, in the embodiments of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to interact with the processing module of the AI ​​chip via the on-chip bus, and to evenly distribute the write requests and read requests generated based on the on-chip interaction to multiple switching paths in the off-chip switching network for transmission.

[0061] For example, in the embodiments of this application, the balancing allocation mechanism based on the bridging module of the source AI chip (e.g., the first AI chip U1) may include the ECMP (Equal-Cost MultiPath routing) mechanism. The ECMP mechanism can be implemented using a hash algorithm. Furthermore, the balancing allocation mechanism based on the bridging module of the source AI chip (e.g., the first AI chip U1) may also include an affinity mechanism based on address granularity intervals, that is: write requests and read requests pointing to the same address granularity interval (i.e., pointing to the same address granularity interval of the on-chip memory of another AI chip, such as the second AI chip U2) are allocated on the same specified switching path.

[0062] For example, in an embodiment of this application, the address granularity range can be a memory page of on-chip memory, for example, the size of a memory page is 4KB.

[0063] For example, in an embodiment of this application, if the on-chip bus of each AI chip in a multi-chip system is an AMBA protocol bus, then the processing module of the source AI chip (e.g., the first AI chip U1) can initiate write requests for cross-chip write operations and read requests for cross-chip read operations in the Outstanding mode of the AMBA protocol. This ensures that, for the same address granularity range (e.g., an address granularity range of 4KB page size) in the same destination AI chip (e.g., the second AI chip U2), a read request corresponding to a read operation with RAW operation timing requirements for the write operation can be initiated before the write operation is completed. That is, the write and read requests allocated by the bridging module of the source AI chip (e.g., the first AI chip U1) to each switching path can be initiated in the Outstanding mode.

[0064] For example, in the embodiments of this application, each switching path of the switching network can implement flow control-based switching forwarding, such as flow control-based switching forwarding based on PDC (Packet Delivery Context). Multiple switching paths between any pair of source AI chips (e.g., the first AI chip U1) and destination AI chips (e.g., the second AI chip U2) in the switching network can implement CBFC (Credit-Based Flow Control) at the granularity of VC (Virtual Channel).

[0065] For example, in an embodiment of this application, if write requests and read requests on the same switching path are transmitted through the same VC, the bridging module of the source AI chip (e.g., the first AI chip U1) can introduce a Hazard detection mechanism for the FIFO (First In First Out) transmission queue maintained by each VC to ensure RAW consistency between read and write requests of the same VC.

[0066] For example, in embodiments of this application, write requests and read requests on the same switching path may also be allowed to be transmitted separately through different VCs to achieve a more granular CBFC. In this case, since different VCs may produce different delays for write requests and read requests on the same switching path in the switching network, this may cause a read request pointing to a certain address granularity range to be switched and forwarded earlier than a write request pointing to the same address granularity range and sent first (e.g., the predetermined delay interval for switching and forwarding is shortened compared to the first-sent write request, or even earlier than the first-sent write request). This results in inconsistencies in the RAW of read operations on the same memory address in the same AI chip before the write operation is completed.

[0067] For example, in an embodiment of this application, in order to ensure RAW consistency of read and write requests when write and read requests on the same exchange path are transmitted through different VCs, the bridging module of the source AI chip (e.g., the first AI chip U1) can also be configured to dynamically maintain the WPT (Write Pending Table) to implement a RAW consistency maintenance mechanism at the granularity of write requests.

[0068] Figure 2 This is a schematic diagram illustrating the principle of the RAW consistency maintenance mechanism at the write request granularity in a multi-chip system according to an embodiment of this application. Please refer to... Figure 2 In embodiments of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to perform dynamic maintenance operations on the WPT. These dynamic maintenance operations may include: adding a record entry corresponding to an arbitrary write request in the WPT in response to the generation of such a write request, and deleting the record entry corresponding to the write request in the WPT in response to the successful transmission of such a write request. Furthermore, the record entry in the WPT corresponding to an arbitrary write request can be used to restrict read requests generated after the write request, and which, along the same specified exchange path, point to the same address granularity range as the write request, to be enabled for transmission only after the successful transmission of the write request.

[0069] For example, in an embodiment of this application, the record entry in WPT corresponding to any write request may include: the interval identifier of the address granularity interval pointed to by the write request (e.g., the page identifier of a 4KB memory page), and the path identifier of the specified swap path allocated when the write request is generated.

[0070] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to restrict the transmission of read requests using records in the WPT in the following manner: in response to a read request pointing to any address granularity range, the WPT is queried for the record corresponding to the write request pointing to that address granularity range. For example, the WPT is queried for a record containing the range identifier of the address granularity range where the read address of the read request is located and the path identifier of the exchange path where the read request is located. If the query is successful, the read request is temporarily stored and the query for the read request is repeated in the WPT until the record that was successful in the query for the read request is deleted, causing the query to fail.

[0071] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to delete the record item in WPT corresponding to the write request in response to the successful transmission of any write request.

[0072] For example, in an embodiment of this application, if the switching forwarding implemented by the switching network is a switching forwarding based on PDC to implement flow control, then the bridging module of the destination AI chip (e.g., the second AI chip U2) can also be configured to: in response to receiving a write request from the source AI chip (e.g., the first AI chip U1) from the switching network, transmit a PDS (Packet Delivery Sublayer) ACK (successful acknowledgment) carrying the PSN (Packet Sequence Number) corresponding to the write request in the PDC to the switching network, so that the PDS ACK is switched and forwarded by the switching network to the source AI chip (e.g., the first AI chip U1) that generated the write request. Accordingly, the bridging module of the source AI chip (e.g., the first AI chip U1) can also be configured to: in response to receiving a PDS ACK carrying the PSN corresponding to any write request in the PDC through the switching network, confirm that the write request has reached the destination AI chip (i.e., successfully transmitted), so as to delete the record entry in WPT corresponding to the write request.

[0073] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) may also be configured to: detect the path status of multiple switching paths during the period of dynamically maintaining WPT; and, in response to detecting that the path status of any of the multiple switching paths is in a fault state, trigger the migration of write requests and read requests to be transmitted that are directed to the same address granularity range with that switching path as the specified switching path to an alternative switching path selected in the switching network.

[0074] For example, in an embodiment of this application, the selection of an alternative exchange path can be achieved by rehashing the currently remaining available exchange paths among a plurality of exchange paths.

[0075] For example, in an embodiment of this application, if the switching forwarding implemented by the switching network is a switching forwarding with flow control based on PDC, then the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to: confirm that the path state of the specified switching path of the write request is in a fault state in response to a timeout in receiving a PDS ACK for a PSN carrying an arbitrary write request or in response to receiving a PDS NACK (failure acknowledgment) for a PSN carrying an arbitrary write request.

[0076] For example, in embodiments of this application, the migrated write request to be transmitted may include: a write request that has been generated but not yet transmitted through the failed specified switching path, and a write request that has been transmitted through the failed specified switching path but has not received a PDS ACK carrying the PSN corresponding to any write request in the PDC. Moreover, the migrated write request will be retransmitted through an alternative switching path with a different path identifier than the specified switching path before migration.

[0077] For example, in an embodiment of this application, the read request to be migrated may include a read request that is temporarily held due to a query hit in the WPT. Since the migrated read request will be transmitted through an alternative exchange path with a different path identifier than the specified exchange path before migration, the query result of the migrated read request using the dual index of the region identifier and the path identifier in the WPT will fail, causing the record entries in the WPT that could have temporarily held the migrated read request to become invalid, which may lead to RAW inconsistency during migration due to exchange path failure.

[0078] For example, in an embodiment of this application, in order to avoid RAW inconsistency during migration caused by a failure of the switching path, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to dynamically maintain the MWT (Migration Window Table) to implement a RAW consistency maintenance mechanism at the address granularity level during fault migration.

[0079] Figure 3 This is a schematic diagram illustrating the principle of the RAW consistency maintenance mechanism at the address granularity level during fault migration in a multi-chip system according to an embodiment of this application. Please refer to... Figure 3 In the embodiments of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to: in response to detecting that the path status of any of the multiple switching paths is in a fault state, that is, while triggering the migration of write and read requests to be transmitted that are directed to the same address granularity range via the specified switching path to an alternative switching path selected in the switching network, add a window entry corresponding to the address granularity range to which the migrated write and read requests are directed in the MWT according to the WPT. The window entry corresponding to the address granularity range is deleted only after all write requests to the address granularity range have been successfully retransmitted, so as to restrict the read requests to be migrated to the address granularity range to be enabled for transmission only after the write requests to be migrated to the address granularity range have been successfully retransmitted through the alternative switching path.

[0080] For example, in an embodiment of this application, the window entry in MWT corresponding to any address granularity interval may include: the interval identifier of the address granularity interval (e.g., the page identifier of a 4KB memory page), the path identifier of the alternative swap path pointing to the write request and read request of the address granularity interval, and the number of retransmissions to be made for the write request pointing to the address granularity interval.

[0081] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to restrict the enabled transmission of read requests using window entries in the MWT in the following manner: For a read request to be retransmitted that points to an arbitrary address granularity range, the window entry corresponding to the address granularity range is queried in the MWT. For example, the record in the MWT is queried for the range identifier of the address granularity range where the read address of the read request is located and the path identifier of the alternative switching path to which the read request is reassigned. If the query is successful, the read request is kept temporarily suspended, and the query for the read request is repeated in the MWT until the window entry that was successful in the query for the read request in the MWT is deleted, causing the query to fail.

[0082] For example, in an embodiment of this application, if the switching forwarding implemented by the switching network is a switching forwarding with flow control based on PDC, then the bridging module of the source AI chip (e.g., the first AI chip U1) can also be configured to: determine that the write request has been successfully retransmitted in response to receiving a PDS ACK carrying an arbitrary write request PSN.

[0083] For example, in an embodiment of this application, the current value of the number of retransmissions to be performed in the window entry corresponding to any address granularity interval in the MWT can decrease in response to the increase of the number of successful retransmissions of write requests pointing to that address granularity interval through the alternative switching path, until the current value of the number of retransmissions to be performed becomes 0. This can confirm that all write requests pointing to that address granularity interval have been successfully retransmitted, and cause the window entry corresponding to that address granularity interval to be deleted in the MWT.

[0084] For example, in an embodiment of this application, when all window items in the MWT are deleted, it can be assumed that the MWT does not exist at this time, and... Figure 2 The illustration showing only WPT is intended to indicate that the presence of only WPT is permissible.

[0085] Figure 4 This is a first example schematic diagram of a RAW consistency maintenance mechanism at the address granularity level during fault migration in a multi-chip system according to an embodiment of this application. Figure 5 This is a second example schematic diagram of a RAW consistency maintenance mechanism at the address granularity level during fault migration in a multi-chip system according to an embodiment of this application.

[0086] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to perform the operation of adding a window item in MWT in the following manner:

[0087] Search in WPT for all records corresponding to write requests that use any switch path currently detected as faulty as the specified switch path, for example... Figure 4 Figure 5 The record entries corresponding to the first write request W1, the second write request W2, and the third write request W3 shown in the figure can each include the interval identifier (e.g., the page identifier Page_x of the Xth 4KB memory page) of the address granularity interval pointed to by the first write request W1, the second write request W2, and the third write request W3, and the path identifier Path_a of the specified swap path allocated when the write request is generated;

[0088] Using the records found in WPT, add window entries in MWT corresponding to the address granularity ranges pointed to by the migrated write and read requests, for example... Figure 4 Figure 5The window item shown includes the interval identifier of the address granularity interval (e.g., the page identifier Page_x of the Xth 4KB memory page), the path identifier Path_b of the alternative swap path to the write and read requests of the address granularity interval, and the number of retransmissions to be made for the write requests of the address granularity interval.

[0089] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to perform the operation of querying record items in WPT by: using the path identifier of any exchange path currently detected as faulty (e.g., Figure 4 and Figure 5 Using the page identifier Page_x of the Xth 4KB memory page shown as an index, search the WPT for all records corresponding to write requests with that swap path as the specified swap path.

[0090] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to perform the operation of adding window items in the MWT using record items searched from the WPT in the following manner:

[0091] Identify the interval identifiers (e.g., in the records searched in WPT) Figure 4 and Figure 5 The page identifier Page_x of the Xth 4KB memory page is shown in the figure.

[0092] Based on the same range identifier found in WPT (e.g.) Figure 4 and Figure 5 For all records of the page identifier Page_x of the Xth 4KB memory page shown in the figure, create a window entry corresponding to an address granularity range represented by that interval identifier.

[0093] For example, in an embodiment of this application, the initial value of the number of retransmissions to be performed (Count) in a window item corresponding to any address granularity interval can be used to characterize the total number of write requests migrated to that address granularity interval, and can be determined by the number of record items on which the window item is based, for example... Figure 4 and Figure 5 The initial value of the number of retransmissions to be shown in the figure can be "3", which is the number of record items corresponding to the first write request W1, the second write request W2, and the third write request W3.

[0094] For example, in an embodiment of this application, the current value of the number of retransmissions to be performed in a window entry corresponding to any address granularity interval can decrease in response to the increase in the number of successful retransmissions via the alternative switching path for a write request pointing to any address granularity interval.

[0095] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to perform the operation of restricting migration read requests in the following manner:

[0096] In response to a successful retransmission via an alternative switching path for an arbitrary write request pointing to an arbitrary address granularity range, the number of retransmissions to be performed in the window entry corresponding to that address granularity range in the MWT is decremented by one. For example... Figure 4 and Figure 5 The current value of the number of retransmissions to be performed as shown in the figure is decremented by one in response to the successful retransmissions of the first write request W1, the second write request W2, and the third write request W3 through the alternative switching path.

[0097] Based on the current value of the number of retransmissions to be performed in the window entry corresponding to any address granularity interval in MWT, determine the timing for releasing the enable transmission of the migrated read request.

[0098] For example, in an embodiment of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can be configured to perform an operation to determine the release timing of enabled transmissions of migrated read requests by: releasing all migrated read requests in batches in response to the current value of the number of retransmissions to be performed in the window entry corresponding to any address granularity interval in the MWT becoming 0. For example Figure 4 The enable transfers for the first read request R1, the second read request R2, and the third read request R3 of the bulk release migration are shown in the figure.

[0099] For example, in embodiments of this application, the bridging module of the source AI chip (e.g., the first AI chip U1) can also be configured to perform the operation of determining the release timing of the enable transmission of the migrated read requests in the following manner: during the change of the current value of the number of retransmissions in the window item corresponding to any address granularity interval in MWT, the enable transmission of each migrated read request is released one by one according to the sorting relationship between the migrated read requests and write requests (e.g., the initiation order of the Outstanding corresponding to the read requests and write requests). For example Figure 5 The order shown is as follows: first write request W1, first read request R1, second write request W2, second read request R2, third write request W3, and third read request R3. In response to the current value of the number of retransmissions to be made in the window item corresponding to any address granularity interval in MWT, the enabled transmission of the first read request R1, second read request R2, and third read request R3 is released one by one.

[0100] Exemplary examples, in embodiments of this application, such as Figure 4 and Figure 5As shown, the dynamic maintenance of WPT by the bridging module of the source AI chip (e.g., the first AI chip U1) also includes: in response to the successful retransmission of any write request pointing to any address granularity range through the alternative switching path, deleting the record entry added to WPT corresponding to the write request.

[0101] Based on embodiments of this application, the bridging module of the AI ​​chip can utilize the record entries corresponding to any write request in the WPT to enable transmission restrictions on read requests that point to the same address granularity range via the same switching path as the write request, thereby maintaining RAW consistency between read and write requests on the switching path corresponding to that address granularity range. Furthermore, when a failure occurs on the switching path corresponding to any address granularity range, the bridging module of the AI ​​chip can immediately trigger the migration of write and read requests pointing to that address granularity range to an alternative switching path. Simultaneously, even if the record entries in the WPT become invalid due to the change in the switching path of the migrated write requests, the bridging module can continue to maintain RAW consistency between read and write requests migrated to the alternative switching path (regardless of whether the write and read requests are transmitted in the same VC on the alternative switching path or separately in different VCs on the alternative switching path) by adding a window entry corresponding to that address granularity range in the MWT.

[0102] In other words, the embodiments of this application utilize WPT to trace rewrite requests pointing to the address granularity range at a dual granularity of the failed switching path and the address granularity range, accurately identifying address granularity ranges with migration risks, rather than migrating or globally pausing all address granularity ranges, thereby maintaining RAW consistency during migration without affecting the normal routing of unmigrated address granularity ranges.

[0103] Other embodiments of this application also provide an operation method for an AI chip, which can be applied to the case where the AI ​​chip is the source AI chip described above.

[0104] Figure 6 This is an exemplary flowchart illustrating the operation method of the AI ​​chip in this application embodiment. Please refer to... Figure 6 Another embodiment of this application provides an operating method for an AI chip that may include the following steps performed by the bridging module of the AI ​​chip:

[0105] S610: During the dynamic maintenance of WPT, detect the path status of multiple switching paths;

[0106] S630: In response to detecting that the path status of any of the multiple switching paths is in a fault state, triggering the migration of write and read requests to be transmitted that are directed to the same address granularity range via that switching path to an alternative switching path selected in the switching network, and adding a window entry corresponding to the address granularity range to be migrated in the MWT according to WPT, the window entry corresponding to the address granularity range is not deleted until all write requests to that address granularity range are successfully retransmitted, so as to restrict the read requests to be migrated to that address granularity range to be enabled for transmission only after the write requests to be migrated to that address granularity range are successfully retransmitted through the alternative switching path.

[0107] For example, in an embodiment of this application, if the switching forwarding implemented by the switching network is a switching forwarding with flow control based on PDC, then S610 may include: in response to a timeout in receiving a PDS ACK for a PSN carrying an arbitrary write request or in response to receiving a PDS NACK (failure acknowledgment) for a PSN carrying an arbitrary write request, confirming that the path state of the specified switching path of the write request is in a fault state.

[0108] For example, in an embodiment of this application, if the switching forwarding implemented by the switching network is a switching forwarding with flow control based on PDC, then S630 can determine that the write request has been successfully retransmitted in response to receiving a PDS ACK carrying an arbitrary write request from a PSN.

[0109] For example, in an embodiment of this application, the operation of S630 adding a window entry in the MWT corresponding to the address granularity range pointed to by the migrated write request and read request based on the WPT may include: searching in the WPT for all records corresponding to write requests with any exchange path currently detected as faulty as the specified exchange path; and using the records searched in the WPT, adding a window entry in the MWT corresponding to the address granularity range pointed to by the migrated write request and read request. The dynamic maintenance of the WPT includes adding a record entry corresponding to the write request in the WPT in response to the generation of any write request, and deleting a record entry corresponding to the write request in the WPT in response to the successful transmission of any write request. Furthermore, the record entry in the WPT corresponding to any write request is used to restrict read requests generated after the write request and pointing to the same address granularity range via the same specified exchange path as the write request to be enabled for transmission only after the successful transmission of the write request.

[0110] For example, in an embodiment of this application, the operation of S630 in searching the WPT for all record entries corresponding to a write request with any exchange path currently detected as faulty as the specified exchange path may include: using the path identifier of any exchange path currently detected as faulty as an index, searching the WPT for all record entries corresponding to a write request with that exchange path as the specified exchange path.

[0111] For example, in an embodiment of this application, the operation of S630 in adding a window entry in the MWT corresponding to the address granularity interval pointed to by the write request and read request of the migration using the record entries searched in the WPT may include: identifying the interval identifier in the record entries searched in the WPT, and creating a window entry corresponding to an address granularity interval represented by the interval identifier based on all record entries in the WPT that include the same interval identifier.

[0112] For example, in an embodiment of this application, the operation of restricting read requests for migrations to the address granularity range to be enabled for transmission after write requests for migrations to the address granularity range are successfully retransmitted via an alternative switching path may include: for a read request to be retransmitted that points to any address granularity range, querying the window entry corresponding to the address granularity range in the MWT, for example, querying the record in the MWT that has the range identifier of the address granularity range where the read address of the read request is located and the path identifier of the alternative switching path to which the read request has been reassigned; if the query is successful, the read request is kept temporarily suspended, and the query for the read request is repeated in the MWT until the window entry that was successful in the query for the read request in the MWT is deleted, causing the query to fail.

[0113] For example, in an embodiment of this application, the operation of restricting the read request for migration to the address granularity range to be enabled for transmission after the write request for migration to the address granularity range is successfully retransmitted through the alternative switching path may include: in response to the successful retransmission of any write request to any address granularity range through the alternative switching path, decrementing the number of retransmissions to be performed in the window entry corresponding to the address granularity range in the MWT; and determining the release timing for enabling transmission of the read request for migration based on the current value of the number of retransmissions to be performed in the window entry corresponding to any address granularity range in the MWT.

[0114] For example, in an embodiment of this application, the operation of S630 determining the release timing of the enable transmission of the migrated read request based on the current value of the number of retransmissions to be performed in the window entry corresponding to any address granularity interval in the MWT may include: releasing all migrated read requests in batches in response to the current value of the number of retransmissions to be performed in the window entry corresponding to any address granularity interval in the MWT becoming 0, or, during the change of the current value of the number of retransmissions to be performed in the window entry corresponding to any address granularity interval in the MWT, releasing the enable transmission of each migrated read request one by one according to the sorting relationship between the migrated read requests and write requests.

[0115] For example, in an embodiment of this application, the operation method of dynamic maintenance of WPT by the bridging module may further include: in response to the successful retransmission of any write request pointing to any address granularity range through an alternative switching path, deleting the record entry in WPT corresponding to the write request.

[0116] Based on the above-described operation method of this application embodiment, the bridging module of the AI ​​chip can utilize the record entry corresponding to any write request in the WPT to enable transmission restrictions on read requests that point to the same address granularity range via the same switching path as the write request, thereby maintaining RAW consistency between read and write requests on the switching path corresponding to that address granularity range. Furthermore, when a failure occurs on the switching path corresponding to any address granularity range, the bridging module of the AI ​​chip can immediately trigger the migration of write and read requests pointing to that address granularity range to an alternative switching path. Simultaneously, even if the record entry in the WPT becomes invalid due to the change in the switching path of the migrated write request, the bridging module can continue to maintain RAW consistency between read and write requests migrated to the alternative switching path by adding a window entry corresponding to that address granularity range in the MWT.

[0117] Another embodiment of this application also provides an electronic device that may include the AI ​​chip (e.g., source AI chip) or multi-chip system described in the foregoing embodiments.

[0118] It is understood that, in the embodiments of this application, the various parts described by example may be related by an "and / or" relationship. In this document, "and / or" means that the contexts connected by it may be a common "and" relationship or an alternative "or" relationship. Therefore, the various parts having an "and / or" relationship can be understood to include different combinations of situations where the "and / or" between each pair of parts represents a common "and" relationship or an alternative "or" relationship, and such combinations of different situations can be considered substantially equivalent to the scope of "at least one of the parts".

[0119] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and mechanism of this application should be included within the scope of protection of this application.

Claims

1. An artificial intelligence chip, characterized in that, include: Processing module; The bridging module is configured to interact with the processing module via an on-chip bus, and to evenly distribute write requests and read requests generated based on the on-chip interaction to multiple switching paths in an off-chip switching network for transmission. The even distribution mechanism includes: write requests and read requests pointing to the same address granularity range are assigned to the same specified switching path. The bridging module is further configured as follows: During the period of dynamically maintaining the write request wait list, the path status of multiple swap paths is checked; In response to the detection that any switching path is in a fault state, a mechanism is triggered to migrate pending write and read requests that are directed to the same address granularity range via that switching path to an alternative switching path selected in the switching network. Furthermore, based on the write request wait list, a window entry corresponding to the address granularity range to which the migrated write and read requests are directed is added to the migration window list. This window entry is deleted only after all write requests to that address granularity range have been successfully retransmitted, thereby restricting the migration of read requests to that address granularity range to be enabled for transmission only after the migration of write requests to that address granularity range has been successfully retransmitted via the alternative switching path.

2. The artificial intelligence chip according to claim 1, characterized in that, The bridging module is specifically configured to perform the operation of adding window items to the migration window list in the following manner: Search the write request wait list for all records corresponding to write requests that use any switch path currently detected as faulty as the specified switch path. Using the record entries found in the write request wait list, add window entries to the migration window list that correspond to the address granularity range pointed to by the migrated write and read requests; The record entry corresponding to any write request in the write request wait list is used to restrict read requests that are generated after the write request and that point to the same address granularity range through the same specified exchange path as the write request to be enabled for transmission after the write request is successfully transmitted.

3. The artificial intelligence chip according to claim 2, characterized in that, The record entries in the write request wait list corresponding to any write request include: the path identifier of the specified exchange path assigned to the write request when it is generated; The bridging module is specifically configured to perform the operation of querying record items in the write request wait list in the following manner: Using the path identifier of any exchange path currently detected as faulty as an index, search the write request wait list for all records corresponding to write requests that specify that exchange path as the exchange path.

4. The artificial intelligence chip according to claim 2, characterized in that, The record item corresponding to any write request in the write request wait list includes: the interval identifier of the address granularity interval pointed to by the write request; The bridging module is specifically configured to perform the operation of adding window items to the migration window list using the searched record items in the following manner: Identify the interval identifier in the record item searched in the write request wait list; Based on all records containing the same interval identifier found in the write request wait list, a window entry corresponding to an address granularity interval represented by that interval identifier is created.

5. The artificial intelligence chip according to claim 4, characterized in that, The window item in the migration window list corresponding to any address granularity interval includes: the interval identifier of the address granularity interval, the path identifier of the alternative switching path pointing to the write request and read request of the address granularity interval, and the number of retransmissions to be made for the write request pointing to the address granularity interval. The initial value of the number of retransmissions to be made in the window entry corresponding to any address granularity interval is determined by the number of record entries on which the window entry is based; and the current value of the number of retransmissions to be made in the window entry corresponding to any address granularity interval decreases in response to the increase in the number of successful retransmissions of the write request to the arbitrary address granularity interval through the alternative switching path.

6. The artificial intelligence chip according to claim 1, characterized in that, The window entries in the migration window list corresponding to any address granularity interval include: the interval identifier of the address granularity interval, the path identifier of the alternative switching path for the write request and read request pointing to the address granularity interval, and the number of retransmissions to be made for the write request pointing to the address granularity interval; wherein, the initial value of the number of retransmissions to be made in the window entries corresponding to any address granularity interval is used to characterize the total number of migrations for the write request pointing to the address granularity interval. The bridging module is specifically configured to perform operations that restrict migration of read requests in the following manner: In response to a successful retransmission of an arbitrary write request pointing to an arbitrary address granularity range via an alternative switching path, the number of retransmissions to be performed in the window item corresponding to the address granularity range in the migration window list is decremented by one. Based on the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity range in the migration window list, the timing for releasing the enable transmission of the migration read request is determined.

7. The artificial intelligence chip according to claim 6, characterized in that, Dynamic maintenance of the write request wait list includes: In response to any write request, add a record entry corresponding to that write request; In response to the successful transmission of any write request, delete the record corresponding to that write request; and, In response to a successful retransmission of an arbitrary write request pointing to an arbitrary address granularity range via an alternative switching path, the record entry corresponding to the write request is removed from the write request wait list.

8. The artificial intelligence chip according to claim 6, characterized in that, The bridging module is specifically configured to perform the operation of releasing the enable transfer timing for determining the read request of the migration in the following manner: In response to the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity range in the migration window list becoming 0, the enabled transmission of all read requests for the migration is released in batches.

9. The artificial intelligence chip according to claim 6, characterized in that, The bridging module is specifically configured to perform the operation of releasing the enable transfer timing for determining the read request of the migration in the following manner: During the change of the current value of the number of retransmissions in the window item corresponding to any address granularity range in the migration window list, the enabled transmission of each read request in the migration is released one by one according to the sorting relationship between the read requests and write requests in the migration.

10. The artificial intelligence chip according to any one of claims 1 to 9, characterized in that, The address granularity range is a memory page, and the size of the memory page is 4KB; And / or, The arbitrary switching path in the switching network is used to implement switching forwarding based on PDC flow control. The bridging module is specifically configured to: determine whether the write request has been successfully transmitted or successfully retransmitted in response to receiving a PDS ACK carrying an arbitrary write request PSN; and confirm that the path status of the specified switching path of the write request is in a fault state in response to a timeout in receiving a PDS ACK carrying an arbitrary write request PSN or in response to receiving a PDS NACK carrying an arbitrary write request PSN.

11. A method for operating an artificial intelligence chip, characterized in that, The artificial intelligence chip includes a processing module and a bridging module. The bridging module interacts with the processing module via an on-chip bus. The bridging module evenly distributes write requests and read requests generated based on the on-chip interaction to multiple switching paths in an off-chip switching network. The even distribution mechanism includes: write requests and read requests pointing to the same address granularity range are assigned to the same specified switching path. The operation method includes the following steps performed by the bridging module: During the period of dynamically maintaining the write request wait list, the path status of multiple swap paths is checked; In response to the detection that any switching path is in a fault state, a mechanism is triggered to migrate pending write and read requests that are directed to the same address granularity range via that switching path to an alternative switching path selected in the switching network. Furthermore, based on the write request wait list, a window entry corresponding to the address granularity range to which the migrated write and read requests are directed is added to the migration window list. This window entry is deleted only after all write requests to that address granularity range have been successfully retransmitted, thereby restricting the migration of read requests to that address granularity range to be enabled for transmission only after the migration of write requests to that address granularity range has been successfully retransmitted via the alternative switching path.

12. The operating method according to claim 11, characterized in that, The step of adding window items to the migration window list according to the write request wait list, corresponding to the address granularity range pointed to by the migrated write and read requests, includes: Search the write request wait list for all records corresponding to write requests that use any switch path currently detected as faulty as the specified switch path. Using the record entries found in the write request wait list, add window entries to the migration window list that correspond to the address granularity range pointed to by the migrated write and read requests; The record entry corresponding to any write request in the write request wait list is used to restrict read requests that are generated after the write request and that point to the same address granularity range through the same specified exchange path as the write request to be enabled for transmission after the write request is successfully transmitted.

13. The operating method according to claim 12, characterized in that, The record entries in the write request wait list corresponding to any write request include: the path identifier of the specified exchange path assigned to the write request when it is generated; The search in the write request wait list includes all records corresponding to write requests that use any switch path currently detected as faulty as the specified switch path, including: Using the path identifier of any exchange path currently detected as faulty as an index, search the write request wait list for all records corresponding to write requests that specify that exchange path as the exchange path.

14. The operating method according to claim 12, characterized in that, The record item corresponding to any write request in the write request wait list includes: the interval identifier of the address granularity interval pointed to by the write request; The step of adding window entries corresponding to the address granularity ranges pointed to by the migrated write and read requests to the migration window list using the record entries searched in the write request wait list includes: Identify the interval identifier in the record item searched in the write request wait list; Based on all records containing the same interval identifier found in the write request wait list, a window entry corresponding to an address granularity interval represented by that interval identifier is created.

15. The operating method according to claim 11, characterized in that, The window entries in the migration window list corresponding to any address granularity interval include: the interval identifier of the address granularity interval, the path identifier of the alternative switching path for the write request and read request pointing to the address granularity interval, and the number of retransmissions to be made for the write request pointing to the address granularity interval; wherein, the initial value of the number of retransmissions to be made in the window entries corresponding to any address granularity interval is used to characterize the total number of migrations for the write request pointing to the address granularity interval. The step of using this window item to restrict the migration of read requests to transmission only after the migration of write requests has been successfully retransmitted via the alternative switching path includes: In response to a successful retransmission of an arbitrary write request pointing to an arbitrary address granularity range via an alternative switching path, the number of retransmissions to be performed in the window item corresponding to the address granularity range in the migration window list is decremented by one. Based on the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity range in the migration window list, the timing for releasing the enable transmission of the migration read request is determined.

16. The operating method according to claim 15, characterized in that, Dynamic maintenance of the write request wait list includes: In response to any write request, add a record entry corresponding to that write request; In response to the successful transmission of any write request, delete the record corresponding to that write request; and, In response to a successful retransmission of an arbitrary write request pointing to an arbitrary address granularity range via an alternative switching path, the record entry corresponding to the write request is removed from the write request wait list.

17. The operating method according to claim 15, characterized in that, The step of determining the release timing of enabling transmission for the read request of the migration based on the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity interval in the migration window list includes: In response to the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity range in the migration window list becoming 0, the enabled transmission of all read requests for the migration is released in batches.

18. The operating method according to claim 15, characterized in that, The step of determining the release timing of enabling transmission for the read request of the migration based on the current value of the number of retransmissions to be performed in the window item corresponding to any address granularity interval in the migration window list includes: During the change of the current value of the number of retransmissions in the window item corresponding to any address granularity range in the migration window list, the enabled transmission of each read request in the migration is released one by one according to the sorting relationship between the read requests and write requests in the migration.

19. A multi-chip system, characterized in that, It includes at least two artificial intelligence chips interconnected via a switching network, and at least one of the at least two artificial intelligence chips is an artificial intelligence chip as claimed in any one of claims 1 to 10.

20. An electronic device, characterized in that, Including the multi-chip system as described in claim 19.