Artificial intelligence device and deadlock avoidance method in aggregate communication
By setting up a request scheduler, response scheduler, request buffer and response buffer on an artificial intelligence device, physical isolation between request and response path is achieved, deadlock problem in multi-device collective communication is solved, and the stability and data transmission capabilities of the system are improved.
Patent Information
- Application Number
- CN202510725806.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-29
AI Technical Summary
The deadlock problem in multi-device collective communication leads to system crashes or data loss, which is difficult to effectively solve in the existing technology.
Set up a request scheduler, response scheduler, request buffer and response buffer on an artificial intelligence device to achieve physical isolation between the request and the response path. The request scheduler processes data requests and generates processing responses. The response scheduler is responsible for returning the response to ensure that the processing flow of the request and the response does not interfere with each other.
It effectively avoids deadlocks in multi-device collective communication, improves the reliability and efficiency of the system, reduces the risk of system stagnation caused by resource competition, and is suitable for parallel processing needs in many-to-many communication scenarios.
Smart Images

Figure CN120560871A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence chip technology, and in particular to an artificial intelligence device and a deadlock avoidance method in collective communication. Background Art
[0002] With the rapid development of artificial intelligence and high-performance computing, AI device clusters, leveraging their powerful parallel computing capabilities, have become a mainstream platform for processing large amounts of data and complex algorithms. Multi-device collective communication, as the core mechanism supporting the efficient operation of these applications, is responsible for enabling rapid data exchange and synchronization between multiple devices, thereby ensuring the smooth progress of computing tasks. However, with the increase in the number of devices and the increasing complexity of communication, deadlock issues in existing multi-device collective communication scenarios have become increasingly prominent, becoming a key bottleneck restricting system performance and stability.
[0003] Specifically, during multi-device collective communication, when multiple devices simultaneously send data or requests to a specific device, the device's receive buffer often becomes the focus of resource competition. Once the receive buffer is occupied or not released in time, the data or requests sent by other devices will not be able to arrive smoothly, causing these devices to fall into a blocked state. More seriously, because the device can no longer receive new requests, the response it should have returned cannot be sent. This two-way blocking state will eventually cause the entire communication system to deadlock. Once a deadlock occurs, it will not only cause the computing task to stagnate, but may also cause serious consequences such as system crashes or data loss. Therefore, there is an urgent need for a method that can solve or avoid the deadlock problem in multi-device collective communication. Summary of the Invention
[0004] The present invention provides an artificial intelligence device and a deadlock avoidance method in collective communication, which is used to solve the defect in the related art that deadlock is prone to occur in a multi-device collective communication scenario, resulting in system crash or data loss.
[0005] The present invention provides an artificial intelligence device, comprising a request scheduler, a response scheduler, a request buffer, and a response buffer; The request buffer is configured to store a received first data request, where the first data request is sent by a first device; The request scheduler is configured to obtain the first data request from the request buffer, process the first data request, and generate a first processing response; The response buffer is used to store the generated first processing response; The response scheduler is configured to obtain the first processing response from the response buffer and return the first processing response to the first device.
[0006] According to an artificial intelligence device provided by the present invention, the response scheduler is further used to: The first data request in the request buffer is released.
[0007] According to an artificial intelligence device provided by the present invention, the request scheduler is further configured to send a second data request to the second device; The response buffer is further used to store a second processing response returned by the second device, where the second processing response is generated after the second device processes the second data request.
[0008] An artificial intelligence device according to the present invention further includes a request queue; The request queue is used to store the second data request generated locally; The request scheduler is further configured to, upon detecting that corresponding resources exist in the second device, obtain the second data request from the request queue and send the second data request to the second device.
[0009] According to an artificial intelligence device provided by the present invention, the request scheduler is further configured to: Obtaining the second processing response from the response buffer, and parsing the second processing response to obtain a target request identifier; The data request matching the target request identifier in the request queue is released, and the second processing response in the response buffer is released.
[0010] According to an artificial intelligence device provided by the present invention, the first processing response is returned to the first device with write semantics, and the second processing response is returned to the artificial intelligence device with write semantics.
[0011] According to an artificial intelligence device provided by the present invention, the first data request is any one of a data read request and a data write request, and the second data request is any one of a data read request and a data write request.
[0012] According to an artificial intelligence device provided by the present invention, the request buffer and the response buffer are both first-in-first-out queues.
[0013] The present invention also provides a deadlock avoidance method in collective communication, the method being applied to an artificial intelligence device, the artificial intelligence device including a request scheduler, a response scheduler, a request buffer, and a response buffer, the method comprising: receiving a first data request sent by a first device, and storing the first data request in the request buffer; Based on the request scheduler, obtain the first data request from the request buffer, and process the first data request to generate a first processing response; The first processing response is stored in the response buffer, and based on the response scheduler, the first processing response in the response buffer is returned to the first device.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the deadlock avoidance method in collective communication as described above is implemented.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the deadlock avoidance method in collective communication as described above is implemented.
[0016] The deadlock avoidance method in the artificial intelligence device and collective communication provided by the present invention realizes the physical isolation of the request and response paths by simultaneously setting a request scheduler, a response scheduler, a request buffer and a response buffer on the artificial intelligence device, thereby effectively solving the deadlock problem in the collective communication of multiple devices. The request scheduler is specifically responsible for extracting and processing the first data request sent by the first device from the request buffer, and the generated first processing response is stored in the response buffer, and then the response scheduler is responsible for returning it to the first device. This separation design ensures that the processing flows of the request and response do not interfere with each other, and avoids resource competition in two-way communication. Even in high concurrency or resource competition scenarios, the blocking of the request channel will not affect the normal operation of the response channel, and vice versa. By decoupling the storage and scheduling logic of requests and responses, the present invention not only avoids deadlocks caused by circular waiting, but also improves the reliability and efficiency of collective communication, so that the artificial intelligence device can maintain stable data transmission and processing capabilities in high-load scenarios such as multi-device collaborative computing and large-scale model training, while reducing the risk of system stagnation caused by resource competition. It is particularly suitable for parallel processing requirements in many-to-many communication scenarios and significantly improves the overall stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 It is a schematic diagram of the structure of the artificial intelligence device provided by the present invention; Figure 2It is a schematic diagram of the structure of the request queue provided by the present invention; Figure 3 It is a flowchart of a deadlock avoidance method in collective communication provided by the present invention; Figure 4 Schematic diagram of the deadlock avoidance system architecture in multi-GPU collective communication provided by the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0020] With the rapid development of AI (artificial intelligence) and high-performance computing, AI device clusters, with their powerful parallel computing capabilities, have become the mainstream platform for processing large amounts of data and complex algorithms. Here, AI devices can include graphics processing units (GPUs), general-purpose graphics processing units (GPGPUs), and tensor processing units (TPUs). For ease of understanding, the following explanation uses a GPU as an example.
[0021] In high-performance computing and deep learning training, multi-GPU systems widely utilize collective communication operations such as AllReduce, AllGather, and Broadcast to achieve efficient global data synchronization. These operations typically involve many-to-many communication between multiple GPUs, where all GPUs act as both data senders (also known as sources) and data receivers (also known as targets), forming complex interdependencies. However, in this many-to-many communication model, deadlock becomes a key challenge affecting system reliability.
[0022] Deadlock occurs when multiple compute nodes or GPUs become blocked indefinitely during communication due to each other waiting for the other to release resources, preventing the entire system from continuing. In multi-GPU collective communication, deadlock is often caused by circular waiting and resource contention. Traditional collective communication implementations typically use a single-channel communication model, where requests and responses share the same scheduler and buffer. When multiple GPUs initiate communication simultaneously, requests can fill up the buffer, preventing responses from being written. The target GPU, unable to release the buffer due to a lack of responses, then forms a circular dependency.
[0023] Taking the AllReduce operation as an example, all GPUs need to send and receive data simultaneously. If the communication buffer is full and some GPUs stop sending new data due to being unable to receive responses, other GPUs, relying on this data, will also be stalled, eventually forming a global deadlock. For example, GPU0 sends a request to GPU1 (such as reading data) and waits for GPU1's response. At this point, GPU0's request has occupied GPU1's buffer resources. While processing the request, GPU1 may need to send another request to GPU2 (for example, to obtain necessary data). Therefore, GPU1 also enters a wait state, relying on GPU2's response. After receiving GPU1's request, GPU2 may need to initiate another request to GPU0 due to computing or communication needs, and thus wait for GPU0's response. At this point, the dependencies among the three GPUs form a closed loop, forming a circular wait chain. Assume that during this process, other GPUs send requests to these three GPUs, causing the buffer of each GPU to be full. This will make GPU1 unable to process GPU0's request because when it requests GPU2, GPU2's buffer is full, causing GPU1 to be blocked; GPU2 is also blocked because GPU0's buffer is full; GPU0 cannot continue because of GPU1's blockage, causing all GPUs to stagnate and form a deadlock.
[0024] Once deadlock occurs, the system's computing tasks will become stagnant and unable to move forward, which will not only lead to a waste of computing resources, but may also cause serious consequences such as system crashes or data loss. In this regard, the present invention proposes a deadlock avoidance method in collective communication. By setting a request scheduler, a response scheduler, a request buffer, and a response buffer on the artificial intelligence device at the same time, the physical isolation of the request and response paths is achieved, fundamentally eliminating the deadlock condition of circular waiting, thereby effectively solving the deadlock problem in multi-device collective communication and overcoming the above-mentioned defects. The technical solution of the present invention will be described in detail below.
[0025] Figure 1 This is a schematic diagram of the structure of the artificial intelligence device provided by the present invention. Figure 1As shown, the device includes a request scheduler 110, a response scheduler 120, a request buffer 130 and a response buffer 140; The request buffer 130 is configured to store a received first data request, where the first data request is sent by the first device; The request scheduler 110 is configured to obtain the first data request from the request buffer 130, process the first data request, and generate a first processing response; The response buffer 140 is used to store the generated first processing response; The response scheduler 120 is configured to obtain the first processing response from the response buffer 140 and return the first processing response to the first device.
[0026] It should be noted that in multi-GPU collective communication, each GPU is equipped with a request scheduler, a response scheduler, a request buffer, and a response buffer. Both the request scheduler and the response scheduler can be hardware or software modules in the communication protocol stack. These two types of schedulers manage the communication process. The request scheduler is primarily responsible for processing data requests (such as data reads, data writes, and synchronization instructions). The response scheduler is primarily responsible for processing responses corresponding to data requests (such as read data, operation completion confirmation, etc.). It should be understood that a data request is a communication instruction initiated by a source device to a target device. It may include the request type (such as data read, data write), request data (i.e., specific data related to the request, such as the data address), and metadata (such as the request identifier, priority, data size, data type, and source and target device identifiers). Requests are typically transmitted in the form of data packets. A response is feedback data generated by the target device after executing an operation on a pending data request. It may include a response identifier, a request identifier (used to identify the corresponding original request), a response type, and result data (such as read data and a status code).
[0027] The request buffer and response buffer refer to two independent physical or virtual storage areas on the device. These two types of buffers are used to isolate request and response data. The request buffer is used to store received data requests; the response buffer is used to store pending and received processed responses. It should be understood that both the request buffer and the response buffer can be high-speed storage areas on the device, such as GPU memory or on-chip cache. The physical isolation between the request and response buffers ensures that responses are not blocked by request traffic.
[0028] Specifically, in multi-device communication, when a first device (e.g., GPU0) sends a data request to an AI device (e.g., GPU1), the first device can be the source device, and the AI device can be the destination device. Here, the source device refers to the device currently sending the data request. The destination device refers to the device currently receiving and processing the data request. It should be understood that, like the AI device, the first device is also equipped with a request scheduler, a response scheduler, a request buffer, and a response buffer. For ease of understanding, the following example uses GPU0 as the first device and GPU1 as the AI device.
[0029] The first device, GPU0 (the source device), can send the first data request to the artificial intelligence device, GPU1 (the destination device), via its local request scheduler. The request scheduler on GPU0 can send the request to GPU1's request buffer via a communication link, such as a PCIe (Peripheral Component Interconnect Express) interface.
[0030] After receiving the first data request, the local request scheduler of the target device GPU1 will take out the data request from the request buffer and process the request to obtain a processing result. Here, when the target device GPU1 receives multiple first data requests, the request scheduler on GPU1 can take out the data requests to be processed from the request buffer in sequence and process them accordingly according to strategies such as FIFO (First In First Out) or priority scheduling. It should be understood that the processing result refers to the output data or status information obtained after the target device (such as GPU1) performs an operation on the data request to be processed. For example, the processing result can be a data read result (i.e., the actual data read from the memory) or status information, i.e., a flag indicating whether the operation is successful (such as a success / failure code).
[0031] Specifically, the request scheduler on the target device GPU1 will first check whether the local request buffer is empty. If it is not empty, it can take out the data request to be processed from the request buffer in FIFO or priority order. Then, the request scheduler will parse the content of the request, extract key information (such as request type, target address, parameters, etc.), and perform corresponding operations according to the request type. For example, if the data request to be processed is a data read request, the data can be read from the specified memory address according to the target address obtained by parsing, and the read data is the processing result. Subsequently, GPU1 will generate a corresponding first processing response based on the processing result, and store the processing response in the local response buffer so that the subsequent response scheduler can obtain the processing response from the response buffer and return it to the source device GPU0.
[0032] It is understandable that for the target-side GPU, its local response buffer can serve as an intermediate temporary storage layer, allowing the target side to quickly release computing resources (such as writing the response to the response buffer immediately after processing the request, and releasing the corresponding request in the request buffer) without waiting for the source-side GPU to actually complete the reception, thereby cutting off the blocking chain. If the target-side GPU sends the response directly to the source side without going through the local buffer, it may be blocked due to instantaneous resource competition. For example, the response buffer of the source side is full and cannot receive a new response; or, the response scheduler of the target side cannot send it immediately due to network congestion. In an embodiment of the present invention, the processed response is first temporarily stored in the local response buffer, and then the response scheduler sends the response on demand according to the network conditions and the source-side buffer space, thereby avoiding deadlock. In addition, the response buffer can be a FIFO queue, which can ensure that the responses are submitted to the response scheduler in the order of generation and then sent in sequence, thereby ensuring that the communication process is reliable and stable.
[0033] Furthermore, after receiving the first processed response returned by the response scheduler of GPU1, the source device GPU0 will store the response in its local response buffer to trigger GPU0 to perform subsequent processing. For example, after the first processed response is stored in the response buffer, an interrupt can be triggered to notify the response scheduler on GPU0. For another example, the response scheduler on GPU0 can adopt a polling mechanism to periodically check whether the number of responses in the response buffer has changed. When a notification is received or a new response is detected, the response scheduler can read the corresponding response data from the response buffer and parse it to extract the response identifier, request identifier, response type, result data, etc.
[0034] The source device can store the extracted result data in local memory (such as GPU registers or video memory) or update local status (such as marking a write task as complete). The response scheduler on the source device can then release the processed response in the response buffer so that a new processed response can be written. Releasing a response here means removing the processed response from the response buffer and reclaiming the storage space it occupied so that a new response can be stored.
[0035] For example, using the AllReduce operation as an example, the request scheduler on the source side, GPU0, can send a data write request to the target side, GPU1. This request is stored in GPU1's request buffer. GPU1's request scheduler then retrieves the request from the request buffer, parses it, and performs a Reduce operation on the parsed data blocks, generating a result. This result is then temporarily stored in the response buffer. Subsequently, GPU1's response scheduler retrieves the result from the response buffer and writes it to GPU0's response buffer. GPU0's response scheduler then reads the result from the response buffer, completing this round of communication. It should be understood that in this embodiment of the present invention, the independent operation of the request and response schedulers ensures that control signals (i.e., responses) are not blocked by data traffic (i.e., requests) during high loads. Even if the request buffer is full (e.g., GPU1's request buffer is overloaded with requests), responses can still be returned through the independent response buffer and response scheduler, avoiding circular waiting and thus preventing deadlock.
[0036] The device provided by the embodiment of the present invention realizes the physical isolation of the request and response paths by simultaneously setting a request scheduler, a response scheduler, a request buffer and a response buffer on the artificial intelligence device, thereby effectively solving the deadlock problem in the collective communication of multiple devices. The request scheduler is specifically responsible for extracting and processing the first data request sent by the first device from the request buffer, and the generated first processing response is stored in the response buffer, and then the response scheduler is responsible for returning it to the first device. This separation design ensures that the processing flows of the request and response do not interfere with each other, and avoids resource competition in two-way communication. Even in high concurrency or resource competition scenarios, the blocking of the request channel will not affect the normal operation of the response channel, and vice versa. By decoupling the storage and scheduling logic of requests and responses, the present invention not only avoids the deadlock caused by circular waiting, but also improves the reliability and efficiency of collective communication, so that the artificial intelligence device can maintain stable data transmission and processing capabilities in high-load scenarios such as multi-device collaborative computing and large-scale model training, while reducing the risk of system stagnation caused by resource competition. It is particularly suitable for parallel processing requirements in many-to-many communication scenarios and significantly improves the overall stability of the system.
[0037] Based on the above embodiment, the response scheduler 120 is further configured to: The first data request in the request buffer is released.
[0038] Specifically, after the request scheduler of the target device (i.e., the AI device) completes processing the first data request and receives the corresponding first processing response, it stores the processing response in a local response buffer. Subsequently, the target device can use the response scheduler to release the first data processing request from the request buffer. Releasing a request here means removing the request from the request buffer and reclaiming the storage space it occupied, so that subsequent requests can be stored.
[0039] In addition, for the target device, after the response scheduler returns the corresponding first processing response in the response buffer to the first device (ie, the source device), the processing response in the response buffer may be released.
[0040] Based on any of the above embodiments, the request scheduler 110 is further configured to send a second data request to the second device; The response buffer 140 is further configured to store a second processing response returned by the second device, where the second processing response is generated after the second device processes the second data request.
[0041] It's important to note that in multi-GPU collective communication, the roles of source and destination devices are dynamic, and each GPU can be both a source and a destination at the same time. For example, when GPU0 sends a data request to GPU1, GPU1 is the destination device; when GPU1 sends a data request to GPU2, GPU1 becomes the source device.
[0042] Specifically, when an AI device sends a data request to a second device, the AI device becomes the source device, and the second device becomes the destination device. Here, the second device refers to the device that receives the data request from the AI device, and the first device refers to the device that sends the data request to the AI device. The first and second devices can be different devices or the same device, and this is not specifically limited in the embodiments of the present invention. It should be understood that the first data request referred to in the present invention refers to a data request sent by the first device to the AI device, and the second data request refers to a data request sent by the AI device to the second device.
[0043] It is understandable that, like the artificial intelligence device, the second device is also provided with a request scheduler, a response scheduler, a request buffer, and a response buffer. For ease of understanding, the following description is based on an example in which the artificial intelligence device is GPU1 and the second device is GPU2.
[0044] AI device GPU1 (the source device) can send the second data request to GPU2 (the target device) via its local request scheduler. The request scheduler on GPU1 can send the request to GPU2's request buffer via a communication link, such as a PCIe interface. After receiving the second data request, GPU2's local request scheduler retrieves the request from the request buffer and processes it to generate a corresponding response (the second response).
[0045] After generating the second processed response, target device GPU2 temporarily stores it in its local response buffer. This allows the response scheduler on GPU2 to retrieve the second processed response from the local response buffer based on network conditions and the buffer space available on source device GPU1, and return it to source device GPU1. After receiving the second processed response from GPU2's response scheduler, source device GPU1 stores the response in its local response buffer to trigger GPU1 to perform subsequent processing.
[0046] Based on any of the above embodiments, the device further includes a request queue; The request queue is used to store the second data request generated locally; The request scheduler 110 is further configured to, upon detecting that corresponding resources exist in the second device, obtain the second data request from the request queue and send the second data request to the second device.
[0047] Specifically, when an AI device acts as a source device, it can also be configured with a request queue. Here, a request queue refers to a local storage structure on the device that temporarily stores pending data requests (i.e., second data requests). This queue mechanism allows the source device to buffer bursts of traffic, preventing requests from being lost or blocked due to insufficient resources on the target device (i.e., second device). Furthermore, the source device can generate requests and queue them in advance, eliminating the need to wait for the target device to become ready.
[0048] Specifically, the request scheduler on the source device can employ a polling mechanism to detect the availability of the target device's corresponding resources (e.g., network bandwidth, request buffer, etc.). When these resources are available, the scheduler retrieves the data request from the request queue and sends it to the target device. Here, the target device's corresponding resources refer to the resources available to the target device to properly receive the source device's data request. It should be understood that the request scheduler on the source device can employ strategies such as FIFO or priority scheduling to dynamically determine when to send requests from the queue.
[0049] Furthermore, if the request scheduler on the source device detects that the target device (i.e., the second device) does not have the corresponding resources, this indicates that the target device is currently unable to process the new request. Sending the request directly may result in request loss or blocking. In this case, the source device can temporarily store the request in a local request queue until the target device's resources become available.
[0050] In an embodiment of the present invention, the source device first stores the data request in a local request queue. Only when it detects that the target device has corresponding resources will it obtain the data request from the local request queue and send it, rather than sending it directly. This avoids request loss or blocking due to insufficient resources of the target device, and buffers burst traffic through the request queue, which helps to reduce the instantaneous impact on the target resources.
[0051] Figure 2 It is a schematic diagram of the structure of the request queue provided by the present invention, such as Figure 2 As shown, considering that collective communication relies on strict timing, for example, GPU1 sends requests A, B, and C, which must be processed in sequence by GPU2. Disorder may lead to incorrect calculation results. Therefore, in order to ensure this timing, the request queue on the source device can be set as a first-in-first-out queue. New requests are stored in the queue in the order of arrival and are taken out in the same order during processing. This ensures that the data requests generated by the source device can be sent to the target device in sequence for processing. Specifically, the queue stores the id of each request (that is, the unique identifier of the request, also called the request identifier) and the request content (recorded using the request field, such as data, instructions, etc.). In the request queue, the end (tail) pointer points to the head of the free area in the queue, that is, the location where the next new request will be stored; the start (head) pointer points to the head of the request to be sent in the queue, that is, the next request to be sent to the downstream (that is, the target end). It should be noted that Figure 2 The gray area shown in the figure represents all data requests waiting to be sent in the request queue. All operations in the request queue (such as enqueue, dequeue, and release) are scheduled serially, meaning only one operation is executed on the queue at a time.
[0052] For example, in a multi-GPU collective communication scenario, the source device can be a GPU (such as GPU1). When the SPC (Stream Processor Cluster) of the GPU generates a new data request, the request can be placed in the request queue through the following steps: first, according to the size of the request content, storage space is allocated in the free area of the queue (marked by the end pointer); then, the request ID and content are stored in the allocated space; finally, the end pointer is moved backward to point to the new free area head.
[0053] The request scheduler on the source device can periodically check whether the corresponding resources (such as network bandwidth, receive buffer, etc.) of the target device are ready. For example, the source request scheduler can adopt a polling strategy, checking the resource status every 100 microseconds. Here, resource detection mechanisms (such as credit mechanisms and status queries) can be used to confirm whether the target device can receive requests. When the source request scheduler detects that the target device has the corresponding resources (such as the target request buffer is idle and the computing resources are ready), it indicates that the target device is currently able to receive and process new requests. At this time, the request scheduler on the source device can remove the request pointed to by the start pointer (including ID and request) from the queue and send it to the target device through the communication interface. Subsequently, the start pointer is moved backward to point to the next data request to be sent.
[0054] Based on any of the above embodiments, the request scheduler 110 is further configured to: Obtaining the second processing response from the response buffer, and parsing the second processing response to obtain a target request identifier; The data request matching the target request identifier in the request queue is released, and the second processing response in the response buffer is released.
[0055] Specifically, in a scenario where an AI device, acting as a source device, sends a second data request to a second device (i.e., a target device), once the response scheduler on the target device writes the second processed response to the response buffer on the source device, it triggers the request scheduler on the source device to release the corresponding request from the request queue, and also triggers the response scheduler on the source device to release the corresponding response from the local response buffer. Releasing a request here means removing the request from the request queue and reclaiming the storage space it occupied so that subsequent new requests can be stored in the queue.
[0056] Specifically, after the target-side response scheduler writes the processed response to the source-side response buffer, the source-side response buffer can send a notification to the source-side request scheduler. Upon receiving the notification, the source-side request scheduler can read the corresponding response data from the source-side response buffer and parse the response to extract the target request identifier, response type, and result data. Based on the extracted target request identifier, the request queue can be searched for a data request whose request ID matches the target request identifier. The data request can then be released, such as by removing it from the queue and reclaiming the storage space it occupies.
[0057] Furthermore, after releasing the corresponding request in the request queue, the request scheduler at the source end may also release the processed response in the response buffer at the source end so that a new response can be written subsequently.
[0058] It can be understood that in the above embodiment, the release of requests and responses is completed by the request scheduler of the source end. In another embodiment, the request scheduler of the source end can release the corresponding request in the request queue, and the response scheduler of the source end can release the corresponding response in the response buffer. Specifically, after the response scheduler of the target end writes the second processing response to the response buffer of the source end, the response buffer of the source end can simultaneously notify the request scheduler and response scheduler of the source end so that the request scheduler of the source end releases the corresponding request in the request queue; at the same time, the response scheduler of the source end will also read the corresponding processing response from the response buffer of the source end and parse the response. The result data obtained by the parsing can be stored in the local memory or the local state can be updated. Subsequently, the processing response in the response buffer of the source end can be released.
[0059] Based on any of the above embodiments, the first processing response is returned to the first device with write semantics, and the second processing response is returned to the artificial intelligence device with write semantics.
[0060] It should be noted that in traditional collective communication scenarios, responses to data requests are typically returned to the sender (i.e., the source device) with read semantics. In other words, after the source device sends a data request to the target device, it must actively read the response from the target device through a polling mechanism or port monitoring mechanism. For example, when GPU0 needs to read data from GPU1, GPU0 typically sends a data read request to GPU1, which then actively pulls the corresponding data from GPU1. During this process, GPU1 must wait for GPU0 to actively read the data, resulting in significant resource overhead. To address this, embodiments of the present invention propose returning the response to the source device with write semantics, eliminating the need for the source device to actively read data and thus reducing resource overhead. It should be understood that returning a response with write semantics means that after processing the source device's request, the target device returns a response using certain write operation behaviors and semantics. This design, which implements read semantics with write semantics, helps reduce resource overhead.
[0061] Specifically, the first processing response is a processing response returned by the artificial intelligence device to the first device. In this case, the artificial intelligence device acts as the destination device, and the first device acts as the source device. The second processing response is a processing response returned by the second device to the artificial intelligence device. In this case, the artificial intelligence device acts as the source device, and the second device acts as the destination device. Therefore, both the first processing response and the second processing response are processing responses returned by the destination device to the source device.
[0062] Specifically, the response scheduler on the target device can obtain the address of the response buffer on the source device through the communication interface and write the response data directly from local memory to the source's response buffer using the DMA (Direct Memory Access) engine. In this embodiment of the present invention, by using write semantics, response data is actively pushed to the sender (i.e., the source device) rather than pulled by the sender. This approach reduces the wait time for the sender to respond and optimizes buffer usage.
[0063] Based on any of the foregoing embodiments, the first data request is any one of a data read request and a data write request, and the second data request is any one of a data read request and a data write request.
[0064] Specifically, a first data request is a data request sent from a first device to an AI device. In this case, the first device acts as a source device, and the AI device acts as a destination device. A second data request is a data request sent from an AI device to a second device. In this case, the AI device acts as a source device, and the second device acts as a destination device. Therefore, both the first and second data requests are data requests sent from a source device to a destination device. These data requests can be either data read or data write requests.
[0065] Here, a data read request is an operation initiated by a source device to a target device, intended to retrieve specified data from the target device's memory. A data write request is another operation initiated by a source device to a target device, intended to write specified data to the target device's memory. It should be understood that in distributed training or computing, data read and write requests are the foundation of inter-GPU communication, supporting operations such as model parameter synchronization and gradient aggregation.
[0066] When a data request is a read request, the corresponding processing response can include the read data; when a data request is a write request, the corresponding processing response can include write status information (such as success or failure). Responses to both data read and data write requests are returned to the source device using write semantics, helping to reduce resource overhead.
[0067] Based on any of the above embodiments, the request buffer and the response buffer are both first-in-first-out queues.
[0068] Specifically, in multi-GPU collective communication, each GPU's request and response buffers are managed using a first-in-first-out (FIFO) data structure. This design is crucial for the reliability, sequentiality, and deadlock avoidance of the communication process. A first-in-first-out (FIFO) queue is a linear data structure where data is processed and retrieved in the order it arrives. Requests and responses that enter the queue first are processed first, while those that enter later are queued, strictly adhering to sequential order.
[0069] Specifically, whether it is a source device or a target device, the request buffer on it is mainly used to store received data requests so that the request scheduler can process the requests in FIFO order; the response buffer is used to store the processing responses to be sent or received so that the response scheduler can take out the responses in FIFO order and notify the upper-level application.
[0070] Based on any of the above embodiments, Figure 3 : is a flow chart of the deadlock avoidance method in collective communication provided by the present invention, such as Figure 3 As shown, the method is applied to an artificial intelligence device, the artificial intelligence device including a request scheduler, a response scheduler, a request buffer, and a response buffer, and the method includes: Step 310: Receive a first data request sent by a first device, and store the first data request in the request buffer; Step 320: Based on the request scheduler, obtain the first data request from the request buffer, process the first data request, and generate a first processing response; Step 330: Store the first processed response in the response buffer, and return the first processed response in the response buffer to the first device based on the response scheduler.
[0071] It should be noted that the methods provided in the embodiments of the present invention can be applied to artificial intelligence devices. Such devices are provided with a request scheduler, a response scheduler, a request buffer, and a response buffer. Both the request scheduler and the response scheduler can be hardware or software modules in the communication protocol stack, and these two types of schedulers are used to manage the communication process. The request scheduler is primarily responsible for processing data requests (such as data reads, data writes, synchronization instructions, etc.). The response scheduler is primarily responsible for processing responses corresponding to data requests (such as read data, operation completion confirmation, etc.).
[0072] The request buffer and response buffer refer to two independent physical or virtual storage areas on the device. These two types of buffers are used to isolate request and response data. The request buffer is used to store received data requests; the response buffer is used to store pending and received processed responses. It should be understood that both the request buffer and the response buffer can be high-speed storage areas on the device, such as GPU memory or on-chip cache. The physical isolation between the request and response buffers ensures that responses are not blocked by request traffic.
[0073] Specifically, in multi-device communication, when a first device (e.g., GPU0) sends a data request to an AI device (e.g., GPU1), the first device can be the source device, and the AI device can be the destination device. Here, the source device refers to the device currently sending the data request. The destination device refers to the device currently receiving and processing the data request. It should be understood that, like the AI device, the first device is also equipped with a request scheduler, response scheduler, request buffer, and response buffer. For ease of understanding, the following example uses GPU0 as the first device and GPU1 as the AI device.
[0074] For the first device GPU0 (i.e., the source device), it can send the first data request to the artificial intelligence device GPU1 (i.e., the target device) through the local request scheduler. After receiving the first data request, the target device GPU1 will store it in the local request buffer. Subsequently, the request scheduler of the target device GPU1 can take out the first data request to be processed from the local request buffer in sequence according to strategies such as FIFO or priority scheduling, and process the request to obtain the processing result. Here, the processing result refers to the output data or status information obtained after the target device performs an operation on the data request to be processed. For example, the processing result can be a data read result (i.e., the actual data read from the memory), or it can be status information, that is, a sign of whether the operation is successful (such as a success / failure code).
[0075] Subsequently, the target device, GPU1, generates a corresponding first processing response based on the processing results and returns it to the response buffer of the source device, GPU0, via the local response scheduler for subsequent processing by the source device. Here, a processing response refers to the feedback data generated by the target after performing an operation on the pending data request. It may include the request ID (used to identify the corresponding original request), the response type, and result data (such as the read data and status code).
[0076] The method provided by the embodiment of the present invention realizes the physical isolation of the request and response paths by simultaneously setting a request scheduler, a response scheduler, a request buffer and a response buffer on the artificial intelligence device, thereby effectively solving the deadlock problem in the collective communication of multiple devices. The request scheduler is specifically responsible for extracting and processing the first data request sent by the first device from the request buffer, and the generated first processing response is stored in the response buffer, and then the response scheduler is responsible for returning it to the first device. This separation design ensures that the processing flows of the request and response do not interfere with each other, and avoids resource competition in two-way communication. Even in high concurrency or resource competition scenarios, the blocking of the request channel will not affect the normal operation of the response channel, and vice versa. By decoupling the storage and scheduling logic of requests and responses, the present invention not only avoids the deadlock caused by circular waiting, but also improves the reliability and efficiency of collective communication, so that the artificial intelligence device can maintain stable data transmission and processing capabilities in high-load scenarios such as multi-device collaborative computing and large-scale model training, while reducing the risk of system stagnation caused by resource competition. It is particularly suitable for parallel processing requirements in many-to-many communication scenarios and significantly improves the overall stability of the system.
[0077] It should be noted that other embodiments or specific implementations of the deadlock avoidance method in collective communication provided by the present invention can refer to the above-mentioned device embodiments and will not be repeated here.
[0078] Based on any of the above embodiments, Figure 4 This is a schematic diagram of the deadlock avoidance system architecture in multi-GPU collective communication provided by the present invention, such as Figure 4 As shown, the system can include multiple GPUs, each of which can be either a source or a destination. GPU0 can send data requests to multiple GPUs (such as GPU1 and other GPUs) or receive data requests from one or more GPUs (such as GPU1 and other GPUs). Similarly, GPU1 can receive data requests from one or more GPUs (such as GPU0 and other GPUs) or send data requests to multiple GPUs (such as GPU0 and other GPUs).
[0079] like Figure 4 As shown, each GPU is equipped with a request scheduler, a response scheduler, a request buffer, and a response buffer. The request scheduler processes data read and write requests, while the response scheduler processes data read and write responses. The request buffer stores received data requests, and the response buffer stores pending or received responses. For ease of understanding, the following description uses GPU0 as the source and GPU1 as the destination.
[0080] Suppose that source GPU0 needs to read a data block from the memory of target GPU1 to perform a reduction calculation. In this scenario, the data reading process steps include: S1, source GPU generates request: Specifically, the SPC of the source GPU0 generates a data read request and specifies relevant information, such as the target device (i.e., GPU1), the data address (i.e., the memory address of GPU1), the data size, etc. The request is transmitted to the request scheduler of GPU0 via the NoC (Network OnChip), and the request scheduler stores it in the local request queue ( Figure 4 not shown).
[0081] S2, the source GPU sends a request: Specifically, when GPU0's request scheduler detects a data read request in the request queue and GPU1 has the corresponding resources, it removes the request from the request queue and sends it to GPU1's request buffer via a communication link (such as a PCIe interface). In this case, GPU0's scheduler acts as the sender (i.e., Tx), and GPU1's request buffer acts as the receiver (i.e., Rx).
[0082] S3, target GPU processing request: Specifically, GPU1's request scheduler retrieves the data read request from its local request buffer and, based on the address in the request, reads the corresponding data from local memory (such as the L2 cache or HBM (High Bandwidth Memory)) via the NoC. The read data is then stored in the local response buffer, which is then written to GPU0's response buffer by the response scheduler. At this point, GPU1's response scheduler acts as the sender (i.e., Tx), and GPU0's response buffer acts as the receiver (i.e., Rx). Subsequently, GPU1's response scheduler releases the corresponding request in its local request buffer and the corresponding response in its local response buffer.
[0083] S4, the source GPU receives the response: Specifically, GPU0's request scheduler retrieves the processing response from the response buffer, parses the data, and stores it in local memory for subsequent processing and calculation. The request scheduler then releases the data read request from the request queue and clears the corresponding response from the response buffer.
[0084] It is understandable that in the aforementioned interaction between GPU0 and GPU1, although GPU0 is the source and GPU1 is the destination, GPU0 can also act as the destination, receiving data requests from other GPUs and storing these requests in a local request buffer. This allows the request scheduler to sequentially retrieve these requests from the request buffer for processing according to the FIFO principle and return the generated processed responses to the other GPUs via the response scheduler. Similarly, GPU1 can also act as the source, sending data requests to other GPUs via the request scheduler and allowing the other GPUs to write the processed responses into the response buffer.
[0085] Clearly, for both GPU0 and GPU1, by separating the request and response channels, the request stream (i.e., data read requests and data write requests) can be independently processed by the request scheduler and request buffer, while the response stream (i.e., data read responses and data write responses) can also be independently processed by the response scheduler and response buffer. This design, by physically isolating the request and response channels, prevents the request and response flows from interfering with each other. Even if the request channel is blocked due to a full buffer, the response can still be delivered through the independent buffer, ensuring that critical control signals (such as ACK) are not blocked, thus preventing the system from entering a waiting loop and resolving deadlock issues in multi-GPU communication.
[0086] In addition, both data read responses and data write responses are returned to the source device with write semantics, that is, read semantics are implemented with write semantics, which helps reduce resource overhead.
[0087] The present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the deadlock avoidance method in collective communication provided by the above-mentioned methods. The method is applied to an artificial intelligence device, which includes a request scheduler, a response scheduler, a request buffer and a response buffer. The method includes: receiving a first data request sent by a first device and storing the first data request in the request buffer; based on the request scheduler, obtaining the first data request from the request buffer, processing the first data request, and generating a first processing response; storing the first processing response in the response buffer, and based on the response scheduler, returning the first processing response in the response buffer to the first device.
[0088] Furthermore, the aforementioned computer program can be implemented as a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the relevant art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0089] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the deadlock avoidance method in collective communication provided by the above-mentioned methods. The method is applied to an artificial intelligence device, and the artificial intelligence device includes a request scheduler, a response scheduler, a request buffer and a response buffer. The method includes: receiving a first data request sent by a first device, and storing the first data request in the request buffer; based on the request scheduler, obtaining the first data request from the request buffer, and processing the first data request to generate a first processing response; storing the first processing response in the response buffer, and based on the response scheduler, returning the first processing response in the response buffer to the first device.
[0090] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0091] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An artificial intelligence device, characterized in that: Includes request scheduler, response scheduler, request buffer and response buffer; The request buffer is configured to store a received first data request, where the first data request is sent by a first device; The request scheduler is configured to obtain the first data request from the request buffer, process the first data request, and generate a first processing response; The response buffer is used to store the generated first processing response; The response scheduler is configured to obtain the first processing response from the response buffer and return the first processing response to the first device.
2. The artificial intelligence device according to claim 1, characterized in that The response dispatcher is further configured to: The first data request in the request buffer is released.
3. The artificial intelligence device according to claim 1, characterized in that The request scheduler is further configured to send a second data request to the second device; The response buffer is further used to store a second processing response returned by the second device, where the second processing response is generated after the second device processes the second data request.
4. The artificial intelligence device according to claim 3, characterized in that Also includes request queue; The request queue is used to store the second data request generated locally; The request scheduler is further configured to, upon detecting that corresponding resources exist in the second device, obtain the second data request from the request queue and send the second data request to the second device.
5. The artificial intelligence device according to claim 4, characterized in that The request scheduler is further configured to: Obtaining the second processing response from the response buffer, and parsing the second processing response to obtain a target request identifier; The data request matching the target request identifier in the request queue is released, and the second processing response in the response buffer is released.
6. The artificial intelligence device according to any one of claims 3 to 5, characterized in that The first processed response is returned to the first device with write semantics, and the second processed response is returned to the artificial intelligence device with write semantics.
7. The artificial intelligence device according to any one of claims 3 to 5, characterized in that The first data request is any one of a data read request and a data write request, and the second data request is any one of a data read request and a data write request.
8. The artificial intelligence device according to any one of claims 1 to 5, characterized in that: The request buffer and the response buffer are both first-in-first-out queues.
9. A deadlock avoidance method in collective communication, characterized in that: The method is applied to an artificial intelligence device, the artificial intelligence device including a request scheduler, a response scheduler, a request buffer, and a response buffer, and the method includes: receiving a first data request sent by a first device, and storing the first data request in the request buffer; Based on the request scheduler, obtain the first data request from the request buffer, and process the first data request to generate a first processing response; The first processing response is stored in the response buffer, and based on the response scheduler, the first processing response in the response buffer is returned to the first device.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the deadlock avoidance method in collective communication according to claim 9 is implemented.