Method and monitoring system for monitoring a direct connection link between computing devices
By reading the advanced error report register information of the direct connection port of the computing device, determining the error type and sending interrupt instructions, the problem of monitoring and diagnosis of the direct connection link between computing devices is solved, real-time monitoring and diagnosis of the direct connection link is realized, and the normal data transmission is ensured.
Patent Information
- Application Number
- CN202410451686.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-04-15
AI Technical Summary
The prior art cannot monitor and diagnose direct link errors between computing devices, making it difficult to diagnose errors occurring at direct links in a timely manner.
By reading the information in the advanced error report register in the direct connection port of the computing device, determining the error type, and sending an interrupt instruction to the interrupt collector based on the error type to indicate that the direct connection link is in an abnormal working state, real-time monitoring and diagnosis of the direct connection link between computing devices is achieved.
Real-time monitoring and diagnosis of direct links between computing devices can be realized, link errors can be discovered and processed in a timely manner, and data transmission is carried out normally.
Smart Images

Figure CN118245328B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the field of chip technology, and more particularly to a method and a monitoring system for monitoring a direct connection link between computing devices. Background Art
[0002] The PCIe (Peripheral Component Interconnect Express) bus is a high-bandwidth expansion bus, commonly used to connect a host (such as a central processing unit (CPU)) to various computing devices (such as a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), etc.) to enable interaction between the host and the computing devices, such as data transmission. In other words, the data transmission between the host and the computing devices implemented via the PCIe bus complies with the PCIe bus standard (also known as the PCIe protocol).
[0003] On this basis, if one of the multiple computing devices connected to the CPU is regarded as a master device, the master device and other computing devices can achieve direct connection between the computing devices based on the PCIe protocol, and data transmission between the computing devices can be performed without going through the CPU.
[0004] However, the prior art only supports monitoring the connection link based on the PCIe protocol between the host and the computing devices, but cannot monitor the direct connection link based on the PCIe protocol between the computing devices. That is to say, the disadvantage of the prior art is that it is impossible to directly monitor the direct connection link between the computing devices, resulting in difficulty in diagnosing the errors occurring at the direct connection link in a timely manner. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method and a monitoring system for monitoring a direct connection link between computing devices, enabling real-time monitoring of each direct connection port located at each computing device to monitor the direct connection link based on the PCIe protocol formed between different ports of each computing device, thereby realizing the monitoring and diagnosis of the direct connection link between the computing devices.
[0006] According to a first aspect of the present invention, there is provided a method for monitoring a direct connection link between computing devices, the method comprising: reading information in an advanced error reporting register in a direct connection port of a computing device, wherein the computing device is at least at one end of the direct connection link; in response to the information read including report information related to an error, determining the type of the error based on the report information related to the error; and sending an interrupt instruction to an interrupt collector of the computing device at least based on the type of the error to indicate that the direct connection link with the direct connection port as one end is in an abnormal working state.
[0007] In some embodiments, sending an interrupt instruction to an interrupt collector of a computing device based at least on the type of error includes: generating an interrupt instruction in response to the type of error being a first type; and sending the interrupt instruction to the interrupt collector of the computing device.
[0008] In some embodiments, sending an interrupt instruction to an interrupt collector of a computing device based at least on the type of error includes: counting the number of errors of a second type in response to the type of error being a second type; generating an interrupt instruction in response to the number of errors of the second type exceeding a threshold number; and sending the interrupt instruction to the interrupt collector of the computing device to indicate that a direct link with a direct port as one end is in an abnormal working state.
[0009] In some embodiments, counting the number of errors of a second type includes: controlling a counter value for counting the number of errors of the second type to increment by 1 in response to a current error of the second type being detected within a threshold time; and controlling the counter to be cleared in response to the threshold time being reached.
[0010] In some embodiments, the method according to the first aspect of the present invention further includes: re-reading information in an advanced error reporting register in a direct port of the computing device after a predetermined time in response to the read information not including report information related to an error.
[0011] In some embodiments, the method according to the first aspect of the present invention further includes: in response to the interrupt collector of the computing device receiving an interrupt instruction, sending an interrupt signal from the computing device to a host to report a port error indication to the host, where the port error indication is used to indicate that there is an error in the current direct port of the computing device and that a direct link with the current direct port as one end is in an abnormal working state.
[0012] According to a second aspect of the present invention, there is provided a monitoring system for monitoring a direct link between computing devices, the monitoring system including: at least one direct port detection module configured to be connected to a direct port of a computing device, where the direct port detection module includes: an advanced error reporting register monitoring module configured to read information in an advanced error reporting register in a direct port of the computing device and determine the type of error based on the report information related to the error in response to the read information including report information related to an error; an interrupt instruction generation and sending module configured to send an interrupt instruction to an interrupt collector of the computing device based at least on the type of error, where the interrupt instruction indicates that a direct link with a direct port as one end is in an abnormal working state.
[0013] In some embodiments, the types of errors include a first type and a second type. In these embodiments, the interrupt instruction generation and sending module is further configured to: generate an interrupt instruction in response to the type of error being the first type; and send the interrupt instruction to an interrupt collector of the computing device. In these embodiments, the interrupt instruction generation and sending module is further configured to: in response to the type of error being the second type, count the number of errors of the second type; generate an interrupt instruction in response to the number of errors of the second type exceeding a threshold number; and send the interrupt instruction to an interrupt collector of the computing device to indicate that a direct connection link with a direct connection port as one end is in an abnormal operating state.
[0014] In some embodiments, the advanced error reporting register monitoring module is further configured to: in response to the information read not including error-related reporting information, reread the information in the advanced error reporting register in the direct connection port of the computing device after a predetermined time.
[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements.
[0017] Figure 1 FIG. shows a schematic diagram of the topology of a system based on the PCIe protocol according to the present invention.
[0018] Figure 2 FIG. shows a schematic diagram of a monitoring system for monitoring a direct connection link between computing devices according to an embodiment of the present invention.
[0019] Figure 3 FIG. shows a flowchart of a method for monitoring a direct connection link between computing devices according to an embodiment of the present invention.
[0020] Figure 4 FIG. shows a framework diagram of an exemplary system according to the present invention, wherein the exemplary system includes a monitoring system for monitoring a direct connection link between computing devices according to an embodiment of the present invention.
[0021] Figure 5 FIG. shows a flowchart of a method for monitoring a direct connection link between computing devices according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The exemplary embodiments of the present invention will be described below with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0023] As used herein, the term "comprising" and its variations mean open inclusion, i.e., "including but not limited to". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included hereinafter.
[0024] As described above, a host (e.g., a CPU) and a computing device (e.g., a GPU, a GPGPU, etc.) can implement the interaction between the host and the computing device, such as data transmission, via a PCIe bus. In addition, data transmission between computing devices can also be achieved through a direct connection link based on the PCIe protocol. Figure 1 FIG. 100 shows a schematic diagram of the topology of a PCIe protocol-based system according to the present invention.
[0025] As Figure 1 shown, the system 100 includes a host (CPU) 110, a root complex (RC) 120, at least one PCIe switch (i.e., the first PCIe switch 130A and the second PCIe switch 130B, which can be collectively referred to as the PCIe switch 130), and at least one computing device (GPGPU or GPU) (i.e., the first computing device 140A, the second computing device 140B, the third computing device 140C, the fourth computing device 140D, which can be collectively referred to as the computing device 140).
[0026] As Figure 1 shown, at least one root port (RP), i.e., the first root port RP1 and the second root port RP2, can be configured on the RC 120.
[0027] As Figure 1As shown, each PCIe switch 130 may be configured with an upstream port (USP) and at least one downstream port (DSP). For example, the first PCIe switch 130A is configured with one upstream port (i.e., the first upstream port USP1) and two downstream ports (i.e., the first downstream port DSP1 and the second downstream port DSP2); similarly, the second PCIe switch 130B is configured with the second upstream port USP2, the third downstream port DSP3, and the fourth downstream port DSP2.
[0028] Generally, the root port of the RC 120 can be connected to the upstream port of the PCIe switch 130 to form a link for data transmission between the RC 120 and the PCIe switch 130. For example, as Figure 1 shown, the first root port RP1 of the RC 120 is connected to the first upstream port USP1 of the first PCIe switch 130A, and the second root port RP2 of the RC 120 is connected to the second upstream port USP2 of the second PCIe switch 130B.
[0029] As Figure 1 shown, each computing device 140 may be configured with an upper port for connecting to the PCIe switch 130. Taking the first computing device 140A as an example, it is configured with the first upper port P1, and the first upper port P1 is connected to the first downstream port DSP1 of the first PCIe switch 130A, thereby forming a link for data transmission between the first computing device 140A and the first PCIe switch 130A (i.e., Link0). Thus, the first computing device 140A can interact with the host CPU and the system memory (not shown) via the link between its first upper port P1 and the first downstream port DSP1 of the first PCIe switch 130A, the first PCIe switch 130A, and the link between the first upstream port USP1 of the first PCIe switch 130A and the first root port RP1 of the RC 120, and the RC 120. Similarly, the second computing device 140B, the third computing device 140C, and the fourth computing device 140D can interact with the host CPU and the system memory (not shown) in a similar manner.
[0030] Furthermore, each computing device 140 may also be configured with a direct connection port for interconnecting with other computing devices 140. As Figure 1As shown, for example, multiple direct connection ports are configured on the first computing device 140A, namely the first direct connection port PIP1, the second direct connection port PIP2, and the third direct connection port PIP3. The first direct connection port PIP1 of the first computing device 140A is connected to the direct connection port PIP4 of the second computing device 140B to form a direct connection link (i.e., Link1) for data transmission between the first computing device 140A and the second computing device 140B. Similarly, the second direct connection port PIP2 of the first computing device 140A is connected to the direct connection port PIP5 of the third computing device 140C to form a direct connection link (i.e., Link2) for data transmission between the first computing device 140A and the third computing device 140C; the third direct connection port PIP3 of the first computing device 140A is connected to the direct connection port PIP6 of the fourth computing device 140D to form a direct connection link (i.e., Link3) for data transmission between the first computing device 140A and the fourth computing device 140D. In other words, in the system 100 as Figure 1 shown, the first computing device 140A can achieve direct connections with the second computing device 140B, the third computing device 140C, and the fourth computing device 140D respectively via Link1, Link2, and Link3 based on the PCIe protocol, and perform interactions between computing devices through Link1, Link2, and Link3, such as accessing the memories of each other's computing devices.
[0031] In the related art, currently only the connection link based on the PCIe protocol between the monitoring host and the computing device is supported. For example, a PCIe protocol analyzer can be used to monitor and diagnose the connection link between the computing device and the host (including Link0 such as Figure 1 ). Specifically, a PCIe protocol analyzer can be used to monitor the data packets transmitted on Link0 and decode them into, for example, ordered sets, transaction layer packets (TLP), data link layer packets (DLLP), etc., to facilitate users in solving problems related to the PCIe protocol. In the related art, the status of each computing device can also be monitored by running software at, for example, the host. For example, PCIcrawler can be run at the host to display, filter, and export information related to the PCIe bus and the connected computing devices, such as PCIe bus errors. However, the above technologies can only monitor and diagnose the errors occurring on the connection link between the computing device and the host.
[0032] Given that the direct connection ports configured on a computing device for interconnecting with the computing device are different from the upstream ports for connecting to a PCIe switch. For example, a computing device is generally inserted into a motherboard through a gold finger (i.e., the upstream port) to achieve interconnection with a host; while for the direct connection port, a computing device usually uses a snap-on board or a Universal Base Board (UBB) to achieve interconnection between computing devices. Therefore, the direct connection port of a computing device cannot be directly inserted into the platform of a PCIe protocol analyzer. That is to say, it is difficult to directly use a PCIe protocol analyzer for monitoring and diagnosing a direct connection link. If a PCIe protocol analyzer is to be used, an additional adapter board needs to be customized so that the direct connection port can be inserted into the platform of the PCIe protocol analyzer via the adapter board, which increases the cost. Moreover, since the direct connection links between computing devices are not included in the topology of a traditional PCIe-based system, the existing software monitoring scope does not currently cover the direct connection links between computing devices. That is to say, it is difficult to monitor and diagnose the direct connection links between computing devices through software running on a host. In summary, the related art cannot monitor and diagnose errors occurring on the direct connection links between computing devices.
[0033] To at least partially solve one or more of the above problems and other potential problems, example embodiments of the present invention propose a solution for monitoring a direct connection link between computing devices. In this solution, information in an Advanced Error Reporting (AER) register in the direct connection port of a computing device at at least one end of the direct connection link is read; in response to the information read including a report related to an error, the type of the error is determined based on the report related to the error; and in response to the type of the error being a first type, an interrupt instruction is sent to an interrupt collector of the computing device to indicate that the direct connection link with the direct connection port as one end is in an abnormal working state, so that real-time monitoring of the direct connection link with the direct connection port as one end can be achieved by monitoring the direct connection port, thereby realizing the monitoring and diagnosis of the direct connection link between computing devices.
[0034] The following will be combined with Figures 2 to 5 to describe in detail a solution for monitoring a direct connection link between computing devices according to an embodiment of the present invention.
[0035] Figure 2 A schematic diagram of a monitoring system 200 for monitoring a direct connection link between computing devices according to an embodiment of the present invention is shown. It should be understood that the monitoring system 200 may further include additional modules not shown and / or modules shown may be omitted, and the scope of the present invention is not limited in this regard.
[0036] According to the inventive concept of the present aspect, the monitoring system 200 may include at least one direct connection port detection module configured to be connected to a direct connection port of a computing device, so as to facilitate real-time monitoring of the direct connection port, thereby enabling real-time monitoring of the direct connection link with the direct connection port as one end, and further enabling monitoring and diagnosis of the direct connection link between computing devices.
[0037] As Figure 2 shown, the monitoring system 200 includes a direct connection port detection module 210, which may be configured to be connected to a direct connection port of a computing device.
[0038] Regarding the computing device, it may be, for example, a GPU, a GPGPU.
[0039] Regarding the direct connection port, it may be a port configured on the computing device for direct interconnection with other computing devices. According to an embodiment of the present invention, the direct connection port may be a PCIe interconnection port (PIP) of the computing device.
[0040] As Figure 2 shown, the direct connection port detection module 210 may further include: an advanced error reporting register monitoring module 212 and an interrupt instruction generation and sending module 214.
[0041] Regarding the advanced error reporting register monitoring module 212, it may be connected to the advanced error reporting register of the direct connection port of the computing device to read the information stored therein. According to an embodiment of the present aspect, the advanced error reporting register monitoring module 212 may be configured to read the information in the advanced error reporting register in the direct connection port of the computing device.
[0042] Regarding the advanced error reporting register (AER register), it may be used to store information related to the occurred errors, for example, completion timeout, etc. According to an embodiment of the present aspect, each direct connection port of the computing device is configured with an advanced error reporting register to store error information related to the direct connection port. For example, when an error occurs on the direct connection link with a certain direct connection port as one end, the information related to the error may be stored in the advanced error reporting register of this direct connection port.
[0043] According to an embodiment of this aspect, the advanced error reporting register monitoring module 212 may also perform the following operations based on the read information, including but not limited to, for example: in response to the read information including error-related reporting information, determining the type of error based on the error-related reporting information; in response to the read information not including error-related reporting information, rereading the information in the advanced error reporting register in the direct connection port of the computing device after a predetermined time, etc.
[0044] Regarding the type of error, according to the definition in the PCIe protocol, errors can be classified into, for example, non-fatal errors and fatal errors. Among them, non-fatal errors refer to errors that do not affect the PCIe link (such as the direct connection link between computing devices) used for data transmission, while fatal errors refer to errors that can affect the PCIe link (such as the direct connection link between computing devices) used for data transmission. For example, non-fatal errors may include but are not limited to: completion timeout, unsupported request, completer abort, etc.; fatal errors may include but are not limited to: receiver overflow, flow control protocol error, etc. It should be understood that the type of error can also be defined according to other common rules in the art, and the present invention does not limit this.
[0045] Regarding the interrupt instruction generation and sending module 214, it can be connected to the advanced error reporting register monitoring module 222 to send instructions or signals to at least one of the computing device and the host to indicate that the direct connection link with the direct connection port connected by the direct connection port detection module 210 as one end is in an abnormal working state. According to an embodiment of this aspect, the interrupt instruction generation and sending module 214 may be configured to send an interrupt instruction to the interrupt collector of the computing device. For example, the interrupt instruction generation and sending module 214 may be configured to send an interrupt instruction to the interrupt collector of the computing device at least based on the type of error.
[0046] Regarding the interrupt instruction, it may indicate that the direct connection link with the current direct connection port as one end is in an abnormal working state.
[0047] Specifically, according to an embodiment of the present invention, the interrupt instruction generation and sending module 214 may be further configured to: generate an interrupt instruction in response to the error type being the first type; send the interrupt instruction to the interrupt collector of the computing device; and in response to the error type being the second type, count the number of errors of the second type; generate an interrupt instruction in response to the number of errors of the second type exceeding a threshold number; and send the interrupt instruction to the interrupt collector of the computing device to indicate that the direct connection link with the direct connection port at one end is in an abnormal working state. How to generate and send interrupt instructions according to different types of errors will be further described in conjunction with method 300 below, and will not be elaborated here for now.
[0048] It should be understood that according to the inventive concept of the present invention, the monitoring system 200 may include several direct connection port detection modules 210. In some embodiments, the number of direct connection port detection modules 210 in the monitoring system 200 is related to the number of direct connection links between the computing devices to be monitored. For example, when there are two direct connection links between computing devices that need to be monitored, the monitoring system 200 may include, for example, 4 direct connection port detection modules 210, where the ports at both ends of each direct connection link are connected to 1 direct connection port detection module 210. In still other embodiments, the number of direct connection port detection modules 210 in the monitoring system 200 may be determined at least based on the number of computing devices and the number of direct connection ports configured on each computing device. For example, in one example, there are 4 computing devices, and each computing device is configured with 4 direct connection ports, then the monitoring system 200 may include 16 direct connection port detection modules 210. Those skilled in the art can easily determine or change the number of direct connection port detection modules 210 in the monitoring system 200 according to the specific applicable situation, and the present invention does not limit this.
[0049] Figure 3 The flowchart of method 300 for monitoring the direct connection link between computing devices according to an embodiment of the present invention is shown. It should be understood that method 300 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.
[0050] In step 302, the monitoring system 200 reads the information in the advanced error reporting register in the direct connection port of the computing device at at least one end of the direct connection link.
[0051] In some embodiments, it may be read by, for example, Figure 2 the advanced error reporting register monitoring module 212 of the direct connection port detection module 210 of the monitoring system 200 as shown, the information in the advanced error reporting register in the direct connection port connected to this direct connection port detection module 210. As shown above, the advanced error reporting register in the direct connection port may store error information related to this direct connection port.
[0052] In step 304, in response to the read information including error-related report information, the monitoring system 200 determines the type of error based on the error-related report information.
[0053] In some embodiments, in response to the read information including error-related report information, it can be based on the Figure 2 advanced error report register monitoring module 212 of the direct connection port detection module 210 of the monitoring system 200 as shown to determine the type of the occurred error based on this report information. For example, if the read information includes report information related to completion timeout, the type of the occurred error (i.e., completion timeout) can be determined to be non-fatal based on this report information related to completion timeout. In another example, if the read information includes report information related to receiver overflow, the type of the occurred error (i.e., receiver overflow) can be determined to be fatal.
[0054] In step 306, the monitoring system 200 sends an interrupt instruction to the interrupt collector of the computing device at least based on the type of the error, to indicate that the direct connection link with the direct connection port as one end is in an abnormal working state.
[0055] Regarding sending an interrupt instruction to the interrupt collector of the computing device at least based on the type of the error, it can include: the monitoring system 200 generates an interrupt instruction in response to the type of the error being the first type; sends the interrupt instruction to the interrupt collector of the computing device; and in response to the type of the error being the second type, counts the number of errors of the second type; in response to the number of errors of the second type exceeding the threshold number, generates an interrupt instruction; and sends the interrupt instruction to the interrupt collector of the computing device to indicate that the direct connection link with the direct connection port as one end is in an abnormal working state.
[0056] Regarding the first type of error, it can refer to an error with a relatively high severity, for example, a fatal error defined in the PCIe protocol standard.
[0057] Regarding the second type of error, it can refer to an error with a relatively low severity, for example, a non-fatal error defined in the PCIe protocol standard.
[0058] Regarding counting the number of errors of the second type, it can include: in response to detecting the current second type of error within the threshold time, controlling the value of the counter for counting the number of errors of the second type to be incremented by 1; and in response to reaching the threshold time, controlling the counter to be cleared. Therefore, according to the inventive concept of the present invention, the direct connection port detection module included in the monitoring system (such as Figure 2The direct connection port detection module 210 of the monitoring system 200 shown may further include a counter and a timer for performing the operation of counting the number of errors of the aforementioned second type of statistical type. For example, the counter may be configured to count the number of errors of the second type, and the timer may be configured to determine whether a threshold time has been reached.
[0059] According to an embodiment of the present invention, after the monitoring system 200 sends an interrupt instruction to the interrupt collector of the computing device, the computing device may further send an interrupt signal to the host to report a port error indication to the host.
[0060] Regarding the port error indication, it may be used to indicate that there is an error in the current direct connection port of the computing device and that the direct connection link with the current direct connection port as one end is in an abnormal working state. For example, the port error indication may include information related to the port number of the direct connection port with the error to indicate that there is an error in the direct connection port corresponding to the port number.
[0061] For example, in some embodiments, it may be Figure 2 The advanced error reporting register monitoring module 212 of the direct connection port detection module 210 of the monitoring system 200 shown sends an interrupt instruction sending indication to the interrupt instruction generation and sending module 214 of the same direct connection port detection module 210 to indicate that the interrupt instruction generation and sending module 214 sends an interrupt instruction to the interrupt collector of the computing device; in response to the received interrupt instruction sending indication, the interrupt instruction generation and sending module 214 of the direct connection port detection module 210 sends an interrupt instruction to the interrupt collector of the computing device to indicate that the direct connection link with the direct connection port as one end is in an abnormal working state.
[0062] The following will be combined with Figure 4 The system 400 shown and Figure 5 The flowchart of the method 500 shown will be used to describe in detail how to monitor the direct connection link between computing devices according to the solution provided by the present invention.
[0063] Figure 4 A block diagram of an exemplary system 400 according to the present invention is shown. It should be understood that the system 400 may further include additional modules not shown and / or the shown modules may be omitted, and the scope of the present invention is not limited in this regard.
[0064] As Figure 4 shown, the system 400 includes a host (CPU) 410, a PCIe bus 420, a first computing device GPGPU1, and a second computing device GPGPU2. Among them, the host (CPU) 410 can interact with the first computing device GPGPU1 and the second computing device GPGPU2 via the PCIe bus 420.
[0065] Regarding the first computing device GPGPU1, a first direct connection port PIP11 and a second direct connection port PIP12 are configured thereon, and corresponding AER registers are configured on each direct connection port. Specifically, a first AER register 422 is configured on the first direct connection port PIP11, and a second AER register 424 is configured on the second direct connection port PIP12.
[0066] Regarding the second computing device GPGPU2, a third direct connection port PIP21 is configured thereon, and a third AER register 426 is configured on the third direct connection port PIP21.
[0067] As Figure 4 shown, the system 400 further includes a monitoring system 430. Further, the monitoring system 430 includes: a first direct connection port detection module 442, configured to be connected to the first direct connection port PIP11; a second direct connection port detection module 444, configured to be connected to the second direct connection port PIP12; and a third direct connection port detection module 446, configured to be connected to the third direct connection port PIP21.
[0068] Further, as Figure 4 shown, the first direct connection port detection module 442 includes a first advanced error reporting register monitoring module 452 and a first interrupt instruction generation and sending module 462, wherein the first advanced error reporting register monitoring module 452 is connected to the first AER register 422 of the first direct connection port PIP11 to read the information in the first AER register 422. The second direct connection port detection module 444 includes a second advanced error reporting register monitoring module 454 and a second interrupt instruction generation and sending module 464, wherein the second advanced error reporting register monitoring module 454 is connected to the second AER register 424 of the second direct connection port PIP12 to read the information in the second AER register 424. The third direct connection port detection module 446 includes a third advanced error reporting register monitoring module 456 and a third interrupt instruction generation and sending module 466, wherein the third advanced error reporting register monitoring module 456 is connected to the third AER register 426 of the third direct connection port PIP21 to read the information in the third AER register 426.
[0069] It should be understood that the system 400 here is only exemplary, and the present invention does not limit the number of computing devices included in the system 400, the number of direct connection ports configured on each computing device, or the number of direct connection port detection modules in the monitoring system 430.
[0070] According to an embodiment of the present invention, in one example, after PCIe link training, a certain direct connection port (e.g., the first direct connection port PIP11) of the first computing device GPGPU1 in the system 400 can be interconnected with the third direct connection port PIP21 of the second computing device GPGPU2 to form a direct connection link (i.e., Link4) between the first computing device GPGPU1 and the second computing device GPGPU2. Assume that the working mode of the first direct connection port PIP11 of the first computing device GPGPU1 is the EP (End Point) mode, and the working mode of the third direct connection port PIP21 of the second computing device GPGPU2 is the RC (Root Complex) mode. Then, the first computing device GPGPU1 can interact with the second computing device GPGPU2 through the direct connection link Link4. For example, the first computing device GPGPU1 can send a read request to the second computing device GPGPU2. If an error occurs during the interaction through the direct connection link Link4, such as the response related to the aforementioned read operation has not been returned, and finally a completion timeout error occurs, then the first direct connection port PIP11 of the first computing device GPGPU1 will then report the completion timeout error to the third direct connection port PIP21 of the second computing device GPGPU2, and the information related to this completion timeout error will be stored in the third AER register 426 of the third direct connection port PIP21 of the second computing device GPGPU2. The following will be combined with Figure 5 explain how to determine the working state of a direct connection link with a direct connection port as one end based on the information stored in the direct connection port of the computing device.
[0071] Figure 5 FIG. shows a flowchart of a method 500 for monitoring a direct connection link between computing devices according to an embodiment of the present invention. It should be understood that the method 500 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.
[0072] In step 502, the monitoring system reads the information in the AER register of the direct connection port.
[0073] According to an embodiment of the present invention, if you want to read the information in the AER register of a certain direct connection port, the information in the AER register of this direct connection port can be read by the direct connection port detection module connected to this direct connection port in the monitoring system. Specifically, the information in the AER register of the direct connection port can be read by the advanced error reporting register monitoring module in the direct connection port detection module. For example, in the above example, the information stored in the third AER register 426 of the third direct connection port PIP21 can be read by the third advanced error reporting register monitoring module 456 of the third direct connection port detection module 446 of the monitoring system 430.
[0074] In step 504, it is determined by the monitoring system whether the read information includes error-related report information.
[0075] According to an embodiment of the present invention, in response to the advanced error report register monitoring module reading information from the corresponding AER register, it can then be determined by the advanced error report register monitoring module whether the read information includes error-related report information. For example, in the above example, in response to the third advanced error report register monitoring module 456 reading the stored information from the third AER register 426, it is determined by the third advanced error report register monitoring module 456 whether the read information includes error-related report information.
[0076] Furthermore, as described above, since the information related to the error of completion timeout is stored in the third AER register 426 of the third direct connection port PIP21 of the second computing device GPGPU2, in this example, the third advanced error report register monitoring module 456 can determine that the read information includes error-related report information, that is, the information related to the error of completion timeout.
[0077] In response to determining that the read information includes error-related report information, proceed to step 506, and it is determined by the monitoring system whether the type of the error is a fatal error.
[0078] According to an embodiment of the present invention, it can be further determined by the advanced error report register monitoring module whether the type of the error in the read report information is a fatal error. For example, in the above example, the third advanced error report register monitoring module 456 can determine that the completion timeout in the read information belongs to a non-fatal error. In other words, the third advanced error report register monitoring module 456 can determine that the type of the error in the read report information is a non-fatal error. In another example, the read information includes report information related to receiver overflow. Since receiver overflow belongs to a fatal error, the third advanced error report register monitoring module 456 can determine that the type of the error in the read report information is a fatal error.
[0079] As can be seen from the above, if the type of the error in the read report information is a fatal error, that is, an error that can affect data transmission on the direct connection link between computing devices, it can be inferred that the direct connection link is in an abnormal working state, and data transmission using the current direct connection link may be abnormal. On this basis, it is possible to proceed to step 508, and the monitoring system sends an interrupt instruction to the computing device.
[0080] For example, when the third high-level error report register monitoring module 456 determines that the type of error in the read report information is a fatal error, the third high-level error report register monitoring module 456 can send an interrupt instruction sending indication to the third interrupt instruction generation and sending module 466; then, in response to the received interrupt instruction sending indication, the interrupt instruction generation and sending module 466 sends an interrupt instruction to the interrupt collector (not shown) of the second computing device GPGPU2 to indicate that the direct link Link4 with the third direct port PIP21 as one end is in an abnormal working state.
[0081] If at step 506, the monitoring system determines that the type of error is not a fatal error, for example, the type of error is a non-fatal error, then it proceeds to step 510 for the monitoring system to determine whether the time has reached the threshold time. For example, in the above example, if the third high-level error report register monitoring module 456 determines that the error (i.e., completion timeout) in the read information is a non-fatal error, then it further determines whether the time has reached the threshold time.
[0082] Regarding the threshold time, the threshold time can be configured by configuring the monitoring system (such as, the monitoring system includes a timer). According to an embodiment of the present invention, the configuration range of the threshold time can be from 1 hour to 72 hours. For example, the threshold time can be such as 70 hours, 52 hours, 4 hours, etc.
[0083] According to an embodiment of the present invention, the timer configured in the direct port detection module of the monitoring system can be used to determine whether the time has reached the threshold time. For example, in the Figure 4 system shown, for example, if the first direct port detection module 442 determines that the type of error in the report information read from the first AER register 422 of the first direct port PIP11 is a non-fatal error, then the first direct port detection module 442 further determines whether the time has reached the threshold time, that is, determines whether the current non-fatal error has been detected within the threshold time.
[0084] In response to the time reaching the threshold time, it proceeds to step 512, where the monitoring system clears the value of the counter used to count the number of non-fatal errors and returns to step 502 again.
[0085] In response to the time not reaching the threshold time, that is, in response to the current non-fatal error being detected within the threshold time, it proceeds to step 514, where the monitoring system increments the value of the counter used to count the number of non-fatal errors by 1.
[0086] For example, in the above example, if the first direct connection port detection module 442 determines that the threshold time has been reached, the value of the counter used to count the number of non-fatal errors is cleared, and the first direct connection port PIP11 is monitored again. For example, after a predetermined time, the information stored in the first AER register 422 of the first direct connection port PIP11 is read again. If the first direct connection port detection module 442 determines that a current non-fatal error is detected within the threshold time, the value of the counter used to count the number of non-fatal errors is incremented by 1, and the process proceeds to step 516.
[0087] In step 516, it is determined by the monitoring system whether the number of non-fatal errors exceeds a threshold number.
[0088] Regarding the threshold number, it can be configured by configuring the monitoring system. According to an embodiment of the present invention, the value range of the threshold number can be from 1 to 11. For example, the threshold number can be configured to 2.
[0089] According to the inventive concept of the present invention, if the number of non-fatal errors occurring within the threshold time exceeds the threshold number, even though the severity of the non-fatal errors themselves is relatively low and does not affect the PCIe link for data transmission, however, due to the overly frequent occurrence of the errors, it can be inferred that the PCIe link for data transmission is affected by these errors. Therefore, according to an embodiment of the present invention, when the number of non-fatal errors occurring within the threshold time exceeds the threshold number, it is considered that the direct connection link with the monitored direct connection port as one end is in an abnormal working state. In an example, the threshold time is configured to 72 hours and the threshold number is configured to 2. Then, when the number of non-fatal errors occurring within 72 hours exceeds 2, it can be considered that the direct connection link with the monitored direct connection port as one end is in an abnormal working state. Thus, when it is determined in step 516 that the number of non-fatal errors exceeds the threshold number, the process proceeds to step 508, and the monitoring system sends an interrupt instruction to the computing device.
[0090] Correspondingly, if it is determined in step 516 that the number of non-fatal errors does not exceed the threshold number, it can be considered that the PCIe link for data transmission has not been affected by these errors, and the direct connection link with the monitored direct connection port as one end is in a normal working state. Therefore, when it is determined in step 516 that the number of non-fatal errors does not exceed the threshold number, the process returns to step 502 again.
[0091] In addition, if in step 504, the monitoring system determines that the read information does not include error-related report information, the process also returns to step 502, for example, to read the information from the AER register of the direct connection port again after a predetermined time.
[0092] Regarding the predetermined time, it can be configured with the waiting time register. According to an embodiment of the present invention, the length of the predetermined time can be related to the working state of the computing device. For example, when the computing device is in a stable working state, since the requirement for data real-time performance is not high, a longer predetermined time length can be configured. In still other embodiments, when the computing device is in the debugging phase, in order to improve the real-time performance of data, the length of the predetermined time is relatively short.
[0093] Return to step 508. When the monitoring system sends an interrupt instruction to the computing device (for example, the interrupt collector of the computing device), since the interrupt collector of the computing device can report the interrupt to the host (such as the CPU) via the PCIe bus, in response to the computing device receiving the interrupt instruction, it can further proceed to 518, and the computing device sends an interrupt signal to the host.
[0094] Regarding the interrupt signal, it can be used to report a port error indication to the host, where the port error indication can indicate that there is an error in the current direct connection port of the computing device and the direct connection link with the current direct connection port as one end is in an abnormal working state.
[0095] For example, if the interrupt collector (not shown) of the second computing device GPGPU2 receives an interrupt instruction sent by the third direct connection port detection module 446, the second computing device GPGPU2 can then report the interrupt to the CPU 410 via the PCIe bus 420, that is, it realizes reporting the error information detected at the direct connection port of the computing device to the host, so that the monitoring and diagnosis of the direct connection link between computing devices can be realized by monitoring the direct connection ports of the computing device.
[0096] According to an embodiment of the present invention, further, in response to the host (for example, Figure 4 the CPU 400 shown in Figure 4 receiving an interrupt reported by the computing device (for example, the first computing device GPGPU1 or the second computing device GPGPU2 shown in
[0097] ), then the host can start an interrupt handler, such as, to determine whether there are abnormalities in other direct connection ports on the computing device that reported the interrupt. For example, whether there are errors in these direct connection ports, whether the direct connection links with these direct connection ports as one end are in a normal working state, etc. The monitoring for each port can refer to the method described above and will not be elaborated here.For example, in one example, if the first computing device GPGPU1 receives an interrupt indication sent by the first direct connection port detection module 442 and reports the interrupt to the CPU 400 by interrupting the signal based on the received interrupt indication, then in response to the received interrupt signal, the CPU 400 can start an interrupt handler to scan the direct connection port detection module (i.e., the second direct connection port detection module 444) connected to other direct connection ports (i.e., the second direct connection port PIP12) configured on the first computing device GPGPU1 except for the first direct connection port PIP11, to determine whether there is an error in the second direct connection port PIP12, so as to determine whether the direct connection link with the second direct connection port PIP12 as one end is in a normal working state.
[0098] On this basis, if based on the interrupt handler, the host determines that the number of abnormal direct connection ports on a certain computing device is too large, for example, exceeding 50%, then this computing device can be regarded as damaged for subsequent replacement of the computing device. If the number of abnormal direct connection ports on a certain computing device is small, for example, only 1 or 2, then the abnormal direct connection ports on this computing device can be masked by software, such as, and the abnormal direct connection ports can be replaced with other direct connection ports on the computing device, so that data transmission and other interactions can be carried out normally between computing devices.
[0099] In summary, according to the solution provided by the present invention, by introducing the monitoring system according to the embodiments of the present invention into the existing PCIe protocol-based system, it is possible to monitor each direct connection port located at each computing device in real time, so as to monitor the direct connection links based on the PCIe protocol formed between different ports of each computing device, thereby realizing the monitoring of the direct connection links between computing devices. And, according to the solution provided by the present invention, the errors detected at each direct connection port can be reported to the host, so that the host can monitor the working state of each direct connection port of the computing device and diagnose the abnormal direct connection link in time. Even if the system includes a large-scale deployment of computing devices, the host can also monitor, collect data, isolate faults and recover for each direct connection port on each computing device.
[0100] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.
[0101] The above are only alternative embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for monitoring a direct connection link between computing devices, characterized in that, including: A direct port detection module connected to a direct port of a computing device through a direct connection port of a monitoring system reads information in an advanced error reporting register in the direct port, where the computing device is at at least one end of the direct link, the direct link is a direct link based on the PCIe protocol, and the direct port is a PCIe interconnect port of the computing device, and the computing device includes a GPGPU; In response to the information read including error-related report information, determining, by the direct port detection module, a type of the error based on the error-related report information; and Sending, by the direct port detection module, an interrupt instruction to an interrupt collector of the computing device at least based on the type of the error, to indicate that a direct link with the direct port as one end is in an abnormal working state.
2. The method according to claim 1, wherein Sending an interrupt instruction to the interrupt collector of the computing device at least based on the type of the error includes: Generating an interrupt instruction in response to the type of the error being a first type; and Sending the interrupt instruction to the interrupt collector of the computing device.
3. The method according to claim 2, wherein Sending an interrupt instruction to the interrupt collector of the computing device at least based on the type of the error includes: Counting the number of errors of a second type in response to the type of the error being the second type; Generating an interrupt instruction in response to the number of errors of the second type exceeding a threshold number; and Sending the interrupt instruction to the interrupt collector of the computing device to indicate that a direct link with the direct port as one end is in an abnormal working state.
4. The method according to claim 3, wherein Counting the number of errors of the second type includes: Controlling a value of a counter for counting the number of errors of the second type to be incremented by 1 in response to detecting a current error of the second type within a threshold time; and Controlling the counter to be cleared in response to reaching the threshold time.
5. The method according to claim 1, characterized in that further including: In response to the information read not including error-related report information, rereading the information in the advanced error reporting register in the direct port of the computing device after a predetermined time.
6. The method according to claim 1, wherein further including: In response to the interrupt collector of the computing device receiving the interrupt instruction, sending, by the computing device, an interrupt signal to a host to report a port error indication to the host, where the port error indication is used to indicate that there is an error in a current direct port of the computing device and that a direct link with the current direct port as one end is in an abnormal working state.
7. A monitoring system for monitoring a direct connection link between computing devices, characterized in that, including: At least one direct port detection module configured to be connected to a direct port of a computing device, where the direct link is a direct link based on the PCIe protocol, and the direct port is a PCIe interconnect port of the computing device, the computing device includes a GPGPU, and the direct port detection module includes: An advanced error reporting register monitoring module configured to read information in an advanced error reporting register in the direct port of the computing device, and determine a type of the error based on the error-related report information in response to the information read including error-related report information; An interrupt instruction generation and sending module, configured to send the interrupt instruction to an interrupt collector of the computing device at least based on the type of the error, where the interrupt instruction indicates that a direct connection link with the direct connection port as one end is in an abnormal working state.
8. The monitoring system according to claim 7, characterized in that, The type of the error includes a first type and a second type, where the interrupt instruction generation and sending module is further configured to: In response to the type of the error being the first type, generate an interrupt instruction; and send the interrupt instruction to an interrupt collector of the computing device.
9. The monitoring system according to claim 8, characterized in that, The interrupt instruction generation and sending module is further configured to: In response to the type of the error being the second type, count the number of errors of the second type; In response to the number of errors of the second type exceeding a threshold number, generate an interrupt instruction; And Send the interrupt instruction to an interrupt collector of the computing device to indicate that a direct connection link with the direct connection port as one end is in an abnormal working state.
10. The monitoring system according to claim 7, characterized in that, The advanced error report register monitoring module is further configured to: In response to the read information not including error-related report information, re-read the information in the advanced error report register in the direct connection port of the computing device after a predetermined time.
Citation Information
Patent Citations
FPGA based EFM OAM processing method and hardware realization device
CN105897446A