Server interconnection exception handling system, method, device and storage medium

By detecting the status of the server's mechanical connection interface, generating interrupt signals and fault logs, the problem of high failure rate caused by loose cable connections during server interconnection is solved, achieving efficient fault location and reducing operation and maintenance costs.

CN115114068BActive Publication Date: 2026-07-21INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2022-06-29
Publication Date
2026-07-21

Smart Images

  • Figure CN115114068B_ABST
    Figure CN115114068B_ABST
Patent Text Reader

Abstract

The application relates to a server interconnection exception processing system, method, equipment and storage medium, the system comprising a connection module, a detection module and a fault log generation module; the connection module connects a first server and a second server; the connection module comprises an interconnection bus and an interconnection interface; the interconnection interface comprises a mechanical connection interface; the detection module is connected to the connection module and is used for detecting the connection state of the mechanical connection interface; when the connection state of the mechanical connection interface is disconnected, a corresponding interruption signal is generated and sent to the fault log generation module; the fault log generation module is connected to the detection module and generates an alarm event in response to the interruption signal; the downlink equipment information of the interconnection interface corresponding to the interruption signal is acquired, and a fault log is generated according to the downlink equipment information. The application generates a unified fault log to avoid a large number of fault work orders of downlink equipment and reduce operation and maintenance costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a server interconnection anomaly handling system, method, device and storage medium. Background Technology

[0002] Currently, traditional rack-mount servers can reduce space usage, but they are still limited by space constraints, making it difficult to simultaneously fit a high-performance CPU (Central Processing Unit) and a high-performance GPU (Graphics Processing Unit) within a limited space while meeting heat dissipation requirements. Using cables to externally interconnect CPU and GPU servers expands the functionality of traditional rack-mount servers, meeting the needs of certain scenarios and applications. The interconnect bus can use PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard). Appropriate PCIe bandwidth is selected based on actual business needs. To ensure high-performance PCIe signal transmission, a PCIe Repeater adapter card is installed at the CPU server expansion port to convert the interface from PCIe or OCP connectors to MiniSAS connectors, while also providing relay functionality for long-distance PCIe signal transmission. Using cables to externally interconnect CPU and GPU servers is a common expansion solution.

[0003] However, the presence of external interconnect cables necessitates on-site cable installation during system racking and relocation. Due to varying levels of product familiarity among operators, this leads to varying probabilities of loose connections. While the CPU and GPU servers' respective BMC management units detect cable connectivity and communicate with each other, existing CPU and GPU server connection detection only checks the electrical continuity of the cable, which cannot guarantee reliable detection. Furthermore, loose external cable connections can cause downlink device connection failures, resulting in numerous work orders for downlink devices. This leads to a high failure rate and provides only fault descriptions without fault location, requiring maintenance personnel to troubleshoot each issue individually, which is time-consuming and labor-intensive. Summary of the Invention

[0004] Based on this, this application provides a server interconnection anomaly handling system, method, device, and storage medium to solve the problems existing in the prior art.

[0005] Firstly, a server interconnection anomaly handling system is provided, which includes: a connection module, a detection module, and a fault log generation module;

[0006] The connection module connects the first server and the second server; the connection module includes an interconnect bus and an interconnect interface; the interconnect interface includes a mechanical connection interface;

[0007] The detection module is connected to the connection module and is used to detect the connection status of the mechanical connection interface. When the connection status of the mechanical connection interface is disconnected, a corresponding interrupt signal is generated and the interrupt signal is sent to the fault log generation module.

[0008] The fault log generation module is connected to the detection module, generates an alarm event in response to the interrupt signal, obtains downlink device information of the interconnection interface corresponding to the interrupt signal, and generates a fault log based on the downlink device information.

[0009] According to one possible implementation method in an embodiment of this application, the system further includes: an alarm module;

[0010] The alarm module responds to the alarm event and issues an alarm based on the audible and visual alarm circuit.

[0011] According to one achievable method in an embodiment of this application, the detection module is further configured to:

[0012] When the connection status of the mechanical connection interface is disconnected, obtain the location information of the abnormal interconnection interface corresponding to the disconnected mechanical connection interface.

[0013] According to one achievable method in an embodiment of this application, the fault log generation module is further configured to:

[0014] Obtain the topology of the interconnect bus, obtain the downlink device information corresponding to the abnormal interconnect interface based on the location information of the abnormal interconnect interface and the topology of the interconnect bus, and generate a fault log based on the downlink device information.

[0015] Secondly, a method for handling server interconnection anomalies is provided, the method comprising:

[0016] Obtain the mechanical connection interface status of the interconnection interface between the first server and the second server;

[0017] When the connection state of the mechanical connection interface is disconnected, a corresponding interrupt signal is generated;

[0018] An alarm event is generated based on the interrupt signal;

[0019] A fault log is generated based on the downlink device information of the interconnect interface corresponding to the interrupt signal.

[0020] According to one possible implementation method in an embodiment of this application, the method further includes:

[0021] In response to the alarm event, an alarm is triggered based on the audible and visual alarm circuit.

[0022] According to one achievable method in an embodiment of this application, when the connection state of the mechanical connection interface is disconnected, generating a corresponding interrupt signal includes:

[0023] When the connection state of the mechanical connection interface is disconnected, the location information of the abnormal interconnection interface corresponding to the disconnected mechanical connection interface is obtained, and an interrupt signal corresponding to the abnormal interconnection interface is generated.

[0024] According to one achievable method in an embodiment of this application, generating a fault log based on the downlink device information of the interconnect interface corresponding to the interrupt signal includes:

[0025] Obtain the topology of the server interconnection bus, obtain the downlink device information corresponding to the abnormal interconnection interface based on the location information of the abnormal interconnection interface and the topology of the interconnection bus, and generate a fault log based on the downlink device information.

[0026] Thirdly, a computer device is provided, comprising:

[0027] At least one processor; and

[0028] A memory communicatively connected to the at least one processor; wherein,

[0029] The memory stores computer instructions that can be executed by the at least one processor to enable the at least one processor to perform the method involved in the first aspect above.

[0030] Fourthly, a computer-readable storage medium is provided, having stored thereon computer instructions, wherein the computer instructions are used to cause a computer to perform the methods involved in the first aspect above.

[0031] According to the technical content provided in the embodiments of this application, this application detects the connection status of the mechanical connection interface between servers. When the mechanical connection interface is disconnected, an interrupt signal is generated and the downlink device information corresponding to the disconnected interconnection interface is obtained, generating a unified fault log. By setting up the mechanical connection interface, it is convenient to predict, detect, and handle the disconnection of the interconnection interface. At the same time, by generating a unified fault log, it avoids the generation of a large number of fault work orders for downlink devices, reducing operation and maintenance costs. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of the server interconnection exception handling system in one embodiment;

[0033] Figure 2 This is another structural diagram of the server interconnection exception handling system in one embodiment;

[0034] Figure 3 This is a schematic diagram of the detection circuit structure of a server interconnection anomaly handling system in one embodiment;

[0035] Figure 4 This is a flowchart illustrating a server interconnection exception handling method in one embodiment;

[0036] Figure 5 This is a schematic structural diagram of a computer device in one embodiment. Detailed Implementation

[0037] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present application and are not intended to limit the scope of the present application.

[0038] Figure 1 This is a schematic diagram of a hard disk temperature testing system provided in an embodiment of this application. The following description first refers to... Figure 1 This application will be described in detail.

[0039] like Figure 1 As shown, this application provides a hard disk temperature testing system 100, which includes: a connection module 110, a detection module 120 and a fault log generation module 130;

[0040] The connection module 110 connects the first server and the second server; the connection module 110 includes an interconnect bus and an interconnect interface; the interconnect interface includes a mechanical connection interface.

[0041] Specifically, the connection module 110 is used to connect a first server and a second server. The first server may include at least one server, and the second server may include at least one server. The servers may include a CPU server and a GPU server. For example, the first server is a CPU server, and the second server is a GPU server. The connection module 110 connects the CPU server and the GPU server. The connection module 110 includes an interconnect bus and an interconnect interface; the interconnect interface includes a mechanical connection interface. The interconnect bus is a cable used to transmit signals between the two servers, for example, a PCIe interconnect bus; the interconnect interface is used to connect the cable, providing both an electrical connection for communication and a mechanical connection for fixation. Figure 2As shown, 111 is the interconnect bus, and 112 is the interconnect interface. Both the CPU server and GPU server sides connect cables 111 via interconnect interface 112. Interconnect interface 112 includes a mechanical connection interface, enabling mechanical connection for fixation. When cable 111 is inserted into the connector, the electrical connection is completed first, followed by the mechanical connection; the mechanical connection occurs later than the electrical connection. When cable 111 is removed from the connector, the mechanical connection is disconnected first, followed by the electrical connection; the mechanical connection is disconnected earlier than the electrical connection. This allows the mechanical connection to be disconnected before the electrical connection is disconnected at the interconnect interface, facilitating the predictive detection and handling of electrical interface disconnections.

[0042] The detection module 120 is connected to the connection module 110 and is used to detect the connection status of the mechanical connection interface. When the connection status of the mechanical connection interface is disconnected, a corresponding interrupt signal is generated and the interrupt signal is sent to the fault log generation module 130.

[0043] Specifically, such as Figure 1 As shown, and in combination Figure 2 The detection module 120 is connected to the connection module 110 and is used to detect the connection status of the mechanical connection interface of the interconnect interface 112 in the connection module 110. When the mechanical connection interface of the interconnect interface 112 in the connection module 110 is disconnected, the mechanical connection interface of the interconnect interface 112 is interrupted, and the mechanical connection interfaces of the interconnect interfaces 112 at both ends of the cable 111 can be interrupted. Based on the interruption generated by the mechanical connection interface, the detection module 120 generates an interrupt signal corresponding to the specific interface position and sends the interrupt signal to the fault log generation module 130.

[0044] The fault log generation module 130 is connected to the detection module 120, and generates an alarm event in response to an interrupt signal; it obtains the downlink device information of the interconnection interface corresponding to the interrupt signal, and generates a fault log based on the downlink device information.

[0045] Specifically, such as Figure 1 As shown, the fault log generation module 130 is connected to the detection module 120, receives the interrupt signal sent by the detection module 120, and generates an alarm event. For example, as... Figure 2 As shown, the fault log generation module 130 consists of the motherboard BMC and PCH. After receiving an interrupt signal, the fault log generation module 130 obtains the location information of the interconnect interface 112 that generated the interrupt corresponding to the interrupt signal. At the same time, the BMC obtains the connection relationship between the system downlink devices and the corresponding interconnect interfaces through the PCH. Based on the location information of the interconnect interface 112 that generated the interrupt signal, the BMC obtains the information of the downlink device corresponding to the interconnect interface 112 that generated the interrupt. Based on the information of the downlink device, a unified fault log is generated and displayed uniformly by the detection system.

[0046] It is worth noting that since the downlink devices are all connected via the corresponding interconnect interface 112, when the corresponding interconnect interface 112 is interrupted, the corresponding downlink devices will all fail, generating a large number of fault work orders. In this embodiment, the fault log generation module 130 obtains the downlink device information of the interconnect interface corresponding to the interrupt signal and generates a unified fault log based on the downlink device information, thereby avoiding the simultaneous generation of a large number of work orders and reducing maintenance costs.

[0047] According to the technical content provided in the embodiments of this application, this application detects the connection status of the mechanical connection interface between servers. When the mechanical connection interface is disconnected, an interrupt signal is generated and the downlink device information corresponding to the disconnected interconnection interface is obtained, generating a unified fault log. By setting up the mechanical connection interface, it is convenient to predict, detect, and handle the disconnection of the interconnection interface. At the same time, by generating a unified fault log, it avoids the generation of a large number of fault work orders for downlink devices, reducing operation and maintenance costs.

[0048] In one embodiment of this application, the hard disk temperature testing system provided in the above embodiment further includes: an alarm module 140; the alarm module 140 responds to an alarm event and alarms based on an audible and visual alarm circuit.

[0049] Specifically, such as Figure 1 As shown, and in combination Figure 2 The alarm module 140 responds to alarm events sent by the fault log generation module 130 by triggering an alarm based on an audible and visual alarm circuit, such as by using an LED indicator or by using an audible alarm. For example, as... Figure 2 As shown, the BMC triggers an alarm event based on an interrupt signal and sends it to the alarm module 140 via the I2C bus. On-site audible and visual alarm indications facilitate the location of abnormal interconnect interfaces for maintenance personnel.

[0050] In one embodiment of this application, the detection module 120 is further configured to: when the connection state of the mechanical connection interface is disconnected, obtain the location information of the abnormal interconnection interface corresponding to the disconnected mechanical connection interface.

[0051] Specifically, when the mechanical connection interface of the interconnect interface 112 in the connection module 110 is disconnected, the interconnect interface 112 generates an abnormal interruption. For example... Figure 2 As shown, since the connection module 110 connecting the two servers includes several cables 111 and corresponding interconnect interfaces 112, when the connection status of a certain mechanical connection interface is disconnected, it is necessary to obtain the location information of the abnormal interconnect interface 112 corresponding to the disconnected mechanical connection interface. This facilitates maintenance personnel in locating the abnormal interconnect interface and also facilitates obtaining the connection relationship of the downlink device corresponding to that interface based on the location information.

[0052] In one embodiment of this application, the fault log generation module 130 is further configured to: obtain the topology of the interconnect bus, obtain downlink device information corresponding to the abnormal interconnect interface based on the location information of the abnormal interconnect interface and the topology of the interconnect bus, and generate a fault log based on the downlink device information.

[0053] Specifically, after receiving an interrupt signal, the fault log generation module 130 can obtain the location information of the abnormal interconnect interface 112 that caused the interrupt. Simultaneously, the fault log generation module 130 further obtains the topology of the interconnect bus 111, which includes the connection relationships of all downlink devices connected to each interconnect interface 112. Based on this topology and the location information of the abnormal interconnect interface 112, the downlink device information corresponding to the abnormal interconnect interface 112 can be obtained, and a unified fault log can be generated based on the downlink device information.

[0054] Based on the above embodiments, in one specific embodiment of this application, the detection module 120 includes a mechanical connectivity detection circuit. For example... Figure 3 As shown, the mechanical continuity detection circuit consists of a TVS diode, a pull-up resistor R, and an inverse Schmitt trigger. The TVS diode is placed close to the connector signal pin for electrostatic discharge protection. The pull-up resistor R ensures a defined signal level. The inverse Schmitt trigger eliminates signal jitter and increases drive capability. The outputs of the inverse Schmitt triggers are connected together and output as an interrupt signal to the main board. When the mechanical connection is disconnected, the input of the inverse Schmitt trigger is pulled up to VCC by the resistor R, and the interrupt signal at the output is low. When the mechanical connection is connected, the input of the inverse Schmitt trigger is connected to GND, and the interrupt signal at the output is high. The state of the mechanical connection can be determined based on the level of the interrupt signal. Changes in the mechanical connection status of the interfaces at both ends of the cable will trigger changes in the interrupt signal. The rising edge of the pin (PIN1) on the interconnect interface indicates that the mechanical connection of the interface on the CPU server side has gone from being connected to being disconnected, and the rising edge of the pin (PIN2) on the interconnect interface indicates that the mechanical connection of the interface on the GPU server side has gone from being connected to being disconnected. The falling edge of the interrupt signal indicates that at least one of the connectors at both ends of the cable has gone from being connected to being disconnected.

[0055] According to the technical content provided in the embodiments of this application, this application detects the connection status of the mechanical connection interface between servers. When the mechanical connection interface is disconnected, an interrupt signal is generated and the downlink device information corresponding to the disconnected interconnection interface is obtained, generating a unified fault log. Setting up a mechanical connection interface facilitates the predictive detection and handling of interconnection interface disconnections; simultaneously, generating a unified fault log avoids the generation of numerous fault work orders for downlink devices, reducing maintenance costs; and on-site audible and visual alarm indications facilitate maintenance personnel in locating the abnormal interconnection interface.

[0056] Figure 4 A flowchart of a server interconnection anomaly handling method provided in this application embodiment is shown below. Figure 4 As shown, the method may include the following steps:

[0057] Step 101: Obtain the mechanical connection interface status of the interconnection interface between the first server and the second server.

[0058] Specifically, the first server may include at least one server, and the second server may include at least one server. The servers may include a CPU server and a GPU server. For example, the first server may be a CPU server, and the second server may be a GPU server. The connection module 110 connects the CPU server and the GPU server. The first server and the second server are connected via an interconnection interface, which includes a mechanical connection interface. The interconnection interface is used to connect cables. It provides both an electrical connection for communication and a mechanical connection for fixation. When a cable is inserted into the connector, the electrical connection is completed first, followed by the mechanical connection; when the cable is removed from the connector, the mechanical connection is disconnected first, followed by the electrical connection; the mechanical connection is disconnected earlier than the electrical connection. This allows the mechanical connection to be disconnected before the electrical connection is disconnected at the interconnection interface, facilitating the predictive detection and handling of electrical interface disconnections.

[0059] Step 102: When the mechanical connection interface is disconnected, generate the corresponding interrupt signal.

[0060] Specifically, when the mechanical connection interface of the interconnection interface is disconnected, such as by force or accident, an interrupt signal is generated based on the interruption caused by the mechanical connection interface. This interrupt signal contains the specific location information of the abnormal interface.

[0061] Step 103: Generate an alarm event based on the interrupt signal.

[0062] Specifically, alarm events are generated based on interrupt signals to issue alarms.

[0063] Step 104: Generate a fault log based on the downlink device information of the interconnect interface corresponding to the interrupt signal.

[0064] Specifically, after receiving an interrupt signal, the location information of the interconnect interface that caused the interrupt can be obtained, and the connection relationship between the system downlink device and the corresponding interconnect interface can be obtained. Based on the location information of the interconnect interface that caused the interrupt, the information of the downlink device corresponding to the interconnect interface that caused the interrupt can be obtained. Based on the information of the downlink device, a unified fault log is generated and the fault log is uniformly displayed by the detection system.

[0065] According to the technical content provided in the embodiments of this application, this application detects the connection status of the mechanical connection interface between servers. When the mechanical connection interface is disconnected, an interrupt signal is generated and the downlink device information corresponding to the disconnected interconnection interface is obtained, generating a unified fault log. By setting up the mechanical connection interface, it is convenient to predict, detect, and handle the disconnection of the interconnection interface. At the same time, by generating a unified fault log, it avoids the generation of a large number of fault work orders for downlink devices, reducing operation and maintenance costs.

[0066] In one embodiment of this application, the server interconnection anomaly handling method further includes: responding to an alarm event by triggering an alarm based on an audible and visual alarm circuit.

[0067] Specifically, in response to the alarm event generated in step 103, an alarm is triggered based on the audible and visual alarm circuit, such as by an LED indicator or by an audible alarm. On-site audible and visual alarm indications facilitate the location of abnormal interconnect interfaces for maintenance personnel.

[0068] In one embodiment of this application, when the connection state of the mechanical connection interface is disconnected, a corresponding interrupt signal is generated, including: when the connection state of the mechanical connection interface is disconnected, obtaining the location information of the abnormal interconnection interface corresponding to the disconnected mechanical connection interface, and generating an interrupt signal corresponding to the abnormal interconnection interface.

[0069] Specifically, when the mechanical connection interface is disconnected, the interconnect interface experiences an abnormal interruption. Since the connection module connecting the two servers includes several cables and corresponding interconnect interfaces, when the connection status of a certain mechanical connection interface is disconnected, it is necessary to obtain the location information of the abnormal interconnect interface corresponding to the disconnected mechanical connection interface. This facilitates maintenance personnel in locating the abnormal interconnect interface and also facilitates obtaining the connection relationship of the downlink device corresponding to that interface based on the location information.

[0070] In one embodiment of this application, generating a fault log based on the downlink device information of the interconnection interface corresponding to the interrupt signal includes: obtaining the topology of the server interconnection bus, obtaining the downlink device information corresponding to the abnormal interconnection interface based on the location information of the abnormal interconnection interface and the topology of the interconnection bus, and generating a fault log based on the downlink device information.

[0071] Specifically, in response to the interrupt signal generated in step 102, after receiving the interrupt signal, the location information of the abnormal interconnect interface that caused the interrupt can be obtained. Simultaneously, the topology of the interconnect bus is further obtained, which includes the connection relationships of all downlink devices connected to each interconnect interface. Based on this topology and the location information of the abnormal interconnect interface, the downlink device information corresponding to the abnormal interconnect interface can be obtained, and a unified fault log is generated based on the downlink device information.

[0072] According to the technical content provided in the embodiments of this application, this application detects the connection status of the mechanical connection interface between servers. When the mechanical connection interface is disconnected, an interrupt signal is generated and the downlink device information corresponding to the disconnected interconnection interface is obtained, generating a unified fault log. Setting up a mechanical connection interface facilitates the predictive detection and handling of interconnection interface disconnections; simultaneously, generating a unified fault log avoids the generation of numerous fault work orders for downlink devices, reducing maintenance costs; and on-site audible and visual alarm indications facilitate maintenance personnel in locating the abnormal interconnection interface.

[0073] The embodiments of this application primarily address the problem of downlink device failures and a large number of work orders caused by the disconnection of the interconnection interface between servers by detecting the mechanical connectivity between servers. It can be used on motherboards, as well as on RISER cards, ReTime cards, redrive cards, node motherboards, etc. The connector can be a MiniSAS connector, and the cable can be a corresponding MiniSAS cable. The detection circuit must have a protective circuit near the connector. PCIe links with the same port can share a single interrupt signal to save pin resources. The audible and visual alarm circuit is placed near the corresponding connector and is visible to the outside, and can use a buzzer and LED indicator lights for alarm.

[0074] It should be understood that, although Figure 4 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this application, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Furthermore, Figure 4At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0075] The same or similar parts among the above embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.

[0076] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., explicit consent from the user, actual notification to the user, explicit authorization from the user, etc.).

[0077] According to embodiments of this application, this application also provides a computer device and a computer-readable storage medium. This application further provides a computer device including at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores computer instructions executable by the at least one processor, the computer instructions being executed by the at least one processor to enable the at least one processor to perform the server interconnection exception handling method described in any of the above embodiments.

[0078] like Figure 5 The diagram shown is a block diagram of a computer device according to an embodiment of this application. The term "computer device" is intended to represent various forms of digital computers or mobile devices. The digital computer may include a desktop computer, a portable computer, a workbench, a personal digital assistant, a server, a mainframe computer, and other suitable computers. The mobile device may include a tablet computer, a smartphone, a wearable device, etc.

[0079] like Figure 5 As shown, the computer device 500 includes a computing unit 501, a ROM 502, a RAM 503, a bus 504, and an input / output (I / O) interface 505. The computing unit 501, ROM 502, and RAM 503 are interconnected via the bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0080] The computing unit 501 can execute various processes in the method embodiments of this application according to computer instructions stored in the read-only memory (ROM) 502 or computer instructions loaded from the storage unit 508 into the random access memory (RAM) 503. The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. The computing unit 501 can include, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. In some embodiments, the methods provided in the embodiments of this application can be implemented as computer software programs, which are tangibly contained in a computer-readable storage medium, such as the storage unit 508.

[0081] RAM 503 can also store various programs and data required for the operation of device 500. Part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509.

[0082] The input unit 506, output unit 507, storage unit 508, and communication unit 509 in computer device 500 can be connected to I / O interface 505. The input unit 506 can be, for example, a keyboard, mouse, touchscreen, or microphone; the output unit 507 can be, for example, a monitor, speaker, or indicator light. Device 500 can exchange information and data with other devices through the communication unit 509.

[0083] It should be noted that the device may also include other components necessary for normal operation. It may also include only the components necessary for implementing the solution of this application, without necessarily including all the components shown in the figures.

[0084] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), payload programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof.

[0085] The computer instructions used to implement the methods of this application may be written in any combination of one or more programming languages. These computer instructions may be provided to the computing unit 501 such that when executed by the computing unit 501, such as a processor, the computer instructions cause the execution of the steps involved in the embodiments of the methods of this application.

[0086] This application also provides a computer-readable storage medium storing computer instructions thereon, the computer instructions being used to cause a computer to execute the server interconnection exception handling method described in any of the above embodiments.

[0087] The computer-readable storage medium provided in this application can be a tangible medium that can contain or store computer instructions for performing the steps involved in the method embodiments of this application. The computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, and other forms of storage media.

[0088] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A server interconnection anomaly handling system, characterized in that, The system includes: a connection module, a detection module, and a fault log generation module; The connection module connects the first server and the second server; the connection module includes an interconnect bus and an interconnect interface; the interconnect interface includes a mechanical connection interface, which realizes both electrical connection for communication and mechanical connection for fixation; when the cable is inserted into the connector, the electrical connection is completed first, and then the mechanical connection is completed; when the cable is pulled out of the connector, the mechanical connection is disconnected first, and then the electrical connection is disconnected. The detection module is connected to the connection module and is used to detect the connection status of the mechanical connection interface. When the connection status of the mechanical connection interface is disconnected, a corresponding interrupt signal is generated and the interrupt signal is sent to the fault log generation module. The interrupt signal is generated by the mechanical connection of the interconnect interface, and the mechanical connection of the interconnect interface at both ends of the cable can generate an interrupt signal. The fault log generation module is connected to the detection module, and generates an alarm event in response to the interrupt signal; it obtains downlink device information of the interconnection interface corresponding to the interrupt signal, and generates a fault log based on the downlink device information; The detection module is further used for: When the connection status of the mechanical connection interface is disconnected, obtain the location information of the abnormal interconnection interface corresponding to the disconnected mechanical connection interface; The fault log generation module is further used for: Obtain the topology of the interconnect bus, obtain the downlink device information corresponding to the abnormal interconnect interface based on the location information of the abnormal interconnect interface and the topology of the interconnect bus, and generate a fault log based on the downlink device information.

2. The server interconnection anomaly handling system according to claim 1, characterized in that, The system also includes: an alarm module; The alarm module responds to the alarm event and issues an alarm based on the audible and visual alarm circuit.

3. A method for handling server interconnection anomalies based on the server interconnection anomaly handling system of claim 1, characterized in that, The method includes: Obtain the mechanical connection interface status of the interconnection interface between the first server and the second server; When the connection state of the mechanical connection interface is disconnected, a corresponding interrupt signal is generated; An alarm event is generated based on the interrupt signal; A fault log is generated based on the downlink device information of the interconnect interface corresponding to the interrupt signal.

4. The server interconnection anomaly handling method according to claim 3, characterized in that, The method further includes: In response to the alarm event, an alarm is triggered based on the audible and visual alarm circuit.

5. The server interconnection anomaly handling method according to claim 3, characterized in that, When the connection state of the mechanical connection interface is disconnected, a corresponding interrupt signal is generated, including: When the connection state of the mechanical connection interface is disconnected, the location information of the abnormal interconnection interface corresponding to the disconnected mechanical connection interface is obtained, and an interrupt signal corresponding to the abnormal interconnection interface is generated.

6. The server interconnection anomaly handling method according to claim 5, characterized in that, The step of generating a fault log based on the downlink device information of the interconnect interface corresponding to the interrupt signal includes: Obtain the topology of the server interconnection bus, obtain the downlink device information corresponding to the abnormal interconnection interface based on the location information of the abnormal interconnection interface and the topology of the interconnection bus, and generate a fault log based on the downlink device information.

7. A computer device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores computer instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 3 to 6.

8. A computer-readable storage medium storing computer instructions thereon, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 3 to 6.