Controller communication system, method, computer product, device and storage medium

By setting multiple pins and connection lines between the controllers, the hardware failure information is directly transmitted, which solves the problem of low service switching efficiency between the controllers, and realizes rapid fault information transmission and service switching, improving the reliability and stability of the system.

CN120353749BActive Publication Date: 2025-08-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510857534.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-08-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In the prior art, the service switching between controllers is inefficient, and it is impossible to timely sense system downtime or restart caused by hardware failures, affecting system reliability.

Method used

By setting multiple pins and connection lines between the controllers, hardware failure information is directly transmitted, and software protocol stack is bypassed to realize the direct connection between the processor and the hardware of complex programmable logic devices, distinguishing between hardware and software failure types, and quickly transmitting fault information.

Benefits of technology

It improves the transmission rate of fault information between controllers, shortens the service switching time, reduces from seconds to milliseconds, and improves the reliability and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353749B_ABST
    Figure CN120353749B_ABST
Patent Text Reader

Abstract

The present application discloses a controller communication system, method, computer product, device, and storage medium, relating to the field of controller communication technology. The controller communication system includes: a first controller provided with a first processor and a first complex programmable logic device; a second controller provided with a second processor and a second complex programmable logic device; the first processor is connected to the first complex programmable logic device via a first connection line, the first complex programmable logic device is connected to the second complex programmable logic device via a second connection line, and the second complex programmable logic device is connected to the second processor via a third connection line, and the complex programmable logic device is provided with multiple pins; fault information transmission between controllers is achieved based on links formed by different target fault information types and their corresponding multiple pins and multiple connection lines. This solves the technical problem of low service switching efficiency between controllers and can improve controller communication efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of controller communication technology, and in particular to a controller communication system, method, computer product, device, and storage medium. Background Art

[0002] With the rapid development of storage systems, they are being rapidly adopted and promoted across various industries, placing higher demands on the stability and reliability of storage systems. In particular, when a single node or controller fails, services must be quickly switched from the failed controller to the backup controller to avoid disruptions to customer services caused by controller failures.

[0003] Related technologies primarily rely on obtaining a controller presence signal to determine whether service handover between controllers is necessary. However, this approach limits its application scenarios because the other controller can only detect the absence of one controller and take over its services when one is unplugged. Furthermore, the presence signal cannot promptly detect system downtime or restarts caused by controller hardware failures, reducing controller system reliability. Summary of the Invention

[0004] The present application provides a controller communication system, method, computer product, device and storage medium to at least solve the technical problem of low service switching efficiency between controllers in the related art.

[0005] The present application provides a controller communication system, comprising: a first controller and a second controller; the first controller comprising a first processor and a first complex programmable logic device; the second controller comprising a second processor and a second complex programmable logic device; the first processor being connected to a first end of the first complex programmable logic device via a first connection line, the second end of the first complex programmable logic device being connected to a first end of the second complex programmable logic device via a second connection line, and the second end of the second complex programmable logic device being connected to the second processor via a third connection line, and a plurality of pins being provided on the first complex programmable logic device and the second complex programmable logic device, respectively; the first processor acquiring target fault information and a target fault information type; in response to the target fault information type being a hardware fault type, the first processor transmitting the target fault information to the second processor via a pin matching the hardware fault type and the second connection line; in response to the target fault information type being a software fault type, the first processor transmitting the target fault information to the second processor via the first connection line, the second connection line, and the third connection line.

[0006] The present application also provides a controller communication method, including: in response to the first processor obtaining target fault information, the first processor parses the target fault information and obtains a target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor via a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor via the first connection line, the second connection line, and the third connection line.

[0007] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the controller communication system in any of the following embodiments. In response to a first processor acquiring target fault information, the first processor parses the target fault information and acquires a target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor via a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor via the first connection line, the second connection line, and the third connection line.

[0008] The present application also provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the steps of the controller communication system in the following embodiments when executing the computer program.

[0009] In response to the first processor obtaining the target fault information, the first processor parses the target fault information and obtains the target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor through the pin matching the hardware fault type and the second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line and the third connection line.

[0010] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the controller communication system in any of the following embodiments are implemented.

[0011] In response to the first processor obtaining the target fault information, the first processor parses the target fault information and obtains the target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor through the pin matching the hardware fault type and the second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line and the third connection line.

[0012] The controller communication system provided by the present application includes: a first controller and a second controller; the first controller includes a first processor and a first complex programmable logic device; the second controller includes a second processor and a second complex programmable logic device; the first processor is connected to the first end of the first complex programmable logic device via a first connection line, the second end of the first complex programmable logic device is connected to the first end of the second complex programmable logic device via a second connection line, and the second end of the second complex programmable logic device is connected to the second processor via a third connection line, and multiple pins are provided on the first complex programmable logic device and the second complex programmable logic device respectively; the first processor obtains target fault information and target fault information type, and in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor via a pin matching the hardware fault type and the second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor via the first connection line, the second connection line, and the third connection line. In this way, the target fault information type is distinguished, and different links formed by multiple pins and multiple connection lines are determined according to the target fault information type obtained by the first processor, thereby realizing fault information transmission between controllers. The technical problem of low service switching efficiency between controllers is solved, and the communication efficiency of controllers can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0014] Figure 1 A schematic diagram of the structure of a controller communication system provided in related technology 1;

[0015] Figure 2 A schematic diagram of the structure of the controller communication system provided by the related technology 2;

[0016] Figure 3A schematic diagram of the structure of a controller communication system provided in an embodiment of the present application;

[0017] Figure 4 A flow chart of a controller communication method provided in an embodiment of the present application;

[0018] Figure 5 A flowchart of a controller communication method provided in another embodiment of the present application;

[0019] Figure 6 A structural block diagram of a controller communication device provided in an embodiment of the present application;

[0020] Figure 7 This is a diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0023] The CPU (central processing unit) emerged in the era of large-scale integrated circuits. Iterative updates to processor architecture design and continuous advancements in integrated circuit technology have driven its continuous development and improvement. Storage devices, on the other hand, require greater stability and high reliability to prevent data loss due to device anomalies. Therefore, during the design and development process, high redundancy is required for multiple components. For example, controllers must be designed with at least two, four, or even more redundant components.

[0024] With the rapid development of storage systems, they are rapidly being used and promoted across various industries. This has led to higher requirements for stability and reliability in storage server design, as well as the implementation of more functions, which poses greater challenges to system stability. In particular, when a single node or controller fails, services can be quickly switched from the failed controller to a backup controller, preventing customer service disruptions caused by controller failures.

[0025] See also Figure 1 In the first related technology, a storage system uses a dual-controller environment, such as controller A and controller B. Under normal circumstances, customer services run on controller A. However, when controller A is unplugged, the customer services will notify controller B of the unplugging of controller A through the NTB channel between the two controllers. After receiving the information that controller A is not in place, controller B begins to take over the services.

[0026] See also Figure 2 In related technology 2, each controller's corresponding presence signal is obtained and transmitted to the CPLD (Complex Programmable Logic Device) of the other controller via the backplane. After receiving the absence signal, the other CPLD directly sends it to the other CPU. If the other CPU receives the absence signal, the NTB (Non-Transparent Bridge) link between the two controllers triggers a LinkDown alarm. When the NTB link is disconnected, controller B receives the LinkDown time. The LinkDown time is estimated to be around 5ms, but the software requires approximately 1.6s for de-jitter processing before reporting. Therefore, the excessive software processing time prevents the fault information from being quickly synchronized to other controllers when a fault occurs.

[0027] It can be seen from this that the related art has the problem that it takes a long time to perform service switching between a faulty controller and other controllers.

[0028] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0029] like Figure 3 As shown, an embodiment of the present application provides a controller communication system, which specifically includes a first controller and a second controller:

[0030] The first controller includes a first processor and a first complex programmable logic device; the second controller includes a second processor and a second complex programmable logic device;

[0031] A controller is a device that controls the starting, speed regulation, braking, and reversing of a motor by changing the wiring of the main or control circuits and the resistance values ​​in a predetermined sequence. Composed of a program counter, instruction register, instruction decoder, timing generator, and operation controller, it is the "decision-making body" that issues commands, effectively coordinating and directing the operations of the entire computer system.

[0032] Both the first and second processors are central processing units (CPUs), serving as the computing and control core of a computer system and the final execution unit for information processing and program execution. Since their inception, CPUs have made significant advancements in logical structure, operational efficiency, and functional extension.

[0033] Both the first and second CPLDs are complex programmable logic devices (CPLDs), digital integrated circuits whose logic functions can be customized by the user. Their basic design approach involves using an integrated development software platform, schematics, hardware description languages, and other methods to generate target files. The code is then transferred to the target chip via a download cable ("in-system" programming), realizing the designed digital system.

[0034] The first processor is connected to a first end of a first complex programmable logic device (CPLD1) via a first connection line, a second end of the first complex programmable logic device (CPLD1) is connected to a first end of a second complex programmable logic device (CPLD2) via a second connection line (ZCS CLK and ZCSSDA lines shown in the figure), and a second end of the second complex programmable logic device (CPLD2) is connected to the second processor via a third connection line. The first complex programmable logic device (CPLD1) and the second complex programmable logic device (CPLD2) are respectively provided with multiple pins.

[0035] The first connection line includes a first sub-connection line, a second sub-connection line, and a third sub-connection line. The first sub-connection line is a fault transmission line, used to transmit target fault information of the software fault type to the first complex programmable logic device. The second sub-connection line is a serial clock line, and the third sub-connection line is a serial data line. The second and third sub-connection lines are used to query fault information. The software fault type here refers to controller software fault information, which can include assert faults, kernel faults, and other fault information. An assert fault typically refers to an error triggered during program execution when an assertion expression is false (i.e., 0). Kernel panic is a critical error state in the Linux operating system, indicating that the processor kernel has encountered an unrecoverable error, causing the system to crash.

[0036] The controller communication system also includes a backplane having a second connection line disposed thereon. A first end of the backplane is connected to a second end of the first complex programmable logic device via a first end of the second connection line, and a second end of the backplane is connected to a first end of the second complex programmable logic device via a second end of the second connection line. The backplane may be a PCB backplane for arranging connection lines between the complex programmable logic devices of different controllers. Using the backplane for arranging lines simplifies wiring and reduces the degree of heat dissipation obstruction from the enclosure.

[0037] The second connection line includes a master connection line and a slave connection line. The master connection line is used to transmit target fault information, and the slave connection line is used to transmit target fault information when the master connection line fails. The master connection line can be the ZCSCLK and ZCS SDA lines shown in the figure, and the slave connection line can be a backup fiber optic channel (not shown). The master connection line is preferentially used for communication between the first complex programmable logic device and the second complex programmable logic device. When the master connection line is unresponsive or fails, communication between the first and second complex programmable logic devices can be switched to the slave connection line. This improves the fault tolerance of data transmission and thereby improves system reliability.

[0038] The third connection line includes a fourth sub-connection line, a fifth sub-connection line and a sixth sub-connection line; the fourth sub-connection line is a fault transmission line, and the fourth sub-connection line is used to transmit target fault information belonging to a software fault type; the fifth sub-connection line is a serial clock line, and the sixth sub-connection line is a serial data line. The fifth sub-connection line and the sixth sub-connection line are used to query fault information.

[0039] The multiple pins include a first pin (PWRGD ALL), a second pin (CPU CATERR / ERR / RST N), and a third pin (PEER PRESET N). The hardware fault types include a first hardware fault type, a second hardware fault type, and a third hardware fault type. The first pin, the second pin, and the third pin match the first hardware fault type, the second hardware fault type, and the third hardware fault type, respectively.

[0040] Pins can be GPIOs (General Purpose Input / Output Ports), which are pins on a chip. As input ports, they can read the pin's state—high or low. As output ports, they can output high or low levels to control connected external devices. GPIO readings can be performed using the system's built-in GPIO driver, eliminating the need for independent development and saving costs.

[0041] The first hardware fault type here is the PWRGD ALL type of fault information, which is also power-related fault information. The second hardware fault type is the CPU CATERR / ERR / RST N type of fault information, which is also processor fatal error / recoverable error / processor reset-related fault information. The third hardware fault type is the PEER PRESET N type of fault information, which is also controller status-related fault information. The hardware fault type in this application is controller hardware fault information. The controller hardware fault information is divided into different subtypes, and different pins are set for different subtypes. The hardware fault information of the subtype corresponding to the pin can be directly transmitted through the pin, which can increase the transmission rate of the hardware fault information.

[0042] This application is not limited to the scenario where the controller is unplugged. It can obtain controller fault information in real time, promptly perceive system crashes or restarts caused by hardware failures, and realize direct hardware connection between the processor and complex programmable logic devices through pins. It does not use software methods such as protocols to perceive hardware fault information. By bypassing the software protocol stack and directly triggering the hardware, it avoids the delay caused by software response, and can reduce the service switching time from seconds to milliseconds, which can increase the fault information transmission rate.

[0043] In one embodiment, a first processor obtains target fault information and a target fault information type. In response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor via a pin matching the hardware fault type and a second connection line. In response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor via the first connection line, the second connection line, and the third connection line.

[0044] The first complex programmable logic device and the second complex programmable logic device are used to control the level states of multiple pins; wherein, in response to the first complex programmable logic device receiving target fault information transmitted by the second processor, the first complex programmable logic device pulls up the level of the pin on the first complex programmable logic device that matches the target fault information type to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the first processor; in response to the second complex programmable logic device receiving the target fault information transmitted by the first processor, the second complex programmable logic device pulls up the level of the pin on the second complex programmable logic device that matches the target fault information type to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor.

[0045] Specifically, the target fault information is the controller fault information currently received by the first processor. After the first processor obtains the target fault information, it can obtain the target fault information type, that is, first determine whether the target fault information belongs to the hardware fault type (controller hardware fault information) or the software fault type (controller software fault information).

[0046] When the target fault information is of a hardware fault type, a target pin matching the target hardware fault type corresponding to the target fault information of the hardware fault type is obtained. For example, assuming the target fault information is power supply-related fault information (the first hardware fault type in the target fault information of the hardware fault type), a first pin matching the first hardware fault type is selected from multiple pins. The first processor transmits the target fault information to the first complex programmable logic device via the first pin on the first complex programmable logic device. In response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device via a second connection line provided on the backplane. In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device pulls up the voltage level of the first pin on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor. In response to the second processor receiving the target fault interrupt signal, the second processor takes over the services of the first processor.

[0047] When the target fault information is of a software fault type, a target software fault type corresponding to the target fault information of the software fault type is further determined, where the target software fault type corresponding to the target fault information of the software fault type is at least one of a first software fault type (assert fault) and a second software fault type (kernel fault). In response to the target software fault type being determined, the first processor transmits the target fault information to the first complex programmable logic device via a first sub-connection line in the first connection line. In response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device via a second connection line provided on the backplane. In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device raises the level of a pin corresponding to a third connection line on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor. In response to the second processor receiving the target fault interrupt signal, the second processor takes over the service of the first processor.

[0048] In this application, the controller software fault information and the controller hardware fault information are distinguished. Through the combination of multiple pins and the second connection line, a direct hardware connection between the processor in the controller and the complex programmable logic is achieved to avoid software stack delay, thereby achieving rapid switching of services between controllers when the controller hardware fails. Through the combination of the first connection line, the second connection line and the third connection line, the service switching between controllers when the controller software fails is achieved, which can solve the technical problem of low efficiency of service switching between controllers in related technologies.

[0049] like Figure 4 As shown, an embodiment of the present application provides a controller communication method, which specifically includes the following steps:

[0050] Step 101: In response to the first processor acquiring target fault information, the first processor parses the target fault information and acquires a target fault information type corresponding to the target fault information.

[0051] Step 102: In response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor via a pin matching the hardware fault type and a second connection line.

[0052] Specifically, the specific process of step 102 is shown in steps 1021 to 1026.

[0053] Step 1021: Determine a target hardware fault type corresponding to the target fault information of the hardware fault type, where the target hardware fault type is at least one of the first hardware fault type, the second hardware fault type, and the third hardware fault type.

[0054] Step 1022: Acquire a target pin that matches the target hardware fault type, where the target pin is at least one of the first pin, the second pin, and the third pin.

[0055] Step 1023 : In response to the target hardware fault type and the target pin being determined, the first processor transmits target fault information to the first complex programmable logic device through the target pin on the first complex programmable logic device.

[0056] Step 1024: In response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device through a second connection line provided on the backplane.

[0057] Step 1025: In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device pulls up the level of the target pin on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor.

[0058] Step 1026: In response to the second processor receiving the target fault interrupt signal, the second processor takes over the service of the first processor.

[0059] Step 103: In response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor via the first connection line, the second connection line, and the third connection line.

[0060] Specifically, when the target software fault type of the target fault information type of the software fault type is the first software fault type, the specific process of step 103 is shown as steps 1031 to 1035 .

[0061] Step 1031: Determine a target software fault type corresponding to target fault information of the software fault type, where the target software fault type is at least one of a first software fault type and a second software fault type.

[0062] Step 1032: In response to the target software fault type being the first software fault type, the first processor transmits target fault information of the first software fault type to the first complex programmable logic device via the first sub-connection line in the first connection line.

[0063] Step 1033: In response to the first complex programmable logic device receiving the target fault information of the first software fault type, the first complex programmable logic device transmits the target fault information of the first software fault type to the second complex programmable logic device through a second connection line provided on the backplane.

[0064] Step 1034: In response to the second complex programmable logic device receiving the target fault information of the first software fault type, the second complex programmable logic device pulls up the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor.

[0065] Step 1035: In response to the second processor receiving the target fault interrupt signal, the second processor takes over the service of the first processor.

[0066] See also Figure 5 When the target software fault type of the target fault information type of the software fault type is the second software fault type, the specific process of step 103 is shown as steps S1031 to S1037.

[0067] Step S1031: In response to the target software fault type corresponding to the target fault information of the software fault type being the second software fault type, the first processor shuts down or exits the service software corresponding to the target fault information of the second software fault type.

[0068] Step S1032: The first processor transmits target fault information of the second software fault type to the first complex programmable logic device through the first sub-connection line in the first connection line.

[0069] Step S1033: In response to the first complex programmable logic device receiving the target fault information of the second software fault type, the first complex programmable logic device broadcasts the target fault information of the second software fault type to the second complex programmable logic device through a second connection line provided on the backplane.

[0070] Step S1034: In response to the second complex programmable logic device receiving the target fault information of the second software fault type, the second complex programmable logic device raises the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device through the second complex programmable logic device to generate a target fault interrupt signal, and sends the target fault interrupt signal and the target fault information of the second software fault type to the communication module, which is communicatively connected to the second processor core.

[0071] Step S1035: In response to the communication module receiving the target fault information of the second software fault type, the communication module obtains the first controller fault status, generates a fault notification based on the first controller fault status, and adds the fault notification to the event priority adjustment queue.

[0072] Step S1036: the event priority adjustment queue sends a link disconnection instruction, and the link disconnection instruction is executed to interrupt the service of the first controller.

[0073] Step S1037: In response to the interruption of the first controller service, the second processor takes over the first controller service.

[0074] Specifically, in the kernel failure scenario (the second software failure type), the service software cannot be broadcast through the MCS (broadband) cluster. The interruption time is usually greater than 15 seconds, which will cause service switching delays.

[0075] In the present application, when the first processor obtains the target fault information and determines that the target fault information is a second software fault type fault, the service software of the first controller fault is closed or exited, and the fault controller is broadcast through the second connection line that a controller core fault has occurred. The fault controller triggers a fault signal to enable other controllers to receive the fault information (triggering the fault signal to enable other nodes to receive the fault information). In response to the second complex programmable logic device receiving the second software fault type, the second complex programmable logic device pulls up the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device through the second complex programmable logic device to generate a target fault interrupt signal, and sends the target fault interrupt signal to the communication module. block, and at the same time sends the second software fault type to the communication module, writes the second software fault type into the fault register, the communication module is connected to the second processor core, the communication module obtains the fault status of the first controller based on the kernel fault information, and generates a fault notification based on the fault status of the first controller, and adds the fault notification to the event priority adjustment queue. The event priority adjustment queue sends a link disconnection instruction, and executes the link disconnection instruction to interrupt the service of the first controller. After the second controller perceives that the service software processing link of the first controller is disconnected, it reads the kernel fault information broadcast through the second connection line and starts the service takeover process. It is expected that the interruption time can be shortened to less than 5s, which greatly reduces the controller switching time and shortens the service interruption time.

[0076] In one embodiment, the present application further includes the following steps before transmitting the target fault information to the second processor: in response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device obtains the fault query interface corresponding to the fifth sub-connection line and the sixth sub-connection line, and integrates the obtained target fault information to obtain integrated fault information; the second complex programmable logic device sends the integrated fault information to the second processor through the fault query interface. Through the integrated processing of the fault information, the efficiency of the second processor in taking over the business of the first controller can be improved.

[0077] In one embodiment, a complex programmable logic device can obtain the input and output interface status information of the controller in real time. After the complex programmable logic device receives the fault information on the local end, it means that the local controller may have a fault. The local complex programmable logic device can then send the input and output interface status information to the complex programmable logic device on the opposite controller for storage, so that the opposite controller saves relevant logs of the local controller, such as the input and output interface status information, so that the cause of the fault of the local controller and other contents can be traced back. In this way, the efficiency of controller fault recovery can be improved.

[0078] In this application, the example of the first controller sending fault information to the second controller is used for description. In actual applications, the second controller may also send fault information to the first controller. It can be understood that the principle of the second controller sending fault information to the first controller is the same as the principle of the first controller sending fault information to the second controller, which will not be repeated here.

[0079] The embodiment of the present application provides a controller communication device, the controller communication device is specifically as follows Figure 6 As shown, the controller communication device includes: an acquisition module 20 and a transmission module 21.

[0080] The acquisition module 20 is configured to, in response to the first processor acquiring the target fault information, parse the target fault information and acquire a target fault information type corresponding to the target fault information.

[0081] Transmission module 21 is used to, in response to the target fault information type being a hardware fault type, transmit the target fault information to the second processor through the pin matching the hardware fault type and the second connection line; in response to the target fault information type being a software fault type, transmit the target fault information to the second processor through the first connection line, the second connection line and the third connection line.

[0082] For the description of the features in the embodiment corresponding to the controller communication device, please refer to the relevant description of the embodiment corresponding to the controller communication system, which will not be repeated here.

[0083] The embodiment of the present application also provides a computer device, such as Figure 7 As shown, it includes a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above controller communication method embodiments.

[0084] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above controller communication method embodiments when running.

[0085] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0086] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0087] The above is a detailed introduction to a controller communication method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A controller communication system, characterized in that: including a first controller, a second controller and a backplane; The first controller includes a first processor and a first complex programmable logic device; The second controller includes a second processor and a second complex programmable logic device; A second connecting circuit is provided on the back panel; The first processor is connected to a first end of a first complex programmable logic device via a first connection line, a second end of the first complex programmable logic device is connected to a first end of the backplane via a first end of a second connection line, a second end of the backplane is connected to a first end of the second complex programmable logic device via a second end of a second connection line, and a second end of the second complex programmable logic device is connected to the second processor via a third connection line, and a plurality of pins are provided on each of the first complex programmable logic device and the second complex programmable logic device; The first connection line includes a first sub-connection line, a second sub-connection line, and a third sub-connection line; the first sub-connection line is a fault transmission line, and the first sub-connection line is used to transmit target fault information belonging to a software fault type; the second sub-connection line is a serial clock line, and the third sub-connection line is a serial data line, and the second sub-connection line and the third sub-connection line are used to query fault information; The third connection circuit includes a fourth sub-connection circuit, a fifth sub-connection circuit and a sixth sub-connection circuit; The fourth sub-connection line is a fault transmission line, and the fourth sub-connection line is used to transmit target fault information belonging to a software fault type; the fifth sub-connection line is a serial clock line, and the sixth sub-connection line is a serial data line, and the fifth sub-connection line and the sixth sub-connection line are used to query fault information; The first processor acquires target fault information and a target fault information type. In response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor via a pin matching the hardware fault type and a second connection line. In response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line, and the third connection line.

2. The controller communication system according to claim 1, wherein: The multiple pins include a first pin, a second pin, and a third pin. The hardware fault types include a first hardware fault type, a second hardware fault type, and a third hardware fault type. The first pin, the second pin, and the third pin match the first hardware fault type, the second hardware fault type, and the third hardware fault type, respectively.

3. The controller communication system according to claim 1, wherein: The first complex programmable logic device and the second complex programmable logic device are used to control the level states of multiple pins; In response to the first complex programmable logic device receiving the target fault information transmitted by the second processor, the first complex programmable logic device pulls up the level of a pin on the first complex programmable logic device that matches the type of the target fault information to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the first processor; In response to the second complex programmable logic device receiving the target fault information transmitted by the first processor, the second complex programmable logic device pulls up the level of a pin on the second complex programmable logic device that matches the target fault information type to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor.

4. The controller communication system according to claim 1, wherein: The second connection line includes a main connection line and a slave connection line. The main connection line is used to transmit target fault information, and the slave connection line is used to transmit target fault information when the main connection line fails.

5. A controller communication method, based on the controller communication system according to any one of claims 1 to 4, characterized in that: The controller communication method includes: In response to the first processor acquiring the target fault information, the first processor parses the target fault information to acquire a target fault information type corresponding to the target fault information; In response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor via a pin matching the hardware fault type and a second connection line; In response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line, and the third connection line.

6. The controller communication method according to claim 5, characterized in that: In response to the target fault information type being a hardware fault type, the first processor transmitting the target fault information to the second processor through a pin matching the hardware fault type and a second connection line includes: Determine a target hardware fault type corresponding to the target fault information of the hardware fault type, where the target hardware fault type is at least one of the first hardware fault type, the second hardware fault type, and the third hardware fault type; Acquire a target pin that matches a target hardware fault type, where the target pin is at least one of a first pin, a second pin, and a third pin; In response to the target hardware fault type and the target pin being determined, the first processor transmits target fault information to the first complex programmable logic device via a target pin on the first complex programmable logic device; In response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device through a second connection line provided on the backplane; In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device pulls up the level of the target pin on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor; In response to the second processor receiving the target fault interrupt signal, the second processor takes over the service of the first processor.

7. The controller communication method according to claim 5, characterized in that: In response to the target fault information type being a software fault type, the first processor transmitting the target fault information to the second processor through the first connection line, the second connection line, and the third connection line includes: Determine a target software fault type corresponding to target fault information of the software fault type, where the target software fault type is at least one of a first software fault type and a second software fault type; In response to the target software fault type being the first software fault type, the first processor transmits target fault information of the first software fault type to the first complex programmable logic device through a first sub-connection line in the first connection line; In response to the first complex programmable logic device receiving the target fault information of the first software fault type, the first complex programmable logic device transmits the target fault information of the first software fault type to the second complex programmable logic device through a second connection line provided on the backplane; In response to the second complex programmable logic device receiving target fault information of the first software fault type, the second complex programmable logic device pulls up the level of a pin corresponding to a fourth sub-connection line on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor; In response to the second processor receiving the target fault interrupt signal, the second processor takes over the service of the first processor.

8. The controller communication method according to claim 5, characterized in that: In response to the target fault information type being a software fault type, the first processor transmitting the target fault information to the second processor through the first connection line, the second connection line, and the third connection line further includes: In response to the target software fault type corresponding to the target fault information of the software fault type being a second software fault type, the first processor shuts down or exits the service software corresponding to the target fault information of the second software fault type; The first processor transmits target fault information of the second software fault type to the first complex programmable logic device through the first sub-connection line in the first connection line; In response to the first complex programmable logic device receiving the target fault information of the second software fault type, the first complex programmable logic device broadcasts the target fault information of the second software fault type to the second complex programmable logic device through a second connection line provided on the backplane; In response to the second complex programmable logic device receiving the target fault information of the second software fault type, the second complex programmable logic device pulls up the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device through the second complex programmable logic device to generate a target fault interrupt signal, and sends the target fault interrupt signal and the target fault information of the second software fault type to the communication module, wherein the communication module is communicatively connected to the second processor core; In response to the communication module receiving the target fault information of the second software fault type, the communication module obtains a first controller fault status, generates a fault notification based on the first controller fault status, and adds the fault notification to an event priority adjustment queue; The event priority adjustment queue sends a link disconnection instruction, and executes the link disconnection instruction to interrupt the service of the first controller; In response to the first controller service being interrupted, the second processor takes over the first controller service.

9. The controller communication method according to claim 5, characterized in that: Before transmitting the target fault information to the second processor, the method further includes: In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device obtains the fault query interface corresponding to the fifth sub-connection line and the sixth sub-connection line, and integrates the obtained target fault information to obtain integrated fault information; The second complex programmable logic device sends the integrated fault information to the second processor through the fault query interface.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 5 to 9 are implemented.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 5 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 5 to 9 are implemented.

Citation Information

Patent Citations

  • Baseboard management controller fault detection device

    CN114691408A

  • Switch reset system and method, storage medium and electronic equipment

    CN115550291A