Controller communication system, method, computer product, device and storage medium
By setting multiple pins and connection lines between the controllers, distinguishing hardware and software fault types, and directly transmitting fault information with hardware, the problem of low service switching efficiency of the controller is solved, rapid fault perception and service switching are achieved, and the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510857534.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In the prior art, the service switching between controllers is inefficient, and the system downtime or restart caused by hardware failure is not possible in time, which reduces the reliability of the system.
By setting multiple pins and connection lines between the controllers, we distinguish hardware failure and software failure types, and directly transmit fault information through direct hardware connection, avoiding software protocol stack delay and achieving fast service switching.
It improves the transmission rate of fault information between controllers, shortens the service switching time, reduces from seconds to milliseconds, and improves the stability and reliability of the system.
Smart Images

Figure CN120353749A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of controller communication, and in particular, to a controller communication system, method, computer product, device, and storage medium. Background Art
[0002] With the rapid development of storage systems, storage systems have been rapidly used and promoted in all walks of life, and higher requirements are placed on the stability and reliability of storage systems. Especially when a single node or controller fails, it is necessary for the service to be quickly switched from the faulty controller to the standby controller to avoid abnormal customer services caused by controller failures.
[0003] In related technologies, it is mainly to determine whether to perform service switching between controllers by obtaining the presence signal of the controller. In this case, only when one end controller is unplugged can the other end controller recognize that one end controller is not present, and then take over the service of that end controller, and the application scenario is limited. And it is impossible to timely sense system downtime or restart caused by hardware failures of the controller through the presence signal. Therefore, the reliability of the controller system will be reduced. Summary of the Invention
[0004] This application provides a controller communication system, method, computer product, device, and storage medium to at least solve the technical problem of low service switching efficiency between controllers in related technologies.
[0005] This application provides a controller communication system, including: a first controller and a second controller; the first controller includes a first processor and a first complex programmable logic device; the second controller includes a second processor and a second complex programmable logic device; the first processor is connected to the first end of the first complex programmable logic device through a first connection line, the second end of the first complex programmable logic device is connected to the first end of the second complex programmable logic device through a second connection line, and the second end of the second complex programmable logic device is connected to the second processor through a third connection line. A plurality of pins are provided on the first complex programmable logic device and the second complex programmable logic device respectively; the first processor obtains target fault information and the type of the target fault information. In response to the type of the target fault information being a hardware fault type, the first processor transmits the target fault information to the second processor through a pin matching the hardware fault type and the second connection line; in response to the type of the target fault information being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line, and the third connection line.
[0006] The present application also provides a controller communication method, including: in response to a first processor obtaining target fault information, the first processor parses the target fault information to obtain a target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to a second processor through a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through a first connection line, a second connection line, and a third connection line.
[0007] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the controller communication system in any of the following embodiments. In response to a first processor obtaining target fault information, the first processor parses the target fault information to obtain a target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to a second processor through a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through a first connection line, a second connection line, and a third connection line.
[0008] The present application also provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor is used to implement the steps of the controller communication system in the following embodiments when executing the computer program.
[0009] In response to a first processor obtaining target fault information, the first processor parses the target fault information to obtain a target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to a second processor through a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through a first connection line, a second connection line, and a third connection line.
[0010] The present application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program, when executed by a processor, implements the steps of the controller communication system in any of the following embodiments.
[0011] In response to the first processor obtaining target fault information, the first processor parses the target fault information to obtain the target fault information type corresponding to the target fault information; in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor through a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through a first connection line, a second connection line, and a third connection line.
[0012] The controller communication system provided in this application includes: a first controller and a second controller; the first controller includes a first processor and a first complex programmable logic device; the second controller includes a second processor and a second complex programmable logic device; the first processor is connected to the first end of the first complex programmable logic device through a first connection line, the second end of the first complex programmable logic device is connected to the first end of the second complex programmable logic device through a second connection line, and the second end of the second complex programmable logic device is connected to the second processor through a third connection line. A plurality of pins are respectively arranged on the first complex programmable logic device and the second complex programmable logic device; the first processor obtains target fault information and the target fault information type. In response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor through a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through a first connection line, a second connection line, and a third connection line. In this way, the target fault information type is distinguished, and different links formed by a plurality of pins and a plurality of connection lines are determined according to the target fault information type obtained by the first processor, so as to realize the transmission of fault information between controllers. The technical problem of low service switching efficiency between controllers is solved, and the communication efficiency of the controllers can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 It is a schematic structural diagram of a controller communication system provided in Related Art 1; Figure 2 It is a schematic structural diagram of a controller communication system provided in Related Art 2; Figure 3 It is a schematic structural diagram of a controller communication system provided in an embodiment of the present application; Figure 4 A schematic flowchart of the controller communication method provided by an embodiment of the present application; Figure 5 A schematic flowchart of the controller communication method provided by another embodiment of the present application; Figure 6 A structural block diagram of the controller communication device provided by an embodiment of the present application; Figure 7 An internal structure diagram of the computer device provided by an embodiment of the present application. Detailed implementation manners
[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0016] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0017] The CPU (Central Processing Unit) emerged in the era of large-scale integrated circuits. The iterative update of the processor architecture design and the continuous improvement of the integrated circuit process have promoted its continuous development and improvement. For storage devices, more requirements are placed on device stability and high reliability to avoid data loss caused by device anomalies. Therefore, in the design and development process, high redundancy design is required for multiple components, etc. For example, the controller needs to be redundantly designed with at least 2, 4 or more controllers.
[0018] With the rapid development of the storage system, the storage system has been rapidly used and promoted in all walks of life. In the design of storage servers, higher requirements are placed on stability and reliability, and more functions are realized, posing higher challenges to the stability of the system. Especially when a single node or controller fails, the service can be quickly switched from the faulty controller to the standby controller to avoid abnormal customer services caused by controller failures.
[0019] Please refer to Figure 1, in Related Art One, the storage system uses a dual - controller environment, such as Controller A and Controller B. Under normal circumstances, the customer's business runs on Controller A. However, when Controller A is unplugged, the customer's business will inform Controller B through the NTB channel between the dual - controllers that Controller A has been unplugged. After receiving information such as Controller A being out - of - place, Controller B starts to take over the business.
[0020] Please refer to Figure 2 , in Related Art Two, the in - place signal corresponding to each controller is obtained, and the in - place signal corresponding to each controller is transmitted to the CPLD (Complex Programmable Logic Device) of the peer controller through the backplane. After the peer CPLD receives the out - of - place information, it directly sends it to the peer CPU. If the peer CPU receives that the peer controller is out - of - place, the NTB (Non - Transparent Bridge) link between the dual - controllers will trigger Linkdown (alarm). When the NTB link is disconnected, Controller B will receive the LinkDown time of the link. The LinkDown time perception is expected to be about 5 ms, but the software needs about 1.6 s of anti - jitter processing before reporting. Therefore, the software processing time is too long, which will cause the failure information not to be quickly synchronized to other controllers when a failure occurs.
[0021] It can be seen that there is a problem in the related art that the time for the faulty controller to switch the business with other controllers is long.
[0022] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0023] As Figure 3 shown, an embodiment of the present application provides a controller communication system, which specifically includes a first controller and a second controller: The first controller includes a first processor and a first complex programmable logic device; the second controller includes a second processor and a second complex programmable logic device; A controller is a master device that changes the wiring of the main circuit or control circuit and changes the resistance value in the circuit in a predetermined order to control the starting, speed regulation, braking, and reverse of the motor. It consists of a program counter, an instruction register, an instruction decoder, a timing generator, and an operation controller. It is the "decision - making body" that issues commands, that is, it completes the coordination and command of the entire computer system operation.
[0024] Both the first processor and the second processor are central processing units (Central Processing Unit, abbreviated as CPU). As the operation and control core of the computer system, they are the final execution units for information processing and program operation. Since the emergence of the CPU, great developments have been made in its logical structure, operation efficiency, and function extension.
[0025] Both the first complex programmable logic device and the second complex programmable logic device are complex programmable logic devices (CPLDs), which are digital integrated circuits that allow users to construct logic functions according to their respective needs. The basic design method is to generate the corresponding target files by means of an integrated development software platform, using schematic diagrams, hardware description languages, etc., and then transfer the code to the target chip through a download cable ("in-system" programming) to implement the designed digital system.
[0026] The first processor is connected to the first end of the first complex programmable logic device (CPLD1) through the first connection line. The second end of the first complex programmable logic device (CPLD1) is connected to the first end of the second complex programmable logic device (CPLD2) through the second connection line (the illustrated ZCS CLK and ZCSSDA lines). The second end of the second complex programmable logic device (CPLD2) is connected to the second processor through the third connection line. The first complex programmable logic device (CPLD1) and the second complex programmable logic device (CPLD2) are respectively provided with multiple pins.
[0027] Among them, the first connection line includes a first sub-connection line, a second sub-connection line, and a third sub-connection line; the first sub-connection line is a fault transmission line, and the first sub-connection line is used to transmit target fault information belonging to the software fault type to the first complex programmable logic device; the second sub-connection line is a serial clock line, and the third sub-connection line is a serial data line. The second sub-connection line and the third sub-connection line are used to query fault information. Here, the software fault type refers to the controller software fault information, and the controller software fault information can include fault information such as assert faults and kernel faults. An assert fault usually refers to an error triggered when the assertion expression is false (i.e., 0) during the program execution. Kernel panic is a serious error state in the Linux operating system, which means that the kernel of the processor has encountered an irrecoverable error, resulting in the system crash.
[0028] The controller communication system further includes a backplane. The second connection line is provided on the backplane. The first end of the backplane is connected to the second end of the first complex programmable logic device through the first end of the second connection line, and the second end of the backplane is connected to the first end of the second complex programmable logic device through the second end of the second connection line. The backplane can be a PCB backplane, which is used to arrange the connection lines between the complex programmable logic devices of different controllers. Using the backplane to arrange the lines can make the wiring more concise and can reduce the degree of heat emission obstruction in the cabinet.
[0029] The second connection line includes a main connection line and a slave connection line, the main connection line is used to transmit target fault information, and the slave connection line is used to transmit target fault information when the main connection line fails. The main connection line can be the ZCSCLK and ZCS SDA lines shown in the figure, and the slave connection line can be a backup optical fiber channel (not shown in the figure). The main connection line is preferentially used to realize the communication between the first complex programmable logic device and the second complex programmable logic device. When the main connection line has no response or the main connection line fails, the slave connection line can be switched to execute the communication between the first complex programmable logic device and the second complex programmable logic device. In this way, the fault tolerance rate of data transmission can be improved, thereby improving the reliability of the system.
[0030] The third connection line includes a fourth sub-connection line, a fifth sub-connection line and a sixth sub-connection line; the fourth sub-connection line is a fault transmission line, and the fourth sub-connection line is used to transmit target fault information belonging to a software fault type; the fifth sub-connection line is a serial clock line, and the sixth sub-connection line is a serial data line. The fifth sub-connection line and the sixth sub-connection line are used to query fault information.
[0031] The multiple pins include a first pin (PWRGD ALL), a second pin (CPU CATERR / ERR / RST N), and a third pin (PEER PRESET N). The hardware fault type includes a first hardware fault type, a second hardware fault type, and a third hardware fault type. The first pin, the second pin, and the third pin match the first hardware fault type, the second hardware fault type, and the third hardware fault type, respectively.
[0032] Pins can be GPIO (General Purpose Input / Output Port), which are pins of the chip. When used as input ports, the pin status can be read through them - high or low level. When used as output ports, we can use them to output high or low level to control the connected external devices. The GPIO reading method can be executed by the GPIO driver that comes with the system bottom layer, without the need for independent development, which can save costs.
[0033] The first type of hardware failure here is the failure information of the PWRGD ALL type, that is, the power supply-related failure information. The second type of hardware failure is the failure information of the CPU CATERR / ERR / RST N type, that is, the failure information related to the processor fatal error / recoverable error / processor reset. The third type of hardware failure is the failure information of the PEER PRESET N type, that is, the failure information related to the controller status. The hardware failure type in this application is the controller hardware failure information. The controller hardware failure information is divided into different subtypes, and different pins are set for different subtypes. The hardware failure information corresponding to the pins can be directly transmitted through the pins, which can improve the transmission rate of the hardware failure information.
[0034] This application is not limited to the scenario of controller removal. It can obtain controller failure information in real time, sense system downtime or restart caused by hardware failures in a timely manner, and achieve direct hardware connection between the processor and the complex programmable logic device through pins. It does not use software methods such as protocols to sense hardware failure information. By bypassing the software protocol stack and directly triggering hardware, it can avoid the delay caused by software responses, improve the service switching time from seconds to milliseconds, and improve the transmission rate of failure information.
[0035] In one embodiment, the first processor obtains the target failure information and the target failure information type. In response to the target failure information type being a hardware failure type, the first processor transmits the target failure information to the second processor through the pin matching the hardware failure type and the second connection line. In response to the target failure information type being a software failure type, the first processor transmits the target failure information to the second processor through the first connection line, the second connection line, and the third connection line.
[0036] The first complex programmable logic device and the second complex programmable logic device are used to control the level states of multiple pins. Among them, in response to the first complex programmable logic device receiving the target failure information transmitted by the second processor, the first complex programmable logic device raises the level of the pin on the first complex programmable logic device that matches the target failure information type to generate a target failure interrupt signal, and transmits the target failure interrupt signal to the first processor. In response to the second complex programmable logic device receiving the target failure information transmitted by the first processor, the second complex programmable logic device raises the level of the pin on the second complex programmable logic device that matches the target failure information type to generate a target failure interrupt signal, and transmits the target failure interrupt signal to the second processor.
[0037] Specifically, the target fault information is the controller fault information received by the current first processor. After the first processor obtains the target fault information, it can obtain the type of the target fault information, that is, first determine whether the target fault information belongs to the hardware fault type (controller hardware fault information) or the software fault type (controller software fault information).
[0038] When the target fault information is of the hardware fault type, obtain the target pin that matches the target hardware fault type corresponding to the target fault information of the hardware fault type. For example, assume that the target fault information is power-related fault information (the first hardware fault type in the target fault information of the hardware fault type), then select the first pin that matches the first hardware fault type from multiple pins. The first processor transmits the target fault information to the first complex programmable logic device through the first pin on the first complex programmable logic device. In response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device through the second connection line provided on the backplane. In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device raises the level of the first pin on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor; in response to the second processor receiving the target fault interrupt signal, the second processor takes over the business of the first processor.
[0039] When the target fault information is of the software fault type, further determine the target software fault type corresponding to the target fault information of the software fault type. The target software fault type corresponding to the target fault information of the software fault type is at least one of the first software fault type (assert fault) and the second software fault type (kernel fault); in response to the determination of the target software fault type being completed, the first processor transmits the target fault information to the first complex programmable logic device through the first sub-connection line in the first connection line; in response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device through the second connection line provided on the backplane; in response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device raises the level of the pin corresponding to the third connection line on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor; in response to the second processor receiving the target fault interrupt signal, the second processor takes over the business of the first processor.
[0040] In this application, controller software fault information and controller hardware fault information are distinguished. Through the combination of multiple pins and the second connection line, direct connection between the processor in the controller and the complex programmable logic device hardware is achieved, avoiding software stack latency. Furthermore, when a controller hardware fault occurs, rapid service switching between controllers is realized. Through the combination of the first connection line, the second connection line, and the third connection line, service switching between controllers is achieved when a controller software fault occurs, which can solve the technical problem of low service switching efficiency between controllers in the related art.
[0041] As Figure 4 shown, an embodiment of this application provides a controller communication method, and this method specifically includes the following steps: Step 101: In response to the first processor obtaining target fault information, the first processor parses the target fault information to obtain the target fault information type corresponding to the target fault information.
[0042] Step 102: In response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor through pins matching the hardware fault type and the second connection line.
[0043] Specifically, the specific process of step 102 is as shown in steps 1021 - 1026.
[0044] Step 1021: Determine the target hardware fault type corresponding to the target fault information of the hardware fault type, where the target hardware fault type is at least one of the first hardware fault type, the second hardware fault type, and the third hardware fault type.
[0045] Step 1022: Obtain target pins matching the target hardware fault type, where the target pins are at least one of the first pin, the second pin, and the third pin.
[0046] Step 1023: In response to the determination of the target hardware fault type and the target pins being completed, the first processor transmits the target fault information to the first complex programmable logic device through the target pins on the first complex programmable logic device.
[0047] Step 1024: In response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device through the second connection line provided on the backplane.
[0048] Step 1025: In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device raises the level of the target pin on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor.
[0049] Step 1026: In response to the second processor receiving the target fault interrupt signal, the second processor takes over the services of the first processor.
[0050] Step 103: In response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line, and the third connection line.
[0051] Specifically, when the target software fault type of the target fault information type of the software fault type is the first software fault type, the specific process of Step 103 is as shown in Steps 1031 - 1035.
[0052] Step 1031: Determine the target software fault type corresponding to the target fault information of the software fault type, where the target software fault type is at least one of the first software fault type and the second software fault type.
[0053] Step 1032: In response to the target software fault type being the first software fault type, the first processor transmits the target fault information of the first software fault type to the first complex programmable logic device through the first sub - connection line in the first connection line.
[0054] Step 1033: In response to the first complex programmable logic device receiving the target fault information of the first software fault type, the first complex programmable logic device transmits the target fault information of the first software fault type to the second complex programmable logic device through the second connection line set on the backplane.
[0055] Step 1034: In response to the second complex programmable logic device receiving the target fault information of the first software fault type, the second complex programmable logic device raises the level of the pin corresponding to the fourth sub - connection line on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor.
[0056] Step 1035: In response to the second processor receiving the target fault interrupt signal, the second processor takes over the services of the first processor.
[0057] Please refer to Figure 5 , when the target software fault type of the target fault information type of the software fault type is the second software fault type, the specific process of Step 103 is as shown in Steps S1031 - S1037.
[0058] Step S1031: In response to the target software fault type corresponding to the target fault information of the software fault type being the second software fault type, the first processor shuts down or exits the service software corresponding to the target fault information of the second software fault type.
[0059] Step S1032: The first processor transmits the target fault information of the second software fault type to the first complex programmable logic device through the first sub-connection line in the first connection line.
[0060] Step S1033: In response to the first complex programmable logic device receiving the target fault information of the second software fault type, the first complex programmable logic device broadcasts the target fault information of the second software fault type to the second complex programmable logic device through the second connection line set on the backplane.
[0061] Step S1034: In response to the second complex programmable logic device receiving the target fault information of the second software fault type, the second complex programmable logic device raises the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device to generate a target fault interrupt signal, and sends the target fault interrupt signal and the target fault information of the second software fault type to the communication module, which is communicatively connected to the second processor core.
[0062] Step S1035: In response to the communication module receiving the target fault information of the second software fault type, the communication module acquires the first controller fault status, generates a fault notification based on the first controller fault status, and adds the fault notification to the event priority adjustment queue.
[0063] Step S1036: The event priority adjustment queue sends a link disconnection command, and executes the link disconnection command to interrupt the first controller service.
[0064] Step S1037: In response to the interruption of the first controller service, the second processor takes over the first controller service.
[0065] Specifically, in the scenario of kernel fault (the second software fault type), the service software cannot broadcast through the MCS (broadband) cluster. Usually, the disconnection time is greater than 15s, which will cause service switching delay.
[0066] In this application, when the first processor obtains the target fault information and determines that the target fault information is a fault of the second software fault type, the business software of the first controller fault is shut down or exited, and the controller kernel fault of the fault controller is broadcast through the second connection line. The fault controller triggers a fault signal so that other controllers can receive the fault information (trigger the fault signal to enable other nodes to receive the fault information). In response to the second complex programmable logic device receiving the second software fault type, the second complex programmable logic device pulls up the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device to generate a target fault interrupt signal, and sends the target fault interrupt signal to the communication module. At the same time, the second software fault type is sent to the communication module and written into the fault register. The communication module is communicatively connected to the second processor core. The communication module obtains the first controller fault status based on the kernel fault information, generates a fault notification based on the first controller fault status, adds the fault notification to the event priority adjustment queue, and the event priority adjustment queue sends a link disconnection instruction. Executing the link disconnection instruction causes the business of the first controller to be interrupted. After the second controller senses that the business software processing link of the first controller is disconnected and reads the kernel fault information broadcast through the second connection line, it starts the service takeover process. It is expected that the disconnection time can be shortened to within 5 s, greatly reducing the controller switching time and shortening the service disconnection time.
[0067] In one embodiment, before setting the target fault information to be transmitted to the second processor in this application, it further includes: in response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device obtains the fault query interfaces corresponding to the fifth sub-connection line and the sixth sub-connection line, and integrates the obtained target fault information to obtain the integrated fault information; the second complex programmable logic device sends the integrated fault information to the second processor through the fault query interface. Through the integrated processing of the fault information, the efficiency of the second processor taking over the business of the first controller can be improved.
[0068] In one embodiment, the complex programmable logic device can obtain the input / output interface status information of the controller in real time. After the complex programmable logic device receives the fault information of its own end, it means that there may be a fault in the controller of its own end. Then the complex programmable logic device of its own end can send the input / output interface status information to the complex programmable logic device on the opposite controller for storage, so that the relevant logs of the controller of its own end, such as the input / output interface status information, are saved in the opposite controller, so that the cause of the fault of the controller of its own end and other contents can be traced back. In this way, the efficiency of controller fault recovery can be improved.
[0069] In this application, the example of the first controller sending a fault message to the second controller is described. In practical applications, it can also be the second controller sending a fault message to the first controller. It can be understood that the principle of the second controller sending a fault message to the first controller is the same as that of the first controller sending a fault message to the second controller, and thus will not be elaborated here.
[0070] An embodiment of this application provides a controller communication device. The controller communication device is specifically as Figure 6 shown. The controller communication device includes: an acquisition module 20 and a transmission module 21.
[0071] The acquisition module 20 is configured to, in response to the first processor obtaining target fault information, the first processor parses the target fault information to obtain the target fault information type corresponding to the target fault information.
[0072] The transmission module 21 is configured to, in response to the target fault information type being a hardware fault type, the first processor transmits the target fault information to the second processor through a pin matching the hardware fault type and a second connection line; in response to the target fault information type being a software fault type, the first processor transmits the target fault information to the second processor through a first connection line, a second connection line, and a third connection line.
[0073] For the description of the features in the corresponding embodiment of the controller communication device, reference can be made to the relevant description of the corresponding embodiment of the controller communication system, which will not be elaborated one by one here.
[0074] An embodiment of this application also provides a computer device, as Figure 7 shown, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the controller communication method.
[0075] An embodiment of this application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the steps in any of the above-mentioned embodiments of the controller communication method when running.
[0076] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs, and other various media that can store computer programs.
[0077] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0078] The above has introduced in detail a controller communication method provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can also be made to this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A controller communication system, characterized in that, It includes a first controller and a second controller; The first controller includes a first processor and a first complex programmable logic device; The second controller includes a second processor and a second complex programmable logic device; The first processor is connected to the first end of the first complex programmable logic device through a first connection line. The second end of the first complex programmable logic device is connected to the first end of the second complex programmable logic device through a second connection line. The second end of the second complex programmable logic device is connected to the second processor through a third connection line. A plurality of pins are provided on the first complex programmable logic device and the second complex programmable logic device respectively; The first processor obtains target fault information and the type of the target fault information. In response to the type of the target fault information being a hardware fault type, the first processor transmits the target fault information to the second processor through a pin matching the hardware fault type and the second connection line; In response to the type of the target fault information being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line and the third connection line.
2. The controller communication system according to claim 1, wherein The controller communication system further includes a backplane. The second connection line is provided on the backplane. The first end of the backplane is connected to the second end of the first complex programmable logic device through the first end of the second connection line. The second end of the backplane is connected to the first end of the second complex programmable logic device through the second end of the second connection line.
3. The controller communication system according to claim 1, wherein The first connection line includes a first sub-connection line, a second sub-connection line and a third sub-connection line; The first sub-connection line is a fault transmission line, and the first sub-connection line is used to transmit target fault information belonging to the software fault type; The second sub-connection line is a serial clock line, and the third sub-connection line is a serial data line. The second sub-connection line and the third sub-connection line are used to query fault information.
4. The controller communication system according to claim 1, wherein The third connection line includes a fourth sub-connection line, a fifth sub-connection line and a sixth sub-connection line; The fourth sub-connection line is a fault transmission line, and the fourth sub-connection line is used to transmit target fault information belonging to the software fault type; The fifth sub-connection line is a serial clock line, and the sixth sub-connection line is a serial data line. The fifth sub-connection line and the sixth sub-connection line are used to query fault information.
5. The controller communication system according to claim 1, characterized in that, The plurality of pins include a first pin, a second pin and a third pin. The hardware fault types include a first hardware fault type, a second hardware fault type and a third hardware fault type. The first pin, the second pin and the third pin respectively match the first hardware fault type, the second hardware fault type and the third hardware fault type.
6. The controller communication system according to claim 1, wherein, The first complex programmable logic device and the second complex programmable logic device are used to control the level states of the plurality of pins; Wherein, in response to the first complex programmable logic device receiving the target fault information transmitted by the second processor, the first complex programmable logic device raises the level of the pin on the first complex programmable logic device that matches the type of the target fault information to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the first processor; In response to the second complex programmable logic device receiving the target fault information transmitted by the first processor, the second complex programmable logic device raises the level of the pin on the second complex programmable logic device that matches the type of the target fault information to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor.
7. The controller communication system according to claim 1, wherein The second connection line includes a main connection line and a slave connection line. The main connection line is used to transmit the target fault information, and the slave connection line is used to transmit the target fault information when the main connection line fails.
8. A controller communication method, based on the controller communication system according to any one of claims 1-7, characterized in that, The controller communication method includes: In response to the first processor obtaining the target fault information, the first processor parses the target fault information to obtain the type of the target fault information corresponding to the target fault information; In response to the type of the target fault information being a hardware fault type, the first processor transmits the target fault information to the second processor through the pin that matches the hardware fault type and the second connection line; In response to the type of the target fault information being a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line, and the third connection line.
9. The controller communication method according to claim 8, wherein The step that in response to the type of the target fault information being a hardware fault type, the first processor transmits the target fault information to the second processor through the pin that matches the hardware fault type and the second connection line includes: Determine the target hardware fault type corresponding to the target fault information of the hardware fault type, where the target hardware fault type is at least one of a first hardware fault type, a second hardware fault type, and a third hardware fault type; Obtain the target pin that matches the target hardware fault type, where the target pin is at least one of a first pin, a second pin, and a third pin; In response to the determination of the target hardware fault type and the target pin being completed, the first processor transmits the target fault information to the first complex programmable logic device through the target pin on the first complex programmable logic device; In response to the first complex programmable logic device receiving the target fault information, the first complex programmable logic device transmits the target fault information to the second complex programmable logic device through the second connection line provided on the backplane; In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device raises the level of the target pin on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor; In response to the second processor receiving the target fault interrupt signal, the second processor takes over the services of the first processor.
10. The controller communication method according to claim 8, wherein, When the target fault information type is a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line, and the third connection line, which includes: Determine the target software fault type corresponding to the target fault information of the software fault type, where the target software fault type is at least one of the first software fault type and the second software fault type; In response to the target software fault type being the first software fault type, the first processor transmits the target fault information of the first software fault type to the first complex programmable logic device through the first sub-connection line in the first connection line; In response to the first complex programmable logic device receiving the target fault information of the first software fault type, the first complex programmable logic device transmits the target fault information of the first software fault type to the second complex programmable logic device through the second connection line provided on the backplane; In response to the second complex programmable logic device receiving the target fault information of the first software fault type, the second complex programmable logic device raises the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device to generate a target fault interrupt signal, and transmits the target fault interrupt signal to the second processor; In response to the second processor receiving the target fault interrupt signal, the second processor takes over the first processor's services.
11. The controller communication method according to claim 8, wherein When the target fault information type is a software fault type, the first processor transmits the target fault information to the second processor through the first connection line, the second connection line, and the third connection line, which further includes: In response to the target software fault type corresponding to the target fault information of the software fault type being the second software fault type, the first processor shuts down or exits the service software corresponding to the target fault information of the second software fault type; The first processor transmits the target fault information of the second software fault type to the first complex programmable logic device through the first sub-connection line in the first connection line; In response to the first complex programmable logic device receiving the target fault information of the second software fault type, the first complex programmable logic device broadcasts the target fault information of the second software fault type to the second complex programmable logic device through the second connection line provided on the backplane; In response to the second complex programmable logic device receiving the target fault information of the second software fault type, the second complex programmable logic device raises the level of the pin corresponding to the fourth sub-connection line on the second complex programmable logic device to generate a target fault interrupt signal, and sends the target fault interrupt signal and the target fault information of the second software fault type to the communication module, where the communication module is communicatively connected to the second processor core; In response to the communication module receiving the target fault information of the second software fault type, the communication module obtains the first controller fault status, generates a fault notification based on the first controller fault status, and adds the fault notification to the event priority adjustment queue; The event priority adjustment queue sends a link disconnection instruction, and executes the link disconnection instruction to interrupt the service of the first controller; In response to the interruption of the service of the first controller, the second processor takes over the service of the first controller.
12. The controller communication method according to claim 8, wherein Before transmitting the target fault information to the second processor, it further includes: In response to the second complex programmable logic device receiving the target fault information, the second complex programmable logic device acquires the fault query interfaces corresponding to the fifth sub-connection line and the sixth sub-connection line, and integrates the acquired target fault information to obtain the integrated fault information; The second complex programmable logic device sends the integrated fault information to the second processor through the fault query interface.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 8-12.
14. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 8 to 12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 8 to 12.
Citation Information
Patent Citations
Baseboard management controller fault detection device
CN114691408A
Switch reset system and method, storage medium and electronic equipment
CN115550291A
Method and device for switching control channels of main and standby main control cards without electronic switch and medium
CN117729099A
Switch reset system and method, non-volatile readable storage medium, and electronic device
WO2024113818A1
Cited By
Hard disk mode control system, server, method, device, medium and product
CN120743201A