Case master control management function redundancy backup method

By deploying a chassis management controller in the VPX and LRM architecture chassis, automatic switching of master-slave working mode and IPMB bus state recovery are achieved, and management function failure caused by abnormal master control modules is solved, ensuring redundant backup of chassis management functions and automation of troubleshooting.

CN120492235AActive Publication Date: 2025-08-15YANGZHOU WANFANG ELECTRONICS TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510618849.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

现有VPX和LRM架构机箱在主控模块异常时无法自动恢复管理功能,导致整机管理功能失效,且在IPMB总线异常情况下,备份模块无法有效接替主控管理功能,影响整机业务功能。

Method used

Deploy the chassis management controller in the chassis, connect the master control and other service modules through the IPMB bus and the serial port to realize automatic switching of the master-slave working mode, and automatically restore the IPMB bus state by using the backup module to force the bus level to restore communication, ensuring redundant backup of the management functions.

Benefits of technology

It realizes redundant backup of management functions without manual intervention when the main control module is abnormal, ensuring the recovery of the health management communication of the whole machine. Users can troubleshoot the faulty module through the WEB interface to avoid interruption of the whole machine's business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492235A_ABST
    Figure CN120492235A_ABST
Patent Text Reader

Abstract

The invention discloses a method for redundant backup of a main control management function of a case, which comprises the following steps of: S1, in the case, realizing signal interconnection of IPMC parts in each module through an IPMB (Intelligent Platform Management Bus) of the case; s2, connecting the main controller and the IPMC in the other business modules carrying the processors with the corresponding processors through a serial port to form respective independent case management controllers; s3, performing data interaction between the WEB management interface and the case management controller; s4, after the case is powered on, the master control module is switched to a host working mode, other service modules serve as backup modules, and the slave working mode is kept unchanged; and S5, when the master control module has a fault and causes the management function to be abnormal, the IPMC in the master control module performs communication interruption judgment through a serial port, the slave working mode is switched to, and the backup module is switched to the host working mode for working. According to the method, an additional module is not needed, manual operation is not needed, the state of the IPMB can be automatically monitored and recovered, the backup module is started to replace the main control management function, and operation is reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent control, and in particular to a method for redundant backup of a master control management function applicable to a VPX or LRM architecture chassis. Background Art

[0002] VPX and LRM chassis are collectively referred to as chassis models designed to meet corresponding management specifications. The VPX architecture is often used in applications requiring high data transmission speed, reliability, and stability. The LRM architecture is commonly used in scenarios such as video transmission, which require high video quality and bandwidth. This architecture offers rapid maintenance and adaptability to harsh environments. Both chassis architectures are designed with standardized management configuration architectures and signal communication interfaces according to specifications, providing hardware support for acquiring system health management information and diagnosing module fault causes.

[0003] Both chassis architectures typically contain a main control module, power supply modules, fan modules, and other service modules. The main control module acquires overall system health information. Each module includes an IPMC (Intelligent Platform Management Controller). The main control module typically uses a CPU to communicate with the IPMC, processing health query commands and exchanging health information with other modules in the chassis via the IPMB (Intelligent Platform Management Bus). Therefore, overall system management functions are primarily integrated into the main control module.

[0004] If the main control module experiences an anomaly during chassis operation, the status of devices mounted on the IPMB bus will be unavailable through the interface or backend, resulting in a failure in health information exchange. In the past, if a main control module problem arose, it was typically reset through the backend or manually unplugged and restarted. During troubleshooting, there was no backup module in the chassis to take over health management functions.

[0005] In some special work environments, the chassis requires extended power supply time and personnel are unable to arrive on-site for troubleshooting. If the main control module malfunctions, the entire system management function may fail and module faults cannot be easily identified through backend troubleshooting. If the IPMB bus becomes unavailable, powering off individual modules is ineffective and the entire system must be powered off and restarted, disrupting overall functionality and leading to more serious consequences.

[0006] Currently, the VPX architecture uses an additional master control module within the chassis for redundant functionality. However, this approach is rarely used in LRM chassis designs. In many scenarios, to prevent resource waste, adding a module for redundant functionality is not adopted. Instead, a backup module is used to take over the primary control management function. This approach only works when the IPMB bus is functioning properly. If the IPMB bus is clamped by any module, even if the backup module takes over the primary control function, the entire system cannot be restored.

[0007] Therefore, it is necessary to design a method for redundant backup of the master management function that can be used in VPX and LRM architecture chassis, automatically restore the IPMB bus status, and does not require the addition of an additional master control module. Summary of the Invention

[0008] In response to the above problems, the present invention provides a method for redundant backup of chassis master management functions, which does not require additional modules or manual operations, can automatically monitor and restore the IPMB bus status and enable a backup module to take over the master management function, and provides a control interface for troubleshooting faulty modules.

[0009] The technical solution of the present invention is: a method for redundant backup of chassis master control management function, comprising the following steps: S1. In a VPX or LRM chassis, install the main control, power supply, fan, and other service modules into the chassis. The IPMC components in each module are interconnected through the chassis IPMB bus. S2. Connect the IPMCs in the main control and other processor-based business modules to the corresponding processors through a serial port to form independent chassis management controllers (CHMCs). S3. Deploy and run the WEB control management interface on the main control or other business modules equipped with processors, and the WEB control management interface exchanges data with the chassis management controller; S4. After the chassis is powered on, the chassis management controller in each module defaults to slave working mode. The IPMC in the module automatically identifies the module type by identifying the current IPMB bus hardware address and switches the management function between master and slave working modes. The master control module switches to the master working mode, and the remaining business modules serve as backup modules and remain in slave working mode. S5. When a failure of the master control module causes abnormal management functions, the IPMC in the master control module determines communication interruption through the serial port and switches to slave mode. Other backup modules including the chassis management controller trigger the management function redundancy backup function by determining the management query timeout on the IPMB channel, and switch the backup module to master mode to operate. S6. After each backup module switches to the host working mode and fails to restore the health management communication of the entire machine, each module in the chassis calculates the maximum IPMB management timeout coefficient based on the maximum number of slots in the chassis, and triggers the IPMB bus recovery mechanism when the coefficient is reached; S7. All modules within the chassis redefine the IPMB bus communication control pins to forced push-pull output mode, forcing the IPMB bus voltage level high. After the voltage level is raised, the module's I2C peripherals are reinitialized, restoring the module's IPMC peripheral communication function. Each backup module re-enables IPMB management bus timeout monitoring and enters the preparation queue for redundant backup of the master management function.

[0010] S8. After the IPMB bus communication returns to normal, the backup module with the shortest IPMB channel detection timeout coefficient takes over the master control management function and resumes operation. Once the entire system is functioning normally, the user can control and troubleshoot the faulty module through the web interface. This provides a redundant backup for the chassis master control management function.

[0011] In step S1, each module in the chassis transmits the IPMI (Intelligent Platform Management Interface) communication protocol via the IPMB management bus.

[0012] In step S2, the chassis management controller communicates with the IPMC via the IPMI protocol to obtain management information of each module.

[0013] In step S2, in order to be compatible with the module design in the VPX and LRM architecture chassis, management software can be deployed in any business module containing a processor in the chassis except the main control module to form a complete chassis management controller.

[0014] In step S3, the WEB control management interface can process the IPMI data of each module in the chassis and display it on the interface, and can send independent control commands to related modules through the graphical buttons on the interface.

[0015] In step S3, the chassis management controller communicates with the WEB control management interface in an active reporting manner, using a many-to-one star communication topology to ensure that the WEB control management interface can work normally when the active reporting source changes.

[0016] In step S4, after the whole machine is powered on, the IPMC of each module detects the hardware address of the IPMB in the current slot. If it is the hardware address of the master module slot, it switches to the master working mode, and the modules in other slots remain in the slave working mode.

[0017] If the backup module switches to master mode, the module's IPMC sends a master mode switch command to the processor via the serial port. After the master module switch is complete, the chassis management controller begins acquiring management information for the entire system. Upon receiving management query commands for this module, the IPMC processes and responds via the serial port, forwarding query commands from other modules via the IPMB bus. In slave mode, the IPMC only responds to query messages received on the IPMB bus and does not communicate with the processor via the serial port.

[0018] In step S5, when the main control module fails, the main control management function communication is interrupted, and the IPMC performs a management query function timeout detection. When the communication interruption time exceeds the preset value, the module is switched to the slave working mode, and the IPMC closes the communication serial port with the processor.

[0019] In step S5, other modules containing the chassis management controller (CMC) obtain a timeout coefficient based on the IPMB bus hardware address. By determining the communication status on the IPMB bus, they sequentially perform management query function timeout checks based on the calculated timeout coefficients. If the timeout exceeds a preset time, the module switches to master mode, and the IPMC opens the serial communication port with the processor, causing the CMC of the corresponding module to enter master mode. If the IPMB bus returns to normal, the other modules cancel the management query function timeout check and continue normal master management function communication.

[0020] In step S5, the main control redundant backup module can be single or multiple. In order to distinguish and monitor the delay coefficient of host communication timeout on the IPMB channel, each module IPMC will perform delay counting according to the IPMB address of the slot where it is located.

[0021] In step S6, if each backup module cannot restore the health management communication of the entire machine after switching the host working mode, the main control, power supply, fan and other business modules in the chassis calculate the IPMB management timeout maximum time coefficient based on the current maximum number of slots in the chassis. After the IPMB cumulative communication interruption time reaches the maximum coefficient, each module containing IPMC triggers the IPMB bus recovery mechanism.

[0022] In step S7, the module containing the IPMC in the chassis redefines the I / O mode of the corresponding control pin for IPMB bus communication to push-pull output mode, forcing the IPMB bus voltage level to rise. During this period, the I / O mode of the configured pin is changed to pull-up input mode, and the current bus voltage level is tested. After the voltage level is raised, the module's I2C peripheral is reinitialized. Otherwise, the configured pin I / O mode is further changed to forced push-pull output mode, raising the IPMB bus voltage level until IPMB bus communication function is restored. Each backup module then re-enables IPMB management bus timeout monitoring and enters the preparation queue for redundant backup of the master management function.

[0023] In step S8, after the entire system's IPMB bus status returns to normal, the backup module with the shortest IPMB channel detection timeout coefficient takes over the master management function and resumes operation. After the master management function's redundant backup is restored, the user can view the status of the faulty module through the web-based control and management interface and troubleshoot the cause of the faulty module using IPMI data. Based on this status information, the user can issue a command to the chassis management controller through the web-based control and management interface, instructing the corresponding module's IPMC to power off or reset the previously faulty module.

[0024] If the health information of the original faulty module returns to normal, the IPMC of the original faulty module maintains the slave working mode and retains the redundant backup function of the master management function. When the master management function of the current backup module becomes abnormal, the original faulty module that has returned to normal can re-enter the preparation queue for the redundant backup of the master management function.

[0025] In operation, this invention utilizes the existing VPX and LRM chassis architecture and deploys a chassis management controller within the master control and other processor-based service modules, achieving redundant backup of the master control management function without requiring additional modules. After an IPMB bus anomaly occurs, each module within the chassis automatically restores the IPMB bus level, a process automated by software without manual intervention. During the redundant switchover of the entire machine health management function, the web-based control and management interface maintains normal operation, and troubleshooting of the previously faulty module can be performed through the web-based control and management interface without affecting the overall machine operation, making it simple and convenient. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a schematic diagram of the connection relationship between the modules in the chassis. Figure 2 This is a diagram showing the normal operation of the master management function. Figure 3 This is a diagram showing the abnormal operation of the main control management function. Figure 4 This is a working diagram of the redundant backup function of the master control management function. Figure 5 This is a flowchart of the chassis master management function redundancy backup. Figure 6 This is a flowchart of the chassis IPMB bus automatic recovery function. DETAILED DESCRIPTION

[0027] The present invention Figure 1-6As shown, the intelligent chassis of VPX or LRM architecture contains a main control, power supply, fan and several business modules. Each module is equipped with an IPMC (i.e., intelligent platform management controller). Each IPMC needs to be connected to two IPMB buses (i.e., intelligent platform management buses), and the IPMC can obtain a unique IPMB communication address based on the different slots in the chassis.

[0028] Each module IPMC has a reserved serial port (UART) for communicating with the processor (CPU) in the module. During use, the main control and other modules containing processors can be selected to deploy CHMC (chassis management controller) software, and data communication is carried out using the serial port through the CHMC software and IPMC. Among them, the IPMC, which communicates in host mode, not only responds to the health query instructions of this module, but is also responsible for forwarding IPMI messages of other boards through the IPMB bus. Figure 1 The power supply, fan, and other business modules of the IPMC remain in slave mode during the operation of the entire machine and do not participate in the redundant backup of management functions.

[0029] like Figure 2 and Figure 3 As shown, the entire machine's health management information can ultimately be presented through an independent web-based control and management interface. Simple Network Management Protocol (SNMP) enables data exchange with the CHMC processor, displaying module information in host mode and acquired IPMI data from other slave devices within the chassis in a graphical interface and corresponding information columns.

[0030] The Chassis Management Controller communicates with the web-based control and management interface through active reporting. Upon receiving correct data, the web-based control and management interface parses and processes it for display. To ensure the proper functioning of the web-based control and management interface's signal source channels, a many-to-one star communication topology is employed. Even if the active reporting source changes, the web-based control and management interface can still access data through other channels. The web-based control and management interface displays information consistently throughout its operation, and redundancy switching of the master control and management functions will not affect the interface display.

[0031] like Figure 2 As shown in the figure, when the entire chassis is first powered on, all modules in the chassis default to slave mode. The IPMC in each module identifies the current module type by obtaining the IPMB bus hardware address of the corresponding slot. The master IPMC opens the serial communication port with the processor and sends a master switch command to the processor, causing the module to enter master mode. The remaining modules operate in slave mode.

[0032] During normal operation of the entire machine, the module working in host mode will communicate data through the IPMB bus. The remaining modules participating in the management function redundancy backup and other slave devices will clear the communication interruption detection flag after receiving the data and re-detect the communication interruption status.

[0033] like Figure 3 As shown in the figure, if the main control module malfunctions during system operation, the IPMC in the module will monitor the serial port communicating with the processor and detect management query function timeouts. If the communication interruption time exceeds a preset value, the IPMC will switch to slave mode and close the serial port communicating with the processor. The chassis management controller in the module will also interrupt information communication with the web-based control and management interface.

[0034] If the master control module malfunctions, no information will be transmitted on the IPMB bus within the chassis, and all IPMCs mounted on the bus will operate in slave mode. At this time, other backup modules including the chassis management controller obtain the timeout wait coefficient based on the IPMB bus hardware address and monitor the communication status on the IPMB bus for timeouts.

[0035] like Figure 4 As shown, the backup module, the IPMC, which includes the chassis management controller (Chassis Management Controller), monitors the IPMB bus communication interruption duration. If it exceeds a preset time, the current module IPMC switches to master mode. The serial port between the IPMC and the processor is opened, and a master mode switch command is sent to the processor, causing the Chassis Management Controller in the corresponding module to enter master mode. If the IPMB bus returns to normal, the other modules cancel the management query function timeout detection and resume normal master management function communication.

[0036] When there are multiple redundant backup modules for the master management function in the chassis, the time it takes for each backup module to trigger a detection timeout is calculated based on the IPMB hardware slot address and the delay coefficient, avoiding communication conflicts caused by multiple backup modules switching to host mode at the same time.

[0037] like Figure 6 As shown in the figure, if the whole machine still cannot restore health management communication after each backup module switches to the host working mode, the main control, power supply, fan and other service modules in the chassis calculate the maximum IPMB management timeout coefficient (number of slots * delay coefficient) based on the current maximum number of slots in the chassis. After the cumulative communication interruption time of the monitored IPMB reaches the maximum coefficient, each module containing IPMC triggers the IPMB bus recovery mechanism.

[0038] Each module in the chassis redefines the I / O mode of the corresponding control pin for IPMB bus communication in the IPMC, changing it to strong push-pull output mode and forcibly pulling the IPMB bus level high. To detect whether the IPMB bus has returned to a high level (idle state), the pin's I / O mode is changed to pull-up input mode during the forced bus pull-up period to test the current bus level. After detecting a high level, the module's I2C peripheral is reinitialized. If the level remains low, the pin's I / O mode is further changed to strong push-pull output mode to pull the IPMB bus level high until IPMB bus communication function is restored.

[0039] The I / O pin mode change mentioned above involves switching from the multiplexed open-drain mode used for communication with the microcontroller's I2C peripheral to push-pull output mode. In push-pull output mode, current flows when the microcontroller pulls the pin high or low, providing improved drive capability. After the communication pin completes the bus pull-up operation, the microcontroller's I / O port operating mode is configured to input detection mode to check the bus status. If the bus status remains abnormal, the pull-up operation is repeated. If the bus returns to normal, the I2C peripheral is reinitialized, resetting the microcontroller's internal I2C control registers and resetting the I2C communication I / O port operating mode.

[0040] After the IPMB bus returns to normal, each backup module will re-monitor the IPMB management bus timeout and enter the preparation queue for redundant backup of the main control management function. The backup module with the shortest IPMB channel timeout coefficient will take over the main control management function and resume work.

[0041] After the backup module takes over the master control function, the chassis management controller receives and processes health information and actively reports it to the web-based control and management interface, which can normally display the status of each slot in the chassis. The web-based control and management interface can be used to view the status of the faulty module, and the cause of the faulty module can be identified through IPMI data information. The web-based control and management interface can also be used to issue instructions to the chassis management controller, causing the corresponding module's IPMC to control the original faulty module to perform a power-off or reset operation.

[0042] If the health information of the original faulty module returns to normal, the IPMC of the module will maintain the slave working mode, retain the redundant backup function of the master management function, and re-enter the preparation queue of the redundant backup of the master management function.

[0043] During operation, the present invention can have a single or multiple redundant backup modules for the master management function, as long as they meet the chassis management controller's component requirements. Because the timeout detection period for each backup module triggering the management query function is determined by the IPMB hardware slot address, communication conflicts caused by multiple backup modules simultaneously switching to host mode after a master management function failure are eliminated. In the event of an IPMB bus failure, the IPMB bus can be automatically restored because each module in the chassis includes an IPMC.

[0044] The above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the same. Although specific embodiments are described in detail herein, those skilled in the art may modify them or replace some of the technical features with equivalents, and such modifications do not deviate from the core concept and scope of protection embodied in the embodiments of the present invention.

Claims

1. A method for redundant backup of chassis master control management function, characterized in that: The following steps are involved: S1. In a VPX or LRM chassis, install the main control, power supply, fan, and other service modules into the chassis. The IPMC components in each module are interconnected through the chassis IPMB bus. S2. Connect the IPMCs in the main control and other processor-equipped business modules to the corresponding processors through a serial port to form independent chassis management controllers. S3. Deploy and run the web control management interface on the main control or other business modules equipped with processors, and the web management interface exchanges data with the chassis management controller; S4. After the chassis is powered on, the chassis management controller in each module defaults to slave working mode. The IPMC in the module automatically identifies the module type by identifying the current IPMB bus hardware address and switches the management function between master and slave working modes. Specifically, the master control module switches to the master working mode, and the remaining business modules serve as backup modules and remain in slave working mode. S5. When the main control module fails and causes abnormal management functions, the IPMC in the main control module determines the communication interruption through the serial port and switches to the slave working mode. Other backup modules including the chassis management controller trigger the management function redundancy backup function by determining the management query timeout on the IPMB channel, and switch the backup module to the host working mode to work.

2. The method for redundant backup of chassis master control management function according to claim 1, characterized in that: Also includes: S6. After each backup module switches to the host working mode and fails to restore the health management communication of the entire machine, each module in the chassis calculates the maximum IPMB management timeout coefficient based on the maximum number of slots in the chassis, and triggers the IPMB bus recovery mechanism when the coefficient is reached; S7, each module redefines the control pin of the IPMB bus communication to a strong push-pull output mode, forcibly pulling up the IPMB bus level, and after the level is pulled high, reinitializes the module I2C peripheral and restores the module IPMC corresponding peripheral communication function; Each backup module re-performs IPMB management bus timeout monitoring and enters the preparation queue for redundant backup of the master control management function; S8. After the IPMB bus communication returns to normal, the backup module with the shortest IPMB channel detection timeout coefficient takes over the main control management function and resumes work; after the health function of the entire machine is operating normally, the user can control and troubleshoot the faulty module through the WEB interface.

3. The method for redundant backup of chassis master control management function according to claim 1, characterized in that: In step S2, the chassis management controller communicates with the IPMC via the IPMI protocol to obtain management information of each module.

4. The method for redundant backup of chassis master control management function according to claim 1, characterized in that: In step S3, the chassis management controller in host mode communication communicates with the WEB control management interface in an active reporting manner, and is connected using a many-to-one star communication topology.

5. The method for redundant backup of chassis master control management function according to claim 1, characterized in that: In step S5, other modules including the chassis management controller obtain the timeout waiting coefficient according to the IPMB bus hardware address, judge the communication status on the IPMB bus, and perform management query function timeout detection in sequence according to the calculated timeout waiting coefficient. If the preset time is exceeded, the module is switched to the host, and the IPMC opens the serial port for communication with the processor, so that the chassis management controller of the corresponding module enters the host working mode.

Citation Information

Patent Citations

  • VPX equipment intelligent cabinet management system

    CN106066821A

  • Intelligent platform management system and fault handling method

    CN108388488A

  • Method for realizing redundant backup of board card in OpenVPX equipment

    CN110995478A

  • Method for upgrading firmware of intelligent platform management controller of VPX case on line

    CN111427602A

  • Automatic recovery method and device for hanging of I2C bus, equipment and medium

    CN116820840A