Multi-case health management system based on advanced telecommunication computing architecture
By adopting a health management system based on an advanced telecommunications computing architecture in a multi-chassis system and utilizing ShMC role competition and authority management, the role competition and fault recovery problems in multi-chassis collaborative work are solved, achieving real-time monitoring of the system and improving security.
Patent Information
- Application Number
- CN202410472443.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-19
- Publication Date
- 2025-10-24
AI Technical Summary
ATCA-based chassis cannot meet the needs of telecommunications services in terms of capacity, performance, and single-point fault disaster recovery. Multi-chassis systems are prone to problems such as role competition, information synchronization, cross-chassis access rights, and fault location and recovery when working together.
A multi-chassis health management system based on an advanced telecom computing architecture is adopted. The master ShMC competes for roles based on timestamps and priorities to achieve alarm information synchronization and fault recovery, ensuring the collaborative operation of multiple chassis. Cross-chassis access is managed through ShMC roles with different permissions.
It improves the maintainability and reliability of multi-chassis systems, ensures real-time monitoring and security of the system, achieves rapid fault detection and recovery, and enhances system capacity and performance.
Smart Images

Figure CN120832284A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to an ATCA-based chassis, in particular, a multi-chassis health management system based on an advanced telecommunications computing architecture. BACKGROUND
[0002] Due to the capacity, performance and single-point fault tolerance backup considerations, a single ATCA-based chassis is difficult to meet the needs of the rapid development of telecommunications services, and when a multi-chassis system is deployed, multiple chassis are prone to confusion in cooperative work. SUMMARY
[0003] In order to solve the common problems of multiple chassis in role competition and switching, information synchronization, cross-chassis access permissions, fault location and recovery, and ensure the cooperation, safety, efficiency and redundancy of multiple chassis, the present application provides a multi-chassis health management system based on an advanced telecommunications computing architecture, which simplifies the daily operation of multi-chassis health management, enhances the maintainability and reliability of the multi-chassis health management system, and makes the multi-chassis health management based on ATCA more convenient, real-time and efficient.
[0004] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0005] In an embodiment of the present application, a multi-chassis health management system based on an advanced telecommunications computing architecture is provided, which comprises:
[0006] The active ShMC of each chassis competes for the roles of the primary group leader ShMC, the deputy group leader ShMC and the group member ShMC of the multi-chassis health management according to the time stamp when it enters the active working state and the set priority;
[0007] The active ShMC of each chassis automatically synchronizes the alarm information of the chassis to the primary group leader ShMC;
[0008] The primary group leader ShMC has the permission to perform all read and write operations for health management of each chassis; the deputy group leader ShMC has the permission to perform read operations for health management of each chassis and write operations for health management of the chassis; and the group member ShMC has the permission to perform all read and write operations for health management of the chassis;
[0009] When the ShMC of the multi-chassis fails or the link between the chassis is interrupted, automatic fault recovery processing is performed.
[0010] Further, the active ShMC of each chassis competes for the roles of the primary group leader ShMC, the deputy group leader ShMC and the group member ShMC of the multi-chassis health management according to the time stamp when it enters the active working state and the set priority, including:
[0011] When the priority of the primary ShMC of each chassis is different, the primary ShMC with the highest priority assumes the role of the primary group leader, the primary ShMC with the second highest priority assumes the role of the deputy group leader, and the other primary ShMCs assume the role of group members;
[0012] When the priority of the primary ShMC of each chassis is the same, the primary ShMCs compete for the roles according to the time stamp when they enter the primary working state, the primary ShMC with the earliest time stamp assumes the role of the primary group leader, the primary ShMC with the second earliest time stamp assumes the role of the deputy group leader, and the other primary ShMCs assume the role of group members.
[0013] Further, when the ShMC of a chassis starts or switches, the time stamp when it enters the primary working state is automatically recorded, and the primary ShMC uses a uniformly allocated floating IP address to communicate externally;
[0014] When configured as multiple chassis, the ShMC of each chassis stores the floating IP addresses of all primary ShMCs of the chassis; after the primary ShMC starts, it automatically initializes its role as a group member ShMC and uses its own floating IP address to actively send a role negotiation request message to the floating IP addresses of all primary ShMCs of other chassis; the role negotiation request message carries the time stamp when the primary ShMC enters the primary working state, the set priority, and the floating IP address of the primary ShMC;
[0015] When the primary group leader ShMC receives the role negotiation request message, it compares the priority and time stamp of the other party with its own internal records to determine the role of the other party and broadcast a notification to the primary ShMC of each chassis;
[0016] When the deputy group leader ShMC and the group member ShMC receive the role negotiation request message, they directly ignore the message and do not perform any processing.
[0017] Further, when configured as multiple chassis, after the primary ShMC of each chassis starts, if it does not receive a response for three consecutive times of sending a role negotiation request message, it automatically changes its role to the primary group leader ShMC, and then initiates a role negotiation request message at a certain interval; when the primary group leader ShMC discovers that there is still a group member ShMC that has not joined the group according to the floating IP address list of the group member ShMC, it initiates a role negotiation request message to the group member ShMC at a certain interval.
[0018] Further, the primary group leader ShMC and the deputy group leader ShMC have fixed primary and deputy group leader floating IP addresses, respectively, when the roles of the primary and deputy group leaders change, the fixed floating IP addresses are automatically bound to the current primary group leader ShMC and deputy group leader ShMC, and the current primary group leader ShMC and deputy group leader ShMC can be accessed by accessing the two fixed floating IP addresses externally, thereby realizing multi-chassis access management.
[0019] Further, when the ShMC under multi-chassis fails, the automatic fault recovery process is performed, including:
[0020] When the primary group leader ShMC fails, the secondary group leader ShMC and the group member ShMC send the alarm synchronization message to the primary group leader ShMC and cannot receive the response; when the secondary group leader ShMC finds that the connection with the primary group leader ShMC fails, and at least one group member ShMC reports the primary group leader ShMC connection failure to the secondary group leader ShMC, the secondary group leader ShMC immediately switches to the new primary group leader ShMC, and the original primary group leader ShMC becomes a lost group member, and the new primary group leader ShMC generates the corresponding alarm;
[0021] When the secondary group leader ShMC fails, the primary group leader ShMC cannot receive the alarm synchronization message from the secondary group leader ShMC, then the primary group leader ShMC selects the new secondary group leader ShMC according to the priority and the time stamp in the stored group member information table, and notifies all group members, finally sends the stored group member record to the secondary group leader ShMC, and the original secondary group leader ShMC becomes a lost group member, and the primary group leader ShMC generates the corresponding alarm;
[0022] When the group member ShMC fails, the primary group leader ShMC cannot receive the alarm synchronization message from the group member ShMC; the group member ShMC becomes a lost group member, and the primary group leader ShMC generates the corresponding alarm.
[0023] Further, when the inter-chassis link under multi-chassis is interrupted, the automatic fault recovery process is performed, including:
[0024] When the network port 0 or the network port 1 of the primary group leader ShMC does not receive the alarm synchronization information from the secondary group leader ShMC or a certain group member ShMC for three consecutive periods, the primary group leader ShMC judges that the link connection of the network port 0 or the network port 1 of the secondary group leader ShMC or the certain group member ShMC is interrupted, and generates the corresponding alarm;
[0025] When the network port 0 and the network port 1 of the primary group leader ShMC do not receive the alarm synchronization confirmation message from a certain group member ShMC or the secondary group leader ShMC for three consecutive periods, the primary group leader ShMC judges that the connection with the group member ShMC or the secondary group leader ShMC is interrupted, and generates the corresponding alarm; at the same time, the primary group leader ShMC initiates a role negotiation request message to the lost group member ShMC or the secondary group leader ShMC every certain period of time;
[0026] When the net port 0 and the net port 1 of a group member ShMC or a deputy group leader ShMC do not receive the alarm synchronization confirmation message returned by the positive group leader ShMC for three continuous periods, it is judged that the connection with the positive group leader ShMC is interrupted, the positive group leader ShMC role is set, and then the role negotiation request message is re-initiated to the main ShMC of other cabinets through the net port 0 and the net port 1 every time interval until the link is recovered and the positive group leader ShMC role is re-competited.
[0027] Beneficial effects:
[0028] 1、The ShMC enters the main state according to the time stamp and the set priority, and competes for the multi-cabinet health management positive group leader ShMC, deputy group leader ShMC and group member ShMC role, so that the positive group leader ShMC and the deputy group leader ShMC are ensured to be in two cabinets.
[0029] 2、The alarm information of all cabinets is automatically synchronized to the positive group leader ShMC, so that the positive group leader ShMC can monitor and handle the alarm of each cabinet in time, and the real-time monitoring of the whole system is ensured.
[0030] 3、The positive group leader ShMC and the deputy group leader ShMC have different permissions when accessing across cabinets in the application: the positive group leader ShMC has read and write permissions when accessing across cabinets, and the deputy group leader ShMC only has read permission when accessing across cabinets, so that the access security of the whole system is ensured.
[0031] 4、The application can automatically detect and discover internal faults of cabinets and connection faults between cabinets, automatically and quickly realize the switching of the positive group leader ShMC by the deputy group leader ShMC, and ensure the continuity of the whole system management and control. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 It is three ATCA-based cabinet interconnection schematic diagrams of the application;
[0033] Figure 2 It is a schematic diagram of sending a role negotiation request message from cabinet 2 to cabinet 1 and cabinet 3;
[0034] Figure 3 It is a role negotiation request message sending flow chart of the application;
[0035] Figure 4 It is a processing flow chart of the positive group leader ShMC after receiving the role negotiation request message;
[0036] Figure 5 It is a processing flow chart of the deputy group leader ShMC and the group member ShMC after receiving the role negotiation request message;
[0037] Figure 6 It is an alarm information synchronization flow chart of the application;
[0038] Figure 7 is a schematic diagram of the ShMC of the primary group executing a human-machine command across chassis according to the present application;
[0039] Figure 8 is a schematic diagram of the ShMC of the secondary group executing a human-machine command across chassis according to the present application;
[0040] Figure 9 is a schematic diagram of the process when the ShMC of the primary group fails according to the present application;
[0041] Figure 10 is a schematic diagram of the process when the ShMC of the secondary group fails according to the present application;
[0042] Figure 11 is a schematic diagram of the process when the ShMC of the group member fails according to the present application;
[0043] Figure 12 is a schematic diagram of a link between chassis being interrupted according to the present application;
[0044] Figure 13 is a schematic diagram of both links between chassis being interrupted according to the present application. DETAILED DESCRIPTION
[0045] The principles and spirits of the present application will be described below with reference to several exemplary embodiments. It should be understood, however, that the embodiments are given solely for the purpose of illustration and are not meant to limit the scope of the present application in any way. On the contrary, the embodiments are provided to make the disclosure more thorough and complete and to convey the scope of the disclosure to those skilled in the art.
[0046] It is to be understood that the embodiments of the present application can be implemented in a variety of forms, including as a system, a device, an apparatus, a method or a computer program product. Therefore, the disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or a combination of hardware and software.
[0047] According to the embodiment of the present application, a multi-chassis health management system based on an advanced telecom computing architecture is provided, further increasing system capacity, improving performance, and enhancing the redundancy backup capability of single-chassis failure. The main ShMC of each chassis competes to determine the role according to the time stamp and preset priority: one is a primary group leader ShMC, one is a deputy group leader ShMC, and the others are group member ShMCs. The primary group leader ShMC and the deputy group leader ShMC have fixed primary and deputy group leader floating IP addresses, respectively. When the primary and deputy group leader roles change, the fixed floating IP addresses are automatically bound to the current primary group leader ShMC and the deputy group leader ShMC. External access to the two floating IP addresses can access the current primary group leader ShMC and the deputy group leader ShMC, thereby realizing multi-chassis access management.
[0048] Each chassis automatically synchronizes its alarm information list to the primary group leader ShMC, so that the primary group leader ShMC can timely monitor the alarm information of each chassis and decide whether to notify the superior network management and operation and maintenance personnel according to the alarm level and alarm content, so as to timely handle.
[0049] The primary group leader ShMC has the permission to perform health management read and write operations on each chassis; the deputy group leader ShMC has the permission to perform health management read and write operations on the chassis, but the deputy group leader can only perform health management read operations on other chassis and cannot perform write operations; the group member ShMC can only perform health management read and write operations on the chassis and cannot perform health management read and write operations on the chassis where the primary and deputy group leaders are located, thereby realizing the allocation of different operation permissions to the primary, deputy group leaders and group members, realizing the collaborative work and safe and convenient health management system of multiple chassis.
[0050] The principles and spirits of the present application will be explained in detail below with reference to several representative embodiments of the present application.
[0051] The present application is suitable for multiple chassis, and the following description is made according to three chassis.
[0052] 1. Multi-chassis health management overall scheme
[0053] As shown in Figure 1 , each ATCA-based chassis contains two switching boards, which adopt a load sharing mode, and are used for Ethernet data exchange of the boards and ShMCs in the chassis and interconnection between the chassis.
[0054] Each chassis has two ShMCs in active-standby mode, and the ShMCs have two network interfaces, of which network interface 0 (ETH0) is connected to switch board 1, and network interface 1 (ETH1) is connected to switch board 2. The active ShMC is responsible for health management of the chassis, mainly including periodic collection of sensor data such as temperature and voltage, hardware working state, hardware alarm management, clock priority management, and the like. When the active ShMC is retired due to failure or human intervention, the standby ShMC immediately switches to the active ShMC to assume the health management task.
[0055] When configured as three chassis, the active ShMCs of the three chassis determine the roles according to the timestamps and priorities: one becomes the health management primary group leader ShMC, one becomes the health management secondary group leader ShMC, and one becomes the health management group member ShMC.
[0056] (1) The primary group leader ShMC has the permission to perform health management read and write operations on each chassis; the active ShMC of each chassis automatically synchronizes the alarm information of the chassis to the primary group leader ShMC, so that the primary group leader ShMC can timely monitor the alarm information of each chassis.
[0057] (2) The secondary group leader ShMC has the permission to perform health management read operations on each chassis, but only has the permission to perform health management write operations on the chassis, and has no permission to perform health management write operations on other chassis.
[0058] (3) The group member ShMC can only perform health management read and write operations on the chassis, and cannot perform health management read and write operations on other chassis.
[0059] Therefore, when there are multiple chassis but only one chassis, the active ShMC of the chassis is the primary group leader, and there is no secondary group leader and group member; when there are two chassis, there are a primary group leader and a secondary group leader, but there is no group member; when there are three or more chassis, there is one primary group leader and one secondary group leader, and the rest are group members.
[0060] 2. Selection of the roles of the primary group leader ShMC, the secondary group leader ShMC, and the group member ShMC in the health management of multiple chassis
[0061] The active ShMC of each chassis is artificially set with a priority, and the priorities of the active ShMCs of the chassis can be the same or different. When the priorities are different, the active ShMC with the highest priority assumes the role of the primary group leader, the active ShMC with the second highest priority assumes the role of the secondary group leader, and the other active ShMCs assume the roles of group members. When the priorities are the same, the active ShMCs compete for the roles according to the timestamps when they enter the active state, the active ShMC with the earliest timestamp assumes the role of the primary group leader, then the secondary group leader, and the others are group members.
[0062] When ShMC starts or switches, it automatically records the time stamp when it enters the master working state, and the master ShMC uses the uniformly allocated floating IP to communicate externally.
[0063] When configured as a multi-chassis, the ShMC of each chassis stores the floating IP address of all chassis master ShMCs. When the master ShMC starts, it automatically initializes its role as a member ShMC and uses its own floating IP address to actively send a role negotiation request message to the floating IP addresses of all other chassis master ShMCs through two network interfaces (ETH0 and ETH1). As shown in Figure 2 When chassis 2 is newly added to the multi-chassis network consisting of chassis 1 and chassis 3, the master ShMC of chassis 2 sends four role negotiation request messages, which are: (1) sent to the master ShMC of chassis 1 through network interface ETH0 and switch 1; (2) sent to the master ShMC of chassis 1 through network interface ETH1 and switch 2; (3) sent to the master ShMC of chassis 3 through network interface ETH0 and switch 1; (4) sent to the master ShMC of chassis 3 through network interface ETH1 and switch 2.
[0064] The role negotiation request message carries the time stamp when the ShMC enters the master working state, the configured priority, and its own floating IP address, and the message flow is as shown in Figure 3 .
[0065] When the primary group leader receives the role negotiation request message, it compares the priority and time stamp of the other party with its own internal records to determine the role of the other party and broadcast a notification to each chassis master ShMC. The processing flow of the primary group leader after receiving the role negotiation request message is as shown in Figure 4 .
[0066] When the deputy group leader ShMC and the member ShMC receive the role negotiation request message, they directly ignore the message and do not perform any processing, and the flow is as shown in Figure 5 .
[0067] When configured as a multi-chassis, after the master ShMC of each chassis is started, if it does not receive a response for three consecutive times (with an interval of 5 seconds) of sending a role negotiation request message, it automatically changes its role to the primary group leader ShMC, and then initiates a role negotiation request message every 10 seconds. When the primary group leader discovers that there are still members who have not joined the group according to the configured member floating IP address list, it initiates a role negotiation request message to the member every 10 seconds.
[0068] The role (GroupRole) of the user in the multi-chassis health management can be queried by using the human-machine command on the active ShMC: group leader (Group_Active), deputy group leader (Group_Standby), or group member (Group_Member).
[0069] The result of querying the active ShMC is as follows:
[0070] / extlog#. / shClient GetShMCInfo
[0071]
[232] [exec_cmd]server_port:9736;(len,cmd)=>12,GetShMCInfo
[0072]
[195] [cli_tcp]GetShMCInfo resp=>
[0073] Location:Rack_1ShMC_1
[0074] Role:Active
[0075] HwAddr:0x10
[0076] OtherShMC:Present
[0077] OtherShMCHealth:Yes
[0078] StartTime:2024-01-3013:46:34
[0079] GroupRole:Group_Active
[0080] GroupActiveShMC:Rack_1ShMC_1
[0081] GroupStandbyShMC:Rack_0ShMC_1
[0082] GroupChannel0Status:InService
[0083] GroupChannel1Status:InService
[0084] Version:V1.3.0
[0085] The result of querying the active ShMC is as follows:
[0086] / extlog#. / shClient GetShMCInfo
[0087]
[232] [exec_cmd] server_port: 9736; (len,cmd) => 12, GetShMCInfo
[0088]
[195] [cli_tcp] GetShMCInfo resp =>
[0089] Location: Rack_2ShMC_1
[0090] Role: Active
[0091] HwAddr: 0x10
[0092] OtherShMC: Present
[0093] OtherShMCHealth: Yes
[0094] StartTime: 2024-02-03 17:01:55
[0095] GroupRole: Group_Standby
[0096] GroupActiveShMC: Rack_1ShMC_1
[0097] GroupStandbyShMC: Rack_0ShMC_1
[0098] GroupChannel0Status: InService
[0099] GroupChannel1Status: InService
[0100] Version: V1.3.0
[0101] Query group member ShMC result as follows:
[0102] / extlog#. / shClient GetShMCInfo
[0103]
[232] [exec_cmd] server_port: 9736; (len,cmd) => 12, GetShMCInfo
[0104]
[195] [cli_tcp] GetShMCInfo resp =>
[0105] Location:Rack_3ShMC_2
[0106] Role:Standby
[0107] HwAddr:0x11
[0108] OtherShMC:Present
[0109] OtherShMCHealth:Yes
[0110] StartTime:2024-01-3013:46:48
[0111] GroupRole:Group_Member
[0112] GroupActiveShMC:Rack_1ShMC_1
[0113] GroupStandbyShMC:Rack_0ShMC_1
[0114] GroupChannel0Status:InService
[0115] GroupChannel1Status:InService
[0116] Version:V1.3.0
[0117] 3. Synchronization of multi-chassis health management alarm information
[0118] After determining the roles of the leader, deputy leader, and team members of the multi-chassis health management, the deputy leader ShMC and each team member ShMC take turns sending all alarm information of the chassis to the leader ShMC via the network port 0 (ETH0) and network port 1 (ETH1) channels every 1 second. After receiving the alarm information, the leader ShMC stores the alarm information in the cache according to the chassis and returns a confirmation message, such as Figure 6 When the Ethernet port 0 (ETH0) channel fails, data will be sent from the Ethernet port 1 (ETH1) channel instead, and vice versa.
[0119] The team leader ShMC decides whether to report the alarm to the operation and maintenance center and the alarm panel for audio and visual presentation based on the urgency and content of the alarm of each chassis in the cache, thereby monitoring the health management alarm of each chassis.
[0120] 4. Execute human-machine commands across multiple chassis in a multi-chassis configuration
[0121] The primary ShMC can query and modify commands of the slave ShMC and the member ShMC in the chassis. As shown in Figure 7 the primary ShMC of the chassis 1 is the primary ShMC of the multi-chassis health management, in addition to being able to execute the query and modification of the owner machine command of the chassis 1, it can also execute the query and modification of the owner machine command of the chassis 2 and the chassis 3.
[0122] The slave ShMC can execute the cross-chassis query command to query the information of the primary ShMC and the member ShMC in the chassis, but cannot execute the cross-chassis modification command. As shown in Figure 8 the primary ShMC of the chassis 1 is the primary ShMC of the multi-chassis health management, in addition to being able to execute the query and modification of the owner machine command of the chassis 1, it can also execute the query and modification of the owner machine command of the chassis 2 and the chassis 3.
[0123] The member ShMC cannot execute the cross-chassis query and modification command, and can only query and modify the information of the local chassis.
[0124] 5. Fault handling under multi-chassis configuration
[0125] (1) Primary ShMC failure
[0126] When the primary ShMC fails, the slave ShMC and the member ShMC send the alarm synchronization message to the primary ShMC and cannot receive the response. When the slave ShMC finds that the primary ShMC connection fails, and at least one member reports to the slave ShMC that the primary ShMC connection fails, the slave ShMC immediately switches to a new primary ShMC, and the original primary ShMC becomes a lost member. The new primary ShMC generates a corresponding alarm, and the detailed process is as shown in Figure 9 .
[0127] (2) Slave ShMC failure
[0128] As shown in Figure 10 when the slave ShMC fails, the primary ShMC cannot receive the alarm synchronization message of the slave ShMC. Then, the primary ShMC selects a new slave ShMC according to the priority and the time stamp in the stored member information table, and notifies all members. Finally, the primary ShMC sends the stored member record to the slave ShMC, and the original slave ShMC becomes a lost member. The primary ShMC generates a corresponding alarm, and the detailed process is as shown in Figure 10 .
[0129] (3) Member failure
[0130] As shown in Figure 11 when the member fails, the primary ShMC cannot receive the alarm synchronization message of the member. The member becomes a lost member, and the primary ShMC generates a corresponding alarm, and the detailed process is as shown in Figure 11 .
[0131] (4) One link between chassis is interrupted
[0132] When the primary group leader does not receive the alarm synchronization information from the secondary group leader (or a group member) through port 0 (or port 1) for 3 consecutive periods (1 second interval), the primary group leader determines that the port 0 (or port 1) link connection of the secondary group leader (or a group member) is interrupted, and generates a corresponding alarm. As shown in Figure 12 When the link between switch 1 and the switching board 1 of chassis 2 is interrupted, the primary group leader triggers the ETH0 link interruption alarm to chassis 2; when the link between switch 1 and the switching board 1 of chassis 1 is interrupted, the primary group leader triggers the ETH0 link interruption alarm to chassis 2 and the ETH0 link interruption alarm to chassis 3.
[0133] (5) Two links between chassis are interrupted
[0134] When the two links of port 0 and port 1 of the primary group leader do not receive the alarm synchronization confirmation message from a group member (or a secondary group leader) for 3 consecutive periods (1 second interval), the primary group leader determines that the connection with the group member (or the secondary group leader) is interrupted, and generates a corresponding alarm. At the same time, the primary group leader initiates a role negotiation request message to the lost group member (or the secondary group leader) every 10 seconds.
[0135] Similarly, when a group member (or a secondary group leader) does not receive the alarm synchronization confirmation message returned by the primary group leader through the two links of port 0 and port 1 for 3 consecutive periods (1 second interval), it is determined that the connection with the primary group leader is interrupted, and it is set as the primary group leader role. Then it initiates a role negotiation request message to the master ShMC of other chassis through port 0 and port 1 every 10 seconds until the link is restored, and competes for the primary group leader role again.
[0136] Therefore, when two links between chassis are interrupted in a multi-chassis configuration environment, two or even multiple primary group leader ShMCs will temporarily appear. After the link is restored, these primary group leader ShMCs will automatically initiate a role negotiation request message to the master ShMC of other chassis according to the timestamp and priority, and finally form a multi-chassis connection network with one primary group leader ShMC, one secondary group leader ShMC, and other group member ShMCs.
[0137] As shown in Figure 13As shown, when the links of chassis 2 to the two switches are interrupted, the primary master will trigger an alarm of the connection interruption of chassis 2, and the master ShMC of chassis 2 automatically switches from the member role to the primary master role. At this time, the master ShMC of chassis 1 and the ShMC of chassis 2 are both primary master roles, but the connection between them is interrupted, so no role confusion will be caused. When the links of chassis 2 to the two switches are restored, according to the time stamp and priority, the master ShMC of chassis 1 will automatically initiate a role negotiation request message to the ShMC of chassis 2, and the master ShMC of chassis 2 will automatically initiate a role negotiation request message to the ShMCs of chassis 1 and chassis 3, and finally form a multi-chassis connection network with one primary master ShMC, one deputy master ShMC, and the others as member ShMCs.
[0138] It should be noted that although several modules of the multi-chassis health management system based on the advanced telecommunications computing architecture are mentioned in the foregoing detailed description, such division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided into several modules.
[0139] The multi-chassis health management system based on ATCA of the present application focuses on solving common problems in role competition and switching, information synchronization, cross-chassis access permission, fault recovery, etc. of multiple chassis, and ensures the cooperation, safety, efficiency and redundancy of multiple chassis. The highlights are as follows:
[0140] 1. Compete for the primary master ShMC, deputy master ShMC and member ShMC of the multi-chassis health management system according to the time stamp of the ShMC entering the master state and the set priority, and ensure that the primary master ShMC and the deputy master ShMC are located in two chassis.
[0141] 2. The alarm information of all chassis is automatically synchronized to the primary master ShMC, so that the primary master ShMC can timely monitor and handle the alarm of each chassis, and ensure real-time monitoring of the entire system.
[0142] 3. The primary master ShMC and the deputy master ShMC have different permissions when accessing across chassis: the primary master has read and write permissions when accessing across chassis, while the deputy master only has read permission when accessing across chassis, thereby ensuring the access security of the entire system.
[0143] 4. Automatically detect and discover internal chassis faults and inter-chassis connection faults, and automatically and quickly switch the primary master ShMC from the deputy master ShMC, thereby ensuring the continuity of the management and control of the entire system.
[0144] The English names and abbreviations involved in the present application are explained as follows in Table 1:
[0145] Table 1
[0146]
[0147] While the principles and uses of the present application have been described in particular embodiments, it will be appreciated that those skilled in the art will readily devise modifications to the implementations and techniques described herein without departing from the spirit and scope of the application. Numerous other changes can be made which will readily suggest themselves to those skilled in the art and which are encompassed in the spirit of the application disclosed and are defined in the appended claims.
[0148] The scope of the protection of this application is set forth in the appended claims. Those skilled in the art will recognize that the claims are not limited to the specific embodiments described herein and that various modifications can be made without departing from the scope of the claims.
Claims
1. A multi-chassis health management system based on an advanced telecom computing architecture, characterized in that, The system comprises: The main ShMC of each chassis competes for the roles of the main ShMC, the deputy group leader ShMC and the group member ShMC according to the time stamp when the main ShMC enters the main working state and the set priority; The main ShMC of each chassis automatically synchronizes the alarm information of the chassis to the main ShMC; The main ShMC has the permission to perform all read and write operations on each chassis; the deputy group leader ShMC has the permission to perform read operations on each chassis and write operations on the chassis; and the group member ShMC has the permission to perform all read and write operations on the chassis; When the ShMC of the multi-chassis fails or the link between the chassis is interrupted, the fault recovery process is automatically performed.
2. The multi-chassis health management system based on advanced telecom computing architecture of claim 1, wherein, The main ShMC of each chassis competes for the roles of the main ShMC, the deputy group leader ShMC and the group member ShMC according to the time stamp when the main ShMC enters the main working state and the set priority, comprising: When the set priorities of the main ShMC of each chassis are different, the main ShMC with the highest priority assumes the role of the main ShMC, the main ShMC with the second highest priority assumes the role of the deputy group leader ShMC, and the other main ShMC assumes the role of the group member ShMC; When the set priorities of the main ShMC of each chassis are the same, the main ShMC competes for the roles according to the time stamp when the main ShMC enters the main working state, the main ShMC with the earliest time stamp assumes the role of the main ShMC, the main ShMC with the second earliest time stamp assumes the role of the deputy group leader ShMC, and the other main ShMC assumes the role of the group member ShMC.
3. The multi-chassis health management system based on advanced telecom computing architecture of claim 2, wherein, When the ShMC of the chassis starts or switches, the time stamp when the ShMC enters the main working state is automatically recorded, and the main ShMC uses a uniformly allocated floating IP address to communicate externally; When configured as a multi-chassis, the ShMC of each chassis stores the floating IP addresses of the main ShMC of all the chassis; After the main ShMC starts, the role of the main ShMC is automatically initialized as the group member ShMC, and the main ShMC actively sends a role negotiation request message to the floating IP addresses of the main ShMC of all the chassis using the floating IP address of the main ShMC; the role negotiation request message carries the time stamp when the main ShMC enters the main working state, the set priority and the floating IP address of the main ShMC; After the main ShMC receives the role negotiation request message, the role of the main ShMC is finally determined according to the priority and the time stamp of the other party and the internal record of the main ShMC, and the main ShMC broadcasts a notification to the main ShMC of each chassis; After the deputy group leader ShMC and the group member ShMC receive the role negotiation request message, the message is directly ignored and no processing is performed.
4. The multi-chassis health management system based on advanced telecom computing architecture of claim 3, wherein, When configured as a multi-chassis, after the main ShMC of each chassis starts, if the main ShMC does not receive a response for three consecutive times of sending the role negotiation request message, the role of the main ShMC is automatically changed to the role of the main ShMC, and then the main ShMC initiates a role negotiation request message at a time interval; after the main ShMC discovers that there is still a group member ShMC that has not joined the group according to the floating IP address list of the group member ShMC, the main ShMC initiates a role negotiation request message to the group member ShMC at a time interval.
5. The multi-chassis health management system based on advanced telecom computing architecture of claim 3, wherein, The primary group leader ShMC and the deputy group leader ShMC have fixed primary group leader floating IP address and deputy group leader floating IP address respectively, when the primary group leader and the deputy group leader change, the fixed floating IP address is automatically bound to the current primary group leader ShMC and deputy group leader ShMC, and the current primary group leader ShMC and deputy group leader ShMC can be accessed by accessing the two fixed floating IP addresses from outside, thereby realizing multi-cabinet access management.
6. The multi-chassis health management system based on advanced telecom computing architecture of claim 1, wherein, When the ShMC under the multi-cabinet fails, automatic fault recovery processing is performed, including: When the primary group leader ShMC fails, the deputy group leader ShMC and the member ShMC send alarm synchronization messages to the primary group leader ShMC and cannot receive responses; when the deputy group leader ShMC finds that the primary group leader ShMC connection fails, and at least one member ShMC reports the primary group leader ShMC connection failure to the deputy group leader ShMC, the deputy group leader ShMC immediately switches to a new primary group leader ShMC, the original primary group leader ShMC becomes a lost member, and the new primary group leader ShMC generates corresponding alarms; When the deputy group leader ShMC fails, the primary group leader ShMC cannot receive the alarm synchronization messages of the deputy group leader ShMC, then the primary group leader ShMC selects a new deputy group leader ShMC according to the priority and the time stamp in the stored member information table, and notifies all members, finally sends the stored member record to the deputy group leader ShMC, the original deputy group leader ShMC becomes a lost member, and the primary group leader ShMC generates corresponding alarms; When the member ShMC fails, the primary group leader ShMC cannot receive the alarm synchronization messages of the member ShMC; the member ShMC becomes a lost member, and the primary group leader ShMC generates corresponding alarms.
7. The multi-chassis health management system based on advanced telecom computing architecture of claim 1, wherein, When the link between the cabinets under the multi-cabinet is interrupted, automatic fault recovery processing is performed, including: When the network port 0 or the network port 1 of the primary group leader ShMC does not receive the alarm synchronization information sent by the deputy group leader ShMC or a certain member ShMC for three consecutive periods, the primary group leader ShMC judges that the network port 0 or the network port 1 of the deputy group leader ShMC or the certain member ShMC is interrupted, and generates corresponding alarms; When the network port 0 and the network port 1 of the primary group leader ShMC do not receive the alarm synchronization confirmation messages of a certain member ShMC or the deputy group leader ShMC for three consecutive periods, the primary group leader ShMC judges that the connection with the member ShMC or the deputy group leader ShMC is interrupted, generates corresponding alarms, and simultaneously, the primary group leader ShMC initiates a role negotiation request message to the lost member ShMC or the deputy group leader ShMC every certain period of time; When the network port 0 and the network port 1 of a certain member ShMC or the deputy group leader ShMC do not receive the alarm synchronization confirmation messages returned by the primary group leader ShMC for three consecutive periods, it is judged that the connection with the primary group leader ShMC is interrupted, the self is set as the primary group leader ShMC role, and then the network port 0 and the network port 1 are used to initiate a role negotiation request message to the standby ShMC of the other cabinet every certain period of time until the link is restored, and the primary group leader ShMC role is re-competited.