A method for fault detection and handling in RapidIO switching networks
By employing a dual-redundant system management node and a primary/backup switching strategy with an independent bus, combined with hardware reset and health status assessment, the problem of fault propagation in RapidIO switching networks is solved, achieving highly reliable fault detection and handling.
Patent Information
- Application Number
- CN202211613068.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-12-15
AI Technical Summary
RapidIO switching networks suffer from complex system topology and random communication conflicts, leading to fault propagation and deterioration that existing technologies struggle to effectively address.
The system employs a dual-redundant RapidIO system management node and an independent system management bus. Through a master-slave switching strategy and hardware reset signal, combined with system health status assessment and scoring strategies, it achieves fault detection and handling.
It improves the reliability of fault detection in RapidIO switching networks, avoids the impact on business data transmission, provides system health management strategies, and promptly prevents fault propagation and deterioration.
Smart Images

Figure CN116260701B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of embedded computer technology, and in particular to a method for fault detection and handling of RapidIO switching networks. Background Technology
[0002] RapidIO switching networks are widely used in embedded computing due to their open and flexible network structure, high system performance and transmission efficiency, and strong scalability. However, the complex system topology of RapidIO switching networks leads to uneven traffic distribution between nodes and unpredictable transmission timing, resulting in random communication collisions. While RapidIO switching chips possess some error detection and recovery capabilities, when the faults caused by communication collisions exceed the RapidIO switch's processing capacity, the fault propagates and deteriorates, leading to short-term network paralysis or even irrecoverable failure. Summary of the Invention
[0003] In view of this, this application provides a fault detection and handling method for RapidIO switching networks, which solves the problems in the prior art and effectively solves the spread and deterioration of faults in RapidIO switching networks.
[0004] The fault detection and handling method for RapidIO switching networks provided in this application adopts the following technical solution:
[0005] A method for fault detection and handling of a RapidIO switching network includes several RapidIO service nodes, a dual-redundant RapidIO system management node, and a RapidIO switch. The RapidIO system management node is connected to the RapidIO switch and the RapidIO service nodes through an independent RapidIO system management bus interface.
[0006] The two RapidIO system management nodes use a primary or backup switching strategy to determine the primary node. The primary and backup nodes are determined at power-up and are periodically checked during operation. After the RapidIO system management primary node is determined, it is responsible for detecting and handling RapidIO switching network faults.
[0007] Optionally, the RapidIO system management node connects to the reset interface of the RapidIO switch and the RapidIO service node via a hard reset signal. The RapidIO system management node resets the hardware resources of the RapidIO service node and the RapidIO switch, and then re-initializes the system.
[0008] Optionally, both RapidIO system management nodes read the register verification information in the RapidIO switch through the system management bus and send the read information to each other. The two parties determine the primary and backup relationship through cross-checking and use hot backup to achieve dual-redundancy system management.
[0009] Optionally, the RapidIO system management node pre-records the correct operating status and historical operating status of the system. After the system is powered on or periodically, it reads the RapidIO switch register through the independent bus interface of the RapidIO system management to check the RapidIO switch status. The RapidIO switch status includes Link status, error status, retransmission status, flow status and bit error rate. By comparing the read RapidIO switch register value with the pre-recorded register value when the system is running correctly, the current operating status of the system is determined.
[0010] Each time a RapidIO switch register value is read, it is recorded to form a system log for subsequent fault analysis and handling.
[0011] Optionally, a scoring-based system health status assessment strategy is used to evaluate the deterioration of system health. When the RapidIO system management node detects a system anomaly, it collects and classifies the anomaly information. Based on the collected system fault information, a total system fault score is calculated, and the faults are divided into the following five levels according to the calculated total fault score:
[0012] Level 5 fault: Total fault score 0 < Ex ≤ T1, at which point a node transmission error occurs;
[0013] Level 4 fault: Total fault score T1 < Ex ≤ T2, at which point the node traffic is too high or the short-term bit error rate is too high.
[0014] Level 3 fault: Total fault score T2 < Ex ≤ T3, at which point the network condition deteriorates, node retransmission occurs, and recovery is impossible;
[0015] Level 2 fault: Total fault score T3 < Ex ≤ T4, indicating a degradation in node link connectivity and a decrease in transmission bandwidth;
[0016] Level 1 fault: Total fault score T4 < Ex ≤ T5, the node link connection is broken, and the corresponding node cannot communicate via the RapidIO bus.
[0017] Optionally, for a level 4 fault, the corresponding node application layer software is notified to slow down transmission; for a level 3 fault, the software process using the affected memory space is killed and the memory space is removed from the system resource pool; for a level 2 fault, the working state is saved, the faulty node is soft reset, and the working state is restored after the reset; for a level 1 fault, the working state is saved, the faulty node is hard reset, and the working state is restored after the reset.
[0018] In summary, this application includes the following beneficial technical effects:
[0019] (1) By using a dedicated RapidIO system management node and a dedicated RapidIO system management bus, the transmission and processing of system business data can be avoided from affecting the system management fault detection and processing functions.
[0020] (2) The adoption of a "master-backup" RapidIO system management node improves the reliability of RapidIO switching network fault detection and processing;
[0021] (3) It provides a system health management strategy based on system log function, which provides a basis for comprehensively and systematically assessing the system health status;
[0022] (4) Provide a scoring-based system health status assessment strategy and adopt a graded fault handling mechanism according to the deterioration of system health. It can adopt corresponding handling plans at different stages of system fault development to prevent the spread and deterioration of faults in a timely manner. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a block diagram illustrating the fault detection and handling principle of the RapidIO switching network in this application.
[0025] Figure 2 This is a structural block diagram of a specific embodiment of this application;
[0026] Figure 3 This application presents a functional block diagram for the RapidIO switching network fault detection and handling design. Detailed Implementation
[0027] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0028] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this application, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0030] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0031] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0032] This application provides a method for fault detection and handling in a RapidIO switching network.
[0033] like Figure 1 As shown, a fault detection and handling method for a RapidIO switching network includes several RapidIO service nodes, a dual-redundant RapidIO system management node, and a RapidIO switch. The RapidIO system management node is connected to the RapidIO switch and the RapidIO service nodes through an independent RapidIO system management bus interface.
[0034] The RapidIO switch is located at the center of the RapidIO switching network structure. RapidIO service nodes run system application programs and connect to the RapidIO switch via RapidIO interfaces. RapidIO service nodes communicate extensively with each other through the RapidIO switch. Two RapidIO system management nodes connect to the RapidIO switch via independent RapidIO system management bus interfaces (such as I2C, SPI, or RapidIO interfaces), and also connect to each RapidIO service node, responsible for fault detection and handling in the RapidIO switching network.
[0035] The two RapidIO system management nodes use a primary or backup switching strategy to determine the primary node. The primary and backup nodes are determined at power-up and are periodically checked during operation. After the RapidIO system management primary node is determined, it is responsible for detecting and handling RapidIO switching network faults.
[0036] Specifically, the two RapidIO system management nodes employ the following "master-slave" failover strategy: After system power-on, both RapidIO system management nodes read the verification information from the registers in the RapidIO switch via the system management bus and send the read information to each other. Then, they cross-check the read information against pre-stored correct information; the RapidIO system management node that matches correctly becomes the master node. If both are correct, any node is selected as the master node, and the other node becomes the backup node. During system operation, the two RapidIO system management nodes periodically perform the above information reading and cross-checking to promptly detect "failed" system management nodes and initiate a "master-slave" failover. If a faulty system management node is repaired, and both nodes have problems, they are repaired; if only one node has a problem, it is used normally, and the faulty node is repaired.
[0037] The RapidIO system management node connects to the reset interfaces of the RapidIO switch and RapidIO service nodes via a hard reset signal to reset the hardware resources of the RapidIO service nodes and RapidIO switch, and then re-initializes the system.
[0038] Both RapidIO system management nodes read the register verification information in the RapidIO switch through the system management bus and send the read information to each other. The two parties determine the primary and backup relationship through cross-checking and use hot backup to achieve dual-redundancy system management.
[0039] The RapidIO switching network comprises two buses: the RapidIO service bus and the dedicated system management bus. The RapidIO service bus is used for communication of service data streams between service nodes within the RapidIO system. The RapidIO system management bus is used by the RapidIO system management node for the management and maintenance of the entire RapidIO switching network. By isolating these two different buses, the RapidIO service data stream and the system management data stream are effectively separated.
[0040] RapidIO service nodes run system service programs. Each RapidIO service node has both a RapidIO service interface and a RapidIO system management interface, and these two interfaces exist independently. The service interface connects to the RapidIO service bus, and the RapidIO system management interface connects to the dedicated system management bus.
[0041] The RapidIO switch has one system management interface and several RapidIO service interfaces. The system management interface is connected to the RapidIO system management bus, through which the RapidIO system management node manages and maintains the RapidIO switch.
[0042] The RapidIO system management node connects to the RapidIO system management bus via the RapidIO system management interface to manage and maintain the RapidIO service nodes and RapidIO switches on the bus. It also connects to the RapidIO service nodes and RapidIO switches via a hard reset signal, allowing it to issue a hard reset signal when necessary to reset the hardware resources of the RapidIO service nodes and RapidIO switches, subsequently re-initializing the system.
[0043] The RapidIO system management node pre-records the correct operating status and historical operating status of the system. After the system is powered on or periodically, it reads the RapidIO switch register through the RapidIO system management independent bus interface to check the RapidIO switch status. The RapidIO switch status includes Link status, error status, retransmission status, flow status and bit error rate. By comparing the read RapidIO switch register values with the pre-recorded register values when the system is running correctly, the current operating status of the system is determined.
[0044] Each time a RapidIO switch register value is read, it is recorded to form a system log for subsequent fault analysis and handling.
[0045] The scoring-based system health status assessment strategy evaluates the deterioration of system health. When the RapidIO system management node detects a system anomaly, it collects and classifies the anomaly information. Based on the collected system fault information, it calculates the total system fault score and classifies the faults into the following five levels:
[0046] Level 5 fault: Total fault score 0 < Ex ≤ T1. At this time, node transmission error occurs, mostly due to channel interference.
[0047] Level 4 fault: The total fault score T1 < Ex ≤ T2. At this time, the node traffic is too high, and it has not yet caused retransmission or caused short-term retransmission; or the short-term bit error rate is too high, but it has not yet caused a fault.
[0048] Level 3 fault: The total fault score T2 < Ex ≤ T3. At this point, due to the deterioration of the network conditions, the node retransmits and cannot be recovered.
[0049] Level 2 fault: The total fault score T3 < Ex ≤ T4, indicating a degradation in node link connection, such as the connection line width decreasing from 4x to 1x, resulting in a decrease in transmission bandwidth.
[0050] Level 1 fault: The total fault score T4 < Ex ≤ T5 indicates that the node link connection is broken, and the corresponding node cannot communicate via the RapidIO bus.
[0051] Specifically, the RapidIO system management node reads the RapidIO switch registers and compares them with normal register values during system operation. When a system anomaly is detected, the node collects and categorizes the anomaly information. The statistical basis is shown in the table below.
[0052]
[0053] Based on the statistical system fault information, the total system fault score is calculated using the following formula:
[0054]
[0055] Faults are classified according to a fault score, and corresponding handling plans are implemented for different levels.
[0056] For level 5 faults, this type of error can generally be recovered automatically immediately without any processing. For level 4 faults, the corresponding node application layer software is notified to slow down transmission. For level 3 faults, the software process using the affected memory space is killed, and the memory space is removed from the system resource pool. For level 2 faults, the working state is saved, the faulty node is soft reset, and the working state is restored after the reset. For level 1 faults, the working state is saved, the faulty node is hard reset, and the working state is restored after the reset.
[0057] In one embodiment: such as Figure 2 As shown, the RapidIO switching network consists of two RapidIO system management nodes, three RapidIO service nodes, and one RapidIO switch. The RapidIO system management nodes are implemented using an STM32F407IGT7 processor, the RapidIO service nodes are implemented using an 80HCPS1848, and the RapidIO service nodes are implemented using a TMS320C6455 processor.
[0058] The 80HCPS1848 switching chip provides an 18-port, 48-wire RapidIO physical interface, supporting the RapidIO V2.1 specification. In a specific embodiment, it is configured with 4 module ports, each connected to the TMS320C6455 processor within the RapidIO processing unit. The 80HCPS1848 switching chip provides I2C as the RapidIO system management interface, connecting to the STM32F407IGT7 processor and the TMS320C6455 processor to enable fault detection and handling in the RapidIO switching network.
[0059] like Figure 3 As shown, the RapidIO system management interface of the STM32F407IGT7 processor is implemented via I2C. Two STM32F407IGT7 processors employ a "master-slave" switching strategy, determining the RapidIO system management master node through cross-checking verification information. The "master-slave" switching is performed periodically upon power-on or during system operation. If the master node fails during operation, it automatically switches to the standby node. During system operation, the RapidIO system management master node periodically reads the RapidIO switch registers through the independent RapidIO system management bus interface to check the RapidIO switch status, including Link status, error status, retransmission status, flow status, and bit error rate. The read register status is compared with pre-recorded register values from correctly running system operation, and system fault information is statistically analyzed. Faults are classified according to fault scoring rules, and corresponding handling plans are implemented based on the five-level fault classification.
[0060] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A fault detection and handling method for a RapidIO switching network, characterized in that, It includes several RapidIO service nodes, a dual-redundant RapidIO system management node, and a RapidIO switch. The RapidIO system management node is connected to the RapidIO switch and the RapidIO service nodes through an independent RapidIO system management bus interface. The two RapidIO system management nodes use a primary or backup switching strategy to determine the primary node. The primary and backup nodes are determined at power-up and are periodically checked during operation. The two RapidIO system management nodes employ the following master-slave failover strategy: After system power-on, both RapidIO system management nodes read the verification information from the registers in the RapidIO switch via the system management bus and send the read information to each other. Then, they cross-check the read information against pre-stored correct information. The RapidIO system management node that matches correctly becomes the master node. If both are correct, any node is selected as the master node, and the other node becomes the backup node. During system operation, the two RapidIO system management nodes periodically perform the above information reading and cross-checking to promptly detect any failed system management nodes and perform master-slave failover. After the RapidIO system management master node is determined, it is responsible for detecting and handling RapidIO switching network faults. The RapidIO system management node pre-records the correct operating status and historical operating status of the system. After the system is powered on or periodically, it reads the RapidIO switch register through the RapidIO system management independent bus interface to check the RapidIO switch status. The RapidIO switch status includes Link status, error status, retransmission status, flow status and bit error rate. By comparing the read RapidIO switch register values with the pre-recorded register values when the system is running correctly, the current operating status of the system is determined. Record the RapidIO switch register values read each time to form a system log for subsequent fault analysis and handling; The scoring-based system health status assessment strategy evaluates the deterioration of system health. When the RapidIO system management node detects a system anomaly, it collects and classifies the anomaly information. Based on the collected system fault information, it calculates the total system fault score and classifies the faults according to the fault score. Corresponding handling plans are implemented for different levels.
2. The fault detection and handling method for RapidIO switching networks according to claim 1, characterized in that, The RapidIO system management node connects to the reset interfaces of the RapidIO switch and RapidIO service nodes via a hard reset signal. The RapidIO system management node resets the hardware resources of the RapidIO service nodes and RapidIO switch, and then re-initializes the system.
3. The fault detection and handling method for RapidIO switching networks according to claim 1, characterized in that, Based on the calculated total fault score, the faults are divided into the following five levels: Level 5 fault: Total fault score 0 < Ex ≤ T1, at which point a node transmission error occurs; Level 4 fault: Total fault score T1 < Ex ≤ T2, at which point the node traffic is too high or the short-term bit error rate is too high. Level 3 fault: Total fault score T2 < Ex ≤ T3, at which point the network condition deteriorates, node retransmission occurs, and recovery is impossible; Level 2 fault: Total fault score T3 < Ex ≤ T4, indicating a degradation in node link connectivity and a decrease in transmission bandwidth; Level 1 fault: Total fault score T4 < Ex ≤ T5, the node link connection is broken, and the corresponding node cannot communicate via the RapidIO bus.
4. The fault detection and handling method for RapidIO switching networks according to claim 3, characterized in that, For a Level 4 fault, the corresponding node's application layer software is notified to slow down transmission. For a Level 3 fault, the software process using the affected memory space is killed, and the memory space is removed from the system resource pool. For a Level 2 fault, the working state is saved, the faulty node is soft reset, and the working state is restored after the reset. For a Level 1 fault, the working state is saved, the faulty node is hard reset, and the working state is restored after the reset.
Citation Information
Patent Citations
Method and device for selecting master control equipment and equipment linkage system
CN110687809A
Dual-computer hot standby deployment method and system based on KVM virtualization system
CN111078352A