A fault location and processing method for relay protection system based on dual-receive and dual-transmit CAN bus

By adding an ECU to the relay protection system and utilizing logical node and repeater technology, the difficulty of locating faults in the dual-receive, dual-transmit redundant CAN bus is resolved, enabling rapid fault location and processing, ensuring the normal operation of the system during faults and improving the reliability and stability of the power system.

CN119583330BActive Publication Date: 2025-09-26BEIJING INFORMATION SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411799157.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-09-26
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

The existing dual-receive, dual-transmit redundant CAN bus technology is difficult to quickly locate the fault location and type when a fault occurs. When both buses are disconnected, the system cannot work properly, affecting the reliability of the relay protection device and the stability of the power system.

Method used

An ECU is added to the relay protection system and combined into logical node 1 and logical node 2 through four communication lines. The CAN bus status is monitored in real time. The ECU is used as a detector and repeater to quickly locate and handle faults in the event of a fault, including relay state transition, heartbeat message detection, and fault type identification.

Benefits of technology

It achieves fast and accurate fault location and processing, ensures that the system continues to operate normally in the event of a fault, reduces the impact of the fault on the system, and improves the reliability of the relay protection device and the safety of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119583330B_ABST
    Figure CN119583330B_ABST
Patent Text Reader

Abstract

A method for locating and handling faults in a relay protection system based on a dual-receive and dual-transmit CAN bus belongs to the technical field of power system automation. The present invention solves the problem that when a CAN bus fails, the existing technology is difficult to quickly determine the location and type of the fault, and the system cannot continue to operate normally when both CAN buses are broken. The present invention can quickly locate the specific location of the system fault, so as to quickly and accurately find the fault point, set corresponding handling plans for different types of faults, and minimize the impact of the fault on the system. The handling plans include timely switching to a backup bus, automatically shielding the faulty node, and adding an ECU as a detector and repeater to handle special situations in a targeted manner. Through such a strategy, the method of the present invention can make a quick and accurate response when a fault occurs, thereby ensuring the reliability of the relay protection device and the safe operation of the power system. The method of the present invention can be applied to relay protection system fault handling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system automation, and in particular relates to a fault locating and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus. Background Art

[0002] Relay protection devices are a crucial piece of equipment in power systems, used to monitor, detect, and protect equipment and lines within them to ensure their safe operation. Their primary function is to rapidly isolate or limit the fault point when an anomaly or fault occurs in the power system, protecting the system's equipment and personnel from the effects of the power failure. Relay protection devices can be applied to a variety of power system equipment and lines, including generators, transformers, switchgear, and lines. Their design and configuration must be tailored to the specific power system requirements and equipment characteristics to ensure system reliability, safety, and stability. Relay protection devices play a crucial role in power systems, effectively protecting the normal operation of power equipment and systems while also providing maintenance personnel with a crucial basis for fault diagnosis and troubleshooting.

[0003] The CAN bus is a communications protocol and system architecture used for communication in the automotive and industrial sectors. Originally developed by Bosch in Germany in the 1980s to address the communication needs between automotive electronic systems, the CAN bus traditionally uses hardwiring for communication and data transmission. However, with the increasing complexity and scale of power systems, this traditional communication approach faces challenges, such as complex wiring and difficulty troubleshooting connection faults. To address these issues and improve the reliability and safety of relay protection devices, dual CAN redundant bus systems have been introduced. Dual CAN redundant bus systems are based on two independent CAN buses and utilize redundant communication methods for data transmission. This means that each critical system is connected to two independent CAN buses, transmitting the same data in parallel to improve system reliability and fault tolerance. As a reliable, efficient, and secure communication method, dual CAN redundant bus systems are widely used in relay protection devices, improving the reliability, stability, and safety of power systems and making significant contributions to the development of the power industry.

[0004] The implementation methods of dual-redundant CAN bus are classified according to the working status of the two buses. There are two main methods to implement dual-redundant CAN bus system. One is the active-standby bus, whose principle is: one bus is working and the other bus is on standby. When the working bus fails, it switches to the standby bus; the other is the dual-receive dual-transmit bus, whose principle is: the two buses work at the same time. When the device sends data, it sends the same data to both buses at the same time. When receiving data, it will receive the same two data from the two buses, and you can select one of them.

[0005] The master-slave bus and the dual-receive dual-transmit bus each have their own advantages and disadvantages. Their main features are summarized as follows:

[0006] (1) Characteristics of the master-slave bus: The master-slave bus is only started when the master node or communication channel fails, so redundant backup can be achieved with fewer redundant nodes and communication channels. Compared with the dual-receive and dual-transmit bus, the master-slave bus is simpler to implement and manage, and does not require real-time synchronization and switching mechanisms; however, the master-slave bus needs to be started after the master node or communication channel fails, so there will be a certain switching delay, which may cause data transmission interruption. Moreover, since the master-slave bus needs to wait for the master node or communication channel to fail before switching, there is a certain time window, and some data may be lost during this time. In the master-slave bus, the reliability of the master node is very important to the normal operation of the entire system. These shortcomings make the master-slave bus method unsuitable for systems with strict real-time requirements.

[0007] (2) Features of the dual-receive, dual-transmit bus: The dual-receive, dual-transmit bus can provide instant redundant backup and fault switching to maintain the continuity and reliability of the system. Since the redundant nodes and redundant communication channels are always active, fault switching can be performed with almost no interruption, with minimal impact on the system. The dual-receive, dual-transmit bus can ensure that the system always has spare nodes and communication channels. Even if the main node or communication channel fails, it can immediately switch to the redundant part; however, the dual-receive, dual-transmit bus requires additional redundant nodes and redundant communication channels, which increases hardware and system costs. The implementation and management of the dual-receive, dual-transmit bus is relatively complex, and it is necessary to ensure the synchronization and consistency of the redundant nodes and redundant communication channels with the main node and communication channel. Moreover, if the dual-receive, dual-transmit bus encounters a situation where both buses are disconnected, the system will not be able to work normally.

[0008] Due to the complexity of power systems and environmental uncertainties, relay protection equipment must possess high reliability and anti-interference capabilities. To this end, a dual-receive, dual-transmit redundant CAN bus is used in the communication systems of relay protection equipment. However, when a fault occurs in the dual-receive, dual-transmit redundant CAN bus system in a relay protection device, the following challenges may arise:

[0009] First, it's difficult to accurately locate faults in a timely manner. Due to the system's complexity and redundant design, when a bus fails, it's difficult to quickly determine the specific component or node that has the problem and take appropriate action. This can lead to delays in troubleshooting and repair, impacting the normal operation of relay protection devices and the reliability of the power system.

[0010] Secondly, when the dual-receive and dual-transmit CAN redundant bus system encounters a situation where both buses are broken, the system cannot continue to operate normally, which in turn affects the safety of power equipment and the stable operation of the power system.

[0011] In summary, the existing dual-receive, dual-transmit redundant CAN bus technology still has some defects. When a CAN bus fault occurs, it is difficult to quickly determine the fault location and fault type, so it is impossible to carry out targeted processing based on the fault of the CAN bus. Moreover, when the dual-receive, dual-transmit redundant CAN bus system encounters a situation where both buses are broken, the system cannot continue to operate normally, thereby limiting the performance and reliability of the relay protection equipment. Summary of the Invention

[0012] The purpose of the present invention is to solve the problem that when a CAN bus fault occurs, the existing dual-receive dual-transmit redundant CAN bus technology is difficult to quickly determine the fault location and fault type, and when the dual-receive dual-transmit redundant CAN bus system encounters a situation where both buses are broken, the system cannot continue to operate normally. A fault location and processing method for a relay protection system based on a dual-receive dual-transmit CAN bus is proposed.

[0013] The technical solution adopted by the present invention to solve the above technical problems is: a fault location and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus. The method is implemented by adding an ECU to the relay protection system, and the ECU has four communication lines. The two communication lines connecting the same end of the two CAN buses are combined as a logical node 1, and the two communication lines connecting the other ends of the two CAN buses are combined as a logical node 2.

[0014] Working cycle of the relay protection system

[0015] When the relay protection system is working normally, the ECU is used as a detector to monitor the messages on the two CAN buses in real time, and the working status of the two CAN buses is monitored according to the messages on the two CAN buses.

[0016] When the following conditions (1) or (2) are met, the ECU switches to the relay state, where:

[0017] Condition (1) is: both CAN buses are broken;

[0018] Condition (2) is: for any node, one CAN bus interface of the node is damaged and the bus to which the other CAN bus interface of the node is connected is disconnected;

[0019] In the relay state, the ECU uses logical nodes 1 and 2 to send messages received on one end of the CAN bus to the other end. The ECU continues to monitor the messages on the two CAN buses as a repeater until the circuit breaker is repaired, at which point the ECU switches to the circuit breaker detection state.

[0020] When the relay protection system reaches the cycle switching condition, or other faults other than conditions (1) and (2) occur, the relay protection system switches to the fault detection cycle;

[0021] Fault detection cycle in relay protection system

[0022] The master node in the relay protection system sends a heartbeat message to each slave node. After receiving the heartbeat message, each slave node detects the communication fault and sends a reply frame to the master node and ECU. The master node and ECU perform fault detection based on the slave node's reply.

[0023] If a fault is detected, the fault is processed according to the fault detection result, and the relay protection system switches to the working cycle;

[0024] If no fault is detected, the relay protection system switches to the working cycle.

[0025] Furthermore, the method for determining whether the two CAN buses are disconnected is as follows:

[0026] For a CAN bus, the ECU uses an internal timer to determine the time interval between adjacent messages on the CAN bus. If the time interval between adjacent messages exceeds a threshold or the ID numbers of the messages received by the two logical nodes of the ECU are different, it is considered that there is no circuit break on the CAN bus. Otherwise, a circuit break occurs on the CAN bus.

[0027] Similarly, determine the working status of another CAN bus.

[0028] Furthermore, the periodic switching condition of the relay protection system is:

[0029] The continuous working time of the relay protection system in the working cycle reaches the set threshold.

[0030] Furthermore, the specific process of performing fault detection based on the response of the slave node is as follows:

[0031] The following fault detection process is performed for each slave node:

[0032] Step 1: Send a test message from the node to the CAN bus;

[0033] If the test message from the node is sent successfully, proceed to step 2;

[0034] If the test message from the slave node is not sent successfully, go to step 3;

[0035] Step 2: Detect whether the self-transmission and self-reception are successful from the node;

[0036] If the slave node successfully receives and sends data, the slave node is in normal state.

[0037] If the slave node fails to transmit and receive messages by itself, it is determined whether the slave node has received messages from other slave nodes. If the slave node has not received messages from other slave nodes, the slave node has a reception failure. If the slave node has received messages from other slave nodes, the slave node has an unknown failure.

[0038] Step 3: Read the bus off flag from the node's controller register to see if it exists.

[0039] If the bus off flag is present, a CAN bus short circuit fault occurs;

[0040] If there is no bus off flag, go to step 4;

[0041] Step 4: The slave node determines whether it satisfies ERRC=3 or ERRBIT=25.

[0042] If ERRC=3 or ERRBIT=25, the CAN bus is disconnected or the plug is dropped;

[0043] If ERRC=3 or ERRBIT=25 is not satisfied, the slave node detects whether the self-transmission and self-reception are successful: if the self-transmission and self-reception are successful, an unknown fault occurs; if the self-transmission and self-reception are unsuccessful, execute step 5;

[0044] Step 5 determines whether the slave node has received messages from other slave nodes;

[0045] If the slave node does not receive messages from other slave nodes, the slave node has a sending and receiving failure. If the slave node receives messages from other slave nodes, the slave node has a sending failure.

[0046] Furthermore, the CAN bus disconnection or plug drop, CAN bus short circuit and unknown faults are permanent faults; node reception fault, node transmission fault and node transceiver fault are temporary faults.

[0047] Furthermore, the method for handling the permanent fault is:

[0048] If both CAN buses are broken, the fault is handled by switching the ECU to the relay state.

[0049] If one CAN bus interface of a node is damaged and the bus connected to the other CAN bus interface is disconnected, the fault is handled by switching the ECU to the relay state.

[0050] Remove nodes with unknown faults from the communication network;

[0051] For a CAN bus short circuit fault, it is necessary to stop the entire relay protection system for further fault handling.

[0052] Furthermore, the method for handling the temporary fault is:

[0053] The number of temporary failures and the continuous duration of each node are counted respectively. If the number of temporary failures and the continuous duration of a node exceed the set threshold, the node will be marked as a permanent failure and removed from the communication network.

[0054] The beneficial effects of the present invention are:

[0055] The fault location method of the present invention can quickly locate the specific location of the dual-bus system fault, so as to find the fault point quickly and accurately. The present invention sets corresponding processing solutions for different types of faults, which can minimize the impact of the fault on the system. The processing solutions include timely switching to the backup bus, automatically shielding the faulty node, and adding the ECU as a detector and repeater to carry out targeted processing of special situations that affect the entire system. Through such a strategy, the method of the present invention can make a fast and accurate response when a fault occurs, ensuring the reliability of the relay protection device and the safe operation of the power system. Because when both buses of the dual-receive and dual-transmit CAN redundant bus system are disconnected, the method of the present invention can convert the ECU to a relay state for communication, thereby ensuring that the system continues to operate normally. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a fault detection flow chart;

[0057] Figure 2 It is a schematic diagram of the circuit breaking process;

[0058] Figure 3 This is the ECU detection flow chart in working mode;

[0059] Figure 4 This is the ECU workflow diagram in fault detection mode;

[0060] Figure 5 It is a schematic diagram of ECU state transfer;

[0061] Figure 6 This is a schematic diagram of the internal structure of the ECU. DETAILED DESCRIPTION

[0062] Specific implementation method 1: Combination Figure 5 This embodiment describes a fault location and handling method for a relay protection system based on a dual-receive, dual-transmit CAN bus. The method is implemented by adding an ECU (electronic control unit) to the relay protection system. The ECU has four communication lines. Two communication lines connecting the same end of two CAN buses are combined as logical node 1, and two communication lines connecting the other ends of the two CAN buses are combined as logical node 2.

[0063] Working cycle of the relay protection system

[0064] When the relay protection system is working normally, the ECU is used as a detector to monitor the messages on the two CAN buses in real time, and the working status of the two CAN buses is monitored according to the messages on the two CAN buses.

[0065] When the following conditions (1) or (2) are met, the ECU switches to the relay state, where:

[0066] Condition (1) is: both CAN buses are broken;

[0067] Condition (2) is: for any node, one of the CAN bus interfaces of the node is damaged (i.e., cannot communicate normally), and the bus to which the other CAN bus interface of the node is connected is disconnected;

[0068] In the relay state, the ECU uses logical nodes 1 and 2 to send messages received on one end of the CAN bus to the other end. The ECU continues to monitor the messages on the two CAN buses as a repeater until the circuit breaker is repaired, at which point the ECU switches to the circuit breaker detection state.

[0069] When the relay protection system reaches the cycle switching condition, or other faults other than conditions (1) and (2) occur, the relay protection system switches to the fault detection cycle;

[0070] Fault detection cycle in relay protection system

[0071] The master node in the relay protection system sends a heartbeat message to each slave node. After receiving the heartbeat message, each slave node detects the communication fault and sends a reply frame to the master node and ECU. The master node and ECU perform fault detection based on the slave node's reply.

[0072] If a fault is detected, the fault is processed according to the fault detection result, and the relay protection system switches to the working cycle;

[0073] If no fault is detected, the relay protection system switches to the working cycle.

[0074] like Figure 2 As shown, both bus CANA and bus CANB have two ends, wherein one end of bus CANA is adjacent to one end of bus CANB, and the other end of bus CANA is adjacent to the other end of bus CANB, as shown in FIG. Figure 6 As shown, the ECU has four communication lines. To maintain the redundancy of the dual CAN buses, the ECU is logically divided into two nodes. Specifically, the two communication lines connecting buses A and B on one end of the bus (CAN controller 0 and CAN controller 1) are combined into logical node 1, and the two communication lines connecting buses A and B on the other end of the bus (CAN controller 2 and CAN controller 3) are combined into logical node 2. This divides the ECU into two logical nodes, each distinguished by an ID. The application layer assigns two IDs as normal nodes, and the ID filter mask is designed so that both logical nodes can receive all messages transmitted by the node. The communication lines connecting the two logical nodes to different CAN buses are redundant. During the fault detection cycle, both logical nodes in the ECU send heartbeat messages according to the algorithm for detection. When the ECU acts as a repeater, logical node 1 sends messages to logical node 2, which then transmits the messages to the other end of the bus. Logical nodes 1 and 2 act as a bridge to ensure normal communication. A circuit breaker detection operation is also performed. Once the circuit breaker is repaired, the repeater state is exited, ensuring normal message transmission.

[0075] The present invention designs a dual-receive dual-transmit redundant CAN bus application layer protocol to ensure the normal operation of the system. This protocol plays a key role in the dual CAN redundant bus system and mainly realizes the following functions:

[0076] (1) The communication functions between the master node and each subsystem node are equal:

[0077] Except for messages related to system management and control, which need to be initiated by the master node, other messages (such as status, data, monitoring, etc.) can be actively sent by each node without the need for the master node to perform communication scheduling, thereby achieving flexible data interaction between the master node and each subsystem node, and between subsystem nodes.

[0078] (2)Point-to-point communication function:

[0079] By setting the message identifier, point-to-point communication between the sending node and a certain node is realized. In order to improve the reliability of the transmitted data, the response-retransmission and message verification-retransmission methods are adopted.

[0080] (3) Broadcast communication function:

[0081] The master node can send broadcast messages to establish connections with each subsystem node simultaneously. This function is mainly used by the master node to send broadcast messages to set the common communication parameters of each node in the network. It can also periodically send broadcast messages to request status information from each subsystem node, thereby performing bus monitoring and identifying offline nodes.

[0082] (4) Response-retransmission mode:

[0083] In point-to-point communication, each time the sending node sends a message frame, the receiving node returns a response message to the master node after receiving the message. If the sending node does not receive any response message within the agreed time, it is considered that the communication has failed and the message is resent.

[0084] (5) Automatic fault detection function:

[0085] Utilizing a data frame message defined by the communication protocol specifically for network monitoring, automatic detection of node failures is achieved. Failures are divided into two categories: "permanent failures" and "temporary failures."

[0086] (6) Fault handling function:

[0087] For "permanent faults," the node disconnects the corresponding communication channel for processing; for "temporary faults," the node does not process them. Regardless of the fault type, the system records the fault information. If "temporary faults" occur multiple times, they are classified as "permanent faults," and the master node stores all fault information.

[0088] (7) The master node has bus management functions:

[0089] It can perform automatic detection of two buses, including power-on detection and timing detection. It can manage online nodes and identify offline nodes based on the detection results; and in certain specific cases, it can send alarm information to the upper layer.

[0090] The protocol uses a scheduling method based on primary and secondary time frames. The primary time frame is the time period during which all periodic messages in the system are transmitted at least once. The secondary time frame is the period of the most frequently transmitted frame and is a portion of the primary time frame. Within the same secondary time frame, CAN messages are sent through multi-master competition, and messages between different secondary time frames are separated in time. By limiting the total number of messages sent by each node in each secondary time frame, nodes in the system can be guaranteed to successfully send their assigned messages. Aperiodic messages have no predetermined schedule and are event-based messages. The interval between consecutive transmissions of messages with the same ID is controlled based on the minimum refresh interval, ensuring that the interval between transmissions of messages with the same ID on the same bus is greater than the minimum refresh interval. On a dual redundant bus, if the interval between transmissions of a message with the same ID on both buses exceeds the minimum refresh interval, the later transmitted copy of the message is canceled to maintain message scheduling consistency on the dual redundant bus.

[0091] Specific embodiment 2: This embodiment differs from specific embodiment 1 in that the method for determining whether the two CAN buses are disconnected is as follows:

[0092] For a CAN bus, the ECU uses an internal timer to determine the time interval between adjacent messages on the CAN bus. If the time interval between adjacent messages exceeds a threshold or the ID numbers of the messages received by the two logical nodes of the ECU are different (because when a CAN bus is disconnected, each logical node can only receive messages from the node before the disconnection point on the bus, so when a CAN bus is disconnected, the ID numbers of the messages received by the two logical nodes are different), then it is considered that there is no disconnection on the CAN bus. Otherwise, there is a disconnection on the CAN bus.

[0093] Similarly, determine the working status of another CAN bus.

[0094] Other steps and parameters are the same as those in the first embodiment.

[0095] ECU state transfer is as follows Figure 5As shown, in the normal working cycle of the system, it starts in the disconnection detection state without setting ID filtering, reads all messages on the bus, monitors the messages on the two CAN buses in real time, sets the main time frame to the disconnection threshold, and if the time interval between adjacent message messages exceeds the disconnection threshold, the ECU's logic node considers that the bus is disconnected. It determines whether the bus is disconnected by timeout and whether the two logic nodes receive the same ID number, thereby preventing the two buses from suddenly disconnecting. If both buses are disconnected, or one CAN bus interface of the node is damaged and the bus connected to the other CAN bus interface is disconnected, the ECU will switch to the relay state, and the logic node 1 inside the ECU will receive the received The message is transferred to logical node 2, which serves as a communication bridge to ensure normal communication. In the fault detection cycle, each node in the system (including the two logical nodes of ECU) uses heartbeat messages to perform fault detection operations and sends the fault message to the bus. The ECU and the master node will read the fault message to make a fault judgment. If a non-open circuit fault occurs, the fault will be processed accordingly and then converted to a working cycle. If there is no fault, the ECU will follow the fault detection cycle to convert to a working cycle. If the two open circuit conditions mentioned above happen to occur in the fault detection cycle, the ECU will convert to a relay state. The open circuit detection function continues in the relay state. If the open circuit is repaired, it will convert to a working cycle.

[0096] like Figure 6 As shown, if a timeout is found, the ECU will determine which bus is broken based on the interface number (such as CAN controller 1 corresponds to bus A).

[0097] Specific embodiment three: This embodiment differs from specific embodiment one or two in that the periodic switching condition of the relay protection system is:

[0098] The continuous working time of the relay protection system in the working cycle reaches the set threshold (the specific value can be set according to actual conditions).

[0099] Other steps and parameters are the same as those in the first or second embodiment.

[0100] There are two triggering conditions for the relay protection system in the present invention to switch from the working cycle to the fault detection cycle. One is that the system switches to the fault detection cycle once after each fixed period of continuous operation. The other is that when a fault other than condition (1) and condition (2) occurs in the system, the system needs to switch to the fault detection cycle for detection.

[0101] Specific embodiment 4: This embodiment differs from any one of specific embodiments 1 to 3 in that the specific process of performing fault detection based on the response status of the slave node is as follows:

[0102] The following fault detection process is performed for each slave node:

[0103] Step 1: Send a test message from the node to the CAN bus;

[0104] If the test message from the node is sent successfully, proceed to step 2;

[0105] If the test message from the slave node is not sent successfully, go to step 3;

[0106] Step 2: Detect whether the self-transmission and self-reception are successful from the node;

[0107] If the slave node successfully receives and sends data, the slave node is in normal state.

[0108] If the slave node fails to transmit and receive messages by itself, it is determined whether the slave node has received messages from other slave nodes. If the slave node has not received messages from other slave nodes, the slave node has a reception failure. If the slave node has received messages from other slave nodes, the slave node has an unknown failure.

[0109] Step 3: Read the bus off flag from the node's controller register to see if it exists.

[0110] If the bus off flag is present, a CAN bus short circuit fault occurs;

[0111] If there is no bus off flag, go to step 4;

[0112] Step 4: The slave node determines whether it satisfies ERRC=3 or ERRBIT=25.

[0113] If ERRC=3 or ERRBIT=25, the CAN bus is disconnected or the plug is dropped; the master node and ECU both perform fault detection based on the slave node's response. If a CAN bus disconnect fault is detected, the ECU switches to the relay state.

[0114] If ERRC=3 or ERRBIT=25 is not satisfied, the slave node detects whether the self-transmission and self-reception are successful: if the self-transmission and self-reception are successful, an unknown fault occurs; if the self-transmission and self-reception are unsuccessful, execute step 5;

[0115] Step 5 determines whether the slave node has received messages from other slave nodes;

[0116] If the slave node does not receive messages from other slave nodes, the slave node has a sending and receiving failure. If the slave node receives messages from other slave nodes, the slave node has a sending failure.

[0117] The other steps and parameters are the same as those in the first to third embodiments.

[0118] A slave node sends test messages to both the CANA and CANB buses, and the fault detection process for each bus follows the method of this embodiment. When a slave node detects a fault on one bus or a fault on the slave node itself, it can send the fault information to the master node via the other, functioning bus. If the master node fails to receive messages from a slave node on both buses but can normally receive messages from most other nodes, the slave node itself is likely faulty and needs to be investigated. Otherwise, the bus itself should be investigated first.

[0119] In order to be able to promptly detect and eliminate communication failures, the redundant CAN bus network must have a complete fault detection mechanism. The present invention uses a data frame message specifically for network monitoring, called a heartbeat message. The frame format of the heartbeat message is shown in Table 1.

[0120] Table 1 Heartbeat message frame format

[0121]

[0122] During the fault detection cycle, the master node sends a fault detection message to all nodes. Each node then detects the communication fault and responds to the master node. The master node obtains the fault detection results based on the responses of each node and stores the fault information. After the fault detection cycle is complete, the system continues to enter the working cycle.

[0123] The process of fault detection algorithm is as follows Figure 1 As shown, ERRC=3 indicates CAN_ERR_CRTL, which means a controller problem. This indicates that there may be a problem inside the controller and further inspection is required. ERRBIT=25 indicates CAN_ERR_CRTL_RX_WARNING, which means that the receive error count has reached the warning level. This indicates that attention should be paid to the error status on the CAN bus and appropriate measures should be taken to prevent further errors.

[0124] Specific embodiment five: This embodiment differs from any one of specific embodiments one to four in that the CAN bus disconnection or plug drop, CAN bus short circuit and unknown fault are permanent faults; node reception fault, node transmission fault and node transceiver fault are temporary faults.

[0125] The other steps and parameters are the same as those in the first to fourth embodiments.

[0126] When a node detects a fault, it uses information such as transmit and receive error counters, bus error status, and error interrupts provided by the CAN controller to categorize the fault type into two types: "permanent fault" and "temporary fault." A "permanent fault" refers to a fault that the node cannot recover from, such as a hardware failure or a permanent communication interruption. A "temporary fault" refers to a temporary problem at the node, such as a temporary communication interruption or a transient fault condition.

[0127] Specific implementation method six: combination Figure 3 and Figure 4 This embodiment differs from any one of the first to fifth embodiments in that the method for handling the permanent fault is as follows:

[0128] If both CAN buses are broken, the fault is handled by switching the ECU to the relay state.

[0129] If one CAN bus interface of a node is damaged and the bus connected to the other CAN bus interface is disconnected, the fault is handled by switching the ECU to the relay state.

[0130] Remove nodes with unknown faults from the communication network;

[0131] For a CAN bus short circuit fault, it is necessary to stop the entire relay protection system for further fault handling.

[0132] The other steps and parameters are the same as those in the first to fifth embodiments.

[0133] By switching the relay status of the ECU, it is ensured that the system can continue to operate normally without causing functional failure due to communication interruption.

[0134] Specific embodiment 7: This embodiment differs from any one of specific embodiments 1 to 6 in that the method for handling the temporary fault is:

[0135] The number of temporary failures and the continuous duration of each node are counted respectively. If the number of temporary failures and the continuous duration of a node exceed the set threshold, the node will be marked as a permanent failure and removed from the communication network.

[0136] The other steps and parameters are the same as those in the first to sixth embodiments.

[0137] Since temporary faults are usually temporary and may be caused by environmental interference or communication fluctuations, the node is likely to resume normal operation on its own within a short period of time. Therefore, the system will record the number of occurrences and continuous duration of such faults and include them in the fault information for monitoring and analysis. Only when the set conditions are met will the system mark it as a permanent fault and take appropriate measures. All fault information will be stored by the master node for subsequent fault analysis and system maintenance. The fault information storage mechanism can help system administrators quickly locate and resolve faults, improving the maintainability and reliability of the system. For such permanent faults, the faulty node is directly removed from the communication network. This can prevent the faulty node from causing unnecessary interference to the entire system, ensure the normal operation of other nodes, and maintain the stability and reliability of the entire network.

[0138] To achieve the above functions, the present invention designs relevant message formats and frame formats, as shown in Tables 2 and 3. Table 2 introduces the division of IDs and corresponding functional descriptions. Table 3 uses IDs to distinguish between fault detection frames and normal operation frames.

[0139] Table 2 Message frame format

[0140]

[0141]

[0142] Table 3 Frame type definition

[0143]

[0144] In Table 2, when ID.20 is 0, ID.19 is 1, and ID.18 is 1, the frame type is a fault detection frame, and all other cases are normal operation frames.

[0145] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; however, any obvious variations or modifications derived from the technical solution of the present invention remain within the scope of protection of the present invention.

Claims

1. A fault location and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus, characterized in that: The method is implemented by adding an ECU to the relay protection system, and the ECU has four communication lines. Two communication lines connecting the same end of two CAN buses are combined as a logical node 1, and two communication lines connecting the other ends of the two CAN buses are combined as a logical node 2. Working cycle of the relay protection system When the relay protection system is working normally, the ECU is used as a detector to monitor the messages on the two CAN buses in real time, and the working status of the two CAN buses is monitored according to the messages on the two CAN buses. When the following conditions (1) or (2) are met, the ECU switches to the relay state, where: Condition (1) is: both CAN buses are broken; Condition (2) is: for any node, one CAN bus interface of the node is damaged and the bus to which the other CAN bus interface of the node is connected is disconnected; In the relay state, the ECU uses logical nodes 1 and 2 to send messages received on one end of the CAN bus to the other end. The ECU continues to monitor the messages on the two CAN buses as a repeater until the circuit breaker is repaired, at which point the ECU switches to the circuit breaker detection state. When the relay protection system reaches the cycle switching condition, or other faults other than conditions (1) and (2) occur, the relay protection system switches to the fault detection cycle; Fault detection cycle in relay protection system The master node in the relay protection system sends a heartbeat message to each slave node. After receiving the heartbeat message, each slave node detects the communication fault and sends a reply frame to the master node and ECU. The master node and ECU perform fault detection based on the slave node's reply. If a fault is detected, the fault is processed according to the fault detection result, and the relay protection system switches to the working cycle; If no fault is detected, the relay protection system switches to the working cycle.

2. A fault location and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus according to claim 1, characterized in that: The specific method for determining whether the two CAN buses are disconnected is as follows: For a CAN bus, the ECU uses an internal timer to determine the time interval between adjacent messages on the CAN bus. If the time interval between adjacent messages exceeds a threshold or the ID numbers of the messages received by the two logical nodes of the ECU are different, it is considered that there is no circuit break on the CAN bus. Otherwise, a circuit break occurs on the CAN bus. Similarly, determine the working status of another CAN bus.

3. A fault location and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus according to claim 2, characterized in that: The periodic switching conditions of the relay protection system are: The continuous working time of the relay protection system in the working cycle reaches the set threshold.

4. A fault location and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus according to claim 3, characterized in that: The specific process of performing fault detection based on the response of the slave node is as follows: The following fault detection process is performed for each slave node: Step 1: Send a test message from the node to the CAN bus; If the test message from the node is sent successfully, proceed to step 2; If the test message from the slave node is not sent successfully, go to step 3; Step 2: Detect whether the self-transmission and self-reception are successful from the node; If the slave node successfully receives and sends data, the slave node is in normal state. If the slave node fails to transmit and receive messages by itself, it is determined whether the slave node has received messages from other slave nodes. If the slave node does not receive messages from other slave nodes, the slave node has a reception failure. If a slave node receives a message from another slave node, the slave node has an unknown fault; Step 3: Read the bus off flag from the node's controller register to see if it exists. If the bus off flag is present, a CAN bus short circuit fault occurs; If there is no bus off flag, go to step 4; Step 4: The slave node determines whether it satisfies ERRC=3 or ERRBIT=25. If ERRC=3 or ERRBIT=25, the CAN bus is disconnected or the plug is dropped; If ERRC=3 or ERRBIT=25 is not satisfied, the slave node detects whether the self-transmission and self-reception are successful: if the self-transmission and self-reception are successful, an unknown fault occurs; if the self-transmission and self-reception are unsuccessful, execute step 5; Step 5 determines whether the slave node has received messages from other slave nodes; If the slave node does not receive messages from other slave nodes, the slave node has a sending and receiving failure. If the slave node receives messages from other slave nodes, the slave node has a sending failure.

5. The method for fault location and processing of a relay protection system based on a dual-receive and dual-transmit CAN bus according to claim 4, characterized in that: The CAN bus disconnection or plug falling, CAN bus short circuit and unknown faults are permanent faults; node reception fault, node transmission fault and node transceiver fault are temporary faults.

6. A fault location and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus according to claim 5, characterized in that: The method for handling the permanent fault is as follows: If both CAN buses are broken, the fault is handled by switching the ECU to the relay state. If one CAN bus interface of a node is damaged and the bus connected to the other CAN bus interface is disconnected, the fault is handled by switching the ECU to the relay state. Remove nodes with unknown faults from the communication network; For a CAN bus short circuit fault, it is necessary to stop the entire relay protection system for further fault handling.

7. A fault location and processing method for a relay protection system based on a dual-receive and dual-transmit CAN bus according to claim 6, characterized in that: The method for handling the temporary fault is: The number of temporary failures and the continuous duration of each node are counted respectively. If the number of temporary failures and the continuous duration of a node exceed the set threshold, the node will be marked as a permanent failure and removed from the communication network.

Citation Information

Patent Citations

  • Bus test system and method

    CN107491055A

  • Dual open CAN bus wire fault identification

    CN118175077A