Configuration-free fault-tolerant reconstruction and fault recording method for synchronous distributed collaborative network
By introducing an autonomous fault-tolerant reconfiguration mechanism into a synchronous distributed collaborative network and implementing link control protocol frames using FPGA, the problem of autonomous fault-tolerant reconfiguration when the port's working state is abnormal is solved, reducing development workload and improving fault location efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN SHUDAO AVIATION TECH CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to enable autonomous fault-tolerant reconfiguration of device ports when they malfunction, and the development workload is enormous.
By introducing an autonomous fault-tolerant reconfiguration mechanism into a synchronous distributed cooperative network, the link control protocol frame is reconfigured using FPGA. This includes generating autonomous fault-tolerant reconfiguration signals, configuring device port status, rebuilding network topology, and generating log files to locate faults.
When a port is in an abnormal working state, it can achieve autonomous fault-tolerant reconstruction of the device port, reduce development workload, improve real-time performance and fault location efficiency, and adapt to complex topology changes without repeated configuration.
Smart Images

Figure CN122053330A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed system technology, and in particular to a method for configuration-free fault-tolerant reconstruction and fault recording of synchronous distributed cooperative networks. Background Technology
[0002] Electrical and electronic architecture is a core element in the design of complex digital systems. Its core components include hardware-software interconnection, software-software interconnection, hardware interconnection, and hardware power supply (electrical architecture). For contemporary complex digital systems such as new energy vehicles, hardware interconnection involves bus technologies such as Ethernet and CAN. With the development of software decoupling and regional electrical and electronic architectures, the hardware architecture of complex digital systems exhibits characteristics of multi-device networking based on backbone networks / buses, "more physically distributed, more centrally managed information," and on-demand software deployment. In this context, the importance of maintaining communication between devices increases significantly. Redundancy mechanisms must be implemented to prevent link failures from causing overall system failure, thereby achieving functional safety. Synchronous distributed cooperative networks / buses are a new generation of strong real-time and deterministic networks. Their characteristic is that the entire system operates based on a pre-set step cycle. The system's global clock is generated by a single device, and all devices connected to the synchronous distributed cooperative network share a global clock based on "step clock + number of step cycles + event offset value," without needing to synchronize the absolute time of all local clocks within the system. For data forwarding, the synchronous distributed cooperative network differs from contemporary Ethernet technology. It adopts a communication strategy of "unified queuing for uplink and synchronous concurrency for downlink", which enables each port to acquire all the data uplinked to the network in each step cycle, but the device only retains the data it needs.
[0003] Existing technologies typically implement system redundancy and fault-tolerant reconfiguration in three separate schemes: topology redundancy: using a set of redundant physical network links, usually requiring physical A / B dual-network implementation and adjustments to the network configuration table to achieve redundancy; link redundancy: monitoring port status on devices and switching to another port when one port fails to achieve redundancy; and software redundancy: developing an independent redundancy mechanism based on the network in software, using a heartbeat packet management mechanism to achieve redundancy in the software part. However, existing systems require extensive development of redundancy mechanisms, leading to significant R&D pressure due to topology complexity. Furthermore, additional adaptation development is needed when the physical network changes, resulting in a huge workload. Because the types of link failures in existing distributed systems are complex and random, it is difficult to apply predefined fault-tolerant reconfiguration schemes to all scenarios when port operating states are abnormal, making it difficult to achieve autonomous fault-tolerant reconfiguration of device ports. Summary of the Invention
[0004] The purpose of this invention is to provide a method for configuration-free fault-tolerant reconstruction and fault recording of synchronous distributed collaborative networks, which solves the problems of existing technologies that make it difficult to achieve autonomous fault-tolerant reconstruction of device ports when the port working state is abnormal and the huge amount of development work required.
[0005] To address the aforementioned technical problems, the embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a method for configuration-free fault-tolerant reconfiguration and fault recording in a synchronous distributed cooperative network, comprising: Based on the periodic timeout count and continuously acquired synchronous distributed collaborative network data, determine whether a timeout event has occurred, and generate an autonomous fault-tolerant reconfiguration signal when a timeout event occurs. In response to the autonomous fault-tolerant reconfiguration signal, the system opens all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration. Based on the link control protocol frames transmitted back by the synchronous distributed cooperative network, the state of the device port is configured to complete the autonomous fault-tolerant reconfiguration process. The state of the device port includes uplink port, downlink port and blocking port. The autonomous fault-tolerant reconfiguration process is implemented based on the hardware logic in the FPGA. The system receives and parses the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process, extracts the protocol frame content, and restores the synchronous distributed cooperative network hierarchy based on the protocol frame content to rebuild the synchronous distributed cooperative network topology and generate log files. The log files are used to locate faults. The protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement.
[0006] A second aspect of the present invention provides a configuration-free fault-tolerant reconfiguration and fault recording system for synchronous distributed cooperative networks, comprising: The synchronous distributed cooperative network module, deployed in FPGA hardware, is used to determine whether a timeout event has occurred based on the periodic timeout count and continuously acquired synchronous distributed cooperative network data, and to generate an autonomous fault-tolerant reconfiguration signal when a timeout event occurs. In response to the autonomous fault-tolerant reconfiguration signal, it controls the opening of all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration. Based on the link control protocol frames returned by the synchronous distributed cooperative network, it configures the status of the device ports to complete the autonomous fault-tolerant reconfiguration process. The status of the device ports includes uplink ports, downlink ports, and blocked ports. Synchronous distributed cooperative soft bus module, FPGA connected, used to receive and transmit link control protocol frames returned by the synchronous distributed cooperative network; The fault recording module, connected to the synchronous distributed cooperative soft bus module, is used to receive and parse the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process, extract the protocol frame content, and restore the synchronous distributed cooperative network hierarchy based on the protocol frame content to rebuild the synchronous distributed cooperative network topology and generate log files. The log files are used to locate faults. The protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement.
[0007] Compared to existing technologies, this invention provides a configuration-free fault-tolerant reconfiguration and fault recording method for synchronous distributed cooperative networks. Based on the periodic timeout count and continuously acquired synchronous distributed cooperative network data, it determines whether a timeout event has occurred and generates an autonomous fault-tolerant reconfiguration signal when a timeout event occurs. In response to the autonomous fault-tolerant reconfiguration signal, it controls the opening of all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration. Based on the link control protocol frames returned by the synchronous distributed cooperative network, it configures the status of the device ports to complete the autonomous fault-tolerant reconfiguration process. The device port status includes uplink ports, downlink ports, and blocked ports. The autonomous fault-tolerant reconfiguration process is implemented based on hardware logic in an FPGA. It receives and parses the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process, extracts the protocol frame content, and reconstructs the synchronous distributed cooperative network hierarchy based on the protocol frame content to rebuild the synchronous distributed cooperative network topology and generate a log file. The log file is used to locate faults. The protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement. In this way, an autonomous fault-tolerant reconfiguration mechanism is added to the communication, and a link control protocol frame is added. The autonomous fault-tolerant reconfiguration process is implemented by FPGA. It can determine the working status of the device port based on the operating status of the synchronous distributed cooperative network. When the working status of the port is abnormal, the device port can be easily reconfigured through the synchronous distributed cooperative network. When the connection relationship between system nodes is relatively complex, the autonomous fault-tolerant reconfiguration of the device port can be achieved by the synchronous distributed cooperative network. The entire process does not require the R&D personnel to configure in advance, and there is no need to repeat the configuration when the physical connection is changed multiple times, which can significantly reduce the development workload. Attached Figure Description
[0008] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, with the same or corresponding reference numerals denoteing the same or corresponding parts, wherein: Figure 1 A flowchart illustrating a configuration-free fault-tolerant reconfiguration and fault recording method for synchronous distributed cooperative networks is shown. Figure 2 The flowchart illustrating the triggering of the autonomous fault-tolerant reconfiguration mechanism is shown schematically. Figure 3 A flowchart illustrating the process of completing autonomous fault-tolerant reconfiguration is shown schematically. Figure 4 A schematic diagram illustrating the process of reconstructing the topology of a synchronous distributed cooperative network and determining faults is provided. Figure 5 A schematic diagram of a synchronous distributed cooperative network configuration-free fault-tolerant reconfiguration and fault recording system is shown. Figure 6 A schematic diagram of a synchronous distributed cooperative network is shown. Figure 7 The overall system architecture diagram is shown schematically. Figure 8 The schematic diagram illustrates the FPGA network section of a synchronous distributed collaborative network computing device. Detailed Implementation
[0009] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the invention and to fully convey the scope of the invention to those skilled in the art.
[0010] It should be noted that, unless otherwise stated, the technical or scientific terms used in this invention should have the ordinary meaning as understood by one of ordinary skill in the art.
[0011] The methods described in the embodiments of the present invention will be explained in detail below.
[0012] Figure 1 A flowchart illustrating a configuration-free fault-tolerant reconfiguration and fault recording method for synchronous distributed cooperative networks according to an embodiment of the present invention is shown. See [link to relevant documentation]. Figure 1 As shown, the configuration-free fault-tolerant reconfiguration and fault recording method for synchronous distributed cooperative networks may include: S101. Based on the periodic timeout count and the continuously acquired synchronous distributed collaborative network data, determine whether a timeout event has occurred, and generate an autonomous fault-tolerant reconfiguration signal when a timeout event occurs.
[0013] The period timeout value ranges from 2 to 10. A period timeout value of 2 is preferred. The synchronous distributed network operates with an adjustable step period of 20 microseconds to 2 seconds. At the beginning of each step period, the port receives a scheduling frame. Therefore, if the port has not received a scheduling frame by the start of the next step period, it can be determined that the port is offline. When the period timeout value is 2, if the synchronous distributed network operates with a step period in the microsecond range, for example, a step period of 20 microseconds, the time for port fault detection is 40 microseconds over two step periods, achieving strong real-time port fault monitoring. Even if the step period value is in the millisecond range, such as 50ms, the time from detecting a port fault to port reconstruction is also in the millisecond range. The period timeout value can be set from 2 to 10. Values higher than 2 allow the system to tolerate occasional faults without triggering network reconstruction.
[0014] Specifically, based on the periodic timeout count and continuously acquired synchronous distributed collaborative network data, it is determined whether a timeout event has occurred, and an autonomous fault-tolerant reconfiguration signal is generated when a timeout event occurs, including: Step A1: Continuously extract the step count of the scheduling frame from the synchronous distributed cooperative network data.
[0015] Before step A1, synchronous distributed cooperative network data is continuously acquired from the uplink port.
[0016] Step A2: Reset the monitoring timer each time a step count is extracted.
[0017] Among them, the monitoring timer is a timer set based on the step period.
[0018] Each time the number of step beats is extracted, the timer of the step cycle monitoring module is reset to zero to continuously monitor the number of step beats.
[0019] Step A3: If the monitoring timer's count exceeds the step cycle and the next step count is not extracted, a step cycle timeout is determined to have occurred, and the step cycle monitoring count is incremented by one.
[0020] If the step count is received, the step cycle monitoring count is reset to zero.
[0021] Step A4: Count the number of consecutive step cycle monitoring. When the number of step cycle monitoring reaches the cycle timeout number, determine that a timeout event has occurred, and generate an autonomous fault-tolerant reconfiguration signal when a timeout event occurs.
[0022] In other words, if the number of step cycle monitoring is greater than or equal to the number of step cycle timeouts, the uplink port link is considered to be faulty, triggering the fault-tolerant reconstruction mechanism.
[0023] Specifically, Figure 2The flowchart illustrating the triggering of the autonomous fault-tolerant reconfiguration mechanism is shown in the image. Figure 2 As shown, the process waits for a scheduling frame, then determines whether a scheduling frame has been acquired. If so, it acquires the scheduling frame's step period and step count, resets the step period monitoring count to zero, and starts timing from zero using the step period as a threshold. If not, it waits for a scheduling frame. It then determines whether the next scheduling frame is received during the timing period. If so, it returns to the steps of acquiring the scheduling frame's step period and step count. If not, it determines whether the timer has exceeded the threshold. If so, it increments the step period monitoring count by 1 and restarts timing. Next, it determines whether the step period monitoring count is greater than or equal to the step period timeout count. If so, it considers the uplink port link faulty; if less, it returns to the step of determining whether the next scheduling frame has been received during the timing period. Finally, it issues an autonomous fault-tolerant reconfiguration signal to trigger the autonomous fault-tolerant reconfiguration mechanism.
[0024] S102. In response to the autonomous fault-tolerant reconfiguration signal, control the opening of all physical ports and send link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration.
[0025] The link control protocol frame includes the best ID, hardware ID, and switching layer.
[0026] Specifically, in response to the autonomous fault-tolerant reconfiguration signal, control opens all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration, including: Step B1: In response to the autonomous fault-tolerant reconfiguration signal, control the opening of the uplink port, downlink port, and blocking port.
[0027] Step B2: Send link control protocol frames to the synchronous distributed cooperative network through the uplink port, downlink port and blocking port to initiate autonomous fault-tolerant reconfiguration based on the fast real-time spanning tree protocol.
[0028] In the link control protocol frame, both the best ID and the hardware ID are set to the ID of this device, and the switching level is initialized to 0.
[0029] Specifically, after step B2, root node arbitration is completed based on the fast real-time spanning tree mechanism, and the root node then sends down the arbitrated link control protocol frame (the best ID is set to the root node ID).
[0030] S103. Configure the status of the device ports according to the link control protocol frames transmitted back by the synchronous distributed cooperative network to complete the autonomous fault-tolerant reconfiguration process.
[0031] The device port status includes uplink port, downlink port and blocked port, and the autonomous fault-tolerant reconfiguration process is implemented based on the hardware logic in the FPGA.
[0032] Specifically, based on the link control protocol frames transmitted back from the synchronous distributed cooperative network, the status of the device ports is configured to complete the autonomous fault-tolerant reconfiguration process, including: Step C1: Parse the first switching layer of this device in the synchronous distributed cooperative network from the link control protocol frame transmitted back by the synchronous distributed cooperative network.
[0033] Step C2: Based on the first switching layer, determine the switching layer of the peer device corresponding to other link control protocol frames besides this device.
[0034] Step C3: Determine the status of the device port based on the first switching layer and the switching layer of the peer device.
[0035] Specifically, step C3 includes: Step C31: If the switching level of the peer device is equal to the first switching level, then configure the port of the peer device as a blocking port.
[0036] Step C32: If the switching level of the peer device is greater than the first switching level, then configure the port of the peer device as a downlink port.
[0037] Step C33: If the switching level of the peer device is lower than the first switching level, then configure the port of the peer device as an uplink port or a blocking port.
[0038] Specifically, step C33 includes: Step C331: If the switching level of the peer device is lower than the first switching level, then select the peer device with the smallest hardware ID value from all peer devices whose switching level is lower than the first switching level. Step C332: Configure the port corresponding to the peer device with the smallest selected hardware ID value as the uplink port; Step C333: Configure the ports corresponding to all peer devices except the peer device with the smallest selected hardware ID value as blocked ports.
[0039] Specifically, Figure 3 The flowchart illustrating the process of completing autonomous fault-tolerant reconfiguration is shown in the image. See [link / reference]. Figure 3As shown, after sending the Link Control Protocol (LCP) to all ports, the LCP frames returned by each port are immediately cached. The hardware ID of the directly connected device is determined by the hardware ID and cached. The Fast Real-Time Spanning Tree Bulletin (FAST) is continuously monitored to obtain the best ID, switching level, and hardware ID information. At the same time, LCP frames with the same hardware ID as the directly connected device are stored. After all the LCP frames obtained have the same best ID (indicating the end of the arbitration process), port reconstruction begins. The device obtains its own switching layer (i.e., the first switching layer) from the link control protocol frame whose hardware ID equals its own ID. Based on its own switching layer information, it determines the switching layers of other protocol frames. If the switching layer equals its own layer, the port is configured as a blocking port. If the switching layer is less than its own layer, the hardware ID in the announcement is compared with the temporary ID value, and the smaller one is retained as the temporary ID value. The initial temporary ID value is the hardware ID of the first frame whose switching layer is less than its own layer during the traversal. During the traversal, the temporary ID value is continuously compared with the hardware ID value in the link control protocol frame. The smallest temporary ID value is obtained during the entire traversal, and the corresponding port is configured as the unique uplink port, while other ports are configured as blocking ports. If the temporary ID value is greater than its own layer, the corresponding port is directly configured as a downlink port. The device completes port determination by traversing all link control protocol frames in the buffer. After the traversal ends, the port configuration information is sent.
[0040] S104. Receive and parse the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process, extract the protocol frame content, and restore the synchronous distributed cooperative network hierarchy based on the protocol frame content to reconstruct the synchronous distributed cooperative network topology and generate log files.
[0041] The log file is used to locate faults, and the protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement.
[0042] Specifically, the system receives and parses link control protocol frames generated during the autonomous fault-tolerant reconfiguration process, extracts the protocol frame content, and reconstructs the synchronous distributed cooperative network hierarchy based on the protocol frame content to rebuild the synchronous distributed cooperative network topology and generate log files, including: Step D1: Analyze the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process and extract the best ID.
[0043] Among them, the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process are generated within a complete synchronous distributed cooperative network reconfiguration event announcement cycle.
[0044] The link control protocol frames generated during the autonomous fault-tolerant reconfiguration process also include hardware ID, switching layer, and Fast Real-Time Spanning Tree (FLS) announcement. The best ID, hardware ID, switching layer, and FLS announcement are written to the database.
[0045] Step D2: When the best ID is not equal to the root node ID in a protocol frame, obtain the hardware ID of the device to which the device is connected, and determine the direct connection relationship between the devices based on the hardware ID of the device to which the device is connected.
[0046] Step D3: When the best ID is equal to the root node in the protocol frame, obtain the switching layer of each device in the synchronous distributed cooperative network.
[0047] Step D4: Based on the direct connection relationships between devices and the exchange level of each device in the synchronous distributed collaborative network, restore the synchronous distributed collaborative network level to reconstruct the synchronous distributed collaborative network topology and generate log files.
[0048] The log files include the number of monitoring cycles, link control protocol frames, synchronous distributed cooperative network layers, and direct connection relationships between devices; the log files can be text files or database files.
[0049] Specifically, Figure 4 The flowchart illustrating the reconstruction of the synchronous distributed cooperative network topology and fault diagnosis is shown in the attached diagram. Figure 4 As shown, all Link Control Protocol (LCP) content in the corresponding Fast Real-Time Spanning Tree (FAST) announcement is retrieved and stored in the cache. The best ID of the LCP is traversed, with comparisons made repeatedly until the smallest best ID is obtained, which corresponds to the ID of the root switch. All LCP content with a best ID equal to the root switch ID is filtered out; these LCP contents represent the LCP generated by the root switch after arbitration in the Synchronous Distributed Cooperative Network (SDN). For the above protocol content, a SDN hierarchy diagram is generated according to the "hardware ID-switch level" correspondence. The earliest LCP content in the cache is retrieved, with a quantity equal to the number of device ports. These frames correspond to the device that first sends back a LCP frame at the start of SDN arbitration, i.e., the device directly connected to this device. Connections between the device hardware and surrounding devices are established based on the hardware ID. By comparing the link logic with the previous FAST announcement, faulty links can be quickly identified.
[0050] This invention significantly reduces the development workload for R&D personnel in complex topology environments: the redundancy reconfiguration function is autonomously implemented by the synchronous distributed cooperative network, eliminating the need for individual configuration of each device, connection to a specific topology, and modeling of the synchronous distributed cooperative network. It also reduces system adaptation costs: when the system is expanded or reduced in size, the redundancy strategy is still autonomously completed by the synchronous distributed cooperative network, eliminating the need for repeated configuration. Furthermore, it significantly reduces repetitive workload when system requirements change multiple times. The autonomous fault-tolerant reconfiguration offers enhanced real-time performance: existing fault-tolerant mechanisms often require software implementation, while this fault-tolerant reconfiguration mechanism is implemented based on FPGA, providing strong real-time performance. Finally, it enables rapid fault location: devices can reconstruct the synchronous distributed cooperative network hierarchy and key information about the device's operating topology based on link control protocol frames, thereby achieving rapid fault location.
[0051] Based on the above Figure 1 As can be seen from the implementation method, this embodiment of the invention determines whether a timeout event has occurred based on the periodic timeout number and the continuously acquired synchronous distributed cooperative network data, and generates an autonomous fault-tolerant reconfiguration signal when a timeout event occurs; in response to the autonomous fault-tolerant reconfiguration signal, it controls the opening of all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration; according to the link control protocol frames returned by the synchronous distributed cooperative network, it configures the status of the device ports to complete the autonomous fault-tolerant reconfiguration process. The status of the device ports includes uplink ports, downlink ports, and blocked ports. The autonomous fault-tolerant reconfiguration process is implemented based on hardware logic in the FPGA; it receives and parses the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process, extracts the protocol frame content, and restores the synchronous distributed cooperative network hierarchy according to the protocol frame content to reconstruct the synchronous distributed cooperative network topology and generate a log file. The log file is used to locate faults. The protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement. In this way, an autonomous fault-tolerant reconfiguration mechanism is added to the communication, and a link control protocol frame is added. The autonomous fault-tolerant reconfiguration process is implemented by FPGA. It can determine the working status of the device port based on the operating status of the synchronous distributed cooperative network. When the working status of the port is abnormal, the device port can be easily reconfigured through the synchronous distributed cooperative network. When the connection relationship between system nodes is relatively complex, the autonomous fault-tolerant reconfiguration of the device port can be achieved by the synchronous distributed cooperative network. The entire process does not require the R&D personnel to configure in advance, and there is no need to repeat the configuration when the physical connection is changed multiple times, which can significantly reduce the development workload.
[0052] Based on the same inventive concept, as an implementation of the above-mentioned method for configuration-free fault-tolerant reconstruction and fault recording of synchronous distributed cooperative networks, this embodiment of the invention also provides a system for configuration-free fault-tolerant reconstruction and fault recording of synchronous distributed cooperative networks. Figure 5This is a structural diagram of the synchronous distributed cooperative network configuration-free fault-tolerant reconfiguration and fault recording system in an embodiment of the present invention. (See also...) Figure 5 As shown, the synchronous distributed cooperative network configuration-free fault-tolerant reconfiguration and fault recording system may include: The synchronous distributed cooperative network module 501, deployed in FPGA hardware, is used to determine whether a timeout event has occurred based on the periodic timeout count and continuously acquired synchronous distributed cooperative network data, and to generate an autonomous fault-tolerant reconfiguration signal when a timeout event occurs; in response to the autonomous fault-tolerant reconfiguration signal, it controls the opening of all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration; according to the link control protocol frames returned by the synchronous distributed cooperative network, it configures the status of the device ports to complete the autonomous fault-tolerant reconfiguration process. The status of the device ports includes uplink ports, downlink ports, and blocked ports; The synchronous distributed cooperative soft bus module 502, connected to the FPGA, is used to receive and transmit link control protocol frames returned by the synchronous distributed cooperative network. The fault recording module 503 is connected to the synchronous distributed cooperative soft bus module. It is used to receive and parse the link control protocol frames generated during the autonomous fault-tolerant reconstruction process, extract the protocol frame content, and restore the synchronous distributed cooperative network hierarchy based on the protocol frame content in order to rebuild the synchronous distributed cooperative network topology and generate a log file. The log file is used to locate faults. The protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement.
[0053] Specifically, in terms of hardware design, this invention uses an FPGA to implement the synchronous distributed collaborative network protocol stack and achieves data interconnection with the host computer (computing chip) via a PCIe link. It supports computing devices in a synchronous distributed collaborative network in the form of an FPGA + CPU / SOC. Multiple computing devices can be interconnected through the synchronous distributed collaborative network, forming the electronic and electrical architecture of the equipment with the synchronous distributed collaborative network as the backbone. The FPGA interacts with the CPU through the CPU chip's high-bandwidth bus (usually a PCIe bus).
[0054] The synchronous distributed collaborative network configuration-free fault-tolerant reconfiguration and fault recording system consists of three main modules: a synchronous distributed collaborative network module, a synchronous distributed collaborative soft bus module, and a fault recording module. The synchronous distributed collaborative network module is implemented using FPGA, the synchronous distributed collaborative soft bus module is implemented using software, and the fault recording module is implemented using software.
[0055] The synchronous distributed cooperative network module comprises a physical port control module, a synchronous distributed cooperative network protocol stack, a scheduling control module, a step cycle monitoring module, and a redundancy reconfiguration control module. The physical port control module is responsible for port startup and shutdown and driving physical port transmission and reception. The synchronous distributed cooperative network protocol stack implements data queuing, framing, and transmission / reception functions. The scheduling control module processes synchronous distributed cooperative network scheduling frames. The step cycle monitoring module processes the synchronous distributed network step cycle, thereby enabling autonomous judgment of port online / offline status. The redundancy reconfiguration control module is the core of the fault-tolerant reconfiguration mechanism. It monitors and processes the link control protocol in the synchronous distributed cooperative network and receives fault signals from the step cycle monitoring module, triggering the synchronous distributed cooperative network reconfiguration mechanism. Upon receiving a fault signal generated by the step cycle monitoring module, the redundancy reconfiguration control module sends information to the physical port control module to open all physical ports; it also sends the link control protocol to the synchronous distributed cooperative network protocol stack, thereby sending the link control protocol to all physical ports and triggering synchronous distributed cooperative network link reconfiguration.
[0056] The synchronous distributed collaborative soft bus module is deployed inside the computing device to realize bidirectional interaction of synchronous distributed collaborative network data between the computing device and the FPGA, as well as the parsing of the synchronous distributed collaborative network protocol.
[0057] The fault recording module is implemented in software and includes a protocol receiving module, a protocol parsing module, a topology recovery module, and a log recording module. The protocol receiving module passively receives synchronous distributed cooperative network scheduling data and link control protocol frames transmitted downlink via the soft bus. The protocol parsing module parses and analyzes the best ID, hardware ID, its own switching layer, and Fast Real-Time Spanning Tree Advertisement (FAST) in the link control protocol frames. The topology recovery module reconstructs the topology of the synchronous distributed cooperative network based on all the link control protocol frames received. The log recording module collects log information and writes it to a txt file or records it in a database.
[0058] Figure 6 A schematic diagram of a synchronous distributed cooperative network is shown below. Figure 6As shown, typically, a device uses only one synchronous distributed cooperative network port for uplink data transmission and downlink data reception (referred to as the uplink port). Other ports are either inactive (referred to as blocking ports) or only forward data (device data is not transmitted to this port, referred to as downlink ports). Through the uplink port, the computing device can receive all data within each step cycle of the synchronous distributed cooperative network, including scheduling frames and link control protocols, and transfer the data to the CPU via a high-bandwidth bus to achieve data interaction. The uplink port is also responsible for sending all data, including data received from the downlink ports and data generated by the computing device itself. The downlink port directly forwards the synchronous distributed cooperative network frames received by the uplink port. This process begins as soon as data is received at the uplink port and does not involve data buffering. Furthermore, data received by the downlink port is queued in the FPGA according to parameters such as time and priority before being uniformly uploaded through the uplink port. Blocking ports do not participate in synchronous distributed cooperative network transmission and reception. Their physical links may connect to other devices, potentially forming loopback links.
[0059] Synchronous distributed cooperative networks can adopt arbitrary physical topology connections, but under the Fast Real-Time Spanning Tree Protocol (FAST), the final working topology is a tree topology. This allows each synchronous distributed cooperative network device's working port to be divided into uplink and downlink ports. When the synchronous distributed network is working normally, a single device is elected as the root node of the entire synchronous distributed cooperative network through the FAST mechanism.
[0060] Figure 7 The overall system architecture diagram is shown schematically. (See attached diagram) Figure 7 As shown, the physical port control module enables the starting and stopping of synchronous distributed collaborative network ports and drives physical port transmission and reception. This process is automatically implemented by the FPGA without configuration. By receiving commands from the redundancy reconfiguration control module, the physical port control module can determine which ports of the device are uplink ports, downlink ports, or blocking ports, achieving port self-management. This process is also automatically implemented by the FPGA. This physical port control module is also responsible for driving the Ethernet Physics.
[0061] The Synchronous Distributed Cooperative Network (SDC) protocol stack is used to implement data framing and transmission / reception functions for SDC. Its functionality is similar to the MAC layer (Layer 2 and Layer 3) of ordinary Ethernet, but it uses a custom protocol stack for SDC, particularly including scheduling frames and link control protocols. Furthermore, this module is responsible for sending received SDC data to computing devices via a heterogeneous bus (typically a PCIe bus).
[0062] The scheduling control module processes scheduling frames for the synchronous distributed cooperative network. Each scheduling frame contains a step period (adjustable from 20 microseconds to 2 seconds), a step count, and a step offset. This scheduling control module continuously acquires the step period from the scheduling frame and provides it to other modules.
[0063] The step cycle monitoring module continuously monitors the number of step beats. This module has a built-in timer that uses the synchronous distributed cooperative network's step cycle as a threshold for timing, thus verifying the step cycle. Specifically, the module starts timing synchronously after a step beat update and restarts timing upon receiving the next step beat. If the module's timing exceeds the synchronous distributed cooperative network's step cycle, it indicates an error in the scheduling frame (no scheduling frame received after one step cycle). When an abnormal step beat is detected (discontinuous or stopped), the step cycle monitoring count is incremented by 1. A step cycle timeout can be configured (two step cycles are monitored by default). When the step cycle monitoring count reaches the timeout limit, the module triggers the redundancy reconfiguration control module's reconfiguration mechanism.
[0064] The redundancy reconfiguration module handles the link control protocol. Based on a fast real-time spanning tree, this module identifies the ports of the device's synchronous distributed cooperative network and configures the physical port control module. After the step-cycle monitoring module triggers the redundancy reconfiguration module, it controls the physical port control module to open all physical ports (allowing blocked ports to send link control protocol frames), thus triggering the fault-tolerant reconfiguration of the entire synchronous distributed cooperative network.
[0065] The synchronous distributed collaborative soft bus module enables bidirectional data interaction between the synchronous distributed collaborative network and the operating system of the computing device.
[0066] The protocol receiving module is capable of receiving synchronous distributed cooperative network scheduling data and link control protocol frames.
[0067] The protocol parsing module performs collaborative parsing of synchronous distributed cooperative network (SDR) scheduling data and link control protocol (LCP) data. Specifically, it records the SDR step cycle, the best ID, hardware ID, switching level, and Fast Real-Time Spanning Tree (FPS) announcement (indicating the number of SDR reconfigurations) from the SDR scheduling frame. Within a SDR reconfiguration cycle, all SDR devices connected to the SDR, including switches and computing devices, send LCP frames on all ports, thereby completing SDR arbitration based on the FPS mechanism. After the SDR arbitrates and determines the root node, the root node fans out the arbitration result (the LCP frame used by the root node to complete the arbitration) to all ports. The protocol parsing module manages information about devices connected to the SDR by recording all LCP frames in a FPS announcement.
[0068] The topology recovery module uses the data recorded by the protocol parsing module to determine the information of each synchronous distributed cooperative network device, including the device's hard ID (the ID corresponding to each device), its own switching layer, and its best ID (the root switch ID, used to achieve fast real-time spanning tree). Based on this, it restores the synchronous distributed cooperative network to operation according to logical relationships.
[0069] The logging module records crucial data from the fault-tolerant reconfiguration process of the synchronous distributed cooperative network, including the number of synchronous distributed cooperative network step cycles, link control protocol frames, the synchronous distributed cooperative network hierarchy generated by the topology recovery module, and device connection relationships, and stores this data in the database. Developers can view the corresponding reconfigured synchronous distributed cooperative network topology based on the rapid real-time spanning tree announcement count, thereby quickly locating faulty links and enabling fault diagnosis.
[0070] The synchronous distributed cooperative network module 501 includes: The physical port control module is used to continuously acquire synchronous distributed cooperative network data from the uplink port; A synchronous distributed collaborative network protocol stack is used to parse scheduling frames and provide them to the scheduling control module; The scheduling control module is used to continuously extract the step count of the scheduling frame; The step cycle monitoring module resets the monitoring timer each time a step count is extracted. If the timer value exceeds the step cycle and the next step count is not extracted, a step cycle timeout is determined, and the step cycle monitoring count is incremented. The module counts consecutive step cycle monitoring counts. When the number of step cycle monitoring counts reaches the cycle timeout count, a timeout event is determined, and an autonomous fault-tolerant reconfiguration signal is generated and sent to the redundant reconfiguration module. The monitoring timer is a timer set based on the step cycle.
[0071] The redundancy reconfiguration control module is used to respond to the autonomous fault-tolerant reconfiguration signal, control the opening of the uplink port, downlink port and blocking port, and send the link control protocol to the synchronous distributed cooperative network protocol stack module. In the link control protocol frame, the best ID and hardware ID are both set to the ID of this device, and the switching level is initialized to 0. The synchronous distributed cooperative network protocol stack is also used to implement link control protocol framing. The physical port control module is also used to send link control protocol frames to the synchronous distributed cooperative network through uplink ports, downlink ports and blocking ports to initiate autonomous fault-tolerant reconfiguration based on the fast real-time spanning tree protocol.
[0072] Once the best IDs of all the Link Control Protocol frames it acquires are consistent (indicating the end of the arbitration process), port reconstruction begins.
[0073] The redundancy reconfiguration control module is also used to parse the first switching layer of the device in the synchronous distributed cooperative network in the link control protocol frames transmitted back from the synchronous distributed cooperative network; based on the first switching layer, determine the switching layer of the peer device corresponding to other link control protocol frames besides the device itself; and determine the status of the device port according to the first switching layer and the switching layer of the peer device. The link control protocol frames transmitted back from the synchronous distributed cooperative network are link control protocol frames that have completed root node arbitration based on the fast real-time spanning tree mechanism and are downlinked from the root node after arbitration, wherein the best ID is set as the root node ID.
[0074] In the redundancy reconfiguration control module, the state of the device port is determined based on the first switching level and the switching level of the peer device, including: if the switching level of the peer device is equal to the first switching level, the port of the peer device is configured as a blocking port; if the switching level of the peer device is greater than the first switching level, the port of the peer device is configured as a downlink port; if the switching level of the peer device is less than the first switching level, the port of the peer device is configured as an uplink port or a blocking port.
[0075] In the redundancy reconfiguration control module, if the switching level of the peer device is lower than the first switching level, the port of the peer device is configured as an uplink port or a blocking port, including: if the switching level of the peer device is lower than the first switching level, select the peer device with the smallest hardware ID value from all peer devices with switching levels lower than the first switching level; configure the port corresponding to the selected peer device with the smallest hardware ID value as an uplink port; configure the ports corresponding to the other peer devices except the selected peer device with the smallest hardware ID value as blocking ports.
[0076] Specifically, the step cycle monitoring module is also used to reset the step cycle monitoring count to zero if the monitoring timer's timing value exceeds the step cycle and the next step beat number is extracted.
[0077] See Figure 2 As shown, the step cycle monitoring module is specifically used to wait for scheduling frames. It acquires the step cycle and step count of the scheduling frame, resets the step cycle monitoring count to zero, and starts timing from zero using the step cycle as a threshold. If the next scheduling frame is received during the timing process, it returns to the step of acquiring the scheduling frame's step cycle and step count; if the timing exceeds the threshold, it increments the step cycle monitoring count by 1 and restarts timing. If the step cycle monitoring count is greater than or equal to the step cycle timeout count, the uplink port link is considered faulty. The step cycle monitoring module sends information to the redundancy reconfiguration module, triggering the redundancy reconfiguration control module's reconfiguration mechanism.
[0078] See Figure 3As shown, the redundancy reconfiguration module is specifically used to immediately cache the link control protocols returned by each port after sending the link control protocol. It determines the hardware ID information of the directly connected device by the hardware ID and caches it. It continuously monitors the Fast Real-Time Spanning Tree Bulletin (FAST) to obtain the best ID, switching level, and hardware ID information; simultaneously, it stores link control protocol frames with the same hardware ID as the directly connected device. Once the best ID of all obtained link control protocol frames is consistent (indicating the end of the arbitration process), port reconfiguration begins. It obtains its own switching level information from link control protocol frames with a hardware ID equal to its own ID. Based on its own switching level information, it determines the switching level of other protocol frames. If the switching level is equal to its own level, the port is configured as a blocking port. If the switching level is lower than its own level, it compares the hardware ID of the announcement with the temporary ID value, retaining the smaller one as the temporary ID value; the initial temporary ID value is the hardware ID of the first frame with a switching level lower than its own level during the traversal process. During the traversal, the temporary ID value is continuously compared with the hardware ID value in the link control protocol frame. The port with the smallest temporary ID value is selected and configured as the unique uplink port; other ports are configured as blocking ports. If the temporary ID value is greater than the one in its own layer, the corresponding port is directly configured as the downlink port. The device's port determination is completed by traversing all link control protocol frames in the buffer. Once the traversal is complete, the port configuration information is sent to the physical port control module.
[0079] Specifically, the device physical port control module can receive link control protocol frames, transmit their data to the synchronous distributed cooperative network protocol stack, and then transmit them to the synchronous distributed cooperative soft bus module via the PCIe link.
[0080] The fault recording module 503 includes: The protocol receiving module is used to receive link control protocol frames generated during the autonomous fault-tolerant reconfiguration process via a soft bus and write them into the database. The protocol parsing module is used to parse the link control protocol frames generated during the autonomous fault-tolerant reconstruction process and extract the best ID. The link control protocol frames generated during the autonomous fault-tolerant reconstruction process are generated within a complete synchronous distributed cooperative network reconstruction event announcement cycle. The link control protocol frames generated during the autonomous fault-tolerant reconstruction process also include hardware ID, switching level and fast real-time spanning tree announcement. The topology recovery module is used to obtain the hardware ID of the device to which the device is connected when the best ID is not equal to the root node ID of the protocol frame, and determine the direct connection relationship between the devices based on the hardware ID of the device to which the device is connected; when the best ID is equal to the root node of the protocol frame, it obtains the switching layer of each device in the synchronous distributed cooperative network; and restores the synchronous distributed cooperative network layer based on the direct connection relationship between the devices and the switching layer of each device in the synchronous distributed cooperative network, so as to reconstruct the synchronous distributed cooperative network topology.
[0081] The logging module is used to generate log files.
[0082] The log files include the number of monitoring cycles, link control protocol frames, synchronous distributed cooperative network layers, and direct connection relationships between devices; the log files can be text files or database files. The log files can manage the database content.
[0083] For details, see Figure 4 As shown, the topology recovery module is specifically used to obtain all Link Control Protocol (LCP) content from the corresponding Fast Real-Time Spanning Tree (FAST) announcement and store it in a cache. It iterates through the best LCP IDs, comparing them repeatedly until the smallest best ID is found, thus obtaining the ID corresponding to the root switch. It then filters all LCP content whose best ID equals the root switch ID; these LCP contents represent the LCP generated by the root switch after arbitration in the Synchronous Distributed Cooperative Network (SDN). For the above protocol content, a SDN hierarchy diagram is generated according to the "hardware ID-switch level" correspondence. The earliest LCP content in the cache is retrieved, with the number equal to the number of device ports. These frames correspond to the device that first sends back a LCP frame at the start of SDN arbitration, i.e., the device directly connected to this device. Based on the hardware ID, the connection relationship between the device hardware and surrounding devices is established. By comparing the link logic with the previous FAST announcement, faulty links can be quickly identified.
[0084] Figure 8 A schematic diagram of the FPGA network section of a synchronous distributed cooperative network computing device is shown. (See attached diagram) Figure 8As shown, in terms of hardware, this invention uses a Xilinx Xc7k160t FPGA chip and an RK3588J SOC to develop a hardware device supporting synchronous distributed cooperative networks. The FPGA selected for this computing device is a Xilinx Xc7k160t, which implements bidirectional data interaction based on the PCIe 3.0×4 bus provided by the RK3588. The FPGA expands to three synchronous distributed cooperative network interfaces, and physical layer transmission and reception are implemented through the Yutai Micro RTL8531 chip. The physical interface uses an M12 8-pin X-encoded connector. The FPGA is connected to 1GB of DDR3 memory as a synchronous distributed network protocol frame buffer. This solution can support gigabit-speed synchronous distributed cooperative networks. Table 1 shows the synchronous distributed cooperative network scheduling protocol frames. Each synchronous distributed cooperative network scheduling frame includes at least the number of step beats (StepBeat) and the step offset value (StepOffset), where the step offset value is in microseconds. The StepTime is the stepping period of the synchronous distributed cooperative network, measured in microseconds, representing the network's operating cycle. The StepTime tick count represents the current number of cycles the synchronous distributed cooperative network has run. This value increments by one for each stepping period elapsed by the root node. At the start of each stepping period of the synchronous distributed cooperative network, the root node sends a scheduling frame.
[0085] Table 1 Synchronous Distributed Cooperative Network Scheduling Protocol Frames
[0086] Table 2 shows the link control protocol frames for the synchronous distributed cooperative network. The designed synchronous distributed cooperative network link control protocol frames include a best ID, switching level, hardware ID, and a Fast Real-Time Spanning Tree (FLS) announcement, used to implement link reconfiguration based on FLS. For computing devices, the best ID indicates the root switch (capable of determining whether it is an uplink); the switching level indicates the level of the synchronous distributed cooperative network to which the device resides; the hardware ID is generated by the FPGA and is unique to each device; the FLS announcement indicates how many times the synchronous distributed cooperative network reconfiguration has been performed, and is a counter field used to uniquely identify and associate all link control protocol frames generated during a complete synchronous distributed cooperative network reconfiguration process.
[0087] Table 2 Link Control Protocol Frames for Synchronous Distributed Cooperative Networks
[0088] This invention leverages the operational characteristics of synchronous distributed collaborative networks (SDRs) to implement autonomous redundancy of SDR port ports on devices using FPGA hardware. The redundancy process is autonomously completed by the SDR, not by the devices themselves. Developers do not need to configure ports or create specific network topologies. Autonomous redundancy is achieved simply by connecting the device's SDR port to the backbone network. Furthermore, based on this redundancy mechanism, proactive topology notification and link error reporting are possible, significantly increasing the reliability of the distributed system.
[0089] This invention aims to utilize the communication characteristics provided by a synchronous distributed collaborative network / bus, namely: the synchronous distributed collaborative network operates according to a step cycle, and a single port of a device can receive all data within one step cycle; and the device exhibits "decentralized, peer-to-peer distributed characteristics." A fault-tolerant reconfiguration mechanism is added to the communication, along with a link control protocol and corresponding mechanism. This mechanism, implemented by an FPGA, can determine the working status of a device port based on the operating status of the synchronous distributed collaborative network. When a port's working status is abnormal, it can reconfigure the device port through the synchronous distributed collaborative network. This invention assigns the port reconfiguration function to the synchronous distributed collaborative network, rather than the device itself, thus enabling autonomous fault-tolerant reconfiguration of device ports even when system node connections are complex. The entire process requires no prior configuration by developers, and no changes to redundancy schemes are needed when physical connections change. This greatly simplifies the system fault-tolerant redundancy workload for developers. Furthermore, the autonomous fault-tolerant reconfiguration process is implemented based on an FPGA, without processing by software within the computing device, possessing strong real-time performance, and meeting the high real-time and high deterministic requirements of complex digital systems. After link fault tolerance reconstruction, the system supports recording the current working topology of the device and forming a log file, which makes it easier for R&D personnel to locate link faults.
[0090] It should be noted that the above description of the configuration-free fault-tolerant reconstruction and fault recording system for synchronous distributed cooperative networks is similar to the description of the above-described method embodiment for configuration-free fault-tolerant reconstruction and fault recording in synchronous distributed cooperative networks, and has similar beneficial effects. For technical details not disclosed in the embodiments of the configuration-free fault-tolerant reconstruction and fault recording system for synchronous distributed cooperative networks of the present invention, please refer to the description of the method embodiment for configuration-free fault-tolerant reconstruction and fault recording in synchronous distributed cooperative networks of the present invention for understanding.
[0091] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for configuration-free fault-tolerant reconfiguration and fault recording in a synchronous distributed cooperative network, characterized in that, include: Based on the periodic timeout count and continuously acquired synchronous distributed collaborative network data, determine whether a timeout event has occurred, and generate an autonomous fault-tolerant reconfiguration signal when a timeout event occurs. In response to the autonomous fault-tolerant reconfiguration signal, control opens all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration; Based on the link control protocol frames transmitted back by the synchronous distributed cooperative network, the state of the device port is configured to complete the autonomous fault-tolerant reconfiguration process. The state of the device port includes uplink port, downlink port and blocked port. The autonomous fault-tolerant reconfiguration process is implemented based on hardware logic in FPGA. The system receives and parses the link control protocol frames generated during the autonomous fault-tolerant reconfiguration process, extracts the protocol frame content, and restores the synchronous distributed cooperative network hierarchy based on the protocol frame content to rebuild the synchronous distributed cooperative network topology and generate a log file. The log file is used to locate faults. The protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement.
2. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 1, characterized in that, The process of determining whether a timeout event has occurred based on the periodic timeout count and continuously acquired synchronous distributed cooperative network data, and generating an autonomous fault-tolerant reconfiguration signal when a timeout event occurs, includes: The number of ticks of the scheduling frame is continuously extracted from the synchronous distributed cooperative network data. Each time the number of step beats is extracted, the monitoring timer is reset. The monitoring timer is a timer set based on the step cycle. If the timing value of the monitoring timer exceeds the step cycle and the next step beat is not extracted, it is determined that a step cycle timeout has occurred, and the step cycle monitoring count is incremented by one. The system counts the number of consecutive step cycle monitoring cycles. When the number of step cycle monitoring cycles reaches the number of cycle timeouts, it determines that a timeout event has occurred and generates the autonomous fault-tolerant reconfiguration signal when the timeout event occurs.
3. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 1, characterized in that, The link control protocol frame includes the best ID, the hardware ID, and the switching layer. In response to the autonomous fault-tolerant reconfiguration signal, controlling the opening of all physical ports and sending a link control protocol frame to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration includes: In response to the autonomous fault-tolerant reconfiguration signal, the uplink port, the downlink port, and the blocking port are controlled to open. The link control protocol frame is sent to the synchronous distributed cooperative network through the uplink port, the downlink port and the blocking port to initiate the autonomous fault-tolerant reconfiguration based on the fast real-time spanning tree protocol. In the link control protocol frame, the best ID and the hardware ID are both set to the ID of this device, and the switching level is initialized to 0.
4. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 1, characterized in that, The process of configuring the device port status based on the link control protocol frames returned by the synchronous distributed cooperative network to complete the autonomous fault-tolerant reconfiguration includes: In the link control protocol frame transmitted back by the synchronous distributed cooperative network, the first switching layer of this device in the synchronous distributed cooperative network is parsed out; Based on the first switching layer, determine the switching layer of the peer device corresponding to other link control protocol frames besides the device itself; The state of the device port is determined based on the first switching level and the switching level of the peer device.
5. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 4, characterized in that, Determining the state of the device port based on the first switching layer and the switching layer of the peer device includes: If the switching level of the peer device is equal to the first switching level, then the port of the peer device is configured as the blocking port; If the switching level of the peer device is greater than the first switching level, then the port of the peer device is configured as the downlink port; If the switching level of the peer device is lower than the first switching level, then the port of the peer device is configured as the uplink port or the blocking port.
6. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 5, characterized in that, If the switching level of the peer device is lower than the first switching level, then configuring the port of the peer device as the uplink port or the blocking port includes: If the switching level of the peer device is lower than the first switching level, then select the peer device with the smallest hardware ID value from all peer devices whose switching level is lower than the first switching level. Configure the port corresponding to the peer device with the smallest selected hardware ID value as the uplink port; Configure the ports corresponding to the peer devices other than the peer device with the smallest selected hardware ID value as the blocking ports.
7. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 1, characterized in that, The process of receiving and parsing the link control protocol frames generated during the autonomous fault-tolerant reconfiguration, extracting the protocol frame content, and reconstructing the synchronous distributed cooperative network hierarchy based on the protocol frame content to rebuild the synchronous distributed cooperative network topology and generate log files includes: The link control protocol frames generated during the autonomous fault-tolerant reconfiguration process are analyzed to extract the best ID. The link control protocol frames generated during the autonomous fault-tolerant reconfiguration process are generated within a complete synchronous distributed cooperative network reconfiguration event announcement period. When the optimal ID is not equal to the root node ID in a protocol frame, the hardware ID of the device to which the device is connected is obtained, and the direct connection relationship between the devices is determined based on the hardware ID of the device to which the device is connected. When the optimal ID is equal to the root node in the protocol frame, obtain the switching layer of each device in the synchronous distributed cooperative network; Based on the direct connection relationships between the devices and the exchange level of each device in the synchronous distributed collaborative network, the synchronous distributed collaborative network level is restored to reconstruct the synchronous distributed collaborative network topology and generate the log file.
8. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 1, characterized in that, The log file includes the number of monitoring steps in the cycle, the link control protocol frames, the synchronous distributed cooperative network hierarchy, and the direct connection relationships between devices; the log file is a text file or a database file.
9. The method for configuration-free fault-tolerant reconfiguration and fault recording of synchronous distributed cooperative networks according to claim 7, characterized in that, The value of the periodic timeout number ranges from 2 to 10.
10. A configuration-free fault-tolerant reconfiguration and fault recording system for a synchronous distributed cooperative network, characterized in that, include: A synchronous distributed cooperative network module, deployed in FPGA hardware, is used to determine whether a timeout event has occurred based on the periodic timeout count and continuously acquired synchronous distributed cooperative network data, and to generate an autonomous fault-tolerant reconfiguration signal when a timeout event occurs; in response to the autonomous fault-tolerant reconfiguration signal, it controls the opening of all physical ports and sends link control protocol frames to the synchronous distributed cooperative network to initiate autonomous fault-tolerant reconfiguration; according to the link control protocol frames returned by the synchronous distributed cooperative network, it configures the state of the device ports to complete the autonomous fault-tolerant reconfiguration process, wherein the state of the device ports includes uplink ports, downlink ports, and blocked ports; A synchronous distributed cooperative soft bus module, connected to the FPGA, is used to receive and transmit link control protocol frames returned by the synchronous distributed cooperative network. The fault recording module, connected to the synchronous distributed cooperative soft bus module, is used to receive and parse the link control protocol frames generated during the autonomous fault-tolerant reconstruction process, extract the protocol frame content, and restore the synchronous distributed cooperative network hierarchy based on the protocol frame content to rebuild the synchronous distributed cooperative network topology and generate a log file. The log file is used to locate faults. The protocol frame content includes the best ID, hardware ID, switching level, and fast real-time spanning tree announcement.