Multi-level resetting method and system for distributed architecture of protection control equipment

By using hardware control words and software configuration words in the distributed architecture of the protection control device, multi-level reset control is achieved, and the problem of single reset strategy and weak global exception handling capabilities in the existing technology is solved, and the reliability and availability of the system are improved.

CN119937433AActive Publication Date: 2025-05-06NANJING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510428429.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the prior art, the distributed architecture reset mechanism of the protection control device has problems such as single reset strategy, insufficient flexibility, and insufficient state feedback, which is difficult to meet the requirements of high-reliability distributed protection control devices.

Method used

The reset enable control of different service CPU boards through hardware control words, and combined with the software configuration words of the entire device, the multi-level reset control of independent single boards, coupled multi-boards and entire systems of distributed protection control devices is realized.

Benefits of technology

It effectively solves the problems of single reset strategy, weak global exception handling capabilities and insufficient applicability of distributed systems, improves the overall reliability and availability of the system, and avoids the failure of non-associated business functions or the failure of the entire system caused by local abnormal spread.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937433A_ABST
    Figure CN119937433A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-level resetting method and system for a distributed architecture of protection control equipment, which are particularly suitable for the protection control equipment of the distributed architecture. And determining exception handling modes of different function service modules of the device by adopting a mode of combining a hardware control word and a software configuration word. The method comprises the following steps: selecting a reset level according to an abnormal influence range, combining hierarchical identification of abnormal board card types and a dynamic reset strategy, preferentially performing local reset of a single board or multiple boards, and ensuring service collaboration and state synchronization among reset modules through communication link reconstruction and data consistency maintenance. And starting system reset of the whole device when local reset is unsuccessful for multiple times or multiple boards are abnormal at the same time. Through multi-level reset control of single board reset, multi-board cooperative reset and whole device system reset, the shutdown risk of the whole machine is effectively reduced, and the system availability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of embedded protection and control equipment for electric power systems, and in particular to a multi-level reset method and system for a distributed architecture of protection and control equipment. Background Art

[0002] With the development of new power systems, the functions of protection and control equipment are becoming increasingly complex, and the requirements for protection and control coordination are becoming higher and higher. Protection and control equipment based on distributed architecture design has become a hot topic of research. Such devices usually adopt a modular design and deploy different protection, control and measurement tasks in a distributed architecture. However, the design of the distributed architecture makes the fault management and reset control of each module complicated. In the prior art, common reset mechanisms mainly include power monitoring reset, module independent reset and multi-core collaborative reset, but these mechanisms are not adaptable enough to distributed architectures, and there are problems such as a single reset strategy, insufficient flexibility, and insufficient status feedback, which makes it difficult to meet the requirements of high-reliability distributed protection and control equipment.

[0003] In view of the above problems, Chinese patent CN114527857A proposes a multi-core system reset method. Its technical solution is as follows: after the multi-core system performs bus reset, the reset information of the target processor is obtained; after judging that the reset of the target processor is successful, the target processor is used to detect whether all the remaining processors of the multi-core system are successfully reset; if so, the multi-core reset is terminated; if it is detected that any processor is not successfully reset, the register corresponding to the processor that is not successfully reset is controlled by the target processor to reset again, until all the remaining processors of the multi-core system are successfully reset, and the multi-core reset is terminated.

[0004] However, the prior art has the following disadvantages:

[0005] (1) Single reset strategy and lack of flexibility

[0006] The existing reset strategy has a single trigger condition, usually relying on simple hardware monitoring signals, and a fixed reset mode, making it difficult to select the reset range according to the type and severity of the abnormality. This "one-size-fits-all" reset method lacks a sophisticated hierarchical design, which is particularly evident in the multi-core business board architecture, and may lead to an expansion of the reset range and an increase in reset time, thereby reducing system availability.

[0007] (2) Lack of priority mechanism for core business modules

[0008] The existing reset technology does not adequately distinguish between core business modules and non-core business modules. In practical applications, if the entire device system is reset after a non-core business exception occurs, the core business operation will be interrupted, which will bring certain risks to the operation of the power system. Whether the entire device system needs to be reset after a non-core business exception occurs remains to be discussed.

[0009] (3) Data status mismatch between boards after reset

[0010] The existing technology lacks effective management of the state synchronization of communication modules between boards in a distributed architecture. During the reset and restart process after an abnormality, data disorder may occur between related modules, which may lead to failure of module collaboration. In severe cases, it may even cause protection malfunction or refusal to operate, affecting the normal operation of the system. Summary of the invention

[0011] The purpose of the present invention is to address the problems existing in the above-mentioned prior art and to provide a multi-level reset method and system for a distributed architecture of protection and control equipment. Reset enable control of different service CPU boards is realized through hardware control words. Combined with the software configuration words of the entire device, multi-level reset control of independent single boards, coupled multiple boards and the entire system of distributed protection and control equipment is realized to prevent local system anomalies from causing failure of non-associated business functions or failure of the entire machine business, thereby ensuring the effectiveness of the distributed deployment of the entire machine business and improving the overall reliability and availability of the device system.

[0012] The technical solution to achieve the purpose of the present invention is: on the one hand, a multi-level reset system of a distributed architecture for protection and control equipment is provided, the system is designed based on a distributed architecture, and includes a management CPU board, a business CPU board, an IO board, a power board and a bus board;

[0013] The management CPU board deploys a management module to implement the issuance of reset instructions and process guidance, and receives status information from the service board through a communication link, and analyzes and makes decisions on abnormal status;

[0014] The business CPU board deploys business modules to implement logical calculation and processing of various real-time tasks;

[0015] The IO board is used to collect input and control output of switch signals, as well as collect and measure small analog signals;

[0016] The power board is used to provide power to the management CPU board and the service CPU board;

[0017] The bus board is used to realize signal interconnection between the management CPU board and the service CPU board.

[0018] Furthermore, the management CPU board and the service CPU board deploy one or more functional modules among an onboard power supply monitoring module, a communication module, a data integrity maintenance module, a functional integrity reset module and an abnormality monitoring module;

[0019] The power monitoring module is used to monitor the abnormal status of the onboard main power supply and local power supply, output a reset signal, and the hardware control word determines whether to reset the board;

[0020] The communication module is deployed in the management CPU board and the service CPU board, and is used for data interaction between the management CPU board and the service CPU board;

[0021] The data integrity maintenance module is deployed in the management CPU board and is used to maintain the consistency of the interaction data between the management CPU and the reset business CPU board;

[0022] The functional complete reset module is deployed in the management CPU board. When a local or entire device system is reset, it restores the entire system through function interface initialization, parameter loading initialization and task scheduling;

[0023] The abnormality monitoring module is deployed in the management CPU board, monitors the system status in real time, captures single-board abnormalities, multi-board abnormalities and system abnormalities, and triggers corresponding reset strategies.

[0024] On the other hand, a multi-level reset method for a distributed architecture of a protection control device is provided, the method comprising:

[0025] Single board reset: For single board abnormalities, the output of the reset signal is controlled by the hardware control word, so that the core business CPU board is reset first;

[0026] Multi-board coordinated reset: For other core business CPU boards that are business-coupled with the abnormal core business CPU board, after the abnormal core business CPU board is reset, the management CPU board actively triggers the reset of the other core business CPU boards according to the software configuration word;

[0027] Reset the entire device system: When the single-board reset and multi-board coordinated reset fail to the upper limit or multiple unrelated core business CPU boards are reset at the same time, the management CPU board actively triggers the reset of all other business CPU boards, and then the management CPU board resets itself to start the power-on reset process of the entire device.

[0028] Furthermore, the single board resetting specifically includes:

[0029] The power monitoring circuit inside the board detects the abnormality of the main power supply and the local power supply. If the voltage threshold is exceeded, a reset signal is sent to the CPU of the board to reset the CPU of the board.

[0030] Monitor the abnormal status of the CPU of this board. When an abnormality is detected, send an alarm or lockout signal to the communication bus. Specifically: if this board is a core business CPU board, send a lockout signal to the backplane bus; if this board is a non-core business CPU board, send an alarm signal to the backplane bus.

[0031] Furthermore, the reset signal is output through a hardware control word control, and the hardware control word is to place a pull-up or pull-down resistor at the bus board position corresponding to the business CPU board, one end of the pull-up or pull-down resistor is connected to the positive power supply or power ground, and the other end is logically ANDed with the reset output of the power monitoring module of this business CPU board and connected to the reset pin of the CPU chip to determine whether to reset this board due to power abnormality.

[0032] Furthermore, the triggering mechanism of the multi-board coordinated reset is: when the core business CPU board is abnormal and the single board reset is completed, the multi-board coordinated reset process is triggered to start.

[0033] Furthermore, the multi-board coordinated resetting specifically includes:

[0034] (1) Multi-board coordinated reset logic judgment:

[0035] Comprehensive logical judgment is performed based on the communication status and the blocking signal: if the communication link of the abnormal business CPU board is interrupted and a blocking signal is generated in the backplane bus, the business CPU board is determined to be a core business CPU board, and then enters the multi-board collaborative reset process; if the abnormal business CPU board only generates an alarm signal and has no blocking status, it is determined to be a non-core business CPU board and the on-site status is retained;

[0036] (2) Dynamic analysis of associated business boards:

[0037] By analyzing the software configuration words, identify other core business CPU boards that have business coupling with the core business CPU board;

[0038] (3) Data consistency maintenance mechanism between boards:

[0039] For business-related boards, execute data synchronization and cleanup strategies to ensure consistency of interaction data between the abnormal core business CPU board and the business-related boards after recovery; the data cleanup goals are: to ensure the integrity and consistency of the DDR external memory and FPGA internal RAM data of the business-related boards to meet the core business requirements of protection and control equipment;

[0040] (4) Communication link reconstruction and mission recovery:

[0041] Coordinate the business-related boards and the abnormal core business CPU board to re-establish the communication link.

[0042] Furthermore, the software configuration word is to inform the management CPU board of the distribution, quantity and deployment of core and non-core businesses of the business CPU boards in the device through a configuration file after the board configuration of the relay protection device is determined. The initial configuration information includes the business board slot, type and business coupling identifier.

[0043] Furthermore, the trigger mechanism for resetting the entire device system is:

[0044] By monitoring the status of the entire device system in real time, the entire device system reset process can be triggered after capturing any of the following trigger conditions:

[0045] (1) Abnormal multi-board coordinated reset: data synchronization between core business CPU boards fails or communication link recovery is invalid;

[0046] (2) Multiple board reset failures: The core business CPU board cannot resume normal function after multiple consecutive board resets due to abnormalities;

[0047] (3) Multiple business CPU boards are abnormal at the same time: multiple business CPU boards in the system have locking signals or communication interruptions at the same time.

[0048] Furthermore, the whole device system reset specifically includes:

[0049] (1) Send the reset command of the whole device system:

[0050] When the whole device system reset process is triggered and started, the management CPU board sends the whole system reset command to all service CPU boards synchronously and independently;

[0051] (2) Reset and initialization of the entire device system:

[0052] The management CPU board guides the system into the reset initialization phase;

[0053] (3) Task reconstruction and system recovery:

[0054] After completing the reset and initialization of the entire device system, all boards start the task scheduling mechanism to restore the running status of the task. At the same time, the abnormality monitoring module starts a new round of status abnormality monitoring.

[0055] Compared with the prior art, the present invention has the following significant advantages:

[0056] (1) The present invention effectively solves the problems of single reset strategy, weak global exception handling capability and insufficient applicability of distributed systems in the prior art by constructing a multi-level reset method and system.

[0057] (2) Through the enable control of the hardware control word and the on-demand configuration of the software configuration word, the system can realize multi-level reset control for independent single boards, coupled multiple boards and the entire device system, flexibly respond to abnormal scenarios, avoid the spread of local abnormalities leading to the failure of non-related business functions or the failure of the entire system, and significantly improve the fault tolerance, reset efficiency and overall reliability and availability of the distributed protection device.

[0058] (3) By combining the board hardware control word and software configuration word, the management CPU board selects the corresponding reset strategy and identifies other core business boards associated with the abnormal core business board. The data integrity maintenance module cleans up the internal cache data after the communication link is rebuilt to ensure the consistency and synchronization of data between boards.

[0059] The present invention is further described in detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 A schematic diagram of the core architecture of a distributed protection and control device in an embodiment.

[0061] Figure 2 The figure is a single board reset logic diagram in one embodiment.

[0062] Figure 3 FIG. 4 is a schematic diagram of a hardware control word circuit in an embodiment.

[0063] Figure 4 The figure is a flowchart of multi-board coordinated reset in one embodiment.

[0064] Figure 5 FIG. 1 is a flowchart of resetting the entire device system in one embodiment. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0066] In one embodiment, a multi-level reset system for a distributed architecture of a protection control device is provided, combining Figure 1 , including management CPU board, business CPU board, IO board, power board and bus board;

[0067] The management CPU board deploys a management module to implement the issuance of reset instructions and process guidance, and receives status information from the service board through a communication link, and analyzes and makes decisions on abnormal status;

[0068] The business CPU board deploys business modules to realize the logical calculation and processing of real-time tasks such as protection, control, and measurement; in abnormal state, it executes board reset or global reset tasks according to the instructions of the management CPU; supports inter-board collaboration under distributed architecture to maintain the consistency of communication data interaction;

[0069] The IO board is used to collect input and control output of switch signals, as well as collect and measure small analog signals;

[0070] The power board is used to provide power to the management CPU board and the service CPU board;

[0071] The bus board is used to realize signal interconnection between the management CPU board and the service CPU board.

[0072] Here, the business board is a flexibly configurable part.

[0073] Here, the protection and control equipment includes but is not limited to relay protection control devices.

[0074] Further, in one of the embodiments, the management CPU board and the service CPU board deploy one or several functional modules among an onboard power supply monitoring module, a communication module, a data integrity maintenance module, a functional integrity reset module and an abnormality monitoring module;

[0075] For the management CPU board: the power monitoring module is used to monitor the abnormal status of the onboard main power supply and local power supply, output a reset signal, and the hardware control word determines whether to reset the board; the communication module is deployed in the management CPU board and the business CPU board, and is used for data interaction between the management CPU board and the business CPU board; the data integrity maintenance module is deployed in the management CPU board, and is used to maintain the consistency of the interaction data between the management CPU and the reset business CPU board; the functional integrity reset module is deployed in the management CPU board, and when the local or entire device system is reset, the entire system is restored to operation through function interface initialization, parameter loading initialization and task scheduling; the abnormality monitoring module is deployed in the management CPU board, monitors the system status in real time, captures single-board abnormalities, multi-board abnormalities and system abnormalities, and triggers corresponding reset strategies.

[0076] For the business CPU board: the power monitoring module is responsible for real-time monitoring of the internal power supply status and outputting a reset signal. The hardware control word determines whether to reset the board; the communication module is responsible for data interaction with the management CPU board or other boards.

[0077] Furthermore, in one of the embodiments, in the distributed protection and control device, the service boards are divided into two types: core service boards and non-core service boards according to the importance of the deployed services. The abnormal classification and dynamic reset strategy can significantly improve the flexibility and efficiency of the reset of the entire device system through accurate identification of abnormal board types, comprehensive analysis of distributed architecture, and flexible adaptation of multi-level reset strategies. See Table 1 below for details.

[0078] Table 1 Abnormal classification and multi-level reset strategy table

[0079]

[0080] Business board classification: According to the functions and importance of each board in the system, the boards are divided into the following two categories: Core business board: a business board that undertakes key functions (such as protection tasks). Its abnormality may cause the loss of key functions of the equipment and bring major risks to the operation of the power system. For example, the protection CPU board belongs to the core business board, and its abnormality will cause the primary equipment to lose its protection function. Non-core business board: a business board that undertakes auxiliary functions (such as measurement and control tasks). Its abnormality has little impact on the overall system. For example, the measurement and control CPU board belongs to the non-core business board.

[0081] Classification of business board abnormalities: According to whether the board in the system is a core business board and whether the business in the distributed architecture is deployed on multiple boards, the types of board abnormalities in the distributed architecture can be divided into the following four types: core business board abnormalities in business single-board deployment, non-core business board abnormalities in business single-board deployment, core business board abnormalities in business multi-board deployment, and non-core business board abnormalities in business multi-board deployment, as shown in Table 1.

[0082] Multi-level reset strategy: In response to the above-mentioned abnormal types and system architecture, a multi-level reset strategy is designed, including level I, level II and level III reset strategies. Level I reset strategy is applicable to the reset management of single-core business boards. Depending on whether the board is deployed with core business, the reset or non-reset strategy is selected; Level II reset strategy is for multi-core business boards. Through coordinated reset or non-reset strategy, the stability of the core business and system consistency are ensured; Level III reset strategy covers system-level reset management. According to the deployment of the system's core business, the system reset or non-reset solution of the entire device can be flexibly selected to comprehensively improve the reliability and fault tolerance of the equipment.

[0083] Reset decision mechanism: This mechanism obtains abnormal business boards in real time by managing the abnormal monitoring module in the CPU, and dynamically analyzes it in combination with the software configuration word to determine the reset strategy of each board in the system architecture. For abnormal objects, first determine the appropriate reset level according to its impact range, and then execute the corresponding reset strategy according to the board attributes to minimize the impact on the normal function of the system.

[0084] In one embodiment, a multi-level reset method of a distributed architecture of a protection control device is provided, the method comprising:

[0085] Single board reset: For single board abnormalities, the output of the reset signal is controlled by the hardware control word, so that the core business CPU board is reset first;

[0086] Multi-board coordinated reset: For other core business CPU boards that are business-coupled with the abnormal core business CPU board, after the abnormal core business CPU board is reset, the management CPU board actively triggers the reset of the other core business CPU boards according to the software configuration word;

[0087] Reset the entire device system: When the single-board reset and multi-board coordinated reset fail to the upper limit or multiple unrelated core business CPU boards are reset at the same time, the management CPU board actively triggers the reset of all other business CPU boards, and then the management CPU board resets itself to start the power-on reset process of the entire device.

[0088] Furthermore, in one of the embodiments, the single-board reset belongs to the level I reset in the multi-level reset method, which is configured on each board in the architecture. On the one hand, the abnormality of the main power supply and the local power supply is detected by the power monitoring circuit inside the board, and a reset signal is sent when the voltage threshold is crossed. If the hardware control word is enabled, the CPU of this board is reset. On the other hand, with the help of a watchdog circuit composed of a monostable trigger, the CPU is monitored for abnormal conditions, and when an abnormality is detected, an alarm or lockout signal is sent to the communication bus. If this board is a core business board, a lockout signal is sent to the backplane bus; if this board is a non-core business board, an alarm signal is sent to the backplane bus, such as Figure 2 shown.

[0089] To assist in deciding the reset strategy of the core business CPU and non-core business CPU, a hardware control word circuit is designed, such as Figure 3 This circuit is an important part of the single-board reset circuit. Its input is the output of the power monitoring circuit. The OE pin in the non-core business CPU board is pulled up to shield the reset signal of this board; the OE pin in the core business CPU board is grounded to enable the reset signal of this board.

[0090] At the same time, the hardware control word circuit enables the core business module to obtain independent local self-reset capability to ensure its reset priority.

[0091] Preferably, in some embodiments, the abnormal reset process includes the following steps:

[0092] Step 1: The abnormal service CPU board enters the reset startup process.

[0093] Step 2: The management CPU board detects abnormal business CPU boards, parses the software configuration word, and performs reset according to the abnormal classification and multi-level reset strategy table. If the business CPU board is deployed with core business, go to step 3; if the business CPU board is deployed with non-core business, go to step 4.

[0094] Step 3: If the service CPU board for deploying core services is an independent board, go to step 5; if the service CPU board for deploying core services has a service-associated board, go to step 6.

[0095] Step 4: Set alarms for non-core services and wait for maintenance.

[0096] Step 5: The abnormal service CPU board that triggers the reset of the management CPU board enters the power-on reset process.

[0097] Step 6: The management CPU board triggers the associated service CPU board to reset actively, and then the coupled multi-core service CPU board enters the power-on reset process.

[0098] Furthermore, in one of the embodiments, the multi-board coordinated reset belongs to the level II reset strategy in the multi-level reset method, which not only resets the abnormal board, but also passively resets the associated service boards with business coupling to ensure the integrity of the functions of the associated service boards. The specific processing flow is as follows: Figure 4 As shown, including:

[0099] (1) Multi-board coordinated reset trigger mechanism: When a service board is abnormal and a single board is reset, the system triggers the multi-board coordinated reset process.

[0100] (2) Multi-board coordinated reset logic judgment: Comprehensive logic judgment is performed based on the communication status and locking signal: If the communication link of the abnormal business board is interrupted and a locking signal is generated in the backplane bus, the board is judged as a core business board and then enters the multi-board coordinated reset process; if the abnormal board only generates an alarm signal and has no locking state, it is judged as a non-core business board, retains the on-site status, and waits for further processing by the operation and maintenance personnel. Finally, the abnormal information is displayed on the human-machine interface.

[0101] (3) Dynamic analysis of associated business boards: By analyzing the software configuration characters, identify other core business boards that are business-coupled with the core business board. The specific steps include:

[0102] Step 1: Determine communication anomaly. The abnormality monitoring module of the management CPU checks the communication link status of each board to determine which board has a communication anomaly.

[0103] Step 2: Control word analysis: Read the coupled board information in the software configuration word, and further confirm the other board information related to the abnormal service board.

[0104] Here, the software configuration word includes the service board slot, type and service coupling identifier.

[0105] Service board slot: used to identify the position of the service CPU board on the bus board.

[0106] Service board type: used to distinguish between core service boards and non-core service boards.

[0107] Business coupling identifier: describes the coupling relationship between business boards.

[0108] Step 3: Reset strategy selection. According to the control word analysis result, the management module of the management CPU board identifies the priority and service coupling status of the abnormal board, selects the corresponding reset strategy according to the abnormal classification and multi-level reset strategy table (such as Table 1 above, which can also be customized and fall within the protection scope of the present invention), and updates the reset strategy information in the software control word.

[0109] (4) Inter-board data consistency maintenance mechanism: For business-related boards, execute data synchronization and cleaning strategies to ensure that after the abnormal core business board is restored, the interactive data between it and the business-related boards remains consistent. Data cleaning goal: Ensure the integrity and consistency of the DDR external memory and FPGA internal RAM data of the business-related boards to meet the core business requirements of protection and measurement and control devices. Protection business data includes raw sampling data such as voltage and current instantaneous values ​​collected in real time, fault-triggered recorded waveform data, protection action setting parameters (such as trip time, action threshold, current transformer and voltage transformer ratio), and phase calculation intermediate data based on fast Fourier transform (FFT), providing accurate input for protection logic. Measurement and control business data involves telesignaling position record (SOE) to record the time series of equipment status changes, telemetered values ​​(such as bus voltage, line current, active power, reactive power and frequency, etc.), circuit breaker and disconnector status signals, and event control data, which are used for the judgment and execution of measurement and control logic. Management business data includes initialization dynamic configuration (such as module parameter loading and communication topology allocation), event-triggered recording data, communication link status information, and clock synchronization and correction parameters in distributed devices to ensure the global consistency and dynamic collaboration capabilities of the system. By cleaning and rebuilding the above key data, the linkage and reliability of each business module after reset are guaranteed to fully meet the operation requirements in complex business scenarios. Scope of impact: The affected business-related boards usually include management boards and other core business boards.

[0110] Here, the data cleaning goal is to ensure the integrity and consistency of the DDR external memory of the business-related board and the FPGA internal RAM data.

[0111] Protection service data: including instantaneous value sampling data of voltage and current, fault recording data, dynamic adjustment parameters of protection setting area, and intermediate data of phase angle calculation used for protection logic, etc.;

[0112] Measurement and control business data: including remote signal change records (SOE), remote measurement values ​​(such as real-time data such as bus voltage and line current), switchgear status signals (such as circuit breaker position signals), and event records and control command feedback of measurement and control devices;

[0113] Management business data: dynamic configuration information related to system initialization, long-term recording data triggered by events, communication link status monitoring information, and configuration information and deviation correction parameters of the synchronous clock between boards.

[0114] (5) Communication link reconstruction and task recovery: The management board issues instructions to coordinate the business-related boards and the abnormal core business board to re-establish the communication link.

[0115] Here, the reconstruction of the communication link includes: 1) Data handshake. Confirm that the basic communication status between the abnormal board and the associated board is restored; 2) Task synchronization. Update and synchronize the task status of the associated board to re-adapt it to the business needs of the abnormal core business board. Through the reconstruction of the communication link, ensure that the data interaction between the abnormal core business board and the business-related board is restored to normal, and complete the inter-board collaborative reset.

[0116] In summary, the purpose of multi-board collaborative reset is to effectively isolate abnormal core business boards through the inter-board collaboration mechanism to prevent them from affecting the normal operation of other business boards; to ensure business collaboration and data consistency between associated boards and abnormal boards through precise data synchronization and communication link reconstruction; and to shorten the reset recovery time between business-associated boards and abnormal core boards to improve user experience.

[0117] Furthermore, in one of the embodiments, the whole device system reset belongs to the level III reset strategy in the multi-level reset method, which is specially used to solve the complex abnormal scenarios that cannot be handled by single-board reset (level I reset) and multi-board coordinated reset (level II reset), aiming to ensure the functional recovery of the entire distributed system. The following is a detailed description of the whole device system reset process, see Figure 5 As shown:

[0118] (1) The whole device system reset trigger mechanism: The management CPU monitors the system status in real time, captures any of the following trigger conditions, and then triggers the functional complete reset module to start the whole system reset process: 1. Multi-board coordinated reset exception: When the data synchronization between the core business boards fails or the communication link recovery is invalid, the whole system reset is triggered; 2. Multiple board reset failures: The core business board cannot restore normal function after multiple consecutive single board resets due to abnormalities; 3. Multiple business boards are abnormal at the same time: Multiple business boards in the system have locking signals or communication interruptions at the same time, which seriously affects the stability of system operation.

[0119] (2) Sending a reset command for the entire device: When the reset conditions for the entire system are met, the management CPU sends a reset command to all service boards independently to enter the reset process. All boards execute the reset operation in parallel, covering the stages of hardware resource reset, cache data cleanup, and task scheduling module loading.

[0120] (3) System reset initialization of the whole device: After the reset of the whole system is triggered, the management CPU guides the system into the reset initialization phase, which specifically includes the following steps: Function interface initialization: Initialize the basic hardware resources of all boards, including communication interfaces, storage modules, external device interfaces, etc. Parameter loading initialization: After the function interface initialization is completed, the management CPU sends parameter initialization commands to each board to ensure that the configuration parameters required for system operation are loaded into each module.

[0121] (4) Task reconstruction and system recovery: After the entire system is reset and initialized, all boards start the task scheduling mechanism and restore the task's running status. At this point, the system enters the normal operation phase, and the abnormality monitoring module starts a new round of status abnormality monitoring to continuously ensure the stability and reliability of the system.

[0122] Preferably, in some embodiments, the power-on reset process includes the following steps:

[0123] Step 1: The management CPU board sends an initialization command to the business CPU board. After receiving the command, the business CPU board starts the initialization work, calls the initialization function interface of all functional modules, and returns an initialization success message after completing the initialization.

[0124] Step 2: The management CPU board sends a parameter initialization command to the business CPU board. After receiving the command, the business CPU board calls the parameter initialization interface of each functional module.

[0125] Step 3: The management CPU board sends an initialization end command to the business CPU board. After receiving the command, the business CPU board completes all initialization processes and starts the task scheduling mechanism to start the task running. If the startup fails, the management CPU board triggers a system reset of the entire device and enters step 1. There is an upper limit on the number of repetitions, which can be configured as needed.

[0126] As a specific example, in one of the embodiments, the present invention is further verified and explained in detail.

[0127] The following takes the specific distributed architecture of the protection, measurement and control integrated device as an example and further describes the specific implementation mode of the present invention in detail in conjunction with the accompanying drawings. However, this example is only used to illustrate the present invention in detail and does not limit the scope of application of the present invention.

[0128] 1. Specific configuration of the integrated protection and measurement device:

[0129] The device is configured from left to right on the backplane slots as management board (slot 1), protection board 1 (slot 2), protection board 2 (slot 3), measurement and control board (slot 4), and protection board 3 (slot 5). The configuration of the management board is as follows: Figure 1As shown, it is a non-core business board, and the hardware control word circuit is configured to not reset; protection board 1 and protection board 2 are core business boards, with common core business, and the hardware control word circuits of the two boards are configured to reset the board; the measurement and control board is a non-core business board, and the hardware control word circuit is configured to not reset; protection board 3 is a core business board, and has no business coupling with protection board 1 and protection board 2, and the hardware control word circuit is configured to reset the board. Assume that the main power supply in each board is +5V, the voltage offset threshold is +4.7V, the local power supply is 3.3V, and the voltage offset threshold is +3.0V. The initial configuration of the software configuration word of each board is as follows:

[0130] Management board (corresponding to Figure 1 Management CPU board in): Business board slot (010: slot 1), business board type (0: non-core business), board coupling identifier (00000: no business coupling slot), reset policy (00: no reset).

[0131] Protection plate 1 (corresponding to Figure 1 Service CPU board 1): Service board slot (010: slot 2), service board type (1: core business), board coupling identifier (00100: coupled with slot 3 business), reset strategy (10: single board reset + multi-board coordinated reset).

[0132] Protection board 2 (based on business CPU board 1, and so on to business CPU board 2): business board slot (011: slot 3), business board type (1: core business), board coupling identifier (01000: coupled with slot 2 business), reset policy (10: single board reset + multi-board coordinated reset).

[0133] Measurement and control board (based on business CPU board 2, and so on to business CPU board 3): business board slot (100: slot 4), business board type (0: non-core business), board coupling identifier (00000: no business coupling slot), reset policy (00: no reset).

[0134] Protection board 3 (based on business CPU board 3, and so on to business CPU board 4): business board slot (101: slot 5), business board type (1: core business), board coupling identifier (00000: no business coupling slot), reset policy (10: single board reset).

[0135] 2. Reset the board under level I policy:

[0136] If the voltage of the main power supply of the protection board 3 drops below 4.7V, or the local power supply drops below 3.0V. At this time, the power monitoring module in the protection board 3 detects that the main power supply is abnormal or the local power supply is abnormal, and generates a reset signal. The reset signal is judged as the core business board reset signal by the hardware control word circuit, triggering the single board reset and resetting the protection board 3. At the same time, the management board determines the exit lock through the backplane bus, combined with the abnormal board configuration word information, the configuration word information of the protection board 3 at this time is: business board type (1: core business), reset strategy (10: single board reset), board coupling identification (00000: no coupling business slot), it can be determined that the protection board 3 is a protection board with no core business coupling, and only the single board reset is performed.

[0137] Then the management CPU guides the protection board 3 to restart: 1. The management board sends an initialization command to the protection board 3. After receiving the command, the protection board 3 calls the initialization function interface of all modules and returns the initialization success message after completion. 2. The management board sends a parameter initialization command to the protection board 3. After receiving the command, the protection board 3 calls the parameter initialization interface of each module. 3. The management board sends an initialization end and formal operation command to the protection board 3. After receiving the command, the protection board 3 completes all initialization processes, starts the task scheduling mechanism, starts the task running, and completes the single board reset.

[0138] 3. No reset under level I strategy:

[0139] If the voltage of the main power supply of the measurement and control board drops below 4.7V, or the local power supply drops below 3.0V. At this time, the power supply monitoring module in the measurement and control board detects that the main power supply is abnormal or the local power supply is abnormal, and generates a reset signal. The reset signal is judged as a non-core business board reset signal by the hardware control word circuit, and no longer triggers the reset of the single board, only an alarm signal is issued. At this time, the configuration word information of the abnormal measurement and control board is: business board type (0: non-core business), reset strategy (00: no reset), board coupling identification (00000: no coupled business slot). It can be determined that the abnormal board is a measurement and control board with non-core business, and no reset is performed under the global strategy. After that, the non-faulty board continues to operate normally, and the device displays the abnormal alarm of the measurement and control board on the human-machine interface, waiting for on-site maintenance personnel to handle.

[0140] 4. Multi-board coordinated reset under Level II strategy

[0141] If the voltage of the main power supply of protection board 1 drops below 4.7V, or the local power supply drops below 3.0V. At this time, the power monitoring module in protection board 1 detects that the main power supply is abnormal or the local power supply is abnormal, and generates a reset signal. The reset signal is judged as the core business board reset signal by the hardware control word circuit, triggering the single board reset and resetting the abnormal protection board 1. At the same time, the management CPU board determines the exit lock through the backplane bus, combined with the abnormal board configuration word information, the configuration word information of the abnormal protection board 1 at this time is: business board type (1: core business), reset strategy (10: single board reset + multi-board collaborative reset), board coupling identification (00100: coupled with slot 3 business). It can be determined that the abnormal board is the protection CPU board 1 with core business, and there is a core business coupling with the protection board 2. Therefore, while the protection board 1 starts the single board reset, the protection board 2 performs a multi-board collaborative reset.

[0142] The management CPU guides the protection board 1 to restart: 1. The management board sends an initialization command to the protection board 1. After receiving the command, the protection board 1 calls the initialization function interface of all modules and returns the initialization success message after completion. 2. The management CPU sends a parameter initialization command to the protection board 1. After receiving the command, the protection board 1 calls the parameter initialization interface of each module. 3. The management board sends an initialization end and formal operation command to the protection board 1. After receiving the command, the protection board 1 completes all initialization processes, starts the task scheduling mechanism, starts the task running, and completes the single board reset of the protection board 1.

[0143] The management CPU guides the protection board 2 to enter the multi-board coordinated reset judgment. If it is judged that the communication link between the protection board 1 and the management CPU board is interrupted, and the protection action exit is locked on the backplane. At this time, it is judged that the protection board 1 is abnormal, and the board coupling mark (00100: coupled with the slot 3 service) is obtained through the software configuration word, indicating that the protection board 2 has business coupling. Then clear the communication data cache of the DDR in the protection board 2 and the RAM in the FPGA to always ensure data synchronization with the reset protection board 1. Finally, guide the protection board 2 after the multi-board coordinated reset and the protection board 1 after the single board reset to restore link communication.

[0144] 5. Reset the entire device system under Level III strategy

[0145] If there are multiple abnormalities in the single-board reset in the Level I strategy and the Level II strategy (such as the protection board 3 restarts and resets multiple times in a short period of time, or the reset fails, etc.), the multi-board coordinated reset is abnormal (such as the data of protection board 1 is always inconsistent within a certain period of time, or the data interaction with protection board 2 is abnormal, etc.), and multiple business boards are abnormal at the same time (such as protection board 1, protection board 2, protection board 3 and the measurement and control board are abnormal at the same time, etc.).

[0146] The functional integrity reset module on the management board sends a system-wide reset instruction. At this time, the reset strategy is upgraded to the level III reset strategy (whole system reset), guiding the system to restart: 1. The management CPU sends an initialization command to all slave boards. After receiving the command, the slave board starts the initialization work and calls the initialization function interface of all modules. After completion, it returns the initialization success message. 2. The management CPU sends a parameter initialization command to all slave boards. After receiving the command, the slave board calls the parameter initialization interface of each module. 3. The management CPU sends an initialization end and formal operation command to all slave boards. After receiving the command, the slave board completes all initialization processes, starts the task scheduling mechanism, starts task running, and the system restart is completed.

[0147] In summary, the present invention adopts a combination of hardware control words and software configuration words to determine the exception handling method of different functional business modules of the device. The reset level is selected according to the impact range of the exception, combined with the hierarchical identification and dynamic reset strategy of the abnormal board type, and the local reset of single or multiple boards is prioritized. Through communication link reconstruction and data consistency maintenance, the business collaboration and status synchronization between modules after reset are ensured. When the local reset fails multiple times or multiple boards are abnormal at the same time, the whole device system reset is initiated. Through the multi-level reset control of single-board reset, multi-board collaborative reset and whole device system reset, the risk of whole machine shutdown is effectively reduced and the system availability is improved.

[0148] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A distributed architecture multi-level reset system for protection and control equipment, characterized in that: The system is designed based on a distributed architecture, including a management CPU board, a business CPU board, an IO board, a power board and a bus board; The management CPU board deploys a management module to implement the issuance of reset instructions and process guidance, and receives status information from the service board through a communication link, and analyzes and makes decisions on abnormal status; The business CPU board deploys business modules to implement logical calculation and processing of various real-time tasks; The IO board is used to collect input and control output of switch signals, as well as collect and measure small analog signals; The power board is used to provide power to the management CPU board and the service CPU board; The bus board is used to realize signal interconnection between the management CPU board and the service CPU board.

2. The distributed architecture multi-level reset system for protection and control equipment according to claim 1, characterized in that: The management CPU board and the service CPU board deploy one or more functional modules among an onboard power supply monitoring module, a communication module, a data integrity maintenance module, a functional integrity reset module and an abnormality monitoring module; The power monitoring module is used to monitor the abnormal status of the onboard main power supply and local power supply, output a reset signal, and the hardware control word determines whether to reset the board; The communication module is deployed in the management CPU board and the service CPU board, and is used for data interaction between the management CPU board and the service CPU board; The data integrity maintenance module is deployed in the management CPU board and is used to maintain the consistency of the interaction data between the management CPU and the reset business CPU board; The functional complete reset module is deployed in the management CPU board. When a local or entire device system is reset, it restores the entire system through function interface initialization, parameter loading initialization and task scheduling; The abnormality monitoring module is deployed in the management CPU board, monitors the system status in real time, captures single-board abnormalities, multi-board abnormalities and system abnormalities, and triggers corresponding reset strategies.

3. A multi-level reset method for a distributed architecture of a protection control device based on the system according to any one of claims 1 to 2, characterized in that: The method comprises: Single board reset: For single board abnormalities, the output of the reset signal is controlled by the hardware control word, so that the core business CPU board is reset first; Multi-board coordinated reset: For other core business CPU boards that are business-coupled with the abnormal core business CPU board, after the abnormal core business CPU board is reset, the management CPU board actively triggers the reset of the other core business CPU boards according to the software configuration word; Reset the entire device system: When the single-board reset and multi-board coordinated reset fail to the upper limit or multiple unrelated core business CPU boards are reset at the same time, the management CPU board actively triggers the reset of all other business CPU boards, and then the management CPU board resets itself to start the power-on reset process of the entire device.

4. The multi-level reset method of the distributed architecture of the protection and control equipment according to claim 3 is characterized in that: The single board resetting specifically includes: The power monitoring circuit inside the board detects the abnormality of the main power supply and the local power supply. If the voltage threshold is exceeded, a reset signal is sent to the CPU of the board to reset the CPU of the board. Monitor the abnormal status of the CPU of this board. When an abnormality is detected, send an alarm or lockout signal to the communication bus. Specifically: if this board is a core business CPU board, send a lockout signal to the backplane bus; if this board is a non-core business CPU board, send an alarm signal to the backplane bus.

5. The multi-level reset method of the distributed architecture of the protection and control equipment according to claim 4 is characterized in that: The reset signal is output through a hardware control word control. The hardware control word is to place a pull-up or pull-down resistor at the bus board position corresponding to the business CPU board. One end of the pull-up or pull-down resistor is connected to the positive power supply or power ground, and the other end is connected to the reset pin of the CPU chip after a logical AND operation with the reset output of the power monitoring module of the business CPU board to determine whether to reset the board due to power abnormality.

6. The multi-level reset method of the distributed architecture of the protection and control equipment according to claim 3 is characterized in that: The trigger mechanism of the multi-board coordinated reset is: when the core business CPU board is abnormal and the single board reset is completed, the multi-board coordinated reset process is triggered to start.

7. The multi-level reset method of the distributed architecture of the protection and control equipment according to claim 3 is characterized in that: The multi-board coordinated resetting specifically includes: (1) Multi-board collaborative reset logic judgment: Comprehensive logical judgment is performed based on the communication status and the blocking signal: if the communication link of the abnormal business CPU board is interrupted and a blocking signal is generated in the backplane bus, the business CPU board is determined to be a core business CPU board, and then enters the multi-board collaborative reset process; if the abnormal business CPU board only generates an alarm signal and has no blocking status, it is determined to be a non-core business CPU board and the on-site status is retained; (2) Dynamic analysis of associated business boards: By analyzing the software configuration words, identify other core business CPU boards that have business coupling with the core business CPU board; (3) Data consistency maintenance mechanism between boards: For business-related boards, execute data synchronization and cleanup strategies to ensure consistency of interaction data between the abnormal core business CPU board and the business-related boards after recovery; the data cleanup goals are: to ensure the integrity and consistency of the DDR external memory and FPGA internal RAM data of the business-related boards to meet the core business requirements of protection and measurement and control devices; (4) Communication link reconstruction and mission recovery: Coordinate the business-related boards and the abnormal core business CPU board to re-establish the communication link.

8. The multi-level reset method of the distributed architecture of the protection and control equipment according to claim 7, characterized in that: The software configuration word is used to inform the management CPU board of the distribution, quantity and deployment of core and non-core businesses of the business CPU board in this device through a configuration file after the board configuration of the protection control device is determined. The initial configuration information includes the business board slot, type and business coupling identifier.

9. The multi-level reset method of the distributed architecture of the protection and control equipment according to claim 3, characterized in that: The trigger mechanism for resetting the entire device system is: By monitoring the status of the entire device system in real time, the entire device system reset process can be triggered after capturing any of the following trigger conditions: (1) Abnormal multi-board coordinated reset: data synchronization between core business CPU boards fails or communication link recovery is invalid; (2) Multiple board reset failures: The core business CPU board cannot resume normal function after multiple consecutive board resets due to abnormalities; (3) Multiple business CPU boards are abnormal at the same time: multiple business CPU boards in the system have locking signals or communication interruptions at the same time.

10. The multi-level reset method of the distributed architecture of the protection and control equipment according to claim 9, characterized in that: The whole device system reset specifically includes: (1) Send the reset command of the whole device system: When the whole device system reset process is triggered and started, the management CPU board sends the whole system reset command to all service CPU boards synchronously and independently; (2) Reset and initialization of the entire device system: The management CPU board guides the system into the reset initialization phase; (3) Task reconstruction and system recovery: After completing the reset and initialization of the entire device system, all boards start the task scheduling mechanism to restore the running status of the task. At the same time, the abnormality monitoring module starts a new round of status abnormality monitoring.

Citation Information

Patent Citations

  • Multi-core system resetting method, device and equipment and readable storage medium

    CN114527857A

  • Method and system for realizing one board forced resetting

    CN101021740A

  • Remote sensing satellite ground station monitoring system and method

    CN115549751A

  • X shield network security emergency disposal server platform

    CN117040887A

  • Method and system for actively protecting and automatically recording reasons for abnormal restart of CPU (Central Processing Unit)

    CN117743025A

Cited By

  • Method for recovering VPX case board card with electronic identity tag

    CN120179468A

  • Design method of novel backboard bus

    CN120743827A

  • Fine-grained multi-level management system for audio frequency integrated signal processing

    CN120803738A

  • Resetting method and system for reading abnormity of SD (Secure Digital) card

    CN121542093A

  • A method and system for resetting SD card read errors

    CN121542093B