A multi-level reset method and system for a distributed architecture of a protection control device

Through the combination of hardware control words and software configuration words, multi-level reset control of distributed protection control devices is realized, solving the problem of single reset strategy and mismatch in the existing technology, and improving the reliability and availability of the system.

CN119937433BActive Publication Date: 2025-07-25NANJING UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510428429.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the prior art, the reset strategy of distributed protection control equipment is single and lacks flexibility, making it difficult to select the reset range according to the type and severity of the exception, resulting in an expansion of the reset range, an increase in time, and reducing system availability; lack of a priority mechanism for core business modules, and non-core business abnormalities may lead to core business interruption; inter-board data status mismatch affects the normal operation of the system.

Method used

The combination of hardware control words and software configuration words is adopted to realize multi-level reset control, including single-board reset, multi-board collaborative reset and complete device system reset. By managing the CPU board, abnormal analysis and decision-making are carried out to ensure priority reset of core business modules and maintain data consistency between boards.

Benefits of technology

Improve the fault tolerance and availability of distributed protection devices, avoid local abnormal spread, and cause failure of non-associated business functions or system-wide failure, ensuring the reliability and availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937433B_ABST
    Figure CN119937433B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-level reset method and system for a protection control device distributed architecture, which is particularly applicable to a protection control device with a distributed architecture. By combining a hardware control word and a software configuration word, the exception handling method for different functional service modules of the device is determined. The reset level is selected according to the scope of the exception impact. Combining the hierarchical identification and dynamic reset strategy of the exception board type, local reset of a single board or multiple boards is preferred. Through communication link reconstruction and data consistency maintenance, the business cooperation and status synchronization between modules after reset are ensured. When the local reset fails multiple times or multiple boards are abnormal at the same time, the entire device system reset is started. Through multi-level reset control of single board reset, multi-board collaborative reset and entire device system reset, the risk of the entire machine downtime is effectively reduced and the system availability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of embedded protection control equipment for power systems, and particularly relates to a multi-level reset method and system for the distributed architecture of protection control equipment. Background Art

[0002] With the development of new power systems, the functions of protection control equipment are becoming increasingly complex, and the requirements for protection control coordination are getting higher and higher. Protection control equipment designed based on a distributed architecture has become a research hotspot. Such devices usually adopt a modular design and deploy different protection, control, and measurement tasks in a distributed architecture. However, the distributed architecture design makes the fault management and reset control of each module complex. In the prior art, common reset mechanisms mainly include power monitoring reset, module-independent reset, and multi-core cooperation reset, etc. However, these mechanisms are insufficient in adapting to the distributed architecture, and there are problems such as a single reset strategy, insufficient flexibility, and insufficient state feedback, making it difficult to meet the requirements of highly reliable distributed protection control equipment.

[0003] In response to the above problems, Chinese Patent CN114527857A proposes a multi-core system reset method. Its technical solution is as follows: After the multi-core system performs a bus reset, obtain the reset information of the target processor; after determining that the target processor has been reset successfully, detect whether all the remaining processors in the multi-core system have been reset successfully through the target processor; if so, end the multi-core reset; if it is detected that any processor has not been reset successfully, control the register corresponding to the processor that has not been reset successfully to perform a re-reset through the target processor until all the remaining processors in the multi-core system have been reset successfully, and then end the multi-core reset.

[0004] However, the prior art has the following disadvantages:

[0005] (1) The reset strategy is single and lacks flexibility

[0006] The trigger condition of the reset strategy in the prior art is single, usually relying on simple hardware monitoring signals, and the reset mode is fixed. It is difficult to select the reset range according to the type and severity of the abnormality. This "one-size-fits-all" reset method lacks a fine hierarchical design, which is particularly obvious in the multi-core business board architecture, and may lead to an enlarged reset range and increased reset time, thereby reducing the system availability.

[0007] (2) Lack of a priority mechanism for core business modules

[0008] The prior reset technology has insufficient discrimination between core business modules and non-core business modules. In practical applications, if the entire device system is reset after a non-core business abnormality, it will cause the interruption of the core business operation, bringing certain risks to the operation of the power system. Whether the entire device system needs to be reset after a non-core business abnormality is debatable.

[0009] (3)Mismatch of data status between boards after reset

[0010] The prior art lacks effective management of the status synchronization of communication modules between boards in a distributed architecture. During the reset and restart process after an anomaly, data disorders may occur between associated modules, which may further lead to the failure of module cooperation. In severe cases, it may even cause incorrect or refused operation of protection, affecting the normal operation of the system. Summary of the Invention

[0011] The object of the present invention is to provide a multi-level reset method and system for the distributed architecture of a protection control device to address the problems existing in the above prior art. Through the hardware control word, the reset enable control of different service CPU boards is realized, and combined with the software configuration word of the entire device, the multi-level reset control of independent single boards, coupled multi-boards and the entire system of the distributed protection control device is realized, preventing the failure of non-associated service functions or the failure of the entire machine service caused by local anomalies in the system, ensuring the effectiveness of the distributed deployment of the entire machine service, and further improving the overall reliability and availability of the device system.

[0012] On the one hand, a technical solution to achieve the object of the present invention is to provide a multi-level reset system for the distributed architecture of a protection control device. The system is designed based on a distributed architecture and includes a management CPU board, service CPU boards, IO boards, power supply boards, and bus boards;

[0013] The management CPU board deploys a management module, which is used to issue reset instructions and guide the process, and receives status information from service boards through a communication link, and analyzes and makes decisions on abnormal states;

[0014] The service CPU boards deploy service modules, which are used to perform logical calculations and processing of various real-time tasks;

[0015] The IO boards are used to collect and input switch quantity signals and control outputs, as well as collect and measure analog small signals;

[0016] The power supply boards are used to provide power for the management CPU board and service CPU boards;

[0017] The bus boards are used to realize signal interconnection between the management CPU board and service CPU boards.

[0018] Further, the management CPU board and service CPU boards deploy one or several functional modules among on-board power supply monitoring modules, communication modules, data integrity maintenance modules, function complete reset modules, and anomaly monitoring modules;

[0019] The power supply monitoring module is used to monitor the abnormal states of the on-board main power supply and local power supply, output a reset signal, and determine whether to reset the board by the hardware control word;

[0020] The communication module is deployed in the management CPU board and the service CPU board, and is used for data interaction between the management CPU board and the service CPU board;

[0021] The data integrity maintenance module is deployed in the management CPU board, and is used for maintaining the consistency of the interaction data between the management CPU and the reset service CPU board;

[0022] The function complete reset module is deployed in the management CPU board. When the local or entire device system is reset, it initializes through function interfaces, initializes parameter loading, and schedules tasks to restore the operation of the entire system;

[0023] The exception monitoring module is deployed in the management CPU board, monitors the system status in real time, captures single-board exceptions, multi-board exceptions, and system exceptions, and triggers corresponding reset strategies.

[0024] On the other hand, a multi-level reset method for the distributed architecture of a protection control device is provided. The method includes:

[0025] Single-board reset: For single-board exceptions, the output of the reset signal is controlled through a hardware control word, so that the core service CPU board is preferentially reset;

[0026] Multi-board collaborative reset: For other core service CPU boards that have service coupling with the exception core service CPU board, after the exception core service CPU board is reset, the management CPU board actively triggers the reset of the other core service CPU boards according to the software configuration word;

[0027] Entire device system reset: When the single-board reset and the multi-board collaborative reset fail to reach the upper limit number of times or when multiple non-associated core service CPU boards are reset simultaneously, the management CPU board actively triggers the reset of all the remaining service CPU boards, and then the management CPU board self-resets to start the power-on reset process of the entire machine.

[0028] Further, the single-board reset specifically includes:

[0029] Detect the abnormalities of the main power supply and the local power supply through the power monitoring circuit inside the board. If the voltage threshold is exceeded, a reset signal is sent to the CPU of this board to reset the CPU of this board;

[0030] Monitor the abnormal state of the CPU of this board. When an abnormality is detected, an alarm or lock signal is sent to the communication bus. Specifically: if this board is a core service CPU board, a lock signal is sent to the backplane bus; if this board is a non-core service CPU board, an alarm signal is sent to the backplane bus.

[0031] Further, the reset signal is controlled and output through a hardware control word. The hardware control word is formed by placing a pull-up or pull-down resistor at the corresponding bus board position of the service CPU board. One end of the pull-up or pull-down resistor is connected to the positive power supply or the power ground, and the other end is logically ANDed with the reset output of the power monitoring module of the service CPU board and then connected to the reset pin of the CPU chip to determine whether to reset the board in case of power anomaly.

[0032] Further, the triggering mechanism for the multi-board cooperative reset is as follows: when the core service CPU board is abnormal and the single-board reset is completed, the multi-board cooperative reset process is triggered and started.

[0033] Further, the multi-board cooperative reset specifically includes:

[0034] (1) Multi-board cooperative reset logic judgment:

[0035] Based on the communication status and the locking signal, comprehensive logical judgment is performed: if the communication link of the abnormal service CPU board is interrupted and a locking signal is generated in the backplane bus, then this service CPU board is determined to be the core service CPU board, and the subsequent multi-board cooperative reset process is entered; if the abnormal service CPU board only generates an alarm signal and has no locked state, then it is determined to be a non-core service CPU board, and the on-site state is retained.

[0036] (2) Dynamic parsing of associated service boards:

[0037] By analyzing the software configuration word, other core service CPU boards that have service coupling with the core service CPU board are identified.

[0038] (3) Inter-board data consistency maintenance mechanism:

[0039] For service-associated board cards, a data synchronization and cleaning strategy is executed to ensure that the interaction data between the abnormal core service CPU board and the service-associated board cards remains consistent after the abnormal core service CPU board is restored; among them, the data cleaning target is to ensure the integrity and consistency of the DDR external memory and the FPGA internal RAM data of the service-associated board cards, meeting the core service requirements of the protection and control device.

[0040] (4) Communication link reconstruction and task recovery:

[0041] Coordinate the service-associated board cards to re-establish a communication link with the abnormal core service CPU board.

[0042] Further, the software configuration word, after the board card configuration of the relay protection device is determined, informs the management CPU board in the form of a configuration file about the distribution, quantity, and the deployment of core services and non-core services of the service CPU boards in this device. The initial configuration information includes the service board slot positions, types, and service coupling identifiers.

[0043] Furthermore, the trigger mechanism for the reset of the entire device system is as follows:

[0044] By continuously monitoring the status of the entire device system, after capturing any of the following trigger conditions, the reset process of the entire device system is triggered and started:

[0045] (1) Abnormal multi-board collaborative reset: Data synchronization failure or ineffective communication link restoration between the core service CPU boards;

[0046] (2) Multiple single-board reset failures: The core service CPU board still cannot restore normal functions after being reset multiple times continuously due to anomalies;

[0047] (3) Simultaneous anomalies of multiple service CPU boards: Locking signals or communication interruptions occur simultaneously on multiple service CPU boards in the system.

[0048] Furthermore, the reset of the entire device system specifically includes:

[0049] (1) Sending the reset instruction of the entire device system:

[0050] After the reset process of the entire device system is triggered and started, the management CPU board synchronously sends the entire system reset instruction independently to all service CPU boards;

[0051] (2) Initialization of the reset of the entire device system:

[0052] The management CPU board guides the system to enter the reset initialization stage;

[0053] (3) Task reconstruction and system restoration:

[0054] After the reset initialization of the entire device system is completed, all boards start the task scheduling mechanism, restore the running state of the tasks, and at the same time, the anomaly monitoring module starts a new round of status anomaly monitoring.

[0055] Compared with the prior art, the remarkable advantages of the present invention are as follows:

[0056] (1) By constructing a multi-level reset method and system, the present invention effectively solves the problems of single reset strategy, weak global anomaly handling ability, and insufficient applicability to distributed systems in the prior art.

[0057] (2) Through the enabling control of the hardware control word and the on-demand configuration of the software configuration word, the system can achieve multi-level reset control for independent single boards, coupled multi-boards, and the entire device system, flexibly handle anomaly scenarios, avoid the spread of local anomalies resulting in the failure of non-related service functions or the failure of the entire system, and significantly improve the fault tolerance, reset efficiency, and overall reliability and availability of the distributed protection device.

[0058] (3) Through the combination of the board hardware control word and the software configuration word, the management CPU board selects the corresponding reset strategy, identifies other core business board cards associated with the abnormal core business board, and clears the internal cache data after the communication link is rebuilt through the data integrity maintenance module to ensure the consistency and synchronization of the data between the boards.

[0059] The present invention will be further described in detail below with reference to the accompanying drawings. Brief Description of the Drawings

[0060] Figure 1 It is a schematic diagram of the core architecture of the distributed protection control device in one embodiment.

[0061] Figure 2 It is a single-board reset logic diagram in one embodiment.

[0062] Figure 3 It is a schematic diagram of the hardware control word circuit in one embodiment.

[0063] Figure 4 It is a multi-board collaborative reset flow chart in one embodiment.

[0064] Figure 5 It is a whole-device system reset flow chart in one embodiment. Specific Embodiments

[0065] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0066] In one embodiment, a multi-level reset system for the distributed architecture of a protection control device is provided. Combining Figure 1 , it includes a management CPU board, a service CPU board, an IO board, a power supply board and a bus board;

[0067] The management CPU board deploys a management module for issuing reset instructions and guiding processes, receiving status information from service boards through a communication link, and analyzing and making decisions on abnormal statuses;

[0068] The service CPU board deploys a service module for performing logical calculations and processing of real-time tasks such as protection, control, and measurement; in the event of an abnormal status, it performs board reset or global reset tasks according to the instructions of the management CPU; it supports board-to-board collaboration in a distributed architecture and maintains the consistency of communication data interaction;

[0069] The IO board is used to implement the acquisition input and control output of switch quantity signals, as well as the acquisition and measurement of analog small signals;

[0070] The power supply board is used to supply power to the management CPU board and the service CPU board;

[0071] The bus board is used to realize signal interconnection between the management CPU board and the service CPU board.

[0072] Here, the service board is a flexible and configurable part.

[0073] Here, the protection control device includes but is not limited to a relay protection control device.

[0074] Further, in one embodiment, the management CPU board and the service CPU board are deployed with one or several functional modules among an on-board power supply monitoring module, a communication module, a data integrity maintenance module, a function complete reset module, and an exception monitoring module;

[0075] For the management CPU board: The power supply monitoring module is used to monitor the abnormal states of the on-board main power supply and the local power supply, output a reset signal, and determine whether to reset this board according to the hardware control word; The communication module is deployed in the management CPU board and the service CPU board and is used for data interaction between the management CPU board and the service CPU board; The data integrity maintenance module is deployed in the management CPU board and is used to maintain the consistency of the interaction data between the management CPU and the reset service CPU board; The function complete reset module is deployed in the management CPU board. When the local or the whole device system is reset, it initializes through a function interface, initializes parameter loading, and schedules tasks to restore the operation of the whole system; The exception monitoring module is deployed in the management CPU board, monitors the system state in real time, captures single-board exceptions, multi-board exceptions, and system exceptions, and triggers corresponding reset strategies.

[0076] For the service CPU board: The power supply monitoring module is responsible for monitoring the internal power supply state in real time, outputting a reset signal, and determining whether to reset this board according to the hardware control word; The communication module is responsible for data interaction with the management CPU board or other boards.

[0077] Further, in one embodiment, in the distributed protection control device, the service board cards are divided into two types: core service boards and non-core service boards according to the importance of the deployed services. The exception grading and dynamic reset strategy can significantly improve the flexibility and efficiency of the whole device system reset through the accurate identification of the abnormal board card type, the comprehensive analysis of the distributed architecture, and the flexible adaptation of the multi-level reset strategy. See Table 1 below.

[0078] Table 1 Exception Grading and Multi-level Reset Strategy Table

[0079]

[0080] Classification of business boards: According to the functions and importance of each board in the system, the boards are classified into the following two categories: Core business boards: Business boards that undertake key functions (such as protection tasks). Their anomalies may lead to the loss of key functions of the device, posing significant risks to the operation of the power system. For example, the protection CPU board belongs to the core business board, and its anomaly will cause the primary equipment to lose its protection function. Non-core business boards: Business boards that undertake auxiliary functions (such as measurement and control tasks). Their anomalies have less impact on the overall system. For example, the measurement and control CPU board belongs to the non-core business board.

[0081] Classification of business board anomalies: According to whether the board in the system is a core business board and whether the business in the distributed architecture is deployed on multiple boards, the types of board anomalies in the distributed architecture can be divided into the following four types: Anomaly of the core business board in single-board deployment, Anomaly of the non-core business board in single-board deployment, Anomaly of the core business board in multi-board deployment, and Anomaly of the non-core business board in multi-board deployment, as shown in Table 1.

[0082] Multi-level reset strategy: For the above-mentioned anomaly types and system architectures, a multi-level reset strategy is designed, including Level I, Level II, and Level III reset strategies. The Level I reset strategy is applicable to the reset management of a single core business board. According to whether the core business is deployed on the board, a reset or non-reset strategy is selected; the Level II reset strategy is for multiple core business boards. By coordinating the reset or non-reset strategies, the stability of the core business and the system consistency are ensured; the Level III reset strategy covers the system-level reset management. According to the core business deployment situation of the system, a flexible selection of the whole device system reset or non-reset scheme is made to comprehensively improve the reliability and fault tolerance of the device.

[0083] Reset decision-making mechanism: This mechanism obtains the abnormal business boards in real time through the anomaly monitoring module in the management CPU, and dynamically analyzes them in combination with the software configuration word, so as to determine the reset strategies of each board in the system architecture. For the abnormal object, first determine the appropriate reset level according to its influence range, and then execute the corresponding reset strategy according to the board attributes to minimize the impact on the normal functions of the system.

[0084] In one embodiment, a multi-level reset method for the distributed architecture of a protection control device is provided. The method includes:

[0085] Single-board reset: For single-board anomalies, the output of the reset signal is controlled through the hardware control word, so that the core business CPU board is preferentially reset;

[0086] Multi-board collaborative reset: For other core business CPU boards that have business coupling with the abnormal core business CPU board, after the abnormal core business CPU board is reset, the management CPU board actively triggers the reset of the other core business CPU boards according to the software configuration word;

[0087] Reset of the entire device system: After the single-board reset and the multi-board collaborative reset fail to reach the upper limit number of times or when multiple non-associated core business CPU boards are reset simultaneously, the management CPU board actively triggers the reset of all other business CPU boards, and then the management CPU board self-resets to start the power-on reset process of the entire machine.

[0088] Further, in one embodiment, the single-board reset belongs to the level-I reset in the multi-level reset method and is configured on each board in the architecture. On the one hand, the power monitoring circuit inside the board detects the abnormalities of the main power supply and the local power supply, and issues a reset signal when the voltage threshold is exceeded. If the hardware control word is enabled, the CPU of this board is reset. On the other hand, with the watchdog circuit composed of a monostable flip-flop, the abnormal state of the CPU is monitored. When an abnormality is detected, an alarm or lock signal is sent to the communication bus. If this board is a core business board, a lock signal is sent to the backplane bus; if this board is a non-core business board, an alarm signal is sent to the backplane bus, as Figure 2 shown.

[0089] To assist in making decisions on the reset strategies of the core business CPU and the non-core business CPU, a hardware control word circuit is designed, as Figure 3 shown. This circuit is an important part of the single-board reset circuit. Its input is the output of the power monitoring circuit. In the non-core business CPU board, the OE pin is pulled up to shield the reset signal of this board; in the core business CPU board, the OE pin is grounded to enable the reset signal of this board.

[0090] At the same time, the hardware control word circuit can enable the core business module to obtain an independent local self-reset ability to ensure its reset priority.

[0091] Preferably, in some embodiments, the abnormal reset process includes the following steps:

[0092] Step 1: The abnormal business CPU board enters the reset startup process.

[0093] Step 2: The management CPU board detects the abnormal business CPU board, parses the software configuration word, and performs a reset according to the abnormal classification and the multi-level reset strategy table. If the business CPU board deploys core business, go to Step 3; if the business CPU board deploys non-core business, go to Step 4.

[0094] Step 3: If the business CPU board deploying core business is an independent single board, go to Step 5; if the business CPU board deploying core business has business associated boards, go to Step 6.

[0095] Step 4: The non-core business sets an alarm and waits for maintenance.

[0096] Step 5: The management CPU board triggers the abnormal business CPU board that is reset to enter the power-on reset process.

[0097] Step 6: The management CPU board triggers the associated service. The CPU board resets actively, and then the coupled multi-core service CPU boards enter the power-on reset process.

[0098] Furthermore, in one embodiment, the multi-board collaborative reset belongs to the level-II reset strategy in the multi-level reset method. It not only resets the board with anomalies but also passively resets the associated service boards that are coupled with it in terms of services, so as to ensure the integrity of the functions of the associated service boards. The specific processing flow is as Figure 4 shown and includes:

[0099] (1) Multi-board collaborative reset trigger mechanism: When a service board is abnormal and the single-board reset is completed, the system triggers the multi-board collaborative reset process.

[0100] (2) Multi-board collaborative reset logic judgment: Comprehensive logic judgment is performed based on the communication status and the locking signal. If the communication link of the abnormal service board is interrupted and a locking signal is generated in the backplane bus, then this board is determined to be the core service board and enters the multi-board collaborative reset process subsequently; if the abnormal board only generates an alarm signal and has no locked state, it is determined to be a non-core service board, the on-site state is retained, and it waits for the operation and maintenance personnel to perform further processing. Finally, the abnormal information is displayed on the human-machine interface.

[0101] (3) Dynamic parsing of associated service boards: By analyzing the software configuration word, other core service boards that are coupled with the core service board in terms of services are identified. The specific steps include:

[0102] Step 1: Communication anomaly judgment. The anomaly monitoring module of the management CPU checks the communication link status of each board to determine which board has a communication anomaly.

[0103] Step 2: Control word parsing. Read the coupled board information in the software configuration word to further confirm the information of other boards related to the abnormal service board.

[0104] Here, the software configuration word includes the service board slot position, type, and service coupling identifier.

[0105] Service board slot position: Used to identify the position of the service CPU board on the bus board.

[0106] Service board type: Used to distinguish between core service boards and non-core service boards.

[0107] Service coupling identifier: Describes the coupling relationship between service boards.

[0108] Step 3: Reset strategy selection. According to the parsing result of the control word, the management module of the management CPU board identifies the priority and service coupling status of the abnormal board. Based on the abnormal classification and multi-level reset strategy table (such as Table 1 above, or it can be customized, all falling within the protection scope of the present invention), the corresponding reset strategy is selected, and at the same time, the information of the reset strategy in the software control word is updated.

[0109] (4) Inter-board data consistency maintenance mechanism: For service-related boards, execute the data synchronization and cleaning strategy to ensure that the interaction data between the abnormal core service board and its service-related boards remains consistent after recovery. Data cleaning objective: Ensure the integrity and consistency of the DDR external memory and FPGA internal RAM data of service-related boards, meeting the core service requirements of protection and measurement and control devices. Protection service data includes original sampling data such as voltage and current instantaneous values collected in real time, fault-triggered waveform data, protection action setting parameters (such as tripping time, action threshold, current transformer and voltage transformer ratio), and intermediate phasor calculation data based on the Fast Fourier Transform (FFT), providing accurate input for protection logic. Measurement and control service data involves the sequence of events record (SOE) for recording the time series of device state changes, measured values (such as bus voltage, line current, active power, reactive power, and frequency), circuit breaker and disconnector status signals, and event control data, used for the judgment and execution of measurement and control logic. Management service data covers initialization dynamic configuration (such as module parameter loading and communication topology allocation), event-triggered waveform data, communication link status information, and clock synchronization and calibration parameters within the distributed device, ensuring the global consistency and dynamic cooperation ability of the system. By cleaning and reconstructing the above key data, ensure the linkage and reliability of each service module after reset, fully meeting the operation requirements in complex service scenarios. Scope of influence: The affected service-related boards usually include the management board and other core service boards.

[0110] Here, the data cleaning objective: Ensure the integrity and consistency of the DDR external memory and FPGA internal RAM data of service-related boards.

[0111] Protection service data: Include instantaneous value sampling data of voltage and current, fault waveform data, dynamic adjustment parameters of the protection setting area, and intermediate phase angle calculation data for protection logic, etc.;

[0112] Measurement and control service data: Cover the sequence of events record (SOE), measured values (such as real-time data like bus voltage and line current), switch device status signals (such as circuit breaker position signal), and event record and control command feedback of the measurement and control device;

[0113] Managing business data: It involves dynamic configuration information for system initialization, long-time waveform recording data triggered by events, communication link status monitoring information, as well as configuration information for the synchronous clock between boards and deviation correction parameters, etc.

[0114] (5) Communication link reconstruction and task recovery: The management board issues instructions to coordinate the business-related boards and the abnormal core business board to re-establish the communication link.

[0115] Here, the communication link reconstruction includes: 1) Data handshake. Confirm the restoration of the basic communication status between the abnormal board and the associated board; 2) Task synchronization. Update and synchronize the task status of the associated board to make it re-adapt to the business requirements of the abnormal core business board. Through the reconstruction of the communication link, ensure that the data interaction between the abnormal core business board and the business-related board returns to normal, and complete the collaborative reset between boards.

[0116] In summary, the multi-board collaborative reset aims to effectively isolate the abnormal core business board through the board collaboration mechanism to prevent it from affecting the normal operation of other business boards; ensure business collaboration and data consistency between the associated board and the abnormal board through precise data synchronization and communication link reconstruction; at the same time, shorten the reset recovery time between the business-related board and the abnormal core board to improve the user experience.

[0117] Further, in one embodiment, the entire device system reset belongs to the level III reset strategy in the multi-level reset method, which is specifically used to solve complex abnormal scenarios that cannot be handled by the single-board reset (level I reset) and the multi-board collaborative reset (level II reset), aiming to ensure the functional recovery of the entire distributed system. The following is a detailed process description of the entire device system reset, see Figure 5 as follows:

[0118] (1) Entire device system reset trigger mechanism: The management CPU triggers the function complete reset module to start the entire system reset process by real-time monitoring the system status and capturing any of the following trigger conditions: 1. Multi-board collaborative reset exception: When data synchronization fails or communication link restoration is ineffective between the core business boards, trigger the entire system reset; 2. Single-board reset fails multiple times: The core business board still cannot recover its normal function after multiple consecutive single-board resets due to abnormalities; 3. Multiple business boards are abnormal simultaneously: Multiple business boards in the system simultaneously show lock signals or communication interruptions, seriously affecting the stability of the system operation.

[0119] (2) Entire device system reset instruction sending: When the entire system reset condition is met, the management CPU independently sends the entire system reset instruction to all business boards to enter the reset process. All boards execute the reset operation in parallel, covering stages such as hardware resource reset, cache data cleaning, and task scheduling module loading.

[0120] (3)Reset and initialization of the entire device system: After the trigger of the entire system reset, the management CPU boots the system into the reset initialization phase, which specifically includes the following steps: Initialization of function interfaces: Initialize the basic hardware resources of all boards, including communication interfaces, storage modules, external device interfaces, etc. Parameter loading initialization: After the initialization of function interfaces is completed, the management CPU sends parameter initialization commands to each board to ensure that the configuration parameters required for system operation are loaded into each module.

[0121] (4)Task reconstruction and system recovery: After the reset initialization of the entire system is completed, all boards start the task scheduling mechanism to restore the running state of tasks. At this time, the system enters the normal operation phase, and at the same time, the anomaly monitoring module starts a new round of status anomaly monitoring to continuously ensure the stability and reliability of the system.

[0122] Preferably, in some embodiments, the power-on reset process includes the following steps:

[0123] Step 1: The management CPU board sends an initialization command to the service CPU board. After receiving this command, the service CPU board starts the initialization work, calls the initialization function interfaces of all functional modules, and returns an initialization success message after completion.

[0124] Step 2: The management CPU board sends a parameter initialization command to the service CPU board. After receiving this command, the service CPU board calls the parameter initialization interfaces of each functional module.

[0125] Step 3: The management CPU board sends an initialization end command to the service CPU board. After receiving this command, the service CPU board completes all initialization processes and starts the task scheduling mechanism to start task operation. If the startup fails, the management CPU board triggers the reset of the entire device system and enters Step 1. There is an upper limit to the number of repetitions, which can be configured as needed.

[0126] As a specific example, in one of the embodiments, the present invention is further verified and described in detail.

[0127] Taking the specific distributed architecture of the protection and measurement integrated device as an example below, the specific implementation manner of the present invention is further elaborated in connection with the accompanying drawings. However, this example is only used to explain the present invention in detail and does not limit the scope of application of the present invention.

[0128] 1. Specific configuration of the protection and measurement integrated device:

[0129] The device is configured on the backplane slots from left to right as the management board (slot 1), protection board 1 (slot 2), protection board 2 (slot 3), measurement and control board (slot 4), and protection board 3 (slot 5) in sequence. The configuration of the management board is as Figure 1As shown in the figure, it is a non-core business board card, and the hardware control word circuit is configured not to reset; Protection Board 1 and Protection Board 2 are core business boards with common core businesses, and the hardware control word circuits of both boards are configured to reset the respective boards; The measurement and control board is a non-core business board card, and the hardware control word circuit is configured not to reset; Protection Board 3 is a core business board and has no business coupling with Protection Board 1 and Protection Board 2, and the hardware control word circuit is configured to reset the respective board. Assume that the main power supply in each board card is +5V, the voltage offset threshold is +4.7V, the local power supply is 3.3V, and the voltage offset threshold is +3.0V. The initial configuration of the software configuration word of each board is as follows:

[0130] Management board (corresponding to the management CPU board in Figure 1 ): Business board card slot (010: Slot 1), business board card type (0: Non-core business), board card coupling identifier (00000: No business coupling slot), reset strategy (00: Do not reset).

[0131] Protection Board 1 (corresponding to Business CPU Board 1 in Figure 1 ): Business board card slot (010: Slot 2), business board card type (1: Core business), board card coupling identifier (00100: Business coupling with Slot 3), reset strategy (10: Single-board reset + Multi-board collaborative reset).

[0132] Protection Board 2 (based on Business CPU Board 1, and by analogy as Business CPU Board 2): Business board card slot (011: Slot 3), business board card type (1: Core business), board card coupling identifier (01000: Business coupling with Slot 2), reset strategy (10: Single-board reset + Multi-board collaborative reset).

[0133] Measurement and control board (based on Business CPU Board 2, and by analogy as Business CPU Board 3): Business board card slot (100: Slot 4), business board card type (0: Non-core business), board card coupling identifier (00000: No business coupling slot), reset strategy (00: Do not reset).

[0134] Protection Board 3 (based on Business CPU Board 3, and by analogy as Business CPU Board 4): Business board card slot (101: Slot 5), business board card type (1: Core business), board card coupling identifier (00000: No business coupling slot), reset strategy (10: Single-board reset).

[0135] 2. Single-board reset under the I-level strategy:

[0136] If the main power supply of the protection board 3 drops below 4.7V, or the local power supply drops below 3.0V. At this time, the power supply monitoring module in the protection board 3 detects the abnormality of the main power supply or the local power supply, generates a reset signal, and the reset signal is judged as the core service board reset signal by the hardware control word circuit, triggering the single-board reset and resetting the protection board 3. At the same time, the management board judges the outlet lock through the backplane bus, combines the configuration word information of the abnormal board card. At this time, the configuration word information of the protection board 3 is: service board card type (1: core service), reset strategy (10: single-board reset), board card coupling identifier (00000: no coupled service slot), and it can be determined that the protection board 3 is a protection board with no core service coupling, and only the single-board reset is performed.

[0137] Then the management CPU boots the protection board 3 to restart: 1. The management board sends an initialization command to the protection board 3. After receiving the command, the protection board 3 calls the initialization function interfaces of all modules, and returns an initialization success message after completion. 2. The management board sends a parameter initialization command to the protection board 3. After receiving the command, the protection board 3 calls the parameter initialization interfaces of each module. 3. The management board sends an end of initialization and a formal operation command to the protection board 3. After receiving the command, the protection board 3 completes all initialization processes, starts the task scheduling mechanism, starts task operation, and completes the single-board reset.

[0138] 3. Do not reset under the I-level strategy:

[0139] If the main power supply of the measurement and control board drops below 4.7V, or the local power supply drops below 3.0V. At this time, the power supply monitoring module in the measurement and control board detects the abnormality of the main power supply or the local power supply, generates a reset signal, and the reset signal is judged as a non-core service board reset signal by the hardware control word circuit, and no longer triggers the single-board reset, only an alarm signal is issued. At this time, the configuration word information of the abnormal measurement and control board is: service board card type (0: non-core service), reset strategy (00: do not reset), board card coupling identifier (00000: no coupled service slot). It can be determined that the abnormal board card is a measurement and control board with non-core services, and no reset is performed under the global strategy. After that, the non-faulty board cards continue to run normally, and the device displays an alarm for the measurement and control board abnormality on the human-machine interface, waiting for on-site maintenance personnel to handle.

[0140] 4. Multi-board collaborative reset under the II-level strategy

[0141] If the main power supply of the protection board 1 drops below 4.7V, or the local power supply drops below 3.0V. At this time, the power supply monitoring module in the protection board 1 detects an abnormal main power supply or an abnormal local power supply, generates a reset signal, and the reset signal is judged as the core service board reset signal by the hardware control word circuit, triggering the single-board reset and resetting the abnormal protection board 1. At the same time, the management CPU board judges the outlet lock through the backplane bus, and combines the configuration word information of the abnormal board. The configuration word information of the abnormal protection board 1 at this time is: service board type (1: core service), reset strategy (10: single-board reset + multi-board collaborative reset), board coupling identifier (00100: coupled with the service in slot 3). It can be determined that the abnormal board is the protection CPU board 1 with core services, and there is a core service coupling with the protection board 2. Therefore, while the protection board 1 starts a single-board reset, the protection board 2 performs a multi-board collaborative reset.

[0142] The management CPU guides the protection board 1 to restart: 1. The management board sends an initialization command to the protection board 1. After receiving the command, the protection board 1 calls the initialization function interfaces of all modules, and returns an initialization success message after completion. 2. The management CPU sends a parameter initialization command to the protection board 1. After receiving the command, the protection board 1 calls the parameter initialization interfaces of each module. 3. The management board sends an initialization end and official operation command to the protection board 1. After receiving the command, the protection board 1 completes all initialization processes, starts the task scheduling mechanism, starts task operation, and completes the single-board reset of the protection board 1.

[0143] The management CPU guides the protection board 2 to enter the multi-board collaborative reset judgment. If it is judged that the communication link between the protection board 1 and the management CPU board is interrupted, and the protection action outlet is locked on the backplane. At this time, it is judged that the protection board 1 is abnormal, and the board coupling identifier (00100: coupled with the service in slot 3) is obtained through the software configuration word, indicating that the protection board 2 is the one with service coupling. Then, the communication data caches in the DDR and FPGA RAM of the protection board 2 are cleared to always ensure data synchronization with the reset protection board 1. Finally, the link communication between the protection board 2 after multi-board collaborative reset and the protection board 1 after single-board reset is restored.

[0144] 5. System reset of the entire device under the III-level strategy

[0145] If there are multiple abnormal single-board resets in the I-level strategy and the II-level strategy (such as the protection board 3 restarts and resets multiple times in a short time, or the reset fails, etc.), multi-board collaborative reset abnormalities (such as the data of the protection board 1 is always inconsistent within a certain time, or the data interaction with the protection board 2 is abnormal, etc.), and multiple service board cards are abnormal at the same time (such as the protection board 1, protection board 2, protection board 3, and measurement and control board are abnormal at the same time, etc.).

[0146] The functional integrity reset module on the management board sends a system-wide reset instruction. At this time, the reset policy is upgraded to the level III reset policy (system-wide reset), and the system is guided to restart: 1. The management CPU sends an initialization command to all slave boards. After receiving this command, the slave boards start the initialization work, call the initialization function interfaces of all modules, and after completion, return an initialization success message. 2. The management CPU sends a parameter initialization command to all slave boards. After receiving this command, the slave boards call the parameter initialization interfaces of each module. 3. The management CPU sends an end-of-initialization and formal operation command to all slave boards. After receiving this command, the slave boards complete all initialization processes, start the task scheduling mechanism, start task operation, and the system restart is completed.

[0147] In summary, the present invention uses a combination of a hardware control word and a software configuration word to determine the exception handling methods of different functional service modules of the device. The reset level is selected according to the scope of the exception impact, combined with the hierarchical identification and dynamic reset policy of the exception board type, and local reset of a single board or multiple boards is preferred. Through communication link reconstruction and data consistency maintenance, the service collaboration and status synchronization between modules after reset are ensured. When the local reset is unsuccessful multiple times or multiple boards are abnormal at the same time, the system-wide reset of the entire device is started. Through multi-level reset control of single-board reset, multi-board collaborative reset, and system-wide reset of the entire device, the risk of the entire machine downtime is effectively reduced, and the system availability is improved.

[0148] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-level reset system for a distributed architecture of a protection control device, characterized in that The system is designed based on a distributed architecture, including a management CPU board, a service CPU board, an IO board, a power supply board, and a bus board; The management CPU board deploys a management module, which is used to issue reset instructions and guide the process, and receives status information from service boards through a communication link, and analyzes and makes decisions on abnormal status; The service CPU board deploys a service module, which is used to implement logical calculations and processing of various real-time tasks; The IO board is used to collect and input digital quantity signals and control outputs, as well as collect and measure small analog quantity signals; The power supply board is used to supply power to the management CPU board and the service CPU board; The bus board is used to realize signal interconnection between the management CPU board and the service CPU board; One or several functional modules among an on-board power supply monitoring module, a communication module, a data integrity maintenance module, a function complete reset module, and an abnormal monitoring module are deployed on the management CPU board and the service CPU board; The power supply monitoring module is used to monitor the abnormal status of the on-board main power supply and local power supply, output a reset signal, and determine whether to reset this board according to the hardware control word; The communication module is deployed on the management CPU board and the service CPU board, and is used for data interaction between the management CPU board and the service CPU board; The data integrity maintenance module is deployed on the management CPU board, and is used to maintain the consistency of the interaction data between the management CPU board and the reset service CPU board; The function complete reset module is deployed on the management CPU board. When the local or the entire device system is reset, it initializes through a function interface, initializes parameter loading, and schedules tasks to restore the operation of the entire system; The abnormal monitoring module is deployed on the management CPU board, monitors the system status in real time, captures single-board abnormalities, multi-board abnormalities, and system abnormalities, and triggers corresponding reset strategies; Based on the protection control device distributed architecture multi-level reset method of the reset system, the method includes: Single-board reset: For single-board abnormalities, control the output of the reset signal through the hardware control word to give priority to resetting the core service CPU board; Multi-board collaborative reset: For other core service CPU boards that have service coupling with the abnormal core service CPU board, after the abnormal core service CPU board is reset, the management CPU board actively triggers the reset of the other core service CPU boards according to the software configuration word; Entire device system reset: When the single-board reset and the multi-board collaborative reset fail to reach the upper limit number of times or multiple non-associated core service CPU boards are reset simultaneously, the management CPU board actively triggers the reset of all the remaining service CPU boards, and then the management CPU board self-resets to start the power-on reset process of the whole machine.

2. The multi-level reset system for a protection control device distributed architecture according to claim 1, wherein, The single-board reset specifically includes: Detect the abnormalities of the main power supply and the local power supply through the power supply monitoring circuit inside the board card. If the voltage threshold is exceeded, send a reset signal to the CPU of this board to reset the CPU of this board; Monitor the abnormal status of the local board CPU. When an abnormality is detected, send an alarm or blocking signal to the communication bus. Specifically: if the local board is the core service CPU board, send a blocking signal to the backplane bus; if the local board is a non-core service CPU board, send an alarm signal to the backplane bus.

3. The multi-level reset system for a protection control device distributed architecture according to claim 1, characterized in that The reset signal is controlled and output through a hardware control word. The hardware control word is to place a pull-up or pull-down resistor at the corresponding bus board position of the service CPU board. One end of the pull-up or pull-down resistor is connected to the positive power supply or the power ground, and the other end is logically ANDed with the reset output of the power supply monitoring module of this service CPU board and then connected to the reset pin of the CPU chip to determine whether the power abnormality resets the local board.

4. The multi-level reset system for a protection control device distributed architecture according to claim 1, wherein The trigger mechanism for multi-board collaborative reset is: when the core service CPU board is abnormal and the single-board reset is completed, trigger and start the multi-board collaborative reset process.

5. The multi-level reset system for a protection control device distributed architecture according to claim 1, characterized in that, The multi-board collaborative reset specifically includes: (1) Multi-board collaborative reset logic judgment: Perform comprehensive logic judgment based on the communication status and blocking signal: if the communication link of the abnormal service CPU board is interrupted and a blocking signal is generated in the backplane bus, then this service CPU board is determined to be the core service CPU board, and the subsequent multi-board collaborative reset process is entered; if the abnormal service CPU board only generates an alarm signal and there is no blocking state, it is determined to be a non-core service CPU board, and the on-site state is retained. (2) Dynamic parsing of associated service boards: Identify other core service CPU boards that have business coupling with the core service CPU board by analyzing the software configuration word; (3) Inter-board data consistency maintenance mechanism: For business-associated board cards, execute a data synchronization and cleaning strategy to ensure that the interaction data between the abnormal core service CPU board and the business-associated board cards remains consistent after the abnormal core service CPU board is restored; among them, the data cleaning target is: to ensure the integrity and consistency of the DDR external memory and the FPGA internal RAM data of the business-associated board cards, and to meet the core business requirements of the protection and measurement devices. (4) Communication link reconstruction and task recovery: Coordinate the business-associated board cards to re-establish a communication link with the abnormal core service CPU board.

6. The multi-level reset system for a protection control device distributed architecture according to claim 5, characterized in that, The software configuration word is, after the board card configuration of the protection control device is determined, to inform the management CPU board in the form of a configuration file about the distribution, quantity, and deployment of core services and non-core services of the service CPU boards in this device. The initial configuration information includes the service board slot positions, types, and business coupling identifiers.

7. The protection control device distributed architecture multi-level reset system according to claim 1, characterized in that The trigger mechanism for the whole device system reset is: By real-time monitoring the status of the whole device system, trigger and start the whole device system reset process after capturing any of the following trigger conditions: (1) Multi-board collaborative reset exception: data synchronization failure or invalid communication link recovery between core service CPU boards; (2) Single-board reset fails multiple times: the core service CPU board still cannot recover normal function after being continuously single-board reset due to an abnormality for multiple times; (3) Multiple service CPU boards are abnormal simultaneously: multiple service CPU boards in the system simultaneously have blocking signals or communication interruptions.

8. The protection control device distributed architecture multi-level reset system according to claim 7, characterized in that, The whole device system reset specifically includes: (1) Sending the whole device system reset instruction: After the overall device system reset process is triggered and started, the management CPU board synchronously sends the overall system reset instruction to all service CPU boards independently; (2) Overall device system reset initialization: The management CPU board boots the system into the reset initialization phase; (3) Task reconstruction and system recovery: After the overall device system reset initialization is completed, all boards start the task scheduling mechanism, restore the running state of the tasks, and at the same time, the anomaly monitoring module starts a new round of status anomaly monitoring.

Citation Information

Patent Citations

  • Multi-core system resetting method, device and equipment and readable storage medium

    CN114527857A

  • Method for monitoring SOC in vehicle and SOC monitoring system

    CN119201599A

  • Monitoring methd, monitoring equipment in system with multiple cores, and multiple cores system

    CN1916858A