A master authority takeover method and device and an industrial system
Patent Information
- Application Number
- CN202610846713.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-12
AI Technical Summary
[0005]本发明提供了一种主控权限的接管方法、装置及工业系统,解决了现有技术中依赖云端仲裁、固定优先级缺乏动态适应性、以及竞争冲突与恢复效率难以平衡的技术问题
[0021]本发明实施例的技术方案,通过引入基于本地健康度评分的单调非线性映射机制,通过随机扰动分量消除同分节点的碰撞风险;通过倒计时折叠机制加速恶劣工况下的接管进程;通过主控节点变更报文的强制广播立即终止全网竞争,消除脑裂风险。整个过程由各从属节点本地独立执行,无需云端仲裁,显著提升了工业控制系统在断网或通信受限环境下的自主容灾能力与恢复效率。
Smart Images

Figure CN122395211B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial automation control technology, and in particular to a method, device and industrial system for taking over master control authority. Background Technology
[0002] Industrial automation systems often employ a "master-slave" architecture. The master node undertakes core functions such as data routing and task scheduling, while slave nodes are coordinated by it. This architecture has a significant risk of single point of failure; if the master node fails, the entire control chain will be paralyzed.
[0003] There are two main existing takeover schemes: First, a cloud-based arbitration mechanism, where a new master controller is appointed in the cloud. However, this relies on external communication links and becomes completely ineffective when the network is isolated or the cloud is unreachable. Second, a fixed-priority takeover, which does not consider the real-time operating status of nodes. If high-priority nodes are resource-constrained or experience significant communication latency, forced takeover can easily trigger secondary failures; simultaneous declarations by nodes of the same priority can also lead to split-brain conflicts. Furthermore, the existing fixed-waiting mechanism does not adequately consider the health status of nodes for differentiated scheduling. A waiting window that is too long results in significant recovery delays, while a window that is too short leads to a high probability of conflict and cannot guarantee the deterministic victory of the node in the optimal state.
[0004] Therefore, under conditions of no central arbitration or limited communication, how to achieve differentiated competitive takeover based on the real-time operating status of nodes, and reduce conflict risks and improve recovery efficiency, is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This invention provides a method, device, and industrial system for taking over master control authority, which solves the technical problems of existing technologies such as reliance on cloud arbitration, lack of dynamic adaptability due to fixed priorities, and difficulty in balancing competition conflicts and recovery efficiency.
[0006] According to one aspect of the present invention, a method for taking over master control authority is provided, applied to a slave node in an industrial system, the method being executed independently and locally by the slave node, comprising:
[0007] When a failure is detected in the master control node, the evaluation indicators of the local real-time operating status are obtained, and the local health score is determined based on the evaluation indicators.
[0008] Based on the local health score, an adaptive waiting time is independently generated and a countdown is started;
[0009] Within the countdown of the adaptive waiting period, monitor whether the industrial system network receives a takeover announcement message broadcast by other slave nodes;
[0010] If so, immediately terminate the local takeover process and confirm that the node that sent the takeover announcement message is the new master node;
[0011] Otherwise, broadcast its own takeover announcement message and take over the master control.
[0012] Optionally, the adaptive waiting time is determined by both the nonlinear time component and the random perturbation component.
[0013] Optionally, the nonlinear time component is generated through a monotonically nonlinear mapping, specifically: the nonlinear time component monotonically increases as the local health score decreases; and the increasing slope of the nonlinear mapping gradually increases as the local health score decreases.
[0014] Optionally, the random disturbance component is a random time value generated within a preset time window.
[0015] Optionally, during the countdown of the adaptive waiting time, the method further includes: continuously counting the duration of the industrial system network maintaining a silent and unannounced state; when the duration reaches a preset acceleration threshold, triggering a countdown folding mechanism to automatically reduce the remaining time in the adaptive waiting time according to a preset compression ratio.
[0016] Optionally, broadcasting its own takeover announcement message includes: generating its own takeover announcement message and sending the takeover announcement message to other slave nodes in the industrial system network, wherein the takeover announcement message is used to identify the current slave node as the new master node, so that other slave nodes that receive the takeover announcement message will terminate their local takeover competition.
[0017] Optionally, the evaluation indicators of the local real-time operating status are obtained, and the local health score is determined based on the evaluation indicators, including: obtaining evaluation indicators that include at least communication latency indicators, memory resource indicators, and processing load indicators; normalizing each evaluation indicator to obtain a normalized score for each evaluation indicator; obtaining preset weight coefficients, and performing a comprehensive calculation on each normalized score and the corresponding weight coefficient to obtain the local health score.
[0018] Optionally, a fault is detected in the master control node, specifically including: not receiving a heartbeat response from the master control node within a preset heartbeat period, and failing to restore communication after reconnection attempts.
[0019] According to another aspect of the present invention, an industrial system is provided, comprising: a master node, at least one slave node, and an industrial system network, wherein the master node and the slave nodes are connected through the industrial system network; the master node is used for access control and scheduling control of the industrial system, and periodically broadcasts heartbeat responses to each slave node; the slave nodes are used to perform local industrial control tasks and monitor the heartbeat responses of the master node in real time; the industrial system network is used for data communication and takeover announcement message broadcasting between the master node and the slave nodes, and among the slave nodes.
[0020] According to another aspect of the present invention, a master control authority takeover device is provided, comprising: a health score determination module, configured to acquire local real-time operating status evaluation indicators and determine local health score based on the evaluation indicators when a failure of the master control node is detected; an adaptive duration generation module, configured to independently generate an adaptive waiting duration based on the local health score and start a countdown; and a declaration message monitoring module, configured to monitor whether the industrial system network receives a takeover declaration message broadcast by other subordinate nodes within the countdown of the adaptive waiting duration; if so, immediately terminate the local takeover process and confirm that the node that sent the takeover declaration message is the new master control node; otherwise, broadcast its own takeover declaration message and take over the master control authority.
[0021] The technical solution of this invention introduces a monotonic nonlinear mapping mechanism based on local health scores to eliminate the collision risk of nodes with the same score through random perturbation components; accelerates the takeover process under harsh operating conditions through a countdown folding mechanism; and immediately terminates network-wide competition by forcibly broadcasting change messages from the master node, eliminating the risk of split-brain. The entire process is executed independently locally by each slave node, without the need for cloud arbitration, significantly improving the autonomous disaster recovery capability and efficiency of the industrial control system in environments with network outages or limited communication.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a method for taking over master control authority according to Embodiment 1 of the present invention;
[0025] Figure 2 This is a schematic diagram of the structure of an industrial system provided according to Embodiment 2 of the present invention;
[0026] Figure 3 This is a timing diagram of master control takeover contention provided in Embodiment 2 of the present invention;
[0027] Figure 4 This is a schematic diagram of a master control authority takeover device provided in Embodiment 3 of the present invention;
[0028] Figure 5This is a schematic diagram of the structure of an electronic device that implements a master control authority takeover method according to an embodiment of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 This document provides a flowchart of a method for taking over master control authority according to Embodiment 1 of the present invention. This method is applied to slave nodes in an industrial system and is executed independently by the slave node locally. This embodiment is applicable to autonomous recovery scenarios after a master controller failure in industrial fieldbus networks (such as Profinet, EtherCAT, Modbus TCP, etc.). The method specifically includes the following steps:
[0033] S110. When a failure is detected in the master control node, obtain the evaluation indicators of the local real-time operating status and determine the local health score based on the evaluation indicators.
[0034] It should be noted that the master control takeover method in this embodiment is applied to the slave nodes in the industrial system. The entire method is executed independently by the slave node locally, without relying on the cloud or external arbitration equipment. It can achieve distributed autonomous master election when the master node fails, avoid the brain split problem caused by multi-node competition, and improve the disaster recovery efficiency and operational reliability of the industrial system.
[0035] In an industrial system, a slave node is a node that normally performs local control tasks and is coordinated by the master node. It possesses the ability to monitor the master node's status and autonomously compete for master control authority, executing takeover logic independently without relying on the cloud. The master node is the core control node of the industrial system, undertaking core functions such as data routing, task scheduling, and access control. It periodically broadcasts heartbeat messages to slave nodes and acts as the control center of the industrial system. Real-time operational status evaluation metrics are used to measure the operational status of slave nodes, including at least communication latency, memory resources, and processing load, serving as the foundational data for calculating health. The local health score is a comprehensive score calculated by normalizing communication latency, memory, and load metrics for the slave node and combining them with preset weights. A higher score indicates that the node is more suitable to take over master control authority.
[0036] Specifically, during normal operation, the slave node continuously monitors the status of the master node. When a failure is detected in the master node, the local takeover preparation process is initiated.
[0037] For example, specific implementation methods for fault monitoring may include: Method 1: If no service scheduling instructions are received from the master control node for N consecutive control cycles, and there is no response to the active request for retransmission, the master control node is determined to be faulty. Method 2: Using distributed joint determination, when more than half of the slave nodes in the network report a loss of communication with the master control node, and the local detection of the node times out, the master control node is confirmed to be faulty. Method 3: If no running status frame is received from the master control node within a specified period, and there is no response to the active query for status, the master control node is determined to be faulty. Method 4: Combining link quality and communication timeout joint determination, when the communication link quality is lower than a preset threshold and heartbeat messages are lost, the master control node is determined to be faulty.
[0038] Optionally, the evaluation indicators of the local real-time operating status are obtained, and the local health score is determined based on the evaluation indicators, including: obtaining evaluation indicators that include at least communication latency indicators, memory resource indicators, and processing load indicators; normalizing each evaluation indicator to obtain a normalized score for each evaluation indicator; obtaining preset weight coefficients, and performing a comprehensive calculation on each normalized score and the corresponding weight coefficient to obtain the local health score.
[0039] The communication latency metric is extracted based on the round-trip time (RTT) of a subordinate node sending a probe message to a standard node in the industrial system network and receiving a response. A shorter RTT indicates better network link quality for that node. The memory resource metric is extracted based on the proportion of currently available memory space to total memory capacity of the subordinate node. A higher proportion of available memory indicates more abundant resource reserves for that node. The processing load metric is extracted based on the current CPU utilization rate of the subordinate node. A lower CPU utilization rate indicates a lighter computational load and greater processing headroom after takeover. The above metric data can be obtained locally in real time through system call interfaces or hardware monitoring commands.
[0040] Specifically, each subordinate node performs normalization on each evaluation indicator to map indicators with different dimensions and numerical ranges to the interval [0,1], eliminating differences in magnitude between indicators and ensuring that all indicators have an equal basis for comparison in the scoring calculation. During normalization, the communication delay indicator is based on the maximum RTT value historically monitored by the subordinate node. The shorter the RTT of a node, the higher its normalized score for the communication delay indicator. The closer to 1, the higher the normalized score of the memory resource metric, which is based on total memory capacity. A node with a higher percentage of available memory receives a higher score. The closer to 1, the higher the normalized score for the processing load metric, with 100% full load as the baseline. Nodes with lower CPU utilization receive a higher processing load metric score. The closer to 1, the better. After completing the indicator normalization process, it is necessary to obtain the preset weighting coefficients. These weighting coefficients are pre-configured based on the actual application needs in the industrial field and correspond to the communication latency indicator, memory resource indicator, and processing load indicator, respectively, denoted as... , and Furthermore, the sum of the three weighting coefficients is 1, which can be flexibly adjusted according to the different emphases on communication real-time performance, equipment resource reserves, and operational stability on site. Finally, the subordinate node calculates each normalized score and its corresponding weighting coefficient according to the preset comprehensive calculation formula, as shown in the following formula (1):
[0041] (1)
[0042] in, This represents the normalized score corresponding to the communication delay metric. This represents the normalized score corresponding to the memory resource metric. This represents the normalized score corresponding to the processing load metric. , and These represent the weight coefficients corresponding to each scoring indicator. The scoring result calculated by this formula directly reflects the overall operational status and takeover capability of each subordinate node.
[0043] Optionally, a fault is detected in the master control node, specifically including: not receiving a heartbeat response from the master control node within a preset heartbeat period, and failing to restore communication after reconnection attempts.
[0044] Specifically, during the normal operation of the industrial system, the master node continuously broadcasts heartbeat response messages to each slave node according to a preset heartbeat cycle. This indicates to the slave nodes that it is in normal working condition. The slave nodes continuously receive and monitor these heartbeat messages to determine whether the master node is online. If a slave node does not receive a heartbeat response from the master node within the preset heartbeat cycle, it will not immediately determine that the master node is faulty. Instead, it will first determine that it is a first heartbeat timeout to avoid misjudgment due to network jitter, brief interference, or other factors. At this time, the slave node will actively initiate a reconnection attempt with the master node, attempting to restore communication by resending connection requests and querying status. If, after the reconnection attempt, no response is received from the master node and normal communication cannot be re-established, the slave node finally confirms that the master node has failed and then initiates subsequent health scoring, adaptive waiting, monitoring, and takeover processes. By combining heartbeat cycle monitoring with active reconnection verification, the impact of network fluctuations can be effectively eliminated, improving the accuracy and reliability of fault diagnosis and providing accurate triggering conditions for subsequent stable master control takeover.
[0045] S120: Based on local health scores, independently generate adaptive waiting times and start a countdown.
[0046] It is known that adaptive waiting time is a key step in achieving distributed autonomous leader election. By converting health scores into waiting time, nodes in better condition are given priority to take over. This eliminates the need for negotiation between nodes and does not rely on central node arbitration. This mechanism avoids the split-brain phenomenon caused by multiple nodes simultaneously taking over, and further improves the efficiency and stability of system takeover.
[0047] The adaptive waiting time is generated by a combination of nonlinear time components and random perturbation components. The higher the local health score, the shorter the adaptive waiting time, enabling differentiated competition by prioritizing high-scoring nodes to trigger takeover. The countdown refers to the timing process started locally by the slave node after generating the adaptive waiting time. During the countdown period, the slave node continuously listens for network announcement messages.
[0048] Specifically, the slave node independently calculates the adaptive waiting time based on its local health score. The calculation process does not interact with other nodes or rely on network data. Once the calculation is complete, a countdown begins immediately, and the node enters the listening and waiting phase.
[0049] For example, the adaptive waiting time generation methods may include: Method 1: Using a linear mapping method, the waiting time increases by a fixed time step for every fixed decrease in the health score. Method 2: Using a non-linear mapping method, the waiting time for high-scoring nodes is drastically compressed, while the waiting time for low-scoring nodes is significantly extended, thus increasing the trigger interval between nodes. Method 3: Using a segmented fixed-time method, the health score is divided into high-score, medium-score, and low-score segments, with different fixed waiting times corresponding to different segments.
[0050] Optionally, the adaptive waiting time is determined by both the nonlinear time component and the random perturbation component.
[0051] The nonlinear time component is generated based on the local health score through a monotonic nonlinear mapping. Its function is to assign shorter base waiting times to nodes with higher health scores and significantly longer base waiting times to nodes with lower health scores. This mechanism ensures that nodes in better condition trigger takeover first, while widening the waiting time difference between nodes with different scores, reducing the probability of multiple nodes triggering simultaneously. The random perturbation component is a random time value generated within a preset small time window. Its main function is to solve the split-brain problem caused by nodes with the same health score potentially ending their countdowns simultaneously. To avoid disrupting the overall priority relationship established by the nonlinear time component, the time span of the random perturbation component is strictly limited, ensuring it is less than the difference between the nonlinear time components corresponding to the minimum score step size. This ensures that the waiting times of each node maintain their original priority order after the addition of random perturbation. Adding the nonlinear time component and the random perturbation component yields the final adaptive waiting time.
[0052] Optionally, the nonlinear time component is generated through a monotonically nonlinear mapping, specifically: the nonlinear time component monotonically increases as the local health score decreases; and the increasing slope of the nonlinear mapping gradually increases as the local health score decreases.
[0053] Specifically, the nonlinear time component exhibits a strict negative correlation with the local health score. As the local health score decreases, the nonlinear time component shows a monotonically increasing trend, meaning that the worse the node's operating status and the lower its health score, the longer its base waiting time. Simultaneously, this nonlinear mapping has a gradually increasing slope. In the range of higher local health scores, the slope of the mapping curve is relatively gentle, and the difference in the nonlinear time component between adjacent scores is small, resulting in highly similar waiting times for high-scoring nodes, enabling rapid and concentrated triggering and shortening the system's masterless vacuum period. Conversely, in the range of lower local health scores, the slope of the mapping curve continuously increases as the score decreases, significantly amplifying the difference in the nonlinear time component between adjacent scores. This drastically lengthens the waiting time of low-scoring nodes, creating a clear temporal isolation from high-scoring nodes. This effectively reduces the possibility of different performance nodes simultaneously triggering their countdowns to end, avoiding the split-brain problem caused by multiple node preemption from the mapping mechanism, and improving the stability and determinism of the master control takeover process.
[0054] In one specific implementation, the nonlinear time component can be calculated using the following exponential decay function formula (2):
[0055] (2)
[0056] in, Represents the nonlinear time component. This represents a preset time stretching constant, used to limit the maximum waiting time of the worst-case node. Represents the natural constant. This represents the preset attenuation coefficient, which is a positive number. For example, a value of 0.05 can be used to control the compression steepness of the nonlinear curve. This represents the local health score, ranging from 0 to 100. This represents the preset minimum waiting time, for example, 10ms. Based on the above formula, the mapping pattern exhibits the following characteristics:
[0057] (1) For high health score nodes, achieve extreme time compression: For high score nodes with excellent status (e.g. >90), due to the score Larger, negative exponent term Approaching 0, the calculated nonlinear time component is extremely compressed, infinitely approaching the limit value. Within this high-resolution interval, the slope of the mapping curve is extremely gentle, allowing the high-resolution node group to form a highly clustered "fast channel" on the time axis, thereby minimizing the vacuum period when the system's main control is missing.
[0058] (2) Implement non-linear large-scale physical isolation for nodes with low health scores: as the health score increases... The decrease, negative exponent term It increases exponentially and rapidly in the opposite direction. At this point, the mapping curve becomes extremely steep, significantly lengthening the corresponding waiting time with a non-linearly increasing slope. This exponential stretching mechanism significantly amplifies the time difference between adjacent scores within the lower score range.
[0059] For example, setting =10ms =5000ms =0.05. When the node score is 100, ≈43.6ms; when the score is 90 points, ≈65.5ms. At this point, the time difference for a 10-point score difference is only about 21.9ms, with the high-scoring node triggering quickly and efficiently. When the node score reaches 70 points, ≈161ms; Conversely, when the node score is 60 points, ≈259ms; when the score is 50, ≈420ms. In the low-score range at this point, a 10-point difference in score results in a drastically amplified time difference to 161ms. In other words, even if two low-score nodes differ by only 1 point, the corresponding time difference will be far greater than the difference corresponding to 1 point in the high-score range. This effectively ensures that low-performance nodes are strictly delayed, greatly reducing the probability of a "split-brain" collision where multiple nodes simultaneously reset to zero in an unhealthy state within the industrial system.
[0060] Optionally, the random disturbance component is a random time value generated within a preset time window.
[0061] Specifically, the random perturbation component is a random time value generated within a preset time window. Its purpose is to address the split-brain conflict caused by multiple nodes with the same health score potentially having identical non-linear time components, leading to simultaneous countdown termination and broadcast of takeover announcements. This further enhances the uniqueness and stability of takeover competition. The preset time window is a small, fixed time range. Its maximum span is strictly limited, requiring it to be less than the difference in the non-linear time component corresponding to the minimum score step size in the local health score. This ensures that adding random perturbation does not alter the overall priority order determined by the non-linear time component, prioritizing higher scores and delaying lower scores, and does not interfere with the competitive ranking between nodes. When generating the random perturbation component, each subordinate node independently generates a random time value within the preset time window using a local random algorithm. This random time value is then superimposed on the non-linear time component to form the final adaptive waiting time. Because the random disturbance component has a small value and is randomly generated within a limited window, it can cause nodes with the same score and consistent nonlinear time components to have a small difference in waiting time, thus avoiding multiple nodes triggering announcement broadcasts at the same time. This mechanism completely eliminates the risk of concurrent competition among nodes with the same score and ensures that the distributed autonomous leader election process is stable and reliable.
[0062] S130. During the countdown of the adaptive waiting time, listen to whether the industrial system network receives a takeover announcement message broadcast by other slave nodes. If yes, execute S140; otherwise, execute S150.
[0063] The industrial system network refers to a dedicated network in an industrial setting used for data communication, heartbeat transmission, and takeover announcement message broadcasting between master nodes and slave nodes, and among slave nodes. A takeover announcement message is a master control change message broadcast to the entire network when a slave node's countdown ends and no other announcements are detected, used to announce its takeover of master control authority.
[0064] Specifically, during the countdown, the slave node continuously listens to the industrial system network, identifies and parses broadcast messages in the network, and determines whether there are takeover announcement messages from other nodes. The listening process is synchronized with the countdown process.
[0065] Optionally, during the countdown of the adaptive waiting time, the method further includes: continuously counting the duration of the industrial system network maintaining a silent and unannounced state; when the duration reaches a preset acceleration threshold, triggering a countdown folding mechanism to automatically reduce the remaining time in the adaptive waiting time according to a preset compression ratio.
[0066] Specifically, during the countdown of the adaptive waiting time, this implementation also sets up network silence monitoring and countdown folding acceleration logic to cope with the severe operating conditions that may occur in industrial sites, where the health of all nodes is generally low and the overall waiting time is too long, resulting in the system being in a masterless vacuum state for a long time. Throughout the countdown process, the slave nodes will continuously count the duration for which the industrial system network remains in a silent, unannounced state. This duration refers to the cumulative time from when the slave node starts the countdown to the current moment, without listening to any other slave node issuing a takeover announcement message. When the counted network silence duration reaches a preset acceleration threshold, it indicates that there is no high-health node in the current network that can quickly trigger takeover, and the system faces the risk of masterless timeout. At this time, the slave nodes automatically trigger the countdown folding mechanism. The countdown folding mechanism will proportionally reduce the remaining time in the adaptive waiting time according to a pre-set compression ratio. It is important to emphasize that this folding mechanism uses a proportional compression method. The remaining waiting time of all nodes that are still executing the countdown will be reduced by the same ratio. Therefore, the relative priority relationship between nodes, which was originally determined by the health score and non-linear mapping, remains unchanged and will not disrupt the competition rule of high score priority.
[0067] S140. Immediately terminate the local takeover process and confirm that the node that sent the takeover announcement message is the new master node.
[0068] It should be noted that this step is a competition backoff mechanism. When a takeover announcement message from another node is heard, it indicates that a better takeover node has emerged in the network. This node actively withdraws from the competition, which can quickly converge the election results, avoid multi-master conflicts, and ensure the rapid stability of the system control structure.
[0069] The takeover process refers to the complete autonomous competition steps of a subordinate node, from detecting a master node failure, calculating its health, generating a waiting time, monitoring the countdown, to finally broadcasting a takeover announcement or terminating the competition. The new master node is the subordinate node that first broadcasts a takeover announcement message within the countdown, and after being confirmed by other nodes, replaces the original failed node as the new control node of the system.
[0070] Specifically, if a slave node hears a takeover announcement message before the countdown ends, it immediately stops the local countdown, stops all takeover-related processing, stops subsequent broadcast operations, and identifies the node that sent the announcement message as the new master node.
[0071] S150, broadcasts its own takeover announcement message and takes over the master control authority.
[0072] It should be noted that this step is the master control function. When the countdown ends and no announcement message is heard, it indicates that this node is the optimal takeover node in the current network. This node will complete the role switch and assume the master control function, so that the system can quickly recover from the fault and ensure that industrial control services are not interrupted.
[0073] Specifically, if the slave node's countdown reaches zero and it does not hear any takeover announcement message, it immediately broadcasts its own takeover announcement message to the entire network, triggering other nodes to stop competing and update the master control information. At the same time, it completes the role switch and officially takes over the master control authority.
[0074] Optionally, broadcasting its own takeover announcement message includes: generating its own takeover announcement message and sending the takeover announcement message to other slave nodes in the industrial system network, wherein the takeover announcement message is used to identify the current slave node as the new master node, so that other slave nodes that receive the takeover announcement message will terminate their local takeover competition.
[0075] The takeover announcement message is a dedicated message used to identify that the current slave node has switched to the new master node. Its function is to announce the role change of this node to the entire network and also has the effect of forcibly terminating competition. When other slave nodes receive this takeover announcement message, they will immediately recognize the new master identity of the sending node, automatically stop all ongoing takeover competition processes such as health calculation, adaptive waiting, and countdown listening, and will no longer attempt to takeover or broadcast their own announcement messages. This quickly converges the network-wide competition state and avoids the split-brain problem of multiple masters coexisting. The takeover announcement message is sent using a network-wide broadcast mode, ensuring that all slave nodes in the industrial system can receive it in a timely manner, guaranteeing the consistency and real-time nature of state synchronization. Through this mechanism, once a node successfully triggers takeover and broadcasts the announcement, the master recognition of all nodes in the network can be quickly unified, allowing the entire system to rapidly recover from a fault state to a stable master-slave operating structure, improving the fault recovery speed and operational reliability of the industrial system.
[0076] The technical solution of this invention introduces a monotonic nonlinear mapping mechanism based on local health scores to eliminate the collision risk of nodes with the same score through random perturbation components; accelerates the takeover process under harsh operating conditions through a countdown folding mechanism; and immediately terminates network-wide competition by forcibly broadcasting change messages from the master node, eliminating the risk of split-brain. The entire process is executed independently locally by each slave node, without the need for cloud arbitration, significantly improving the autonomous disaster recovery capability and efficiency of the industrial control system in environments with network outages or limited communication.
[0077] Example 2
[0078] Figure 2This is a schematic diagram of an industrial system according to Embodiment 2 of the present invention. This embodiment provides a detailed description of the industrial system based on Embodiment 1 described above. The system includes: a master control node, at least one slave node, and an industrial system network. The master control node and the slave node are connected through the industrial system network.
[0079] Specifically, the master node is used for access control and scheduling of the industrial system, and periodically broadcasts heartbeat responses to each slave node; slave nodes are used to execute local industrial control tasks and monitor the heartbeat responses of the master node in real time; the industrial system network is used for data communication and takeover announcement message broadcasting between the master node and slave nodes, as well as between each slave node.
[0080] The master control node is deployed within the industrial system network and serves as the core control unit of the entire system. Its primary functions include routing and coordinating industrial control data, scheduling and allocating tasks, and managing master control permissions. During normal system operation, the master control node continuously broadcasts heartbeat response messages to all slave nodes within the network according to a preset heartbeat cycle. This heartbeat signal indicates to each slave node that it is in normal working condition, providing a basis for slave nodes to monitor the master control node's operational status. Slave nodes establish communication connections with the master control node through the industrial system network. During normal system operation, slave nodes execute local industrial control tasks according to the unified coordination of the master control node, while continuously monitoring the heartbeat responses sent by the master control node to determine its online status. Slave nodes possess complete hardware and software infrastructure, enabling them to calculate a health score based on their local real-time operating status when a master control node failure is detected. They can then autonomously participate in the master control permission competition and complete the takeover process. The entire competition and takeover logic is executed independently locally, without relying on the cloud or external arbitration equipment. As the communication support for the entire system, the industrial system network is responsible for realizing data interaction, heartbeat message transmission, and takeover announcement message broadcasting between the master node and slave nodes, and among the slave nodes. It can ensure that when a fault occurs, the takeover announcement message can be delivered to all slave nodes quickly and reliably, ensuring that each node can promptly detect the emergence of a new master node and terminate the competition, avoiding the split-brain phenomenon and maintaining the stability of the system control structure.
[0081] Specific application scenarios: Figure 3 A timing diagram for master control takeover competition is provided for Embodiment 2 of the present invention. Figure 3This invention demonstrates the entire process of distributed autonomous election of a new master node and takeover of master control authority by subordinate nodes after a failure of the master control node in an industrial system. The process is divided into five parts according to operational phases: normal operation, fault triggering, health assessment and adaptive waiting, announcement broadcasting and backoff convergence, and new master node resuming operations. The participating entities include the master control node, subordinate node A (health score 95), subordinate node B (health score 92), subordinate node C (health score 60), and the industrial system network. First, in the normal operation phase, the industrial system is in normal working condition. The master control node, as the system control center, periodically broadcasts heartbeat response messages to all subordinate nodes according to a preset heartbeat cycle. After receiving the heartbeat message, subordinate nodes A, B, and C confirm that the master control node is online and synchronously execute local industrial control tasks. The entire process is handled by the industrial system network, which is responsible for message transmission and communication assurance. Next, in the fault triggering phase, the master control node fails and can no longer periodically broadcast heartbeat response messages, causing the heartbeat reception of subordinate nodes A, B, and C to be interrupted. Slave nodes do not immediately diagnose a failure based on a single lost heartbeat. Instead, they first attempt to actively reconnect to the master node. Only after multiple failed reconnections and a continued lack of heartbeat responses is the master node's failure confirmed, thus avoiding misjudgments caused by network fluctuations or brief interference. After failure confirmation, a health assessment and adaptive waiting phase begins. All slave nodes independently calculate their health scores locally: slave node A scores 95, slave node B scores 92, and slave node C scores 60. Higher scores indicate better node performance and a stronger ability to take over master control. Subsequently, each node generates an adaptive waiting time based on its own health score using a non-linear mapping method. Higher health scores result in shorter waiting times. Therefore, slave node A has the shortest waiting time T_A, slave node B has the next shortest T_B, and slave node C has the longest T_C. Furthermore, a random perturbation component is superimposed on each node's waiting time to prevent nodes with the same score from triggering simultaneously, thus preventing split-brain issues. After generating the waiting time, each node immediately starts a countdown locally and continuously listens for takeover announcement messages broadcast by other nodes in the industrial system network. When the countdown ends, the slave node A with the highest health is the first to complete its countdown to zero. At this time, it has not heard any other node's takeover announcement message, so it immediately broadcasts a takeover announcement message to the industrial system network, which is the master node change message, announcing its takeover master authority. The industrial system network will simultaneously broadcast this announcement message to slave nodes B and C. Slave node B receives the announcement message during the countdown T_B, and slave node C receives the announcement message during the countdown T_C. Both will immediately terminate their local takeover competition process, stop the countdown and subsequent broadcast operations, and mark slave node A as the new master node, achieving rapid convergence of the network-wide competition state and avoiding multi-master conflicts.Finally, the system enters the new master node recovery phase. Slave node A completes its identity switch, transitioning from slave node mode to master node mode, officially taking over the system's master control authority. It then assumes the core functions of data routing, task scheduling, access control, and periodically broadcasting heartbeat messages to slave nodes. Slave nodes B and C then re-establish communication connections with the new master node A, resuming heartbeat subscriptions and data interaction. The industrial system as a whole recovers from a fault state to a stable master-slave structure. The entire process is conducted independently by each slave node without cloud or external arbitration equipment, ensuring that the optimal node takes over first and preventing split-brain conflicts, thus significantly improving the disaster recovery efficiency and operational reliability of the industrial system.
[0082] The technical solution of this invention provides a stable and reliable basis for judging the master control status by periodically broadcasting heartbeat responses to each slave node through the master control node, ensuring that slave nodes can perceive the online status of the master control node in a timely and accurate manner. The slave nodes execute local industrial control tasks and monitor the heartbeat responses of the master control node in real time, which not only ensures the continuous and stable operation of industrial field services, but also quickly triggers the takeover process when the master control node fails, providing basic support for distributed autonomous master election. The industrial system network realizes data communication and takeover announcement message broadcasting between the master control node and slave nodes, and between each slave node, which can ensure the efficient and reliable transmission of heartbeat messages, control commands and takeover announcement messages, and ensure that the new master control announcement can be quickly synchronized to all nodes in the network, providing communication guarantee for quickly converging the competition state and avoiding multi-master conflicts.
[0083] Example 3
[0084] Figure 4 This is a schematic diagram of a master control takeover device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: a health score determination module 310, used to obtain local real-time operating status evaluation indicators when a failure of the master node is detected, and determine the local health score based on the evaluation indicators; an adaptive duration generation module 320, used to independently generate an adaptive waiting duration based on the local health score and start a countdown; and an announcement message listening module 330, used to listen for whether the industrial system network receives a takeover announcement message broadcast by other subordinate nodes within the countdown of the adaptive waiting duration; if so, immediately terminate the local takeover process and confirm that the node that sent the takeover announcement message is the new master node; otherwise, broadcast its own takeover announcement message and take over the master control authority.
[0085] Optionally, the device also includes: a remaining time shortening module, used to: continuously count the duration of the industrial system network maintaining a silent and unannounced state; when the duration reaches a preset acceleration threshold, trigger a countdown folding mechanism to automatically reduce the remaining time in the adaptive waiting time according to a preset compression ratio.
[0086] Optionally, the announcement message listening module 330 specifically includes: a takeover announcement message broadcasting unit, used to: generate its own takeover announcement message and send the takeover announcement message to other slave nodes in the industrial system network, wherein the takeover announcement message is used to identify the current slave node as the new master node, so that other slave nodes that receive the takeover announcement message will terminate their local takeover competition.
[0087] Optionally, the health score determination module 310 specifically includes: an evaluation index acquisition unit, used to: acquire evaluation indicators including at least communication latency indicators, memory resource indicators, and processing load indicators; normalize each evaluation index to obtain a normalized score for each evaluation index; acquire preset weight coefficients, and perform comprehensive calculations on each normalized score and the corresponding weight coefficients to obtain a local health score.
[0088] Optionally, the health score determination module 310 specifically includes: a master node fault monitoring unit, used to: if no heartbeat response is received from the master node within a preset heartbeat cycle, and communication is still not restored after reconnection attempts.
[0089] The technical solution of this invention introduces a monotonic nonlinear mapping mechanism based on local health scores to eliminate the collision risk of nodes with the same score through random perturbation components; accelerates the takeover process under harsh operating conditions through a countdown folding mechanism; and immediately terminates network-wide competition by forcibly broadcasting change messages from the master node, eliminating the risk of split-brain. The entire process is executed independently locally by each slave node, without the need for cloud arbitration, significantly improving the autonomous disaster recovery capability and efficiency of the industrial control system in environments with network outages or limited communication.
[0090] The master control authority takeover device provided in the embodiments of the present invention can execute a master control authority takeover method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0091] Example 4
[0092] Figure 5A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0093] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) or random access memory (RAM), communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from the storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. Input / output (I / O) interfaces are also connected to the bus 14.
[0094] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0095] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a master control takeover method.
[0096] In some embodiments, a master control takeover method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the master control takeover method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a master control takeover method by any other suitable means (e.g., by means of firmware).
[0097] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0098] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0099] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0100] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0101] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0102] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0103] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0104] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for taking over master control authority, applied to slave nodes in an industrial system, characterized in that, The method is executed independently locally by the slave node, including: When a failure is detected in the master control node, the evaluation index of the local real-time operating status is obtained, and the local health score is determined based on the evaluation index. Based on the local health score, an adaptive waiting time is independently generated and a countdown is started. The adaptive waiting time is determined by a nonlinear time component and a random disturbance component. The random disturbance component is a random time value generated within a preset time window. The preset time window is smaller than the difference of the nonlinear time component corresponding to the minimum score step size in the local health score. During the countdown of the adaptive waiting time, monitor whether the industrial system network receives a takeover announcement message broadcast by other slave nodes; If so, immediately terminate the local takeover process and confirm that the node that sent the takeover announcement message is the new master node; Otherwise, broadcast its own takeover announcement message and take over the master control; The nonlinear time component is generated through a monotonic nonlinear mapping, specifically: The nonlinear time component increases monotonically as the local health score decreases; and the increasing slope of the nonlinear mapping gradually increases as the local health score decreases. During the countdown of the adaptive waiting time, the method further includes: continuously counting the duration of the industrial system network maintaining a silent and unannounced state; when the duration reaches a preset acceleration threshold, triggering a countdown folding mechanism to automatically reduce the remaining time in the adaptive waiting time according to a preset compression ratio; The nonlinear time component is calculated using the following exponential decay function formula: ; in, Represents the nonlinear time component. This represents a preset time stretching constant, used to limit the maximum waiting time of the worst-case node. Represents the natural constant. This represents the preset attenuation coefficient, which is a positive number and is used to control the compression steepness of the nonlinear curve. Indicates the local health score. This indicates the preset minimum waiting time.
2. The method according to claim 1, characterized in that, The broadcast's own takeover announcement message includes: The node generates its own takeover announcement message and sends the takeover announcement message to other slave nodes in the industrial system network. The takeover announcement message is used to identify the current slave node as the new master node, so that other slave nodes that receive the takeover announcement message will terminate their local takeover competition.
3. The method according to claim 1, characterized in that, The step of obtaining evaluation indicators of local real-time operating status and determining local health scores based on the evaluation indicators includes: Obtain evaluation metrics that include at least communication latency metrics, memory resource metrics, and processing load metrics; The evaluation indicators are normalized to obtain the normalized scores of each evaluation indicator. Obtain preset weight coefficients, perform comprehensive calculations on each normalized score and its corresponding weight coefficients to obtain a local health score.
4. The method according to claim 1, characterized in that, The detection of a fault in the master control node specifically includes: not receiving a heartbeat response from the master control node within a preset heartbeat period, and failing to restore communication after reconnection attempts.
5. An industrial system applied to the method of any one of claims 1-4, characterized in that, include: A master node, at least one slave node, and an industrial system network, wherein the master node and the slave node are connected through the industrial system network; The master control node is used for permission management and scheduling control of the industrial system, and periodically broadcasts heartbeat responses to each subordinate node. The slave node is used to perform local industrial control tasks and monitor the heartbeat response of the master node in real time; The industrial system network is used for data communication and takeover announcement message broadcasting between the master node and slave nodes, as well as between each slave node.
6. A master control takeover device, applied to the method according to any one of claims 1-4, characterized in that, include: The health score determination module is used to obtain evaluation indicators of the local real-time operating status when a failure is detected in the master control node, and to determine the local health score based on the evaluation indicators. An adaptive waiting time generation module is used to independently generate an adaptive waiting time based on the local health score and start a countdown. The announcement message listening module is used to listen for whether the industrial system network receives a takeover announcement message broadcast by other subordinate nodes during the countdown of the adaptive waiting time. If so, immediately terminate the local takeover process and confirm that the node that sent the takeover announcement message is the new master node; otherwise, broadcast its own takeover announcement message and take over the master control authority.
Citation Information
Patent Citations
Destructive method and device based on articulated naturality web autonomous cloud network architecture
CN113489601A
Random medium access methods with backoff adaptation to traffic
US20020154653A1
Method and system for facilitating channel measurements in a communication network
US20090257413A1