A device control method and apparatus, electronic device, storage medium, and product

By dynamically electing a new master node through the timing analysis of slave nodes and the cluster broadcast mechanism, the problem of automatic switching when the master node fails is solved, and the reliability of device management is improved.

CN120567658BActive Publication Date: 2025-10-14INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511066106.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-14
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional device management methods lack an automatic switching mechanism when the main node fails, resulting in insufficient reliability of device management.

Method used

When the master node is offline, the slave node determines the risk level and device priority through timing analysis, adopts a cluster broadcast mechanism for election, and dynamically determines the new master node based on hardware resources, performance indicators, and historical data.

Benefits of technology

This effectively reduces the number of elected nodes, ensures that the new master node can stably and reliably take over the work of the original master node, and improves the reliability of device management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567658B_ABST
    Figure CN120567658B_ABST
Patent Text Reader

Abstract

The application discloses a device control method and device, electronic equipment, storage medium and product, and relates to the technical field of device management. When the master node is offline, the slave node evaluates its own condition, and participates in the election only when its state information and node level meet the election conditions. In order to reflect the adaptability of the node to undertake the master node duty in real time, the device priority of the node is determined according to the hardware resources, performance indexes, historical data and index weight matched with the current scene of the node. The device priority, risk level, performance index and online duration contained in each election message in the region are analyzed to determine the supported node, so that the selected node can stably and reliably undertake the management work of the master node. The election mechanism is determined according to the number of the election nodes, and the most suitable new master node is determined. The new master node can smoothly replace all the work of the original master node, and the reliability of the device management is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of device management, and in particular to a device control method and device, an electronic device, a storage medium and a product. BACKGROUND

[0002] With the rapid development of information technology, the network scale is continuously expanding, the device types are increasingly complex, and the demand for device management is also growing. In data centers, cloud computing environments and Internet of Things scenarios, efficient, secure and reliable device management is the key to ensuring stable system operation.

[0003] The traditional device management method is based on a Simple Service Discovery Protocol (SSDP) to implement a device discovery and management scheme. The protocol is part of the Universal Plug and Play (UPnP) architecture and is used for device discovery and service announcement. The traditional device management method usually adopts a static configuration master-slave architecture, and the master node is responsible for managing and controlling the slave node device. Once the master node fails, manual intervention is required to restore service, and there is a lack of automatic switching mechanism.

[0004] Therefore, how to improve the reliability of device management is a problem to be solved by those skilled in the art. SUMMARY

[0005] The present application provides a device control method and device, an electronic device, a storage medium and a product to at least solve the problem of lack of automatic switching mechanism when the master node fails in the related art.

[0006] The present application provides a device control method, comprising:

[0007] In the case that the master node is offline, the performance indicators of the node are analyzed in time sequence to determine the risk level of the node;

[0008] In the case that the state information of the node and the node level meet the selection conditions, the device priority of the node is determined according to the hardware resources, performance indicators, historical data and indicators weight matched with the current scene of the node; wherein, the node level is adjusted based on the risk level;

[0009] The first election message is sent to the cluster head node in the region, so that the cluster head node broadcasts the first election message to other candidate nodes of the same node level in the region;

[0010] Upon receiving at least one second election message broadcast by the cluster head node, analyzing the device priority, risk level, performance indicator, and online duration contained in each election message according to the set election rules, and sending an election response message to the cluster head node, so that the cluster head node generates a region selection result based on all election response messages in the region, and broadcasts the region selection result to all participating nodes;

[0011] Receive the regional selection results sent by the cluster head nodes in each region, and analyze all regional selection results to determine the new master node based on the election mechanism that matches the number of all current candidate nodes.

[0012] The present application also provides a device control apparatus, comprising a timing analysis unit, a priority determination unit, a sending unit, a node analysis unit, and an election unit;

[0013] The timing analysis unit is used to perform timing analysis on the performance indicators of its own node when the master node is offline to determine the risk level of its own node;

[0014] A priority determination unit is used to determine the device priority of its own node based on its own node's hardware resources, performance indicators, historical data, and indicator weights matching the current scenario, if its own node's status information and node level meet the selection criteria; wherein the node level is adjusted based on the risk level;

[0015] a sending unit, configured to send a first election message to a cluster head node in the area, so that the cluster head node broadcasts the first election message to other candidate nodes of the same node level in the area;

[0016] a node analysis unit configured to, upon receiving at least one second election message broadcast by the cluster head node, analyze the device priority, risk level, performance indicator, and online duration contained in each election message according to a set election rule, and send an election response message to the cluster head node, so that the cluster head node generates a region selection result based on all election response messages in the region, and broadcasts the region selection result to all participating nodes;

[0017] The election unit is used to receive the regional selection results sent by the cluster head nodes in each region, and analyze all regional selection results to determine the new master node based on the election mechanism that matches the number of all current candidate nodes.

[0018] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned device control methods when executing the computer program.

[0019] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned device control methods are implemented.

[0020] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned device control methods when executed by a processor.

[0021] Through this application, when the master node is offline, each slave node first evaluates its own situation, and each slave node can perform a time series analysis on the performance indicators of its own node to determine the risk level of its own node; the node level can be adjusted based on the risk level. Only when the status information and node level of its own node meet the conditions for participation in the election can it participate in the election. In order to reflect the adaptability of the node to assume the responsibilities of the master node in real time, the device priority of its own node can be determined based on the hardware resources, performance indicators, historical data of its own node and the weight of indicators that match the current scenario. In order to reduce the amount of data processing, the election message is processed in a cluster broadcast manner. The participating node can send a first election message to the cluster head node in the area to which it belongs, so that the cluster head node can broadcast the first election message to other participating nodes of the same node level in the area. Upon receiving at least one second election message broadcast by a cluster head node, the node can analyze the device priority, risk level, performance indicator, and online time contained in each election message according to predefined election rules. The node then sends an election response message to the cluster head node, informing it of its supported nodes. The cluster head node then generates a regional selection result based on all election response messages within the region and broadcasts the regional selection result to all participating nodes. The regional selection result includes the node with the most votes in the region. Given that regional selection results may vary for each region, to ultimately determine a unique new master node, upon receiving the regional selection results from the cluster head nodes in each region, the node can analyze all regional selection results based on an election mechanism that matches the current number of participating nodes to determine the new master node. In this technical solution, when the master node is offline, the slave node evaluates its own status and only participates in the election if its status information and node level meet the requirements. This effectively reduces the number of nodes participating in the election and avoids the heavy workload associated with a large number of slave nodes participating in the election. Each candidate node determines the nodes it supports by analyzing the device priority, risk level, performance indicators, and online time contained in each election message within the region. This ensures that the selected node can stably and reliably assume the management of the master node. Furthermore, the election mechanism is determined based on the number of candidate nodes, ensuring that the election mechanism is more in line with actual needs. This allows the most suitable new master node to be determined, ensuring that the new master node can smoothly take over all tasks of the original master node, thereby improving the reliability of device management. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 A flow chart of a device control method provided in an embodiment of the present application;

[0024] Figure 2 A flowchart of a method for selecting candidate nodes provided in an embodiment of the present application;

[0025] Figure 3 A flowchart of a method for determining a candidate node provided in an embodiment of the present application;

[0026] Figure 4 A schematic diagram of the structure of a device control device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0029] Device management based on the SSDP protocol requires manual intervention to switch nodes when the master node fails. Current solutions for automated master node switching typically rely on heartbeat detection detecting a master node anomaly. A slave node becomes a candidate node and sends a vote request message to the other slave nodes in the cluster. Upon receiving the vote request message, the other slave nodes vote and return a vote response message containing the voting results and their own priority to the candidate node. Upon receiving the vote response messages from the other slave nodes and satisfying the voting criteria, the candidate node becomes the new master node. This implementation allows all slave nodes to participate in the election when the master node fails. However, for large-scale management systems with a large number of nodes, requiring all slave nodes to participate in the election results in a significant amount of data processing, hindering the efficiency of selecting a new master node. Furthermore, each slave node typically votes based on the candidate node's load, selecting the candidate node with the lowest current load as the new master node. Given that node load is a dynamically changing variable and that master node management requires high hardware resources, performance, and device stability, relying solely on node load to select a new master node may result in the new master node being unable to effectively handle master node management tasks. It can be seen that how to select a more suitable new master node to improve the reliability of device management is a problem that those skilled in the art need to solve.

[0030] Therefore, the embodiments of the present application provide a device control method, apparatus, electronic device, storage medium, and product. When the master node is offline, the slave node will evaluate its own situation and only participate in the election when its status information and node level meet the conditions for participation in the election. This effectively reduces the number of nodes participating in the election and avoids the heavy workload caused by a large number of slave nodes participating in the election. In order to reflect the adaptability of the node to assume the responsibilities of the master node in real time, the device priority of the node itself can be determined based on the hardware resources, performance indicators, historical data, and indicator weights that match the current scenario. The higher the device priority of the node, the better the adaptability of the node to the responsibilities of the master node. Considering that the device priority of each participating node is the same, when each participating node performs voting, in addition to considering the device priority, it can also consider parameters such as the risk level, performance indicators, and online time of each participating node. That is, each participating node can analyze the device priority, risk level, performance indicators, and online time contained in each election message in the area to determine the node it supports, ensuring that the selected node can stably and reliably assume the management work of the master node. The election mechanism is determined based on the number of participating nodes, ensuring that the election mechanism can better meet actual needs, thereby determining the most suitable new master node, ensuring that the new master node can smoothly take over all the work of the original master node, and improving the reliability of device management.

[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] Figure 1 A flow chart of a device control method provided in an embodiment of the present application, the method comprising:

[0033] S101: When the master node is offline, a time series analysis is performed on the performance indicators of the own node to determine the risk level of the own node.

[0034] There are several ways to identify when a master node is offline. The first method is heartbeat detection timeout detection. A slave node sends heartbeat detection messages to the master node at set intervals. If no heartbeat response message is received from the master node within a set number of consecutive time intervals, the master node is considered offline.

[0035] The time interval and the set number can be pre-set by the developer based on actual needs, or can be set by the user through the human-computer interaction interface.

[0036] For example, if the default interval is set to 30 seconds and the number of heartbeats is set to 3, the slave node can send a heartbeat detection message to the master node every 30 seconds. If the master node does not reply to the heartbeat response message within three consecutive heartbeat cycles, the slave node can determine that the master node is offline and start the election process.

[0037] The second identification method can be intra-group communication anomaly identification. In a data exchange scenario, if a slave node fails to receive a response from the master node within a set number of consecutive times within the timeout period, the master node is considered offline.

[0038] Data interaction scenarios may include reporting device status, receiving management instructions, and other scenarios.

[0039] The timeout period and the number of settings can be flexibly set based on actual needs.

[0040] Taking the timeout period as 1 minute and the number of times set as 3 as an example, when the slave node attempts to interact with the master node for data, if it does not receive a response from the master node within the set 1 minute and communication is still not restored after 3 retries, the master node is considered offline and the election process is started.

[0041] The third identification method is that the master node actively declares itself offline. When a slave node receives an offline notification message broadcast by the master node, it determines that the master node is offline.

[0042] In certain foreseeable circumstances, such as planned maintenance or upgrades of the master node, the master node will proactively broadcast an offline notification message to the slave nodes in the group, informing them of its impending offline. Upon receiving this offline notification message, the slave nodes immediately initiate an election mechanism to ensure that a new master node is elected before the master node goes offline, thus achieving a smooth transfer of management authority.

[0043] Performance indicators are important parameters that influence the selection of new master nodes. Considering that the performance indicators obtained based on real-time static data may have the risk of current performance being qualified but about to be overloaded due to task scheduling or sudden increase in external load, for example, the current CPU utilization rate of a slave node is 30%, but a large-scale firmware upgrade task is scheduled to start in 10 minutes, at which time the load will soar to 90%.

[0044] In order to compensate for the lag of real-time performance indicators and identify potential overloaded nodes in advance, a time series analysis can be performed on the performance indicators of the node itself to determine the risk level of the node itself.

[0045] The risk level reflects the risk of a node experiencing a sudden load surge over a period of time. In practice, each slave node can obtain its own performance metrics over a set time period. Using a trained time series model, it analyzes these performance metrics over that time period to determine the node's load trend. Based on the load peaks included in the load trend, the node's risk level is determined.

[0046] The time period can be set flexibly, such as 30 minutes.

[0047] The Managed Discovery Protocol (MDP) adopted in this application allows for a flexible message format, allowing users to customize message types and extended attributes based on specific application scenarios. A version number field is included in the message header to ensure interoperability between devices with different versions. When new service features are required, new fields or attributes can be added to the message body without modifying the basic structure of the existing message format, facilitating the long-term development and functional evolution of the system.

[0048] In an embodiment of the present application, performance indicators can be synchronized within the group through the load monitoring frame newly added to the MDP protocol. In addition to CPU usage, memory occupancy, and network bandwidth utilization, performance indicators can also include some auxiliary features such as local task scheduling plans and historical load fluctuation periods. Among them, the local task scheduling plan can include the scheduled firmware upgrade timestamp, data backup task timestamp, etc. The historical load fluctuation period can be determined by counting daily load changes, such as the load peak from 9:00 to 11:00 every day.

[0049] S102: In the case that the state information of the self node and the node level meet the election condition, the device priority of the self node is determined according to the hardware resource, performance index, historical data and index weight matched with the current scene of the self node.

[0050] The more the types of parameters contained in the state information, the better the slave node participating in the election can meet the functional requirements of the new master node. Therefore, in the embodiments of the present application, the state information including the node type, hardware resource, performance index and online duration is taken as an example for introduction.

[0051] The node level can be determined by the master node evaluating the state information and management ability of each slave node.

[0052] The node level can be divided into three levels of S, A and B.

[0053] The S-level node indicates that the state information of the node meets the election condition and meets the requirements of the management ability of the master node, such as supporting concurrent management of more than 50 slave nodes, message forwarding delay in the group being less than or equal to 20 ms, and historical management instruction execution success rate being more than 99%. The election condition can be referred to the related introduction of Figure 2 , which will not be described here.

[0054] The A-level node indicates that the state information of the node meets the election condition, but does not meet the requirements of the management ability of the master node. The A-level node only receives the election start notification, has no election right and voting right, but needs to participate in data synchronization, such as reporting the locally stored management data to the new master node.

[0055] The B-level node indicates that the performance or stability of the node does not meet the election condition. The B-level node does not receive any election message and is only passively updated after the configuration is confirmed by the new master node.

[0056] In order to effectively control the number of nodes participating in the master node election, only the S-level node has the election and voting right. The device management system contains a large number of nodes, in order to facilitate the distinction, the node whose state information and node level meet the election condition can be called the election node.

[0057] In actual application, the node level of each slave node can be dynamically evaluated and updated by the current master node or the historical master node through the capability detection frame newly added by the MDP protocol, and the result is synchronized to all S-level and A-level nodes in the group.

[0058] After determining the risk level of the node, the node level can also be adjusted in combination with the risk level.

[0059] For example, the node is an S-level node, and the risk level is high risk, then it can be downgraded to an A-level node, at this time the node no longer participates in the election of the new master node.

[0060] When electing a new master node, the node with the highest device priority can be selected as the new master node. In traditional solutions, device priorities are often pre-set based on hardware configuration and purchase time. In the embodiments of the present application, in order to reflect the adaptability of a node to assume the master node function in real time, device priority can be dynamically determined based on indicators in multiple dimensions.

[0061] In specific implementations, the indicator weights that match the current scenario can be determined based on the correspondence between the set scenario type and the indicator weights. Based on the quantitative scores corresponding to each indicator, the corresponding indicator scores for the node's hardware resources, performance indicators, and historical data are determined. The weighted sum of each indicator score based on the indicator weight is then used to determine the node's device priority.

[0062] In an embodiment of the present application, the quantitative score corresponding to the hardware resources of the node can be called the basic hardware score, the quantitative score corresponding to the performance indicator can be called the real-time performance score, and the quantitative score corresponding to the historical data can be called the historical reliability score.

[0063] Table 1 shows the correspondence between scenario types and indicator weights.

[0064]

[0065] Users can preset three scenario types through the device's visual page, i.e., web page, or command line tool. The system automatically adjusts the weight of each indicator to ensure that the election results are in line with actual business goals.

[0066] The user can select the application scenario in the "Election Policy Configuration" section of the web interface. The system can also adjust the application scenario. For example, if an emergency situation is detected, such as "the master node is offline for ≥ 5 minutes" or "the number of device failures in the group is ≥ 3", the system will automatically switch to the "high performance priority mode" scenario.

[0067] Table 2 is a list of quantitative scores corresponding to each indicator

[0068]

[0069] Taking balanced mode as an example, suppose node A has a 16-core CPU and 32GB of memory. Its basic hardware score is 80 points. Its CPU utilization is 10%, its memory usage is 20%, and its latency is 5ms. Its real-time performance score is 90 points. It has not been offline for 30 days, and its historical reliability score is 100 points. The final score is 80 × 30% + 90 × 50% + 100 × 20% = 24 + 45 + 20 = 89 points. Suppose node B has a 32-core CPU and 64GB of memory. Its basic hardware score is 100 points. Its CPU utilization is 40% + its memory usage is 50% + its latency is 20ms. Its real-time performance score is 60 points. It has been offline once in 30 days, and its historical reliability score is 80 points. The final score is 100 × 30% + 60 × 50% + 80 × 20% = 30 + 30 + 16 = 76 points. This final score can be used as the device priority; higher scores indicate higher device priority.

[0070] In an embodiment of the present application, by upgrading the device priority from traditional static fixation to dynamic intelligent determination, the device priority can reflect the adaptability of the node to assume the main node function in real time, which is more in line with the complex needs of device management in a large-scale network environment.

[0071] S103: Sending a first election message to the cluster head node in the area, so that the cluster head node broadcasts the first election message to other candidate nodes of the same node level in the area.

[0072] In order to reduce the amount of data processing, the embodiment of the present application adopts a cluster broadcast method for election. For large-scale networks such as cross-computer room and multi-segment deployment, a "region cluster head" mechanism is introduced to reduce cross-region broadcast.

[0073] In practical applications, nodes within a device management system can be grouped based on group identification information. Each group has a corresponding master node, with all other nodes acting as slaves. A group can contain a large number of nodes. Each group can be divided into regions based on Internet Protocol (IP) segments, physical locations, and other factors. Within each region, one S-level node is elected as the cluster head. The initiating node first sends an election message to each cluster head node, which then broadcasts the election message to all S-level nodes within the region.

[0074] The election message can include parameters such as device priority, risk level, performance indicators, and online time.

[0075] There are often multiple nodes participating in the election in each area. To facilitate distinction, the election message sent by the node itself can be called the first election message, and the election message broadcast by other nodes can be called the second election message.

[0076] S104: Upon receiving at least one second election message broadcast by the cluster head node, analyze the device priority, risk level, performance indicators and online time contained in each election message according to the set election rules, and send an election response message to the cluster head node, so that the cluster head node can generate an area selection result based on all election response messages in the area, and broadcast the area selection result to all participating nodes.

[0077] The number of nodes participating in the election is often large. To facilitate distinction, the nodes supported by each node can be called candidate nodes.

[0078] Each election response message can include the device identifier of the node sending the election response message and the device identifiers of the candidate nodes supported by the node. Different nodes can be distinguished based on the device identifier, which can include an IP address and a hardware address (Media Access Control, MAC).

[0079] In the embodiment of the present application, when electing candidate nodes, the device priority can be the main factor, followed by the risk level and performance index, and finally the online time can be used as an auxiliary judgment. For example, the node with the highest device priority can be selected as the candidate node; if there are multiple nodes with the highest device priority, the candidate node can be selected based on the risk level and performance index; if there are multiple nodes with the same risk level and similar performance index, the node with the longest online time can be selected as the candidate node. A feasible election method can be found in Figure 3 The introduction of , will not be repeated here.

[0080] After determining the nodes it supports, the cluster head node can send an election response message to the cluster head node. The cluster head node can then generate a region selection result based on all election response messages in the region and broadcast the region selection result to all participating nodes. It should be noted that the cluster head node can also participate in elections and voting.

[0081] The region selection result records the majority voting tendency in the current region, that is, the node with the largest number of votes.

[0082] For example, if there are 10 S-level nodes in a region, 7 of which support node A and the remaining 3 support node B, the region selection result includes the device identifier of node A.

[0083] S105: receiving the region selection results sent by the cluster head nodes of each region, and analyzing all the region selection results to determine a new master node based on an election mechanism that matches the number of all current candidate nodes.

[0084] After receiving the regional selection results sent by the cluster head nodes of each region, the support rate of each candidate node can be determined based on the device identification of the candidate nodes contained in all regional selection results; the voting threshold is determined based on the number of all current candidate nodes; and the candidate node with a support rate greater than or equal to the voting threshold is selected as the new master node.

[0085] Assume that there are 10 regions in total, where regions 1 to 6 all support node A, and regions 7 to 10 all support node B. Then the support rate of node A is 60%, and the support rate of node B is 40%.

[0086] When the number of all participating nodes is small, a higher voting threshold can be set; when the number of all participating nodes is large, the voting threshold can be set lower.

[0087] All S-level nodes can participate in elections and vote. The voting threshold can be determined according to the following formula: voting threshold = 20% + number of S-level nodes * 0.5%.

[0088] When the support rate of the candidate node with the highest support rate is less than the voting threshold, a notification message to continue participating in the election is sent to the top N candidate nodes with the highest support rate, so that the top N candidate nodes with the highest support rate can re-broadcast the election message to the cluster head nodes in the area until a new master node is selected.

[0089] The value of N can be set based on the number of nodes participating in the election, and the default setting is 3.

[0090] For example, when there are ≤10 S-level nodes, the voting threshold can be set to 60% to prevent a small number of nodes from dominating the results. When there are >10 S-level nodes, the voting threshold can be set to 30%. Assuming the support rate of the candidate node with the highest support rate in the first round of voting is less than 30%, a new round of elections can be carried out for the top three candidate nodes to select the new master node.

[0091] It can be seen from the above technical solution that when the master node is offline, each slave node first evaluates its own situation. Each slave node can perform a time series analysis on the performance indicators of its own node to determine the risk level of its own node; the node level can be adjusted based on the risk level. Only when the status information and node level of the node itself meet the conditions for participation in the election can it participate in the election. In order to reflect the adaptability of the node to assume the responsibilities of the master node in real time, the device priority of the node itself can be determined based on the hardware resources, performance indicators, historical data of the node itself and the indicator weights that match the current scenario. In order to reduce the amount of data processing, the election message is processed in a cluster broadcast manner. The participating node can send the first election message to the cluster head node in the area to which it belongs, so that the cluster head node can broadcast the first election message to other participating nodes of the same node level in the area. Upon receiving at least one second election message broadcast by a cluster head node, the node can analyze the device priority, risk level, performance indicator, and online time contained in each election message according to predefined election rules. The node then sends an election response message to the cluster head node, informing it of its supported nodes. The cluster head node then generates a regional selection result based on all election response messages within the region and broadcasts the regional selection result to all participating nodes. The regional selection result includes the node with the most votes in the region. Given that regional selection results may vary for each region, to ultimately determine a unique new master node, upon receiving the regional selection results from the cluster head nodes in each region, the node can analyze all regional selection results based on an election mechanism that matches the current number of participating nodes to determine the new master node. In this technical solution, when the master node is offline, the slave node evaluates its own status and only participates in the election if its status information and node level meet the requirements. This effectively reduces the number of nodes participating in the election and avoids the heavy workload associated with a large number of slave nodes participating in the election. Each candidate node determines the nodes it supports by analyzing the device priority, risk level, performance indicators, and online time contained in each election message within the region. This ensures that the selected node can stably and reliably assume the management of the master node. Furthermore, the election mechanism is determined based on the number of candidate nodes, ensuring that the election mechanism is more in line with actual needs. This allows the most suitable new master node to be determined, ensuring that the new master node can smoothly take over all tasks of the original master node, thereby improving the reliability of device management.

[0092] In the above introduction, we mentioned using a trained time series model to analyze performance indicators within a set time period to determine the load trend of the node itself. Next, we will introduce the training process of the time series model. In this embodiment of the application, the Autoregressive Integrated Moving Average Model (ARIMA) can be used as the time series model.

[0093] In a specific implementation, the timing model can be integrated into the baseboard management controller (BMC) of the slave node as an independent module, and run using the idle computing power of the BMC without occupying the main CPU resources, thus achieving resource isolation and avoiding the occupation of management resources.

[0094] Model training utilizes offline pre-training and online fine-tuning. The initial time series model is trained based on 30 days of historical load data. This load data can be stored on the master node's Network File System (NFS). A lightweight model instance can be deployed locally on each slave node, updating model parameters every hour using the latest 10 minutes of data, effectively avoiding the network overhead of global training.

[0095] After the master node goes offline, each slave node uses the locally deployed lightweight model instance to analyze the currently acquired performance indicators to output load trends.

[0096] The typical time for a master node election is 2-5 minutes, with a 5-minute buffer period reserved to ensure a stable takeover by the new master node. Therefore, the load trend can include load forecast curves for CPU, memory, and bandwidth over the next 10 minutes.

[0097] In order to quantify the prediction results, the risk level of the node itself can be determined based on the load peak contained in the load trend.

[0098] For example, if the peak CPU usage is 60%, the corresponding score is 100 - 60 = 40. A weighted sum of the scores for each indicator yields a predicted score. When the risk is low: a predicted score ≥ 70 points corresponds to a low risk level; a predicted score ≤ 50 points < 70 points corresponds to a medium risk level; and a predicted score < 50 points corresponds to a high risk level.

[0099] In the embodiments of the present application, by performing time series analysis on the original performance indicators, load trends can be predicted, resulting in a more comprehensive assessment. A key function of the master node is to respond to group management instructions in real time. Short-term load stability is more important than historical average performance. By predicting load trends through time series analysis and considering the node's risk level when electing a new master node, the newly elected master node can be more closely aligned with the key functional requirements of the master node.

[0100] Figure 2 A flowchart of a method for selecting candidate nodes provided in an embodiment of the present application, the method comprising:

[0101] S201: Determine whether the node type is compatible with the application scenario of the master node.

[0102] Node type refers to the server type, such as storage server, computing server, or general-purpose server. Different server types are suitable for different application scenarios.

[0103] When a slave node wants to take over the work of the master node, the device type of the slave node must be compatible with the application scenario of the federated management system and be able to support all key functions of the MDP protocol, such as device discovery, group management, and data reporting. For example, in a cloud computing scenario, the node type can be a computing server; in a data storage scenario, the node type can be a storage server.

[0104] If the node type is compatible with the application scenario of the master node, S202 is executed; if the node type is incompatible with the application scenario of the master node, it indicates that the node is not eligible to participate in the election and cannot participate in the election.

[0105] S202: Determine whether the hardware resources of the own node meet the hardware resources required for the master node management.

[0106] Hardware resources can include information such as the number of central processing units (CPUs) of a node, memory capacity, and network bandwidth.

[0107] When a slave node wants to take over the work of the master node, it needs to have the hardware resources required to perform the master node's management responsibilities, such as sufficient CPU processing power, memory capacity, and network bandwidth.

[0108] If the hardware resources of the node itself meet the hardware resources required for master node management, S203 is executed; if the hardware resources of the node itself do not meet the hardware resources required for master node management, it means that the node is not eligible to run for election and cannot participate in the election.

[0109] S203: Determine whether the performance index of the own node meets the set index requirements.

[0110] Performance indicators can include average CPU usage, memory usage, and network bandwidth utilization. In addition to evaluating average CPU usage, memory usage, and network bandwidth utilization, disk basic input and output (I / O) operation response time can also be evaluated.

[0111] In a specific implementation, it is possible to determine whether the performance indicators included in the average CPU usage are less than the set first threshold, whether the memory occupancy is less than the set second threshold, whether the network bandwidth utilization is less than the set third threshold, and whether the disk basic input and output operation response time is less than the set time threshold.

[0112] When the performance indicators include an average CPU usage rate less than a set first threshold, a memory occupancy rate less than a set second threshold, a network bandwidth utilization rate less than a set third threshold, and a disk basic input and output operation response time less than a set time threshold, the performance indicators of the node itself are determined to meet the set indicator requirements.

[0113] For example, the slave node's average CPU utilization should be less than 70%, memory usage less than 60%, and network bandwidth utilization less than 80% within a specified time period. The specified time period can be the last hour. Furthermore, the slave node's disk I / O response time should be less than 100 milliseconds to ensure it can smoothly handle various management tasks when assuming master node responsibilities.

[0114] If the performance index of the node itself meets the set index requirement, S204 is executed; if the performance index of the node itself does not meet the set index requirement, it means that the node is not eligible to participate in the election and cannot participate in the election.

[0115] S204: Determine whether the online duration of the own node is greater than or equal to the set duration.

[0116] The online duration refers to the length of time a node has been running continuously and stably. For example, the online duration can be set to 24 hours.

[0117] The slave node must have been running online stably for a certain period of time to prove that it has a reliable network connection and stable hardware status and is capable of continuously managing the master node.

[0118] When the online time of the node itself is greater than or equal to the set time, S205 is executed; when the online time of the node itself is less than the set time, it indicates that the node is not eligible to participate in the election and cannot participate in the election.

[0119] S205: Determine whether the node level of the own node meets the level requirements for participation.

[0120] The node level can be dynamically set based on the node's status information, management capability information, and risk level.

[0121] The determination of the node level can be found in the introduction of S102 and will not be repeated here.

[0122] By setting node levels for nodes, it is possible to effectively avoid the problem of data processing volume surge caused by nodes that do not have the eligibility to participate in the election. Combined with the introduction of S102, in this embodiment of the application, only S-level nodes meet the election level requirements and have the right to participate in the election and vote.

[0123] When the node level of the node itself meets the election level requirements, it means that the node can participate in the election of the new master node, and S206 can be executed at this time.

[0124] In addition to the above-mentioned judgments on node type, hardware resources, performance indicators, online time, and node level, you can also judge whether the node's firmware version is compatible.

[0125] The firmware version of the slave node must be compatible with the current master node firmware version of the federal management system, or its firmware version has passed system verification and can support all necessary management functions and protocol features. If the firmware version of the slave node is too low, it may cause abnormal management functions or poor protocol communication, and therefore it is not eligible to participate in the election. Therefore, when the node level of the own node meets the level requirements for participation in the election, it can be further determined whether the firmware version of the own node is compatible with the firmware version of the master node. When the firmware version of the own node is compatible with the firmware version of the master node, execute S206 to combine the hardware resources, performance indicators, historical data of the own node and the indicator weights that match the current scenario to determine the device priority of the own node.

[0126] S206: Determine the device priority of the own node based on the hardware resources, performance indicators, historical data of the own node and the indicator weight matching the current scenario.

[0127] The method for determining the device priority can be found in the introduction of S102 and will not be repeated here.

[0128] It should be noted that the execution order of the above S201 to S205 is only one feasible implementation method. In the embodiment of the present application, the execution order of the corresponding judgment steps is not limited. As long as the slave node meets the node type and application scenario compatibility of the master node, the hardware resources meet the hardware resources required for master node management, the performance indicators meet the set indicator requirements, the online time is greater than or equal to the set time, and the node level meets the election level requirements, the slave node can participate in the election.

[0129] In an embodiment of the present application, before a node participates in the election of a new master node, the node type, hardware resources, performance indicators, online time, node level and firmware version of the node are judged through the set election conditions to ensure that the nodes participating in the election have sufficient reliability and stability.

[0130] Considering that in actual applications, a node may still send an election message to the cluster head node even though its own status information and node level do not meet the election conditions, the node can verify the received second election message before participating in the vote to ensure that each node sending the second election message is eligible to participate in the election.

[0131] In a specific implementation, when at least one second election message broadcast by the cluster head node is received, the status information and node level of each candidate node that sends each second election message can be verified according to the election conditions. The implementation method of verifying the status information and node level of each candidate node can be found in Figure 2 The introduction of , will not be repeated here.

[0132] If the status information and node level of each candidate node that sent each second election message are verified, the step of analyzing the device priority, risk level, performance indicator, and online time contained in each election message according to the set election rules can be executed to send an election response message to the cluster head node. If the status information and node level of one or more target candidate nodes that sent the second election message fail to pass the verification, the candidate node's right to participate in the election is revoked.

[0133] In the embodiment of the present application, when each node receives the second election message sent by other candidate nodes, it verifies the status information and node level of the candidate node that sent the second election message, thereby effectively ensuring that all candidate nodes are eligible to participate in the election. This prevents unqualified nodes from participating in the election.

[0134] In actual applications, the number of nodes participating in the election is usually large. When a new master node is elected, the device priority, risk level, performance index and online duration of all candidate nodes can be compared to determine the supported node. The node supported by each node can be referred to as a candidate node.

[0135] Figure 3 A flowchart of a method for determining a candidate node is provided in the embodiments of the present application. The method comprises:

[0136] S301: Selecting a candidate node with the highest priority according to the device priority included in each election message.

[0137] In the embodiments of the present application, the real-time device priority of each candidate node can be determined through weighted calculation of multi-dimensional indexes. The higher the device priority, the higher the adaptability of the candidate node to the responsibilities of the master node.

[0138] S302: In the case where the candidate node with the highest priority is one, the candidate node with the highest priority is selected as the candidate node.

[0139] Each node can compare the device priority of the candidate node, and the node with the highest priority can obtain the most votes. For example, there are three candidate nodes A, B and C, and their priorities are 80, 70 and 90 respectively. In the comparison of device priority, node C has an advantage, and node C can be selected as the candidate node.

[0140] S303: In the case where the candidate node with the highest priority is multiple, the performance index and risk level corresponding to each of the multiple candidate nodes with the highest priority are analyzed to determine the performance score of each of the multiple candidate nodes with the highest priority.

[0141] In the case where the candidate node with the highest priority is multiple, the performance index and risk level of the candidate node can be further combined for voting.

[0142] In actual applications, the residual CPU average usage rate, residual memory usage rate and residual network bandwidth utilization rate of each candidate node can be determined according to the CPU average usage rate, memory usage rate and network bandwidth utilization rate included in the performance index of each candidate node with the highest priority. The residual CPU average usage rate, residual memory usage rate and residual network bandwidth utilization rate of each candidate node are weighted and summed to obtain the performance score of each candidate node. Based on the risk level of each candidate node, the performance score of each candidate node is adjusted to vote for the candidate node with the highest performance score.

[0143] A performance score is calculated by comprehensively considering the average CPU usage, memory usage, and network bandwidth utilization. For example, the performance score = (100 - CPU usage) × 0.4 + (100 - memory usage) × 0.3 + (100 - network bandwidth utilization) × 0.3.

[0144] The risk level reflects the changing trend of a node's performance indicators over the coming period. A higher risk level indicates greater load pressure on the node. After calculating the performance score of each candidate node, the performance score can be further adjusted based on the node's risk level. For example, if the risk level is high, the performance score will be reduced by 20%, if the risk level is medium, the performance score will be reduced by 10%, and if the risk level is low, the performance score will remain unchanged.

[0145] S304: When there is only one candidate node with the highest performance score, the candidate node with the highest performance score is selected as a candidate node.

[0146] If there is only one candidate node with the highest performance score, you can directly vote for the candidate node with the highest performance score.

[0147] S305: When there are multiple candidate nodes with the highest performance scores, the candidate node with the longest online time among the candidate nodes with the highest performance scores is selected as a candidate node.

[0148] In the case where there are multiple candidate nodes with the highest performance scores, voting can be further combined with the online time of the candidate nodes.

[0149] A longer online time indicates that the node has been stable in the past and is more likely to continue to play a stable role in future management. In rare cases, there may be multiple candidate nodes with the same device priority and performance score. In this case, you can vote for the candidate node with the longer online time.

[0150] S306: Send an election response message including the device identification of the candidate node to the cluster head node.

[0151] The election response message sent by each node may include the device identifier of the node and the device identifiers of the candidate nodes supported by the node.

[0152] By dividing the area into different regions, each slave node in each region can send an election response message to the cluster head node of the region after determining the candidate node it supports. The cluster head node will summarize the voting results of each slave node and determine the candidate node with the most votes.

[0153] In the embodiment of the present application, by comparing the device priority, performance score and online time of all candidate nodes in turn, it is ensured that the elected candidate node can adapt well to the work of the master node.

[0154] Once the new master node is confirmed, the new master node immediately broadcasts a New Master Announcement Message within the group to announce its election to all slave nodes, and attaches new group management configuration information such as the new master node's IP address, port number, etc., requiring the slave nodes to update their management configuration and direct subsequent management communications to the new master node.

[0155] Considering that in actual applications, there may be multiple slave nodes claiming to be elected as the new master node. In order to ensure the uniqueness of the new master node, a legitimacy verification can be performed for this situation.

[0156] For ease of distinction, the node declaring itself the new master node is referred to as the target slave node. Upon receiving election messages from multiple target slave nodes, the node can send confirmation requests to each of them, allowing them to perform validation on each of them. The target slave node that passes validation and receives the most support from the slave nodes will be elected as the new master node.

[0157] In practice, network latency may cause some nodes to not receive election response messages from other nodes in a timely manner, resulting in multiple target slave nodes claiming to be the new master. When this happens, other nodes in the group suspend their own election activities and instead send Master Confirmation Request messages to these target slave nodes.

[0158] The target slave node must verify its legitimacy through two-way communication with the candidate nodes within the group within a preset time, proving that it has received votes exceeding the voting threshold. The node that first completes legitimacy verification and receives approval from the majority of the candidate nodes becomes the final new master node. The other target slave nodes claiming to be elected automatically relinquish their status as the new master node and revert to slave status. The preset time can be 30 seconds.

[0159] In an embodiment of the present application, by setting up legitimacy verification, each target slave node verifies that it has obtained support votes exceeding the voting threshold through two-way communication with the candidate nodes in the group. The node that completes the legitimacy verification first and is recognized by the majority of the candidate nodes will be used as the final new master node, ensuring the uniqueness of the master node in the group.

[0160] In traditional master-slave node switching, after a new master node is determined, the other slave nodes upload their locally stored management data to the new master node. Considering that a slave node has a higher probability of being elected as the new master node when it is eligible for election, in an embodiment of the present application, in order to enable the new master node to quickly take over the work of the original master node, each slave node, upon receiving at least one second election message broadcast by the cluster head node, can send its locally stored management data to the cluster head node that broadcast the second election message in a predetermined data format.

[0161] In the embodiment of the present application, the MDP protocol is used to implement device management, and the set data format is the data format corresponding to the MDP protocol.

[0162] Each slave node sends the locally stored management data to the candidate node in advance. After the candidate node is elected as the new master node, the management data of each slave node can be directly summarized, which saves data transmission time and enables the new master node to quickly take over the work of the original master node.

[0163] In a federated management system, multiple devices exist, each of which is considered a node. Group identification information can include group name, tag, username, and password. Initially, group identification information can be set by the user or by default in the system.

[0164] The group name is used to distinguish different groups.

[0165] Tags can be used to identify the type of group. For example, if a group is primarily used for cloud computing, the tag can be set to a compute tag; if a group is primarily used for data storage, the tag can be set to a storage tag. Tags can also be customized, and there are no restrictions here.

[0166] Usernames and passwords can be set by the user when creating a group, or they can be automatically generated by the system.

[0167] A group can contain one master node and multiple slave nodes. Each node supports the MDP protocol. MDP is disabled by default. Users can manually select a device as a master node (controller) or a slave node (device) and enable MDP through the device's web interface. If MDP is enabled without explicitly selecting a controller or device, the device defaults to being a slave node.

[0168] Users can set the same group name, tag, user name, and password on the device's web page as the credentials for the device to join the federated management group.

[0169] After the MDP protocol is enabled, the baseboard management controller of the master node discovers other slave nodes through the MDP protocol based on the group name, label, user name, and password.

[0170] If the node is a slave node, it determines whether it has received a networking request that matches its own group identification information. A networking request is a network request broadcast by the master node to the system containing the master node's device identification and group identification information. The group identification information includes at least the group name, tag, username, and password. If a networking request matches its own group identification information, it returns a response message to the master node based on the master node's device identification included in the networking request, thereby joining the master node's group.

[0171] After receiving the network request broadcast by the master node, if the slave node has the same group name, tag, username and password, it will automatically join the group that the master node is in. This application automatically forms a network based on group identification information instead of entering the discovery pool, which simplifies the device management process.

[0172] Slave nodes can join a specific management group by group name, but the current device management mechanism lacks a mechanism to detect group name conflicts. This may cause multiple master nodes to have the same group name, causing confusion in device management and preventing devices from correctly joining the intended management group.

[0173] Therefore, when setting a group name for the first time, a group name registration request may be sent to other master nodes in the system; wherein the group name registration request includes the group name and the device identifier of the own node.

[0174] When a confirmation message of successful registration is received from other master nodes, it means that there is no group name conflict. At this time, the master node can execute the step of broadcasting a networking request containing the device identification and group identification information of the master node to the system; wherein, the confirmation message is generated by other master nodes when confirming that the locally stored group name list does not contain the same group name as the group name registration request.

[0175] When receiving a conflict notification from other master nodes, it indicates that there is a group name conflict. At this time, the master node can generate a new group name and send a group name registration request containing the new group name to other master nodes until receiving a confirmation message of successful registration from other master nodes.

[0176] When modifying a group name, to avoid duplication, the master node can send a group name modification request to other master nodes in the system. Once the other master nodes confirm that their locally stored group name lists do not contain the same group name as the one in the modification request, they generate a confirmation message. Upon receiving this confirmation message, the master node understands that there is no group name conflict and can proceed with the group name modification.

[0177] In the embodiment of the present application, when a group name is registered for the first time or when a group name is modified, group name conflict detection is performed to ensure that the group name registered for the first time or modified does not conflict with other registered group names.

[0178] In traditional device management systems, devices broadcast notification messages on the local network using the SSDP protocol. Other devices then listen to these messages to discover devices and services on the network. The SSDP protocol lacks effective authentication and data encryption mechanisms, and communications between devices are transmitted in plain text, making it vulnerable to security threats such as man-in-the-middle attacks and counterfeit devices. Therefore, in the present embodiment, a two-way authentication mechanism is added.

[0179] When a node receives a communication request from another node, it verifies the validity of the digital certificate included in the request. If the digital certificate passes the verification, it determines whether the device identifier bound to the digital certificate matches the device identifier included in the request. If so, it sends a verification success response message to the node that sent the request.

[0180] Unlike traditional authentication methods, this application introduces a deep binding mechanism between device identification and digital certificates. This ensures that when each node connects to the federated management system, its identity is not only verified by the digital certificate, but also matches the node's hardware characteristics. This deep binding mechanism effectively prevents the theft or misuse of digital certificates, further enhancing system security.

[0181] In addition to digital certificate authentication, multi-factor authentication mechanisms can also be introduced, such as username / password, SMS verification code, hardware token, etc. When a node first connects to the system, in addition to verifying the digital certificate, the user can also be required to enter a verification code received on a pre-bound mobile phone.

[0182] In addition, a dynamic permission adjustment mechanism is set in an embodiment of the present application to adjust the permission level according to the behavior pattern of the node itself; wherein the behavior pattern may include data access location, data access frequency and / or resource utilization.

[0183] For example, when a node frequently accesses sensitive resources or CPU usage increases sharply, the node's permission level can be lowered.

[0184] In the embodiments of this application, by adjusting the device's permission level in real time based on its behavior pattern, anomalous behavior can be promptly warned and blocked, effectively identifying and preventing potential malicious behavior, and further enhancing system security. Even if a device passes initial authentication, its subsequent behavior will be strictly monitored and restricted.

[0185] In an embodiment of the present application, in order to increase the security of the message, the message content can be encrypted by combining symmetric encryption (AES) and asymmetric encryption (RSA), and a hash algorithm (SHA-256) can be used to generate a message digest, which is attached to the message for verifying the integrity of the message.

[0186] In a specific implementation, each node can periodically update its symmetric encryption key based on its own state parameters, which may include device temperature and / or network traffic. When transmitting data, the key is encrypted using asymmetric encryption, the data is encrypted using the encrypted key, and a message digest of the data is generated using a hash algorithm. The encrypted data and message digest are then transmitted to the destination node.

[0187] Unlike traditional encryption methods, this application introduces the device status parameters as encryption factors during the encryption process, so that the encrypted content is closely bound to the real-time status of the device, which not only ensures the confidentiality of the data, but also effectively prevents the data from being tampered with during transmission.

[0188] The SSDP protocol used in traditional solutions is mainly based on the User Datagram Protocol (UDP) for communication. Although this connectionless transmission method is simple and efficient, it cannot guarantee the reliable transmission of messages and is prone to message loss or duplication.

[0189] Therefore, in an embodiment of the present application, transmission resources can be allocated to the message according to its priority; wherein the priority is set based on the importance and urgency of the message.

[0190] In practical applications, messages can be assigned different priorities based on their importance and urgency. For example, emergency alerts and firmware upgrade instructions can be set to high priority, and high-priority messages will be transmitted first. When necessary, transmission resources for lower-priority messages can be preempted to ensure the timely delivery of critical messages.

[0191] When the message type matches a set first type library, the message is transmitted using transmission resources according to the Transmission Control Protocol (TCP); wherein the first type library includes at least device registration, group management, and firmware upgrade.

[0192] When the message type matches the set second type library, the message is transmitted using the transmission resources according to the User Datagram Protocol; wherein the second type library includes at least device status update and heartbeat detection.

[0193] Table 3 is a list of message transmission mechanisms

[0194]

[0195] Message types can be divided into critical messages and non-critical messages. Critical messages can include messages such as device registration, group management, and firmware upgrade instructions. Non-critical messages can include discovery messages and device status update messages. Discovery messages can include heartbeat detection.

[0196] Critical messages between the master and slave nodes can be transmitted using the TCP protocol to ensure reliable message delivery. After sending a message, the sender starts a retransmission timer, waiting for an acknowledgment (ACK) from the receiver. If the receiver successfully receives and verifies the message, it sends an ACK to the sender. If the sender does not receive an ACK within the timer's timeout, it resends the message until it receives an ACK or reaches the maximum number of retransmissions.

[0197] For critical messages, a redundant transmission mechanism can be added. In other words, in addition to TCP transmission, critical messages can also be redundantly transmitted via UDP. After receiving a TCP message, the receiver will check whether it has also received a UDP message. If the content of the two is consistent, the message confirmation is more reliable.

[0198] Non-critical messages can be transmitted using the UDP transport protocol to improve communication efficiency, and a mechanism for detecting packet loss using sequence numbers can be introduced. The receiver detects packet loss based on the message sequence number and requests the sender to retransmit the lost message. The sender then retransmits the message upon request, ensuring message integrity and accuracy.

[0199] In this embodiment of the present application, to ensure that the new master node can quickly and accurately take over management responsibilities within the group, after the cluster head node broadcasts the election message, all slave nodes can report locally stored management data, such as device status information, performance history, and unhandled alarm events, to the candidate nodes in accordance with a preset data format and transmission protocol. During the process of collecting this data, the candidate node temporarily stores and integrates the data so that it can be used immediately after being elected as the new master node.

[0200] When the node itself is the new master node, the management data reported by each slave node in the group can be integrated to obtain policy configuration information; the election message carrying the policy configuration information can be broadcast to each slave node in the group so that each slave node in the group can update the local management data according to the policy configuration information.

[0201] Managing data for fusion processing includes operations such as resolving data conflicts and checking data integrity.

[0202] For data conflicts, if different slave nodes have inconsistent descriptions of the same device status, judgment and correction are made based on the data reporting timestamp or data source reliability weight. Data integrity checks ensure that key data from all slave nodes has been received and integrated.

[0203] The new master node updates its policy configuration information based on the integrated data to ensure the continuity and consistency of management work.

[0204] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0205] The new master node broadcasts a State Update Message to the slave nodes, notifying them of its successful election as master. This message includes new management instructions and policy configurations (such as adjusting device polling cycles and updating alarm thresholds). Upon receiving this message, the slave nodes immediately update their local management state information and proceed with subsequent operations and data reporting according to the new instructions and policies, ensuring a smooth transition of the entire federated management system during the master node switchover process.

[0206] After the original master node comes back online, it is automatically downgraded to a slave node and integrated into the current federal management group to avoid management conflicts and maintain system stability.

[0207] Through the master node election mechanism described above, the federated management system based on the MDP protocol can quickly and stably elect a new master node in abnormal situations such as when the master node goes offline, ensuring that the management and collaboration of devices within the group are not affected, maintaining efficient system operation and business continuity.

[0208] In the embodiment of the present application, the above-mentioned functions such as master node switching, group name conflict management, identity authentication, and reliable data transmission are implemented based on the federal management system. These functions can be collectively referred to as federal management functions.

[0209] To enhance the flexibility of device control solutions, federation management can be designed as an independent functional module that interacts with other BMC functional modules through standardized interfaces. This modular design allows federation management to be loaded and unloaded independently of other functions. Users can manually enable and disable federation management through the BMC web management interface or command-line tools. By default, federation management is disabled; users must explicitly enable it.

[0210] The federation management function consumes a certain amount of system resources when running. The resource isolation mechanism ensures that the operation of the federation management function does not negatively impact other BMC functions.

[0211] For example, a separate resource pool can be allocated for federation management functions, or other critical functions can be prioritized when resources are scarce.

[0212] The federation management feature is designed with full consideration for compatibility with other BMC functions. For example, when federation management is disabled, other monitoring and management functions of the device continue to operate normally and are not affected by the disabling of federation management.

[0213] Federation management functions, including master node switching, node networking management, identity authentication, and reliable data transmission, can be encapsulated as independent functional modules. Modules interact with each other through standardized interfaces, allowing for the dynamic addition or upgrade of functional modules based on actual needs without impacting the overall system architecture. For example, when support for new device types or encryption algorithms is required, simply develop the corresponding functional module and integrate it into the existing system via a plug-in, enabling smooth system expansion.

[0214] Figure 4 A schematic diagram of the structure of a device control apparatus provided in an embodiment of the present application, comprising a timing analysis unit 41, a priority determination unit 42, a sending unit 43, a node analysis unit 44, and an election unit 45;

[0215] The timing analysis unit 41 is used to perform timing analysis on the performance indicators of the own node when the master node is offline to determine the risk level of the own node;

[0216] The priority determination unit 42 is configured to determine the device priority of the node itself based on its hardware resources, performance indicators, historical data, and indicator weights matching the current scenario, if the node's status information and node level meet the selection criteria;

[0217] The sending unit 43 is configured to send a first election message to the cluster head node in the area, so that the cluster head node broadcasts the first election message to other candidate nodes of the same node level in the area;

[0218] The node analysis unit 44 is configured to, upon receiving at least one second election message broadcast by the cluster head node, analyze the device priority, risk level, performance indicator, and online duration contained in each election message according to a set election rule, and send an election response message to the cluster head node, so that the cluster head node generates a region selection result based on all election response messages in the region, and broadcasts the region selection result to all participating nodes;

[0219] The election unit 45 is configured to receive the region selection results sent by the cluster head nodes of the regions, and analyze all the region selection results according to an election mechanism matched with the number of all the candidate nodes to determine a new master node.

[0220] In some embodiments, the priority determination unit comprises an index weight determination subunit, a quantification subunit and a summation subunit.

[0221] The index weight determination subunit is configured to determine the index weight matched with the current scene based on a set correspondence between scene types and index weights.

[0222] The quantification subunit is configured to determine the index scores of the hardware resources, performance indicators and historical data of the self node respectively according to the quantification scores corresponding to the indicators.

[0223] The summation subunit is configured to weight and sum the index scores according to the index weights to determine the device priority of the self node.

[0224] In some embodiments, the timing analysis unit comprises an acquisition subunit, a prediction subunit and a grade determination subunit.

[0225] The acquisition subunit is configured to acquire the performance indicators of the self node in a set time period.

[0226] The prediction subunit is configured to analyze the performance indicators in the set time period by using the trained timing model to determine the load trend of the self node.

[0227] The grade determination subunit is configured to determine the risk grade of the self node based on the load peak value contained in the load trend.

[0228] In some embodiments, for the identification of the offline master node, the device further comprises a heartbeat detection unit, a response detection unit and an offline notification unit.

[0229] The heartbeat detection unit is configured to send heartbeat detection messages to the master node at a set time interval, and determine that the master node is offline if no heartbeat response message is received from the master node for a continuous set number of time intervals.

[0230] The response detection unit is configured to determine that the master node is offline if no response is received from the master node for a continuous set number of times within a timeout time in a data interaction scene.

[0231] The offline notification unit is configured to determine that the master node is offline if an offline notification message broadcast by the master node is acquired.

[0232] In some embodiments, the priority determination unit comprises a first judgment subunit, a second judgment subunit, a third judgment subunit, a fourth judgment subunit, a fifth judgment subunit, and a determination subunit.

[0233] The first judgment subunit is configured to determine whether the node type is compatible with the application scenario of the master node.

[0234] The second judgment subunit is configured to determine whether the hardware resources of the node meet the hardware resources required for the master node management, when the node type is compatible with the application scenario of the master node.

[0235] The third judgment subunit is configured to determine whether the performance indicators of the node meet the set indicator requirements, when the hardware resources of the node meet the hardware resources required for the master node management.

[0236] The fourth judgment subunit is configured to determine whether the online duration of the node is greater than or equal to the set duration, when the performance indicators of the node meet the set indicator requirements.

[0237] The fifth judgment subunit is configured to determine whether the node level of the node meets the participation level requirements, when the online duration of the node is greater than or equal to the set duration; wherein the node level is dynamically set based on the state information, management capability information, and risk level of the node.

[0238] The determination subunit is configured to determine the device priority of the node according to the hardware resources, performance indicators, historical data, and indicator weights matching the current scenario, when the node level of the node meets the participation level requirements.

[0239] In some embodiments, the third judgment subunit is configured to determine whether the CPU average usage contained in the performance indicators is less than a set first threshold, the memory occupancy is less than a set second threshold, the network bandwidth utilization is less than a set third threshold, and the disk basic input / output operation response time is less than a set time threshold.

[0240] When the CPU average usage contained in the performance indicators is less than the set first threshold, the memory occupancy is less than the set second threshold, the network bandwidth utilization is less than the set third threshold, and the disk basic input / output operation response time is less than the set time threshold, it is determined that the performance indicators of the node meet the set indicator requirements.

[0241] In some embodiments, it further comprises a version judgment unit.

[0242] The version judging unit is configured to judge whether the firmware version of the node itself is compatible with the firmware version of the master node when the online duration of the node itself is greater than or equal to the set duration; and trigger the determination sub-unit to perform the step of determining the device priority of the node itself according to the hardware resource, the performance index, the historical data and the index weight matched with the current scene when the firmware version of the node itself is compatible with the firmware version of the master node.

[0243] In some embodiments, the verification unit is further included;

[0244] The verification unit is configured to verify the state information and the node level of each candidate node sending each second election message according to the candidate condition when at least one second election message broadcast by the cluster head node is received; and trigger the node analysis unit to perform the step of analyzing the device priority, the risk level, the performance index and the online duration contained in each election message according to the set election rule to send an election response message to the cluster head node when the state information and the node level of each candidate node sending each second election message are verified.

[0245] In some embodiments, the node analysis unit includes a selection sub-unit, a first sub-unit, a scoring sub-unit, a second sub-unit, a third sub-unit and a sending sub-unit;

[0246] The selection sub-unit is configured to select the candidate node with the highest priority according to the device priority contained in each election message.

[0247] The first sub-unit is configured to take the candidate node with the highest priority as the candidate node when the candidate node with the highest priority is one.

[0248] The scoring sub-unit is configured to comprehensively analyze the performance index and the risk level corresponding to each of the candidate nodes with the highest priority to determine the performance score corresponding to each of the candidate nodes with the highest priority when the candidate nodes with the highest priority are multiple.

[0249] The second sub-unit is configured to take the candidate node with the highest performance score as the candidate node when the candidate node with the highest performance score is one.

[0250] The third sub-unit is configured to take the candidate node with the longest online duration among the candidate nodes with the highest performance score as the candidate node when the candidate nodes with the highest performance score are multiple.

[0251] The sending sub-unit is configured to send an election response message containing the device identifier of the candidate node to the cluster head node.

[0252] In some embodiments, the scoring unit is configured to determine, according to the performance indicators contained in the CPU average usage rate, the memory occupancy rate and the network bandwidth utilization rate of each candidate node with the highest priority, the residual CPU average usage rate, the residual memory occupancy rate and the residual network bandwidth utilization rate of each candidate node;

[0253] The residual CPU average usage rate, the residual memory occupancy rate and the residual network bandwidth utilization rate of each candidate node are weighted and summed to obtain the performance score of each candidate node.

[0254] Based on the risk level of each candidate node, the performance score of each candidate node is adjusted.

[0255] In some embodiments, the election unit comprises a support rate determination subunit, a threshold determination subunit and a new master node determination subunit.

[0256] The support rate determination subunit is configured to determine the support rate of each candidate node according to the device identifier of the candidate node contained in the all-region selection result.

[0257] The threshold determination subunit is configured to determine the voting threshold according to the number of all candidate nodes at present.

[0258] The new master node determination subunit is configured to take the candidate node with a support rate greater than or equal to the voting threshold as the new master node.

[0259] In some embodiments, the election unit further comprises a candidate selection unit.

[0260] The candidate selection unit is configured to, in the case that the support rate of the candidate node with the highest support rate is less than the voting threshold, send a notification message of continuing to participate to the first N candidate nodes with the highest support rate, so that the first N candidate nodes with the highest support rate rebroadcast the election message to the cluster head nodes in the region until the new master node is selected.

[0261] In some embodiments, after analyzing all the region selection results to determine the new master node according to the election mechanism matched with the number of all candidate nodes at present, the election unit further comprises a legality verification unit.

[0262] The legality verification unit is configured to, in the case that the elected message broadcast by the plurality of target slave nodes is obtained, send a confirmation request to the plurality of target slave nodes so that the plurality of target slave nodes respectively perform legality verification; and take the target slave node that passes the legality verification and obtains the most slave node support as the new master node.

[0263] In some embodiments, the election unit further comprises a data transmission unit.

[0264] The data transmission unit is configured to, in a case where at least one second election message broadcast by a cluster head node is received, send locally stored management data to the cluster head node that broadcasts the second election message in a set data format.

[0265] In some embodiments, in a case where the master node is offline, the performance indicators of the self node are subjected to time series analysis to determine the risk level of the self node before the network formation determination unit and the feedback unit are further included;

[0266] The network formation determination unit is configured to, in a case where the self node is a slave node, determine whether a network formation request with the same group identification information as that of the self node is obtained; wherein the network formation request is a network formation request broadcast by the master node to the system, and the network formation request contains the device identification and the group identification information of the master node; and wherein the group identification information at least includes a group name, a label, a user name, and a password.

[0267] The feedback unit is configured to, in a case where the network formation request with the same group identification information as that of the self node is obtained, feed back response information to the master node according to the device identification of the master node contained in the network formation request, so as to join the group in which the master node is located.

[0268] In some embodiments, before the network formation request containing the device identification and the group identification information of the master node is broadcast to the system, a registration unit and a modification unit are further included;

[0269] The registration unit is configured to, in a case where the group name is set for the first time, send a group name registration request to other master nodes in the system; wherein the group name registration request contains the group name and the device identification of the self node; and in a case where a confirmation message of successful registration sent by other master nodes is received, the step of broadcasting the network formation request containing the device identification and the group identification information of the master node to the system is executed; wherein the confirmation message is a message generated by other master nodes when it is confirmed that the same group name as that contained in the group name registration request does not exist in a locally stored group name list;

[0270] The modification unit is configured to, in a case where a conflict notification sent by other master nodes is received, send a group name registration request containing a new group name to other master nodes until a confirmation message of successful registration sent by other master nodes is received.

[0271] In some embodiments, a certificate verification unit, an identification verification unit, and a response unit are further included;

[0272] The certificate verification unit is configured to, in a case where a communication request sent by other nodes is received, perform legality verification on a digital certificate carried in the communication request;

[0273] The identification verification unit is configured to, in a case where the digital certificate passes the legality verification, determine whether the device identification bound to the digital certificate is consistent with the device identification carried in the communication request;

[0274] The response unit is used to feed back a verification-passed response message to the node that sent the communication request when the device identification bound to the digital certificate is consistent with the device identification carried in the communication request.

[0275] In some embodiments, an adjustment unit is further included;

[0276] The adjustment unit is used to adjust the permission level according to the behavior pattern of its own node; wherein the behavior pattern includes data access location, data access frequency and / or resource usage.

[0277] In some embodiments, an update unit and an encryption transmission unit are further included;

[0278] An updating unit, configured to periodically update the symmetric encryption key based on the state parameters of its own node; wherein the state parameters include device temperature and / or network traffic;

[0279] The encryption transmission unit is used to encrypt the key using asymmetric encryption when transmitting data, encrypt the data using the encrypted key, and generate a message digest of the data based on the hash algorithm; and transmit the encrypted data and the message digest to the destination node.

[0280] In some embodiments, further comprising a distribution unit, a first transmission unit, and a second transmission unit;

[0281] an allocation unit, configured to allocate transmission resources to messages according to their priorities, wherein the priorities are set based on the importance and urgency of the messages;

[0282] A first transmission unit is configured to transmit the message using transmission resources according to a transmission control protocol when a message type of the message matches a set first type library; wherein the first type library includes at least device registration, group management, and firmware upgrade;

[0283] The second transmission unit is used to transmit the message according to the user datagram protocol using transmission resources when the message type matches the set second type library; wherein the second type library at least includes device status update and heartbeat detection.

[0284] In some embodiments, further comprising a fusion unit and a broadcast unit;

[0285] The fusion unit is used to fuse the management data reported by each slave node in the group to obtain policy configuration information when the node itself is the new master node;

[0286] The broadcast unit is used to broadcast the election message carrying the policy configuration information to each slave node in the group, so that each slave node in the group can update the local management data according to the policy configuration information.

[0287] It can be seen from the above technical solution that when the master node is offline, each slave node first evaluates its own situation. Each slave node can perform a time series analysis on the performance indicators of its own node to determine the risk level of its own node; the node level can be adjusted based on the risk level. Only when the status information and node level of the node itself meet the conditions for participation in the election can it participate in the election. In order to reflect the adaptability of the node to assume the responsibilities of the master node in real time, the device priority of the node itself can be determined based on the hardware resources, performance indicators, historical data of the node itself and the indicator weights that match the current scenario. In order to reduce the amount of data processing, the election message is processed in a cluster broadcast manner. The participating node can send the first election message to the cluster head node in the area to which it belongs, so that the cluster head node can broadcast the first election message to other participating nodes of the same node level in the area. Upon receiving at least one second election message broadcast by a cluster head node, the node can analyze the device priority, risk level, performance indicator, and online time contained in each election message according to predefined election rules. The node then sends an election response message to the cluster head node, informing it of its supported nodes. The cluster head node then generates a regional selection result based on all election response messages within the region and broadcasts the regional selection result to all participating nodes. The regional selection result includes the node with the most votes in the region. Given that regional selection results may vary for each region, to ultimately determine a unique new master node, upon receiving the regional selection results from the cluster head nodes in each region, the node can analyze all regional selection results based on an election mechanism that matches the current number of participating nodes to determine the new master node. In this technical solution, when the master node is offline, the slave node evaluates its own status and only participates in the election if its status information and node level meet the requirements. This effectively reduces the number of nodes participating in the election and avoids the heavy workload associated with a large number of slave nodes participating in the election. Each candidate node determines the nodes it supports by analyzing the device priority, risk level, performance indicators, and online time contained in each election message within the region. This ensures that the selected node can stably and reliably assume the management of the master node. Furthermore, the election mechanism is determined based on the number of candidate nodes, ensuring that the election mechanism is more in line with actual needs. This allows the most suitable new master node to be determined, ensuring that the new master node can smoothly take over all tasks of the original master node, thereby improving the reliability of device management.

[0288] For the description of the features in the embodiment corresponding to the device control apparatus, reference can be made to the relevant description of the embodiment corresponding to the device control method, and no further details will be given here.

[0289] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned device control method embodiments.

[0290] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned device control method embodiments when run.

[0291] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0292] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned device control method embodiments are implemented.

[0293] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned device control method embodiments are implemented.

[0294] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0295] The above is a detailed introduction to a device control method, apparatus, electronic device, storage medium and product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A device control method, characterized in that: include: When the master node is offline, perform time series analysis on the performance indicators of the node itself to determine the risk level of the node itself; When the node's status information and node level meet the eligibility criteria, the node's device priority is determined based on the node's hardware resources, performance indicators, historical data, and indicator weights that match the current scenario. The status information includes node type, hardware resources, performance indicators, and online time. The node level is dynamically set based on the node's status information, management capability information, and risk level. Sending a first election message to a cluster head node in the area, so that the cluster head node broadcasts the first election message to other candidate nodes of the same node level in the area; Upon receiving at least one second election message broadcast by the cluster head node, analyzing the device priority, risk level, performance indicator, and online duration contained in each election message according to a set election rule, and sending an election response message to the cluster head node, so that the cluster head node generates a region selection result based on all election response messages in the region, and broadcasts the region selection result to all participating nodes; Receive the region selection results sent by the cluster head nodes of each region, and analyze all the region selection results to determine a new master node based on an election mechanism that matches the number of all current candidate nodes.

2. The device control method according to claim 1, wherein: Determine the device priority of your own node based on its own hardware resources, performance indicators, historical data, and indicator weights that match the current scenario, including: Based on the correspondence between the set scenario type and indicator weight, determine the indicator weight that matches the current scenario; According to the quantitative scores corresponding to each indicator, determine the corresponding indicator scores of the hardware resources, performance indicators, and historical data of the node itself; The scores of each indicator are weighted and summed according to the indicator weights to determine the device priority of the own node.

3. The device control method according to claim 1, wherein: Perform time series analysis on the performance indicators of the node itself to determine the risk level of the node itself, including: Obtain the performance indicators of its own node within the set time period; Analyze the performance indicators within the set time period using the trained time series model to determine the load trend of the node itself; Based on the load peak value included in the load trend, the risk level of the own node is determined.

4. The device control method according to claim 1, wherein: For identifying that the master node is offline, the method includes: Send heartbeat detection messages to the master node at set time intervals; If no heartbeat response message fed back by the master node is received within a set number of consecutive time intervals, the master node is determined to be offline; Or in a data interaction scenario, if no response is received from the master node within a set number of consecutive times within a timeout period, the master node is determined to be offline; Alternatively, when an offline notification message broadcast by the master node is obtained, it is determined that the master node is offline.

5. The device control method according to claim 1, wherein: When the node's status information and node level meet the selection criteria, the node's device priority is determined based on its hardware resources, performance indicators, historical data, and indicator weights that match the current scenario. This includes: Determine whether the node type is compatible with the application scenario of the master node; If the node type is compatible with the application scenario of the master node, determine whether the hardware resources of the own node meet the hardware resources required for management of the master node; If the hardware resources of the own node meet the hardware resources required for the master node management, determine whether the performance indicators of the own node meet the set indicator requirements; When the performance index of the node itself meets the set index requirements, determine whether the online time of the node itself is greater than or equal to the set time; If the online time of the own node is greater than or equal to the set time, determine whether the node level of the own node meets the candidate level requirements; wherein the node level is dynamically set based on the node's status information, management capability information and risk level; When the node level of its own node meets the requirements for the election level, the device priority of its own node is determined based on its own node's hardware resources, performance indicators, historical data, and indicator weights that match the current scenario.

6. The device control method according to claim 5, characterized in that: Determining whether the performance indicators of the node meet the set indicator requirements includes: Determine whether the performance indicators include an average CPU usage rate less than a set first threshold, a memory occupancy rate less than a set second threshold, a network bandwidth utilization rate less than a set third threshold, and a disk basic input / output operation response time less than a set time threshold; When the performance indicators include an average CPU usage rate less than a set first threshold, a memory occupancy rate less than a set second threshold, a network bandwidth utilization rate less than a set third threshold, and a disk basic input and output operation response time less than a set time threshold, it is determined that the performance indicators of the node itself meet the set indicator requirements.

7. The device control method according to claim 5, characterized in that: After determining whether the online duration of the node itself is greater than or equal to the set duration, the following steps are also included: When the online time of the own node is greater than or equal to the set time, determine whether the firmware version of the own node is compatible with the firmware version of the master node; When the firmware version of the own node is compatible with the firmware version of the master node, the device priority of the own node is determined based on the hardware resources, performance indicators, historical data of the own node and the indicator weight matching the current scenario.

8. The device control method according to claim 1, wherein: Before analyzing the device priority, risk level, performance index, and online duration contained in each election message according to the set election rules to send an election response message to the cluster head node, the method further includes: Upon receiving at least one second election message broadcast by the cluster head node, verifying the status information and node level of each candidate node that sends each second election message according to the election condition; When the status information and node level of each participating node that sends each second election message are verified, the device priority, risk level, performance index and online time contained in each election message are analyzed according to the set election rules to send an election response message to the cluster head node.

9. The device control method according to claim 1, wherein: Analyze the device priority, risk level, performance index, and online duration contained in each election message according to the set election rules, and send an election response message to the cluster head node, including: Based on the device priority contained in each election message, the node with the highest priority is selected; When there is only one candidate node with the highest priority, the candidate node with the highest priority is selected as the candidate node; In the case where there are multiple candidates with the highest priority, a comprehensive analysis is conducted on the performance indicators and risk levels corresponding to the multiple candidates with the highest priority to determine the performance scores corresponding to the multiple candidates with the highest priority; When there is only one candidate node with the highest performance score, the candidate node with the highest performance score is selected as the candidate node; If there are multiple candidate nodes with the highest performance scores, the candidate node with the longest online time among the candidate nodes with the highest performance scores is selected as the candidate node; An election response message including the device identification of the candidate node is sent to the cluster head node.

10. The device control method according to claim 9, characterized in that: Comprehensively analyze the performance indicators and risk levels of the highest priority candidate nodes to determine the performance scores of the highest priority candidate nodes, including: Based on the performance indicators of each candidate node with the highest priority, including the average CPU usage, memory usage, and network bandwidth usage, the remaining average CPU usage, remaining memory usage, and remaining network bandwidth usage of each candidate node are determined; The weighted sum of the remaining average CPU usage, remaining memory usage, and remaining network bandwidth utilization of each candidate node is taken to obtain the performance score of each candidate node. Based on the risk level of each candidate node, the performance score of each candidate node is adjusted.

11. The device control method according to claim 1, wherein: Based on an election mechanism that matches the number of all currently participating nodes, the results of all regional selections are analyzed to determine the new master node, including: Determining the support rate of each candidate node based on the device identifiers of the candidate nodes included in all the region selection results; The voting threshold is determined based on the number of all currently participating nodes; The candidate node whose support rate is greater than or equal to the voting threshold is selected as the new master node.

12. The device control method according to claim 11, characterized in that: Also includes: When the support rate of the candidate node with the highest support rate is less than the voting threshold, a notification message of continuing to participate in the election is sent to the top N candidate nodes with the highest support rate, so that the top N candidate nodes with the highest support rate can re-broadcast the election message to the cluster head nodes in the area until a new master node is selected.

13. The device control method according to claim 1, wherein: After analyzing all the regional selection results to determine the new master node based on an election mechanism that matches the number of all currently participating nodes, the following steps are also included: In the case of obtaining election messages broadcast by multiple target slave nodes, sending confirmation requests to the multiple target slave nodes so that the multiple target slave nodes can perform legitimacy verification respectively; The target slave node that has passed the legitimacy verification and has the most support from the slave nodes will be used as the new master node.

14. The device control method according to claim 1, wherein: Also includes: In the case of receiving at least one second election message broadcasted by the cluster head node, locally stored management data is sent to the cluster head node that broadcasts the second election message according to a set data format.

15. The device control method according to claim 1, wherein: When the master node is offline, the time series analysis of the node's performance indicators is performed to determine the node's risk level. This also includes: When the node is a slave node, determining whether a networking request identical to its own group identification information has been obtained; wherein the networking request is a networking request broadcast by the master node to the system containing the master node's device identification and group identification information; wherein the group identification information includes at least a group name, a tag, a user name, and a password; In the case of obtaining a networking request with the same group identification information as its own, it feeds back response information to the master node according to the device identification of the master node included in the networking request, so as to join the group where the master node belongs.

16. The device control method according to claim 15, characterized in that: Before broadcasting the networking request containing the master node's device identification and group identification information to the system, it also includes: When setting a group name for the first time, a group name registration request is sent to other master nodes in the system; wherein the group name registration request includes the group name and the device identifier of the node itself; Upon receiving a confirmation message of successful registration from another master node, performing a step of broadcasting a networking request including the device identification and group identification information of the master node to the system; wherein the confirmation message is generated by the other master node when confirming that the same group name as that included in the group name registration request does not exist in the locally stored group name list; In the case of receiving a conflict notification sent by other master nodes, a group name registration request containing a new group name is sent to other master nodes until a confirmation message of successful registration is received from other master nodes.

17. The device control method according to claim 1, wherein: Also includes: Upon receiving a communication request from another node, verify the legitimacy of the digital certificate carried in the communication request; If the digital certificate passes the legitimacy verification, determining whether the device identifier bound to the digital certificate is consistent with the device identifier carried in the communication request; In the case that the device identification bound to the digital certificate is consistent with the device identification carried in the communication request, a response message indicating that the verification is successful is fed back to the node that sent the communication request.

18. The device control method according to claim 1, wherein: Also includes: The permission level is adjusted according to the behavior pattern of the node itself; wherein the behavior pattern includes data access location, data access frequency and / or resource usage.

19. The device control method according to claim 1, wherein: Also includes: Regularly update the symmetric encryption key based on the node's own state parameters; wherein the state parameters include device temperature and / or network traffic; In the case of transmitting data, encrypting the key using an asymmetric encryption method, encrypting the data using the encrypted key, and generating a message digest of the data according to a hash algorithm; The encrypted data and the message digest are transmitted to the destination node.

20. The device control method according to claim 1, wherein: Also includes: Allocating transmission resources to messages based on their priority; wherein the priority is set based on the importance and urgency of the message; When a message type of the message matches a set first type library, transmitting the message using the transmission resource in accordance with the transmission control protocol; wherein the first type library includes at least device registration, group management, and firmware upgrade; When the message type of the message matches a set second type library, the message is transmitted using the transmission resource in accordance with the User Datagram Protocol; wherein the second type library includes at least device status update and heartbeat detection.

21. The device control method according to claim 1, wherein: Also includes: When the own node is the new master node, the management data reported by each slave node in the group is integrated to obtain policy configuration information; Broadcasting the election message carrying the policy configuration information to each slave node in the group, so that each slave node in the group updates local management data according to the policy configuration information.

22. A device control device, characterized in that: It includes a timing analysis unit, a priority determination unit, a sending unit, a node analysis unit and an election unit; The timing analysis unit is used to perform timing analysis on the performance indicators of its own node when the master node is offline to determine the risk level of its own node; The priority determination unit is configured to determine the device priority of its own node based on its own node's hardware resources, performance indicators, historical data, and indicator weights matching the current scenario, when the node's own status information and node level meet the selection criteria; the status information includes node type, hardware resources, performance indicators, and online time; and the node level is dynamically set based on the node's status information, management capability information, and risk level. The sending unit is configured to send a first election message to a cluster head node in the area, so that the cluster head node broadcasts the first election message to other candidate nodes of the same node level in the area; The node analysis unit is configured to, upon receiving at least one second election message broadcast by the cluster head node, analyze the device priority, risk level, performance indicator, and online duration contained in each election message according to a set election rule, and send an election response message to the cluster head node, so that the cluster head node generates a region selection result based on all election response messages in the region, and broadcasts the region selection result to all participating nodes; The election unit is configured to receive the region selection results sent by the cluster head nodes of each region, and analyze all the region selection results to determine a new master node based on an election mechanism that matches the number of all currently participating nodes.

23. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the device control method according to any one of claims 1 to 21 when executing the computer program.

24. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the device control method according to any one of claims 1 to 21 are implemented.

25. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the device control method according to any one of claims 1 to 21 are implemented.

Citation Information

Patent Citations

  • Main monitoring node selection method and device in distributed cluster

    CN107948260A

  • Node election method and device, storage medium and electronic equipment

    CN116170289A