A method and system for managing a cluster of server cryptomachines

CN121690616BActive Publication Date: 2026-09-25SHANDONG YUANDUN NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610116098.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-28
Publication Date
2026-09-25
Estimated Expiration
2046-01-28

AI Technical Summary

Technical Problem

由于管理节点承担着整个集群的核心管理功能,一旦管理节点出现故障,如硬件损坏、软件崩溃或遭受网络攻击等,将导致整个集群的管理功能瘫痪,无法进行任务分配、状态监控和故障恢复等操作,进而严重影响业务的正常运行和数据安全保障

Benefits of technology

1.本申请通过设置多个管理节点组成管理节点组,即使部分管理节点出现故障,剩余的管理节点仍能通过共识算法重新选举出新的主节点,提高了集群管理系统的可靠性以及集群管理的连续性。从节点实时监控主节点的状态信息,一旦主节点状态不符合预期,能够迅速重新选举新的主节点,减少了因管理节点故障导致的服务中断时间,进一步提升了集群的高可用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121690616B_ABST
    Figure CN121690616B_ABST
Patent Text Reader

Abstract

The application provides a kind of management method and system of server cryptomachine cluster, it is related to the technical field of cryptomachine management, the method comprises: first, in cluster, multiple management nodes are set to constitute management node group, the state information of each management node is collected and shared to the rest of management node by heartbeat protocol;Afterwards, according to state information, master node is selected in management node group using consensus algorithm, the rest is slave node, slave node real-time monitoring master node state, re-election when abnormal;Afterwards, when new task arrives, master node decides and executes locally according to state information using preset distributed task allocation algorithm.The application can improve the reliability of server cryptomachine cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cryptographic machine management, and in particular to a management method and system for a server cryptographic machine cluster. Background Technology

[0002] In today's digital information age, data security is of paramount importance. Server cryptographic machines, as core devices for ensuring data security, are widely used in numerous fields such as finance and enterprises, undertaking critical security operations such as data encryption, decryption, digital signatures, and identity authentication. With the continuous expansion of business scale and the rapid growth of data volume, a single server cryptographic machine can no longer meet the demands of high concurrency and high reliability. Therefore, server cryptographic machine clusters have emerged, achieving load balancing, fault tolerance, and performance scalability by having multiple server cryptographic machines work collaboratively to cope with increasingly complex and demanding security tasks.

[0003] Currently, the management of server cryptographic machine clusters primarily employs a centralized management architecture. This architecture features a dedicated management node responsible for collecting status information from each server cryptographic machine within the cluster, such as device operating status, load, and key storage status. Based on this information, the management node performs unified scheduling and resource allocation for the cluster. The management node establishes connections with each server cryptographic machine via a specific communication protocol, periodically sending query commands to retrieve status data. Simultaneously, according to preset rules and algorithms, such as round-robin and least-connections algorithms, it assigns new security tasks to appropriate server cryptographic machines for processing.

[0004] The above solution has a significant risk of single point of failure. Since the management node is responsible for the core management functions of the entire cluster, if the management node fails, such as due to hardware damage, software crash, or network attack, the management functions of the entire cluster will be paralyzed, making it impossible to perform operations such as task allocation, status monitoring, and fault recovery. This will seriously affect the normal operation of business and data security. Summary of the Invention

[0005] To improve the reliability of server cryptographic machine clusters, this application provides a management method and system for server cryptographic machine clusters.

[0006] Firstly, this application provides a management method for a server cryptographic machine cluster, employing the following technical solution: A method for managing a server cryptographic machine cluster includes the following steps: Multiple management nodes are set up in the cluster to form a management node group. The status information of each management node in the management node group is collected and shared with the remaining management nodes through a heartbeat protocol. Based on the status information, a consensus algorithm is used to elect a master node in the management node group, and the remaining management nodes in the management node group are used as slave nodes. The slave nodes monitor the status information of the master node in real time, and re-execute this step when the status information of the master node does not meet expectations. When a new task arrives, the master node makes a task allocation decision locally based on the status information using a preset distributed task allocation algorithm, obtains a preliminary decision result, and allocates the new task according to the preliminary decision result.

[0007] This application establishes a management node group with multiple management nodes. Even if some management nodes fail, the remaining management nodes can still re-elect a new master node through a consensus algorithm, improving the reliability and continuity of the cluster management system. Slave nodes monitor the master node's status information in real time. If the master node's status does not meet expectations, a new master node can be quickly re-elected, reducing service interruption time caused by management node failures and further enhancing the cluster's high availability.

[0008] When making task allocation decisions, the master node can leverage the shared state information of all nodes in the management node group to make its initial decisions, based on a pre-defined distributed task allocation algorithm, more scientific and reasonable. This better matches the actual capabilities of each node in the cluster, minimizing uneven task allocation that could lead to some nodes being overloaded while others are idle, thereby improving the overall resource utilization and task processing efficiency of the cluster. When a new task arrives, the master node directly makes the task allocation decision and executes it locally, minimizing intermediate steps and communication overhead in the task processing process. This allows tasks to be processed faster, improving the cluster's response speed and processing capacity.

[0009] Optionally, after obtaining the preliminary decision results, the method further includes: The preliminary decision results are shared to at least one slave node through a distributed coordination service. The slave node that receives the preliminary decision results verifies them. If the verification passes, a new task is assigned according to the preliminary decision results. If the verification fails, the master node and the slave node that received the preliminary decision results re-determine the task allocation through a consensus algorithm to obtain the final decision results, which are then updated to reflect the preliminary decision results.

[0010] This application shares the preliminary decision results with at least one slave node for verification through a distributed coordination service, essentially introducing a multi-node auditing mechanism. When a single master node makes a decision, any errors could negatively impact the entire cluster's task execution. However, with a multi-node auditing mechanism, even if the master node's preliminary decision is flawed, it can be promptly detected and corrected during the slave node verification phase, minimizing the large-scale execution of erroneous decisions within the cluster. This reduces the risks and losses caused by decision-making errors and enhances the reliability of cluster task processing. Only verified preliminary decisions are executed by slave nodes, ensuring that all nodes participating in task execution operate based on the same verified decision results. This minimizes the confusion and inconsistency in task execution caused by different nodes interpreting or executing decisions differently, improving the uniformity and coordination of cluster task execution.

[0011] Optionally, the method further includes: Obtain the physical location, network topology, and load of each server's cryptographic machine, and use the K-means algorithm to divide the logical regions, with each logical region containing at least one management node; Obtain the server cryptographic machine that is closest to the physical location of the new task initiator, and denote it as the target cryptographic machine. Denote the logical area where the target cryptographic machine is located as the target area. If the load of the server cryptographic machine in the target area does not reach the preset load threshold, then the new task is assigned to the server cryptographic machine in the target area. If the load of all server cryptographic machines in the target area has reached the preset load threshold and the load of server cryptographic machines in at least one adjacent logical area of ​​the target area has not reached the preset load threshold, then new tasks will be assigned to server cryptographic machines in the adjacent logical areas of the target area in order of load from low to high.

[0012] This application employs the K-means algorithm to divide logical regions, comprehensively considering the physical location of server cryptographic machines and network topology information. This allows the master node to more rationally utilize server cryptographic machine resources in different regions when allocating tasks, minimizing resource idleness or over-concentration. This application designates the server cryptographic machine physically closest to the initiator of a new task as the target cryptographic machine and prioritizes assigning tasks to the region where this target machine is located (if load allows). Due to the shorter physical distance and data transmission path, network latency is effectively reduced, accelerating task processing.

[0013] This application also sets up at least one management node in each logical region. The management node can be responsible for tasks such as monitoring and configuring the cryptographic machines of the servers in its logical region, which reduces the complexity and difficulty of management and improves the efficiency and accuracy of system management.

[0014] Optionally, in the process of sequentially assigning new tasks to server cryptographic machines within adjacent logical areas of the target area, the method further includes: The network topology changes of the server cryptographic machines in each adjacent logical area are monitored in real time. If the task allocation path is abnormal due to the network topology change, the master node will re-make the task allocation decision based on the current physical location of each logical area and the status information of each server cryptographic machine, and obtain the re-allocation decision result.

[0015] This application monitors network topology changes of server cryptographic machines within adjacent logical regions in real time, enabling timely detection of factors that may lead to abnormal task allocation paths. Upon detection of an anomaly, the master node reacts swiftly, re-evaluating task allocation decisions to minimize the risk of tasks failing to allocate or execute properly due to network topology changes. During task allocation, network topology changes may disrupt the allocation path, affecting task progress. This application, through real-time monitoring and timely re-allocation decisions, can quickly reassign tasks to suitable server cryptographic machines, improving task reliability.

[0016] Optionally, when a server cryptographic machine malfunctions, the method further includes: The logical area where the faulty server cryptographic machine is located is designated as the fault area. The faulty server cryptographic machine sends the fault information to the management node in the fault area via a heartbeat protocol. The management node in the fault area uses a preset distributed task allocation algorithm to reassign the tasks assigned to the faulty server cryptographic machine to the remaining server cryptographic machines in the fault area.

[0017] By adopting the above scheme, after the server cryptographic machine fails, this application can quickly send the fault information to the management node in the fault area through the heartbeat protocol, so as to avoid the task being suspended for a long time due to the failure to detect the fault in time, and improve the reliability of the system.

[0018] Optionally, after verification and before assigning new tasks based on preliminary decision results, the method further includes: Extract the server cryptographic machine that will execute the new task from the preliminary decision results and denote it as the execution node. Select m execution nodes, inject abnormal data into the selected execution nodes, use the execution nodes after injecting abnormal data to execute the new task, and obtain the execution result and execution log of the new task. Calculate the success rate based on the execution result of the new task. If the success rate is less than the preset success rate threshold, the execution log of the new task is parsed, and the distribution data of error codes is obtained and reported back; otherwise, a new task is assigned based on the preliminary decision results.

[0019] This application simulates possible abnormal operating environments or data interference situations by injecting abnormal data into selected execution nodes and executing new tasks. It can detect the performance, processing capabilities, and potential defects of the server cryptographic machine in the face of anomalies in advance, thereby conducting a comprehensive detection of potential problems and making risk assessments in advance before the formal execution of tasks.

[0020] This application tests by injecting abnormal data in advance, which can identify and resolve problems that may affect the normal execution of the task in advance, and minimize serious consequences such as task failure and data leakage caused by encountering similar anomalies during actual operation.

[0021] When the calculated success rate of a new task is less than a preset success rate threshold, the execution log is parsed to obtain and report the distribution data of error codes. By analyzing the distribution of error codes, the link and cause of the problem can be quickly located. Based on the test results after injecting abnormal data, this application can evaluate and optimize the preliminary decision results. If it is found that some execution nodes perform poorly in handling abnormal data, important tasks can be avoided in subsequent task allocation, or these execution nodes can be further checked and optimized, thereby improving the overall success rate and efficiency of task execution.

[0022] Optionally, selecting m execution nodes includes: Set an initial selection probability for each server cipher machine; Obtain historical execution logs, parse and process the historical execution logs to obtain the parsing results, calculate the failure rate of each execution node based on the parsing results, reduce the probability of selecting execution nodes with failure rates greater than a preset failure threshold, and obtain the final selection probability; The top m execution nodes are selected based on the final selection probability, and the cipher machines of servers whose final selection probability is lower than the preset selection probability threshold are fed back to the operation and maintenance personnel.

[0023] This application adjusts the selection probability by setting an initial selection probability for each server cryptographic machine and calculating the failure rate based on historical execution logs. This allows execution nodes with good historical performance and low failure rates to receive a higher selection probability. When selecting the top m execution nodes based on the final selection probability, it is more likely to choose stable and reliable nodes to execute new tasks. For execution nodes with a failure rate greater than a preset failure threshold, their selection probability is reduced, effectively avoiding the frequent selection of these potentially problematic nodes that could easily lead to task failures. This further reduces the probability of errors due to unreliable nodes during the execution of new tasks, improving the stability and success rate of task execution.

[0024] This application adjusts the selection probability and chooses execution nodes based on their historical performance and failure rate, enabling a more reasonable allocation of task load. Tasks are distributed more to stable and reliable nodes, while fewer tasks are assigned to error-prone nodes. This ensures that each node can handle an appropriate workload within its capacity, improving the overall cluster resource utilization efficiency.

[0025] Optionally, if the master node fails during the execution of a new task by the execution node after injecting abnormal data, the execution status data of the new task is recorded, and after a new master node is elected, the execution status data of the new task is sent to the new master node.

[0026] Optionally, after sending the execution status data of the new task to the new master node, the method further includes: The new master node parses the received execution status data to determine whether the new task is in a recoverable execution state. If so, the new task is recovered based on the execution status data; otherwise, the new task is marked as a failed task, and a task failure notification is sent to the new task initiator.

[0027] This application records the execution status data of new tasks when the master node fails, and sends the data to the new master node after a new master node is elected. This enables the new master node to take over the work of the original master node, reduces the risk of task interruption or loss due to master node failure, and enhances the fault tolerance capability of the entire system for master node failure.

[0028] When a new master node determines that a new task is in a recoverable execution state, it can recover the new task based on the execution state data, avoiding the need to start the task from scratch as much as possible, thus saving computing resources, time, and network bandwidth.

[0029] If the new master node determines that a new task is unrecoverable, it will mark it as a failed task and send a task failure notification to the task initiator, allowing the task initiator to know the execution result of the task in a timely manner. Users can take other measures in a timely manner, such as re-initiating the task or choosing other transmission methods, thus improving the user experience.

[0030] Secondly, this application provides a management system for a server cryptographic machine cluster, which adopts the following technical solution: A management system for a server cryptographic machine cluster includes: a processor, and a memory communicatively connected to the processor; The memory is provided with a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the processor processes a computer program stored on the computer-readable storage medium, it implements the method as described in the first aspect.

[0031] In summary, this application includes at least one of the following beneficial technical effects: 1. This application establishes a management node group with multiple management nodes. Even if some management nodes fail, the remaining management nodes can still re-elect a new master node through a consensus algorithm, improving the reliability and continuity of the cluster management system. Slave nodes monitor the master node's status information in real time. If the master node's status does not meet expectations, a new master node can be quickly re-elected, reducing service interruption time caused by management node failures and further enhancing the cluster's high availability.

[0032] 2. When making task allocation decisions, the master node can leverage the shared status information of all nodes in the management node group. This allows the master node to make more scientific and reasonable preliminary decisions using a pre-defined distributed task allocation algorithm, better matching the actual capabilities of each node in the cluster. This minimizes situations where uneven task allocation leads to some nodes being overloaded while others are idle, thereby improving the overall resource utilization and task processing efficiency of the cluster. When a new task arrives, the master node directly makes the task allocation decision and executes it locally, minimizing intermediate steps and communication overhead in the task processing process. This enables tasks to be processed faster, improving the cluster's response speed and processing capacity. Attached Figure Description

[0033] Figure 1 This is a flowchart of Embodiment 1 of this application; Figure 2 This is a flowchart of Embodiment 2 of this application; Figure 3 This is a flowchart of Embodiment 3 of this application. Detailed Implementation

[0034] The following combination Figures 1 to 3 This application will be described in further detail.

[0035] Example 1: This example discloses a management method for a server cryptographic machine cluster, referring to... Figure 1 The method includes: S11 setting up management nodes, S12 status monitoring, and S13 task allocation. First, multiple management nodes are set up in the cluster to form a management node group. The status information of each management node is collected and shared with the other management nodes through a heartbeat protocol. Then, based on the status information, a consensus algorithm is used to elect a master node from the management node group, and the rest are slave nodes. The slave nodes monitor the status of the master node in real time, and a re-election is held in case of an anomaly. Afterward, when a new task arrives, the master node makes a local decision and executes it based on the status information using a preset distributed task allocation algorithm. The execution process of each step in this embodiment is as follows: S11 sets up management nodes, selects multiple server cryptographic machines in the server cryptographic machine cluster as management nodes to form a management node group, collects the status information of each management node in the management node group, and each management node will send a heartbeat message to other management nodes at a certain time interval (such as every 5 seconds). The heartbeat message contains the status information of the management node. After receiving the heartbeat message, other management nodes will update the status information of the management node stored locally.

[0036] The status information includes hardware status, software status, and network connectivity status. Hardware status includes CPU utilization, memory utilization, disk space utilization, and network bandwidth utilization. Software status includes the running status of the management service program, whether processes are functioning normally, and whether services are running normally. Network connectivity status refers to whether the current management node's network connections with other management nodes and other computing nodes in the cluster are normal. Network connectivity can be detected by sending network probe packets (such as the ping command).

[0037] In other embodiments, the status information can also be set as needed.

[0038] S12 Status Monitoring: Based on the status information, a consensus algorithm (such as Paxos or Raft) is used to elect a master node in the management node group, and the remaining management nodes in the management node group are used as slave nodes. The slave nodes receive the status information sent by the master node in real time through the heartbeat protocol and compare it with the historical status information of the master node stored locally to determine whether the status information sent by the master node in real time meets the expectations. If the status information sent by the master node in real time does not meet the expectations, this step is re-executed.

[0039] Taking the Raft algorithm as an example, the election of the master node includes steps 1 to 4, specifically the following: Step 1: At the start of each election cycle, each node in the management node group is in a follower state. If a follower does not receive a heartbeat message from the leader within a certain period of time, it will become a candidate node and initiate an election.

[0040] Step 2: The candidate node sends a voting request to other management nodes, requesting them to vote for it. The voting request includes the candidate node's term number and its own identification information.

[0041] Step 3: After receiving the voting request, other management nodes will make a voting decision according to certain rules. If the term number of a candidate node is greater than the term number currently recorded by itself, and the candidate node has not voted in previous terms, then the other management nodes will vote for the candidate node.

[0042] Step 4: If a candidate node receives more than half of the votes from the management nodes, it will become the new master node. If no candidate node receives more than half of the votes during an election cycle, the election cycle ends.

[0043] In this embodiment, the real-time status information sent by the master node not conforming to expectations refers to any of the following situations: Scenario 1: The hardware components of the master node are damaged or abnormal, such as the CPU overheating and causing it to run at a reduced frequency, the memory module malfunction causing data read and write errors, the hard drive having bad sectors causing problems with data storage and retrieval, and the unstable power supply causing the node to restart frequently.

[0044] Scenario 2: The master node's hardware resources are overused, reaching or nearing their limits. For example, CPU utilization remains above 90% for 10 consecutive seconds, leaving the master node with insufficient computing power to handle new management requests; memory usage is too high (greater than 80%), causing frequent memory swapping, severely impacting performance; disk space utilization exceeds 95%, making it impossible to store new log information or temporary data, etc.

[0045] Scenario 3: The management service program running on the master node crashes or terminates abnormally, or the management service program on the master node gets stuck in a deadlock or infinite loop, causing the program to be unable to continue responding to external requests, or the version of the management software installed on the master node is incompatible with the versions of other management nodes or software that the cluster depends on.

[0046] Deadlock occurs when multiple threads or processes are competing for resources and waiting for each other, preventing all related threads or processes from continuing execution. An infinite loop occurs when an improperly set loop condition in the code causes the program to remain in a loop indefinitely.

[0047] Scenario 4: The network connection between the master node and other management nodes in the cluster is interrupted, or the network latency between the master node and other nodes is greater than 50ms, resulting in excessively long data transmission and communication times, or the network bandwidth between the master node and other nodes is heavily occupied, resulting in slow data transmission speed.

[0048] Scenario 5: No heartbeat packets are received from the master node for 5 consecutive times, or the slave node sends a status query request to the master node, but the master node does not respond.

[0049] In other embodiments, rules can be set to prevent the master node status information from deviating from expectations, as needed.

[0050] S13 Assigns Tasks: When a new task arrives, the master node makes a task assignment decision locally based on the status information using a preset distributed task assignment algorithm, obtains a preliminary decision result, and assigns the new task based on the preliminary decision result.

[0051] In this embodiment, the preset distributed task allocation algorithm adopts a load balancing algorithm.

[0052] The process by which the master node makes task allocation decisions locally based on status information using a load balancing algorithm, and obtains preliminary decision results, is as follows: The master node calculates the remaining resource rate of each server cryptographic machine in the server cryptographic machine cluster based on the status information. The calculation model is as follows:

[0053]

[0054] Where R is the remaining resource rate; n is the type of data contained in the status information. In this embodiment, the value of n is 8. The weight of the i-th state information; The calculation model for the i-th state information is as follows in this embodiment:

[0055] in, CPU utilization; Memory usage; Disk space usage; Network bandwidth utilization; This indicates that the management service program is running normally; This indicates that the process is normal; This indicates that the service is running normally; This indicates that the network connection is normal, meaning the network latency is no more than 50ms. For about The truth function, when When true, The value of is 1, when When it is false, The value of is 0.

[0056] The master node assigns new tasks to the cipher machines on the server with the highest remaining resource rate. If a new task requires multiple cipher machines to work together, the subtasks are assigned in descending order of remaining resource rate to balance the load across the cipher machines as much as possible.

[0057] In other embodiments, a priority scheduling algorithm is used to make task allocation decisions locally, obtaining preliminary decision results. Specifically, the master node first sorts tasks according to their priorities, prioritizing the allocation of high-priority tasks. For tasks of the same priority, load balancing is performed based on the remaining resource rates of each server's cryptographic machines. If a low-priority task is being executed and a high-priority task arrives, the master node can schedule slave nodes to pause the low-priority task and prioritize processing the high-priority task. The low-priority task is then resumed after the high-priority task is completed. In this embodiment, the remaining resources of each server's cryptographic machine can be calculated using the above calculation model.

[0058] Example 2: Refer to Figure 2 The difference between this embodiment and Embodiment 1 is that, after obtaining the preliminary decision result, the method further includes: S21 verification involves sharing the preliminary decision results with at least one slave node through a distributed coordination service (i.e., etcd coordination service), and the slave node that receives the preliminary decision results verifies the preliminary decision results.

[0059] Extract the server cryptographic machine that will execute the new task from the preliminary decision results and denote it as the execution node.

[0060] In this embodiment, the verification rule is as follows: retrieve the cluster node status table cached locally from the node, compare the resource quota allocated in the preliminary decision result with the real-time remaining resources of the execution node, and the resource quota is less than or equal to 80% of the real-time remaining resources of the execution node.

[0061] In other embodiments, the verification rules can also be set according to requirements, such as verifying whether the task priority matches the permission level of the target node, and that the task priority is not higher than the permission level of the execution node. Another example is retrieving the historical task logs of the execution node to check for records of concurrent execution failures of the same type of task, and ensuring that the concurrent execution failure rate of the same type of task is less than a certain value.

[0062] If more than half of the slave nodes report successful verification, then S22 simulated allocation is executed.

[0063] Conversely, if the initial decision is not received, the master node and the slave nodes that received the initial decision result will re-determine the task allocation through a consensus algorithm to obtain the final decision result. The process is as follows: The slave nodes upload the reasons for the failed verification to the etcd coordination service, and the master node summarizes all dissenting information. The master node and the slave nodes that participated in the verification form a temporary consensus group. An improved Raft algorithm is used to divide the temporary consensus group into three core roles: proposer (played by the master node), voter (played by the slave nodes that participated in the verification), and observer (optional configuration, played by slave nodes that did not participate in the verification).

[0064] Based on the reasons for the verification failure, the proposer generates a new task allocation proposal. The proposal must include the task ID, the adjusted list of execution nodes, resource quota parameters, and a correction explanation for the verification failure. The proposal content and basis are broadcast to all voters through the distributed coordination service (etcd coordination service), the voting results are collected, and the proposal is determined to pass. If the proposal passes, it is used as the final decision result, and the final decision result is synchronized to the local cache.

[0065] Voters receive the proposal from the proposer, verify the proposal's rationality (including verifying the proposal's legality (such as digital signature, task ID uniqueness) and validity (verifying whether the correction description corresponds to the previous failure reason) based on the status information of the local node and the execution node, and provide feedback to the proposer to vote in favor or against, along with the reasons for the vote.

[0066] The observer listens to the proposal and voting process of the consensus group, does not participate in voting, and synchronizes the consensus process log in real time for cluster auditing and fault backtracking.

[0067] Update the final decision result to the preliminary decision result and execute the S22 simulation allocation.

[0068] S22 simulates the allocation, setting an initial selection probability for each server cipher machine. In this embodiment, all server cipher machines have the same initial selection probability, for example, all initial selection probabilities are set to 0.5. In other embodiments, the value of the initial selection probability and the method of setting the initial selection probability can be set according to requirements, such as setting the initial selection probability based on the hardware performance of the server cipher machine (such as processing speed and concurrent processing capability).

[0069] Collect historical execution logs from each execution node for the past 30, 60, or 90 days, filter out historical execution logs that are the same type as the new task, parse the filtered historical execution logs to obtain the parsing results, and calculate the failure rate of each execution node based on the parsing results. Failure rate = number of failed historical execution logs ÷ number of filtered historical execution logs × 100%.

[0070] The final selection probability is obtained by reducing the selection probability of execution nodes with a failure rate greater than a preset failure threshold (e.g., 80%). That is, the initial selection probability of execution nodes with a failure rate not greater than the preset failure threshold remains unchanged, and their final selection probability is equal to the initial selection probability; the selection probability of execution nodes with a failure rate greater than the preset failure threshold is reduced, and their final selection probability is less than the initial selection probability.

[0071] In this embodiment, the formula for calculating the probability of selecting execution nodes with a failure rate greater than a preset failure threshold (e.g., 80%) is as follows:

[0072] in, This is the reduced selection probability, i.e., the final selection probability; This represents the failure rate of the current execution node. The default failure threshold is set to 80% in this embodiment.

[0073] The execution nodes are sorted in descending order of their final selection probability. The first m execution nodes in the sequence are selected, and the server cryptographic machines whose final selection probability is lower than the preset selection probability threshold (10%) are fed back to the operation and maintenance personnel.

[0074] In other embodiments, m execution nodes may be randomly selected.

[0075] S23 simulates the execution of a task and injects abnormal data into the selected execution node. The abnormal data includes: abnormal resource usage (simulating the CPU usage of the execution node to suddenly rise to 90% and the memory usage to suddenly rise to 85%), abnormal key data (injecting keys with incorrect format, expired keys, and mismatched keys), abnormal network transmission (simulating data packet loss (packet loss rate greater than 10%), delay (delay time greater than 50ms)), and abnormal input parameters (injecting parameters that are out of range, empty parameters, and incorrectly formatted parameters).

[0076] This embodiment uses a hybrid injection method to inject abnormal data into the execution node.

[0077] The master node distributes the new task to the selected m execution nodes, and simultaneously distributes the injected abnormal data. During the execution of the task, the execution nodes collect the following data in real time: Execution result: Success, Failure, Timeout.

[0078] Execution log: contains timestamps, status codes, and error messages (if failure) for each step of the operation.

[0079] Resource consumption: CPU usage, memory usage, and key operation time during task execution.

[0080] The success rate is calculated based on the execution results of the new task. Success rate = number of execution nodes with successful execution results ÷ m.

[0081] S24 Success Rate Judgment: Determines whether the success rate is less than a preset success rate threshold (e.g., 80%). If so, the master node aggregates the execution logs of all simulated failed execution nodes, extracts error codes (such as key mismatch 001, insufficient resources 002, network timeout 003), counts the occurrence frequency and percentage of each error code, and generates error code distribution data.

[0082] The error code distribution data and the list of failed nodes are fed back to the operations and maintenance personnel. The operations and maintenance personnel can locate the cause of the failure based on the error code type (such as error code 001 corresponding to incorrect key configuration, and error code 002 corresponding to insufficient resources) and perform targeted optimizations (such as reconfiguring the key or upgrading the node hardware).

[0083] If not, new tasks will be assigned based on the preliminary decision results.

[0084] In other embodiments, after performing the simulated task in S23 and before performing the success rate determination in S24, the method further includes: If the master node fails during the execution of a new task by the execution node after injecting abnormal data, i.e. the master node's status information does not meet expectations, the execution status data of the new task is recorded. The management node group immediately triggers the master-slave node election process in S12 status monitoring, and after a new master node is elected, the execution status data of the new task is sent to the new master node.

[0085] In this embodiment, the execution status data includes: basic task information (task ID, task type, priority, resource quota (CPU usage / memory usage)), execution progress information (current execution stage of the task, completed steps), abnormal data injection information (injected abnormal data type (resource usage or key data or network transmission or input parameters), injection time), and execution node status information.

[0086] The new master node broadcasts an execution status data reporting instruction to all execution nodes participating in the simulation. The execution status data reporting instruction includes its own node ID, communication port, and data reception timeout (e.g., 2 seconds).

[0087] The new master node parses the received execution status data and determines whether the new task is in a recoverable execution state based on three aspects: task execution progress, the current status information of the execution node, and the degree of impact of abnormal data. If all three requirements are met, the new task is determined to be in a recoverable execution state.

[0088] The new task being in a recoverable execution state includes the following three rules: The new task has completed more than 30% of its steps, and the currently executing step is a step that can be resumed from a breakpoint (e.g., the algorithm calculation stage can be resumed from a breakpoint, but the initialization stage cannot).

[0089] The current status information of the execution node is in line with expectations, and its judgment rules are the same as those of the master node.

[0090] The injected abnormal data did not cause irreversible damage to the core process of the task. For example, abnormal resource usage and abnormal network transmission are reversible effects (execution can be resumed after troubleshooting); abnormal key data and incorrect input parameter format are irreversible effects (which will invalidate the task results).

[0091] If so, then a new task will be resumed based on the execution status data; If not, mark the new task as a failed task and send a task failure notification to the task initiator.

[0092] Example 3: Reference Figure 3 The difference between this embodiment and Embodiment 1 is that the method further includes: S31 divides the region and obtains the physical location (i.e., the x and y coordinates, which can be obtained by converting the latitude and longitude of each server's cryptographic machine into the x and y coordinates of a Cartesian coordinate system), network topology (i.e., the one-way network latency between the current server's cryptographic machine and the management node), and load (calculated by weighted summation of CPU utilization (weight 0.4), memory utilization (weight 0.3), and key operation load (weight 0.3)).

[0093] The K-means algorithm is used to divide logical regions, with each logical region containing at least one management node. The process is as follows: Management nodes are selected as the initial cluster centers, so that each initial cluster center corresponds to one management node; Calculate the weighted Euclidean distance from each server's cipher machine to each cluster center (distance weights are consistent with feature weights: physical location 0.3, network topology 0.4, load 0.3), using the following formula:

[0094] in, Let be the weighted Euclidean distance between the i-th server cipher machine and the j-th cluster center; Let x be the x-coordinate of the physical location of the i-th server's cryptographic machine; The x-coordinate of the physical location of the j-th cluster center; Let y be the ordinate of the physical location of the i-th server's cryptographic machine; The ordinate of the physical location of the j-th cluster center; Let be the network latency value between the i-th server cryptographic machine and the management node; Let j be the average network latency value for the j-th cluster center; The load for the j-th cluster center; Let be the load of the j-th cluster center.

[0095] The nodes are assigned to the logical regions with the smallest Euclidean distance. Then, the feature mean of each logical region is calculated iteratively, the cluster center is updated, and the iteration is repeated until the cluster center is stable.

[0096] After clustering is completed, it is checked whether each logical region contains at least one management node. If there is a logical region without a management node, the logical region is merged into its adjacent logical region. The merging rule is: the Euclidean distance between the cluster center of the logical region and the cluster center of the merged adjacent logical region is minimized.

[0097] S32 assigns tasks based on proximity, obtains the server cipher machine that is closest to the physical location of the new task initiator, and records it as the target cipher machine. Based on the ownership identifier of the target cipher machine, determines the logical area where it is located as the target area, and prioritizes attempting to complete the task assignment within the target area.

[0098] If the load on the server cryptographic machines within the target area does not reach the preset load threshold (e.g., 0.7), the new task will be assigned to the server cryptographic machines within the target area. The assignment method may be: New tasks are assigned to the cipher machines on the server with the highest remaining resource rate. If a new task requires multiple cipher machines to work together, the task is split into multiple subtasks and distributed to multiple low-load (less than 0.2) adjacent logical regions according to the load ratio of adjacent logical regions. Within each adjacent logical region, subtasks are distributed in descending order of remaining resource rate. This balances the load of multiple cipher machines on the server while minimizing the occurrence of sudden load increases in a single logical region.

[0099] If the load of the server cryptographic machines in the target area has reached the preset load threshold (e.g., 0.7), then further analyze whether the load of the server cryptographic machines in the adjacent logical areas of the target area has reached the preset load threshold (e.g., 0.7). If the load of the server cryptographic machine in at least one adjacent logical region of the target region does not reach the preset load threshold (e.g., 0.7), then new tasks will be assigned to the server cryptographic machines in the adjacent logical regions of the target region in order of load from low to high.

[0100] If the server cryptographic machine load in the adjacent logical regions of the target area all reaches the preset load threshold (e.g., 0.7), then an overload message will be sent to the new task initiator, asking you to try again later.

[0101] In the process of sequentially assigning new tasks to server cryptographic machines in adjacent logical regions of the target region, the method further includes: Real-time monitoring of network topology changes of server cryptographic machines in adjacent logical areas. If a task allocation path anomaly occurs due to a network topology change, the task allocation path anomaly includes at least one abnormal event, which includes: Abnormal event 1: A sudden increase in network latency between adjacent logical regions (e.g., from 20ms to 100ms). Abnormal event 2: The link packet loss rate is greater than 5% and the duration is greater than 5 seconds; Abnormal event 3: Inter-node communication link interruption.

[0102] The master node suspends task allocation on the abnormal path and, based on the current physical location of each logical region and the status information of each server's cryptographic machine, makes a task reallocation decision using a preset distributed task allocation algorithm, obtaining the reallocation decision result. The process is as follows: Based on the physical location of the logical region on the non-abnormal path and the status information of each server's cryptographic machine, the allocation probability is calculated. The calculation model is as follows:

[0103] in, The probability of assigning a value to the j-th logical region; M is the number of server cryptographic machines in the j-th logical region; The remaining resources of the cipher machine of the ith server in the j-th logical region; Let be the Euclidean distance between the cluster center of the j-th logical region and the cluster center of the k-th logical region; It is the minimum Euclidean distance from the cluster center of the j-th logical region to the logical regions on the non-abnormal path; It represents the maximum Euclidean distance from the cluster center of the j-th logical region to the logical regions on the non-abnormal path; This represents the average remaining resources of the j-th logical region.

[0104] In other embodiments, when a server cryptographic machine malfunctions, the method further includes: The logical area where the faulty server cryptographic machine is located is designated as the fault area. The faulty server cryptographic machine sends the fault information (including unexpected status information, fault occurrence time, and currently executing task ID) to the management node in the fault area via a heartbeat protocol. The management node in the fault area uses a preset distributed task allocation algorithm (load balancing algorithm) to redistribute the tasks assigned to the faulty server cryptographic machine to the remaining server cryptographic machines in the fault area. The allocation method is the same as the task allocation in S13.

[0105] Example 4: This example discloses a management system for a server cryptographic machine cluster. The system includes a processor and a memory communicatively connected to the processor. The memory is provided with a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the processor processes a computer program stored on the computer-readable storage medium, it implements the method.

[0106] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A management method for a server cryptographic machine cluster, characterized in that, include: Multiple management nodes are set up in the cluster to form a management node group. The status information of each management node in the management node group is collected and shared with the remaining management nodes through a heartbeat protocol. Based on the status information, a consensus algorithm is used to elect a master node in the management node group, and the remaining management nodes in the management node group are used as slave nodes. The slave nodes monitor the status information of the master node in real time, and re-execute this step when the status information of the master node does not meet expectations. When a new task arrives, the master node makes a task allocation decision locally based on the status information using a preset distributed task allocation algorithm, and obtains a preliminary decision result. The preliminary decision results are shared to at least one slave node through a distributed coordination service. The slave node that receives the preliminary decision results verifies the preliminary decision results. If the verification is successful, a new task is assigned according to the preliminary decision results. If the verification fails, the master node and the slave node that received the preliminary decision result will re-determine the task allocation through the consensus algorithm to obtain the final decision result, and update the preliminary decision result with the final decision result. The master node and the slave nodes that received the preliminary decision results re-determine the task allocation through a consensus algorithm. This process includes: the master node and the participating slave nodes forming a temporary consensus group, within which three core roles are defined: the master node acts as the proposer, the participating slave nodes act as voters, and the non-participating slave nodes act as observers. The proposer generates a new task allocation proposal based on the reasons for the verification failure, broadcasts the proposal content and basis to all voters through a distributed coordination service, collects voting results, and determines whether the proposal passes. If the proposal passes, it is used as the final decision result. Voters receive the proposal content from the proposer, verify the rationality of the proposal based on the status information of their local nodes and the execution node, and provide feedback to the proposer indicating whether they agree or disagree. Observers listen to the proposal and voting process of the consensus group, do not participate in voting, and synchronize the consensus process log in real time. New tasks were assigned based on the preliminary decision-making results.

2. The management method for a server cryptographic machine cluster according to claim 1, characterized in that, The method further includes: Obtain the physical location, network topology, and load of each server's cryptographic machine, and use the K-means algorithm to divide the logical regions, with each logical region containing at least one management node; Obtain the server cryptographic machine that is closest to the physical location of the new task initiator, and denote it as the target cryptographic machine. Denote the logical area where the target cryptographic machine is located as the target area. If the load of the server cryptographic machine in the target area does not reach the preset load threshold, then the new task is assigned to the server cryptographic machine in the target area. If the load of all server cryptographic machines in the target area has reached the preset load threshold and the load of server cryptographic machines in at least one adjacent logical area of ​​the target area has not reached the preset load threshold, then new tasks will be assigned to server cryptographic machines in the adjacent logical areas of the target area in order of load from low to high.

3. The management method for a server cryptographic machine cluster according to claim 2, characterized in that, In the process of sequentially assigning new tasks to server cryptographic machines in adjacent logical regions of the target region, the method further includes: The network topology changes of the server cryptographic machines in each adjacent logical area are monitored in real time. If the task allocation path is abnormal due to the network topology change, the master node will re-make the task allocation decision based on the current physical location of each logical area and the status information of each server cryptographic machine, and obtain the re-allocation decision result.

4. The management method for a server cryptographic machine cluster according to claim 3, characterized in that, When a server cryptographic machine malfunctions, the method further includes: The logical area where the faulty server cryptographic machine is located is designated as the fault area. The faulty server cryptographic machine sends the fault information to the management node in the fault area via a heartbeat protocol. The management node in the fault area uses a preset distributed task allocation algorithm to reassign the tasks assigned to the faulty server cryptographic machine to the remaining server cryptographic machines in the fault area.

5. The management method for a server cryptographic machine cluster according to claim 1, characterized in that, After verification but before assigning new tasks based on preliminary decision results, the method further includes: Extract the server cryptographic machine that will execute the new task from the preliminary decision results and denote it as the execution node. Select m execution nodes, inject abnormal data into the selected execution nodes, use the execution nodes after injecting abnormal data to execute the new task, and obtain the execution result and execution log of the new task. Calculate the success rate based on the execution result of the new task. If the success rate is less than the preset success rate threshold, the execution log of the new task is parsed, and the distribution data of error codes is obtained and reported back; otherwise, a new task is assigned based on the preliminary decision results.

6. The management method for a server cryptographic machine cluster according to claim 5, characterized in that, The selection of m execution nodes includes: Set an initial selection probability for each server cipher machine; Obtain historical execution logs, parse and process the historical execution logs to obtain the parsing results, calculate the failure rate of each execution node based on the parsing results, reduce the probability of selecting execution nodes with failure rates greater than a preset failure threshold, and obtain the final selection probability; The top m execution nodes are selected based on the final selection probability, and the cipher machines of servers whose final selection probability is lower than the preset selection probability threshold are fed back to the operation and maintenance personnel.

7. The management method for a server cryptographic machine cluster according to claim 5 or 6, characterized in that, If the master node fails during the execution of a new task using the injected abnormal data, the execution status data of the new task is recorded, and after a new master node is elected, the execution status data of the new task is sent to the new master node.

8. The management method for a server cryptographic machine cluster according to claim 7, characterized in that, After sending the execution status data of the new task to the new master node, the method further includes: The new master node parses the received execution status data to determine whether the new task is in a recoverable execution state. If so, the new task is recovered based on the execution status data; otherwise, the new task is marked as a failed task, and a task failure notification is sent to the new task initiator.

9. A management system for a server cryptographic machine cluster, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory is provided with a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the processor processes a computer program stored on the computer-readable storage medium, it implements the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Cipher equipment cluster management method and device, equipment and storage medium

    CN118944883A