Cryptographic analysis task breakpoint continuation type fault-tolerant scheduling method

By constructing a breakpoint-resumption model and a fault-tolerant scheduling rule base, precise breakpoint management and fault recovery of cryptanalysis tasks in a distributed computing cluster were achieved, solving the task interruption problem and improving the continuity of task execution and resource utilization.

CN121907928APending Publication Date: 2026-04-21HUNAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUNAN UNIV
Filing Date
2025-12-29
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing cryptanalysis task scheduling methods lack precise breakpoint management mechanisms in distributed computing clusters, resulting in the need to restart tasks after interruption, which wastes computing resources and prolongs task completion cycles. Furthermore, traditional fault tolerance strategies have limited coverage and non-standard recovery processes.

Method used

By collecting task attribute information and node running status data, a breakpoint resume model and fault-tolerant scheduling rule base are constructed. An initial scheduling scheme is generated and a fault simulation test is conducted. The scheme is then optimized and adjusted to form the final fault-tolerant scheduling scheme, achieving accurate breakpoint location and resume status recovery.

Benefits of technology

It improves the continuity of cryptanalysis task execution, reduces time and resource consumption, enhances resource utilization and fault recovery efficiency, adapts to large-scale distributed resource scenarios, and fully leverages the parallel computing power advantages of exascale supercomputing clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121907928A_ABST
    Figure CN121907928A_ABST
Patent Text Reader

Abstract

The invention discloses a cryptanalysis task breakpoint continuation type fault-tolerant scheduling method, which relates to the technical field of distributed computing scheduling, and is characterized by comprising the following steps of: acquiring task attribute information, historical scheduling data and distributed node running state data of a cryptanalysis task; performing feature recognition on the task attribute information to obtain a task feature result, and matching the task feature result with tasks and nodes in the node operation state data to determine fault-tolerant scheduling core parameters; the method has the advantages that accurate breakpoint positioning and continuous calculation state recovery are achieved in combination with the breakpoint continuous calculation model, repeated calculation after task interruption is avoided, time and resource consumption is greatly reduced, the execution continuity of a cryptographic analysis task is improved, the method is convenient to adapt to a large-scale distributed resource scene of an E-level super-calculation cluster, and the method is suitable for large-scale distributed resource scenes of the E-level super-calculation cluster. The parallel computing power advantage of the E-level super computing cluster can be fully exerted, and the execution continuity of the cryptographic analysis task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed computing scheduling technology, and more specifically, to a fault-tolerant scheduling method for resuming interrupted computations in cryptanalysis tasks. Background Technology

[0002] With the rapid development of cryptanalysis technology, the complexity of encryption algorithms and the scale of ciphertext data to be processed continue to rise. A single computing node can no longer support the efficient execution of large-scale cryptanalysis tasks, and distributed computing clusters are gradually becoming the core environment for cryptanalysis tasks. However, the scheduling of cryptanalysis tasks under a distributed architecture faces multiple technical bottlenecks.

[0003] However, existing scheduling schemes for cryptanalysis tasks are characterized by long execution cycles and intensive resource consumption. For example, cracking hash algorithms often requires hours or even days of parallel computation. During this process, frequent failures occur, such as the downtime of distributed nodes, network transmission interruptions, and exhaustion of computing resources. Existing scheduling methods lack precise breakpoint management mechanisms. Most schemes simply record the approximate progress of the task without integrating key information such as execution steps and data verification values. After a task is interrupted, it usually needs to be restarted, which not only causes a serious waste of computing resources but also significantly extends the task completion cycle. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the present invention aims to provide a fault-tolerant scheduling method for resuming interrupted computations in cryptanalysis tasks.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A fault-tolerant scheduling method for resuming interrupted computations in cryptanalysis tasks, comprising the following steps:

[0007] Collect task attribute information, historical scheduling data, and distributed node running status data for cryptanalysis tasks;

[0008] Task feature results are obtained by performing feature recognition on task attribute information. The task feature results are then matched with the tasks and nodes in the node running status data to determine the core parameters of fault-tolerant scheduling.

[0009] Based on breakpoint records, fault recovery cases, and core parameters of fault-tolerant scheduling in historical scheduling data, a breakpoint continuation model and a fault-tolerant scheduling rule base are constructed.

[0010] Based on the real-time status data and node load of the current cryptanalysis task, the breakpoint continuation model and fault-tolerant scheduling rule base are invoked to generate an initial scheduling scheme.

[0011] The scheduling quality assessment results are obtained by conducting fault simulation tests and scheduling efficiency evaluations on the initial scheduling scheme.

[0012] Based on the scheduling quality assessment results, the initial scheduling scheme is optimized and adjusted to form the final fault-tolerant scheduling scheme. The final fault-tolerant scheduling scheme is executed, and the breakpoint information set and scheduling log are collected in real time to generate a breakpoint-resumption fault-tolerant scheduling for the cryptographic analysis task.

[0013] Preferably, the task attribute information includes the complexity level, data volume, computational accuracy requirements, and task priority of the cryptanalysis task, and the distributed node operating status data includes node computing power resources, storage resources, network bandwidth, and historical fault records.

[0014] Preferably, the task attribute information is used to perform feature recognition to obtain task feature results, and the task feature results are matched with the tasks and nodes in the node running status data to determine the core parameters of fault-tolerant scheduling. Specifically, this includes the following steps:

[0015] Key features in the task attribute information are extracted and classified to obtain task complexity features, data processing features and priority features. The task complexity features, data processing features and priority features are then summarized to form the task feature results.

[0016] Statistically analyze the resource availability, failure rate, and failure recovery speed of each distributed node, and establish a node performance evaluation index system.

[0017] Based on the task feature results and node performance evaluation index system, the fitness scores of the task and each node are calculated to form a fitness dataset.

[0018] Based on the fit dataset and task priority, key parameters affecting scheduling stability and fault tolerance efficiency are selected, and core parameters for fault-tolerant scheduling are determined. These core parameters include breakpoint recording frequency, fault detection threshold, node switching threshold, and task shard size.

[0019] Preferably, based on breakpoint records, fault recovery cases, and core parameters of fault-tolerant scheduling in historical scheduling data, a breakpoint continuation model and a fault-tolerant scheduling rule base are constructed, specifically including the following steps:

[0020] Extract breakpoint location information, breakpoint time task progress data, and fault type records of cryptanalysis tasks from historical scheduling data to establish a breakpoint information dataset;

[0021] Based on the fault handling process, node replacement scheme, and task continuation scheme in historical fault recovery cases, generate fault-tolerant scheduling experience rules.

[0022] Using core parameters as constraints, and combining breakpoint information datasets, a task progress backtracking model is trained to construct a breakpoint continuation model.

[0023] The fault-tolerant scheduling rule library is obtained by inputting the fault-tolerant scheduling empirical rules into the breakpoint continuation calculation model.

[0024] Preferably, based on the real-time status data and node load of the current cryptanalysis task, an initial scheduling scheme is generated by calling the breakpoint continuation model and the fault-tolerant scheduling rule base, specifically including the following steps:

[0025] Collect real-time status data on the current cryptanalysis task's execution progress, incomplete computation modules, and current resource usage;

[0026] Real-time load data of nodes is obtained by monitoring the real-time load rate, resource idleness and network connectivity status of each distributed node.

[0027] The real-time status data and node real-time load data are matched with the fault-tolerant scheduling rule base to obtain the scheduling scheme and node allocation scheme;

[0028] The breakpoint resume calculation model is invoked to verify the feasibility of breakpoint resume calculation of the scheduling scheme and node allocation scheme. The resource allocation ratio is adjusted in combination with task priority to generate an initial scheduling scheme.

[0029] Preferably, the scheduling quality assessment result is obtained by conducting fault simulation tests and scheduling efficiency evaluations on the initial scheduling scheme, specifically including the following steps:

[0030] Simulate typical failure scenarios such as distributed node offline, network interruption, and exhaustion of computing resources to test the fault detection response speed and breakpoint recording accuracy of the initial scheduling scheme;

[0031] The task allocation balance, node resource utilization, and task execution efficiency of the initial scheduling scheme are statistically analyzed, and a scheduling efficiency score is calculated.

[0032] The fault tolerance performance evaluation value is obtained by judging the fault recovery success rate, continuation data consistency and recovery time of the initial scheduling scheme in the fault simulation test;

[0033] The scheduling quality assessment result is determined by combining the scheduling efficiency score and the fault tolerance performance evaluation value.

[0034] Preferably, the initial scheduling scheme is optimized and adjusted based on the scheduling quality assessment results to form the final fault-tolerant scheduling scheme, specifically including the following steps:

[0035] Based on the scheduling quality assessment results, identify unreasonable node allocation, inappropriate breakpoint recording frequency, and insufficient adaptability of fault tolerance scheme in the initial scheduling scheme, and determine the optimization impact weights.

[0036] Based on task priority and optimization impact weight, determine optimization priority, and use optimization priority to identify the core issues affecting fault tolerance stability;

[0037] Based on the core issues affecting fault tolerance stability, the node allocation scheme was adjusted, the breakpoint recording frequency and fault recovery strategy were optimized, and the feasibility and effectiveness of the scheme were re-verified to form the final fault-tolerant scheduling scheme.

[0038] Preferably, the execution of the final fault-tolerant scheduling scheme and the real-time collection of breakpoint information sets and scheduling logs to generate a breakpoint-resumption fault-tolerant scheduling for the cryptanalysis task specifically includes the following steps:

[0039] The final fault-tolerant scheduling scheme is broken down into several sub-tasks, which are then assigned to the corresponding distributed nodes and started for execution.

[0040] According to the breakpoint recording frequency set in the core parameters, the execution progress, data processing status and node running parameters of each subtask are collected in real time to generate a breakpoint information set, which includes the current execution step, data verification value and resource usage snapshot;

[0041] The task scheduling process collects node allocation changes, fault occurrence time, fault handling measures, and restart time to form a complete scheduling log.

[0042] When a fault is detected, the task execution state is restored through the breakpoint information set and scheduling logs, and the node switching or task reassignment process is initiated according to the fault-tolerant scheduling rule base to realize the breakpoint resume computing.

[0043] Compared with existing technologies, this invention has the following advantages: By collecting a set of breakpoint information including execution steps, data verification values, and resource usage snapshots, and combining it with a breakpoint continuation model, it achieves precise breakpoint location and continuation state recovery, avoiding repeated calculations after task interruption, significantly reducing time and resource consumption, and improving the execution continuity of cryptanalysis tasks; By extracting task complexity and data volume characteristics and combining them with a node performance evaluation index system to calculate adaptability, it achieves precise matching between tasks and nodes, avoiding resource overflow failures and improving the resource utilization and task computation efficiency of the distributed cluster; Based on historical failure cases, a fault-tolerant scheduling rule base is constructed, and combined with failure simulation test optimization strategies, it can adapt to failure scenarios such as node crashes and network interruptions; At the same time, by dynamically adjusting core parameters such as node switching thresholds and task shard size, it shortens the failure recovery time, ensures the consistency of continuation data, and improves the adaptability and failure recovery efficiency of fault-tolerant strategies. Addressing the problems of traditional fault-tolerant strategies having limited coverage scenarios and non-standard recovery processes, it is easy to adapt to the large-scale distributed resource scenarios of E-level supercomputing clusters, and can fully leverage the parallel computing power advantages of E-level supercomputing clusters to improve the execution continuity of cryptanalysis tasks. Attached Figure Description

[0044] Figure 1This invention presents a schematic diagram illustrating the steps of a fault-tolerant scheduling method for resuming interrupted computations in cryptanalysis tasks.

[0045] Figure 2 This is a schematic diagram illustrating the steps of generating an initial scheduling scheme in a fault-tolerant scheduling method for resuming breakpoint calculations in cryptanalysis tasks proposed in this invention.

[0046] Figure 3 This is a schematic diagram illustrating the steps in forming the final fault-tolerant scheduling scheme in the fault-tolerant scheduling method for breakpoint continuation of cryptographic tasks proposed in this invention. Detailed Implementation

[0047] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0048] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0049] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0050] Reference Figures 1-3 As shown.

[0051] This embodiment further illustrates the fault-tolerant scheduling method for resuming interrupted computations in cryptanalysis tasks proposed in this invention.

[0052] A fault-tolerant scheduling method for resuming interrupted computations in cryptanalysis tasks, comprising the following steps:

[0053] Collect task attribute information, historical scheduling data, and distributed node running status data for cryptanalysis tasks;

[0054] Task feature results are obtained by performing feature recognition on task attribute information. The task feature results are then matched with the tasks and nodes in the node running status data to determine the core parameters of fault-tolerant scheduling.

[0055] Based on breakpoint records, fault recovery cases, and core parameters of fault-tolerant scheduling in historical scheduling data, a breakpoint continuation model and a fault-tolerant scheduling rule base are constructed.

[0056] Based on the real-time status data and node load of the current cryptanalysis task, the breakpoint continuation model and fault-tolerant scheduling rule base are invoked to generate an initial scheduling scheme.

[0057] The scheduling quality assessment results are obtained by conducting fault simulation tests and scheduling efficiency evaluations on the initial scheduling scheme.

[0058] Based on the scheduling quality assessment results, the initial scheduling scheme is optimized and adjusted to form the final fault-tolerant scheduling scheme. The final fault-tolerant scheduling scheme is executed, and the breakpoint information set and scheduling log are collected in real time to generate a breakpoint-resumption fault-tolerant scheduling for the cryptographic analysis task.

[0059] The task attribute information includes the complexity level of the cryptanalysis task, the data volume, the required computational precision, and the task priority. The distributed node operating status data includes the node's computing power resources, storage resources, network bandwidth, and historical fault records.

[0060] The complexity level in the task attribute information is an indicator based on the type of encryption algorithm and key length corresponding to the cryptanalysis task. For example, the complexity level of a task cracking a 128-bit symmetric encryption algorithm will be higher than that of a task cracking a 64-bit algorithm, thus directly relating to the scale of computing resources required for the task. The data volume refers to the total amount of encrypted data to be processed by the task. For example, a cryptanalysis task needs to process 100GB of ciphertext data. A larger data volume will place higher demands on the storage resources and data transmission capabilities of the nodes. The computational precision requirement is the task's tolerance for errors in the computation results. For example, some cryptanalysis tasks require precise matching of every character of the key, which has a high computational precision requirement, increasing the task's computation time and resource consumption. The task priority is a sorting criterion set according to the urgency of the task. For example, cryptanalysis tasks involving security emergencies have a higher priority than regular testing tasks, and tasks with higher priority will receive priority in resource allocation.

[0061] In distributed node runtime status data, node computing power resources refer to the hardware performance indicators of a node, such as CPU processing speed and DSP energy efficiency. For example, a node equipped with a multi-core processor and supporting parallel computing has computing power resources that are more suitable for high-complexity cryptanalysis tasks. Storage resources refer to the local or distributed storage capacity that a node can provide. When the task data volume is large, nodes with sufficient storage resources can avoid task interruptions caused by insufficient data storage. Network bandwidth refers to the data transmission rate between a node and other nodes or data centers. Nodes with high network bandwidth are more suitable for tasks with large data volumes and can accelerate the transmission and synchronization efficiency of encrypted data. Historical fault records refer to the types of faults, frequency of faults, and recovery time information that have occurred in the past operation of a node. For example, if a node's historical fault records show that it has experienced two downtime faults per month, such nodes will be preferentially excluded from the allocation list of high-priority tasks by scheduling rules to reduce the risk of task interruption.

[0062] The system integrates and judges task attribute information and node running status data. For example, when the complexity level of the cryptanalysis task is high, the data volume is large and the task priority is high, the system prioritizes matching nodes with strong computing power resources, sufficient storage resources, high network bandwidth and few historical fault records. Based on these characteristics, the system determines the corresponding fault-tolerant scheduling core parameters, such as setting a higher breakpoint recording frequency for the task to ensure that the task can quickly resume computing when an anomaly occurs.

[0063] The task attribute information is used to perform feature recognition to obtain task feature results. The task feature results are then matched with the tasks and nodes in the node running status data to determine the core parameters for fault-tolerant scheduling. This process includes the following steps:

[0064] Key features in the task attribute information are extracted and classified to obtain task complexity features, data processing features and priority features. The task complexity features, data processing features and priority features are then summarized to form the task feature results.

[0065] Statistically analyze the resource availability, failure rate, and failure recovery speed of each distributed node, and establish a node performance evaluation index system.

[0066] Based on the task feature results and node performance evaluation index system, the fitness scores of the task and each node are calculated to form a fitness dataset.

[0067] Based on the fit dataset and task priority, key parameters affecting scheduling stability and fault tolerance efficiency are selected, and core parameters for fault-tolerant scheduling are determined. These core parameters include breakpoint recording frequency, fault detection threshold, node switching threshold, and task shard size.

[0068] Key features in the task attribute information are extracted and classified, and finally summarized to form the task feature results. The task complexity feature is extracted from the encryption algorithm type and key length information of the task. For example, for a task that cracks a hash algorithm, its complexity feature will be marked as high level. The data processing feature is based on the data volume and computational precision requirements of the task. For example, if a task needs to process 200GB of ciphertext and requires the computational error to be less than 0.01%, its data processing feature will be defined as a large data volume and high precision type. The priority feature is determined according to the urgency of the task. For example, for a cryptanalysis task corresponding to an emergency security event, the priority feature will be set to the highest level. After classification, the complexity feature, data processing feature, and priority feature are integrated into a complete task feature result, which serves as the basis for subsequent adaptation.

[0069] By analyzing the operational data of each distributed node, a node performance evaluation index system is established. Resource margin refers to the remaining computing power, storage, and network resources of a node. For example, if a node's CPU utilization is only 30% and its remaining storage capacity is 500GB, its resource margin is considered sufficient. Failure rate is the number of times a node experiences a failure per unit of time. For example, if a node experiences three downtimes in the past month, its failure rate is marked as high. Failure recovery speed refers to the average time it takes for a node to recover from a failure to normal operation. For example, if a node recovers in an average of 10 minutes after a failure, its recovery speed is considered fast. By statistically analyzing resource margin, failure rate, and failure recovery speed, the performance levels of each node can be clearly defined, forming a quantifiable node performance evaluation index system.

[0070] Based on the task feature results and node performance evaluation index system, the fit score between the task and each node is calculated. For example, when the task feature results show that the task is highly complex, has a large amount of data, and is of high priority, nodes with sufficient resource reserves, low failure rate, and fast recovery speed are matched first. The fit score is assigned according to the degree of matching between the task features and the node index. For example, the fit score of a fully matched node is 90 points, and the score of a partially matched node is 60 points, ultimately forming a fit dataset covering all nodes.

[0071] Based on the suitability dataset and task priorities, key parameters affecting scheduling stability and fault tolerance efficiency are identified, thus determining the core parameters for fault-tolerant scheduling. For example, when assigning tasks to high-priority nodes with high suitability, the breakpoint recording frequency is set to once every 5 minutes to ensure rapid location of breakpoints after task interruption; the fault detection threshold is set to a lower level to promptly detect abnormal node states; the node switching threshold is matched to the task priority, with higher-priority tasks having a more lenient switching threshold to ensure rapid switching to a backup node when a node fails; and the task shard size is determined based on the node's resource availability, with nodes with sufficient resources receiving larger task shards to improve computational efficiency. These parameters collectively constitute the core parameters for fault-tolerant scheduling, providing a basis for subsequent scheduling scheme construction.

[0072] Based on breakpoint records, fault recovery cases, and core parameters of fault-tolerant scheduling from historical scheduling data, a breakpoint continuation model and a fault-tolerant scheduling rule base are constructed, specifically including the following steps:

[0073] Extract breakpoint location information, breakpoint time task progress data, and fault type records of cryptanalysis tasks from historical scheduling data to establish a breakpoint information dataset;

[0074] Based on the fault handling process, node replacement scheme, and task continuation scheme in historical fault recovery cases, generate fault-tolerant scheduling experience rules.

[0075] Using core parameters as constraints, and combining breakpoint information datasets, a task progress backtracking model is trained to construct a breakpoint continuation model.

[0076] The fault-tolerant scheduling rule library is obtained by inputting the fault-tolerant scheduling empirical rules into the breakpoint continuation calculation model.

[0077] Key information is extracted from historical scheduling data to establish a breakpoint information dataset. Breakpoint location information refers to the computational stage at which the task was interrupted. For example, if a cryptanalysis task is interrupted during the 3333rd operation of key cracking, the corresponding breakpoint location will be precisely recorded. Task progress data at the breakpoint time represents the percentage of computation already completed at the time of interruption. For instance, if the task has completed 60% of the key matching computation at the breakpoint time, this progress data will be synchronously retained. Fault type records indicate the specific fault category that caused the task interruption. For example, if the interruption was caused by insufficient network bandwidth, the corresponding fault type will be marked as a network fault. This allows for the creation of a breakpoint information dataset covering different task scenarios, providing foundational data for subsequent model training.

[0078] Historical fault recovery cases are analyzed to generate fault-tolerant scheduling rules of experience. The fault handling process is the operational steps from fault detection to recovery in the case study. For example, after a node crashes, the fault type is first detected, and then a backup node is switched. This process is refined into standardized steps. The node replacement scheme is the logic for selecting a replacement node for the failed node. For example, when a high-performance node fails, a node with the same computing power level and sufficient resource reserves is prioritized as a replacement. The task continuation scheme is the restart plan for tasks after fault recovery. For example, for a task with a 60% completion point, the continuation scheme explicitly resumes computation from the breakpoint rather than restarting it, thus forming reusable fault-tolerant scheduling rules of experience.

[0079] Using defined fault-tolerant scheduling core parameters as constraints, a task progress backtracking model is trained using a breakpoint information dataset, thereby constructing a breakpoint-resumption model. For example, if the core parameters specify a breakpoint recording frequency of once every 5 minutes, then during model training, the breakpoint location and task progress data will be matched at a 5-minute time granularity to ensure that the model can accurately backtrack the task progress status at any breakpoint. Once the model training is complete, the starting point for resumption of computation can be quickly located based on the real-time breakpoint information of the task; this model is then known as the breakpoint-resumption model.

[0080] The generated fault-tolerant scheduling rules are input into the breakpoint continuation model, enabling the model to output corresponding scheduling schemes based on these rules. These scheduling schemes are then organized into a structured set, namely the fault-tolerant scheduling rule library. For example, when a task is interrupted due to a node failure, the model will call the corresponding node replacement and continuation schemes from the rule library to automatically complete the fault handling and task continuation.

[0081] Based on the real-time status data and node load of the current cryptanalysis task, an initial scheduling scheme is generated by calling the breakpoint continuation model and the fault-tolerant scheduling rule base. The specific steps include:

[0082] Collect real-time status data on the current cryptanalysis task's execution progress, incomplete computation modules, and current resource usage;

[0083] Real-time load data of nodes is obtained by monitoring the real-time load rate, resource idleness and network connectivity status of each distributed node.

[0084] The real-time status data and node real-time load data are matched with the fault-tolerant scheduling rule base to obtain the scheduling scheme and node allocation scheme;

[0085] The breakpoint resume calculation model is invoked to verify the feasibility of breakpoint resume calculation of the scheduling scheme and node allocation scheme. The resource allocation ratio is adjusted in combination with task priority to generate an initial scheduling scheme.

[0086] The system collects real-time status data for the current cryptanalysis task, specifically covering the execution progress, incomplete computation modules, and current resource usage. The execution progress represents the percentage of the total task's computations completed so far; for example, if a key-breaking task has completed 70% of the character matching computations, this progress will be recorded in real time. Incomplete computation modules refer to the remaining computational stages of the task; for example, if the task still has the final three rounds of key iteration computation, the computational complexity and resource requirements of these modules will be simultaneously statistically analyzed. Current resource usage data shows the node's computing power and storage resources currently consumed by the task; for example, if the task has already used 40% of the node's CPU computing power and 200GB of storage resources, this data, when integrated, forms the real-time status data of the task, clearly reflecting its current execution status.

[0087] Monitoring the real-time load data of each distributed node mainly includes real-time load rate, resource idleness, and network connectivity status. Real-time load rate is the proportion of resources currently used by a node relative to its total resources. For example, a node's CPU real-time load rate of 60% indicates that its computing resources are largely utilized. Resource idleness is the proportion of remaining allocable resources for a node. For example, a node with 30% of its storage resources and 25% of its network bandwidth remaining demonstrates its resource redundancy. Network connectivity status is the stability of the connection between the node and other nodes or data centers. For example, a stable network connectivity status indicates that data transmission is less likely to be interrupted. By monitoring these indicators, the operational capacity of each node can be assessed in real time.

[0088] The real-time status data of tasks and the real-time load data of nodes are matched with the previously built fault-tolerant scheduling rule base to obtain a preliminary scheduling scheme and node allocation scheme. For example, if the rule base stipulates that high-progress tasks should be matched with nodes with high resource idleness, then a task that is currently 70% completed will be assigned to a node with a resource idleness of more than 40%. At the same time, the complexity of the unfinished computing modules of the task is combined to determine the proportion of computing resources allocated to the nodes. This process will generate the corresponding scheduling and node allocation scheme.

[0089] The breakpoint resumption model was invoked to verify the feasibility of breakpoint resumption under the above scheduling and node allocation schemes. For example, the model simulates whether the breakpoint can be accurately located and resumed quickly when a task is interrupted under this node allocation scheme. If the model verification results show that the resumption delay is too high, the resource allocation ratio will be adjusted according to the task priority. For example, if the task has the highest priority, the resource idle ratio of its allocated node will be increased to reduce the risk of resumption. After verification and adjustment, the initial scheduling scheme is finally formed, providing a foundation for subsequent optimization tests.

[0090] The initial scheduling scheme is subjected to fault simulation testing and scheduling efficiency evaluation to obtain scheduling quality evaluation results, which specifically includes the following steps:

[0091] Simulate typical failure scenarios such as distributed node offline, network interruption, and exhaustion of computing resources to test the fault detection response speed and breakpoint recording accuracy of the initial scheduling scheme;

[0092] The task allocation balance, node resource utilization, and task execution efficiency of the initial scheduling scheme are statistically analyzed, and a scheduling efficiency score is calculated.

[0093] The fault tolerance performance evaluation value is obtained by judging the fault recovery success rate, continuation data consistency and recovery time of the initial scheduling scheme in the fault simulation test;

[0094] The scheduling quality assessment result is determined by combining the scheduling efficiency score and the fault tolerance performance evaluation value.

[0095] Simulating typical failure scenarios such as distributed node offline, network interruption, and exhaustion of computing resources, the fault detection response speed and breakpoint recording accuracy of the initial scheduling scheme are tested. For example, simulating a scenario where an allocation node suddenly goes offline, the time it takes for the initial scheduling scheme to trigger fault detection is observed. If fault identification is completed within 10 seconds, the fault detection response speed meets the standard. Simultaneously, the content of the breakpoint recording is verified. For instance, if a task is at the 3332nd operation stage of key cracking when the node is offline, the breakpoint position recorded by the initial scheduling scheme must be completely consistent with the actual execution stage to determine the accuracy of the breakpoint recording. Through scenario simulation, the basic ability of the scheme to handle faults can be preliminarily verified.

[0096] The scheduling efficiency score is calculated by statistically analyzing multiple indicators of the initial scheduling scheme. Task allocation balance refers to the difference in the amount of tasks undertaken by each node. For example, if the scheme splits a large task and distributes it to three nodes, and the task load of the three nodes is 34%, 33%, and 33% respectively, then the balance is high. Node resource utilization is the ratio of the actual resources used by a node to the total resources. For example, if the computing power resource utilization rate of a node reaches 85%, it indicates that the resources are not idle. Task execution efficiency is the ratio of the task completion time to the expected time. For example, if the scheme makes the actual task completion time only 90% of the expected time, then the execution efficiency is excellent. The scheduling efficiency score is calculated by combining the weights of these indicators.

[0097] Based on the results of fault simulation tests, fault tolerance performance evaluation values ​​are obtained. Fault recovery success rate refers to the proportion of times the initial scheduling scheme successfully recovers tasks in fault scenarios out of the total number of tests. For example, if the scheme successfully recovers 9 times out of 10 offline node tests, the success rate is 90%. Resumed computation data consistency refers to the degree of matching between the results of task resumption and the data before the breakpoint. For example, if the intermediate result of key matching after resumption is completely consistent with the breakpoint record, the consistency standard is met. Recovery time refers to the time from the occurrence of the fault to the task's resumption. For example, if the recovery time in a fault scenario is only 2 minutes, the time performance is good. These indicators are integrated into a fault tolerance performance evaluation value to quantify the fault tolerance capability of the scheme.

[0098] The final scheduling quality assessment result is determined by combining the scheduling efficiency score and the fault tolerance performance evaluation value. For example, if the scheduling efficiency score is 90 points and the fault tolerance performance evaluation value is 88 points, the comprehensive evaluation result is calculated by combining the importance weights of the two. This result is used to determine whether the initial scheduling scheme meets the requirements and provides a basis for subsequent optimization and adjustment.

[0099] Based on the scheduling quality assessment results, the initial scheduling scheme is optimized and adjusted to form the final fault-tolerant scheduling scheme, which includes the following steps:

[0100] Based on the scheduling quality assessment results, identify unreasonable node allocation, inappropriate breakpoint recording frequency, and insufficient adaptability of fault tolerance scheme in the initial scheduling scheme, and determine the optimization impact weights.

[0101] Based on task priority and optimization impact weight, determine optimization priority, and use optimization priority to identify the core issues affecting fault tolerance stability;

[0102] Based on the core issues affecting fault tolerance stability, the node allocation scheme was adjusted, the breakpoint recording frequency and fault recovery strategy were optimized, and the feasibility and effectiveness of the scheme were re-verified to form the final fault-tolerant scheduling scheme.

[0103] Based on the scheduling quality assessment results, problems in the initial scheduling scheme are identified, and their impact weights are determined. For example, if the assessment results show that some nodes are undertaking tasks far exceeding their resource capacity, this indicates an unreasonable node allocation problem; if the breakpoint recording frequency is set too high, leading to redundant resource consumption, or too low, resulting in missing continuation data, this indicates an inappropriate breakpoint recording frequency; if the fault recovery scheme only adapts to network interruption faults but cannot cope with scenarios where node computing power is exhausted, this indicates insufficient adaptability of the fault tolerance scheme. Combining the degree of impact of these problems on scheduling quality, each problem is assigned a corresponding weight. For example, unreasonable node allocation has the greatest impact on scheduling stability, with a weight of 40%; inappropriate breakpoint recording frequency has a weight of 30%; and insufficient adaptability of the fault tolerance scheme has a weight of 30%.

[0104] By combining task priority and optimization impact weight, the optimization priority is determined, and the core issues affecting fault tolerance stability are identified accordingly. For example, if the current task priority is the highest, issues with high weight are addressed first: if an unreasonable node allocation has the highest weight, it is listed as the primary optimization priority and is identified as the core issue affecting fault tolerance stability; if the task priority is normal, the optimization order of multiple issues may be balanced, but issues with high weight are still the core issues.

[0105] To address the core issues affecting fault tolerance stability, the corresponding solutions are adjusted, and their feasibility and effectiveness are re-verified to form the final fault-tolerant scheduling scheme. For example, if the core issue is unreasonable node allocation, the node allocation scheme will be readjusted, migrating some tasks from high-load nodes to nodes with higher resource idleness. If the core issue is inappropriate breakpoint recording frequency, the breakpoint recording frequency will be optimized, for example, from once every 5 minutes to once every 3 minutes. If the core issue is insufficient adaptability of the fault tolerance scheme, a fault recovery scheme will be added, along with new node switching rules for scenarios where computing power is exhausted. After adjustments are made, fault scenarios are simulated again, and scheduling efficiency is evaluated. Once it is confirmed that the scheme has no obvious defects, an executable fault-tolerant scheduling scheme is finally formed.

[0106] The final fault-tolerant scheduling scheme is executed, and the breakpoint information set and scheduling log are collected in real time to generate a breakpoint-resumption fault-tolerant scheduling for the cryptanalysis task. This process includes the following steps:

[0107] The final fault-tolerant scheduling scheme is broken down into several sub-tasks, which are then assigned to the corresponding distributed nodes and started for execution.

[0108] According to the breakpoint recording frequency set in the core parameters, the execution progress, data processing status and node running parameters of each subtask are collected in real time to generate a breakpoint information set, which includes the current execution step, data verification value and resource usage snapshot;

[0109] The task scheduling process collects node allocation changes, fault occurrence time, fault handling measures, and restart time to form a complete scheduling log.

[0110] When a fault is detected, the task execution state is restored through the breakpoint information set and scheduling logs, and the node switching or task reassignment process is initiated according to the fault-tolerant scheduling rule base to realize the breakpoint resume calculation.

[0111] The final fault-tolerant scheduling scheme is broken down into several sub-tasks and assigned to corresponding distributed nodes for execution. For example, a cryptanalysis task that requires 10,000 encryption operations to crack the key can be broken down into three sub-tasks. Based on the resource adaptability of the nodes, the first sub-task is assigned to node A with sufficient computing power, the second to node B with abundant storage resources, and the third to node C with high network bandwidth. Then, the execution process of the sub-tasks on each node is started synchronously.

[0112] According to the breakpoint recording frequency set in the core parameters, the system collects information from each subtask in real time to generate a breakpoint information set. For example, if the core parameters specify a breakpoint recording frequency of once every 3 minutes, the system will record the execution progress of the subtask every 3 minutes, such as the subtask of node A having completed 60% of its computation; it will also record the data processing status, such as currently performing the 3334th operation; and it will also record the node's running parameters, such as the CPU utilization of node A being 75%. The breakpoint information set formed by integrating this information specifically includes the data verification value of the current execution step and a resource usage snapshot. The current execution step clarifies the computational stage of the subtask, the data verification value ensures the accuracy of the data when continuing the computation, and the resource usage snapshot records the current resource consumption of the node.

[0113] Key information collected during the task scheduling process is compiled into a complete scheduling log. This includes node allocation changes, such as the operation record of migrating a subtask from node A to node D; the time of failure occurrence, such as node B experiencing a computing power exhaustion failure at 14:20 while executing a subtask; failure handling measures, such as the system initiating a switchover operation on backup node E after detecting the failure; and the time of resumed computation, such as node E completing preparation and starting the subtask resumed computation at 14:25. These recorded information provide a complete basis for fault tracing and solution optimization.

[0114] When the system detects a fault, it will resume execution based on the breakpoint information set and scheduling logs. For example, if node B experiences a computing power exhaustion fault, the system first retrieves the breakpoint information set to determine that the subtask's execution progress is 50% complete, the data verification value is a specific code, and the resource usage snapshot shows that node B's computing power has reached its limit. Then, combined with the fault handling measures recorded in the scheduling logs, the system restores the execution state of the subtask through the breakpoint resumption model, allowing it to continue running from the breakpoint position of the 3333rd operation. Simultaneously, based on the fault-tolerant scheduling rule base, the system initiates a node switching process, assigning the subtask to the backup node E, completing the task reallocation, and ultimately achieving breakpoint resumption of the subtask, ensuring the continuous progress of the entire cryptanalysis task.

[0115] It also includes subsequent processing of scheduling logs and breakpoint information sets, specifically including the following steps:

[0116] Regularly analyze the scheduling logs, extract patterns of failures and scheduling optimization opportunities, and update the fault-tolerant scheduling rule base.

[0117] The breakpoint information set is encrypted, stored, and backed up to ensure data security and integrity, providing reference data for scheduling similar tasks in the future.

[0118] The scheduling logs are analyzed and processed regularly to update the fault-tolerant scheduling rule base. For example, the system summarizes and judges the scheduling logs weekly, extracting patterns of failure occurrences. If it is found that the occurrence rate of node computing power exhaustion failures on Wednesday afternoons is three times that of other times, the resource scheduling risk corresponding to this pattern is identified. At the same time, scheduling optimization space is extracted. For example, if the logs show that the fault recovery time is 50% longer when a certain type of subtask is assigned to a low-computing-power node than when it is assigned to a high-computing-power node, it indicates that there is room for optimization in the node allocation scheme for this type of task. These patterns and optimization directions are transformed into new scheduling rules, such as adding a node computing power resource reservation strategy for Wednesday afternoons, adjusting the node allocation priority of corresponding subtasks, and thus updating the fault-tolerant scheduling rule base to improve the rationality of subsequent scheduling schemes.

[0119] The breakpoint information set is encrypted, stored, and backed up to ensure data security and provide a reference for subsequent tasks. The breakpoint information set contains critical content such as task execution progress data and verification values, which directly affect the accuracy of the continuation of the cryptanalysis task and the confidentiality of the data. Therefore, the system uses encryption algorithms that meet security standards to encrypt it, preventing information leakage or tampering. Multiple backups are also made on different storage nodes to prevent the loss of breakpoint information due to the failure of a single node. These encrypted and backed-up breakpoint information sets serve as reference data for similar tasks. For example, when a cryptanalysis task of the same complexity level occurs later, the system can retrieve resource usage snapshots from the historical breakpoint information set to plan node resource allocation in advance, improving scheduling efficiency and fault tolerance.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A fault-tolerant scheduling method for resuming breakpoint computation in cryptanalysis tasks, characterized in that, The method includes the following steps: Collect task attribute information, historical scheduling data, and distributed node running status data for cryptanalysis tasks; Task feature results are obtained by performing feature recognition on task attribute information. The task feature results are then matched with the tasks and nodes in the node running status data to determine the core parameters of fault-tolerant scheduling. Based on breakpoint records, fault recovery cases, and core parameters of fault-tolerant scheduling in historical scheduling data, a breakpoint continuation model and a fault-tolerant scheduling rule base are constructed. Based on the real-time status data and node load of the current cryptanalysis task, the breakpoint continuation model and fault-tolerant scheduling rule base are invoked to generate an initial scheduling scheme. The scheduling quality assessment results are obtained by conducting fault simulation tests and scheduling efficiency evaluations on the initial scheduling scheme. Based on the scheduling quality assessment results, the initial scheduling scheme is optimized and adjusted to form the final fault-tolerant scheduling scheme. The final fault-tolerant scheduling scheme is executed, and the breakpoint information set and scheduling log are collected in real time to generate a breakpoint-resumption fault-tolerant scheduling for the cryptographic analysis task.

2. The fault-tolerant scheduling method for breakpoint continuation of cryptanalysis tasks according to claim 1, characterized in that, The task attribute information includes the complexity level of the cryptanalysis task, the data volume, the required computational accuracy, and the task priority. The distributed node operating status data includes the node's computing power resources, storage resources, network bandwidth, and historical fault records.

3. The fault-tolerant scheduling method for breakpoint continuation of cryptanalysis tasks according to claim 2, characterized in that, The task attribute information is used to perform feature recognition to obtain task feature results. The task feature results are then matched with the tasks and nodes in the node running status data to determine the core parameters for fault-tolerant scheduling. This process includes the following steps: Key features in the task attribute information are extracted and classified to obtain task complexity features, data processing features and priority features. The task complexity features, data processing features and priority features are then summarized to form the task feature results. Statistically analyze the resource availability, failure rate, and failure recovery speed of each distributed node, and establish a node performance evaluation index system. Based on the task feature results and node performance evaluation index system, the fitness scores of the task and each node are calculated to form a fitness dataset. Based on the fit dataset and task priority, key parameters affecting scheduling stability and fault tolerance efficiency are selected, and core parameters for fault-tolerant scheduling are determined. These core parameters include breakpoint recording frequency, fault detection threshold, node switching threshold, and task shard size.

4. The fault-tolerant scheduling method for breakpoint continuation of cryptanalysis tasks according to claim 3, characterized in that, Based on breakpoint records, fault recovery cases, and core parameters of fault-tolerant scheduling from historical scheduling data, a breakpoint continuation model and a fault-tolerant scheduling rule base are constructed, specifically including the following steps: Extract breakpoint location information, breakpoint time task progress data, and fault type records of cryptanalysis tasks from historical scheduling data to establish a breakpoint information dataset; Based on the fault handling process, node replacement scheme, and task continuation scheme in historical fault recovery cases, generate fault-tolerant scheduling experience rules. Using core parameters as constraints, and combining breakpoint information datasets, a task progress backtracking model is trained to construct a breakpoint continuation model. The fault-tolerant scheduling rule library is obtained by inputting the fault-tolerant scheduling empirical rules into the breakpoint continuation calculation model.

5. The fault-tolerant scheduling method for breakpoint continuation of cryptanalysis tasks according to claim 4, characterized in that, Based on the real-time status data and node load of the current cryptanalysis task, an initial scheduling scheme is generated by calling the breakpoint continuation model and the fault-tolerant scheduling rule base. The specific steps include: Collect real-time status data on the current cryptanalysis task's execution progress, incomplete computation modules, and current resource usage; Real-time load data of nodes is obtained by monitoring the real-time load rate, resource idleness and network connectivity status of each distributed node. The real-time status data and node real-time load data are matched with the fault-tolerant scheduling rule base to obtain the scheduling scheme and node allocation scheme; The breakpoint continuation model is invoked to verify the feasibility of breakpoint continuation of the scheduling scheme and node allocation scheme. The resource allocation ratio is adjusted in combination with task priority to generate an initial scheduling scheme.

6. The fault-tolerant scheduling method for breakpoint continuation of cryptanalysis tasks according to claim 5, characterized in that, The initial scheduling scheme is subjected to fault simulation testing and scheduling efficiency evaluation to obtain scheduling quality evaluation results, which specifically includes the following steps: Simulate typical failure scenarios such as distributed node offline, network interruption, and exhaustion of computing resources to test the fault detection response speed and breakpoint recording accuracy of the initial scheduling scheme; The task allocation balance, node resource utilization, and task execution efficiency of the initial scheduling scheme are statistically analyzed, and a scheduling efficiency score is calculated. The fault tolerance performance evaluation value is obtained by judging the fault recovery success rate, continuation data consistency and recovery time of the initial scheduling scheme in the fault simulation test; The scheduling quality assessment result is determined by combining the scheduling efficiency score and the fault tolerance performance evaluation value.

7. The fault-tolerant scheduling method for breakpoint continuation of cryptanalysis tasks according to claim 6, characterized in that, Based on the scheduling quality assessment results, the initial scheduling scheme is optimized and adjusted to form the final fault-tolerant scheduling scheme, which includes the following steps: Based on the scheduling quality assessment results, identify unreasonable node allocation, inappropriate breakpoint recording frequency, and insufficient adaptability of fault tolerance scheme in the initial scheduling scheme, and determine the optimization impact weights. Based on task priority and optimization impact weight, determine optimization priority, and use optimization priority to identify the core issues affecting fault tolerance stability; Based on the core issues affecting fault tolerance stability, the node allocation scheme was adjusted, the breakpoint recording frequency and fault recovery strategy were optimized, and the feasibility and effectiveness of the scheme were re-verified to form the final fault-tolerant scheduling scheme.

8. The fault-tolerant scheduling method for breakpoint continuation of cryptanalysis tasks according to claim 7, characterized in that, The final fault-tolerant scheduling scheme is executed, and the breakpoint information set and scheduling log are collected in real time to generate a breakpoint-resumption fault-tolerant scheduling for the cryptanalysis task. This process includes the following steps: The final fault-tolerant scheduling scheme is broken down into several sub-tasks, which are then assigned to the corresponding distributed nodes and started for execution. According to the breakpoint recording frequency set in the core parameters, the execution progress, data processing status and node running parameters of each subtask are collected in real time to generate a breakpoint information set, which includes the current execution step, data verification value and resource usage snapshot; The task scheduling process collects node allocation changes, fault occurrence time, fault handling measures, and restart time to form a complete scheduling log. When a fault is detected, the task execution state is restored through the breakpoint information set and scheduling logs, and the node switching or task reassignment process is initiated according to the fault-tolerant scheduling rule base to realize the breakpoint resume computing.

Citation Information

Cited By

  • Adaptive fault-tolerant multi-party secure computation (SMPC) dynamic scheduling method and device

    CN122195685A

  • Adaptive Fault-Tolerant Multi-Party Secure Computation (SMPC) Dynamic Scheduling Method and Apparatus

    CN122195685B