Cryptographic operation task scheduling method and device, electronic equipment and storage medium

By acquiring task characteristics and node status information, calculating dynamic trust scores, and optimizing scheduling decisions using a scheduling model, the problem of insufficient consideration of task heterogeneity and trust in distributed CPU clusters is solved, and efficient and secure scheduling of cryptographic tasks is achieved.

CN122152480APending Publication Date: 2026-06-05GUANGDONG CERTIFICATE AUTHORITY
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG CERTIFICATE AUTHORITY
Filing Date
2026-04-30
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

In a distributed CPU cluster environment, traditional cryptographic task scheduling methods fail to adequately consider task heterogeneity, insufficient awareness of node dynamic states, and lack of consideration for node trust levels. This leads to unreasonable task allocation, reduced cluster throughput and response speed, and increased security risks.

Method used

By acquiring task feature information and node status information of cryptographic tasks, dynamic trust scores are calculated, and scheduling is performed using a pre-set cryptographic scheduling model. By combining Markov decision processes and reinforcement learning algorithms, scheduling decisions are optimized to achieve globally optimal scheduling.

Benefits of technology

It achieves globally optimal scheduling of cryptographic tasks, significantly reducing time costs, improving computational efficiency and security, and ensuring efficient and secure task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122152480A_ABST
    Figure CN122152480A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of cryptographic operation task scheduling, and discloses a cryptographic operation task scheduling method and device, an electronic device and a storage medium. The method comprises the following steps: obtaining task characteristic information of a cryptographic operation task and node state information of each to-be-allocated node, performing calculation according to the node state information, obtaining a dynamic trust score of each to-be-allocated node, performing scheduling processing by using a preset cryptographic operation scheduling model based on the dynamic trust score, in combination with the task characteristic information and the node state information, and obtaining cryptographic operation scheduling result information corresponding to the cryptographic operation task. Through the above method, global optimal scheduling of the cryptographic operation task is realized, and the operation efficiency of the cryptographic operation task is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cryptographic task scheduling, and more specifically, to a cryptographic task scheduling method, apparatus, electronic device, and storage medium. Background Technology

[0002] In a distributed central processing unit (CPU) cluster environment, the execution efficiency and resource utilization of cryptographic tasks are key performance indicators. With the rapid development of technologies such as cloud computing, big data, and the Internet of Things, the demand for cryptographic computation is increasing, and the complexity and diversity of cryptographic tasks are also increasing.

[0003] Traditional scheduling methods often fail to adequately consider the heterogeneity of cryptographic tasks, such as the varying demands of different tasks on computing resources, storage resources, and network bandwidth. Furthermore, these methods struggle to perceive and respond in real-time to the dynamic states of individual computing nodes in a distributed CPU cluster, including real-time node load, available resources, network latency, and potential failure risks. This information asymmetry and decision-making lag lead to inappropriate task allocation, potentially causing some nodes to be overloaded while others remain idle, thereby reducing the overall cluster throughput and response speed.

[0004] Furthermore, traditional cryptographic task scheduling methods often lack consideration for node trust levels, failing to effectively prevent sensitive cryptographic tasks from being scheduled to nodes with security vulnerabilities, thereby increasing the security risks of cryptographic tasks.

[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0006] The purpose of this application is to provide a method, apparatus, electronic device, and storage medium for scheduling cryptographic tasks. By inputting the task feature information of the cryptographic task, the node state information of each node to be assigned, and the dynamic trust score calculated based on the node state information into a preset cryptographic task scheduling model, the cryptographic task scheduling result information corresponding to the cryptographic task is calculated. This solves the problems of existing cryptographic task scheduling methods, such as insufficient consideration of task heterogeneity, insufficient awareness of node dynamic state, and lack of consideration of node trust. It can achieve globally optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and improve the computational efficiency of cryptographic tasks.

[0007] Firstly, this application provides a method for scheduling cryptographic computation tasks, comprising the following steps: Obtain the task characteristic information of the cryptographic operation task and the node status information of each node to be assigned; The dynamic trust score of each node to be assigned is calculated based on the node status information. Based on the dynamic trust score, combined with the task feature information and the node status information, a preset cryptographic operation scheduling model is used for scheduling processing to obtain the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

[0008] The cryptographic task scheduling method provided in this application can schedule cryptographic tasks. By inputting the task feature information of the cryptographic task, the node state information of each node to be assigned, and the dynamic trust score calculated based on the node state information into a preset cryptographic scheduling model, the cryptographic scheduling result information corresponding to the cryptographic task is calculated. This method solves the problems of existing cryptographic task scheduling methods, such as insufficient consideration of task heterogeneity, insufficient awareness of node dynamic state, and lack of consideration of node trust. It can achieve globally optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and improve the computational efficiency of cryptographic tasks.

[0009] Optionally, the node status information includes monitoring indicator information; based on the node status information, a dynamic trust score is calculated for each of the nodes to be assigned, including: The monitoring indicator information of each node to be assigned is normalized to obtain the current monitoring indicator parameters of each node to be assigned. For each node to be assigned, a dynamic trust score is obtained by dynamically scoring based on the current monitoring indicator parameters and the preset indicator weights corresponding to the current monitoring indicator parameters.

[0010] Optionally, the cryptographic operation scheduling model is constructed through the following steps: Obtain pre-set cryptographic operation scheduling evaluation parameters and their corresponding parameter weights; the cryptographic operation scheduling evaluation parameters include a security reward score calculated based on the current value of the dynamic trust score; Based on the cryptographic operation scheduling evaluation parameters and the parameter weights, a mathematical model containing the security reward score is constructed to obtain the initial scheduling reward function; With the aim of maximizing the cumulative scheduling reward parameter, a cryptographic operation scheduling model is constructed based on the initial scheduling reward function using a Markov decision process.

[0011] The cryptographic task scheduling method provided in this application can schedule cryptographic tasks. By introducing security reward scores and combining them with Markov decision processes to construct a scheduling model, the scheduling decision not only considers performance but also security, thereby effectively reducing security risks while ensuring efficient task execution.

[0012] Optionally, with the aim of maximizing the cumulative scheduling reward parameter, a Markov decision process is used to construct a cryptographic operation scheduling model based on the initial scheduling reward function, including: Using a Markov decision process, the initial scheduling reward function is set with a quintuple parameter, a learning parameter, and a cumulative scheduling reward parameter. Based on the quintuple parameters and the learning parameters, the correlation between the initial scheduling reward function and the cumulative scheduling reward parameters is constructed to obtain the calculation expression for the cumulative scheduling reward parameters; With the aim of maximizing the cumulative scheduling reward parameter, the calculation expression of the cumulative scheduling reward parameter is adjusted to obtain the cryptographic operation scheduling model.

[0013] Optionally, based on the dynamic trust score, combined with the task feature information and the node state information, a preset cryptographic operation scheduling model is used for scheduling processing to obtain cryptographic operation scheduling result information corresponding to the cryptographic operation task, including: The task feature information, the node status information, and the dynamic trust score are encoded to obtain a status identifier vector; The state identifier vector is input into the preset cryptographic operation scheduling model to obtain the target scheduling model; Based on the target scheduling model, the cryptographic operation tasks are allocated to obtain the optimal scheduling scheme; The optimal scheduling scheme is determined as the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

[0014] Optionally, the task feature information, the node state information, and the dynamic trust score are encoded to obtain a state identifier vector, including: The cryptographic operation task is divided into multiple subtasks, and the task feature information of each subtask is converted into a corresponding task feature vector. The node state information and dynamic trust score of each node to be assigned are converted into a corresponding node state vector. The task feature vector and the node state vector are concatenated and mapped to obtain the state identifier vector.

[0015] Optionally, based on the target scheduling model, the cryptographic operation tasks are allocated to obtain an optimal scheduling scheme, including: Using the target scheduling model, each subtask in the cryptographic operation task is matched with each node to be assigned, resulting in a set of scheduling schemes; Calculate the cumulative scheduling reward value for each scheduling scheme in the set of scheduling schemes; The optimal scheduling scheme is obtained by selecting the scheduling scheme with the largest cumulative reward value from the set of scheduling schemes.

[0016] The cryptographic task scheduling method provided in this application can schedule cryptographic tasks. By calculating and comparing the cumulative reward values ​​of different scheduling schemes and selecting the one with the largest reward as the optimal scheme, the scheduling result is optimized in terms of both performance and security, thereby further improving the decision quality of scheduling.

[0017] Secondly, this application provides a cryptographic task scheduling device, comprising: The acquisition module is used to acquire the task characteristic information of the cryptographic operation task and the node status information of each node to be assigned; The calculation module is used to calculate the dynamic trust score of each node to be assigned based on the node status information. The scheduling module is used to perform scheduling processing based on the dynamic trust score, combined with the task feature information and the node status information, using a preset cryptographic operation scheduling model to obtain the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

[0018] This cryptographic task scheduling device inputs the task feature information of the cryptographic task, the node state information of each node to be assigned, and the dynamic trust score calculated based on the node state information into a preset cryptographic task scheduling model to calculate the cryptographic task scheduling result information. It solves the problems of existing cryptographic task scheduling methods, such as failing to fully consider task heterogeneity, insufficient awareness of node dynamic state, and lack of consideration of node trust. It can achieve global optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and improve the computational efficiency of cryptographic tasks.

[0019] Thirdly, this application provides an electronic device, including a processor and a memory, wherein the memory stores a computer program executable by the processor, and when the processor executes the computer program, it performs the steps in the cryptographic task scheduling method described above.

[0020] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the cryptographic task scheduling method described above.

[0021] Beneficial effects: The cryptographic task scheduling method, apparatus, electronic device, and storage medium provided in this application input the task feature information of the cryptographic task, the node state information of each node to be assigned, and the dynamic trust score calculated based on the node state information into a preset cryptographic task scheduling model to calculate the cryptographic task scheduling result information. This solves the problems of existing cryptographic task scheduling methods, such as insufficient consideration of task heterogeneity, insufficient awareness of node dynamic state, and lack of consideration of node trust. It can achieve global optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and improve the computational efficiency of cryptographic tasks. Attached Figure Description

[0022] Figure 1 A flowchart of a cryptographic task scheduling method provided in an embodiment of this application.

[0023] Figure 2 This is a schematic diagram of the structure of the cryptographic task scheduling device provided in the embodiments of this application.

[0024] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0025] Labeling Explanation: 1. Acquisition Module; 2. Calculation Module; 3. Scheduling Module; 301. Processor; 302. Memory; 303. Communication Bus. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0027] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0028] Please refer to Figure 1 , Figure 1This application provides a method for scheduling cryptographic computation tasks in some embodiments, used for scheduling cryptographic computation tasks, including: Step S101: Obtain the task feature information of the cryptographic operation task and the node status information of each node to be assigned; Step S102: Calculate the dynamic trust score of each node to be assigned based on the node status information. Step S103: Based on dynamic trust scoring, combined with task feature information and node status information, a preset cryptographic operation scheduling model is used for scheduling processing to obtain the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

[0029] This cryptographic task scheduling method inputs the task feature information of the cryptographic task, the node state information of each node to be assigned, and the dynamic trust score calculated based on the node state information into a preset cryptographic task scheduling model to calculate the cryptographic task scheduling result information. It solves the problems of existing cryptographic task scheduling methods, such as failing to fully consider task heterogeneity, insufficient awareness of node dynamic state, and lack of consideration of node trust. It can achieve global optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and improve the computational efficiency of cryptographic tasks.

[0030] Specifically, in step S101, the task characteristic information of the cryptographic operation task and the node status information of each node to be assigned are obtained. The cryptographic operation task refers to a computational task requiring operations such as encryption, decryption, signing, and signature verification. Task characteristic information includes task type, computational load, data volume, priority, and deadline. This information can be obtained through user input during task submission or automatic system identification. Each node to be assigned refers to a computational unit in a distributed CPU cluster that can be used to execute the cryptographic operation task. Its node status information may include data such as CPU utilization, memory usage, network bandwidth, available storage space, current load, length of the pending queue within the TEE, network latency matrix, and monitoring metrics. This node status information can be collected in real-time by monitoring agents deployed on each node to be assigned.

[0031] Among them, monitoring indicator information refers to the indicator information of basic operation dimension (such as node liveness status, CPU utilization, memory utilization, etc., which are used to reflect node stability), security compliance dimension (such as task success rate, cryptographic operation error rate, TEE integrity proof verification results, abnormal access audit logs, etc., which are used to reflect security events and compliance), environmental security dimension (such as data center security level, network isolation domain, etc., which are used to reflect physical and network environment), and behavioral reputation dimension (such as key management audit logs, historical task completion records, etc., which are used to reflect historical collaboration reliability).

[0032] Specifically, in step S102, a dynamic trust score for each node to be assigned is calculated based on the node status information, including: The monitoring metric information of each node to be assigned is normalized to obtain the current monitoring metric parameters of each node to be assigned. For each node to be assigned, a dynamic trust score is obtained by dynamically scoring it based on the current monitoring indicator parameters and the preset indicator weights corresponding to the current monitoring indicator parameters.

[0033] In step S102, existing methods such as min-max normalization and Z-score normalization can be used to normalize the monitoring indicator information of each node to be assigned, thereby obtaining the current monitoring indicator parameters of each node. This eliminates the influence of different dimensions between indicators and ensures the comparability of various indicators during the scoring process. Therefore, the current monitoring indicator parameters of each node can objectively reflect its performance in various dimensions.

[0034] Dynamic trust score is a dynamic numerical value derived from a quantitative assessment of the reliability, security, or performance of each node. It is calculated using a weighted fusion algorithm. This dynamic trust score reflects the trust level of a node in real-time or near real-time; a higher score generally indicates a more trustworthy node. By dynamically scoring each node using current monitoring indicator parameters and their corresponding preset indicator weights, the overall trustworthiness of the nodes can be quantified. The dynamic trust score can be calculated using the following formula: ; in, Let J be the dynamic trust score of node j in the m-th period (i.e., the trust score in the current period). Let J be the dynamic trust score of node j in the (m-1)th period (i.e., the trust score of the previous period). This is the historical attenuation coefficient, which can be set according to actual needs; The preset indicator weight is the corresponding indicator weight for the kth monitoring indicator information; This refers to the current monitoring indicator parameter (i.e., the indicator value after normalization) for the k-th monitoring indicator information in the m-th period. When m=1, This represents the default value of the dynamic trust score for node j, which is 0.

[0035] Specifically, the pre-defined cryptographic operation scheduling model is constructed through the following steps: Obtain the pre-set cryptographic operation scheduling evaluation parameters and their corresponding parameter weights; the cryptographic operation scheduling evaluation parameters include the security reward score calculated based on the current value of the dynamic trust score; Based on cryptographic operation scheduling evaluation parameters and parameter weights, a mathematical model containing security reward scores is constructed to obtain the initial scheduling reward function; With the aim of maximizing the cumulative reward parameter of the scheduling, a pre-defined cryptographic operation scheduling model is obtained by using a Markov decision process and constructing it based on the initial scheduling reward function.

[0036] Before constructing a cryptographic operation scheduling model, it is necessary to measure various metrics to evaluate the quality and effectiveness of cryptographic operation task scheduling. These evaluation parameters may include, but are not limited to, task completion time (representing the performance of the scheduling scheme), the standard deviation of the load on each node upon task completion (representing the load balancing performance of the scheduling scheme), security reward score (representing the security performance of the scheduling scheme), and data transmission volume between nodes (representing the communication cost performance of the scheduling scheme). For example, the shorter the task completion time, the better the scheduling effect. Parameter weights are used to represent the relative importance of each evaluation parameter in the overall scheduling objective. For example, in some scenarios, task completion time may be more critical than load balancing, so its weight will be set higher. These parameters and weights can be pre-set and adjusted according to actual application needs, system characteristics, and business strategies.

[0037] After defining the evaluation parameters and weights, these qualitative or quantitative indicators need to be transformed into a mathematical expression that includes a safety reward score, i.e., the initial scheduling reward function. This function aims to quantify the immediate benefits of each scheduling decision. For example, the initial scheduling reward function can be constructed as a weighted sum model, where each evaluation parameter is multiplied by its corresponding weight, and then all weighted parameters are summed to obtain a comprehensive reward value. The higher the reward value, the better the current scheduling decision. This mathematical model provides a basic evaluation framework for subsequent optimization. The safety reward score serves as a key bonus factor to ensure that scheduling decisions fully consider safety factors while pursuing efficiency.

[0038] The initial scheduling reward function is as follows: ; in, The immediate scheduling reward for scheduling scheme i; Let i be the task completion time for scheduling scheme i; Let be the standard deviation of the load on each node when scheduling scheme i is completed. The load on each node when scheduling scheme i is completed; The safety reward score for scheduling scheme i; This represents the amount of data transmission between nodes when scheduling scheme i is completed. The parameter weights corresponding to the task completion time; The parameter weights are the standard deviations of the load. The parameter weights corresponding to the security rewards; The parameter weights are the data transmission volume.

[0039] The security reward score can be calculated based on the node's current trust score (i.e., the current value of the dynamic trust score). The specific formula for calculating the security reward score is as follows: ; Where n is the total number of nodes to be assigned, j≤n; For safety reward weighting coefficient; This is a boolean value indicating whether node j executes the task within scheduling scheme i. When node j executes the task within scheduling scheme i... When node j is not executing a task within scheduling scheme i, ; The penalty base represents the security violation penalty item, which is used to impose additional penalties when the scheduling scheme violates security constraints. It can be set according to the actual situation. For example, it can be set according to the dynamic trust score, comparing the trust score of the current period with the preset penalty threshold to determine the number of nodes with a trust score lower than the preset penalty threshold. Different threshold ranges are set according to this number, and the corresponding penalty base is set for the corresponding threshold range.

[0040] Specifically, with the aim of maximizing the cumulative scheduling reward parameter, a Markov decision process is used to construct a pre-defined cryptographic operation scheduling model based on the initial scheduling reward function, including: Using Markov decision processes, we set quintuple parameters, learning parameters, and cumulative scheduling reward parameters for the initial scheduling reward function; Based on the quintuple parameters and the learned parameters, the relationship between the initial scheduling reward function and the cumulative scheduling reward parameters is constructed, and the calculation expression for the cumulative scheduling reward parameters is obtained. With the aim of maximizing the cumulative reward parameter of the scheduling, the calculation expression of the cumulative reward parameter of the scheduling is adjusted to obtain the cryptographic operation scheduling model.

[0041] In constructing a cryptographic task scheduling model, the quintuple parameters are a core component of the Markov Decision Process (MDP), typically including the state set (S), action set (A), state transition probabilities (P), reward function (R), and discount factor (γ). Specifically, the state set S represents all possible states during the cryptographic task scheduling process, such as a combination of the current task's features and the node state information of the node to be assigned; the action set A represents all scheduling actions that can be taken in each state, such as assigning a subtask to a node; the state transition probability P describes the probability that the system will transition to the next state after taking an action in a given state; the reward function R defines the immediate reward obtained by taking a specific action in a specific state, corresponding to the initial scheduling reward function mentioned above; and the discount factor γ is used to measure the importance of future rewards. The aim is to provide a clear decision-making environment model for reinforcement learning algorithms by clarifying these fundamental elements.

[0042] Learning parameters are key parameters used to control the learning process in reinforcement learning algorithms, such as the learning rate (α) and the exploration rate (ε). The learning rate α determines the degree to which new information corrects old information each time the Q-table is updated; the exploration rate ε controls the balance between exploring unknown actions and utilizing known optimal actions. By setting these parameters appropriately, reinforcement learning algorithms can be effectively guided to achieve a balance between exploration and utilization, thereby accelerating the learning process and improving convergence.

[0043] The table of cumulative scheduling reward parameters (i.e., the Q-table) is used in reinforcement learning to store the value of state-action pairs. Each cumulative scheduling reward value Q(s, a) in the Q-table represents the expected future cumulative reward that can be obtained by taking action a in state s. The purpose is that, by continuously updating the Q-table, reinforcement learning algorithms can learn which actions to take in different states to maximize long-term cumulative rewards, thereby forming an optimal scheduling strategy. In practical applications, the cumulative scheduling reward parameters can be obtained by training tabular Q-Learning and Deep Q-Networks (DQN) based on historical data or real-time data (i.e., historical records based on task feature information and node state information, or using existing real-time task feature information and real-time node state information of nodes).

[0044] Based on the quintuple parameters and the learned parameters, a correlation is constructed between the initial scheduling reward function and the cumulative scheduling reward parameters, resulting in an expression for calculating the cumulative scheduling reward parameters. For example, this expression can be constructed based on the Bellman equation to describe the iterative relationship between Q values, i.e.: ; in, In the m-th period, the state is... Select action The achievable cumulative scheduling reward parameter (i.e., the cumulative scheduling reward value, usually referred to as the Q value); In the (m-1)th period, the state is... Select action The cumulative scheduling reward parameters that can be obtained; The learning rate; Select action for the m-th cycle The immediate reward that can be obtained afterward; This is a discount factor (which controls the weight of future rewards and indicates the degree of importance attached to future rewards). This refers to the new state entered in the m-1th cycle after performing the action in the m-1th cycle. This refers to a new action that is performed in the m-1th cycle after the action in the m-1th cycle. To enter a new state in the m-th period Then take the maximum Q value from all possible new actions; when m=1, This is the initial value for scheduling the accumulated reward value. To enter a new state at the initial moment Then take the maximum Q value among all possible new actions. and It can be calculated using tabular Q-Learning and deep Q-networks. The purpose of the expression for calculating the cumulative reward parameter for scheduling is to gradually approximate the optimal Q value through iterative calculation.

[0045] The goal is to maximize the cumulative reward parameter for scheduling. This is achieved by iteratively updating the Q-values ​​in the Q-table until each Q-value eventually converges to the maximum cumulative reward obtainable by taking the corresponding action in the corresponding state. The expression for calculating the cumulative reward parameter is then adjusted to obtain a cryptographic operation scheduling model that guides optimal scheduling. The cryptographic operation scheduling model is as follows: ; in, This is a cryptographic operation scheduling model; for The function is used to select the action that maximizes the Q value from all possible actions as the scheduling scheme for this time.

[0046] Specifically, in step S103, based on dynamic trust scoring, and combining task feature information and node status information, a preset cryptographic operation scheduling model is used for scheduling processing to obtain cryptographic operation scheduling result information corresponding to the cryptographic operation task, including: Encode the task feature information, node status information, and dynamic trust score to obtain a status identifier vector; The state identifier vector is input into the preset cryptographic operation scheduling model to obtain the target scheduling model; Based on the target scheduling model, task allocation for cryptographic operations is performed to obtain the optimal scheduling scheme. The optimal scheduling scheme is determined by the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

[0047] Specifically, in step S103, the task feature information, node state information, and dynamic trust score are encoded to obtain a state identifier vector, including: The cryptographic task is divided into multiple subtasks, and the task feature information of each subtask is transformed into a corresponding task feature vector. The node status information and dynamic trust score of each node to be assigned are converted into the corresponding node status vector. The task feature vector and node state vector are concatenated and mapped to obtain the state identifier vector.

[0048] In step S103, a large cryptographic task is decomposed into several smaller, more manageable subtasks based on attributes such as complexity, computational load, security level, or execution stage. For example, a complex encryption process may be divided into multiple subtasks such as key generation, data encryption, and signature verification. This division aims to achieve finer-grained scheduling and resource allocation, improving the parallelism and efficiency of task execution. After dividing into multiple subtasks, the task feature information of each subtask is quantized and vectorized to obtain the corresponding task feature vector. This vector form facilitates subsequent processing and analysis.

[0049] The node status information and dynamic trust score of each node to be assigned are quantified and vectorized to integrate and transform them into a unified node status vector. This node status vector can comprehensively characterize the current operating status and trustworthiness of the node to be assigned.

[0050] The concatenation process combines the feature vector of a subtask with the node state vector of the node to be assigned, forming a longer vector that comprehensively reflects the potential state of a specific subtask executed on a specific node. The mapping process performs dimensionality transformation, normalization, or feature extraction on the concatenated vector to generate a state identifier vector in a unified format. This state identifier vector can serve as input to the target scheduling model, accurately describing the complete state of the current scheduling environment.

[0051] In a preferred embodiment, the above-described concatenation and mapping processes can be implemented, but are not limited to, in the following ways, to generate state label vectors suitable for processing by deep reinforcement learning models (such as deep Q-networks): First, perform the concatenation process. Combine the task feature vector v of any subtask u. uWith the node state vector v of a node j to be assigned j Perform direct splicing. Assume v u It is a p-dimensional vector, v j If a vector is q-dimensional, then concatenating them yields a (p+q)-dimensional original combined vector. This vector integrates all the raw environmental information required for the decision to "schedule subtask u to allocation node j". For a scheduling scenario with w subtasks and n nodes to be allocated, w*n such raw combination vectors will be generated, each corresponding to a possible (subtask, node) allocation pair.

[0052] Then, perform the mapping process. The original combined vector generated in the previous step... The input is fed into a pre-defined multi-layer fully connected neural network (i.e., a mapping network) for feature transformation. This mapping network can share parameters with the feature extraction layer in the cryptographic operation scheduling model, or it can be pre-trained as an independent preprocessing network.

[0053] Performing mapping processing serves at least the following two purposes: (1) Feature Extraction and Dimensionality Reduction: The original combined vector may have a high dimension (p+q) and contain a large number of redundant or nonlinear features. The mapping network automatically learns and extracts the high-order abstract features most relevant to the scheduling decision through the nonlinear activation function (such as ReLU) of its hidden layers, while compressing the vector dimension to a preset low-dimensional space (e.g., from 200 dimensions to 64 dimensions, the specific dimension can be adjusted according to the actual scenario). This can effectively alleviate the "curse of dimensionality" problem faced by subsequent reinforcement learning models during training, and improve learning efficiency and generalization ability.

[0054] (2) Numerical Normalization and Scale Unification: The original task features (such as computational cost) and node states (such as CPU utilization) often have different dimensions and numerical ranges. The mapping network can contain a batch normalization layer to uniformly adjust the numerical range of all input features to a standard normal distribution with a mean of 0 and a variance of 1, or within the interval [0,1]. This ensures that features of different magnitudes (e.g., a delay in nanoseconds and a trust score between 0 and 1) are treated fairly in subsequent Q-value calculations, avoiding model training instability or slow convergence caused by differences in numerical scale.

[0055] After the above mapping process, each original combined vector Ultimately, it is transformed into a state label vector with uniform dimensions and stable values. At any given scheduling moment, the system generates corresponding state identifier vectors for all possible (subtasks, nodes) assignment pairs, which together constitute the state space at the current scheduling moment, for the cryptographic operation scheduling model to make decisions.

[0056] In another embodiment, the concatenation and mapping processes described above can also be implemented using traditional machine learning methods: First, the concatenation process is performed to combine the task feature vector with the node state vector to form the original combined vector. Then, during the mapping process, Principal Component Analysis (PCA) is used to reduce the dimensionality of the original combined vector, extracting the top k principal components whose cumulative variance contribution rate reaches a preset threshold (e.g., 85%) to form the dimensionality-reduced feature vector. Simultaneously, Min-Max normalization is used to linearly transform all feature values ​​to the [0,1] interval to eliminate the influence of dimensions. After the above processing, each original combined vector is transformed into a dimensionally uniform and numerically stable state identifier vector for the cryptographic operation scheduling model to make decisions.

[0057] In step S103, the state identifier vector is substituted into the preset cryptographic operation scheduling model to form an instance scheduling model with specific parameters and variables for the current specific scheduling scenario, namely the target scheduling model.

[0058] Specifically, in step S103, based on the target scheduling model, the cryptographic operation tasks are allocated to obtain the optimal scheduling scheme, including: Using the target scheduling model, each subtask in the cryptographic operation task is matched with each node to be assigned, resulting in a set of scheduling schemes; Calculate the cumulative scheduling reward value for each scheduling scheme in the set of scheduling schemes; Using the target scheduling model, the optimal scheduling scheme is obtained by selecting the scheduling scheme with the largest cumulative reward value from the set of scheduling schemes.

[0059] In step S103, after the cryptographic task is divided into multiple subtasks and there are multiple nodes to be assigned, the target scheduling model is used, along with existing machine learning methods such as reinforcement learning algorithms or other improved reinforcement learning algorithms, to match all subtasks with all nodes to be assigned, generating allocation combinations between all subtasks and all nodes to be assigned, thus forming a set containing all potential scheduling strategies. Each element in this set represents a complete subtask allocation scheme.

[0060] For each specific scheduling scheme in the set of scheduling schemes, the cumulative reward that the scheme may obtain during execution is calculated according to the reward mechanism defined by the preset cryptographic operation scheduling model. This cumulative scheduling reward is a key indicator for evaluating the quality of a scheduling scheme, and its calculation typically considers node state information (especially dynamic trust scores), task characteristic information, and various evaluation parameters set in the scheduling model. For example, in a scheduling scheme, when a computationally intensive subtask is matched to a node with strong computational power and a high dynamic trust score, or when a subtask with extremely high security requirements is matched to a node with an extremely high dynamic trust score (i.e., a dynamic trust score greater than or equal to the average or greater than or equal to a preset penalty threshold), the scheduling scheme will calculate a high cumulative scheduling reward value; conversely, when a subtask with extremely high security requirements is matched to a node with a low dynamic trust score (i.e., a dynamic trust score less than the average or less than a preset penalty threshold), the scheduling scheme will calculate a low cumulative scheduling reward value.

[0061] After calculating the cumulative scheduling reward values ​​for all candidate scheduling schemes, the scheduling scheme with the highest cumulative scheduling reward value is identified by comparing these values. This scheme with the maximum reward value is considered the optimal scheduling scheme for the current situation. The purpose is to ensure that the selected scheme maximizes the cumulative scheduling reward parameter, thereby achieving efficient and secure scheduling of cryptographic tasks.

[0062] In step S103, after optimization calculation, the optimal scheduling scheme is finally selected as the cryptographic operation scheduling result information corresponding to the cryptographic operation task. This cryptographic operation scheduling result information has been rigorously optimized and verified, and can significantly improve the overall performance, security and resource utilization of cryptographic operation task scheduling.

[0063] As shown above, this cryptographic task scheduling method acquires the task feature information of the cryptographic task and the node state information of each node to be assigned. Based on the node state information, it calculates the dynamic trust score of each node to be assigned. Based on the dynamic trust score, combined with the task feature information and node state information, it uses a preset cryptographic task scheduling model for scheduling processing to obtain the cryptographic task scheduling result information. Thus, by inputting the task feature information of the cryptographic task, the node state information of each node to be assigned, and the dynamic trust score calculated based on the node state information into the preset cryptographic task scheduling model, the cryptographic task scheduling result information is calculated. This solves the problems of existing cryptographic task scheduling methods, such as insufficient consideration of task heterogeneity, inadequate awareness of node dynamic states, and lack of node trust considerations. It can achieve globally optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and increase the computational efficiency of cryptographic tasks.

[0064] refer to Figure 2 This application provides a cryptographic operation task scheduling device for scheduling cryptographic operation tasks, including: Module 1 is used to acquire task feature information of the cryptographic operation task and node status information of each node to be assigned; Calculation module 2 is used to calculate the dynamic trust score of each node to be assigned based on the node status information. The scheduling module 3 is used to perform scheduling processing based on dynamic trust scoring, combined with task feature information and node status information, using a preset cryptographic operation scheduling model to obtain the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

[0065] This cryptographic task scheduling device inputs the task feature information of the cryptographic task, the node state information of each node to be assigned, and the dynamic trust score calculated based on the node state information into a preset cryptographic task scheduling model to calculate the cryptographic task scheduling result information. It solves the problems of existing cryptographic task scheduling methods, such as failing to fully consider task heterogeneity, insufficient awareness of node dynamic state, and lack of consideration of node trust. It can achieve global optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and improve the computational efficiency of cryptographic tasks.

[0066] Specifically, during execution, module 1 acquires the task characteristic information of the cryptographic computation task and the node status information of each node to be assigned. The cryptographic computation task refers to a computational task requiring operations such as encryption, decryption, signing, and verification. Task characteristic information includes task type, computational load, data volume, priority, and deadline. This information can be obtained through user input during task submission or automatic system identification. Each node to be assigned refers to a computational unit in a distributed CPU cluster that can be used to execute the cryptographic computation task. Its node status information may include data such as CPU utilization, memory usage, network bandwidth, available storage space, current load, length of the pending queue within the TEE, network latency matrix, and monitoring metrics. This node status information can be collected in real-time by monitoring agents deployed on each node to be assigned.

[0067] Among them, monitoring indicator information refers to the indicator information of basic operation dimension (such as node liveness status, CPU utilization, memory utilization, etc., which are used to reflect node stability), security compliance dimension (such as task success rate, cryptographic operation error rate, TEE integrity proof verification results, abnormal access audit logs, etc., which are used to reflect security events and compliance), environmental security dimension (such as data center security level, network isolation domain, etc., which are used to reflect physical and network environment), and behavioral reputation dimension (such as key management audit logs, historical task completion records, etc., which are used to reflect historical collaboration reliability).

[0068] Specifically, when calculation module 2 calculates the dynamic trust score of each node to be assigned based on the node status information, it executes: The monitoring metric information of each node to be assigned is normalized to obtain the current monitoring metric parameters of each node to be assigned. For each node to be assigned, a dynamic trust score is obtained by dynamically scoring it based on the current monitoring indicator parameters and the preset indicator weights corresponding to the current monitoring indicator parameters.

[0069] During execution, calculation module 2 can employ existing methods such as min-max normalization and Z-score normalization to normalize the monitoring indicator information of each node to be assigned, obtaining the current monitoring indicator parameters for each node. This eliminates the influence of different dimensions between indicators and ensures the comparability of various indicators during the scoring process. Therefore, the current monitoring indicator parameters of each node can objectively reflect its performance across various dimensions.

[0070] Dynamic trust score is a dynamic numerical value derived from a quantitative assessment of the reliability, security, or performance of each node. It is calculated using a weighted fusion algorithm. This dynamic trust score reflects the trust level of a node in real-time or near real-time; a higher score generally indicates a more trustworthy node. By dynamically scoring each node using current monitoring indicator parameters and their corresponding preset indicator weights, the overall trustworthiness of the nodes can be quantified. The dynamic trust score can be calculated using the following formula: ; in, Let J be the dynamic trust score of node j in the m-th period (i.e., the trust score in the current period). Let J be the dynamic trust score of node j in the (m-1)th period (i.e., the trust score of the previous period). This is the historical attenuation coefficient, which can be set according to actual needs; The preset indicator weight is the corresponding indicator weight for the kth monitoring indicator information; This refers to the current monitoring indicator parameter (i.e., the indicator value after normalization) for the k-th monitoring indicator information in the m-th period. When m=1, This represents the default value of the dynamic trust score for node j, which is 0.

[0071] Specifically, the pre-defined cryptographic operation scheduling model is constructed through the following steps: Obtain the pre-set cryptographic operation scheduling evaluation parameters and their corresponding parameter weights; the cryptographic operation scheduling evaluation parameters include the security reward score calculated based on the current value of the dynamic trust score; Based on cryptographic operation scheduling evaluation parameters and parameter weights, a mathematical model containing security reward scores is constructed to obtain the initial scheduling reward function; With the aim of maximizing the cumulative reward parameter of the scheduling, a pre-defined cryptographic operation scheduling model is obtained by using a Markov decision process and constructing it based on the initial scheduling reward function.

[0072] Before constructing a cryptographic operation scheduling model, it is necessary to measure various metrics to evaluate the quality and effectiveness of cryptographic operation task scheduling. These evaluation parameters may include, but are not limited to, task completion time (representing the performance of the scheduling scheme), the standard deviation of the load on each node upon task completion (representing the load balancing performance of the scheduling scheme), security reward score (representing the security performance of the scheduling scheme), and data transmission volume between nodes (representing the communication cost performance of the scheduling scheme). For example, the shorter the task completion time, the better the scheduling effect. Parameter weights are used to represent the relative importance of each evaluation parameter in the overall scheduling objective. For example, in some scenarios, task completion time may be more critical than load balancing, so its weight will be set higher. These parameters and weights can be pre-set and adjusted according to actual application needs, system characteristics, and business strategies.

[0073] After defining the evaluation parameters and weights, these qualitative or quantitative indicators need to be transformed into a mathematical expression that includes a safety reward score, i.e., the initial scheduling reward function. This function aims to quantify the immediate benefits of each scheduling decision. For example, the initial scheduling reward function can be constructed as a weighted sum model, where each evaluation parameter is multiplied by its corresponding weight, and then all weighted parameters are summed to obtain a comprehensive reward value. The higher the reward value, the better the current scheduling decision. This mathematical model provides a basic evaluation framework for subsequent optimization. The safety reward score serves as a key bonus factor to ensure that scheduling decisions fully consider safety factors while pursuing efficiency.

[0074] The initial scheduling reward function is as follows: ; in, The immediate scheduling reward for scheduling scheme i; Let i be the task completion time for scheduling scheme i; Let be the standard deviation of the load on each node when scheduling scheme i is completed. The load on each node when scheduling scheme i is completed; The safety reward score for scheduling scheme i; This represents the amount of data transmission between nodes when scheduling scheme i is completed. The parameter weights corresponding to the task completion time; The parameter weights are the standard deviations of the load. The parameter weights corresponding to the security rewards; The parameter weights are the data transmission volume.

[0075] The security reward score can be calculated based on the node's current trust score (i.e., the current value of the dynamic trust score). The specific formula for calculating the security reward score is as follows: ; Where n is the total number of nodes to be assigned, j≤n; For safety reward weighting coefficient; This is a boolean value indicating whether node j executes the task within scheduling scheme i. When node j executes the task within scheduling scheme i... When node j is not executing a task within scheduling scheme i, ; The penalty base represents the security violation penalty item, which is used to impose additional penalties when the scheduling scheme violates security constraints. It can be set according to the actual situation. For example, it can be set according to the dynamic trust score, comparing the trust score of the current period with the preset penalty threshold to determine the number of nodes with a trust score lower than the preset penalty threshold. Different threshold ranges are set according to this number, and the corresponding penalty base is set for the corresponding threshold range.

[0076] Specifically, with the aim of maximizing the cumulative scheduling reward parameter, a Markov decision process is used to construct a pre-defined cryptographic operation scheduling model based on the initial scheduling reward function, including: Using Markov decision processes, we set quintuple parameters, learning parameters, and cumulative scheduling reward parameters for the initial scheduling reward function; Based on the quintuple parameters and the learned parameters, the relationship between the initial scheduling reward function and the cumulative scheduling reward parameters is constructed, and the calculation expression for the cumulative scheduling reward parameters is obtained. With the aim of maximizing the cumulative reward parameter of the scheduling, the calculation expression of the cumulative reward parameter of the scheduling is adjusted to obtain the cryptographic operation scheduling model.

[0077] In constructing a cryptographic task scheduling model, the quintuple parameters are a core component of the Markov Decision Process (MDP), typically including the state set (S), action set (A), state transition probabilities (P), reward function (R), and discount factor (γ). Specifically, the state set S represents all possible states during the cryptographic task scheduling process, such as a combination of the current task's features and the node state information of the node to be assigned; the action set A represents all scheduling actions that can be taken in each state, such as assigning a subtask to a node; the state transition probability P describes the probability that the system will transition to the next state after taking an action in a given state; the reward function R defines the immediate reward obtained by taking a specific action in a specific state, corresponding to the initial scheduling reward function mentioned above; and the discount factor γ is used to measure the importance of future rewards. The aim is to provide a clear decision-making environment model for reinforcement learning algorithms by clarifying these fundamental elements.

[0078] Learning parameters are key parameters used to control the learning process in reinforcement learning algorithms, such as the learning rate (α) and the exploration rate (ε). The learning rate α determines the degree to which new information corrects old information each time the Q-table is updated; the exploration rate ε controls the balance between exploring unknown actions and utilizing known optimal actions. By setting these parameters appropriately, reinforcement learning algorithms can be effectively guided to achieve a balance between exploration and utilization, thereby accelerating the learning process and improving convergence.

[0079] The table of cumulative scheduling reward parameters (i.e., the Q-table) is used in reinforcement learning to store the value of state-action pairs. Each cumulative scheduling reward value Q(s, a) in the Q-table represents the expected future cumulative reward that can be obtained by taking action a in state s. The purpose is that, by continuously updating the Q-table, reinforcement learning algorithms can learn which actions to take in different states to maximize long-term cumulative rewards, thereby forming an optimal scheduling strategy. In practical applications, the cumulative scheduling reward parameters can be obtained by training tabular Q-Learning and Deep Q-Networks (DQN) based on historical data or real-time data (i.e., historical records based on task feature information and node state information, or using existing real-time task feature information and real-time node state information of nodes).

[0080] Based on the quintuple parameters and the learned parameters, a correlation is constructed between the initial scheduling reward function and the cumulative scheduling reward parameters, resulting in an expression for calculating the cumulative scheduling reward parameters. For example, this expression can be constructed based on the Bellman equation to describe the iterative relationship between Q values, i.e.: ; in, In the m-th period, the state is... Select action The achievable cumulative scheduling reward parameter (i.e., the cumulative scheduling reward value, usually referred to as the Q value); In the (m-1)th period, the state is... Select action The cumulative scheduling reward parameters that can be obtained; The learning rate; Select action for the m-th cycle The immediate reward that can be obtained afterward; This is a discount factor (which controls the weight of future rewards and indicates the degree of importance attached to future rewards). This refers to the new state entered in the m-1th cycle after performing the action in the m-1th cycle. This refers to a new action that is performed in the m-1th cycle after the action in the m-1th cycle. To enter a new state in the m-th period Then take the maximum Q value from all possible new actions; when m=1, This is the initial value for scheduling the accumulated reward value. To enter a new state at the initial moment Then take the maximum Q value among all possible new actions. and It can be calculated using tabular Q-Learning and deep Q-networks. The purpose of the expression for calculating the cumulative reward parameter for scheduling is to gradually approximate the optimal Q value through iterative calculation.

[0081] The goal is to maximize the cumulative reward parameter for scheduling. This is achieved by iteratively updating the Q-values ​​in the Q-table until each Q-value eventually converges to the maximum cumulative reward obtainable by taking the corresponding action in the corresponding state. The expression for calculating the cumulative reward parameter is then adjusted to obtain a cryptographic operation scheduling model that guides optimal scheduling. The cryptographic operation scheduling model is as follows: ; in, This is a cryptographic operation scheduling model; for The function is used to select the action that maximizes the Q value from all possible actions as the scheduling scheme for this time.

[0082] Specifically, when scheduling module 3 obtains the cryptographic operation scheduling result information corresponding to the cryptographic operation task by performing scheduling processing based on dynamic trust scoring, combined with task feature information and node status information, and using a preset cryptographic operation scheduling model, it executes: Encode the task feature information, node status information, and dynamic trust score to obtain a status identifier vector; The state identifier vector is input into the preset cryptographic operation scheduling model to obtain the target scheduling model; Based on the target scheduling model, task allocation for cryptographic operations is performed to obtain the optimal scheduling scheme. The optimal scheduling scheme is determined by the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

[0083] Specifically, when scheduling module 3 encodes task feature information, node status information, and dynamic trust score to obtain a status identifier vector, it executes: The cryptographic task is divided into multiple subtasks, and the task feature information of each subtask is transformed into a corresponding task feature vector. The node status information and dynamic trust score of each node to be assigned are converted into the corresponding node status vector. The task feature vector and node state vector are concatenated and mapped to obtain the state identifier vector.

[0084] During execution, scheduling module 3 decomposes a large cryptographic task into several smaller, more manageable subtasks based on attributes such as complexity, computational load, security level, or execution stage. For example, a complex encryption process may be divided into multiple subtasks such as key generation, data encryption, and signature verification. This division aims to achieve finer-grained scheduling and resource allocation, improving the parallelism and efficiency of task execution. After obtaining multiple subtasks, the task feature information of each subtask is quantified and vectorized to obtain the corresponding task feature vector. This vector form facilitates subsequent processing and analysis.

[0085] The node status information and dynamic trust score of each node to be assigned are quantified and vectorized to integrate and transform them into a unified node status vector. This node status vector can comprehensively characterize the current operating status and trustworthiness of the node to be assigned.

[0086] The concatenation process combines the feature vector of a subtask with the node state vector of the node to be assigned, forming a longer vector that comprehensively reflects the potential state of a specific subtask executed on a specific node. The mapping process performs dimensionality transformation, normalization, or feature extraction on the concatenated vector to generate a state identifier vector in a unified format. This state identifier vector can serve as input to the target scheduling model, accurately describing the complete state of the current scheduling environment.

[0087] In a preferred embodiment, the above-described concatenation and mapping processes can be implemented, but are not limited to, in the following ways, to generate state label vectors suitable for processing by deep reinforcement learning models (such as deep Q-networks): First, perform the concatenation process. Combine the task feature vector v of any subtask u. u With the node state vector v of a node j to be assigned j Perform direct splicing. Assume v u It is a p-dimensional vector, v j If a vector is q-dimensional, then concatenating them yields a (p+q)-dimensional original combined vector. This vector integrates all the raw environmental information required for the decision to "schedule subtask u to allocation node j". For a scheduling scenario with w subtasks and n nodes to be allocated, w*n such raw combination vectors will be generated, each corresponding to a possible (subtask, node) allocation pair.

[0088] Then, perform the mapping process. The original combined vector generated in the previous step... The input is fed into a pre-defined multi-layer fully connected neural network (i.e., a mapping network) for feature transformation. This mapping network can share parameters with the feature extraction layer in the cryptographic operation scheduling model, or it can be pre-trained as an independent preprocessing network.

[0089] Performing mapping processing serves at least the following two purposes: (1) Feature Extraction and Dimensionality Reduction: The original combined vector may have a high dimension (p+q) and contain a large number of redundant or nonlinear features. The mapping network automatically learns and extracts the high-order abstract features most relevant to the scheduling decision through the nonlinear activation function (such as ReLU) of its hidden layers, while compressing the vector dimension to a preset low-dimensional space (e.g., from 200 dimensions to 64 dimensions, the specific dimension can be adjusted according to the actual scenario). This can effectively alleviate the "curse of dimensionality" problem faced by subsequent reinforcement learning models during training, and improve learning efficiency and generalization ability.

[0090] (2) Numerical Normalization and Scale Unification: The original task features (such as computational cost) and node states (such as CPU utilization) often have different dimensions and numerical ranges. The mapping network can contain a batch normalization layer to uniformly adjust the numerical range of all input features to a standard normal distribution with a mean of 0 and a variance of 1, or within the interval [0,1]. This ensures that features of different magnitudes (e.g., a delay in nanoseconds and a trust score between 0 and 1) are treated fairly in subsequent Q-value calculations, avoiding model training instability or slow convergence caused by differences in numerical scale.

[0091] After the above mapping process, each original combined vector Ultimately, it is transformed into a state label vector with uniform dimensions and stable values. At any given scheduling moment, the system generates corresponding state identifier vectors for all possible (subtasks, nodes) assignment pairs, which together constitute the state space at the current scheduling moment, for the cryptographic operation scheduling model to make decisions.

[0092] In another embodiment, the concatenation and mapping processes described above can also be implemented using traditional machine learning methods: First, the concatenation process is performed to combine the task feature vector with the node state vector to form the original combined vector. Then, during the mapping process, Principal Component Analysis (PCA) is used to reduce the dimensionality of the original combined vector, extracting the top k principal components whose cumulative variance contribution rate reaches a preset threshold (e.g., 85%) to form the dimensionality-reduced feature vector. Simultaneously, Min-Max normalization is used to linearly transform all feature values ​​to the [0,1] interval to eliminate the influence of dimensions. After the above processing, each original combined vector is transformed into a dimensionally uniform and numerically stable state identifier vector for the cryptographic operation scheduling model to make decisions.

[0093] When the scheduling module 3 is executed, it substitutes the state identifier vector into the preset cryptographic operation scheduling model to form an instance scheduling model with specific parameters and variables for the current specific scheduling scenario, namely the target scheduling model.

[0094] Specifically, when scheduling module 3 allocates tasks for cryptographic operations based on the target scheduling model and obtains the optimal scheduling scheme, it executes: Using the target scheduling model, each subtask in the cryptographic operation task is matched with each node to be assigned, resulting in a set of scheduling schemes; Calculate the cumulative scheduling reward value for each scheduling scheme in the set of scheduling schemes; Using the target scheduling model, the optimal scheduling scheme is obtained by selecting the scheduling scheme with the largest cumulative reward value from the set of scheduling schemes.

[0095] When scheduling module 3 executes, after the cryptographic task is divided into multiple subtasks and there are multiple nodes to be assigned, it uses the target scheduling model and existing machine learning methods such as reinforcement learning algorithms or other improved reinforcement learning algorithms to match all subtasks with all nodes to be assigned, thereby generating all allocation combinations between all subtasks and all nodes to be assigned, thus forming a set containing all potential scheduling strategies. Each element in this set represents a complete subtask allocation scheme.

[0096] For each specific scheduling scheme in the set of scheduling schemes, the cumulative reward that the scheme may obtain during execution is calculated according to the reward mechanism defined by the preset cryptographic operation scheduling model. This cumulative scheduling reward is a key indicator for evaluating the quality of a scheduling scheme, and its calculation typically considers node state information (especially dynamic trust scores), task characteristic information, and various evaluation parameters set in the scheduling model. For example, in a scheduling scheme, when a computationally intensive subtask is matched to a node with strong computational power and a high dynamic trust score, or when a subtask with extremely high security requirements is matched to a node with an extremely high dynamic trust score (i.e., a dynamic trust score greater than or equal to the average or greater than or equal to a preset penalty threshold), the scheduling scheme will calculate a high cumulative scheduling reward value; conversely, when a subtask with extremely high security requirements is matched to a node with a low dynamic trust score (i.e., a dynamic trust score less than the average or less than a preset penalty threshold), the scheduling scheme will calculate a low cumulative scheduling reward value.

[0097] After calculating the cumulative scheduling reward values ​​for all candidate scheduling schemes, the scheduling scheme with the highest cumulative scheduling reward value is identified by comparing these values. This scheme with the maximum reward value is considered the optimal scheduling scheme for the current situation. The purpose is to ensure that the selected scheme maximizes the cumulative scheduling reward parameter, thereby achieving efficient and secure scheduling of cryptographic tasks.

[0098] During execution, scheduling module 3, after optimization calculation, finally selects the optimal scheduling scheme as the cryptographic operation scheduling result information corresponding to the cryptographic operation task. This cryptographic operation scheduling result information has been rigorously optimized and verified, and can significantly improve the overall performance, security and resource utilization of cryptographic operation task scheduling.

[0099] As can be seen from the above, this cryptographic task scheduling device acquires the task feature information of the cryptographic task and the node status information of each node to be assigned. Based on the node status information, it calculates the dynamic trust score of each node to be assigned. Based on the dynamic trust score, combined with the task feature information and node status information, it uses a preset cryptographic task scheduling model for scheduling processing to obtain the cryptographic task scheduling result information. Thus, by inputting the task feature information of the cryptographic task, the node status information of each node to be assigned, and the dynamic trust score calculated based on the node status information into the preset cryptographic task scheduling model, the cryptographic task scheduling result information is calculated. This solves the problems of existing cryptographic task scheduling methods, such as insufficient consideration of task heterogeneity, inadequate awareness of node dynamic status, and lack of node trust consideration. It can achieve globally optimal scheduling of cryptographic tasks, significantly reduce time costs, improve computational security, and increase the computational efficiency of cryptographic tasks.

[0100] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device includes a processor 301 and a memory 302. The processor 301 and the memory 302 are interconnected and communicate with each other via a communication bus 303 and / or other connection mechanisms (not shown). The memory 302 stores a computer program executable by the processor 301. When the electronic device is running, the processor 301 executes the computer program to perform the cryptographic task scheduling method in any optional implementation of the above embodiments, to achieve the following functions: obtaining task feature information of the cryptographic task and node status information of each node to be assigned, the node status information including the dynamic trust score of the node to be assigned; obtaining a pre-constructed cryptographic scheduling model, the cryptographic scheduling model being constructed with the aim of maximizing the cumulative scheduling reward parameter, the cumulative scheduling reward parameter being determined based on the dynamic trust score; and, based on the dynamic trust score, combined with the task feature information and node status information, performing scheduling processing using the pre-constructed cryptographic scheduling model to obtain the cryptographic scheduling result information corresponding to the cryptographic task.

[0101] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it executes the cryptographic task scheduling method in any optional implementation of the above embodiments to achieve the following functions: obtaining task feature information of the cryptographic task and node status information of each node to be assigned, the node status information including the dynamic trust score of the node to be assigned; obtaining a pre-constructed cryptographic scheduling model, the cryptographic scheduling model being constructed with the aim of maximizing the cumulative scheduling reward parameter, the cumulative scheduling reward parameter being determined based on the dynamic trust score; and, based on the dynamic trust score, combined with the task feature information and node status information, performing scheduling processing using the pre-constructed cryptographic scheduling model to obtain the cryptographic scheduling result information corresponding to the cryptographic task. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0102] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0103] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0104] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0105] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0106] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for scheduling cryptographic computation tasks, characterized in that, Including the following steps: Obtain the task characteristic information of the cryptographic operation task and the node status information of each node to be assigned; The dynamic trust score of each node to be assigned is calculated based on the node status information. Based on the dynamic trust score, combined with the task feature information and the node status information, a preset cryptographic operation scheduling model is used for scheduling processing to obtain the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

2. The cryptographic task scheduling method according to claim 1, characterized in that, The node status information includes monitoring indicator information; based on the node status information, a dynamic trust score is calculated for each node to be assigned, including: The monitoring indicator information of each node to be assigned is normalized to obtain the current monitoring indicator parameters of each node to be assigned. For each node to be assigned, a dynamic trust score is obtained by dynamically scoring based on the current monitoring indicator parameters and the preset indicator weights corresponding to the current monitoring indicator parameters.

3. The cryptographic operation task scheduling method according to claim 1, characterized in that, The preset cryptographic operation scheduling model is constructed through the following steps: Obtain pre-set cryptographic operation scheduling evaluation parameters and their corresponding parameter weights; the cryptographic operation scheduling evaluation parameters include a security reward score calculated based on the current value of the dynamic trust score; Based on the cryptographic operation scheduling evaluation parameters and the parameter weights, a mathematical model containing the security reward score is constructed to obtain the initial scheduling reward function; With the aim of maximizing the cumulative scheduling reward parameter, a Markov decision process is used to construct the preset cryptographic operation scheduling model based on the initial scheduling reward function.

4. The cryptographic operation task scheduling method according to claim 3, characterized in that, With the aim of maximizing the cumulative scheduling reward parameter, a Markov decision process is used to construct the preset cryptographic operation scheduling model based on the initial scheduling reward function, including: Using a Markov decision process, the initial scheduling reward function is set with a quintuple parameter, a learning parameter, and a cumulative scheduling reward parameter. Based on the quintuple parameters and the learning parameters, the correlation between the initial scheduling reward function and the cumulative scheduling reward parameters is constructed to obtain the calculation expression for the cumulative scheduling reward parameters; With the aim of maximizing the cumulative scheduling reward parameter, the calculation expression of the cumulative scheduling reward parameter is adjusted to obtain the preset cryptographic operation scheduling model.

5. The cryptographic operation task scheduling method according to claim 1, characterized in that, Based on the dynamic trust score, combined with the task feature information and the node status information, a preset cryptographic operation scheduling model is used for scheduling processing to obtain the cryptographic operation scheduling result information corresponding to the cryptographic operation task, including: The task feature information, the node status information, and the dynamic trust score are encoded to obtain a status identifier vector; The state identifier vector is input into the preset cryptographic operation scheduling model to obtain the target scheduling model; Based on the target scheduling model, the cryptographic operation tasks are allocated to obtain the optimal scheduling scheme; The optimal scheduling scheme is determined as the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

6. The cryptographic operation task scheduling method according to claim 5, characterized in that, The task feature information, the node status information, and the dynamic trust score are encoded to obtain a status identifier vector, including: The cryptographic operation task is divided into multiple subtasks, and the task feature information of each subtask is converted into a corresponding task feature vector. The node state information and dynamic trust score of each node to be assigned are converted into a corresponding node state vector. The task feature vector and the node state vector are concatenated and mapped to obtain the state identifier vector.

7. The cryptographic operation task scheduling method according to claim 6, characterized in that, Based on the target scheduling model, the cryptographic operation tasks are allocated to obtain the optimal scheduling scheme, including: Using the target scheduling model, each subtask in the cryptographic operation task is matched with each node to be assigned, resulting in a set of scheduling schemes; Calculate the cumulative scheduling reward value for each scheduling scheme in the set of scheduling schemes; The optimal scheduling scheme is obtained by selecting the scheduling scheme with the largest cumulative reward value from the set of scheduling schemes.

8. A cryptographic operation task scheduling device, characterized in that, include: The acquisition module is used to acquire the task characteristic information of the cryptographic operation task and the node status information of each node to be assigned; The calculation module is used to calculate the dynamic trust score of each node to be assigned based on the node status information. The scheduling module is used to perform scheduling processing based on the dynamic trust score, combined with the task feature information and the node status information, using a preset cryptographic operation scheduling model to obtain the cryptographic operation scheduling result information corresponding to the cryptographic operation task.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program executable by the processor, and when the processor executes the computer program, it performs the steps of the cryptographic task scheduling method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it performs the steps in the cryptographic task scheduling method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Password resource allocation method and system based on attention mechanism and residual network

    CN119728107A

  • Distributed computing task scheduling security guarantee method and system based on dynamic trust evaluation

    CN121050858A

  • Task scheduling method and system based on dynamic trust degree and data security level self-adaption

    CN121349712A

  • Method and apparatus for task scheduling based on deep reinforcement learning, and device

    US20210081787A1