Workflow adaptive multistage cluster scheduling method and system based on privacy perception

Through the privacy-aware workflow adaptive multi-level cluster scheduling method, the problems of data privacy security and resource scheduling in the workflow cluster are solved, and secure and efficient computing is achieved on distributed nodes, ensuring the security and computing efficiency of private data.

CN120371475APending Publication Date: 2025-07-25TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510460786.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

During the workflow cluster processing process, how to achieve efficient execution of computing tasks while ensuring data privacy and security, especially reasonable scheduling of resources on distributed nodes to avoid server load imbalance and resource waste.

Method used

Adaptive multi-level cluster scheduling method based on privacy awareness is adopted to achieve safe and efficient execution of tasks on distributed nodes through task splitting, dependency establishment, task priority determination with the earliest completion time minimized, load balancing scheduling between servers and within servers, and dynamic adjustment of simulated annealing algorithm.

Benefits of technology

It realizes 100% secure domain computing for privacy tasks, maximizes adaptive load balancing, reduces task execution time, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371475A_ABST
    Figure CN120371475A_ABST
Patent Text Reader

Abstract

The invention discloses a workflow adaptive multistage cluster scheduling method and system based on privacy perception, which comprehensively considers the privacy priority, dependency relationship and computing resource availability of a workflow task, coordinates and manages the adaptive multistage operation of the computing task on distributed nodes, and ensures safety and high efficiency. The workflow adaptive multi-stage cluster scheduling method based on privacy perception comprises the steps of task splitting and modeling, and adaptive multi-stage cluster load balance scheduling based on task priority determination of minimization of earliest completion time. Practice proves that the multi-level self-adaptive efficient scheduling of the computing task can be realized on the premise of ensuring privacy security by applying the embodiment of the invention.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the field of computer data processing and distributed computing technology, and in particular to the scheduling of computing tasks and privacy protection and trusted computing technologies in the process of workflow cluster processing. [Background technology]

[0002] In today's digital age, data has become a key production factor that drives the development of various industries. With the explosive growth of data volume and the increasing demand for computing efficiency, distributed computing technology has emerged and has been widely used. With its powerful parallel computing capabilities, the workflow cluster processing mode can efficiently handle large-scale complex computing tasks and plays an important role in many fields such as scientific research, financial analysis, and industrial manufacturing.

[0003] However, workflow cluster processing faces many challenges. On the one hand, data privacy and security issues are becoming increasingly serious. The data involved in computing tasks contains a large amount of sensitive information, such as personal identity data, business secrets, medical records, etc. Once these private data are leaked, it will not only cause huge losses to the data owner, but may also cause serious social problems. Traditional security protection methods are difficult to fully guarantee data security in a distributed, multi-node workflow cluster environment. How to ensure that private data is not leaked during the computing process has become a problem that needs to be solved urgently.

[0004] On the other hand, the reasonable scheduling of computing resources is crucial. Different computing tasks have very different requirements for resources, while the server resources in the cluster are limited and heterogeneous. If the scheduling is unreasonable, it is easy for some servers to be overloaded and inefficient, while some server resources are idle. This will not only reduce the computing performance of the entire cluster, but also cause resource waste and increase operating costs. In addition, there are complex dependencies between workflow tasks. How to comprehensively consider task dependencies, resource allocation, and privacy protection to achieve efficient execution of computing tasks on distributed nodes is a research hotspot and difficulty in the field of workflow cluster scheduling.

[0005] The existing workflow scheduling methods have certain limitations when dealing with the above problems. Some scheduling methods only focus on the execution efficiency of tasks and ignore privacy protection, which increases data security risks; while some solutions that focus on privacy protection often sacrifice computing performance and cannot meet the needs of practical applications for efficient computing. In terms of server resource scheduling, most of the existing load balancing algorithms are static or based on simple rules, which are difficult to adapt to dynamically changing computing tasks and complex cluster environments, and cannot achieve optimal resource allocation. Therefore, it is of great practical significance and application value to develop a method that can coordinate and manage the safe, efficient and adaptive execution of computing tasks on distributed nodes.

Summary of the Invention

[0006] In view of the above problems, the present invention realizes a privacy-aware workflow adaptive multi-level cluster scheduling method, which can coordinate and manage the secure, efficient and adaptive execution of workflow tasks on distributed nodes.

[0007] The present invention provides a privacy-aware workflow adaptive multi-level cluster scheduling method, including the following steps:

[0008] S1. Obtain workflow tasks, split the workflow tasks into a series of privacy tasks and regular tasks, and establish the dependency relationship of the split subtasks according to the workflow task dependencies;

[0009] S2. Based on the task execution time and the dependency relationship of the tasks, generate a task execution list based on the minimization of the earliest completion time;

[0010] S3. Obtain the highest-priority task in the task execution list, perform privacy protection comprehensive scheduling among servers, and allocate a target server for the highest-priority task;

[0011] S4. Perform fine-grained load balancing scheduling inside the server according to the load balancing conditions of the TEE side and the REE side of the target server and the privacy of the highest-priority task;

[0012] S5. Based on the simulated annealing algorithm, perform multi-level linkage dynamic adjustment of the highest-priority task among servers and inside the server according to the fine-grained load balancing scheduling result to ensure load balancing and privacy protection, delete the highest-priority task that has completed scheduling from the task execution list, and repeat steps S3 - S5 until the task execution list is empty.

[0013] For the identification of privacy tasks: This step mainly relies on the methods of manual marking and keyword recognition. Either relevant professionals can mark whether the tasks in the workflow cluster are privacy tasks in turn, or machine recognition can be performed by setting keywords containing privacy information, and mark the tasks covering privacy information as privacy tasks.

[0014] For the splitting and modeling of tasks: This step first splits the workflow into a series of privacy tasks and regular tasks according to the privacy recognition results, and then establishes a DAG graph of the split subtasks according to the workflow task dependencies.

[0015] Preferably, a part of the present invention provides a method for generating a task execution list based on the minimization of the earliest completion time, including the following steps:

[0016] Obtain the computing costs of different tasks on different servers;

[0017] Obtain the communication cost between tasks with dependencies;

[0018] Based on the calculated cost and communication cost, construct a task priority function and generate a task execution list.

[0019] Preferably, a part of the present invention provides a method for determining task priority based on minimizing the earliest finish time (EFT), and the method includes the following steps:

[0020] Establish a calculation cost matrix: Assume that the goal of this scheduling is to schedule n tasks to p servers, then a matrix of size n×p needs to be established to estimate the running time of different tasks on different servers, so as to represent the calculation cost by the running time.

[0021] Establish a list of communication costs between computing tasks: Assume that there are n tasks in the scheduling process, and a list is used for each task to store the time required for data transfer of tasks with dependencies on it, and its data transfer time Use the inter-task data transfer time c i,j to represent the communication cost between tasks.

[0022] The c i,j is the time required for cross-server transmission between task i and task j, data i,j refers to the amount of data transferred between task i and task j, and bandwidth refers to the average bandwidth between servers.

[0023] Determine task priority: Comprehensively consider the time cost required for tasks to be transmitted to the server and the communication time cost between tasks with dependencies, that is, the task calculation cost and the task communication cost, and use the task priority calculation formula to calculate the ranking of tasks, and then determine the task priority according to the descending order of the sorting results to generate a task execution list H.

[0024] The u i is the i-th task, p k is the k-th server, rank(u i ) represents the priority of task u i , the larger the rank, the higher the priority, represents the running time of task u i on server p k , P represents the set of servers, p represents the number of servers, reflects the average running time of task u i in the server cluster P, u j is the successor node of u i , c i,j represents task u iand u j The time required for data transfer indicating the selected task u i All successor nodes u j in c ij + the maximum value of rank(u j )

[0025] Another part of the present invention lies in providing an inter - server privacy - protected comprehensive scheduling, including the following steps:

[0026] Calculate the data transfer time, calculation time, and / or encryption / decryption time for the highest - priority task running on different servers to obtain the calculation resource affinity value;

[0027] Calculate the waiting time for the highest - priority task to start running on different servers to obtain the server load - balancing value;

[0028] Based on user requirements and the current computing environment, construct a priority function according to the calculation resource affinity value and the load - balancing value and allocate a target server for the highest - priority task.

[0029] Preferably, obtain the highest - priority task and perform server allocation on it, so that the task with the greatest impact on the final completion time runs on the server first, thereby ensuring the shortest final completion time.

[0030] Another part of the present invention lies in providing an inter - server privacy - protected comprehensive scheduling, including the following steps:

[0031] Calculation of the resource affinity value: Calculate the sum of the data transfer time and the calculation time for the current highest - priority task u i running on different servers p k and / or the data encryption / decryption time t which represents the resource affinity value between the i - th task and the k - th server; s

[0032] Calculation of the load - balancing value: Calculate the waiting time required for the current highest - priority task u i to start execution on different servers p k which represents the load - balancing value of the k - th server;

[0033] Selecting the target server: Based on user requirements and the current computing environment, calculate the comprehensive evaluation value of the current highest - priority task u for the k - th server according to the formula, select the server corresponding to the minimum AMEM value as the target server, and allocate the highest - priority task u i to it i ​​, where α + β = 1, and the initial values of α and β are both 0.5.

[0034] Another part of the present invention lies in providing a fine-grained load balancing scheduling within a server, including the following steps:

[0035] Obtain the comprehensive scheduling calculation tasks for privacy protection among servers, the task queues and load conditions of the TEE side and REE side within the server;

[0036] Determine whether the current calculation task is a privacy task;

[0037] If so, the current calculation task enters the TEE side for calculation; if not, the current calculation task enters the REE side for calculation;

[0038] Collect the calculation results after the calculation ends.

[0039] Another part of the present invention lies in providing a method for determining whether the current calculation task is a privacy task; if so, the current calculation task enters the TEE side for calculation, including the following steps:

[0040] Determine the load L of the TEE side of the current server TEE whether it exceeds the maximum value T TEEt ;

[0041] If L TEE < T TEEt , then the current calculation task enters the TEE side of the current server for calculation;

[0042] If L TEE > T TEEt , obtain the waiting time T for calculation on the current server wait , and the data encryption and decryption time and data replication time T required to copy it to the server with the lowest load com ;

[0043] If T wait < T com , select to execute on the current server and add it to the TEE side task queue of the current server; if T wait > T com , allocate the current calculation task to the server with the lowest load for calculation, and encrypt and copy the required data to the TEE side of the server with the lowest load.

[0044] Another part of the present invention lies in providing a method for determining whether the current calculation task is a privacy task; if not, the current calculation task enters the REE side for calculation, including the following steps:

[0045] Determine the load L of the REE side of the current server REE whether it exceeds the maximum value T REEt ;

[0046] If L REE > T REEt , determine whether the current server TEE side exceeds the maximum load value T TEEt ;

[0047] If L TEE < T TEEt , then the current computing task enters the current server TEE side for computing; If L TEE > T TEEt , then allocate the data required for the current computing task to the processor with the lowest current load through data replication, and re - allocate the task;

[0048] If L REE < T REEt , then determine whether the current server TEE side and REE side meet the load balance;

[0049] If the load balance is met, that is, |L REE - L TEE | ≤ T upper , then the computing task enters the REE side for computing;

[0050] Otherwise, the computing task enters the TEE side for computing;

[0051] Among them, T upper is the maximum load balance difference.

[0052] Preferably, use the simulated annealing algorithm to achieve parameter adaptive adjustment: Use the simulated annealing algorithm to achieve adaptive adjustment of parameters α and β, determine the optimal parameters α and β to adapt to the load situation, and avoid the extension of the execution time due to load imbalance.

[0053] Preferably, the method of using the simulated annealing algorithm to achieve parameter adaptive adjustment includes the following steps:

[0054] Generate a new solution by perturbation: Generate a new solution by adding a random number that follows a normal distribution with a mean of 0 and a standard deviation of 0.1, that is, α new = α + ε1, β new = β + ε2.

[0055] The ε1 and ε2 are random numbers that satisfy the above normal distribution, and at the same time ensure that α new + β new = 1.

[0056] Calculate the new fitness function value: Calculate the new fitness function value F new and β new based on the parameters α new (α new , β new ).

[0057] Judge the new fitness function value F new (α new , β new ) whether it is greater than the old fitness function value

[0058] F old (α, β): The fitness function F(α, β) reflects the load balancing situation. The larger the fitness function value, the more balanced the load of the server cluster is.

[0059] If the judgment is yes, it means that the new parameters are more inclined to load balancing, then adjust α to α new , and adjust β to β new .

[0060] If the judgment is no, it means that the old parameters are more inclined to load balancing, but in order to jump out of the local optimum, decide whether to accept the new solution according to the probability , generate a random number r between (0, 1), and judge whether r is less than P;

[0061] If r < P, then adjust ɑ to α new , and adjust β to β new .

[0062] If r > P, then ɑ and β remain unchanged.

[0063] Adjust T = ωT: Reduce the probability of selecting the wrong solution by cooling down.

[0064] Repeat the above steps until T < T end .

[0065] Preferably, based on the simulated annealing algorithm, the highest-priority tasks perform multi-level linkage dynamic adjustment between and within servers to ensure load balancing and privacy protection, including the following steps:

[0066] Define the fitness function Calculate the fitness function value F old (α, β); where p is the total number of servers, L i represents the load situation of the i-th server, represents the average load of p servers, set the optimal initial temperature T, termination temperature T end , cooling rate ω, 0 < ω < 1, initially α = 0.5, β = 0.5;

[0067] Perturbation to generate a new solution: Add a random number that follows a normal distribution with a mean of 0 and a standard deviation of 0.1 to generate a new solution, that is, α new = α + ε1, β new = β + ε2, where ε1 and ε2 are random numbers that satisfy the above normal distribution, α new + βnew = 1;

[0068] Calculate the new fitness function value: Based on the parameters α new and β new Calculate the new fitness function value F new (α new , β new );

[0069] Judge whether the new fitness function value F new (α new , β new ) is greater than the old fitness function value F old (α, β);

[0070] If so, adjust α to α new , and adjust β to β new ;

[0071] If not, decide whether to accept the new solution according to the probability Generate a random number r between (0, 1), and judge whether r is less than P;

[0072] If r < P, adjust α to α new , and adjust β to β new ;

[0073] If r > P, α and β remain unchanged;

[0074] Adjust T = ωT;

[0075] Repeat the above steps until T < T end .

[0076] Preferably, in the server multi-level scheduling strategy, the scheduling selections between two levels affect each other. After each fine scheduling within the server is completed, it is necessary to dynamically adjust the inter-server scheduling method according to the replication and transfer situation of the data within the server. By changing the proportion of load balancing in the AMEM formula between servers, the overall running time extension caused by load imbalance can be effectively avoided.

[0077] The inter-server privacy protection comprehensive scheduling, this method includes the following steps:

[0078] Calculate the resource affinity value: Estimate the sum of the data transfer time and execution time required for the currently selected highest-priority task u i to run on different servers p k respectively. If this task is a privacy task, the encryption and decryption time t of the data needs to be additionally added s , and use the calculation result to represent the resource affinity value of the highest-priority task for different servers.

[0079] Preferably, for the fine-grained load balancing in the server, the method includes the following steps:

[0080] Obtain the privacy protection comprehensive scheduling calculation task between servers: After the privacy protection comprehensive scheduling between servers is completed, this task is successfully allocated to the best server and waits to be executed. Obtain the task queue of the best server: Obtain the task queues of the TEE side and the REE side of this server respectively, so that after the calculation is completed, the task can be added to the appropriate queue.

[0081] Calculate the server resource situation: Obtain the CPU utilization rate and memory occupancy rate of the TEE side and the REE side at this time, evaluate the occupancy of server resources in a weighted manner, and further compare with the maximum load value and the maximum load balancing difference, so as to determine the scheduling selection of this task within the server.

[0082] Judge whether this calculation task is a privacy task: If so, use the trusted task processing algorithm for scheduling; if not, use the conventional task processing algorithm for scheduling.

[0083] Collect the calculation results after the calculation is completed: After the task is executed in the TEE or REE environment of the specified server, collect the calculation results to complete this round of scheduling.

[0084] Preferably, for the calculation of the server resource situation, the specific steps are as follows:

[0085] Define thresholds: Set the maximum loads T TEEt and T REEt of the TEE side and the REE side, and the maximum load balancing difference T upper . When the load difference between the TEE side and the REE side is greater than T upper , it is determined that the TEE side and the REE side of the server are unbalanced at this time.

[0086] Obtain the load situation of the TEE side and the REE side: Obtain the server load situation of the TEE side and the REE side at this time, specifically including the CPU utilization rates C TEE and C REE , the memory occupancy rates M TEE and M REE , so as to further complete the judgment of the load situation and perform in-server balanced scheduling of the task.

[0087] Calculate the comprehensive load value: Use the formula L = ω1C + ω2M (ω1 + ω2 = 1) to calculate the comprehensive load value.

[0088] The ω1 and ω2 are artificially specified weighting coefficients, and two values are set by relevant technical personnel according to the specific situation of the server.

[0089] Preferably, the trusted task processing algorithm is characterized by including the following steps:

[0090] Judge whether the current trusted resource exceeds the maximum TEE-side load: Trusted tasks are usually executed in the TEE environment.

[0091] If not, that is, L TEE <T TEEt , then allocate trusted resources for the trusted task, place it at the end of the current TEE-side task queue, and wait for execution.

[0092] If so, that is, L TEE >T TEEt , then further estimate the task waiting time T wait and the data encryption and decryption time and data replication time T required to steal it to the server with the lowest load com to further determine the task scheduling situation.

[0093] If T wait <T com , then choose to place it at the end of the current TEE-side task queue and allocate trusted resources for execution after there are idle resources.

[0094] If T wait >T com , then encrypt and transmit the data required for the trusted task to the TEE environment of the current server with the lowest load, and add the trusted task to the TEE-side task queue of the selected server with the lowest load.

[0095] Preferably, the general task processing algorithm is characterized by including the following steps:

[0096] Judge whether the general resource exceeds the maximum load T REEt : General tasks are usually executed in the general environment, but when the server load is unbalanced, consider allocating trusted resources for general tasks to reduce the task waiting time.

[0097] If L REE >T REEt and L TEE <T TEEt , the current general resources are insufficient but the trusted resources are relatively sufficient, or L REE <T REEt and |L REE -L TEE |>T upper , the general resources are sufficient but the current load is unbalanced, then allocate trusted resources for the general task, add it to the end of the current TEE-side task queue, and wait for execution.

[0098] If L REE<T REEt and |L REE -L TEE | ≤ T upper , it is determined that the current load balancing is achieved and the conventional resources are sufficient. Allocate conventional resources to the conventional tasks and add them to the end of the current REE task queue to wait for execution.

[0099] If L REE > T REEt and L TEE > T TEEt , it is determined that the resources of both the TEE side and the REE side of the current server are insufficient. Allocate the data required for this task to the processor with the lowest current load through data replication and re - allocate the task.

[0100] The TEE and REE execution environments involved in this application are specifically: the secure domain TEE side refers to a trusted execution environment based on hardware security, and the non - secure domain REE side refers to a conventional execution environment.

[0101] Finally, after each computing task is completed, the collection of computing results will be carried out. The collection process involves the integration, verification, and storage of the result data to ensure the integrity and accuracy of the computing results, so that subsequent data analysis or other related processing processes can be carried out smoothly.

[0102] Experiments prove that after adopting this method, the cross - partition transmission of privacy tasks realizes encryption protection, ensures 100% secure domain computing of privacy data, and at the same time achieves adaptive load balancing to the greatest extent, reduces the task execution time, meets the requirements of emergency tasks, and the computing efficiency is significantly improved compared with the traditional method.

Description of the Drawings

[0103] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments of this application.

[0104] Figure 1 is the flowchart of the privacy - aware workflow adaptive multi - level cluster scheduling method of this application.

[0105] Figure 2 is the architecture diagram of the privacy - aware workflow adaptive multi - level cluster scheduling system of this application.

[0106] Figure 3 is the specific description of the task priority determination algorithm based on the minimization of the earliest completion time in the embodiments of this application.

[0107] Figure 4 is the specific description of the adaptive multi - level cluster load balancing scheduling in the embodiments of this application.

[0108] Figure 5Specific description of the comprehensive scheduling of privacy protection among servers in the embodiments of this application.

[0109] Figure 6 Specific description of the fine-grained scheduling of load balancing within the server in the embodiments of this application.

[0110] Figure 7 Specific description of the trusted task processing algorithm in the embodiments of this application.

[0111] Figure 8 Specific description of the conventional task processing algorithm in the embodiments of this application.

[0112] Figure 9 Specific description of the multi-level linkage dynamic adjustment algorithm in the embodiments of this application.

[0113] Figure 10 Specific description of the realization of parameter adaptive adjustment using the simulated annealing algorithm in the embodiments of this application.

Detailed implementation manners

[0114] A workflow adaptive multi-level cluster scheduling method based on privacy perception disclosed by the present invention has the characteristics of strong security and high execution efficiency, and can be widely applied to various aspects of production and life. This system is based on domestic servers and their architectures and runs in the secure environment of domestic servers, which can effectively protect the security of domestic servers and their software.

[0115] To enable those skilled in the art to better understand the technical solutions of this disclosure, the following provides a detailed description of a workflow adaptive multi-level cluster scheduling method based on privacy perception provided by this disclosure in conjunction with the accompanying drawings.

[0116] First, several terms related to this application are introduced and explained:

[0117] Keyword matching: A traditional sensitive information recognition technology that identifies text that may contain privacy information through a predefined keyword list.

[0118] Manual marking for privacy recognition: Manual review and identification of privacy information in data.

[0119] DAG: A directed graph is a directed acyclic graph if starting from any vertex, it is impossible to return to that point after passing through several edges.

[0120] Subsequent node: For a certain node u in the DAG graph, if there is a directed edge starting from node u pointing to another node v, that is, there is a directed edge (u, v), then node v is called the successor node of node u.

[0121] Resource affinity value: The degree of affinity or similarity between the data involved in a task and the data stored in a computing node.

[0122] Load balancing: Evenly distribute workflow tasks across multiple servers to improve the system's response speed and overall availability.

[0123] TEE side: Trusted Execution Environment, a secure world with hardware and software resource partitioning implemented based on ARM TrustZone technology, used for secure computing of privacy data.

[0124] REE side: Regular Execution Environment, with lower security, where various applications can run.

[0125] Task reallocation: In a distributed computing environment, the process of dynamically reallocating tasks from one computing node to another according to the current system state and task requirements.

[0126] Dynamic data replication: In data storage and management, the process of dynamically replicating data between different storage nodes according to data access patterns, system load, or fault recovery requirements.

[0127] Data encryption: The technology of converting plaintext into ciphertext through encryption algorithms and encryption keys, which is one of the most reliable methods for computer systems to protect information security.

[0128] Confidential computing: Isolate sensitive data by performing computations in a hardware-based trusted execution environment, and only authorized programming code can access it.

[0129] The following describes the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application through the description of several exemplary embodiments.

[0130] Figure 1 The flowchart of the workflow adaptive multi-level cluster scheduling method based on privacy awareness of the present invention is shown and described as follows:

[0131] In step R1, obtain the workflow tasks, split the workflow tasks into a series of privacy tasks and regular tasks, and establish the dependency relationship of the split subtasks according to the workflow task dependency relationship.

[0132] In step R2, generate a task execution list based on minimizing the earliest completion time according to the task execution time and the dependency relationship of the tasks.

[0133] In step R3, obtain the highest-priority task in the task execution list, perform comprehensive privacy protection scheduling among servers, and allocate a target server for the highest-priority task.

[0134] In step R4, perform fine-grained load balancing within the server according to the load balancing conditions of the TEE side and REE side of the target server and the privacy of the highest-priority task.

[0135] In step R5, based on the fine-grained load balancing scheduling result and the simulated annealing algorithm, the highest-priority task performs multi-level linkage dynamic adjustment between and within servers to ensure load balancing and privacy protection until the highest-priority task is completed. Then, the highest-priority task is deleted from the task execution list, and steps S3 - S5 are repeated until the task execution list is empty.

[0136] In step R1, the preprocessing module extracts the privacy information components in the workflow computing tasks according to the privacy data definition, and completes privacy identification by combining keyword matching and manual identification. In this mode, the privacy information is initially screened by keyword matching, and then accurately judged by manual identification to complete the identification marking. Then, according to the privacy identification result, the workflow is split into a series of privacy tasks and regular tasks, and a DAG graph of the split subtasks is established according to the workflow task dependencies.

[0137] Figure 2 The specific architecture diagram of this system is shown as follows:

[0138] The part from T1 to T2 is the task preprocessing module, which is responsible for privacy data identification, task division, and modeling of workflow tasks.

[0139] The part from T3 to T4 is the task priority determination module, which is used to determine the task priority based on the minimization of the earliest finish time (EFT), and comprehensively consider the execution time and task dependencies to determine the priority of the workflow tasks and generate a task execution list.

[0140] The part from T5 to T7 is the inter-server privacy protection comprehensive scheduling module, which selects a suitable target server for the task by comprehensively considering resource affinity and server load conditions to minimize the workflow running time.

[0141] The part from T8 to T17 is the in-server load balancing fine-grained scheduling module, which is responsible for dynamically scheduling tasks based on the load conditions of trusted resources and regular resources to improve resource utilization.

[0142] The part from T18 to T19 is the server multi-level collaborative scheduling module, which is responsible for dynamically adjusting the scheduling strategy between the two levels of the server based on the existing task execution situation and the current status of the server resources to achieve efficient scheduling.

[0143] The part of T1 is the privacy identification module, which is responsible for completing task privacy identification and marking according to the privacy data definition.

[0144] The part of T2 is the task division and modeling module, which is responsible for dividing the workflow tasks into a series of privacy tasks and regular tasks according to the privacy identification result, and establishing a DAG graph according to the task dependencies.

[0145] Part T3 is the computing tasks divided according to the running logic and privacy tags.

[0146] Part T4 is the task execution list, storing the results after determining the task priorities based on the minimization of the earliest finish time (EFT).

[0147] Part T5 is the server cluster, allowing parallel processing of tasks.

[0148] Part T6 is the task priority queue maintained by each server, processing tasks according to the task priorities.

[0149] Part T7 is the distributed computing node, i.e., the server.

[0150] Part T8 is the interior of the server.

[0151] Part T9 is the TEE side, representing the feasible execution environment.

[0152] Part T10 is the REE side, representing the regular execution environment.

[0153] Part T11 is the computing task reading module, responsible for reading the tasks assigned to each server.

[0154] Part T12 is the resource management module, responsible for recording the load conditions of the TEE side and the REE side of the current server.

[0155] Part T13 is the TEE side trusted computing module, completing task processing using confidential computing.

[0156] Part T14 is the REE side regular computing module, responsible for regular execution of processing tasks.

[0157] Part T15 is the encryption module, responsible for converting plaintext to ciphertext using encryption algorithms or encryption keys when replicating privacy data to protect the security of privacy data.

[0158] Part T16 is the data replication module, responsible for transferring data from one server to another server.

[0159] Part T17 is the result output module, responsible for returning the server computing results to the superior.

[0160] Part T18 is the scheduler module, responsible for completing the comprehensive privacy protection scheduling between servers and the fine-grained load balancing scheduling within the server according to the resource status, and performing multi-level linkage dynamic adjustment according to the data replication and transfer situation.

[0161] Part T19 is the log recording module, responsible for recording and calculating the task execution time and the server load conditions to better serve the scheduling.

[0162] The relationships among the parts shown in the figure are as follows:

[0163] Figure 3 The basic process of the task priority determination algorithm based on the minimization of the earliest finish time (EFT) is shown:

[0164] In step S1, a computational cost matrix is established.

[0165] In step S2, a list of communication costs between computational tasks is established.

[0166] In step S3, the task priorities are determined.

[0167] In step S1, based on the task scale and logging, the running times of different tasks in the workflow on different servers are estimated. Assuming there are n tasks and p servers, a matrix of size n×p is established to estimate the running time of different tasks u i on different servers p k for the running time using the running time to represent the computational cost.

[0168] In step S2, for each task, a list is used to store the time required for data transfer with the tasks that have dependencies on it. Using the inter-task data transfer time c i,j to represent the communication cost, the time required for data transfer between dependent tasks is calculated according to the formula where data i,j refers to the amount of data transferred between task i and task j, and bandwidth refers to the average bandwidth between servers.

[0169] In step S3, the task rank is determined using the formula and a task execution list H is generated in descending order of rank. Here, rank(u i ) represents the priority of task u i , and the higher the rank, the higher the priority. represents the running time of task u i on server p k , P represents the set of servers, p represents the number of servers, reflects the average running time of task u i in the server cluster P, u j is the successor node of u i , c i,j represents the time required for data transfer between task u i and u j , indicating that when selecting task u i for all successor nodes u j where c ij +rank(uj ) The maximum value.

[0170] Figure 4 The basic process of performing adaptive multi-level cluster load balancing scheduling is shown:

[0171] Step Q1 is the initial stage. In this state, the task priorities have been determined through a task priority determination algorithm based on minimizing the earliest finish time (EFT), and a task execution list H has been generated.

[0172] Step Q2 is to obtain the task with the highest priority in the task execution list H.

[0173] Step Q3 is to determine whether the task execution list H is empty.

[0174] Step Q4 is to determine that it is empty and collect the calculation results.

[0175] Step Q5 is to enter the end stage after the calculation results are collected. In this state, all steps of the adaptive multi-level cluster load balancing scheduling are completed.

[0176] Step Q6 is to determine that the task execution list H is not empty and perform comprehensive scheduling for privacy protection between servers.

[0177] Step Q7 is to perform fine-grained load balancing scheduling within the server.

[0178] Step Q8 is to perform multi-level linkage dynamic adjustment.

[0179] Step Q9 is to update the list, delete the current task with the highest priority from the task execution list H, and re-enter Step Q3.

[0180] Figure 5 The basic process of comprehensive scheduling for privacy protection between servers is shown:

[0181] Step M1 is to calculate the resource affinity value.

[0182] Step M2 is to calculate the load balancing value.

[0183] Step M3 is to select the target server.

[0184] In Step M1, calculate the current task u with the highest priority in the task execution list H i on different servers p k The sum of the data transmission time and the calculation time spent on running Taking the sum of the data transmission time and the calculation time represents the resource affinity value between the i-th task and the k-th server. If it is private data, the data encryption and decryption time t needs to be added. s .

[0185] In step M2, calculate the current highest-priority task u i Start the required waiting time for execution on different servers p k With the waiting time represent the load balancing value of the k-th server.

[0186] In step M3, considering the user requirements and the current computing environment, according to the formula Calculate the comprehensive evaluation value of the current highest-priority task u i for the k-th server, select the server corresponding to the minimum AMEM value as the target server, and allocate the highest-priority task u to it i , and the initial values of α and β are both 0.5;

[0187] Figure 6 Shows the basic process of fine-grained load balancing within the server:

[0188] Step 101 is the initial stage, at this time the task has been allocated to the server through the comprehensive scheduling of privacy protection between servers.

[0189] Step 102 is to obtain the comprehensive scheduling calculation task of privacy protection between servers.

[0190] Step 103 is to obtain the task queues at the TEE end and the REE end within the server.

[0191] Step 104 is to calculate the resource situation of this server.

[0192] Step 105 is to determine whether this calculation task is a privacy task.

[0193] Step 106 is to determine that if so, use the trusted task processing algorithm to process.

[0194] Step 107 is to determine that if not, use the conventional task processing algorithm to process.

[0195] Step 108 is to collect the calculation results.

[0196] Step 109 is the end step, under this step, the fine-grained load balancing within the server has been completely completed.

[0197] Figure 7 Shows the basic process of the trusted task processing algorithm:

[0198] Step 201 is the initial stage, at this time the calculation of the server resource situation has been completed, and the task queues at the TEE end and the REE end of the server have been obtained.

[0199] Step 202 is to determine whether the trusted resources exceed the maximum load of the TEE end.

[0200] ​Step 203: If the judgment result is no, allocate trusted resources and the task enters the trusted execution environment for computing.

[0201] Step 204: If the judgment result is yes, then judge the estimated waiting time T for computing on this server. wait Is it greater than the data encryption / decryption time and data replication time T required to steal the server with the lowest load? com .

[0202] Step 205: If the judgment result is no, select to execute on this server, add it to the TEE-side task queue, and re-enter Step 202.

[0203] Step 206: If the judgment result is yes, encrypt and copy the required data to the TEE side of the server with the lowest load.

[0204] Step 207: Allocate this task to the server with the lowest load for computing.

[0205] Step 208: Collect the computing results.

[0206] Step 209: End step. All tasks under this step have been completed.

[0207] Figure 8 Shows the basic process of the conventional task processing algorithm:

[0208] Step 301: Initial stage. At this time, the calculation of the server resource situation has been completed, and the task queues of the TEE side and REE side of the server have been obtained.

[0209] Step 302: Judge whether the conventional resources exceed the maximum load.

[0210] Step 303: If the judgment result is yes, then judge whether the trusted resources exceed the maximum load.

[0211] Step 304: If the judgment result is no, allocate trusted resources and the computing task enters the trusted execution environment (TEE) for computing.

[0212] Step 305: If the judgment result is yes, allocate the data required for this task to the currently lowest-load server through data replication.

[0213] Step 306: After completing the data replication, re-allocate this task.

[0214] Step 307: Judge that the conventional resources do not exceed the maximum load, and then judge whether the loads of the TEE side and REE side are balanced.

[0215] Step 308: If the judgment result is yes, allocate conventional resources and the computing task enters the conventional execution environment (REE) for computing.

[0216] Step 309: If the judgment result is negative, allocate trusted resources and the computing task enters the trusted execution environment (TEE) for computing.

[0217] Step 310: Collect the computing results.

[0218] Step 311: End step. All tasks have been completed under this step.

[0219] Figure 9 The basic process of the multi-level linkage dynamic adjustment algorithm is shown:

[0220] Step 401: Initialize parameters. Under this step, the optimal initial temperature T, termination temperature T end , and cooling rate ω are set manually, where 0 < ω < 1. Initially, α = 0.5 and β = 0.5.

[0221] Step 402: Define the fitness function where p is the total number of servers, L i represents the load condition of the i-th server, represents the average load of p servers.

[0222] Step 403: Calculate the fitness value F old (α, β) under the current parameters.

[0223] Step 404: Use the simulated annealing algorithm to achieve parameter self-adaptive adjustment and determine the optimal parameters α and β.

[0224] Figure 10 The basic process of using the simulated annealing algorithm to achieve parameter self-adaptive adjustment is shown:

[0225] Step 501: Initial stage. At this time, the initialization of parameters, the definition of the fitness function, and the calculation of the fitness function value F old (α, β) based on the original parameters α and β have been completed.

[0226] Step 502: Perturbation to generate a new solution. A new solution is generated by adding a random number that follows a normal distribution with a mean of 0 and a standard deviation of 0.1, that is, α new = α + ε1, β new = β + ε2, where ε1 and ε2 are random numbers that satisfy the above normal distribution, and at the same time ensure that α new + β new = 1.

[0227] Step 503: Calculate the new fitness function value F new and β new based on the new parameters α new (α new , β new ).

[0228] Step 504 is to judge the new fitness function value F new (α new , β new ) is greater than the old fitness function value F old (α, β).

[0229] Step 505 is to judge that if so, adjust α to α new , and adjust β to β new .

[0230] Step 506 is to judge that if not, then calculate the probability P value according to the formula and generate a random number r between (0, 1).

[0231] Step 507 is to then judge whether r is less than P

[0232] Step 508 is to judge that if so, adjust α to α new , and adjust β to β new .

[0233] Step 509 is to judge that if not, then α and β remain unchanged

[0234] Step 510 is to adjust T = ωT

[0235] Step 511 is to judge whether T is less than T end , judge that if not, repeat steps 502 - 511

[0236] Step 512 is to judge that if so, enter the end step, and all tasks have been completed in this state

[0237] This example of this application is a typical case, that is, there are many other examples available for use or testing in addition to this example, and this example is only for illustrative description

[0238] The data sets used in this example are all from the Internet, and testing its performance using the corresponding data sets is not within the scope of the claims of this patent

[0239] Except for the items without relevant claims in this patent statement, all other activities related to the use, development, and design of this patent shall be permitted by the inventor before implementation. The content related to the intellectual property of this invention that is not mentioned in this application shall also be protected accordingly

[0240] This application embodiment provides a non - transient computer - readable storage medium, and the non - transient computer - readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the methods provided by the above - mentioned method embodiments, for example, including: setting the scheduler developed by this invention in any field of Spark and TrustZone, using the algorithms and related programs of this application, etc

[0241] Finally, it should be noted that although the present invention mainly shows specific cases in the example display and design, each component involved in this method should still be a part to be protected, including but not limited to modules such as preprocessing, resource management, task scheduler and their codes. Therefore, any modifications, improvements, uses, etc. of the implementation of the present invention should be covered by the content of the claims and the specification of the present invention.

Claims

1. A privacy-aware workflow adaptive multi-level cluster scheduling method, characterized in that, It includes the following steps: S1. Split the obtained workflow tasks into a series of privacy tasks and regular tasks, and establish the dependency relationships of the split subtasks according to the workflow task dependencies; S2. Based on the task execution time and the task dependencies, generate a task execution list by minimizing the earliest completion time; S3. Obtain the highest-priority task in the task execution list, perform comprehensive privacy protection scheduling among servers, and allocate a target server for the highest-priority task; S4. Perform fine-grained load balancing scheduling inside the server according to the load balancing conditions of the TEE side and REE side of the target server and the privacy of the highest-priority task; S5. Based on the fine-grained load balancing scheduling result and the simulated annealing algorithm, the highest-priority task performs multi-level linkage dynamic adjustment among servers and within the server to ensure load balancing and privacy protection until the highest-priority task is completed. Delete the highest-priority task from the task execution list, and repeat steps S3 - S5 until the task execution list is empty.

2. The method according to claim 1, wherein Step S1 obtains workflow tasks, splits the workflow tasks into a series of privacy tasks and regular tasks, and establishes the dependency relationships of the split subtasks according to the workflow task dependencies. Specifically, it includes the following steps: Privacy identification: Obtain workflow tasks, and use a combination of keyword matching and manual marking to obtain the privacy identification result of the workflow tasks; Task splitting: According to the privacy identification result, split the workflow tasks into a series of privacy tasks and regular tasks; Task modeling: Based on the dependencies of the workflow tasks, establish a DAG graph of the split subtasks.

3. The method according to claim 1, characterized in that, The generation of the task execution list based on minimizing the earliest completion time includes the following steps: Obtain the computing costs of different tasks on different servers; Obtain the communication costs between tasks with dependencies; Based on the computing costs and communication costs, construct a task priority function and generate a task execution list; Preferably, the generation of the task execution list based on minimizing the earliest completion time includes the following steps: Build a computational cost matrix: Build an n×p matrix and calculate the running time of different tasks u i on different servers p k The running time Take the running time to represent the computational cost, where n is the number of tasks and p is the number of servers; Establish a list of communication costs between computing tasks: calculate the time required for data transfer between dependent tasks c i,j represents task u i and u j The time required for data transfer, with the data transfer time c between tasks i,j represents the communication cost, where data i,j refers to the amount of data transferred between task i and task j, and bandwidth refers to the average bandwidth between servers; Determine task priorities and generate a task execution list: Using the formula Determine task priorities and generate a task execution list according to the descending order of rank. Among them, rank(u i ) represents the priority of task u i . The larger the rank, the higher the priority. P represents the set of servers reflects the average time for task u i to run in the server cluster P. u j is the successor node of u i , indicating the value of c i + rank(u j ) that is the largest among all successor nodes u ij of u j .

4. The method according to claim 1, characterized in that, The comprehensive privacy protection scheduling among servers includes the following steps: Calculate the data transmission time, computing time, and / or encryption and decryption time for the highest-priority task to run on different servers, representing the computing resource affinity value; Calculate the waiting time for the highest-priority task to start running on different servers, representing the load balancing value of the servers; Based on user requirements and the current computing environment, construct a priority function according to the computing resource affinity value and the load balancing value and allocate a target server for the highest-priority task.

5. The method according to claim 4, characterized in that The comprehensive privacy protection scheduling among servers includes the following steps: Calculate the resource affinity value: Calculate the current highest-priority task u according to the task execution list i On different servers p k The sum of the data transfer time and the computing time spent on running And / or the data encryption / decryption time t s , representing the resource affinity value between the i-th task and the k-th server; Calculate the load balancing value: Calculate the current highest-priority task u i Start the waiting time required for execution on different servers p k The waiting time required for execution on different servers p Represents the load balancing value of the k-th server; Select the target server: Based on user requirements and the current computing environment, calculate the current highest-priority task u according to the formula Calculate the current highest-priority task u i Regarding the comprehensive evaluation value of the k-th server, select the server corresponding to the minimum AMEM value as the target server and allocate the highest-priority task u to it i , where α + β = 1, and the initial values of α and β are both 0.

5.

6. The method according to claim 1, characterized in that, The fine-grained load balancing scheduling inside the server includes the following steps: Obtain the comprehensive privacy protection scheduling computing tasks among servers, the task queues and load conditions of the TEE side and REE side inside the server; Judge whether the current computing task is a privacy task; If so, the current computing task enters the TEE side for computing; if not, the current computing task enters the REE side for computing; Collect the computing results after the computing ends.

7. The method according to claim 6, wherein Determine whether the current computing task is a privacy task; if so, the current computing task enters the TEE-side calculation, including the following steps: Determine the load L on the TEE side of the current server TEE whether it exceeds the maximum value T TEEt ; If L TEE <T TEEt , then the current computing task enters the computing on the TEE side of the current server; If L TEE >T TEEt , obtain the waiting time T calculated on the current server wait , and the data encryption / decryption time and data replication time T required to copy it to the server with the lowest load com ; If T wait <T com , select to execute on the current server and add it to the task queue at the TEE end of the current server; if T wait >T com , allocate the current computing task to the server with the lowest load for computing, and encrypt and copy the required data to the TEE end of the server with the lowest load; Determine whether the current computing task is a privacy task; If not, the current computing task enters the REE-side calculation, including the following steps: Judge whether the load L at the REE end of the current server REE exceeds the maximum value T REEt ; If L REE > T REEt , determine whether the current server TEE end exceeds the maximum load value T TEEt ; If L TEE <T TEEt , the current computing task enters the computing on the TEE side of the current server; if L TEE >T TEEt , the data required for the current computing task is allocated to the processor with the lowest current load through data replication, and the task allocation is redone; If L REE <L REEt , then determine whether the current server's TEE side and REE side meet the load balancing; If the load balancing is satisfied, i.e., |L REE -L TEE | ≤ T upper , then the computing task enters the REE side for computing; Otherwise, the computing task enters the TEE-side calculation; Among them, T upper is the maximum difference of load balancing.

8. The method according to claim 1, characterized in that, According to the fine-grained load balancing scheduling result, based on the simulated annealing algorithm, the highest-priority task performs multi-level linkage dynamic adjustment between and within servers to ensure load balancing and privacy protection, including the following steps: Define the fitness function: Define the fitness function Calculate the fitness function value F under the current parameters old (α,β); where p is the total number of servers, and L i represents the load condition of the i-th server, represents the average load of p servers, set the optimal initial temperature T, termination temperature T end , cooling rate ω, 0 < ω < 1, initially α = 0.5, β = 0.5; New solutions are generated by perturbation: A random number that follows a normal distribution with a mean of 0 and a standard deviation of 0.1 is added to generate a new solution, i.e., α new = α + ε1, β new = β + ε2, where ε1 and ε2 are random numbers that satisfy the above normal distribution, and α new + β new = 1; Calculate the value of the new fitness function: Based on the parameters α new and β new Calculate the value of the new fitness function F new (α new , β new ); Judge the new fitness function value F new (α new ,β new ) is greater than the old fitness function value F old (α, β); If so, adjust α to α new , and adjust β to β new ; Otherwise, according to the probability decide whether to accept the new solution, generate a random number r between (0, 1), and determine whether r is less than P; If r < P, then adjust α to α new , and adjust β to β new ; If r > P, then α and β remain unchanged; Adjust T = ωT; Repeat the above steps until T < T end .

9. A privacy-aware workflow adaptive multi-level cluster scheduling system, characterized in that, Including: A task preprocessing module, which is used to obtain workflow tasks, split the workflow tasks into a series of privacy tasks and regular tasks, and establish the dependency relationship of the split subtasks according to the workflow task dependency relationship; A task determination module, which is used to generate a task execution list based on the earliest completion time minimization according to the task execution time and the dependency relationship of the tasks; An inter-server privacy protection comprehensive scheduling module, which is used to obtain the highest-priority task in the task execution list, perform inter-server privacy protection comprehensive scheduling, and allocate a target server for the highest-priority task; An intra-server load balancing fine-grained scheduling module, which performs intra-server load balancing fine-grained scheduling according to the load balancing conditions of the TEE side and the REE side of the target server and the privacy of the highest-priority task; A multi-level collaborative server scheduling module, which is used to perform multi-level linkage dynamic adjustment of the highest-priority task between and within servers based on the simulated annealing algorithm according to the fine-grained load balancing scheduling result to ensure load balancing and privacy protection until the highest-priority task is completed, and delete the highest-priority task from the task execution list until the task execution list is empty.

10. A privacy-aware workflow adaptive multi-level cluster scheduling system according to claim 9, characterized in that, The task preprocessing module includes a privacy identification sub-module and a task partitioning and modeling sub-module; the privacy identification sub-module is used to complete task privacy identification and marking according to the privacy data definition; The task partitioning and modeling sub-module is responsible for partitioning the workflow tasks into a series of privacy tasks and regular tasks according to the privacy identification result, and establishing a DAG graph according to the task dependency relationship; the intra-server load balancing fine-grained scheduling module includes a computing task reading module, a resource management module, a TEE-side trusted computing module, a REE-side regular computing module, an encryption module, a data replication module, and a result output module. The computing task reading module is used to read the tasks assigned to each server; The resource management module is used to record the load conditions of the TEE side and the REE side of the current server; the TEE-side trusted computing module is used to complete task processing using confidential computing; The REE side is used for the regular computing module, which is responsible for regular execution of processing tasks; the encryption module is used to convert plaintext to ciphertext using an encryption algorithm or encryption key when copying privacy data to protect the security of privacy data; the data replication module is used to transfer data from one server to another server; The result output module is used to return the server calculation result to the upper level; The server multi-level collaborative scheduling module includes a scheduler module and a log recording module. The scheduler module is used to complete the comprehensive privacy protection scheduling between servers and the fine-grained load balancing scheduling within the server according to the current resource status, and perform multi-level linkage dynamic adjustment according to the data replication and transfer situation. The log recording module is used to record and calculate the task execution time and the server load situation to better serve the scheduling.