Multi-task scheduling method, system and device based on trusted computing and storage medium
By constructing a computational sandbox state channel and a blockchain scheduling model in federated learning, the problems of trusted supervision and privacy leakage in federated learning are solved, achieving data privacy protection and system reliability, and improving resource utilization and task efficiency.
Patent Information
- Application Number
- CN202310216223.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-03-02
AI Technical Summary
Existing technologies lack effective and trustworthy oversight mechanisms in federated learning, making it impossible to prevent privacy breaches caused by malicious node attacks. Furthermore, reliance on a central server poses a single point of failure risk, failing to guarantee data privacy and security.
By constructing a computing sandbox state channel, a blockchain scheduling model is used to manage participating nodes and training nodes, perform global model aggregation and parameter calculation, and transmit training results in the computing sandbox state channel. Combined with digital signatures and smart contracts, trusted supervision is carried out to prevent malicious behavior.
It achieves trusted supervision in the federated learning process, prevents malicious node attacks, ensures user data privacy and security, and improves the reliability and efficiency of the system.
Smart Images

Figure CN116360939B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a multi-task scheduling method, system, device and storage medium based on trusted computing, and belongs to the technical field of multi-task scheduling. BACKGROUND
[0002] Traditional centralized industry chain finance generally has problems such as high verification cost, incomplete information, repeated financing, difficulty in monitoring false financing, and increased financing cost. As a distributed accounting technology, blockchain has become a way to provide secure data sharing in data systems, and can realize peer-to-peer transmission of digital assets without any intermediaries, thereby eliminating the process of verifying transactions by any intermediary institution, and can provide high-quality data and secure data sharing for participants.
[0003] Current research combines industry chain finance with blockchain and integrates blockchain technology into the complex business scenarios of industry chain finance. This also brings the problem of data privacy security of industry chain finance in a multi-party data sharing environment. Although multi-party data sharing based on blockchain can make data open and transparent and ensure the trusted circulation of data, it cannot guarantee the data privacy and security of users. How to protect data privacy in the case of data sharing is a hot issue in current research.
[0004] Federated learning (FL, federated learning) as a new data value sharing method aims to solve the problem of data silos while protecting data privacy. Unlike traditional machine learning, federated learning refers to the coordination of joint learning training by one or more central servers and multiple clients, and each participant does not need to upload its own local data, but only needs to upload the parameters and models trained using local data to the central server, and aggregate the global model at the central server, thereby protecting the privacy and security of the data.
[0005] Traditional federated learning relies on a single central server, and if this fragile central server fails, it will lead to inaccurate global models and even hinder the entire federated learning process. In addition, if the central server is attacked or the central server itself has malicious behavior, the data privacy and security of each participant cannot be guaranteed. The secure and reliable decentralized data sharing solution provided by blockchain perfectly compensates for such defects, and the protection of data privacy by federated learning also avoids the shortcoming of blockchain technology in protecting data privacy.
[0006] There are many studies combining blockchain and federated learning, using the characteristics of blockchain to replace the central server, and providing corresponding rewards to nodes with good performance to encourage nodes with better data to participate more actively in training. These mechanisms and schemes use blockchain technology to verify and supervise the original data and final calculation results of federated learning. Although such an approach can share data value while protecting user privacy to some extent by not exporting local data, it does not consider the trusted supervision of the federated learning model training process and the privacy leakage caused by malicious node attacks, and ignores the trusted supervision of the federated learning model training and aggregation process. SUMMARY
[0007] In view of the defects of the prior art, a first object of the present application is to provide a method for establishing a computing sandbox state channel by constructing a participant node, a blockchain scheduling model, and a training node. In the computing sandbox state channel, the training node performs global model aggregation and parameter calculation on relevant parameters to obtain training results and state information of the federated learning task. The training results and state information are transmitted to the blockchain scheduling model, and the computing sandbox state channel is closed to release the occupied computing power resources. The blockchain scheduling model sends the training results to the participant node, updates the state information of the computing power resources, completes the scheduling of the federated learning multi-task, realizes the trusted supervision of the federated learning model training and aggregation process, and effectively avoids privacy leakage caused by malicious node attacks. The scheme is scientific, reasonable, and feasible, and is a multi-task scheduling method based on trusted computing.
[0008] In view of the defects of the prior art, a second object of the present application is to provide a multi-task scheduling system based on trusted computing, which establishes a computing sandbox state channel by constructing a participant node, a blockchain scheduling module, and a training node. In the computing sandbox state channel, the training node performs global model aggregation and parameter calculation on relevant parameters, so that the model training and parameter transmission process of the federated learning task is performed in the computing sandbox state channel, ensuring the trustworthiness of the calculation and the privacy security of the user data in the federated learning training process, and realizing the trusted supervision of the federated learning model training and aggregation process.
[0009] In view of the defects of the prior art, a third object of the present application is to provide a multi-task scheduling method, system, device, and storage medium based on trusted computing, which establishes a federated learning training trusted supervision framework based on a computing sandbox state channel, and performs the model training and parameter transmission process of the federated learning task in the computing sandbox state channel, ensuring the trustworthiness of the calculation and the privacy security of the user data in the federated learning training process, and realizing the trusted supervision of the federated learning model training and aggregation process.
[0010] To achieve the above object, the first technical solution of the present application is:
[0011] A federated learning multi-task scheduling method based on a trusted computing sandbox, comprising the following contents:
[0012] Obtain one or more federated learning tasks to be requested through a pre-constructed participant node;
[0013] The participant node sets relevant parameters according to the characteristics of the federated learning task, and signs an intelligent contract with the pre-constructed blockchain scheduling model using a digital signature technology;
[0014] After signing, the relevant parameters are transmitted to the blockchain scheduling model;
[0015] The blockchain scheduling model selects a training node according to the relevant parameters of the federated learning task, and allocates and schedules computing power resources, and establishes a computing sandbox state channel;
[0016] The training node performs global model aggregation and parameter calculation on the relevant parameters in the computing sandbox state channel, and obtains the training result and state information of the federated learning task;
[0017] The training result and state information are transmitted to the blockchain scheduling model, and the computing sandbox state channel is closed to release the occupied computing power resources;
[0018] The blockchain scheduling model sends the training result to the participant node, updates the state information of the computing power resources, and waits for a new federated learning task to arrive to allocate and schedule the computing power resources, thereby realizing the federated learning multi-task scheduling based on the trusted computing sandbox.
[0019] Through continuous exploration and testing, the present application constructs a participant node, a blockchain scheduling model and a training node; then the blockchain scheduling model establishes a computing sandbox state channel according to the task request of the participant node; in the computing sandbox state channel, the training node performs global model aggregation and parameter calculation on the relevant parameters, and obtains the training result and state information of the federated learning task; the training result and state information are transmitted to the blockchain scheduling model, and the computing sandbox state channel is closed to release the occupied computing power resources; the blockchain scheduling model sends the training result to the participant node, updates the state information of the computing power resources, completes the scheduling of the federated learning multi-task, realizes the trusted supervision in the federated learning model training and aggregation process, and can effectively avoid the privacy leakage caused by malicious node attacks, and the scheme is scientific, reasonable and feasible.
[0020] Further, the application constructs a federated learning training trusted supervision framework based on a computing sandbox state channel, and places the model training and parameter transmission process of the federated learning task in the computing sandbox state channel, thereby ensuring the trustworthiness of the computation and the privacy and security of the user data during the federated learning training process, and achieving trusted supervision during the federated learning model training and aggregation process.
[0021] As a preferred technical measure:
[0022] The blockchain scheduling model is used to manage the state and verification information of each participant node, integrate and virtualize various computing power resources registered by each participant node, and coordinate and supervise each participant node for federated learning training through a smart contract. The method for selecting a training node is as follows:
[0023] Since different training nodes have resource heterogeneity, the computing power resources of the training nodes are matched according to the computing power resources possessed by the training nodes.
[0024] When a training node does not have the ability to independently complete a model training task, the resources of a plurality of training nodes are integrated and reasonably allocated and scheduled to construct a composite training node, so as to coordinate and solve the problem of resource heterogeneity.
[0025] The method for establishing a computing sandbox state channel by the blockchain scheduling model is as follows:
[0026] In the process of constructing the computing sandbox state channel, the blockchain scheduling model reasonably schedules resources to the federated learning task according to the requirements of different federated learning tasks, so that the entire system works more efficiently.
[0027] Meanwhile, according to the requirements of the current federated learning task, the blockchain scheduling model dynamically schedules and allocates computing power resources to establish a computing sandbox state channel, so as to fully utilize the computing power resources and release the used computing power resources immediately after the task is completed. The federated learning task does not own or retain the allocated computing power resources.
[0028] After that, the training and parameter transmission of the federated learning task will be performed in the computing sandbox state channel. Any malicious behavior will be supervised by all the training nodes and reported to the blockchain scheduling model, so as to ensure the trusted computation of data and the security of the model training and aggregation process.
[0029] The malicious behavior includes data poisoning and model poisoning.
[0030] Data poisoning refers to the pollution of task data.
[0031] Model poisoning refers to sending incorrect model parameters or damaged model parameters.
[0032] The real physical address of the computing sandbox state channel is not disclosed to all participant nodes and training nodes, and the blockchain scheduling model is responsible for shielding the real physical address of the computing sandbox state channel from the participant nodes and training nodes, and only providing virtual addresses and interfaces to the participant nodes and training nodes;
[0033] The federal learning task is completed in the computing sandbox state channel, and each participant node and training node cannot obtain the model and parameters of other participant nodes and training nodes, so as to prevent attacks on the original data, thereby protecting the privacy and security of user data.
[0034] As a preferred technical measure:
[0035] The participant node is the requester of the federal learning task, and can also serve as a training node participating in the federal learning task, and has user data and resources with training value for representing an industry or enterprise company;
[0036] The participant node applies to the blockchain scheduling model to become a training node and cooperates with other participant nodes to perform the federal learning task; and uses local data to jointly train a global model through a federal learning method to share the value of data, but does not share user data with other participant nodes;
[0037] The related parameters at least include the number of training rounds or / and the initialized model or / and the total budget overhead;
[0038] The computing resource at least includes a virtual resource or / and a storage space.
[0039] As a preferred technical measure:
[0040] The blockchain scheduling model can schedule and distribute multiple federal learning tasks to multiple training nodes to achieve the goal of minimizing the completion time of the entire system under the constraints of task cost and completion time, and the construction process is as follows:
[0041] A node set of several training nodes is constructed, and the expression is as follows:
[0042]
[0043] Where Nodes represents a node set providing computing resources, N i represents the i-th training node, and N node is the number of training nodes;
[0044] The expression of the training node is as follows:
[0045]
[0046] Where PU j represents the j-th processing unit, This represents the number of processing units in the i-th training node, j = 1, 2, ..., N. pu ;
[0047] The expression for the processing unit PU is as follows:
[0048] PU = {SIDP, E, COST, TD}
[0049] Where SIDP is the processing unit number, E is the execution capacity of the processing unit (the amount of work it can process per unit time), COST is the cost incurred in using the processing unit to execute per unit time, and TD represents the communication latency from the training node where the PU is located to other training nodes. i,j This represents the communication delay from training node i to training node j;
[0050] The set of federated learning tasks is constructed as follows:
[0051]
[0052] Where Tasks is a set of federated learning tasks, T i Representing the i-th federated learning task, N task The number of federal learning tasks;
[0053] The expression for the federated learning task T is as follows:
[0054] T = {SIDT, WorkLoad, ...}
[0055] MaxT, MaxC, PI, N train}
[0056] SIDT is the number of the federated learning task, WorkLoad represents its workload, MaxT represents the maximum completion time that the task can tolerate, MaxC represents the maximum cost that the task can bear, MaxT and MaxC are used as indicators of task service quality requirements, PI represents the priority of the task, and Ntrain represents the set of nodes participating in this federated learning task.
[0057] Construct a task resource allocation matrix to represent the task resource allocation scheme. Its calculation formula is as follows:
[0058]
[0059] Where X is N task *N pu The size of the task resource allocation matrix, x i,j The expression representing the allocation of task i on computation unit j is as follows:
[0060]
[0061] The completion time of each federated learning task is calculated according to the task resource allocation matrix, and the calculation formula is as follows
[0062]
[0063] Where ECT i represents the completion time of the i-th federated learning task, WorkLoad i represents the workload of the i-th federated learning task, N pu represents the number of computing units, E j represents the execution capability of the j-th computing unit, x i,j represents the allocation of the i-th federated learning task to the j-th computing unit.
[0064] The cost of each federated learning task is calculated, and the calculation formula is as follows
[0065]
[0066] TCost i represents the cost of the i-th federated learning task, COST j represents the execution cost of the j-th computing unit.
[0067] The maximum completion time maxMakespan of the computing system is calculated, and the calculation formula is as follows:
[0068]
[0069]
[0070] Finally, the resource scheduling is finally abstracted as a target optimization problem, that is, to solve the resource allocation matrix X, which satisfies
[0071]
[0072]
[0073] and TCost i ≤ MaxC i ;
[0074] Thus, the construction of the blockchain scheduling model is completed.
[0075] Further, the federated learning task includes an anomaly behavior detection task or / and a risk assessment task or / and a customer behavior analysis task or / and a product intelligent recommendation task.
[0076] As a preferred technical measure:
[0077] The method for scheduling and allocating computing resource by the blockchain scheduling model is as follows
[0078] The blockchain scheduling model selects an optimal resource allocation scheme according to the current node resource state and the change of the federated learning task to be allocated, constructs a computing sandbox state channel, and works more efficiently while meeting the service quality requirements of the task to the greatest extent, which includes an environment supervision unit and an agent decision unit.
[0079] The environment supervision unit is used for supervising the state of each node and managing the computing resources possessed by the nodes, maintaining a task list waiting for resource allocation, and providing the current state information to the agent decision unit, including the current resource state, the state of the task to be allocated, etc.
[0080] The environment supervision unit returns the reward of the current operation and the state at the next moment after the agent decision unit makes an action.
[0081] The agent decision unit is a resource scheduler, which can make corresponding decisions according to the state given by the environment supervision unit, make corresponding actions, and allocate computing resources to the task according to the action.
[0082] The agent decision unit is composed of a value network Critic and a policy network Actor. The value network Critic network can score the current state and evaluate the state. The policy network Actor network selects the optimal action according to the state of the environment supervision unit.
[0083] The experience replay method is adopted, and an experience buffer pool R is set. The state of the environment supervision unit, the executed action, the obtained reward, and the state at the next moment are stored as experience samples in the buffer pool. When training network parameters, a small batch of data is randomly extracted from the buffer pool for training and parameter updating, so as to better utilize past experience and break the correlation of parameter updating.
[0084] Before resource scheduling, the federated learning tasks waiting for resource allocation at the current moment are sorted according to priority, and resources are allocated to them in turn.
[0085] The priority of each federated learning task The priority of the i-th federated learning task is represented as:
[0086] PI i =initial_pr i +wait_time i
[0087] Where initial_pr i is the initial priority of the federated learning task, and W iThe less the workload is, the higher the priority is, which takes advantage of the short job first, and is beneficial to reduce the average waiting time, wait_time i The time for waiting for resource allocation for the federated learning task is a dynamically changing value, and the longer the waiting time is, the higher the priority is, which is beneficial to reduce the problem of job starvation;
[0088] When training the strategy network Actor and the value network Critic, according to the current environment regulatory unit state s t , an action a t is made based on the strategy network, the reward r t is obtained by interacting with the environment regulatory unit, and the new environment regulatory unit state s t+1 is obtained, (s t , a t , r t , s t+1 ) is put into the buffer pool R as a sample;
[0089] Then N groups of data samples are randomly extracted from the buffer pool, and the average value of the gradient is calculated to update the network parameters of the strategy network Actor and the value network Critic.
[0090] The application aims at the resource heterogeneity problem of the federated learning participant node, proposes a participant resource management method, models and analyzes the resource scheduling problem in the process of establishing the computing sandbox state channel, and constructs a resource scheduling strategy for the process of the computing sandbox state channel.
[0091] As a preferred technical measure:
[0092] The blockchain scheduling model solves the resource scheduling optimization by using a deep reinforcement learning method, which includes the following contents:
[0093] A state space S is constructed, which includes a learning task resource allocation matrix of each computing resource at t time, and a specific state of the federated learning task to be allocated at t time;
[0094] The specific state of the federated learning task to be allocated includes the workload, budget, and maximum completion time that can be tolerated of the federated learning task.
[0095] The expression of the state space is as follows:
[0096] s t ={X t , WorkLoad i , MaxT i , MaxC i}
[0097] Where s t represents the state at t time, WorkLoadi , MaxT i , MaxC i respectively represent the workload of the i-th federated learning task to be allocated, the maximum tolerable completion time, and the maximum budget overhead;
[0098] Action space A, the agent performs the corresponding action in state s t , the highest priority federated learning task is taken out from the waiting queue and resources are allocated for it, and the expression of the action is as follows:
[0099]
[0100] Where a t represents the action at time t, N pu is the number of resources of the current system, γ i ∈{0,1} represents the allocation of federated learning tasks on the i-th resource, 0 represents unallocated, and 1 represents allocated;
[0101] Reward R, the reward of the action performed is calculated according to whether the service quality of the federated learning task is met after the resources are allocated, and the system completion time;
[0102] If the action of resource allocation causes the completion time and cost of the federated learning task to be greater than the maximum value that the federated learning task can tolerate, a huge penalty will be obtained, and the maximum completion time of the system is used as the standard of the reward;
[0103] The expression of the reward is as follows:
[0104]
[0105] Where r t represents the reward obtained by performing a t action in state s t , maxMakespan represents the maximum completion time of the system, ECT i , Tcost i , MaxT i , MaxC i respectively represent the completion time, cost overhead, maximum tolerable completion time and budget overhead of the federated learning task i, and ε is a parameter set by the system.
[0106] The deep reinforcement learning method is adopted to solve the resource scheduling problem in the process of constructing the computing sandbox state channel, the maximum completion time of the system is minimized under the condition of meeting the service quality requirements of the task, and the resource utilization rate and the overall performance of the system are improved.
[0107] As a preferred technical measure:
[0108] The specific steps of solving include the following contents:
[0109] A total of M generations are trained, T times of network parameter updates are performed in each generation, a federal learning task waiting queue is randomly initialized at the beginning of each generation, and an initial state is obtained from the environment supervision unit;
[0110] Before updating the parameters each time, the system obtains a plurality of (s t , a t , r t , s t+1 ) samples from the environment supervision unit and puts them into the buffer pool R; when the amount of data in R reaches a set value, N groups of samples are randomly selected from R to perform batch gradient update on the network parameters;
[0111] The update process calculates the target value label and the target error, and calculates the gradient of the policy network, and calculates the loss function according to the least square difference;
[0112] Thus, the gradient calculation formula of the value network is obtained; the batch update method is used to update the parameters of the policy network and the value network respectively; and the parameters of the target value network are updated after k iterations.
[0113] To achieve one of the above purposes, the second technical scheme of the present application is:
[0114] A federal learning multi-task scheduling system based on a trusted computing sandbox includes a plurality of participant nodes, a blockchain scheduling module, and training nodes:
[0115] The participant node is used to obtain one or more requested federal learning tasks, set relevant parameters according to the characteristics of the federal learning task, and sign an intelligent contract with the blockchain scheduling module using digital signature technology;
[0116] The blockchain scheduling module selects a training node according to the relevant parameters of the federal learning task, allocates and schedules computing resource, and establishes a computing sandbox state channel;
[0117] The training node performs global module aggregation and parameter calculation on the relevant parameters in the computing sandbox state channel, obtains the training result and state information of the federal learning task, and transmits the training result and state information to the blockchain scheduling module, and closes the computing sandbox state channel to release the occupied computing resource;
[0118] The blockchain scheduling module sends the training result to the participant node, updates the state information of the computing resource, and waits for a new federal learning task to arrive to allocate and schedule the computing resource, thereby realizing federal learning multi-task scheduling based on a trusted computing sandbox.
[0119] Through continuous exploration and experiments, the application constructs a participant node, a blockchain scheduling module and a training node, then the blockchain scheduling module establishes a computing sandbox state channel according to a task request of the participant node, in the computing sandbox state channel, the training node performs global model aggregation and parameter calculation on related parameters, thereby the model training and parameter transmission process of the federated learning task is performed in the computing sandbox state channel, the credibility of the calculation in the federated learning training process and the privacy security of user data are ensured, and the trusted supervision in the federated learning model training and aggregation process is realized.
[0120] To achieve one of the above purposes, the third technical solution of the application is:
[0121] A computer device comprises:
[0122] One or more processors;
[0123] A storage device for storing one or more programs;
[0124] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned federated learning multi-task scheduling method based on a trusted computing sandbox.
[0125] To achieve one of the above purposes, the fourth technical solution of the application is:
[0126] A computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the above-mentioned federated learning multi-task scheduling method based on a trusted computing sandbox.
[0127] Through continuous exploration and experiments, the application constructs a federated learning training trusted supervision framework based on a computing sandbox state channel, the model training and parameter transmission process of the federated learning task is performed in the computing sandbox state channel, the credibility of the calculation in the federated learning training process and the privacy security of user data are ensured, and the trusted supervision in the federated learning model training and aggregation process is realized.
[0128] Compared with the prior art, the application has the following beneficial effects:
[0129] The present application constructs a participant node, a blockchain scheduling model and a training node through continuous exploration and experiments; then the blockchain scheduling model establishes a computing sandbox state channel according to the task request of the participant node; in the computing sandbox state channel, the training node performs global model aggregation and parameter calculation on the related parameters to obtain the training result and state information of the federated learning task; the training result and state information are transmitted to the blockchain scheduling model, and the computing sandbox state channel is closed to release the occupied computing power resources; the blockchain scheduling model sends the training result to the participant node, updates the state information of the computing power resources, completes the scheduling of the federated learning multi-task, realizes the trusted supervision in the federated learning model training and aggregation process, and can effectively avoid the privacy leakage caused by malicious node attacks, and the scheme is scientific, reasonable and feasible.
[0130] Further, the present application constructs a federated learning training trusted supervision framework based on a computing sandbox state channel, and puts the model training and parameter transmission process of the federated learning task into the computing sandbox state channel, thereby ensuring the trustworthiness of the calculation and the privacy security of the user data in the federated learning training process, and realizing the trusted supervision in the federated learning model training and aggregation process.
[0131] Further, the present application constructs a federated learning training trusted supervision framework based on a computing sandbox state channel, and puts the model training and parameter transmission process of the federated learning task into the computing sandbox state channel, thereby ensuring the trustworthiness of the calculation and the privacy security of the user data in the federated learning training process, and realizing the trusted supervision in the federated learning model training and aggregation process. BRIEF DESCRIPTION OF DRAWINGS
[0132] Figure 1 The first flowchart of the federated learning multi-task scheduling method of the present application;
[0133] Figure 2 The system block diagram of the federated learning multi-task scheduling system of the present application;
[0134] Figure 3 The second flowchart of the federated learning multi-task scheduling method of the present application;
[0135] Figure 4 The flowchart of the resource scheduling algorithm of the present application. DETAILED DESCRIPTION
[0136] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and are not intended to limit the present application.
[0137] On the contrary, the present application covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present application as defined by the claims. Further, in order to make the public have a better understanding of the present application, some specific details are described in detail in the following detailed description of the present application. The present application can also be completely understood without the description of these details by those skilled in the art.
[0138] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present application. The term "or / and" used herein includes any and all combinations of one or more related listed items.
[0139] As shown in Figure 1 The first specific embodiment of the present application based on the federated learning multi-task scheduling method of the trusted computing sandbox is as follows:
[0140] A federated learning multi-task scheduling method based on a trusted computing sandbox, comprising the following steps:
[0141] Firstly, one or more requested federated learning tasks are obtained through a pre-constructed participant node;
[0142] Secondly, the participant node sets relevant parameters according to the characteristics of the federated learning task, and uses a digital signature technology to sign an intelligent contract with a pre-constructed blockchain scheduling model;
[0143] After signing, the relevant parameters are transmitted to the blockchain scheduling model;
[0144] Thirdly, the blockchain scheduling model selects a training node according to the relevant parameters of the federated learning task, allocates and schedules computing power resources, and establishes a computing sandbox state channel;
[0145] Fourthly, the training node performs global model aggregation and parameter calculation on the relevant parameters in the computing sandbox state channel, and obtains the training result and state information of the federated learning task;
[0146] Fifthly, the training result and state information are transmitted to the blockchain scheduling model, and the computing sandbox state channel is closed to release the occupied computing power resources;
[0147] In the sixth step, the blockchain scheduling model sends the training result to the participant node, updates the state information of the computing resource, and waits for a new federated learning task to arrive to allocate and schedule the computing resource, so as to realize the federated learning multi-task scheduling based on the trusted computing sandbox.
[0148] As shown in Figure 2 , a specific embodiment of the federated learning multi-task scheduling system based on the trusted computing sandbox of the present application is as follows:
[0149] A federated learning multi-task scheduling system based on a trusted computing sandbox introduces the concept of a computing sandbox in trusted computing, and constructs a federated learning training framework based on a computing sandbox state channel in a blockchain scenario. The model parameter transmission and aggregation of the federated learning task are carried out in the computing sandbox state channel and are supervised, and the occupied resources are released after the task is completed. Then the result is transmitted to the requester in a trusted manner. This process supports the trusted computing and sharing of data, and is invisible to the participants of the training. The whole system framework mainly includes three parts: participant node, blockchain scheduling module and computing sandbox state channel.
[0150] The participant node can act as a requester of the federated learning task, or as a local training node participating in the federated learning task. Each participant has user data and resources with training value, and represents an industry or enterprise company such as a financial institution or a credit investigation department. The participant node does not want to share user data with other parties, but wants to use local data to jointly train a global model through federated learning and other methods to share the value of data. The participant node can apply to the blockchain scheduling module to register as a training node, and apply to cooperate with other participant nodes to perform a federated learning task and train a global model.
[0151] In the traditional federated learning training process, the locally trained model parameters need to be sent to a central server or an aggregation node, and the global model is aggregated on the aggregation node. In this case, a single-point attack on the aggregation node can easily cause the collapse of the entire federated learning system, the leakage of local data privacy and other security problems.
[0152] To this end, the present application introduces a computing sandbox as a computing sandbox state channel, and the global model aggregation and parameter calculation of the federated learning task rely on the computing resources and space provided by each participant, and the computing process is placed in the computing sandbox state channel for execution under the supervision of the blockchain scheduling module. In the process of constructing the computing sandbox state channel, the blockchain scheduling module reasonably schedules resources to the federated learning task according to the needs of different tasks, so that the entire system works more efficiently. The pros and cons of the scheduling strategy determine the efficiency and resource utilization of the entire system. The federated learning task does not own or retain the allocated resources, but the system dynamically allocates resources using a scheduling algorithm according to current needs to fully utilize resources and release the used resources immediately after the task is completed.
[0153] The blockchain scheduling module is used to manage the states and verification information of each participant, integrate and virtualize various resources registered by each participant, such as computing power resources and storage space, and coordinate and supervise each participant for federated learning training through a smart contract. Considering the resource heterogeneity of different participant nodes, the computing power and storage resources of different participant nodes are not the same, and some participant nodes do not have the ability to independently complete model training tasks. Integrating and reasonably allocating and scheduling the resources of all nodes to construct a computing sandbox state channel can well coordinate and solve the problem of resource heterogeneity. When a task request arrives, the blockchain scheduling module schedules resources, allocates resources to the task, and establishes a computing sandbox state channel. The training of the federated learning task model and the transmission of the parameters will all be carried out in the computing sandbox state channel, and any malicious behavior will be supervised by all nodes and reported to the blockchain scheduling module to ensure the trusted computing of data and the security of the model training aggregation process.
[0154] In order to protect the privacy and security of user data, the real physical address of the computing sandbox state channel should not be disclosed to all participant nodes, and the blockchain scheduling module is responsible for shielding the real physical address of the computing sandbox state channel from the participant nodes and only providing virtual addresses and interfaces to the participant nodes. The federated learning task is completed in the computing sandbox state channel, and each participant node cannot obtain the models and parameters of other participant nodes, so it cannot attack the original data.
[0155] As shown in Figure 3 The second specific embodiment of the federated learning multi-task scheduling method based on the trusted computing sandbox is as follows:
[0156] The federated learning multi-task scheduling method based on the trusted computing sandbox comprises the following steps:
[0157] First, the participant who has user data wants to instantiate a federated learning task, aggregate a global model, and share the value of data. First, the participant needs to apply for registration as a training node to the blockchain scheduling model, submit identity information, and describe the data owned and the resources that can be provided, including the calculation unit operation rate or / and the storage space size or / and the cost, etc., to facilitate the blockchain scheduling model to allocate and schedule.
[0158] Second, the blockchain scheduling model verifies the identity of the applied node and integrates the resources that the node can provide, establishes a virtual resource pool, maintains a virtual resource state table, records the allocation and use of resources, and the mapping of virtual ports to physical addresses, etc.
[0159] Third, the participant node applies to instantiate a federated learning task, and the blockchain scheduling model selects the training node according to the data description information provided by each training node during registration.
[0160] Fourth, the participants negotiate the related parameters of the federated learning task, including the number of training rounds, the initial model, the total budget, etc., and use digital signature technology to sign a smart contract to instantiate a federated learning task.
[0161] Fifth, the blockchain scheduling model determines the initialization priority of the task and places it in the task queue, waiting for the allocation and scheduling of computing resources.
[0162] Sixth, the blockchain scheduling model reasonably allocates the federated learning task in the waiting state to different virtual computing units. The blockchain scheduling model obtains the physical address of the virtual resource by looking up the resource state table, and encrypts the information of the resource physical address for building a computing sandbox state channel, establishes a trusted computing sandbox as a computing sandbox state channel and updates the resource usage state. The participant node only knows the information of the allocated virtual resource and cannot obtain the real physical address of the computing sandbox, so it cannot maliciously attack the model in training.
[0163] Seventh, the training node completes the federated learning training task in the computing sandbox state channel, aggregates and updates the model, and the behavior and state of each node are supervised by all nodes. Once malicious behavior is found, it is immediately reported to the blockchain scheduling model and the malicious node is punished accordingly.
[0164] Eighth, after completing the federated learning task, the training node writes the final result and state information to the blockchain scheduling model and immediately releases the occupied resources and closes the computing sandbox state channel.
[0165] Ninth, the blockchain scheduling model sends the final result to the task requester and updates the virtual resource state information, waiting for new tasks to come for resource allocation and scheduling.
[0166] The resource scheduling method of the application can make the whole system work more efficiently while meeting the quality of service requirements of the task as much as possible.
[0167] A specific embodiment of the application for modeling the resource scheduling problem is as follows:
[0168] In the process of constructing the computing sandbox state channel for the federated learning task, different resource scheduling schemes may affect the working efficiency of the system and the service completion quality of the task. Unreasonable scheduling strategy may cause some nodes to have too large computing load and thus fail, while the resource utilization of other nodes decreases, causing resource waste and greatly reducing the system efficiency. It may also appear that unreasonable resource allocation scheme makes it difficult to meet the service quality of the task, such as exceeding the budget and the expected completion time. The application constructs a resource scheduling optimization problem model under the federated learning training supervision mechanism based on the computing sandbox state channel.
[0169] The application defines the resource scheduling problem as how to schedule and allocate multiple federated tasks to multiple computing nodes, such as anomaly behavior detection tasks, risk assessment tasks, customer behavior analysis tasks, product intelligent recommendation tasks, etc., to obtain a scheduling scheme that minimizes the completion time of the whole system under the constraints of task cost and completion time.
[0170] Suppose the set of training nodes of the participant is
[0171]
[0172] where Nodes represents the set of nodes providing computing resources, N i represents the i-th node, and N node is the number of nodes. The nodes are represented as
[0173]
[0174] where PU j (j=1, 2,..., N pu ) represents the j-th processing unit, represents the number of processing units of the i-th node. The processing unit PU is represented as
[0175] PU={SIDP, E, COST, TD}
[0176] where SIDP is the number of the processing unit, E is the execution capability of the processing unit, which is the amount of work that can be processed per unit time, COST is the cost of using the processing unit to execute per unit time, and TD represents the communication delay of the node where the PU is located to other nodes, TD i,jdenotes the communication latency from node i to node j.
[0177] Let the set of federated learning tasks be Tasks.
[0178]
[0179] where Tasks is the set of federated learning tasks, T i denotes the i-th task, N task denotes the number of tasks. A task T is represented as
[0180] T = {SIDT, WorkLoad,
[0181] MaxT, MaxC, PI, Ntrain}
[0182] where SIDP is the number of the federated learning task, WorkLoad represents the amount of its task, MaxT represents the maximum completion time that the task can tolerate, MaxC represents the maximum cost that the task can bear, MaxT and MaxC are taken as indicators of the quality of service requirements of the task in the present application, PI represents the priority of the task, N train denotes the set of training nodes participating in the task.
[0183] Therefore, the present application can use a matrix to represent the task resource allocation scheme, and the allocation matrix is defined as follows
[0184]
[0185] where X is an N task *N pu size task resource allocation matrix, x i,j denotes the allocation of the i-th task on the j-th computing unit, which is defined as
[0186]
[0187] According to the task resource allocation matrix, the completion time of each task can be calculated, and the calculation formula is as follows
[0188]
[0189] where ECT i denotes the completion time of the i-th task, WorkLoad i represents the amount of work of the i-th task, M pu represents the number of computing units, E j represents the execution capability of the j-th computing unit, x i,j represents the allocation of the i-th task on the j-th computing unit. Similarly, the cost of each task can be calculated, and the calculation formula is as follows
[0190]
[0191] TCost i represent the cost overhead of the ith task, COST j represent the execution cost of the jth computing unit. The formula of the system maximum completion time maxMakespan can be obtained by the present application as follows
[0192]
[0193]
[0194] Finally, the present application abstracts the resource scheduling into a target optimization problem, i.e. solving the resource allocation matrix X, satisfying
[0195] min
[0196]
[0197] and TCost i ≤MaxC i
[0198] The specific system parameters are shown in Table 1.
[0199] Table 1 System parameters
[0200]
[0201] A specific embodiment of the resource scheduling algorithm based on A2C of the present application is as follows:
[0202] In a complex and variable federated learning training environment based on a blockchain scheduling model, the resource scheduling algorithm needs to select the optimal resource allocation scheme according to the current node resource state and the change of the federated learning task to be allocated, construct a computing sandbox state channel, and make the system work more efficiently while meeting the task service quality requirements to the greatest extent. The present application designs an Actor-Critic resource scheduling algorithm based on DRL, and the algorithm framework is as shown in Figure 4 The environment is mainly divided into two parts of environment and agent.
[0203] The environment is responsible for supervising the state of each node and managing the computing resources they have, and maintaining the task list waiting for resource allocation, and providing the current state information to the agent, including the current resource state, the state of the task to be allocated, etc. The environment also needs to return the reward of the current operation, the next time state, etc. after the agent makes an action.
[0204] The agent is a resource scheduler, which can make corresponding decisions and actions according to the state given by the environment, and the system allocates computing resources to the task according to the action. The agent is composed of a Critic value network and an Actor policy network. The Critic network can score the current state and evaluate the goodness of the state, and the Actor network selects the optimal action according to the environment state.
[0205] The traditional Q learning has the disadvantages of wasting previous experience and the correlation of parameter updating, and the method of experience replay is adopted in the application, an experience buffer pool R is set, the environment state, the executed action, the obtained reward and the next time state are stored as experience samples in the buffer pool, and when training the network parameters, a small batch of data can be randomly extracted from the buffer pool for training and parameter updating, so that the previous experience is better utilized and the correlation of parameter updating is broken.
[0206] Before resource scheduling, the system needs to sort the tasks waiting for resource allocation according to the priority at the current time, and allocate resources to them in turn. It is assumed that the priority of each task is The priority of the i-th task is represented as:
[0207] PI i =initial_pr i +wait_time i
[0208] Where initial_pr i is the initial priority of the task, which is related to the workload W i of the task. The less the workload, the higher the priority, which takes advantage of the short job priority and is beneficial to reduce the average waiting time, and wait_time i is the time of the task waiting for resource allocation, which is a dynamic value. The longer the waiting time, the higher the priority, which is beneficial to reduce the job starvation problem.
[0209] When training the Actor-Critic algorithm network, the system makes action a t based on the policy network according to the current environment state s t , obtains reward r t through interaction with the environment, and gets new environment state s t+1 , and (s t , a t , r t , s t+1 ) is put into the buffer pool R as a sample. Then N groups of data samples are randomly extracted from the buffer pool, and the average value of the gradient is calculated to update the Actor-Critic network parameters.
[0210] To solve the proposed resource scheduling optimization problem using deep reinforcement learning method, the problem is further described as an MDP model, which is specifically designed as follows.
[0211] State space S, the state space should include the allocation of each computing resource at time t (task resource allocation matrix), the specific state of the task to be allocated at time t, such as the workload, budget, and maximum completion time that can be tolerated. Therefore, the state space is defined as
[0212] s t ={X t ,WorkLoad i ,MaxT i ,MaxC i}
[0213] Where s t represents the state at time t, WorkLoad i ,
[0214] MaxT i , MaxC i respectively represent the workload of the i-th task to be allocated, the maximum completion time that can be tolerated, and the maximum budget.
[0215] Action space A, under the s t state, the agent performs the corresponding action, takes out the task with the highest priority from the waiting queue and allocates resources for it, and the action is defined as follows:
[0216]
[0217] Where a t represents the action at time t, N pu is the number of resources of the current system, γ i ∈{0,1} represents the allocation of the task on the i-th resource, 0 represents no allocation, and 1 represents allocation.
[0218] Reward R, the reward of the action performed is calculated according to whether the quality of service of the task is met after the resource is allocated, and the system completion time. If the action of resource allocation causes the completion time and cost of the task to be greater than the maximum value that can be tolerated by the task, a huge punishment will be obtained, in addition, the maximum completion time of the system is taken as the standard of the reward. The reward can be defined as
[0219]
[0220] Where r t represents the reward obtained by performing a t action under s t state, maxMakespan represents the maximum completion time of the system, and ECTi , Tcost i , MaxT i , MaxC i respectively represent the completion time, cost overhead, maximum completion time that can be tolerated and budget overhead of task i, and epsilon is a parameter set by the system.
[0221] The specific steps of the resource scheduling algorithm can be seen from Table 2.
[0222] Table 2
[0223]
[0224]
[0225] As shown in Table 2, the specific algorithm trains M generations in total, performs T times of network parameter updates in each generation, randomly initializes a task waiting queue at the beginning of each generation, and observes the environment to obtain a random initial state. Before each parameter update, the system obtains a plurality of (s t , a t , r t , s t+1 ) samples from the environment and puts them into the buffer pool R. When the amount of data in R reaches a set value, N groups of samples are randomly extracted from R to perform batch gradient update on the network parameters.
[0226] The update process first uses the formula y i =r i +γv(s i+1 ;w′) and the formula δ i =v(s i ;w)-y i to calculate TDtarget and TD error, and according to the formula , the gradient of the policy network is calculated, y i is used as the target value label of the value network, and the loss function is calculated according to the least square difference Thus, the gradient calculation formula of the value network is obtained The batch update method is used to update the parameters of the policy network and the value network respectively. After iteration k times, the parameters of the target value network are updated.
[0227] An embodiment of a device applying the method of the application is provided.
[0228] A computer device comprises:
[0229] one or more processors;
[0230] a storage device for storing one or more programs;
[0231] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned federated learning multi-task scheduling method based on a trusted computing sandbox.
[0232] A computer medium embodiment of the method of the application is as follows:
[0233] A computer readable storage medium having stored thereon a computer program, the program being executed by a processor to implement the above-mentioned federated learning multi-task scheduling method based on a trusted computing sandbox.
[0234] Those skilled in the art will understand that the embodiments of the present application can be provided as methods, systems, computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) having computer-usable program code embodied therein.
[0235] The present application is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowcharts and / or block diagrams block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified in the flowchart block or blocks.
[0236] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions means which implement the function specified in the flowcharts and / or block diagrams block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified in the flowchart block or blocks.
[0237] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowcharts and / or block diagrams block or blocks. Figure 1 one or more flows and / or blocks Figure 1steps of the functions specified in the one or more blocks.
[0238] It should be noted that the above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the technical solutions of the present application can still be modified or equivalent replaced without departing from the spirit and scope of the present application, and any modification or equivalent replacement should be covered in the protection scope of the claims of the present application.
Claims
1. A method for multi-task scheduling of federated learning based on a trusted computing sandbox, comprising the following steps: obtaining one or more federated learning tasks to be requested by a pre-constructed participant node; setting relevant parameters according to the characteristics of the federated learning task, and signing an intelligent contract with a pre-constructed blockchain scheduling model using digital signature technology; after signing, transmitting the relevant parameters to the blockchain scheduling model; the blockchain scheduling model selects training nodes according to the relevant parameters of the federated learning task, and allocates and schedules computing resources to establish a computing sandbox state channel; the blockchain scheduling model is used to manage the state and verification information of each participant node, integrate and virtualize various computing resources registered by each participant node, and coordinate and supervise each participant node for federated learning training through an intelligent contract; the method for establishing a computing sandbox state channel by the blockchain scheduling model is as follows: in the process of constructing the computing sandbox state channel, the blockchain scheduling model reasonably schedules resources to the federated learning task according to the requirements of different federated learning tasks; at the same time, the blockchain scheduling model dynamically schedules and allocates computing resources according to the requirements of the current federated learning task to establish a computing sandbox state channel, and releases the used computing resources immediately after the task is completed, and the federated learning task does not own or retain the allocated computing resources; then the training of the federated learning task and the transmission of the parameters are carried out in the computing sandbox state channel, and any malicious behavior will be supervised by all training nodes and reported to the blockchain scheduling model; malicious behavior includes data poisoning and model poisoning; data poisoning refers to data pollution; model poisoning refers to sending incorrect model parameters or damaged model parameters; the real physical address of the computing sandbox state channel is not disclosed to all participant nodes and training nodes, and the blockchain scheduling model is responsible for shielding the real physical address of the computing sandbox state channel from the participant nodes and the training nodes, and only provides virtual addresses and interfaces to the participant nodes and the training nodes; the training nodes perform global model aggregation and parameter calculation in the computing sandbox state channel to obtain the training results and state information of the federated learning task; transmit the training results and state information to the blockchain scheduling model, and close the computing sandbox state channel to release the occupied computing resources; the blockchain scheduling model sends the training results to the participant nodes, updates the state information of the computing resources, and waits for new federated learning tasks to allocate and schedule computing resources, thereby realizing multi-task scheduling of federated learning based on a trusted computing sandbox. 2.The method of claim 1, wherein the blockchain scheduling model selects training nodes as follows: different training nodes have resource heterogeneity, and are matched according to the computing resources possessed by the training nodes; when the training nodes do not have the ability to independently complete the model training task, the resources of several training nodes are integrated and reasonably allocated and scheduled to construct a composite training node. 3. The federated learning multi-task scheduling method based on the trusted computing sandbox according to claim 1, wherein the participant node is a requester of the federated learning task and can also serve as a training node participating in the federated learning task, and has user data and resources with training value for representing an industry or a company; The participant node applies to the blockchain scheduling model to register as a training node and apply to cooperate with other participant nodes to perform the federated learning task; The participant node trains a global model using local data in a federated learning manner to share the value of the data, but does not share user data with other participant nodes.
4. The federated learning multi-task scheduling method based on the trusted computing sandbox according to claim 1, wherein the blockchain scheduling model can schedule and distribute multiple federated learning tasks to multiple training nodes to minimize the completion time of the entire system under the constraints of task cost and completion time, and the construction process is as follows: A node set of training nodes is constructed, and the expression is as follows: The expression of the training node is as follows: A federated learning task set is constructed, and the expression is as follows: wherein represent a set of nodes providing computing resources, represent a first training node, is the number of training nodes; A task resource allocation matrix is constructed to represent the task resource allocation scheme, and the calculation formula is as follows: wherein represents the first block processing unit, represents the number of processing units of the first training node, ; Processing unit The expression of the formula is as follows: wherein is the number of processing units, is the execution capacity of the processing units, i.e. the amount of work that can be processed per time unit, is the cost that has to be paid for using the processing units per time unit, denotes the communication latency from the training node to the other training nodes, denotes the communication latency from the training node to the training node ; The completion time of each federated learning task is calculated according to the task resource allocation matrix, and the calculation formula is as follows wherein is a set of federated learning tasks, represents the i-th federated learning task, represents the number of federated learning tasks; Federated learning task The expression of the above is as follows: Wherein is the number of the federal learning task, represents the amount of its task, represents the maximum completion time that the task can tolerate, represents the maximum cost that the task can bear, and and as an index of the quality of service requirements of the task, represents the priority of the task, represents the node set participating in the federal learning task this time; The cost of each federated learning task is calculated, and the calculation formula is as follows wherein is a matrix of task resource allocation of size represents the allocation of the th task on the th block computing unit, which is expressed as follows: Thus, the construction of the blockchain scheduling model is completed. (1) wherein denotes the completion time of the federated learning task, denotes the workload of the federated learning task, denotes the number of computing units, denotes the execution capability of the block computing unit, denotes the allocation of the federated learning task to the block computing unit; 5. The federated learning multi-task scheduling method based on the trusted computing sandbox according to claim 4, wherein the method for the blockchain scheduling model to schedule and allocate computing resources is as follows (2) Representing the The cost of the Federal Learning Mission Indicates the first The execution cost of a block computing unit; Computing system maximum completion time The formula for this is as follows: (3) Finally, the resource scheduling is abstracted as a target optimization problem, i.e. solving the resource allocation matrix which satisfies minimum; ; The blockchain scheduling model selects the optimal resource allocation scheme according to the current node resource state and the change of the federated learning task to be allocated, constructs a computing sandbox state channel, which includes an environment supervision unit and an agent decision unit; The environment supervision unit is used to supervise the state of each node and manage the computing resources they have, and maintains a task list waiting for resource allocation, and provides the current state information to the agent decision unit, including the current resource state and the state of the task to be allocated; The environment supervision unit returns the reward of the current operation and the next time state after the agent decision unit makes a move; The agent decision unit is a resource scheduler that can make corresponding decisions according to the state provided by the environment supervision unit, make corresponding actions, and allocate computing resources to tasks according to the actions; The agent decision unit is composed of a value network Critic and a policy network Actor, the value network Critic network can score the current state to evaluate the state, and the policy network Actor network selects the optimal action according to the state of the environment supervision unit; Before resource scheduling, the federated learning tasks waiting for resource allocation at the current time are sorted according to priority, and resources are allocated to them in turn; Then N groups of data samples are randomly extracted from the buffer pool, and the average value of the gradient is calculated to update the parameters of the policy network Actor and the value network Critic. Priorities of individual federated learning tasks wherein the priority of the th federated learning task is represented as: wherein is the initial priority of the federated learning task, and the workload of the federated learning task is related, the less the workload, the higher the priority, is the time for which the federated learning task has been waiting for resource allocation, which is a dynamically changing value, the longer the waiting time, the higher the priority, which is conducive to reducing the problem of job starvation; When training the policy network Actor and the value network Critic, a current environment regulatory cell state is based on the policy network to make an action , a reward is obtained by interacting with the environment regulatory cell , and a new environment regulatory cell state is obtained as a sample into the buffer pool R; 6. The federated learning multi-task scheduling method based on trusted computing sandbox according to claim 5, characterized in that, The blockchain scheduling model uses a deep reinforcement learning method to solve resource scheduling optimization, which includes the following contents: Constructing a state space , the state space comprises a learning task resource allocation matrix of each computing resource at a moment, a specific state of a federated learning task to be allocated at a moment; The specific state of the federated learning task to be allocated includes the workload, budget and maximum completion time that can be tolerated of the federated learning task; The expression of the state space is as follows: wherein denotes the state at the moment, , , denote the workload, the maximum tolerable completion time, the maximum budget expenditure of the current to be assigned federated learning task of the number . Action space In the state, the agent performs the corresponding action, takes out the federated learning task with the highest priority from the waiting queue and allocates resources for it, and the expression of the action is as follows: wherein represents the action at the moment, is the number of resources of the current system, represents the allocation of the federated learning task to the resources, 0 represents unallocated, and 1 represents allocated; Return , calculate the return of the action performed according to whether the service quality of the federated learning task is met after the allocation of resources, and the system completion time If the action of resource allocation causes the completion time and cost of the federated learning task to be greater than the maximum value that can be tolerated by the federated learning task, a huge penalty will be obtained, and the maximum completion time of the system is used as the standard of reward; The expression of the reward is as follows: (17) in Indicates in Execute in state The reward obtained from the action Indicates the system's maximum completion time. These represent federal learning tasks. The completion time, cost, maximum tolerable completion time, and budget. Parameters set for the system.
7. The federated learning multi-task scheduling method based on trusted computing sandbox according to claim 6, characterized in that, The specific steps of solving include the following contents: A total of M generations are trained, T times of network parameter updates are performed in each generation, a federated learning task waiting queue is randomly initialized at the beginning of each generation, and an initial state is obtained by observing the environment supervision unit; Before updating the parameters each time, the system obtains a plurality of The data in the buffer pool R is put into the buffer pool R; when the amount of data in R reaches a set value, N groups of samples are randomly selected from R to perform batch gradient update on the network parameters; The target value label and target error are calculated in the update process, the gradient of the policy network is calculated, and the loss function is calculated according to the least square difference; The gradient calculation formula of the value network is obtained; the batch update method is used to update the parameters of the policy network and the value network respectively; and the parameters of the target value network are updated after k iterations.
8. A federated learning multi-task scheduling system based on trusted computing sandbox, characterized in that, The federated learning multi-task scheduling method based on trusted computing sandbox according to any one of claims 1-7 includes a number of participant nodes, a blockchain scheduling module and training nodes: The participant node is used to obtain one or more federated learning tasks to be requested, set relevant parameters according to the characteristics of the federated learning task, and sign an intelligent contract with the blockchain scheduling module using digital signature technology; The blockchain scheduling module selects a training node according to the relevant parameters of the federated learning task, allocates and schedules computing power resources, and establishes a computing sandbox state channel; The training node aggregates global modules and calculates parameters in the computing sandbox state channel to obtain the training result and state information of the federated learning task, and transmits the training result and state information to the blockchain scheduling module and closes the computing sandbox state channel to release the occupied computing power resources; The blockchain scheduling module sends the training result to the participant node, updates the state information of the computing power resources, and waits for new federated learning tasks to allocate and schedule computing power resources, thereby realizing federated learning multi-task scheduling based on trusted computing sandbox.
9. A computer device, characterized in that, It includes: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the federated learning multi-task scheduling method based on trusted computing sandbox according to any one of claims 1-7.
10. A computer readable storage medium, characterized in that, A computer program is stored thereon, and the program is executed by a processor to implement a federated learning multi-task scheduling method based on a trusted computing sandbox according to any one of claims 1-7.