Distributed training task scheduling method and system based on central node management, terminal and storage medium

Through the distributed training task scheduling method based on central node management, dynamic allocation and efficient execution of training tasks are solved, and the problems of uneven task allocation and low resource utilization efficiency in the existing system are achieved, and the optimal utilization of resources and the guarantee of training quality are achieved.

CN119988014APending Publication Date: 2025-05-13SHANDONG UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510074553.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing distributed training system has problems in uneven task allocation, low resource utilization efficiency, and difficult to ensure training quality. Especially in an environment where computing power resources are dispersed and highly heterogeneous, how to efficiently schedule and manage training tasks and ensure the optimal utilization of training quality and resources has become the focus of technicians.

Method used

The distributed training task scheduling method based on central node management is adopted, and the training tasks are dynamically allocated through modules such as identity authentication, task release and points management to ensure efficient task execution and result verification. This method incorporates fixed computing nodes and idle computing nodes into the distributed training network, uses different authentication methods to protect the rights and interests of nodes, introduces intelligent task scheduling and flexible points reward and punishment rules, and encourages nodes to actively participate.

Benefits of technology

The dynamic allocation and efficient execution of training tasks are realized, the resource utilization is maximized, the training quality and active participation of nodes are ensured, and the problem of uneven task allocation is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988014A_ABST
    Figure CN119988014A_ABST
Patent Text Reader

Abstract

The invention relates to a distributed training task scheduling method and system based on central node management, a terminal and a storage medium. The method comprises the following steps: completing an identity authentication process according to a node type and a training node; broadcasting a training task execution request to the training nodes, receiving a task receiving request from the training nodes in a task receiving window period, selecting the training node with the highest weight integral as a target node, and issuing a training task; training results from the training nodes are received, each training result corresponds to one training task, a reward and punishment integral of the training task is generated according to the sending time of the training results and the basic information of the training tasks, and the reward and punishment integral is used for adjusting and updating the weight integral of the current training node. Through intelligent management of the center node, dynamic allocation, efficient execution and result verification of the training tasks are realized, and meanwhile, a point reward and punishment mechanism motivates positive participation and high-quality contribution of the training nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of industrial Internet and industrial intelligent applications, and in particular to a distributed training task scheduling method, system, terminal and storage medium based on central node management for large-scale distributed training. Background Art

[0002] Currently, with the rapid development of artificial intelligence technology, large-scale distributed training has become an important means to improve model performance and efficiency. However, existing distributed training systems often have problems such as uneven task distribution, low resource utilization efficiency, and difficulty in ensuring training quality. Especially in an environment where computing resources are dispersed and heterogeneous, how to efficiently schedule and manage training tasks, ensure training quality and optimal resource utilization has become a problem of widespread concern to technicians in this field.

[0003] A specific analysis of the task scheduling schemes followed by existing distributed training systems reveals that the more typical ones include the following:

[0004] First, the task scheduling scheme based on static allocation. This scheme usually statically allocates training tasks to each node according to preset rules or strategies. For example, training tasks are allocated based on hardware indicators such as the computing power and network bandwidth of the node. However, this scheme ignores the real-time status and dynamic changes of the node, which can easily lead to uneven resource utilization and inefficient task execution.

[0005] Second, the polling-based task scheduling scheme. This scheme assigns tasks to each node in turn according to a fixed order. Although the specific implementation process of this scheme is relatively simple, it also does not take into account the real-time status and dynamic changes of the nodes, and therefore cannot achieve optimal resource utilization.

[0006] Third, the priority-based task scheduling scheme. This scheme assigns tasks based on the priority of the task and the real-time status of the node. Tasks with high priority will be assigned to nodes with better performance to ensure timely completion of the task. However, this scheme still has certain limitations. For example, it may not fully utilize the computing power of all nodes, especially when the task priorities vary greatly.

[0007] It can be seen that although the above typical task scheduling schemes have solved the problems of resource utilization and task execution efficiency to a certain extent, they still have many shortcomings and limitations. Therefore, how to propose a new training task scheduling scheme for large-scale distributed training to achieve dynamic allocation, efficient execution and result verification of training tasks has become an urgent problem to be solved by technicians in this field. Summary of the invention

[0008] In order to achieve dynamic allocation, efficient execution and result verification of training tasks in large-scale distributed training, the present application provides a distributed training task scheduling method, system, terminal and storage medium based on central node management.

[0009] In a first aspect, the present application provides a distributed training task scheduling method based on central node management, adopting the technical solution described below.

[0010] A distributed training task scheduling method based on central node management, characterized in that, for training nodes of different node types, the node types include fixed computing power nodes and idle computing power nodes, the method comprises the following steps:

[0011] Complete the identity authentication process with the corresponding training nodes one by one according to the node type, and maintain the communication connection with the training nodes;

[0012] Broadcasting a training task execution request to the training node according to the training task basic information, receiving a task collection request from the training node within a task collection window period, wherein the task collection request is generated by the training node according to the training task execution request and its own computing power load status, selecting a training node with the highest weight score from all the training nodes that sent the task collection request as the target node, and sending the training task corresponding to the training task basic information to the target node;

[0013] Receive training results from any of the training nodes, each of the training results corresponds to a training task, generate reward and punishment points for the training task based on the sending time of the training results and the basic information of the training task corresponding to the training task, and use the reward and punishment points to adjust and update the weight points of the current training node.

[0014] By adopting the above technical solution, fixed computing power nodes and idle computing power nodes are used as training nodes and included in the distributed training network. During the authentication process of training nodes, different node types use different authentication methods, which maximizes the protection of the rights and interests of all training nodes participating in the training, encourages individual idle computing power to actively participate in training, and greatly improves the overall resource utilization.

[0015] At the same time, the use of intelligent task scheduling and flexible points reward and punishment rules can achieve effective allocation and efficient execution of training tasks, fully mobilizing the enthusiasm of each training node to participate in training.

[0016] Preferably, completing the identity authentication process with the training nodes one by one according to the types of the training nodes comprises the following steps:

[0017] When the training node is the fixed computing power node,

[0018] Receive a first identity authentication request from the training node, the first identity authentication request comprising a first request header and a first request signature, the first request header comprising an access key and a first signature value, the first request signature comprising a first signature content obtained by processing with a confidentiality key matching the access key, the access key and the confidentiality key matching and corresponding to each other and both being stored in a first key library;

[0019] The corresponding confidential key is searched in the first key library according to the access key, and the signature value of the first request signature is calculated in combination with the search result and the preset first signature verification rule. The calculation result is compared with the first signature value contained in the first request packet header for consistency. If the comparison result is consistent, it is determined that the identity authentication of the current training node is successful, and a communication connection is established with the current training node using the preset encrypted transmission channel.

[0020] By adopting the above technical solution, identity authentication for fixed computing power nodes is realized based on the AK / SK solution, flexible participation of cross-domain training nodes in the distributed training process is realized, and the security of data transmission when accessed through the public network is effectively guaranteed.

[0021] Preferably, completing the identity authentication process with the training nodes one by one according to the types of the training nodes comprises the following steps:

[0022] When the training node is the idle computing power node,

[0023] Receive a second identity authentication request from the training node, the second identity authentication request includes a second request header and a second request signature, the second request header includes a public key and a second signature value, the second request signature includes a second signature content obtained after processing with a private key matching the public key, the public key and the private key match and correspond to each other and both are generated offline by the current training node and stored locally, and the public key submitted for registration by the training node is stored in a second key library;

[0024] Confirm whether the public key exists in the second key library. If so, calculate the signature value of the second request signature in combination with the public key and the preset second name verification rule, and compare the calculation result with the second signature value contained in the second request packet header for consistency. If the comparison results are consistent, it is determined that the identity authentication of the current training node is successful, and a communication connection is established with the current training node.

[0025] By adopting the above technical solution, identity authentication for fixed computing power nodes is realized based on the public-private key solution. Since the public-private key solution adopts asymmetric encryption, even if the public key is leaked, it will not affect the security of the private key, thereby realizing secure network communication and data transmission.

[0026] Preferably, broadcasting a training task execution request to the training node according to the training task basic information, and receiving a task receiving request from the training node within a task receiving window period, comprises the following steps:

[0027] Generate a corresponding training task execution request according to the training task basic information of the current training task, wherein the training task basic information includes the task algorithm type, the task data volume, the task priority and the task deadline, and broadcast the training task execution request to the training node;

[0028] A task collection request is received from the training node within a task collection window period, the task window period matches the task deadline, and the task collection request is generated by the training node based on its own computing power load status and corresponds to the training task execution request.

[0029] By adopting the above technical solution, the entire training task is released to all training nodes, ensuring the effective participation of each independent training node in the distributed training network. The training node needs to respond to the training task execution request after comprehensive evaluation, which is also to ensure that each training task can be completed efficiently after being received, ensuring the overall execution efficiency of the solution.

[0030] Preferably, the step of selecting a training node with the highest weight score from all the training nodes that send the task claim request as the target node and sending the training task corresponding to the training task basic information to the target node comprises the following steps:

[0031] All the training nodes that sent the task claim request are recorded, and the weight integral corresponding to each training node is determined. The calculation formula of the weight integral is:

[0032]

[0033] Wherein, Sw represents the weight integral of the current training node, Kn represents the node type weight integral of the current training node, Sh represents the historical integral of the current training node, which is related to the historical completion of the training task of the current training node, and Qt represents the number of training tasks received by the current training node within the preset task management cycle;

[0034] Sort all the recorded training nodes, and the sorting rule is to arrange them in descending order according to the weight integral corresponding to each training node, select the training node with the highest weight integral as the target node, and send the training task corresponding to the basic information of the training task to the target node.

[0035] By adopting the above technical solution, the weight integral calculation rules for each independent training node are clarified, which provides technical support for prioritizing the allocation of training tasks to training nodes with high weight integrals in subsequent processes. It also avoids the problem of uneven distribution of training tasks from the execution level and realizes the dynamic allocation of training tasks.

[0036] Preferably, the receiving of training results from any of the training nodes, each of the training results corresponding to a training task, generating reward and punishment points for the training task according to the sending time of the training results and the basic information of the training task corresponding to the training task, and using the reward and punishment points to adjust and update the weight points of the current training node, comprises the following steps:

[0037] Receiving a training result from any of the training nodes, each of the training results corresponding to one of the training tasks, and comparing a sending time of the training result with a task deadline corresponding to the training task;

[0038] If the sending time of the training result does not exceed the task deadline, the reward points are calculated and generated, and the historical points of the current training node are updated with the reward points. The calculation formula of the reward points is:

[0039]

[0040] Dt=Dtb*Vt,

[0041] Among them, Sr represents the reward points of the current training node in the training task, Pt represents the task priority of the training task, Dt represents the difficulty value of the training task, Dtb represents the basic difficulty value of the training task, which is related to the task algorithm type of the training task, Vt represents the corresponding score of the task data volume of the training task, and Tt represents the time taken to complete the training task.

[0042] Preferably, the receiving of training results from any of the training nodes, each of the training results corresponding to a training task, generating reward and punishment points for the training task according to the sending time of the training results and the basic information of the training task corresponding to the training task, and using the reward and punishment points to adjust and update the weight points of the current training node, also includes the following steps:

[0043] If the sending time of the training result exceeds the task deadline, a penalty point is calculated and generated, and the penalty point is used to devalue the historical points of the current training node. The calculation formula of the penalty point is:

[0044]

[0045] Wherein, Sr represents the penalty points of the current training node in the current training task.

[0046] By adopting the above technical solution, the specific implementation process of the point reward and punishment mechanism is clarified, which incentivizes the active participation and high-quality contribution of each independent training node in the distributed training network, encourages training nodes to complete the most difficult training tasks as quickly as possible, and fully taps the resource potential of each training node.

[0047] In the second aspect, the present application provides a distributed training task scheduling system based on central node management, which adopts the technical solution described below.

[0048] A distributed training task scheduling system based on central node management is designed for training nodes of different node types, including fixed computing power nodes and idle computing power nodes. The system includes the following modules:

[0049] An identity authentication module is configured to complete the identity authentication process with the corresponding training nodes one by one according to the node type and maintain a communication connection with the training nodes;

[0050] The task publishing module is configured to broadcast a training task execution request to the training node according to the training task basic information, receive a task collection request from the training node within a task collection window period, the task collection request is generated by the training node according to the training task execution request and its own computing power load status, select a training node with the highest weight score from all the training nodes that sent the task collection request as the target node, and send the training task corresponding to the training task basic information to the target node;

[0051] The points management module is configured to receive training results from any of the training nodes, each of which corresponds to a training task, generate reward and punishment points for the training task based on the sending time of the training results and the basic information of the training task corresponding to the training task, and use the reward and punishment points to adjust and update the weight points of the current training node.

[0052] By adopting the above technical solution, the aforementioned distributed training task scheduling method based on central node management is implemented at the hardware level. It can be seen that its core lies in the intelligent management of the central node to achieve dynamic allocation, efficient execution and result verification of training tasks. At the same time, the introduced point reward and punishment mechanism also encourages the active participation and high-quality contribution of training nodes, maximizing the use of idle computing resources.

[0053] In a third aspect, the present application provides a smart terminal, which adopts the following technical solution:

[0054] An intelligent terminal includes a memory and a processor, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the distributed training task scheduling method based on central node management as described above.

[0055] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution:

[0056] A computer-readable storage medium, wherein at least one instruction, at least one program, code set or instruction set is stored in the readable storage medium, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the distributed training task scheduling method based on central node management as described above.

[0057] In summary, the present application includes at least one of the following beneficial technical effects:

[0058] 1. The solution of this application uses both fixed computing nodes and idle computing nodes as training nodes and incorporates them into a distributed training network. During the authentication process of training nodes, different node types use different authentication methods, which maximizes the protection of the rights and interests of all training nodes participating in the training, encourages individual idle computing power to actively participate in distributed training, and greatly improves the overall resource utilization.

[0059] 2. The solution of the present application realizes intelligent task scheduling, ensures the effective participation of each independent training node in the distributed training network, avoids the problem of uneven distribution of training tasks from the execution level, and realizes dynamic allocation and efficient execution of training tasks.

[0060] 3. The solution of this application introduces flexible points reward and punishment rules to motivate the active participation and high-quality contribution of each independent training node in the distributed training network, encourage training nodes to complete the most difficult training tasks as quickly as possible, and fully tap the resource potential of each training node. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, they can also refer to these drawings to obtain other drawings without paying creative work.

[0062] Figure 1 is a schematic diagram of the architecture of a distributed training network according to an embodiment of the present application;

[0063] Figure 2 It is a flow chart of a distributed training task scheduling method based on central node management in an embodiment of the present application;

[0064] Figure 3 It is a schematic diagram of the architecture of a distributed training task scheduling system based on central node management in an embodiment of the present application. DETAILED DESCRIPTION

[0065] The present application provides a distributed training task scheduling method, system, terminal and storage medium based on central node management to achieve dynamic allocation, efficient execution and result verification of training tasks in large-scale distributed training. In order to make the purpose, technical solution and advantages of the present application clearer, the implementation method of the present application will be further described below.

[0066] The following is a distributed training task scheduling method, system, terminal and storage medium based on central node management of the present application in conjunction with the drawings of the specification.

[0067] First, a distributed training task scheduling method based on central node management in an embodiment of the present application is introduced, which is based on the central node and is oriented to training nodes of different node types. The node types include fixed computing power nodes and idle computing power nodes. The fixed computing power nodes can be training nodes in a self-built computing power center, and the idle computing power nodes can be training nodes that are usually idle, such as personal computers and smart phones. The specific architecture of the distributed training network is as follows: Figure 1 As shown, the training task scheduling center plays the role of the central node.

[0068] like Figure 2 As shown, the method comprises the following steps:

[0069] S1. Complete the identity authentication process with the corresponding training nodes one by one according to the node type, and maintain the communication connection with the training nodes. Depending on the node type of the training nodes, this step specifically includes the following process.

[0070] S111. When the training node is the fixed computing power node, in order to support the flexible participation of cross-domain training nodes, in order to ensure the security of transmission when accessed through the public network, the AK / SK scheme is adopted in this embodiment, and a preset encrypted transmission channel, here the https encrypted channel, is used for transmission.

[0071] Receive a first identity authentication request from the training node, where the first identity authentication request is uniformly sent using the POST method and the content is in application / json format, and the first identity authentication request includes a first request header and a first request signature.

[0072] The first request packet header contains information such as an access key (Access Key, AK) and a first signature value for identifying the identity of the training node. The specific structure of the first request packet header is as follows.

[0073] Contents of the header Contents X-Access Key Access Key X-Date Current millisecond timestamp X-Random Key Randomly generated string, 32 bytes long X-Signature1 First signature value

[0074] The first request signature includes a first signature content, which is the original signature content processed by a secret access key (SK) that matches the access key and is used for signing and verifying the signature.

[0075] The specific form of the original signature content is as follows:

[0076] String To Sign=URI+“\n”+MD5(Body)+“\n”+Time stamp+“\n”+Random key;

[0077] Among them, URI represents the unique identifier of the POST request, Body represents the serial number string of the json format content of the request sorted in ascending order in lexicographical order, Time stamp represents the millisecond timestamp of X-Date in the first request header, and Random key represents the random string of X-Random Key in the first request header.

[0078] It should be noted that the access key and the confidential key match each other and are both generated by the central node and stored in the first key library in the central node. At the same time, the access key and the confidential key are also sent to the training node and stored confidentially by the training node. The access key and the confidential key are uniformly managed by the central node, and the management authority includes operations such as creation, retrieval, update and deletion.

[0079] The specific form of the first request signature is as follows:

[0080] Signature1=Base64(HMAC-SHA256(SK,String To Sign)).

[0081] S112. Search for the corresponding confidential key in the first key library based on the access key, calculate the signature value of the first request signature in combination with the search result and the preset first signature verification rule, and compare the calculation result with the first signature value included in the first request packet header for consistency. If the comparison results are consistent, it is determined that the identity authentication of the current training node is successful, and the encrypted transmission channel is used to establish a communication connection with the current training node.

[0082] S121. When the training node is the idle computing node, a public-private key scheme is adopted, which is an authentication method based on asymmetric encryption, mainly used to achieve secure communication and data transmission.

[0083] Receive a second identity authentication request from the training node, the second identity authentication request includes a second request header and a second request signature, the second request header includes information such as a public key (Public Key) used to verify the signature and a second signature value, and the specific structure of the second request header is as follows.

[0084] Contents of the header Contents X-Public Key Public Key X-Date Current millisecond timestamp X-Random Key Randomly generated string, 32 bytes long X-Signature2 Second signature value

[0085] The second request signature includes second signature content, and the second signature content is the original signature content processed by a private key (Private Key) that matches the public key and is used to generate the signature.

[0086] It should be noted that the public key and the private key are matched and correspond to each other in the form of a key pair and both are generated offline by the current training node and saved locally. The public key submitted for registration by the training node is stored in the second key library in the central node.

[0087] The specific form of the second request signature is as follows:

[0088] Signature2=Base64(RSA-sign(Private Key, String To Sign)).

[0089] S122. Confirm whether the public key exists in the second key library. If so, calculate the signature value of the second request signature in combination with the public key and the preset second name verification rule, and compare the calculation result with the second signature value contained in the second request packet header for consistency. If the comparison results are consistent, it is determined that the identity authentication of the current training node is successful, and a communication connection is established with the current training node.

[0090] Since the public-private key scheme uses asymmetric encryption, even if the public key is leaked, it will not affect the security of the private key. And when the training node participates in the training task, it does not need to submit the private key to the central node, which further ensures the confidentiality of the private key.

[0091] S2, broadcast the training task execution request to the training node according to the training task basic information, receive the task collection request from the training node within the task collection window period, the task collection request is generated by the training node according to the training task execution request and its own computing power load status, select the training node with the highest weight score from all the training nodes that sent the task collection request as the target node, and send the training task corresponding to the training task basic information to the target node. This step specifically includes the following process.

[0092] S21. Generate a corresponding training task execution request based on the basic training task information of the current training task, wherein the basic training task information includes the task algorithm type, task data volume, task priority and task deadline, and broadcast the training task execution request to the training node.

[0093] The task algorithm types include a variety of neural network models and deep learning algorithms, including at least basic machine learning algorithms, traditional neural networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep reinforcement learning algorithms, generative adversarial networks (GANs) and large language models (LLMs). Each of the task algorithm types corresponds to a different basic difficulty value. In this embodiment, the corresponding relationship is as follows.

[0094]

[0095]

[0096] The task data volume varies according to different data types, and different task data volumes have different corresponding scores, and the corresponding relationship is as follows.

[0097]

[0098]

[0099] The task priority is associated with the task deadline. The calculation rule of the task priority is to take the time interval between the current moment and the task deadline. The smaller the time interval, the higher the task priority. The value of the task priority is basically related to the number of days in the time interval. The corresponding score value is up to 100, and the corresponding relationship is as follows.

[0100] The time interval between the current time and the task deadline Task Priority Corresponding score 1 day 1 1 1 week 2 7 January 3 30 100 days and above 4 100

[0101] S22. Receive a task collection request from the training node within a task collection window period, wherein the task window period matches the task deadline and can be freely configured based on the task deadline, and the task collection request is generated by the training node based on its own computing power load status and corresponds to the training task execution request.

[0102] Here, the task collection request is generated only when the computing power load status of the training node itself is idle. The specific judgment can be followed. If the training node only includes a CPU, then the CPU load is lower than 80% and it is considered idle. If the training node also includes a GPU, in addition to the aforementioned CPU load being lower than 80%, there must also be an idle channel in the GPU to be considered idle.

[0103] S23, recording all the training nodes that sent the task claiming request, and determining the weight integral corresponding to each training node. The calculation formula of the weight integral is:

[0104]

[0105] Among them, Sw represents the weight integral of the current training node, Kn represents the node type weight integral of the current training node, Sh represents the historical integral of the current training node, which is related to the historical completion status of the training task of the current training node. Each training node has an independent integral account for managing its own historical integral. Qt represents the number of training tasks received by the current training node within the preset task management cycle. The task management cycle in this embodiment is 24 hours.

[0106] The node type weight score of the training node has only two values, corresponding to the fixed computing power node and the idle computing power node, wherein the node type weight score of the fixed computing power node is higher than that of the idle computing power node. In this embodiment, the node type weight score of the fixed computing power node is assigned a value of 5, and the node type weight score of the idle computing power node is assigned a value of 2.

[0107] Here, the number of training tasks received within the task management cycle is introduced as an influencing factor of the weight integral. The more tasks received, the lower the weight integral. This is to prevent uneven task distribution. The training nodes that receive fewer tasks have priority in allocation.

[0108] S24. Sort all the recorded training nodes according to the sorting rule of arranging them in descending order according to the weight score corresponding to each training node, select the training node with the highest weight score as the target node, and send the training task corresponding to the basic information of the training task to the target node.

[0109] S3, receiving training results from any of the training nodes, each of which corresponds to a training task, generating reward and punishment points for the training task according to the sending time of the training results and the basic information of the training task corresponding to the training task, wherein the reward and punishment points include two types: reward points and penalty points, and using the reward and punishment points to adjust and update the weight points of the current training node. This step specifically includes the following process.

[0110] S31, receiving a training result from any of the training nodes, each of the training results corresponding to a training task, comparing the sending time of the training result with the task deadline corresponding to the training task, and determining whether the current training task is completed in timeout. Considering the possible delay in data transmission, the sending time of the training result is used as the comparison basis here.

[0111] S32: If the sending time of the training result does not exceed the task deadline, then the reward points are calculated and generated, and the historical points in the points account of the current training node are updated with the reward points. The calculation formula of the reward points is:

[0112]

[0113] Dt=Dtb*Vt,

[0114] Among them, Sr represents the reward points of the current training node in the training task, Pt represents the corresponding score of the task priority of the training task, Dt represents the difficulty value of the training task, Dtb represents the basic difficulty value of the training task, which is related to the task algorithm type of the training task, Vt represents the corresponding score of the task data volume of the training task, and Tt represents the time taken to complete the training task.

[0115] From the above calculation formula, it can be seen that the rule of the reward points is that the higher the task priority, the higher the points; the greater the training difficulty, the higher the points; the shorter the training time, the higher the points; the higher the task completion rate, the higher the points. Through the above rules, the training nodes can be encouraged to complete difficult training tasks as quickly as possible.

[0116] S33: If the sending time of the training result exceeds the task deadline, penalty points are calculated and generated, and the historical points in the points account of the current training node are devalued and updated using the penalty points. The calculation formula of the penalty points is:

[0117]

[0118] Wherein, Sr represents the penalty points of the current training node in the current training task.

[0119] Based on the same inventive concept as above, the embodiment of the present application also discloses a distributed training task scheduling system based on central node management, which is based on the central node and faces training nodes of different node types, and the node types include fixed computing power nodes and idle computing power nodes, such as Figure 3 As shown, the system includes the following modules:

[0120] An identity authentication module is configured to complete the identity authentication process with the corresponding training nodes one by one according to the node type and maintain a communication connection with the training nodes;

[0121] The task publishing module is configured to broadcast a training task execution request to the training node according to the training task basic information, receive a task collection request from the training node within a task collection window period, the task collection request is generated by the training node according to the training task execution request and its own computing power load status, select a training node with the highest weight score from all the training nodes that sent the task collection request as the target node, and send the training task corresponding to the training task basic information to the target node;

[0122] The points management module is configured to receive training results from any of the training nodes, each of which corresponds to a training task, generate reward and punishment points for the training task based on the sending time of the training results and the basic information of the training task corresponding to the training task, and use the reward and punishment points to adjust and update the weight points of the current training node.

[0123] In a possible implementation, the identity authentication module includes the following units:

[0124] The fixed computing power first authentication unit is configured to receive a first identity authentication request from the training node when the training node is the fixed computing power node, the first identity authentication request comprising a first request header and a first request signature, the first request header comprising an access key and a first signature value, the first request signature comprising a first signature content obtained after processing with a confidential key matching the access key, the access key and the confidential key matching and corresponding to each other and both being stored in a first key library;

[0125] The second authentication unit with fixed computing power is configured to search for the corresponding confidential key in the first key library based on the access key, calculate the signature value of the first request signature based on the search result and the preset first signature verification rule, and compare the calculation result with the first signature value contained in the first request packet header for consistency. If the comparison result is consistent, it is determined that the identity authentication of the current training node is successful, and a communication connection is established with the current training node using a preset encrypted transmission channel.

[0126] In a possible implementation, the identity authentication module includes the following units:

[0127] The idle computing power first authentication unit is configured to receive a second identity authentication request from the training node when the training node is the idle computing power node, the second identity authentication request includes a second request header and a second request signature, the second request header includes a public key and a second signature value, the second request signature includes a second signature content obtained after processing with a private key matching the public key, the public key and the private key match and correspond to each other and both are generated offline by the current training node and stored locally, and the public key submitted for registration by the training node is stored in a second key library;

[0128] The second authentication unit of idle computing power is configured to confirm whether the public key exists in the second key library. If so, the signature value of the second request signature is calculated in combination with the public key and the preset second name verification rule, and the calculation result is compared with the second signature value contained in the second request packet header for consistency. If the comparison results are consistent, it is determined that the identity authentication of the current training node is successful, and a communication connection is established with the current training node.

[0129] In a possible implementation, the task issuing module includes the following units:

[0130] an execution request broadcasting unit, configured to generate a corresponding training task execution request according to basic training task information of a current training task, wherein the basic training task information includes a task algorithm type, a task data volume, a task priority, and a task deadline, and broadcast the training task execution request to the training node;

[0131] The task collection request receiving unit is configured to receive a task collection request from the training node within a task collection window period, wherein the task window period matches the task deadline, and the task collection request is generated by the training node based on its own computing power load status and corresponds to the training task execution request.

[0132] In a possible implementation, the task issuing module further includes the following units:

[0133] A weight integral determination unit is configured to record all the training nodes that send the task claim request and determine the weight integral corresponding to each training node;

[0134] The training node selection unit is configured to sort all the recorded training nodes, and the sorting rule is to arrange them in descending order according to the weight integral corresponding to each training node, select the training node with the highest weight integral as the target node, and send the training task corresponding to the basic information of the training task to the target node.

[0135] In a possible implementation, the points management module includes the following units:

[0136] A task timeout confirmation unit is configured to receive a training result from any of the training nodes, each of the training results corresponding to a training task, and compare a sending time of the training result with a task deadline corresponding to the training task;

[0137] The reward point updating unit is configured to calculate and generate reward points if the sending time of the training result does not exceed the task deadline, and use the reward points to increase the historical points of the current training node.

[0138] In a possible implementation, the points management module further includes the following units:

[0139] The penalty point updating unit is configured to calculate and generate penalty points if the sending time of the training result exceeds the task deadline, and use the penalty points to devalue and update the historical points of the current training node.

[0140] Based on the same inventive concept mentioned above, an embodiment of the present application also discloses an intelligent terminal, including a memory and a processor, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement any one of the distributed training task scheduling methods based on central node management as described above.

[0141] It should be understood that the "plurality" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0142] Also based on the same inventive concept above, an embodiment of the present application also discloses a computer-readable storage medium, in which at least one instruction, at least one program, code set or instruction set is stored, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement any one of the distributed training task scheduling methods based on central node management as described above.

[0143] Those skilled in the art should understand that all or part of the steps to implement the above embodiments may be completed by hardware, or may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium. The above-mentioned storage medium may include, for example: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0144] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A distributed training task scheduling method based on central node management, characterized in that: For training nodes of different node types, the node types include fixed computing power nodes and idle computing power nodes, and the method includes the following steps: Complete the identity authentication process with the corresponding training nodes one by one according to the node type, and maintain the communication connection with the training nodes; Broadcasting a training task execution request to the training node according to the training task basic information, receiving a task collection request from the training node within a task collection window period, wherein the task collection request is generated by the training node according to the training task execution request and its own computing power load status, selecting a training node with the highest weight score from all the training nodes that sent the task collection request as the target node, and sending the training task corresponding to the training task basic information to the target node; Receive training results from any of the training nodes, each of the training results corresponds to a training task, generate reward and punishment points for the training task based on the sending time of the training results and the basic information of the training task corresponding to the training task, and use the reward and punishment points to adjust and update the weight points of the current training node.

2. The distributed training task scheduling method based on central node management according to claim 1 is characterized in that: The identity authentication process is completed with the training nodes one by one according to the type of the training nodes, including the following steps: When the training node is the fixed computing power node, Receive a first identity authentication request from the training node, the first identity authentication request comprising a first request header and a first request signature, the first request header comprising an access key and a first signature value, the first request signature comprising a first signature content obtained by processing with a confidentiality key matching the access key, the access key and the confidentiality key matching and corresponding to each other and both being stored in a first key library; The corresponding confidential key is searched in the first key library according to the access key, and the signature value of the first request signature is calculated in combination with the search result and the preset first signature verification rule. The calculation result is compared with the first signature value contained in the first request packet header for consistency. If the comparison result is consistent, it is determined that the identity authentication of the current training node is successful, and a communication connection is established with the current training node using the preset encrypted transmission channel.

3. The distributed training task scheduling method based on central node management according to claim 1 is characterized in that: The identity authentication process is completed with the training nodes one by one according to the type of the training nodes, including the following steps: When the training node is the idle computing power node, Receive a second identity authentication request from the training node, the second identity authentication request includes a second request header and a second request signature, the second request header includes a public key and a second signature value, the second request signature includes a second signature content obtained after processing with a private key matching the public key, the public key and the private key match and correspond to each other and both are generated offline by the current training node and stored locally, and the public key submitted for registration by the training node is stored in a second key library; Confirm whether the public key exists in the second key library. If so, calculate the signature value of the second request signature in combination with the public key and the preset second name verification rule, and compare the calculation result with the second signature value contained in the second request packet header for consistency. If the comparison results are consistent, it is determined that the identity authentication of the current training node is successful, and a communication connection is established with the current training node.

4. The distributed training task scheduling method based on central node management according to claim 1 is characterized in that: The method of broadcasting a training task execution request to the training node according to the training task basic information and receiving a task receiving request from the training node within a task receiving window period comprises the following steps: Generate a corresponding training task execution request according to the training task basic information of the current training task, wherein the training task basic information includes the task algorithm type, the task data volume, the task priority and the task deadline, and broadcast the training task execution request to the training node; A task collection request is received from the training node within a task collection window period, the task window period matches the task deadline, and the task collection request is generated by the training node based on its own computing power load status and corresponds to the training task execution request.

5. The distributed training task scheduling method based on central node management according to claim 4 is characterized in that: The step of selecting a training node with the highest weight score from all the training nodes that sent the task claim request as the target node, and sending the training task corresponding to the training task basic information to the target node comprises the following steps: All the training nodes that sent the task claim request are recorded, and the weight integral corresponding to each training node is determined. The calculation formula of the weight integral is: Wherein, Sw represents the weight integral of the current training node, Kn represents the node type weight integral of the current training node, Sh represents the historical integral of the current training node, which is related to the historical completion of the training task of the current training node, and Qt represents the number of training tasks received by the current training node within the preset task management cycle; Sort all the recorded training nodes, and the sorting rule is to arrange them in descending order according to the weight integral corresponding to each training node, select the training node with the highest weight integral as the target node, and send the training task corresponding to the basic information of the training task to the target node.

6. The distributed training task scheduling method based on central node management according to claim 5 is characterized in that: The receiving of training results from any of the training nodes, each of the training results corresponding to a training task, generating reward and punishment points for the training task according to the sending time of the training results and the basic information of the training task corresponding to the training task, and using the reward and punishment points to adjust and update the weight points of the current training node, includes the following steps: Receiving a training result from any of the training nodes, each of the training results corresponding to one of the training tasks, and comparing a sending time of the training result with a task deadline corresponding to the training task; If the sending time of the training result does not exceed the task deadline, the reward points are calculated and generated, and the historical points of the current training node are updated with the reward points. The calculation formula of the reward points is: Dt=Dtb*Vt, Among them, Sr represents the reward points of the current training node in the training task, Pt represents the task priority of the training task, Dt represents the difficulty value of the training task, Dtb represents the basic difficulty value of the training task, which is related to the task algorithm type of the training task, Vt represents the corresponding score of the task data volume of the training task, and Tt represents the time taken to complete the training task.

7. The distributed training task scheduling method based on central node management according to claim 6 is characterized in that: The receiving of training results from any of the training nodes, each of the training results corresponding to a training task, generating reward and punishment points for the training task according to the sending time of the training results and the basic information of the training task corresponding to the training task, and using the reward and punishment points to adjust and update the weight points of the current training node, also includes the following steps: If the sending time of the training result exceeds the task deadline, a penalty point is calculated and generated, and the penalty point is used to devalue the historical points of the current training node. The calculation formula of the penalty point is: Wherein, Sr represents the penalty points of the current training node in the current training task.

8. A distributed training task scheduling system based on central node management, characterized in that: Training nodes for different node types, including fixed computing power nodes and idle computing power nodes, the system includes the following modules: An identity authentication module is configured to complete the identity authentication process with the corresponding training nodes one by one according to the node type and maintain a communication connection with the training nodes; The task publishing module is configured to broadcast a training task execution request to the training node according to the training task basic information, receive a task collection request from the training node within a task collection window period, the task collection request is generated by the training node according to the training task execution request and its own computing power load status, select a training node with the highest weight score from all the training nodes that sent the task collection request as the target node, and send the training task corresponding to the training task basic information to the target node; The points management module is configured to receive training results from any of the training nodes, each of which corresponds to a training task, generate reward and punishment points for the training task based on the sending time of the training results and the basic information of the training task corresponding to the training task, and use the reward and punishment points to adjust and update the weight points of the current training node.

9. An intelligent terminal, characterized in that: It includes a memory and a processor, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the distributed training task scheduling method based on central node management as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The readable storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the distributed training task scheduling method based on central node management as described in any one of claims 1 to 7.