Cross-domain task scheduling strategy determination method and device, equipment and storage medium

By adopting the cross-domain task scheduling strategy determination method optimized by genetic algorithm in cross-geographic computing, the problem of failure to fully consider the finiteness and differences of computing resources in the prior art is solved, and efficient task allocation and execution in complex scenarios is achieved.

CN120179358APending Publication Date: 2025-06-20GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510246864.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing scheduling methods for cross-geographic domain computing, such as the scheduling algorithm of the Irrdim framework, fail to fully consider the finiteness and differences in the computing resources of the cluster, resulting in the inability to find the optimal task allocation plan under complex and diverse cluster conditions, affecting the task execution efficiency.

Method used

A cross-domain task scheduling strategy determination method is adopted. By obtaining the parameters of each cluster, simulating task allocation processing, and optimizing the scheduling strategy using genetic algorithms to ensure that the differences in computing resources, network bandwidth and data volume are considered at each task stage, and then a global optimal scheduling plan is made in complex scenarios.

Benefits of technology

Effectively reduce task execution time and improve task execution efficiency. It can propose the optimal task allocation plan in different scenarios to balance task execution time and transmission time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179358A_ABST
    Figure CN120179358A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain task scheduling strategy determination method and device, equipment and a storage medium. The method comprises the steps of obtaining a first calculation task and a first cluster parameter; and based on the initial scheduling strategy, simulating task allocation processing between clusters according to the first cluster parameter and the first calculation task, and determining a fitness function required by adopting the initial scheduling strategy. And for each task stage, evolutionary operation is performed on the task scheduling strategy corresponding to the minimum fitness function based on a genetic algorithm, the evolutionary task scheduling strategy of each task stage is obtained, the fitness function is determined, and the evolutionary operation of the task scheduling strategy is circularly executed. And when the condition is met, determining that the circulation is ended, and determining a target scheduling strategy according to the task scheduling strategy of the minimum fitness function of each task stage. According to the technical scheme, the optimal task load can be accurately allocated to each cluster, the task execution time can be effectively shortened, and the task execution efficiency is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed computing, and particularly to a method, apparatus, device and storage medium for determining a cross-domain task scheduling strategy. Background Art

[0002] Many companies and organizations build data center clusters in corresponding regions or use cloud service platforms to construct big data clusters to process business data according to the geographical attributes of their businesses. This deployment method results in data usually being distributed in different geographical regions, and the data required for a business may be scattered in multiple clusters. To avoid the high time and money costs brought about by concentrating scattered data in a single cluster, cross-geographical domain computing has gradually become an efficient solution. Cross-domain computing achieves the overall computing goal by allocating computing tasks to multiple distributed clusters and summarizing the computing results in a single cluster after the tasks are completed.

[0003] Currently, the commonly used scheduling method for cross-geographical domain computing is the scheduling algorithm of the Irrdim framework. Irrdim focuses on improving data transmission efficiency and reducing data transmission time to achieve the goal of improving task execution efficiency. It formulates the task scheduling problem as a linear programming problem, takes into account the bandwidth differences of the uplink and downlink of different clusters, and uses the Gurobi linear solver to solve it to find the optimal task scheduling scheme.

[0004] Irrdim improves task execution efficiency by reducing data transmission time, but it only focuses on a single resource, i.e., network bandwidth, defaults that the computing resources of the cluster are infinite, and ignores the finiteness and differences of computing resources. In reality, the computing resources of the cluster are not infinite, but vary significantly due to different conditions. Moreover, cross-domain tasks cannot be infinitely decomposed, and the linear solver solves continuous problems, so this does not meet the actual needs and will cause certain errors. In addition, the factors affecting the execution efficiency of cross-domain computing tasks are not limited to data transmission time. When dealing with complex computing tasks, computing time may become the key factor restricting the overall efficiency. For this reason, Irrdim cannot find the optimal task allocation scheme under complex and diverse cluster conditions. Due to the insufficient consideration of the impact of computing time on task execution efficiency, its applicable scenarios are relatively single, and it is difficult to meet the task scheduling requirements under multiple resource conditions, thus restricting the breadth and effect of practical applications. Summary of the Invention

[0005] The present invention provides a method, apparatus, device and storage medium for determining a cross-domain task scheduling strategy, which can accurately allocate the optimal task amount to each cluster, effectively reduce the task execution time, and further improve the task execution efficiency.

[0006] According to one aspect of the present invention, there is provided a method for determining a cross-domain task scheduling strategy, including:

[0007] Obtain a first computing task that needs to perform cross-domain computing, and first cluster parameters respectively corresponding to each cluster in the cross-domain cluster system for processing the first computing task, where the cross-domain cluster system performs computing tasks based on the Map-Reduce task model, the Map-Reduce task model includes a Map task stage and a Reduce task stage, and the first cluster parameters at least include a transmission bandwidth parameter and a node computing power parameter;

[0008] For each task stage, based on a preset initial scheduling strategy, according to the first cluster parameters and the first computing task, simulate the task allocation process between each cluster, and determine the fitness function corresponding to each task stage when adopting the initial scheduling strategy, where the initial scheduling strategy includes a Map task scheduling strategy and a Reduce task scheduling strategy, and the fitness function is determined according to the task processing time of each task stage;

[0009] For each task stage, based on the genetic algorithm, perform an evolutionary operation on the task scheduling strategy corresponding to the minimum fitness function to obtain the evolved task scheduling strategy for each task stage, and determine the fitness function corresponding to adopting the evolved task scheduling strategy, where in the initial loop state of each task stage, the task scheduling strategy corresponding to the minimum fitness function is the initial scheduling strategy of each task stage;

[0010] Loop to execute the task scheduling strategy evolutionary operation. When a preset convergence condition is met, determine that the loop execution operation ends. Wherein, when the loop execution of the Map task scheduling strategy ends, execute the Reduce task scheduling strategy evolutionary operation;

[0011] Determine the target scheduling strategy according to the Map task scheduling strategy corresponding to the minimum fitness function and the Reduce task scheduling strategy corresponding to the minimum fitness function.

[0012] According to another aspect of the present invention, there is provided a device for determining a cross-domain task scheduling strategy, including:

[0013] A cross-domain task acquisition module, configured to obtain a first computing task that needs to perform cross-domain computing, and first cluster parameters respectively corresponding to each cluster in the cross-domain cluster system for processing the first computing task, where the cross-domain cluster system performs computing tasks based on the Map-Reduce task model, the Map-Reduce task model includes a Map task stage and a Reduce task stage, and the first cluster parameters at least include a transmission bandwidth parameter and a node computing power parameter;

[0014] A cross - domain task allocation module, which is used for each task phase, based on a preset initial scheduling policy, according to the first cluster parameter and the first computing task, to simulate and perform task allocation processing among each cluster, and determine the fitness function corresponding to each task phase when adopting the initial scheduling policy, wherein the initial scheduling policy includes a Map task scheduling policy and a Reduce task scheduling policy, and the fitness function is determined according to the task processing time of each task phase;

[0015] A scheduling policy optimization module, which is used for each task phase, based on a genetic algorithm, to perform an evolutionary operation on the task scheduling policy corresponding to the minimum fitness function, obtain the evolved task scheduling policy for each task phase, and determine the fitness function corresponding to adopting the evolved task scheduling policy, wherein in the initial loop state of each task phase, the task scheduling policy corresponding to the minimum fitness function is the initial scheduling policy of each task phase;

[0016] A policy loop optimization module, which is used to loop and execute the task scheduling policy evolutionary operation. When a preset convergence condition is met, it is determined that the loop execution operation ends. Among them, when the loop execution of the Map task scheduling policy ends, the Reduce task scheduling policy evolutionary operation is executed;

[0017] A scheduling policy determination module, which is used to determine the target scheduling policy according to the Map task scheduling policy corresponding to the minimum fitness function and the Reduce task scheduling policy corresponding to the minimum fitness function.

[0018] According to another aspect of the present invention, an electronic device is provided. The electronic device includes:

[0019] At least one processor; and

[0020] A memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the cross - domain task scheduling policy determination method according to any embodiment of the present invention.

[0022] According to another aspect of the present invention, a computer - readable storage medium is provided. The computer - readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the cross - domain task scheduling policy determination method according to any embodiment of the present invention when executed.

[0023] In the technical solution of the embodiment of the present invention, by obtaining a first computing task that needs to perform cross-domain computing, and first cluster parameters corresponding to each cluster in the cross-domain cluster system that processes the first computing task, it is possible to know the upload data and download data times of each cluster and the time for processing data. For each task stage, based on a preset initial scheduling policy, according to the first cluster parameters and the first computing task, simulate the task allocation process among each cluster, and determine the fitness function corresponding to each task stage in the case of adopting the initial scheduling policy. For each task stage, based on the genetic algorithm, perform an evolutionary operation on the task scheduling policy corresponding to the minimum fitness function to obtain an evolved task scheduling policy, and determine the fitness function required for adopting the evolved task scheduling policy, and perform a loop to execute the task scheduling policy evolutionary operation, fully considering the differences in computing resources, network bandwidth, and initial data volume of each cluster, and being able to make a globally optimal scheduling plan in complex situations. When the preset convergence condition is met, determine that the loop execution operation ends, and determine the Map task scheduling policy corresponding to the minimum fitness function and the Reduce task scheduling policy corresponding to the minimum fitness function as the target scheduling policy, and determine it as the final target scheduling policy. The cross-domain task scheduling policy determination method of the present invention can propose an optimal task allocation plan in different scenarios, can effectively balance the task execution time and transmission time, and thus can improve the cross-domain task execution efficiency.

[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0026] Figure 1 is a flowchart of a method for determining a cross-domain task scheduling policy according to Embodiment 1 of the present invention;

[0027] Figure 2 is a flowchart of a method for determining a cross-domain task scheduling policy according to Embodiment 2 of the present invention;

[0028] Figure 3 is a structural diagram of a device for determining a cross-domain task scheduling policy according to Embodiment 3 of the present invention;

[0029] Figure 4 It is a schematic structural diagram of an electronic device for implementing the cross-domain task scheduling strategy determination method of the embodiments of the present invention. Detailed implementation manners

[0030] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment 1

[0033] Figure 1 It is a flowchart of a cross-domain task scheduling strategy determination method provided in Embodiment 1 of the present invention. This embodiment is applicable to the scenario of cross-geographical domain task scheduling, especially in a complex multi-cluster environment where the task execution efficiency is affected by multiple factors such as computing power resources, network bandwidth, and the difference in the distribution of original data. This method can be executed by a cross-domain task scheduling strategy determination device, which can be implemented in the form of hardware and / or software, and the cross-domain task scheduling strategy determination device can be configured in an electronic device. As Figure 1 shown, the method includes:

[0034] S101. Obtain a first computing task that needs to perform cross-domain computing, and first cluster parameters respectively corresponding to each cluster in the cross-domain cluster system for processing the first computing task.

[0035] Among them, the first computing task may be a task that requires cross-domain computing. It should be noted that a cross-domain cluster system refers to a distributed system composed of computing clusters distributed in multiple different domains. These domains can be different geographical regions, different network domains, different management domains, or even different organizations or data centers. The cross-domain cluster system collaborates to jointly complete large-scale computing tasks or data processing tasks. In a cross-domain cluster system, there are at least two individual clusters in different domains.

[0036] The task scheduling of the present invention is based on the amount of data processed by each cluster in the cross-domain cluster system. The result of the scheduling is the amount of data processed by each cluster, that is, the amount of tasks that each cluster should undertake. Cross-domain computing tasks perform data processing through distribution. Such tasks are based on the Map-Reduce task model. The Map-Reduce task model is mainly divided into the Map task stage and the Reduce task stage. The overall task can be specifically divided into the Map task, the Shuffle stage, and the Reduce task. It can also be said that the Reduce task stage includes the Shuffle stage and the Reduce task. In the Map stage, the Map function converts the input data into key-value pairs; in the Shuffle stage, the intermediate key-value pairs output by the Map stage are distributed to the Reduce tasks according to the key value. In the Reduce stage, the Reduce function performs aggregation operations (such as summation, counting, etc.) on the values with the same key to generate the final result. The Map-Reduce task model simplifies the complexity of large-scale data processing through the Map and Reduce stages, has good scalability and fault tolerance, and is suitable for various big data scenarios.

[0037] In the present invention, the first cluster parameter may refer to the actual cluster information of each cluster in the cross-domain cluster system. The first cluster parameter at least includes the transmission bandwidth parameter and the node computing power parameter. Exemplarily, the transmission bandwidth parameter includes the uplink bandwidth and the downlink bandwidth, and the node computing power parameter at least includes the number of cluster nodes and the computing power of a single node.

[0038] Specifically, obtain the to-be-executed cross-domain computing task from the task queue, obtain the network bandwidth information between clusters through the cluster monitoring system, and obtain the number of nodes in the cluster and the computing power of each node through the cluster configuration management tool, so as to obtain the first computing task and the first cluster parameter.

[0039] S102. For each task stage, based on a preset initial scheduling policy, according to the first cluster parameter and the first computing task, simulate the task allocation process between each cluster, and determine the fitness function corresponding to each task stage when adopting the initial scheduling policy.

[0040] Among them, the initial scheduling policy may refer to a preset default scheduling policy. In fact, since the Map-Reduce task model is mainly divided into the Map task stage and the Reduce task stage, correspondingly, the initial scheduling policy includes the Map task scheduling policy and the Reduce task scheduling policy. Each task stage of the present invention is optimized based on its respective initial scheduling policy to obtain the optimal scheduling policy. In the technical solution of the embodiment of the present invention, the task execution time of each task stage is used as its respective fitness function.

[0041] It should be noted that the essence of the present invention is to determine the optimal scheduling policy, that is, the task execution time of the first computing task is minimized, so that the cross-domain cluster system can complete the execution of the computing task fastest according to the optimal scheduling policy. Therefore, before determining the optimal scheduling policy, all scheduling work is performed for virtual scheduling on the scheduling server. Exemplarily, a virtual cross-domain cluster system can be established on the scheduling server according to the real-time parameters of the actual cross-domain cluster system, and virtual scheduling is performed based on the virtual cross-domain cluster system. It is also possible to call a virtual execution instruction on the basis of not establishing a virtual cross-domain cluster system, so that the scheduling server determines that virtual scheduling has been completed.

[0042] Specifically, in the Map task stage, according to the Map task scheduling policy, the first computing task is subjected to simulated partitioning processing to obtain each sub-task after partitioning. In the Reduce task stage, each sub-task after partitioning is assigned to the corresponding clusters, and task calculation processing is performed. At the same time, according to the transmission bandwidth parameter and the node computing power parameter in the first cluster parameter corresponding to each cluster, the task transmission time and the task calculation time in each task stage are calculated, and then the sum of the task transmission time and the task calculation time is determined as the fitness function of the corresponding task stage.

[0043] S103. For each task stage, based on the genetic algorithm, an evolutionary operation is performed on the task scheduling policy corresponding to the minimum fitness function to obtain the evolved task scheduling policy for each task stage, and the fitness function corresponding to the evolved task scheduling policy is determined.

[0044] Among them, the genetic algorithm is a heuristic search algorithm based on the principles of natural selection and genetic mechanisms, and belongs to a kind of evolutionary algorithm. It simulates operations such as selection, crossover (recombination), and mutation in the biological evolution process and is used to solve optimization and search problems.

[0045] It should be noted that during the evolution of the scheduling policy, it is preferably to further evolve the relatively reasonable scheduling policy in the current scheduling scheme to obtain a more reasonable scheduling policy. In the present invention, the fitness function is used as an evaluation index for the genetic algorithm to evaluate the pros and cons of the task scheduling policy.

[0046] Specifically, for the Map task stage or the Reduce task stage, the task scheduling policy corresponding to the minimum fitness function among the current task scheduling policies in this task stage is evolved according to the genetic algorithm, so as to obtain the evolved task scheduling policy for this task stage. After obtaining the evolved task scheduling policy for this task stage, when it is determined to adopt the evolved task scheduling policy, calculate the task processing time required for the corresponding task as the fitness function.

[0047] It should be noted that in the initial loop state, there are no other scheduling policies that can be compared. Therefore, in the initial loop state, the initial scheduling policy for each task stage is determined as the task scheduling policy corresponding to the minimum fitness function for the corresponding task stage, that is, in the initial loop state, the initial scheduling policy is evolved according to the genetic algorithm to obtain the evolved task scheduling policy.

[0048] Exemplarily, based on the genetic algorithm, evolving the task scheduling policy corresponding to the minimum fitness function to obtain the evolved task scheduling policy includes: for each task stage, randomly exchange the assigned task amounts of two of the clusters in the task scheduling policy corresponding to the minimum fitness function to obtain the initial evolved policy; randomly change the assigned task amount of one of the clusters in the initial evolved policy by a random quantity, and make corresponding changes to the assigned task amounts of other clusters so that the total assigned task quantity remains unchanged, to obtain the evolved task scheduling policy for this task stage.

[0049] That is to say, whether it is the Map task stage or the Reduce task stage, in the present invention, the task scheduling policy corresponding to the minimum fitness function is used as the best allocation plan and the parent plan for subsequent operations, and evolution operations are performed on the basis of the parent plan. First, a crossover operation is performed on the basis of the parent plan. The crossover operation is to explore more solutions. The specific operation is to perform position crossover in the parent plan, that is, exchange the task amounts of two randomly selected clusters in the plan to obtain multiple new initial evolved policies. Then, a mutation operation is performed on the multiple new initial evolved policies. Each cluster in the initial evolved policy has a certain probability of mutating. Mutation means randomly increasing or decreasing the task amount of one cluster, and at the same time decreasing or increasing the corresponding task amount of another cluster. For example, if the task amount corresponding to Cluster 1 increases by 3 data blocks, then the task amount corresponding to Cluster 3 decreases by 3 data blocks. This can ensure that the total task amount of the plan remains unchanged. In this way, multiple evolved task scheduling policies can be obtained. It should be noted that the number of evolved task scheduling policies can be set according to the actual situation, and the present invention does not make specific limitations on this.

[0050] S104. Perform the evolutionary operation of the task scheduling strategy in a loop. When the preset convergence condition is met, determine the end of the loop execution operation.

[0051] Specifically, for the evolutionary operation of each task stage, it is necessary to repeatedly execute the process of evolution - allocation - determination of the fitness function in a loop until the preset convergence condition is met, and then end the loop execution operation.

[0052] Among them, the preset convergence condition can be that the number of loop iterations reaches a preset number, or the task execution time is less than a preset threshold, or other limiting conditions, etc.

[0053] It should be noted that in the Map - Reduce task model, the Map task stage is executed first, and the output result of the Map task stage is used as the output of the Reduce task stage. That is to say, the task processing of the Reduce task stage is affected by the Map task stage. Therefore, in the technical solution of the embodiment of the present invention, when the loop execution of the Map task scheduling strategy ends, the evolutionary operation of the Reduce task scheduling strategy is executed. This is because after the loop execution of the Map task scheduling strategy ends, the optimal Map task scheduling strategy can be obtained. After executing the optimal Map task scheduling strategy, the Map task output result is unique, that is, the input data of the Reduce task stage is unique. In the case of unique input data, the efficiency of the evolutionary operation of the Reduce task scheduling strategy can be improved.

[0054] S105. According to the Map task scheduling strategy corresponding to the minimum fitness function, and the Reduce task scheduling strategy corresponding to the minimum fitness function.

[0055] Specifically, the combined scheduling strategy of the Map task scheduling strategy corresponding to the minimum fitness function and the Reduce task scheduling strategy corresponding to the minimum fitness function is determined as the target scheduling strategy.

[0056] In the technical solution of the embodiment of the present invention, by obtaining a first computing task that needs to perform cross-domain computing, and first cluster parameters corresponding to each cluster in the cross-domain cluster system for processing the first computing task, it is possible to know the upload data and download data times of each cluster and the time for processing data. For each task stage, based on a preset initial scheduling policy, according to the first cluster parameters and the first computing task, simulate the task allocation process among the clusters, and determine the fitness function corresponding to each task stage when adopting the initial scheduling policy. For each task stage, based on the genetic algorithm, perform an evolutionary operation on the task scheduling policy corresponding to the minimum fitness function to obtain an evolved task scheduling policy, and determine the fitness function required for adopting the evolved task scheduling policy, and perform a loop to execute the task scheduling policy evolutionary operation, fully considering the differences in computing resources, network bandwidth, and initial data volume of each cluster, and being able to make a globally optimal scheduling plan in complex situations. When the preset convergence condition is met, determine that the loop execution operation ends, and determine the Map task scheduling policy corresponding to the minimum fitness function and the Reduce task scheduling policy corresponding to the minimum fitness function as the target scheduling policy, and determine it as the final target scheduling policy. The cross-domain task scheduling policy determination method of the present invention can propose an optimal task allocation plan in different scenarios, effectively balance the task execution time and transmission time, and thus can improve the cross-domain task execution efficiency.

[0057] Embodiment 2

[0058] Figure 2 It is a flowchart of a method for determining a cross-domain task scheduling policy provided by Embodiment 2 of the present invention. Based on the above embodiments, the task allocation process simulation among the clusters is further refined, and the task processing of the Map-Reduce task model including the Map task stage, the Shuffle stage, and the Reduce task stage is introduced in detail. As Figure 2 shown, the method includes:

[0059] S201. Obtain a first computing task that needs to perform cross-domain computing, and first cluster parameters corresponding to each cluster in the cross-domain cluster system for processing the first computing task.

[0060] In big data processing, tasks are usually based on the Map-Reduce model, involving the distributed processing of a large amount of data. The data may be stored in different clusters, and the task execution efficiency is affected by the data distribution and network bandwidth. By optimizing the allocation of Map tasks and Reduce tasks, reducing the data transmission time in the Shuffle phase, the execution efficiency of big data processing tasks can be improved, and the overall processing time can be reduced. Here, by way of example, the calculation tasks based on the Map-Reduce model in a cross-domain cluster system in an actual scenario are introduced to understand how to simulate the task allocation and processing between clusters in the present invention:

[0061] Wordcount calculation task: Count the frequency of words appearing in the dataset.

[0062] 1. Input dataset:

[0063] The input of the Wordcount program is a set of files or a single file. Each file contains text content. For example, assume that the text file input.txt contains the following content: HelloHadoop; HelloMapReduce.

[0064] 2. The Wordcount task is divided into a map phase and a reduce phase.

[0065] In the Map phase, the cluster in the cross-domain cluster system that executes the map phase (a certain cluster can be specified according to the situation as the cluster that executes the map task) reads the input text data, decomposes them into words, and then generates a key-value pair (word, 1) for each word. Here, the key is the word itself, and the value is the number 1, representing that the word appears once. For example: (Hello, 1); (Hadoop, 1); (Hello, 1); (MapReduce, 1).

[0066] 3. At the end of the Map phase, a Shuffle operation (which can be understood as a distribution and transmission operation) needs to be performed on the key-value pairs output in the Map phase.

[0067] Shuffle operation: The first step is to divide the key-value pairs output by the Map operation into multiple partitions according to the key value. The task allocation strategy will specify which clusters execute the tasks of which partitions.

[0068] For example: In Cluster 1 during the Map task stage, the output data of its Map tasks is: (Hello,1); (Hadoop,1); (Hello,1); (MapReduce,1); The keys are divided into three partitions according to the Hash value, which are: Partition 1: (Hello,1),(Hello,1) Partition 2: (Hadoop,1) Partition 3: (MapReduce,1).

[0069] In Cluster 2 during the Map task stage, the output data of its Map tasks is: (Hello,1); (Hadoop,1); (Hello,1); (MapReduce,1); The keys are divided into three partitions according to the Hash value, which are: Partition 1: (Hello,1),(Hello,1) Partition 2: (Hadoop,1) Partition 3: (MapReduce,1).

[0070] In Cluster 3 during the Map task stage, the output data of its Map tasks is: (Hello,1); (Hadoop,1); (Hello,1); (MapReduce,1); The keys are divided into three partitions according to the Hash value, which are: Partition 1: (Hello,1),(Hello,1) Partition 2: (Hadoop,1) Partition 3: (MapReduce,1).

[0071] According to the task allocation strategy, Cluster 1 is responsible for Partition 1, Cluster 2 is responsible for Partition 2, and Cluster 3 is responsible for Partition 3. Then, during the shuffle process, it is necessary to transfer the data of Partition 2 from Cluster 1 to Cluster 2, and transfer the data of Partition 3 from Cluster 1 to Cluster 3. The data transfer of other clusters also follows this task allocation strategy.

[0072] 4. Reduce task stage.

[0073] The Reduce function takes the Shuffle output as input. In this example, the Reduce function aggregates the counts and outputs the total number of each word. For example:

[0074] Cluster 1 calculates the data of Partition 1 from all clusters: (Hello,1),(Hello,1)(Hello,1),(Hello,1)(Hello,1),(Hello,1), and finally gets the result: (Hello, 6);

[0075] Cluster 2 calculates the data of Partition 2 from all clusters, and finally gets the result: (Hadoop,3);

[0076] Cluster 3 processes the data of Partition 3 from all clusters, and finally gets the result: (MapReduce,3).

[0077] Finally, the results of Reduce for each cluster are aggregated to the main cluster, and the Wordcount computing task is completed.

[0078] It should be noted that a data center (cluster) with the largest data ratio required for the task is randomly specified or selected as the main cluster for submitting the task, and finally the results output by all clusters are also aggregated to this main cluster. The data required for the cross-geographical domain (cluster) big data computing task is stored in multiple data centers (clusters). The main cluster is selected to submit the task and aggregate the results to the cross-domain cluster system.

[0079] It should be noted that between each cluster in the cross-domain cluster system of the present invention, a point-to-point network model is adopted, that is, there is a dedicated data transmission route between each cluster and other clusters for data transmission. Its cluster scalability is strong. Compared with the core network, expanding the number of clusters will not affect the network bandwidth and will not reduce the task execution efficiency.

[0080] S202. Based on the Map-Reduce task model, in the Map task stage, according to a preset mechanism, determine whether the first computing task performs Map task scheduling, and simulate splitting the first computing task into task key-value pairs.

[0081] It should be noted that for the Map task, task scheduling is not performed by default. The computing tasks are allocated proportionally according to the original data volume of the cluster. The Map task data is processed in the cluster where it is located, and no data transmission is required. This follows the data locality principle, and the movement of the original data may involve privacy issues. However, the present invention provides an optional parameter that allows users to independently select whether to perform task scheduling for the Map task.

[0082] Among them, the preset mechanism for determining whether to perform Map task scheduling at least includes the task confidentiality type, and may also include whether the first computing task needs to use global data.

[0083] Specifically, in the Map task stage, according to a preset mechanism, determine whether the first computing task performs Map task scheduling, and according to whether Map task scheduling needs to be performed, determine the clusters participating in the execution of the Map task, and simulate splitting the first computing task into task key-value pairs through the clusters participating in the execution of the Map task.

[0084] Exemplarily, the simulating splitting the first computing task into task key-value pairs includes:

[0085] In the case where the first computing task does not perform Map task scheduling, determine the target clusters where the cluster data related to the first computing task is located, and for each of the target clusters, simulate splitting the first computing task into task key-value pairs through the target clusters; in the case where the first computing task performs Map task scheduling, based on the Map task scheduling policy, simulate allocating the first computing task to each cluster in the cross-domain cluster system, so that each cluster simulates splitting the first computing task into task key-value pairs.

[0086] That is to say, in the case where Map task scheduling is not performed, the clusters where the relevant cluster data is located are determined as target clusters. Each target cluster performs simulating splitting the first computing task into task key-value pairs.

[0087] In the case where Map task scheduling is performed, according to the initial scheduling policy, allocate the first computing task data, and after the allocation, each cluster performs simulating splitting the first computing task into task key-value pairs.

[0088] Exemplarily, in the Map task stage, if the original data of the computing task cannot be transmitted, then perform the Map task locally in the cluster where the data is located. (Here, it is possible to perform the Map task locally because Map operations are all about transforming the data and do not require global data, so the Map task can be directly executed in the cluster where the data is located without necessarily transmitting the data.) If the involved original data can be transmitted, perform task scheduling for the Map task. Task scheduling is to allocate which clusters process which data. For example: the original data block distribution is: (20, 20, 30), and the allocation scheme after task scheduling is: (30, 15, 25).

[0089] S203. In the Shuffle stage, based on the Reduce task scheduling policy, simulate allocating the task key-value pairs to each cluster in the cross-domain cluster system.

[0090] Specifically, the data transmission of cross-domain tasks is carried out in the Shuffle phase. In the Shuffle phase, according to the Reduce task scheduling strategy, the same set of data that the Reduce task needs to process is pulled to the same cluster to ensure the correct execution of subsequent Reduce tasks. Therefore, the task transmission involved in the Reduce task is the Shuffle phase. During the data transmission process of the Reduce task, data interaction needs to be carried out between clusters. The amount of tasks assigned to the Reduce task affects the amount of data that each cluster needs to pull from other clusters and the amount of data that the cluster needs to transmit. For example, if a cluster has a small initial data volume and a large amount of tasks are assigned to this cluster, then this cluster needs to pull more data from other clusters, and correspondingly, the data transmitted by this cluster to other clusters is less. Pulling data requires occupying the downlink, and transmitting data requires occupying the uplink. Therefore, task scheduling needs to consider the initial data volume of the cluster, the uplink and downlink bandwidths between clusters. At the same time, the computing time of the task and the amount of task allocation are related to the computing resources of the cluster. For example, one cluster has 3 nodes and 10 tasks are assigned, and another cluster has 10 nodes and 3 tasks are assigned. Then the execution time of the entire task is ultimately limited by the cluster with the slowest computing speed. The cluster with fewer computing resources is always in a high-load state, and the cluster with more cluster resources is always in an idle state. For this reason, task scheduling should comprehensively consider the data transmission time and the computing time to achieve efficient utilization of resources and meet the requirements of complex scenarios.

[0091] Exemplarily, in the Reduce task phase, task scheduling is carried out with reference to the Reduce task scheduling strategy, and the task key-value data pairs are assigned to each cluster in the cross-domain cluster system. After determining the task allocation plan here, the Shuffle operation, that is, data transmission, needs to be carried out. For example: after the Map task ends, the data block distribution is (20, 20, 20), and the allocation plan of the task scheduling strategy is (25, 10, 25). In the shuffle operation, first, the data (K, V values) of each cluster is divided into 60 (20 + 20 + 20) partitions according to the Hash value of K. Cluster 1 processes the first 25 partitions, Cluster 2 processes the middle 10 partitions, and Cluster 3 processes the data of the last 25 partitions. Therefore, Cluster 1 needs to send the data of the middle 10 partitions of Cluster 1's data to Cluster 2 and send the data of the last 25 partitions of Cluster 1's data to Cluster 3. Similarly, Cluster 2 sends the data of the first 25 partitions of Cluster 2's data to Cluster 1 and sends the data of the last 25 partitions of Cluster 2's data to Cluster 3. Cluster 3 sends the data of the first 25 partitions of Cluster 3's data to Cluster 1 and sends the data of the middle 10 partitions of Cluster 3's data to Cluster 2.

[0092] Exemplarily, based on the Reduce task scheduling policy, simulating the allocation of the task key-value pairs to each cluster in the cross-domain cluster system includes: converting the task key-value pairs into data blocks based on the size of the data volume included in a predefined single data block; and allocating the data blocks to each cluster in the cross-domain cluster system based on the Reduce task scheduling policy.

[0093] It should be noted that during the allocation of task key-value pairs and / or during the allocation of the first computing task data in the Map phase, data conversion is required. In a cluster, data is generally stored in HDFS (Hadoop Distributed FileSystem). In HDFS, data is divided into data blocks for storage, and by default, a data block is 128MB. Therefore, define the size D of a data block st with a default value of 128MB, that is, D st = 128. Convert the data, so that one data block corresponds to one subtask when the subsequent cluster executes tasks.

[0094] The initial number of data blocks B of cluster i i is:

[0095]

[0096] where B i represents the initial number of data blocks of cluster i; S i represents the initial data volume of cluster i; D st represents the size of the data volume of a single data block.

[0097] After that, during task scheduling, the data volume allocated to each cluster is also allocated in units of blocks. Assume that the data volume allocated to the i-th cluster is r i data blocks. Select the Map task for task scheduling and define two sets DT and DN. For cluster i,

[0098]

[0099] where DT is the set of data blocks that cluster i needs to transfer to other clusters; DN is the set of data blocks that cluster i needs to receive from other clusters; r i represents the number of data blocks allocated to cluster i; B i represents the initial number of data blocks of cluster i.

[0100] Based on the two sets DN and DT, a data transmission scheme is formulated. For each cluster in the set DN, the set DT is sorted according to the bandwidth between it and the clusters in the set DT, and data is first obtained from the cluster with the largest bandwidth. In this way, a new set DC can be obtained, which represents the data transmission scheme. The individuals in the set are (i, j, Datanum), indicating that Datanum blocks of data are transmitted from cluster i to cluster j.

[0101] S204. In the Reduce task stage, for each of the clusters, using the computing resources of the cluster, task calculations are performed on all task key-value pairs in the cluster to determine the calculation results.

[0102] Among them, all task keys in the cluster include its own task key-value pairs and the assigned task key-value pairs.

[0103] Specifically, after the Shuffle operation, each cluster needs to execute the Reduce task. Each cluster uses its own computing resources to perform task calculations on the task key-value pairs in itself to determine the calculation results, and then the calculation results of the Reduce are aggregated to the master cluster.

[0104] S205. Determine the fitness function corresponding to each task stage in the case of adopting the initial scheduling strategy.

[0105] Exemplarily, according to the first cluster parameter, determine the maximum first time for processing tasks in the Map task stage and the maximum second time for processing tasks in the Reduce task stage; determine the fitness function of the Map task stage as the maximum first time, and determine the fitness function of the Reduce task stage as the maximum second time.

[0106] That is to say, by calculating the maximum first time for processing tasks in the Map task stage, the fitness function of the Map task stage can be determined. By calculating the maximum second time for processing tasks in the Reduce task stage, the fitness function of the Reduce task stage can be determined.

[0107] It should be noted that in the Map task stage, different processing schemes are adopted due to whether the Map task scheduling is executed. Therefore, in the process of determining the task processing time, it is necessary to distinguish whether the Map task scheduling is executed in the Map task stage.

[0108] In the case where the first computing task does not execute the Map task scheduling, the maximum first time is the maximum splitting time for each of the clusters to split the first computing task in the Map task stage.

[0109] In the case where the first computing task does not perform Map task scheduling, there is no need to perform data transmission in the Map task phase, and the maximum first time T C can be calculated from the data volume r i , the number of cluster nodes E i and the time t required for calculating a single data block:

[0110]

[0111] Among them, T C represents the maximum first time; r i represents the number of data blocks allocated by cluster i; B i represents the initial number of data blocks of cluster i; E i represents the number of computing power nodes of cluster i; t represents the time required for calculating a single data block.

[0112] However, in the case where the first computing task performs Map task scheduling, the maximum first time is the sum of the maximum transmission time for transmitting the first computing task between the clusters in the Map task phase and the maximum splitting time for splitting the first computing task.

[0113] In the case where the first computing task performs Map task scheduling, data transmission needs to be performed in the Map task phase, and the maximum first time T S can be jointly determined by the maximum transmission time and the maximum splitting time of the task.

[0114] The maximum transmission time is the longest data transmission time among the clusters, that is:

[0115]

[0116] Among them, T t represents the maximum transmission time; T i represents the transmission time of cluster i; Datanum represents the number of data blocks that cluster i needs to transmit to other clusters; U i,j represents the downlink bandwidth from cluster to cluster; DT is the set of data that cluster i needs to transmit to other clusters.

[0117] The maximum splitting time T C is calculated from the data volume r i , the number of cluster nodes E i and the time t required for calculating a single data block:

[0118]

[0119] Among them, T C represents the maximum splitting time; r iIndicates the number of data blocks allocated to cluster i; B i Indicates the initial number of data blocks of cluster i; E i Indicates the number of computing nodes in cluster i; t represents the time required for computing a single data block.

[0120] Finally, the total time required for the execution of the Map task can be obtained, that is, the maximum first time:

[0121] T S = T t + T c ;

[0122] Among them, T S Indicates the maximum first time; T t Indicates the task transmission time; T C Indicates the task computing time.

[0123] The maximum second time is the sum of the maximum transmission time for transmitting the task key-value pairs between the clusters in the Shuffle stage and the maximum computing time for computing the results of each cluster in the Reduce task stage.

[0124] The maximum transmission time for transmitting the task key-value pairs between the clusters in the Shuffle stage can be expressed as:

[0125]

[0126] Among them, T t Indicates the maximum transmission time; S i Indicates the initial data volume of cluster i; r i Indicates the number of data blocks allocated to cluster i; D st Indicates the data volume size of a single data block; S represents the overall data volume that needs to be transmitted across the cluster system; U i,j Indicates the uplink bandwidth or downlink bandwidth between clusters.

[0127] The maximum computing time for computing the results of each cluster in the Reduce task stage can be expressed as:

[0128]

[0129] Among them, T d Indicates the maximum computing time; r i Indicates the number of data blocks allocated to cluster i; B i Indicates the initial number of data blocks of cluster i; E i Indicates the number of computing nodes in cluster i; t represents the time required for computing a single data block.

[0130] Then, the maximum second time can be expressed as:

[0131] T v = T t + T d ;

[0132] Wherein, T v represents the maximum second time; T t represents the maximum transmission time; T d represents the maximum calculation time.

[0133] It should be noted that the method for determining the task processing time required for any scheduling strategy is the same as the method for determining the task processing time required for the initial scheduling strategy, and can be applied to the technical solutions of any embodiment of the present invention. The present invention will not repeat the description here.

[0134] S206. For each task stage, based on the genetic algorithm, perform an evolutionary operation on the task scheduling strategy corresponding to the minimum fitness function to obtain the evolved task scheduling strategy for each task stage, and determine the fitness function corresponding to the evolved task scheduling strategy.

[0135] S207. Loop to execute the task scheduling strategy evolutionary operation. When the preset convergence condition is met, determine that the loop execution operation ends.

[0136] In the actual scenario of loop executing the evolutionary task scheduling strategy, it is necessary to first determine whether to execute the Map task scheduling.

[0137] When it is not necessary to execute the Map task scheduling, perform the Map task in the default manner, and there is no need to execute the Map task scheduling strategy in the initial scheduling strategy. Of course, there is also no need to execute the evolutionary operation of the task scheduling strategy in the Map task stage. According to the result output by the Map task stage, loop to execute the evolutionary operation of the task scheduling strategy in the Reduce task stage.

[0138] When it is necessary to execute the Map task scheduling, first perform the loop evolution of the task scheduling strategy in the Map task stage. After completing the loop evolution of the Map task stage, the optimal Map task scheduling strategy for the Map task stage can be obtained. After obtaining the optimal Map task scheduling strategy, the loop evolution of the task scheduling strategy in the Reduce task stage can be performed.

[0139] When the loop evolution of the Reduce task stage ends, the optimal Map task scheduling strategy and the optimal Reduce task scheduling strategy can be obtained.

[0140] S208. According to the Map task scheduling strategy corresponding to the minimum fitness function, and the Reduce task scheduling strategy corresponding to the minimum fitness function.

[0141] In the technical solution of the embodiment of the present invention, by segmenting the first computing task based on the Map-Reduce task model, the complexity of processing the first computing task is simplified, and it has good scalability and fault tolerance. At the same time, a selection strategy for whether to execute the Map task scheduling is added in the Map task stage, which adapts to different scenario requirements, improves the flexibility of task scheduling, improves resource utilization, avoids resource idleness or overload, and at the same time supports privacy and security requirements, reducing the risk of data leakage.

[0142] Embodiment III

[0143] Figure 3 FIG. is a schematic structural diagram of a cross-domain task scheduling strategy determination device provided in Embodiment III of the present invention. As Figure 3 shown, the device includes:

[0144] A cross-domain task acquisition module 301, configured to acquire a first computing task that needs to perform cross-domain computing, and first cluster parameters corresponding to each cluster in a cross-domain cluster system for processing the first computing task, where the cross-domain cluster system performs computing tasks based on the Map-Reduce task model, and the Map-Reduce task model includes a Map task stage and a Reduce task stage, and the first cluster parameters at least include a transmission bandwidth parameter and a node computing power parameter;

[0145] A cross-domain task allocation module 302, configured to, for each task stage, based on a preset initial scheduling strategy, perform task allocation processing between each cluster by simulating according to the first cluster parameters and the first computing task, and determine a fitness function corresponding to each task stage when adopting the initial scheduling strategy, where the initial scheduling strategy includes a Map task scheduling strategy and a Reduce task scheduling strategy, and the fitness function is determined according to the task processing time of each task stage;

[0146] A scheduling strategy optimization module 303, configured to, for each task stage, based on a genetic algorithm, perform an evolution operation on the task scheduling strategy corresponding to the minimum fitness function to obtain an evolved task scheduling strategy for each task stage, and determine a fitness function corresponding to adopting the evolved task scheduling strategy, where in the initial loop state of each task stage, the task scheduling strategy corresponding to the minimum fitness function is the initial scheduling strategy of each task stage;

[0147] A strategy loop optimization module 304, configured to loop and execute the task scheduling strategy evolution operation, and when a preset convergence condition is met, determine that the loop execution operation ends, where when the loop execution of the Map task scheduling strategy ends, the Reduce task scheduling strategy evolution operation is executed;

[0148] A scheduling policy determination module 305, configured to determine a target scheduling policy according to the Map task scheduling policy corresponding to the minimum fitness function and the Reduce task scheduling policy corresponding to the minimum fitness function.

[0149] The technical solution of the embodiment of the present invention can know the upload data time, download data time, and data processing time of each cluster by obtaining the first computing task that needs to perform cross-domain computing and the first cluster parameters corresponding to each cluster in the cross-domain cluster system that processes the first computing task. For each task stage, based on a preset initial scheduling policy, according to the first cluster parameters and the first computing task, simulate the task allocation process between each cluster, and determine the fitness function corresponding to each task stage when adopting the initial scheduling policy. For each task stage, based on the genetic algorithm, perform an evolutionary operation on the task scheduling policy corresponding to the minimum fitness function to obtain an evolved task scheduling policy, and determine the fitness function required for adopting the evolved task scheduling policy, and perform a loop to execute the task scheduling policy evolutionary operation, fully considering the differences in computing resources, network bandwidth, and initial data volume of each cluster, and being able to make a globally optimal scheduling plan in complex situations. When the preset convergence condition is met, determine that the loop execution operation ends, and determine the Map task scheduling policy corresponding to the minimum fitness function and the Reduce task scheduling policy corresponding to the minimum fitness function as the target scheduling policy, and determine it as the final target scheduling policy. The cross-domain task scheduling policy determination method of the present invention can propose an optimal task allocation plan in different scenarios, effectively balance the task execution time and transmission time, and thus can improve the cross-domain task execution efficiency.

[0150] Optionally, the task processing of the Map-Reduce task model includes a Map task stage, a Shuffle stage, and a Reduce task stage;

[0151] The cross-domain task allocation module 302 includes:

[0152] A Map task execution unit, configured to, based on the Map-Reduce task model, in the Map task stage, according to a preset mechanism, determine whether to execute the Map task scheduling for the first computing task, and simulate splitting the first computing task into task key-value pairs, where the preset mechanism at least includes the task confidentiality type;

[0153] A Shuffle stage execution unit, configured to, in the Shuffle stage, based on the Reduce task scheduling policy, simulate allocating the task key-value pairs to each cluster in the cross-domain cluster system;

[0154] The Reduce task execution unit is used to perform task calculations on all task key-value pairs in the cluster using the computing resources of the cluster during the Reduce task phase, and determine the calculation results, where all task keys in the cluster include its own task key-value pairs and the assigned task key-value pairs.

[0155] Optionally, the Shuffle phase execution unit is specifically used for:

[0156] Convert the task key-value pairs into data blocks based on the size of the data volume included in a predefined single data block;

[0157] Based on the Reduce task scheduling policy, it is planned to allocate the data blocks to each cluster in the cross-domain cluster system.

[0158] Optionally, the Map task execution unit is specifically used for:

[0159] In the case where the first computing task does not execute the Map task scheduling, determine the target cluster where the cluster data related to the first computing task is located, and for each target cluster, simulate splitting the first computing task into task key-value pairs through the target cluster;

[0160] In the case where the first computing task executes the Map task scheduling, based on the Map task scheduling policy, simulate allocating the first computing task to each cluster in the cross-domain cluster system, so that each cluster simulates splitting the first computing task into task key-value pairs.

[0161] Optionally, the scheduling policy optimization module 303 is further specifically used for:

[0162] According to the first cluster parameter, determine the maximum first time for processing tasks in the Map task phase and the maximum second time for processing tasks in the Reduce task phase;

[0163] Determine the maximum first time as the fitness function for the Map task phase and the maximum second time as the fitness function for the Reduce task phase.

[0164] Optionally, in the case where the first computing task does not execute the Map task scheduling, the maximum first time is the maximum splitting time for each cluster to split the first computing task in the Map task phase;

[0165] When the first computing task executes the Map task scheduling, the maximum first time is the sum of the maximum transmission time for transmitting the first computing task between the clusters in the Map task stage and the maximum splitting time for splitting the first computing task.

[0166] The maximum second time is the sum of the maximum transmission time for transmitting the task key-value pairs between the clusters in the Shuffle stage and the maximum computing time for computing the results of each cluster in the Reduce task stage.

[0167] Optionally, the scheduling policy optimization module 303 is specifically configured to:

[0168] For each task stage, in the task scheduling policy corresponding to the minimum fitness function, randomly exchange the assigned task amounts of two of the clusters to obtain an initial evolutionary policy.

[0169] Randomly vary the assigned task amount of one of the clusters in the initial evolutionary policy by a random amount, and make corresponding changes to the assigned task amounts of other clusters so that the overall assigned task quantity remains unchanged, to obtain the task scheduling policy after evolution for this task stage.

[0170] The cross-domain task scheduling policy determination device provided by the embodiments of the present invention can execute the cross-domain task scheduling policy determination method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0171] Embodiment 4

[0172] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0173] As Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0174] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0175] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the cross-domain task scheduling policy determination method.

[0176] In some embodiments, the cross-domain task scheduling policy determination method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the cross-domain task scheduling policy determination method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the cross-domain task scheduling policy determination method by any other appropriate means (e.g., by means of firmware).

[0177] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0178] The computer program for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0179] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0181] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0182] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0183] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0184] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for determining a cross-domain task scheduling strategy, characterized in that: include: Obtaining a first computing task that needs to be cross-domain computing, and first cluster parameters corresponding to each cluster in a cross-domain cluster system that processes the first computing task, wherein the cross-domain cluster system performs computing tasks based on a Map-Reduce task model, the Map-Reduce task model includes a Map task stage and a Reduce task stage, and the first cluster parameters include at least a transmission bandwidth parameter and a node computing power parameter; For each task stage, based on a preset initial scheduling strategy, according to the first cluster parameter and the first computing task, simulate the task allocation process between the clusters, and determine the fitness function corresponding to each task stage when the initial scheduling strategy is adopted, wherein the initial scheduling strategy includes a Map task scheduling strategy and a Reduce task scheduling strategy, and the fitness function is determined according to the task processing time of each task stage; For each task stage, based on the genetic algorithm, the task scheduling strategy corresponding to the minimum fitness function is evolved to obtain the evolved task scheduling strategy of each task stage, and the fitness function corresponding to the evolved task scheduling strategy is determined, wherein, in the initial cycle state of each task stage, the task scheduling strategy corresponding to the minimum fitness function is the initial scheduling strategy of each task stage; cyclically executing the task scheduling strategy evolution operation, and when a preset convergence condition is met, determining that the cyclic execution operation is ended, wherein, when the cyclic execution of the Map task scheduling strategy is ended, executing the Reduce task scheduling strategy evolution operation; The target scheduling strategy is determined according to the Map task scheduling strategy corresponding to the minimum fitness function and the Reduce task scheduling strategy corresponding to the minimum fitness function.

2. The method according to claim 1, characterized in that The Map-Reduce task model task processing includes a Map task stage, a Shuffle stage and a Reduce task stage; the task allocation processing between the clusters is simulated based on the preset initial scheduling strategy according to the first cluster parameters and the first computing task, including: Based on the Map-Reduce task model, in the Map task phase, according to a preset mechanism, determining whether the first computing task executes Map task scheduling, and simulating splitting the first computing task into task key-value pairs, wherein the preset mechanism at least includes a task confidentiality type; In the Shuffle phase, based on the Reduce task scheduling strategy, simulate allocating the task key-value pairs to each cluster in the cross-domain cluster system; In the Reduce task phase, for each cluster, the computing resources of the cluster are used to perform task calculations on all task key-value pairs in the cluster to determine calculation results, wherein all task key values ​​in the cluster include their own task key-value pairs and assigned task key-value pairs.

3. The method according to claim 2, characterized in that The simulating the distribution of the task key-value pairs to each cluster in the cross-domain cluster system based on the Reduce task scheduling strategy includes: Based on the predefined data size contained in a single data block, convert the task key-value pair into a data block; Based on the Reduce task scheduling strategy, the data blocks are planned to be distributed to each cluster in the cross-domain cluster system.

4. The method according to claim 2, characterized in that: The simulation splits the first computing task into task key-value pairs, including: In the case where the first computing task does not perform Map task scheduling, determining a target cluster where cluster data related to the first computing task is located, and for each of the target clusters, simulating splitting the first computing task into task key-value pairs through the target cluster; When the first computing task executes Map task scheduling, based on the Map task scheduling strategy, the first computing task is simulated to be distributed to each cluster in the cross-domain cluster system, so that each cluster simulates splitting the first computing task into task key-value pairs.

5. The method according to claim 2, characterized in that: The determining of the fitness function corresponding to each task stage when the initial scheduling strategy is adopted includes: Determine, according to the first cluster parameter, a maximum first time for processing a task in the Map task phase and a maximum second time for processing a task in the Reduce task phase; The maximum first time is determined as the fitness function of the Map task stage, and the maximum second time is determined as the fitness function of the Reduce task stage.

6. The method according to claim 5, characterized in that The method comprises: In the case where the first computing task does not perform Map task scheduling, the maximum first time is the maximum splitting time of the first computing task split by each of the clusters in the Map task phase; In the case where the first computing task executes Map task scheduling, the maximum first time is the sum of the maximum transmission time of transmitting the first computing task between the clusters in the Map task phase and the maximum splitting time of splitting the first computing task; The maximum second time is the sum of the maximum transmission time of the task key-value pairs between the clusters in the Shuffle phase and the maximum calculation time of the calculation results of the clusters in the Reduce task phase.

7. The method according to claim 1, characterized in that Based on the genetic algorithm, the task scheduling strategy corresponding to the minimum fitness function is evolved to obtain the evolved task scheduling strategy of each task stage, including: For each task stage, in the task scheduling strategy corresponding to the minimum fitness function, the assigned task amounts of the two clusters are randomly exchanged to obtain an initial evolutionary strategy; The amount of assigned tasks of one of the clusters in the initial evolution strategy is randomly changed, and the amount of assigned tasks of other clusters is changed accordingly, so that the overall number of assigned tasks remains unchanged, and the task scheduling strategy after evolution of the task stage is obtained.

8. A device for determining a cross-domain task scheduling strategy, characterized in that: include: A cross-domain task acquisition module, used to acquire a first computing task that needs to be cross-domain computing, and first cluster parameters corresponding to each cluster in a cross-domain cluster system that processes the first computing task, wherein the cross-domain cluster system performs computing tasks based on a Map-Reduce task model, the Map-Reduce task model includes a Map task stage and a Reduce task stage, and the first cluster parameters include at least a transmission bandwidth parameter and a node computing power parameter; A cross-domain task allocation module, for each task stage, based on a preset initial scheduling strategy, according to the first cluster parameter and the first computing task, simulates the task allocation processing between the clusters, and determines the fitness function corresponding to each task stage when the initial scheduling strategy is adopted, wherein the initial scheduling strategy includes a Map task scheduling strategy and a Reduce task scheduling strategy, and the fitness function is determined according to the task processing time of each task stage; The scheduling strategy optimization module is used to perform an evolution operation on the task scheduling strategy corresponding to the minimum fitness function for each task stage based on a genetic algorithm, obtain the evolved task scheduling strategy of each task stage, and determine the fitness function corresponding to the evolved task scheduling strategy, wherein, in the initial cycle state of each task stage, the task scheduling strategy corresponding to the minimum fitness function is the initial scheduling strategy of each task stage; A strategy cycle optimization module, used for cyclically executing the task scheduling strategy evolution operation, and determining the end of the cyclic execution operation when a preset convergence condition is met, wherein, when the Map task scheduling strategy cyclic execution ends, the Reduce task scheduling strategy evolution operation is executed; The scheduling strategy determination module is used to determine the target scheduling strategy according to the Map task scheduling strategy corresponding to the minimum fitness function and the Reduce task scheduling strategy corresponding to the minimum fitness function.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the cross-domain task scheduling strategy determination method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for determining a cross-domain task scheduling strategy according to any one of claims 1 to 7 when executed.

Citation Information

Cited By

  • Repartitioning method and system based on cross-domain data skew, terminal and storage medium

    CN122064913A

  • A re-partitioning method and system based on cross-domain data skew, a terminal and a storage medium

    CN122064913B

  • Method and apparatus for determining cross-domain task scheduling strategy, and device and storage medium

    WO2026184055A1