DDQN-based multi-target remote sensing task scheduling method and system
By refining the remote sensing task into multiple subtasks and using the DDQN algorithm, the problems of long delay time and uneven resource allocation in traditional scheduling methods are solved, and efficient resource utilization and task scheduling are achieved.
Patent Information
- Application Number
- CN202510090591.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional remote sensing task scheduling methods are difficult to cope with the problems of long delay time, uneven resource allocation, high node energy consumption and low resource utilization efficiency at the same time, especially under high-density task loads.
The multi-objective remote sensing task scheduling method based on DDQN is adopted, and the remote sensing task is refined into multiple subtasks, an algorithm collection is constructed, and the optimal execution node is determined using the policy network, combining the pruning strategy and the target network soft update strategy, model training and resource allocation are optimized.
It effectively reduces the delay time, achieves the uniformity of resource allocation, reduces node energy consumption, and improves resource utilization efficiency.
Smart Images

Figure CN120066766A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of task scheduling, and particularly to a multi-objective remote sensing task scheduling method and system based on DDQN. Background Art
[0002] With the rapid development of remote sensing technology, remote sensing products play an increasingly important role in many fields such as disaster monitoring, environmental change analysis, urban planning, and agricultural yield estimation. However, whether in the disaster monitoring or environmental analysis scenarios, the production of remote sensing products is usually accompanied by high concurrency and resource competition, which poses a huge challenge to the efficiency and accuracy of the remote sensing task scheduling system, prompting scheduling optimization to become the core issue in remote sensing task management.
[0003] Some traditional scheduling methods, such as the first-come-first-served scheduling algorithm, the shortest job first scheduling algorithm, the round-robin scheduling algorithm, the dynamic scheduling algorithm, the ant colony algorithm, the genetic algorithm, the particle swarm algorithm, etc., in the remote sensing task scheduling scenario, mainly have the following two problems: (1) These methods are difficult to simultaneously address problems such as long delay time, uneven resource allocation, and high node energy consumption during the scheduling process, especially showing great limitations when facing high-density task loads. (2) Remote sensing tasks usually have a high degree of diversity, with different task scales, priorities, execution time requirements, etc. These methods lack an effective mechanism to allocate resources according to the diversity of tasks, resulting in reduced resource utilization efficiency or resource waste, and cannot meet the requirements of efficient scheduling of actual remote sensing tasks. Summary of the Invention
[0004] In order to at least partially solve the problems of long delay time, uneven resource allocation, high node energy consumption, and low resource utilization efficiency of traditional scheduling methods, the present invention provides a multi-objective remote sensing task scheduling method and system based on DDQN. By processing multiple remote sensing tasks, each remote sensing task is refined into multiple subtasks (algorithms) to obtain an algorithmic representation of the remote sensing task and calculate its value. On this basis, a scheduling model that can perceive state features and algorithm features is designed to construct an algorithm subset. A reward function is designed by comprehensively considering the production time of the algorithm and the node resource status, and a policy network is used to determine the optimal execution node of the algorithm. The DDQN algorithm is used to train the model, and a pruning strategy is adopted during the training process to effectively reduce the search space and reduce the computational complexity. At the same time, a target network soft update strategy is adopted to accelerate the model convergence speed and improve the algorithm stability. The present invention can reduce the delay time, make the resource allocation uniform, reduce the node energy consumption, and improve the resource utilization efficiency.
[0005] To achieve the above object, the technical solution of the present invention is:
[0006] The first aspect of the present invention proposes a multi-objective remote sensing task scheduling method based on DDQN, including:
[0007] Step 1: Input multiple given remote sensing tasks into a preset task processing unit to obtain algorithmic representations of the multiple remote sensing tasks; wherein, the preset task processing unit includes multiple virtual machines and multiple physical machines, and the multiple virtual machines are respectively arranged on the multiple physical machines; facilitating the execution of remote sensing tasks according to the algorithmic representations of the remote sensing tasks.
[0008] Step 2: Construct an algorithm set based on the algorithmic representations of the multiple remote sensing tasks, and obtain the resources required by each algorithm in the algorithm set and the value of the algorithm, facilitating the scheduling of the algorithms.
[0009] Step 3: Input the algorithm set into a preset scheduling model to obtain multiple algorithm subsets, facilitating the parallel processing of the algorithms.
[0010] Step 4: Allocate the multiple algorithm subsets to multiple nodes respectively according to the policy network and the resources required by each algorithm; wherein, the nodes include multiple virtual machines and multiple physical machines; facilitating the execution of the algorithms according to the optimal nodes.
[0011] Step 5: Operate on the algorithm subsets allocated to the nodes to obtain the results of the given multiple remote sensing tasks.
[0012] Further, the inputting of the multiple given remote sensing tasks into the preset task processing unit to obtain algorithmic representations of the multiple remote sensing tasks specifically includes:
[0013] Split each remote sensing task into multiple algorithms required by it, and sort the multiple algorithms according to the running order to obtain the algorithmic representation E i ={a i , a 2 , a 3 ,..., a n}; wherein, E i is the i-th remote sensing task, and a n is the n-th algorithm; facilitating the subsequent construction of algorithm subsets according to the algorithmic representations of the remote sensing tasks.
[0014] Further, the value of the algorithm is expressed by the following formula:
[0015] V ai =w l ×P ai +w r ×(CT - R ai ) + w d ×ID ai ×Dep ai
[0016] Wherein, Vai For algorithm a i The value, w l For algorithm a i The weight of the priority, P ai For algorithm a i The priority, w r The weight of the reception time, CT is the current time, R ai The reception time, w d The weight of the input data, ID ai The size of the input data, Dep ai For algorithm a i The dependencies.
[0017] Furthermore, the resources required for each algorithm in the algorithm set include the unique identifier of the algorithm, the unique identifier of the affiliated remote sensing task, the algorithm name, the algorithm priority, millions of instructions per second, the size of the input data, the production operating system, the CPU configuration, the memory size, the storage size, the GPU requirements, the reception time, and the preset completion time.
[0018] Furthermore, the preset scheduling model includes a Markov decision process;
[0019] The states of the scheduling model include: the state space is defined as the state feature S according to the input multiple algorithms and task processing units rp And the algorithm feature Among them, the feature matrix S rp Includes the identifier for distinguishing physical machines and virtual machines, the unique identifier of the node, the current number of cores of the virtual machine and the physical machine, the current memory size of the virtual machine and the physical machine, the current storage size of the virtual machine and the physical machine, the broadband of the virtual machine and the physical machine, the performance of the virtual machine processing unit and the performance of the physical machine processing unit, the percentage of the remaining CPU amount in the total amount, the percentage of the remaining memory amount in the total amount, the percentage of the remaining storage amount in the total amount, the number of products running on the virtual machine and the number of products running on the physical machine, and the number of virtual machines on the physical machine; the algorithm feature Includes CPU configuration, millions of instructions per second, storage size, input data size, memory size, reception time, and preset completion time;
[0020] The movement of the scheduling model is expressed by the following formula:
[0021] action={0,1,2,...,n - 1}
[0022] Among them, action is the action process, and n represents the total number of virtual machines;
[0023] The state transition of the scheduling model is expressed by the following formula:
[0024] Srp+1 = U(S rp , action i )
[0025] where S rp+1 is the state feature of the next state, and U is the state transition process;
[0026] The reward function of the scheduling model is expressed by the following formula:
[0027] R total = α 3 · R t + β 3 · R util_p + γ 3 · R util_v
[0028]
[0029] where R total is the value of the reward function, α 3 is the weight coefficient of the reward value for the task production time, R t is the reward value for the task production time, β 3 is the weight coefficient of the reward value for the physical machine, R util_p is the reward value for the physical machine, γ 3 is the weight coefficient of the reward value for the virtual machine, R util_v is the reward value for the virtual machine, σ cpu_p , σ mem_p and σ str_p are the standard deviations of the remaining CPU, memory, and storage on the physical machine respectively, u cpu_p , u mem_p and u str_p are the average values of the remaining CPU, memory, and storage on the physical machine respectively, σ cpu_v , σ mem_v and σ str_v are the standard deviations of the remaining CPU, memory, and storage on the virtual machine respectively, u cpu_v , u mem_v and u str_v are the average values of the remaining CPU, memory, and storage on the virtual machine respectively, α 1 , β 1 and γ 1 are the weights of the remaining CPU, memory, and storage on the physical machine respectively, α 2 , β 2 and γ 2 are the weights of the remaining CPU, memory, and storage on the virtual machine respectively, is the execution time of algorithm a i , is algorithm ai The total production time, where ε is a number close to 0 and is a number close to 0.
[0030] Furthermore, inputting the algorithm set into a preset scheduling model to obtain multiple algorithm subsets specifically includes:
[0031] Constructing multiple algorithm subsets A according to the algorithms in the same order in the algorithm representation of each remote sensing task j ={a 1j ,a 2j ,a 3j ,...,a nj}; where A j is the j-th algorithm subset, and a nj is the j-th algorithm in the algorithm representation of the n-th remote sensing task.
[0032] Furthermore, the policy network adopts a multi-layer fully connected structure, and the policy network includes an input layer, multiple hidden layers, and an output layer;
[0033] The input layer includes state_dim + task_dim neurons for receiving the state feature S rp and the algorithm feature where task_dim is the algorithm feature dimension and state_dim is the state feature dimension;
[0034] Multiple hidden layers include the same number of neurons;
[0035] The output layer includes action_dim neurons, where action_dim is the number dimension of virtual machines.
[0036] Furthermore, the policy network is optimized and trained according to the DDQN algorithm. The DDQN algorithm is used for action selection and value evaluation, and a loss function is constructed based on the value evaluation to update the policy network parameters;
[0037] The action selection is represented by the following formula:
[0038] a' = arg max Q(S rp+1 , action i ; θ)
[0039] where a' is the action selected by the current policy network Q at S rp+1 , θ is the current policy network parameter, and arg max is the set of maximum independent variable points;
[0040] The value evaluation is represented by the following formula:
[0041] y = r rp + γQtarget (S rp+1 , a'; θ - )
[0042] Among them, y is the Q value calculated by the target network Q target , r rp is the immediate reward, γ is the discount factor, and θ - are the target network parameters, and Q target is the target network;
[0043] The loss function is expressed by the following formula:
[0044]
[0045] Among them, ω* is the value of the loss function, and N is the sample batch size randomly selected from the experience pool.
[0046] Furthermore, when the policy network is trained, a pruning threshold is also set, and the target network parameters in the DDQN algorithm are softly updated. The specific process is expressed by the following formula:
[0047] τ = percentile(S, p)
[0048]
[0049] Among them, are the updated target network parameters, θ is the current policy network parameter, are the current parameters of the target network, τ is the pruning threshold, percentile(S, p) is the maximum value of the weights of the smallest p% in the return set, % is the modulo operation, S is the set of all weights, and p is the pruning ratio.
[0050] The first aspect of the present invention proposes a multi-objective remote sensing task scheduling system based on DDQN, including:
[0051] A task processing module, which is used to input multiple given remote sensing tasks into a preset task processing unit to obtain algorithm representations of the multiple remote sensing tasks and the algorithm values of the remote sensing tasks; among them, the preset task processing unit includes multiple virtual machines and multiple physical machines, and the multiple virtual machines are respectively arranged on the multiple physical machines; it is convenient to execute remote sensing tasks according to the algorithm representations of the remote sensing tasks;
[0052] An algorithm set module, which is used to construct an algorithm set according to the algorithm representations of multiple remote sensing tasks, and obtain the resources required by each algorithm in the algorithm set and the value of the algorithm, so as to facilitate the scheduling of the algorithm;
[0053] A scheduling module, which is used to input the algorithm set into a preset scheduling model to obtain multiple algorithm subsets, so as to facilitate the parallel processing of algorithms;
[0054] A policy network module is used to allocate multiple algorithm subsets to multiple nodes respectively according to the policy network and the resources required by each algorithm; wherein, the nodes include multiple virtual machines and multiple physical machines; this facilitates executing algorithms based on the optimal nodes.
[0055] An execution module is used to perform operations on the algorithm subsets allocated to the nodes to obtain the results of a given multiple remote sensing tasks.
[0056] Advantages of the present invention:
[0057] The present invention refines remote sensing tasks into algorithms, improves resource adaptability, and ensures the correctness and consistency of the task dependency chain in the time dimension by relying on the perception mechanism. The task scheduling process is abstractly modeled as a Markov decision process, and a dynamic decision-making framework is constructed to achieve precise control of the scheduling process. By training a double deep Q-network (DDQN), the model parameters are gradually optimized, and the association between remote sensing tasks and state characteristics is deeply learned, so as to make optimal task scheduling decisions. Experimental results show that the present invention not only effectively improves the task scheduling efficiency, but also significantly reduces the production time of remote sensing tasks and realizes the balanced allocation of node resources, providing an efficient and feasible solution for the scheduling of remote sensing tasks. Brief Description of the Drawings
[0058] Figure 1 It is a flowchart of a multi-objective remote sensing task scheduling method based on DDQN provided by an embodiment of the present invention.
[0059] Figure 2 It is a schematic diagram of a multi-objective remote sensing task scheduling method based on DDQN provided by an embodiment of the present invention.
[0060] Figure 3 It is a schematic diagram of the algorithm representation of a vegetation index product provided by an embodiment of the present invention.
[0061] Figure 4 It is a schematic diagram of a policy network provided by an embodiment of the present invention.
[0062] Figure 5 It is a schematic diagram of the change curve of the DDQN training loss with the number of training steps provided by an embodiment of the present invention.
[0063] Figure 6 It is a schematic diagram of the change curve of the DDQN training reward with the number of training steps provided by an embodiment of the present invention.
[0064] Figure 7 It is a schematic diagram of the change curve of the training loss of different policy models with the number of training steps provided by an embodiment of the present invention.
[0065] Figure 8Schematic diagram of the change curve of the training rewards of different strategy models provided by the embodiments of the present invention with the number of training steps.
[0066] Figure 9 Schematic diagram of the total task completion time of different algorithms provided by the embodiments of the present invention.
[0067] Figure 10 Schematic diagram of the physical machine resource load balancing value under different numbers of tasks provided by the embodiments of the present invention.
[0068] Figure 11 Schematic diagram of the virtual machine resource load balancing value under different numbers of tasks provided by the embodiments of the present invention.
[0069] Figure 12 Architecture diagram of a multi-objective remote sensing task scheduling system based on DDQN provided by the embodiments of the present invention. Detailed implementation manners
[0070] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0071] Embodiment 1
[0072] As Figure 1 and Figure 2 shown, a multi-objective remote sensing task scheduling method (MORS) based on DDQN includes:
[0073] S101: Input a plurality of given remote sensing tasks (remote sensing products) into a preset task processing unit to obtain algorithmic representations of the plurality of remote sensing tasks; wherein, the preset task processing unit includes a plurality of virtual machines and a plurality of physical machines, and the plurality of virtual machines are respectively arranged on the plurality of physical machines.
[0074] Specifically, the reception and production of remote sensing tasks is a real-time process. Within a time window T, T = [T i , T i+1 , where T i is the i-th time point. The task processing unit receives a set of remote sensing tasks and places these remote sensing tasks in a task queue, and defines the remote sensing task queue as W = {w 1 , w 2 , w 3 ,..., w n}, where W is the remote sensing task queue and w nis the nth remote sensing task. The virtual machine set is defined as VM = {vm 1 , vm 2 , vm 3 ,..., vm n}, where VM is the virtual machine set and vm n is the nth virtual machine. Each virtual machine is distributed in a different physical machine. The physical machine set is defined as PM = {pm 1 , pm 2 , pm 3 ,..., pm m}, where PM is the physical machine set and pm m is the mth physical machine.
[0075] Each remote sensing product consists of several algorithms, which are the basic execution units of the task and are responsible for completing specific data processing functions. There is a strict order among the algorithms of the remote sensing product, that is, the execution of subsequent algorithms depends on the calculation results of the previous algorithms. Therefore, it can be represented by a flowchart. Taking the vegetation index product NDVIs as an example, the flowchart is as Figure 3 shown.
[0076] Figure 3 The nodes from a 1 to a 4 in
[0077] represent algorithms, namely the geometric rectification algorithm, the apparent radiance algorithm, the surface reflectance algorithm, and the vegetation index algorithm respectively. The edges represent the dependency relationships between the algorithms. i Split each remote sensing task into multiple remote sensing algorithms required by it. Sort the multiple remote sensing algorithms according to the running order to obtain the algorithm representation E i = {a 2 , a 3 , a n ,..., a i}, where E n is the ith remote sensing task and a
[0078] S102: Construct an algorithm set based on the algorithm representations of multiple remote sensing tasks, and obtain the resources required by each algorithm in the algorithm set and the value of the algorithm.
[0079] Specifically, since the production of remote sensing products consumes node resources, it is necessary to record the attribute information of physical machines and virtual machines, as shown in Table 1.
[0080] Table 1 Physical machine and virtual machine node attribute information
[0081]
[0082] The execution of an algorithm usually includes steps such as data processing, image generation, and information extraction. These steps have a relatively fixed logical structure. Therefore, each algorithm clearly presets the required resource information, including computing resources, memory requirements, storage space, etc.
[0083] Specifically, the resources required for each algorithm in the algorithm set include the unique identifier of the algorithm, the unique identifier of the affiliated remote sensing task, the algorithm name, the algorithm priority, millions of instructions per second, the size of the input data, the operating system in production, the CPU configuration, the memory size, the storage size, the GPU requirements, the reception time, and the preset completion time. Specifically as shown in Table 2.
[0084] Table 2 Resource information required by the algorithm
[0085]
[0086] MORS follows a timing scheduling mechanism, periodically updates W according to the preset T, and calculates w after the update is completed i the value of each algorithm in, and the value calculation is expressed by the following formula:
[0087]
[0088] Among them, is the value of algorithm a i w l is the weight of the priority of algorithm a i , is the priority of algorithm a i w r is the weight of the reception time, CT is the current time, is the reception time, w d is the weight of the input data, is the size of the input data, is algorithm a i 's dependency.
[0089] Preferably, in order to ensure the correct execution of algorithms with dependency constraints, the present invention proposes a dependency-aware batch scheduling mechanism. First, the dependency-aware mechanism determines the processing status of an algorithm by judging whether the value of the algorithm is 0. For algorithms with a value of 0, they will not be processed to ensure the correctness of the dependency relationship. Specifically, when the value of an algorithm is 0, it means that its previous dependent algorithm has not been completed and the execution condition of the current algorithm is not satisfied, thus avoiding incorrect execution caused by unmet dependency conditions. For algorithms with a value not equal to 0, the mechanism will screen out the set for scheduling according to the algorithm value size. The algorithm set is defined as A = {a i , a 2 , a 3 ,..., a n}. This value - priority strategy can ensure that algorithms with higher priorities, earlier reception times, and larger data volumes are given priority to obtain resources for execution, thereby improving the overall efficiency and response ability of the system.
[0090] Specifically, according to the value - priority strategy, a set of algorithms available for the scheduler to execute is selected and defined as \(A=\{a\) i ,a 2 ,a 3 ,\(\cdots,a\) n \}. There are no dependency constraints among the algorithms in \(A\), so they can be executed simultaneously through parallel processing. To ensure the correct execution of algorithms with dependency constraints, the present invention proposes a dependency - aware batch scheduling mechanism. This mechanism organizes remote - sensing tasks through topological sorting and constructs a hierarchical structure based on the dependency relationship. It preferentially schedules algorithms with clear upstream dependencies to execute in earlier batches, so as to ensure that subsequent algorithms can enter the next batch for execution in a timely manner after the dependency relationship is satisfied.
[0091] The algorithms of remote - sensing tasks belong to atomic algorithms, and their execution is "non - preemptive". To ensure the smooth production of remote - sensing products, the remaining resources of the virtual machine (including the processing power of idle CPUs, the available capacity of memory, and the remaining storage space) should be greater than or equal to the resources required by the algorithm to be executed. Before the algorithm execution, the system checks the resource status of the virtual machine. If the remaining resources of the virtual machine are insufficient to meet the algorithm requirements, the execution will be delayed, and \(a\) i is stored in the set of algorithms waiting to run \(A\) wait =\(\{a\) 1 ,a\) 2 ,a\) 3 ,\(\cdots,a\) m \}, until the virtual machine has idle resources.
[0092] S103: Input the set of algorithms into a preset scheduling model to obtain multiple subsets of algorithms.
[0093] Specifically, the preset scheduling model constructs multiple subsets of algorithms \(A\) j =\(\{a\) 1j ,a\) 2j ,a\) 3j ,\(\cdots,a\) nj \}; where \(A\) j is the \(j\) - th subset of algorithms, and \(a\) nj is the \(j\) - th algorithm in the algorithm representation of the \(n\) - th remote - sensing task.
[0094] Model the remote - sensing task scheduling process as a Markov decision process. The following is a specific description of the modeling process, including several basic elements such as state, action, state transition, and reward.
[0095] Status: MORS defines the state space as state feature S rp and algorithm feature T ai , specifically as follows: S rp The size is m×n, where m is the total number of nodes (the total number of physical machines and virtual machines), and n is the number of features of each node. S rp Each row is [p_v, ip, c, m, s, B, P, c_util, m_util, s_util, rt, vm_c], where p_v is the identifier for differentiating physical machines and virtual machines, ip is the unique identifier of the node, c is the current number of cores of the virtual machine and the physical machine, m is the current memory size of the virtual machine and the physical machine, s is the current storage size of the virtual machine and the physical machine, B is the broadband of the virtual machine and the physical machine, P is the performance of the virtual machine processing unit and the physical machine processing unit, c_util is the percentage of the remaining CPU amount in the total amount, m_util is the percentage of the remaining memory amount in the total amount, s_util is the percentage of the remaining storage amount in the total amount, rt is the number of products running on the virtual machine and the number of products running on the physical machine, and vm_c is the number of virtual machines on the physical machine. It can be expressed as: where, is the CPU configuration, is millions of instructions per second, is the memory size, is the storage size, is the size of the input data, is the receiving time, is the preset completion time.
[0096] Action: The action is performed on the set of virtual machines. Therefore, the action space can be clearly expressed in a discrete and finite form, and the movement is represented by the following formula:
[0097] action = {0, 1, 2,..., n - 1}
[0098] where action is the action process and n represents the total number of virtual machines. Each action has a unique integer identifier, and this integer corresponds to a specific virtual machine. The execution of the action follows a timing trigger mechanism.
[0099] State transition: After the agent selects the action action i , the node state and the next algorithm waiting to be scheduled are updated. The state transition function is as follows:
[0100] S rp+1 = U(S rp , action i )
[0101] where Srp+1 is the state feature for the next state, and U is the state transition process.
[0102] Reward: The reward function is not only used to evaluate the agent's behavior but also guides the agent to learn the optimal policy. The optimization goal of MORS is to reduce the production time of remote sensing products while achieving resource load balancing. Based on this idea, the design of the reward function is as follows:
[0103]
[0104] Among them, R util_p is the reward value of the physical machine. σ cpu_p , σ mem_p and σ str_p are the standard deviations of the remaining CPU, memory, and storage on the physical machine, respectively. The smaller the standard deviation, the larger the reward value. u cpu_p , u mem_p and u str_p are the averages of the remaining CPU, memory, and storage on the physical machine, respectively. α 1 , β 1 and γ 1 are the weights of the remaining CPU, memory, and storage on the physical machine, respectively. ∈ is a number close to 0 to prevent the denominator from being 0.
[0105]
[0106] Among them, R util_v is the reward value of the virtual machine. σ cpu_v , σ mem_v and σ str_v are the standard deviations of the remaining CPU, memory, and storage on the virtual machine, respectively. The smaller the standard deviation, the larger the reward value. u cpu_v , u mem_v and u str_v are the averages of the remaining CPU, memory, and storage on the virtual machine, respectively. α 2 , β 2 and γ 2 are the weights of the remaining CPU, memory, and storage on the virtual machine, respectively.
[0107] The production time of each algorithm consists of data transfer time, node execution time, and latency time.
[0108] The data transfer time is as shown in the following formula.
[0109]
[0110] Among them, is the data transfer time between algorithm a i and the previous algorithm a j and (VMi ) j is the i-th virtual machine on the physical machine j, is the input data, and B is the bandwidth of the i-th virtual machine on the physical machine j.
[0111] The execution time of the algorithm is determined by dividing its total number of instructions by the unit running ability of the virtual machine where it is located, as shown in the following formula.
[0112]
[0113] Among them, is the execution time (unit: s) of algorithm a i on the i-th virtual machine, is millions of instructions per second, and P is the priority.
[0114] From the above formula, the total production time of algorithm a i can be obtained, as shown in the following formula.
[0115]
[0116] Among them, is the total production time of algorithm a i , is the time when the virtual machine starts to execute a i , is the receiving time of a i .
[0117] The reward value of the task production time is shown in the following formula.
[0118]
[0119] Among them, the reward value R t is represented in a normalized form, used to measure the execution efficiency of the algorithm. The less the total production time, the greater the reward value, and ε is a number close to 0.
[0120] The reward function R total takes into account the comprehensive utilization of physical machine resources, virtual machine resources, and task completion time, as shown in the following formula.
[0121] R total = α 3 ·R t + β 3 ·R util_p + γ 3 ·R util_v
[0122] Among them, R total is the value of the reward function, α 3α is the weight coefficient of the reward value for the task production time, R t β is the reward value for the task production time 3 γ is the weight coefficient of the reward value for the physical machine, R util_p γ is the reward value for the physical machine 3 δ is the weight coefficient of the reward value for the virtual machine, R util_v δ is the reward value for the virtual machine
[0123] The above weight coefficients α i , β i , γ i reflect the importance of each goal and can be adjusted according to specific needs, so as to flexibly optimize different scheduling goals
[0124] S104: Allocate multiple algorithm subsets to multiple nodes respectively according to the policy network and the resources required by each algorithm; among them, the nodes include multiple virtual machines and multiple physical machines
[0125] S105: Operate on the algorithm subsets allocated to the nodes to obtain the results of the given multiple remote sensing tasks
[0126] The present invention splits the remote sensing task into multiple algorithms, constructs algorithm subsets according to the order, and then allocates them to the optimal nodes for operation according to the policy network and the resources required by the algorithms to obtain the results of the remote sensing task. The present invention effectively improves the task scheduling efficiency, significantly reduces the production time of the remote sensing task, and realizes the balanced allocation of node resources, providing an efficient and feasible solution for the scheduling of the remote sensing task
[0127] Embodiment 2
[0128] On the basis of the above embodiment, the embodiment of the present invention provides the structure and training process of the policy network, specifically including
[0129] As Figure 4 shown, the policy network proposed by the present invention adopts a multi-layer fully connected structure, including an input layer, six hidden layers and an output layer. The input layer contains state_dim + task_dim neurons, which are used to receive the state feature S rp and the algorithm feature state_dim is the dimension of the node feature matrix, and the dimension of the state feature is state_dim = m × n. The state feature includes the physical machine-virtual machine relationship feature and the node resource feature. The physical machine-virtual machine relationship feature is obtained by extracting the first column of the matrix, and its dimension is m. task_dim is the algorithm feature dimension, and the dimension of the node resource feature is m × (n - 1). The algorithm feature dimension is task_dim = 7
[0130] All six hidden layers are set with 64 neurons, and fully connected layers and ReLU activation functions are used to extract high-level abstract features. The output layer consists of action_dim neurons, which are connected to the hidden layer through fully connected connections, where action_dim is the number of dimensions of virtual machines.
[0131] In the output layer, no non-linear activation function is used, and the Q-values of each action in the action space are directly mapped and output. Through multi-layer non-linear transformations, the entire network architecture fully extracts complex features in the input data to support decision-making.
[0132] The policy network is optimized and trained according to the DDQN algorithm. DDQN introduces a dual-network structure (i.e., two Q-networks) for action selection and value evaluation. The current network is responsible for selecting the action that is most likely to bring high rewards in the current state, and the target network is responsible for evaluating the action value. The network parameter update strategy is as follows:
[0133] Action selection: Use the current network to select the action corresponding to the maximum estimated value of the next state, as shown in the following formula:
[0134] a' = arg max Q(S rp+1 , action i ; θ)
[0135] where a' is the action selected by the current policy network Q at S rp+1 , θ is the parameter of the current policy network, and arg max is the set of independent variable points of the maximum value.
[0136] Value evaluation: The target network calculates the Q-value of the selected action a′, as shown in the following formula:
[0137] y = r rp + γQ target (S rp+1 , a'; θ - )
[0138] where y is the Q-value calculated by the target network Q target , r rp is the immediate reward, γ is the discount factor, θ - is the parameter of the target network, and Q target is the target network.
[0139] Loss function: The loss function of the current network is constructed in the form of mean square error as shown in the following formula:
[0140]
[0141] where ω* is the value of the loss function, and N is the size of the sample batch randomly selected from the experience pool.
[0142] The policy network designed in the present invention adopts a fully connected neural network. During the training process, it is found that the model converges slowly and there is a phenomenon of gradient explosion. To accelerate convergence and solve the problem of gradient explosion, an unstructured pruning method combined with the L1 norm weight importance evaluation method is adopted. By this method, the connections that have little impact on the model output can be effectively removed, thereby reducing the network complexity and alleviating gradient explosion. The L1 norm evaluation strategy is used to measure the importance of each connection in the network. Specifically, the importance score of a weight can be defined as the absolute value of the weight, as shown in the following formula:
[0143] S ij =|w ij |
[0144] where w ij is the weight between the i-th input unit and the j-th output unit, and S ij is the importance score of the weight between the i-th input unit and the j-th output unit.
[0145] Set a pruning threshold τ 1 , and prune the connections with importance scores lower than the threshold. The pruning operation is as shown in the following formula:
[0146]
[0147] where is the pruned weight and τ is the threshold.
[0148] The threshold can be obtained by calculating the percentile of the weight set, as shown in the following formula:
[0149] τ=percentile(S,p)
[0150] where τ is the pruning threshold, percentile(S,p) is the maximum value of the smallest p% of the weights in the returned set, % is the modulo operation, S is the set of all weights, and p is the pruning ratio.
[0151] During the training of DDQN, the target network parameters are not immediately made exactly the same as the current policy network parameters, but are updated gradually and smoothly. This can reduce the fluctuations during the training process and help avoid the overfitting of the policy network to the current target value estimation. The soft update of the target network is as shown in the following formula:
[0152]
[0153] where is the updated target network parameter, θ is the current policy network parameter, is the current parameter of the target network, and τ is the pruning threshold, which determines the update speed of the target network parameters.
[0154] Embodiment 3
[0155] Based on the above embodiments, an embodiment of the present invention provides a verification method for a multi-objective remote sensing task scheduling method based on DDQN, which specifically includes:
[0156] The experiment uses the Python language to build a simulation environment and selects the PyTorch deep learning framework to train the double deep reinforcement learning algorithm. In order to improve the generality of the online scheduling framework for different production environments, the present invention configures different physical machine and virtual machine parameters to simulate different experimental environments, and details the settings of node parameters in Table 3.
[0157] Table 3 Node Parameters
[0158]
[0159] Remote sensing products themselves have extremely high complexity and huge data volumes, often occupying a large amount of computing and storage resources. However, for specific tasks, the required information is usually limited, far lower than the scale of the complete remote sensing product. Therefore, in order to efficiently process and analyze data, remote sensing products are randomly generated through preset parameters, as shown in Table 4 in detail.
[0160] Table 4 Remote Sensing Product Parameters
[0161]
[0162]
[0163] Combined with the settings of node parameters and product parameters, a preliminary analysis and adjustment are made on the scheduling algorithm, and the detailed training parameters are shown in Table 5.
[0164] Table 5 Training Parameters
[0165]
[0166] The learning rate is a crucial hyperparameter of DDQN. It determines the step size of parameter updates during the training process of the model. Selecting the most appropriate learning rate has a crucial impact on the performance of the model. Figure 5 and Figure 6 are the curves of the loss and reward of the training model with different learning rates varying with the number of training steps when the number of remote sensing tasks is 30, respectively.
[0167] From Figure 5 it can be seen that the training model with a learning rate of 0.001 converges faster, but according to Figure 6From the reward values, it can be seen that the model oscillates severely near the optimal solution. When the learning rate is 0.00001, the model converges slowly and has not started to converge even at the 5000th iteration. In contrast, a learning rate of 0.0001 neither causes excessive oscillation nor is too small to result in an overly slow convergence rate.
[0168] (1) Model ablation experiments
[0169] To deeply verify the effectiveness of applying pruning strategies and target network soft updates (smoothing factor τ = 0.001) during training, as well as the superiority of DDQN compared to DQN, a series of ablation experiments were designed. The following are the experimental results when the number of remote sensing tasks is 30, as Figure 7 and Figure 8 shown.
[0170] From Figure 7 it can be seen that DQN and DDQN show a steady downward trend in terms of loss values, indicating that these two models are continuously optimized and gradually converge. In contrast, DDQN without pruning shows a significant decrease at the beginning but fails to reach the convergence level of other models. DDQN without soft updates shows large fluctuations in loss values and remains unstable even in the later stage of training, indicating that the absence of soft updates makes it difficult for the model to converge effectively.
[0171] Figure 8 shows the stability of the reward values of DQN and DDQN. The reward values of both are close to 25 in the later stage of training, indicating that the model has gradually converged and achieved a good decision-making effect. In contrast, DDQN without pruning and soft updates performs less stably. Especially when soft updates are not used, the reward values fluctuate greatly, indicating a poor decision-making effect and a slow convergence process.
[0172] Since it is not possible to clearly distinguish the performance of DQN and DDQN, further analysis is needed in the comparative experiment.
[0173] (2) Comparative experiment
[0174] MORS aims to solve two problems: minimizing the total production time of remote sensing tasks and achieving resource load balancing among virtual machines and physical machines. Through intelligent scheduling strategies, it reduces the waiting time of remote sensing algorithms in the queue, the data transfer time and execution time on virtual machines, thus ensuring that remote sensing products can be produced and delivered as soon as possible. Secondly, it can dynamically adjust the allocation of remote sensing algorithms on virtual machines to balance the resource load of each machine and avoid the situation where some machines are overloaded while others are idle. To achieve the above two goals, corresponding objective formulas were designed.
[0175] Production Time: Due to the limitations of algorithmic dependencies, the production time of each remote sensing task is the sum of the production times of its subtasks, as shown in the following formula:
[0176]
[0177] Among them, is the total production time of task w i Ta j is the production time of the subtask. Since the algorithms of different remote sensing tasks are usually executed in parallel, the total production time can be expressed as the maximum value of the completion times of each task. To improve production efficiency, the maximum completion time should be minimized as much as possible, as shown in the following formula:
[0178]
[0179] Among them, T total is the total production time, and W is the total task.
[0180] Resource Load: When constructing a multi-objective remote sensing task scheduling algorithm, the usage of resources such as CPU, memory, and storage of physical machines and virtual machines is comprehensively considered. The weighted average standard deviation method is used to sum the standard deviations of each resource dimension weighted by weights, as shown in the following formula:
[0181]
[0182] Among them, w cpu 、w mem and w str are the weights of CPU, memory, and storage respectively, RL pm is the resource load of the physical machine, and RL vm is the resource load of the virtual machine.
[0183] When discussing the remote sensing task scheduling problem, several existing task scheduling algorithms were selected for comparison with the scheduling algorithm (DDQN) proposed in the present invention. These algorithms include the Trade-Off algorithm, the polling algorithm, and the DQN (Deep Q-Network) algorithm. Four groups of comparative simulation experiments were designed, in which the total number of remote sensing tasks was 30, 60, 90, and 120 respectively, and the total number of virtual machine nodes was 15. To ensure the reliability of the experimental results, the Monte Carlo experimental method was used to test each group of experiments, and the number of test rounds was 100.
[0184] From Figure 9It can be seen that as the number of remote sensing products increases, the production time of the DDQN algorithm is always relatively low, especially showing obvious advantages when the number is large. When the number of tasks is 90, the production time gaps of each algorithm are relatively small; while when it is 120, the production time of the Trade-Off algorithm is significantly the highest, and RR and DQN also gradually lag behind DDQN. This indicates that DDQN has more advantages in high-concurrency scenarios and can better handle the scheduling pressure brought by the increase in task load.
[0185] Figure 10 It shows the standard deviation of the physical machine node load under different numbers of tasks. The lower the standard deviation, the more balanced the resource allocation. In all cases of the number of tasks, the load standard deviation of DDQN is generally lower than that of other algorithms, indicating that DDQN can allocate resources more effectively, maintain the balance of the system, and reduce the volatility of resource utilization.
[0186] Figure 11 It shows that the advantage of DDQN in the standard deviation of virtual machine resource load is not particularly obvious. Compared with other algorithms, it only shows a certain degree of improvement. The reason for this situation is that when training the model, the reward weight of the physical machine resource load is set higher than that of the virtual machine, resulting in the model paying more attention to the efficient utilization of physical machine resources and ignoring the optimization of virtual machine load balance to a certain extent.
[0187] Embodiment 4
[0188] Based on the above embodiments, as Figure 12 shown, an embodiment of the present invention provides a multi-objective remote sensing task scheduling system based on DDQN, including:
[0189] A task processing module, configured to input multiple given remote sensing tasks into a preset task processing unit to obtain the algorithm representation of multiple remote sensing tasks and the algorithm value of the remote sensing tasks; wherein, the preset task processing unit includes multiple virtual machines and multiple physical machines, and the multiple virtual machines are respectively arranged on the multiple physical machines.
[0190] An algorithm set module, configured to construct an algorithm set according to the algorithm representations of multiple remote sensing tasks, and obtain the resources required by each algorithm in the algorithm set and the value of the algorithm.
[0191] A scheduling module, configured to input the algorithm set into a preset scheduling model to obtain multiple algorithm subsets.
[0192] A policy network module, configured to allocate the multiple algorithm subsets to multiple nodes respectively according to the policy network and the resources required by each algorithm; wherein, the nodes include multiple virtual machines and multiple physical machines.
[0193] An execution module, configured to perform operations on the algorithm subsets allocated to nodes to obtain results of a given plurality of remote sensing tasks.
[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-objective remote sensing task scheduling method based on DDQN, characterized in that: include: Step 1: inputting a plurality of given remote sensing tasks into a preset task processing unit to obtain algorithm representations of the plurality of remote sensing tasks; wherein the preset task processing unit includes a plurality of virtual machines and a plurality of physical machines, and the plurality of virtual machines are respectively arranged on the plurality of physical machines; Step 2: Construct an algorithm set based on the algorithm representations of multiple remote sensing tasks, and obtain the resources required by each algorithm in the algorithm set and the value of the algorithm; Step 3: Input the algorithm set into the preset scheduling model to obtain multiple algorithm subsets; Step 4: Allocate multiple algorithm subsets to multiple nodes according to the policy network and the resources required by each algorithm; wherein the nodes include multiple virtual machines and multiple physical machines; Step 5: Operate the algorithm subsets assigned to the nodes to obtain the results of the given multiple remote sensing tasks.
2. According to the DDQN-based multi-objective remote sensing task scheduling method of claim 1, it is characterized in that: The step of inputting a plurality of given remote sensing tasks into a preset task processing unit to obtain algorithm representations of the plurality of remote sensing tasks specifically includes: Each remote sensing task is divided into multiple algorithms required by it, and multiple algorithms are sorted in the order of operation to obtain the algorithm representation E of the remote sensing task i ={a i ,a2,a3,...,a n }; where E i is the i-th remote sensing task, a n is the nth algorithm.
3. According to the DDQN-based multi-objective remote sensing task scheduling method of claim 1, it is characterized in that: The value of the algorithm is expressed as follows: V ai =w l ×P ai +w r ×(CT-R ai )+w d ×ID ai ×Dep ai Among them, V ai For algorithm a i The value of w l For algorithm a i The priority weight, P ai For algorithm a i The priority of w r is the weight of the receiving time, CT is the current time, R ai is the receiving time, w d is the weight of the input data, ID ai is the size of the input data, Dep ai For algorithm a i Dependencies.
4. The multi-objective remote sensing task scheduling method based on DDQN according to claim 1 is characterized in that: The resources required for each algorithm in the algorithm set include the algorithm's unique identifier, the remote sensing mission's unique identifier, the algorithm name, the algorithm priority, millions of instructions per second, the size of the input data, the operating system of the production, the CPU configuration, the memory size, the storage size, the GPU requirements, the receiving time and the preset completion time.
5. The multi-objective remote sensing task scheduling method based on DDQN according to claim 1 is characterized in that: The preset scheduling model includes a Markov decision process; The state of the scheduling model includes: defining the state space as state characteristics S according to the input of multiple algorithms and task processing units rp and algorithm features Among them, the feature matrix S rp Including identification to distinguish physical machines and virtual machines, unique identifier of the node, current number of cores of virtual machines and physical machines, current memory size of virtual machines and physical machines, current storage size of virtual machines and physical machines, bandwidth of virtual machines and physical machines, performance of virtual machine processing units and performance of physical machine processing units, percentage of CPU remaining in the total, percentage of memory remaining in the total, percentage of storage remaining in the total, number of products running on virtual machines and number of products running on physical machines, and number of virtual machines on physical machines; Algorithm feature T ai Including CPU configuration, millions of instructions per second, storage size, input data size, memory size, receiving time and preset completion time; The movement of the scheduling model is expressed as follows: action={0,1,2,...,n-1} Among them, action is the action process, n represents the total number of virtual machines; The state transition of the scheduling model is expressed as follows: S rp+1 =U(S rp ,action i ) Among them, S rp+1 is the state characteristic of the next state, and U is the state transition process; The reward function of the scheduling model is expressed as follows: R total =α3·R t +β3·R util_p +γ3·R util_v Among them, R total is the value of the reward function, α3 is the weight coefficient of the reward value of the task production time, R t is the reward value of the task production time, β3 is the weight coefficient of the reward value of the physical machine, R util_p is the reward value of the physical machine, γ3 is the weight coefficient of the reward value of the virtual machine, R util_v is the reward value of the virtual machine, σ cpu_p , σ mem_p and σ str_p are the standard deviations of the remaining CPU, memory, and storage on the physical machine, respectively, and u cpu_p 、u mem_p and u str_p are the average values of CPU, memory and storage remaining on the physical machine, σ cpu_v , σ mem_v and σ str_v are the standard deviations of the remaining CPU, memory, and storage on the virtual machine, respectively, and u cpu_v 、u mem_v and u str_v are the average values of the remaining CPU, memory, and storage on the virtual machine, α1, β1, and γ1 are the weights of the remaining CPU, memory, and storage on the physical machine, α2, β2, and γ2 are the weights of the remaining CPU, memory, and storage on the virtual machine, For algorithm a i The execution time, For algorithm a i The total production time, ε is a number close to 0, is a number close to 0.
6. The multi-objective remote sensing task scheduling method based on DDQN according to claim 2 is characterized in that: The algorithm set is input into a preset scheduling model to obtain multiple algorithm subsets, specifically including: Construct multiple algorithm subsets A according to the same order of algorithms in the algorithm representation of each remote sensing task j ={a 1j ,a 2j ,a 3j ,...,a nj }; Among them, A j is the jth algorithm subset, a nj is the jth algorithm in the algorithm representation of the nth remote sensing task.
7. The multi-objective remote sensing task scheduling method based on DDQN according to claim 5 is characterized in that: The strategy network adopts a multi-layer fully connected structure, and the strategy network includes an input layer, multiple hidden layers and an output layer; The input layer includes state_dim+task_dim neurons for receiving state features S rp And algorithm characteristics T ai ; Among them, task_dim is the algorithm feature dimension, and state_dim is the state feature dimension; The plurality of hidden layers include the same number of neurons; The output layer includes action_dim neurons, where action_dim is the number dimension of the virtual machine.
8. The multi-objective remote sensing task scheduling method based on DDQN according to claim 5 is characterized in that: The policy network is optimized and trained according to the DDQN algorithm, which is used for action selection and value evaluation, and a loss function is constructed according to the value evaluation to update the policy network parameters; The action selection is expressed by the following formula: a'=arg max Q(S rp+1 ,action i ; θ) Among them, a' is the current strategy network Q in S rp+1 The action selected, θ is the current policy network parameter, arg max is the maximum independent variable point set; The value assessment is expressed as follows: y=r rp +γQ target (S rp+1 ,a';θ - ) Among them, y is the target network Q target Calculated Q value, r rp is the immediate reward, γ is the discount factor, θ - is the target network parameter, Q target For the target network; The loss function is expressed as follows: Among them, ω* is the value of the loss function, and N is the batch size of samples randomly selected from the experience pool.
9. The multi-objective remote sensing task scheduling method based on DDQN according to claim 8 is characterized in that: The policy network also sets a pruning threshold during training and soft updates the target network parameters in the DDQN algorithm. The specific process is expressed as follows: τ=percentile(S,p) in, is the updated target network parameter, θ is the current strategy network parameter, is the current parameter of the target network, τ is the pruning threshold, percentile(S,p) is the maximum value of the minimum p% weights in the returned set, % is the modulo operation, S is the set of all weights, and p is the pruning ratio.
10. A multi-objective remote sensing task scheduling system based on DDQN, characterized in that: include: A task processing module, used for inputting a plurality of given remote sensing tasks into a preset task processing unit, and obtaining algorithm representations of the plurality of remote sensing tasks and algorithm values of the remote sensing tasks; wherein the preset task processing unit includes a plurality of virtual machines and a plurality of physical machines, and the plurality of virtual machines are respectively arranged on the plurality of physical machines; The algorithm collection module is used to construct an algorithm collection according to the algorithm representations of multiple remote sensing tasks, and obtain the resources required by each algorithm in the algorithm collection and the value of the algorithm; A scheduling module, used to input the algorithm set into a preset scheduling model to obtain multiple algorithm subsets; A policy network module, used to allocate multiple algorithm subsets to multiple nodes respectively according to the policy network and the resources required by each algorithm; wherein the nodes include multiple virtual machines and multiple physical machines; The execution module is used to operate the algorithm subsets assigned to the nodes to obtain the results of the given multiple remote sensing tasks.
Citation Information
Cited By
Pruning double-Q learning echo state network-based effluent ammonia nitrogen concentration prediction method
CN121051680A