Unloading decision-making method based on task collaboration

By proposing a task collaboration decision-making method based on task collaboration in the agricultural machinery multi-machine collaboration scenario, the Markov decision-making process and DQNDE model algorithm are used to optimize the delay and energy consumption of task offloading, and the delay sensitivity and energy consumption constraints during task offloading in the existing technology are solved, significantly improving the operating efficiency and stability of the system.

CN120075900APending Publication Date: 2025-05-30JIANGSU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510234131.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the multi-machine cooperation scenario of agricultural machinery, it is difficult for the existing technology to effectively optimize the delay sensitivity and energy consumption constraints during task unloading. At the same time, it ignores the coordination between tasks, resulting in a decline in service quality of edge computing nodes and affects system operation efficiency.

Method used

A method of offload decision-making based on task coordination is proposed, through task decomposition, establish task coordination model and calculation model, and construct the final optimization goal, and adopt Markov decision-making process and DQNDE model algorithm to solve the problem. In the multi-machine cooperation scenario of agricultural machinery, this method optimizes the delay and energy consumption of task unloading, and fully considers the coordination between tasks.

Benefits of technology

The task offloading efficiency in the multi-machine cooperative scenario of agricultural machinery has been significantly improved. Through the combination of the global optimization framework and the DQNDE algorithm, the amount of repeated calculations is reduced by more than 30%, the utilization rate of computing resources is improved, and the overall stability of the system is improved by 40% in a dynamic farmland environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005292295080000021
    Figure BDA0005292295080000021
  • Figure BDA0005292295080000031
    Figure BDA0005292295080000031
  • Figure BDA0005292295080000061
    Figure BDA0005292295080000061
Patent Text Reader

Abstract

The invention discloses an unloading decision-making method based on task collaboration. The method comprises the following steps: step 1, task decomposition: decomposing an agricultural machinery collaboration task into a plurality of collaboration sub-tasks; step 2, establishing a model: step 2a, constructing a task collaboration model in combination with the dependency relationship among the collaboration subtasks; step 2b, constructing a calculation model according to the local execution energy consumption, the edge node execution energy consumption, the local execution time delay and the time delay required for transmission to the edge node; step 3, problem construction; step 4, making a Markov decision based on the problem constructed in the step 3; 5, solving the problem constructed in the step 3 by adopting a DQNDE model algorithm according to a Markov decision; through the establishment of the model and the improvement of the algorithm, the task unloading efficiency in the multi-machine cooperation scene of the agricultural machinery is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-machine cooperation in agricultural machinery, and particularly relates to an unloading decision-making method based on task collaboration. Background Art

[0002] In edge computing, computing offloading is a common strategy for optimizing task execution and resource utilization. It transfers computing tasks from terminal devices to edge servers or the cloud for processing, reducing the burden on devices, lowering energy consumption, and improving task execution efficiency. Computing offloading is particularly suitable for scenarios that require a large amount of computing resources or low latency, such as video processing and machine learning inference. Intelligent computing offloading decisions help optimize resource management and task scheduling, enhancing system performance and user experience.

[0003] However, multi-machine collaborative operation in agriculture is a complex multi-stage process. From task allocation to agricultural machinery scheduling and then to collaborative operation, a high degree of collaboration between devices is required at each stage. In this process, the collaboration between tasks and the complexity of resource requirements pose unique challenges to edge computing. For example, in the collaborative operation of a harvester and a grain carrier, it is necessary to reasonably allocate tasks and plan paths, while predicting the grain unloading time, location, and controlling the unloading sequence. These tasks require maintaining communication and data synchronization, and have collaboration in terms of time and space. If this collaboration is ignored during task offloading, it may lead to additional communication overhead due to the lack of collaborative tasks in edge computing nodes and abnormal connection between tasks, affecting the operation efficiency of the entire system.

[0004] Therefore, there is an urgent need for a comprehensive optimization solution that can not only meet the time-delay sensitivity and energy consumption constraints of tasks but also fully consider the collaboration between tasks, thereby improving the performance of the entire collaborative system. Summary of the Invention

[0005] The purpose of the present invention is to provide an unloading decision-making method based on task collaboration for the deficiencies of the prior art.

[0006] The technical solution of the present invention to solve the above problems is as follows: An unloading decision-making method based on task collaboration includes the following steps:

[0007] Step 1, task decomposition: Decompose the agricultural machinery collaboration task into several collaborative subtasks, and the collaborative subtasks include path planning, obstacle detection, and collaborative control.

[0008] Step 2, establish models, including: Step 2a, construct a task collaboration model by combining the dependency relationships between the collaborative subtasks; Step 2b, construct a computing model based on the local execution energy consumption, the edge node execution energy consumption, the local execution time delay, and the time delay required for transmission to the edge node.

[0009] Step 3: Problem construction, constructing the final optimization objective as the weighted sum of latency and energy consumption, which is min λA_T 1:n +(1 - λ)E;

[0010] Step 4: Formulate a Markov decision based on the problem constructed in Step 3;

[0011] Step 5: Use the DQNDE model algorithm according to the Markov decision to solve the problem constructed in Step 3.

[0012] Furthermore, the final optimization objective constructed in Step 3 includes the following constraint conditions:

[0013] Constraint condition 1 stipulates that the task can only be executed locally or offloaded to a certain edge server for operation;

[0014] Constraint condition 2 restricts that the CPU frequency of the user equipment shall not exceed its maximum value;

[0015] Constraint condition 3 restricts that the CPU frequency of the user equipment needs to be selected from a discrete set;

[0016] Constraint condition 4 requires that the upload power of the user equipment also belongs to a discrete set;

[0017] Constraint condition 5 stipulates that the value range of the weight factor λ of latency and energy consumption is between 0 and 1;

[0018] Constraint condition 6 requires that the completion time of the last task of the application shall not exceed the maximum tolerable latency.

[0019] Furthermore, the specific method of Step 4 is as follows: MDP can be defined as a 4-tuple <S, A, T, R>, where S is the state space, A is the action space, T is the state transition function, and R is the reward function;

[0020] The state space is represented as S = {D i , C i , Q j,i , P up}: The definition of the state space needs to comprehensively consider the environment of the edge computing network and the relevance of tasks. At time t, the state of the entire system includes the amount of data D i generated by computing task i, the number of CPU cycles required to process each bit of data task C i , the data volume Q j,i between the predecessor task j and task i, and the communication gain P up of the uplink channel;

[0021] The action space is represented as A = {0, 1}: which respectively represent executing the current task to be decided on the agricultural machinery mobile device and offloading the current task to be decided to the edge server for execution;

[0022] The state transition function is expressed as: T = (s t , a t ): Its return is the next state after executing action a t under state s t at time t;

[0023] The reward function R is expressed as: r t = -λAT 1:n - (1 - λ)E: The reward function should be related to the objective to be optimized. It is necessary to minimize both the task delay and the energy consumption of the user equipment. Therefore, the reward function setting should consider both delay and energy consumption at the same time.

[0024] Furthermore, the specific method of step 5 is: The DQNDE model algorithm combines the prioritized experience replay pool and the uniform probability random experience replay pool;

[0025] The double experience replay balances the learning of different experiences by adjusting the size of the first-in-first-out experience replay pool and the sampling ratios of the two experience replay pools respectively; All experience data is first stored in the uniform probability random experience replay pool, and then the data with higher value is selected from it and stored in the prioritized experience replay pool, where the high-value experiences account for the top 70% of the reward values sorted from large to small; When sampling, an experience replay pool is selected with a certain probability; The probability of sampling from the prioritized experience replay pool is 60%, and the probability of sampling from the uniform probability random experience replay pool is 40%; The uniform probability random experience replay pool adopts a simple sequential storage method, and samples are randomly drawn according to the batch size during sampling;

[0026] The network structure of DQNDE includes a main network and a target network; The main network is responsible for using the policy π as the decision criterion, calculating the Q values of each action through the target network, and then using the ε-greedy method to generate a decision action a t with probability ε, as follows;

[0027]

[0028] The network parameters of the main network are initially set to θ, and the structure of the target network is exactly the same as that of the main network. The parameters of the target network are initially set to θ - , and every once in a while, the parameters of the main network are assigned to the target network to synchronize their parameters; The parameters of the neural network are updated by minimizing the loss function. The loss function uses the mean square error function, which is defined as follows:

[0029] J(θ) = E[(y - Q(s t , a t ; θ)) 2

[0030] ​Among them, y is the target value. To reduce the overestimation bias, a double estimator is used to decouple the action selection and action evaluation of the target Q function, reducing the correlation between action selection and the calculation of the target value y. The update of the Q value can be expressed as:

[0031] y = r t + γQ(s t+1 , argmaxQ(s t+1 , a t+1 ; θ); θ - )

[0032] Among them, γ represents the discount factor. Between the last hidden layer and the output layer, the value function V(s) and the action advantage function A(s, a) in Dueling DQN are added. Among them, V(s) only depends on the state s and is used to evaluate the overall value of the current state; while the action advantage function A(s, a) is related to the state s and the action a and is used to measure the advantage of choosing a certain action compared to other actions in a specific state. The optimal action value function is obtained through the linear combination of the value function V(s) and the action advantage function A(s, a), and it is expressed as:

[0033]

[0034] The present invention has beneficial effects:

[0035] The present invention provides an offloading decision-making method based on task collaboration. Through the establishment of a model and the improvement of an algorithm, a significant improvement in the task offloading efficiency in the scenario of multi-machine collaboration of agricultural machinery is achieved. The task collaboration model based on the topological structure deeply analyzes the data transmission order, computational dependency relationship, and spatio-temporal collaboration constraints among tasks, constructs an objective of minimizing the weighted total cost by quantifying key parameters such as delay and energy consumption, and provides a global optimization framework for complex collaboration scenarios. The proposed DQNDE algorithm innovatively adopts a double experience pool mechanism, accelerates the learning of key samples through the prioritized experience pool, enhances the exploration ability through the equiprobable random experience pool, effectively alleviates the overfitting problem while improving the training efficiency. By introducing the linear combination optimization of the action advantage function and the value function, the algorithm reduces the repeated calculation amount by more than 30% in the action decision-making stage, significantly improves the utilization rate of computing resources, especially in the dynamic farmland environment, and the overall stability of the system is improved by 40%. Specific implementation manners

[0036] An offloading decision-making method based on task collaboration specifically includes the following steps:

[0037] Step 1, task decomposition: Decompose the agricultural machinery collaboration task into several collaborative subtasks, and the collaborative subtasks include path planning, obstacle detection, and collaborative control;

[0038] Step 2, establish a model, including:

[0039] (a) Task collaboration model

[0040] Each collaborative subtask can be represented as a mutual dependency relationship within a certain time period. We can represent it using a directed acyclic graph G=(V, E), where V={1, 2, …, n} represents the set of collaborative subtasks, and n is the number of subtasks. E={e i,j |i, j∈N, i<j} represents the set of directed dependency edges between subtasks. For a directed edge e i,j ∈E, subtask i is the direct predecessor task of subtask j, and subtask j is the direct successor task of subtask i.

[0041] (b) Computation model

[0042] Considering the local execution energy consumption and the edge node execution energy consumption, the energy consumption required to execute task i is expressed as: where x m,i is used to describe the location of task execution. When x m,i =0, it means the task is executed locally. When x m,i =1, it means the task is offloaded to edge server m for execution. represents the energy consumption required to execute task i locally. represents the energy consumption required to execute task i on the edge server. Therefore, the total energy consumption required for all agricultural machinery equipment to execute all tasks is

[0043] Combining the local execution delay and the delay required to transmit to the edge node, the delay required to execute task i is expressed as: where represents the delay required to execute task i locally. represents the delay required for task i on the edge server. Therefore, the total delay generated by all tasks is defined as the execution completion time AT of the last task in the task set 1:n .

[0044] Step 3: Problem construction

[0045] Construct the final optimization objective of the weighted sum of delay and energy consumption as minλAT 1:n +(1 - λ)E, including the following constraints:

[0046] Constraint 1 stipulates that a task can only be executed locally or offloaded to a certain edge server for operation.

[0047] Constraint 2 restricts that the CPU frequency of the user equipment shall not exceed its maximum value.

[0048] Constraint 3 restricts the CPU frequency of the user equipment to be selected from a discrete set;

[0049] Constraint 4 requires that the upload power of the user equipment also belongs to a discrete set;

[0050] Constraint 5 stipulates that the weight factor λ of delay and energy consumption ranges from 0 to 1;

[0051] Constraint 6 requires that the completion time of the last task of the application shall not exceed the maximum tolerable delay.

[0052] Step 4: Formulate a Markov decision. Due to the variability of the state of the agricultural machinery mobile system and the correlation between tasks, a Markov decision is formulated based on the problem constructed in Step 3. Specifically, the MDP can be defined as a 4-tuple <S, A, T, R>, where S is the state space, A is the action space, T is the state transition function, and R is the reward function;

[0053] The state space is represented as S = {D i , C i , Q j,i , P up}: The definition of the state space needs to comprehensively consider the environment of the edge computing network and the correlation of tasks. At time t, the state of the entire system includes the amount of data D generated by computing task i i , the number of CPU cycles required to process each bit of data task C i , the data volume Q between the predecessor task j and task i j,i , and the communication gain P of the uplink channel up ;

[0054] The action space is represented as A = {0, 1}: which respectively represent executing the current task to be decided on the agricultural machinery mobile device and offloading the current task to be decided to the edge server for execution.

[0055] The state transition function is represented as: T = (s t , a t ): It returns the next state after executing action a t in the state s t at time t.

[0056] The reward function R is represented as: r t = -λAT 1:n - (1 - λ)E: The reward function should be related to the objective to be optimized. It is necessary to simultaneously minimize the delay of the task and the energy consumption of the user equipment. Therefore, the reward function setting should consider both delay and energy consumption.

[0057] Step 5: Use the DQNDE model algorithm according to the Markov decision to solve the problem constructed in Step 3.

[0058] Experience replay is one of the crucial techniques in the Deep Q-Network. When the agent interacts with the environment and collects experiences, these experiences are stored in the experience replay buffer. In each training iteration, DQN randomly samples a batch of data from the experience replay buffer for training the Q-network. This training method can make the training data more representative, reduce the correlation during training, and thus improve the stability and convergence speed of the algorithm.

[0059] Prioritized experience replay is an improved way of experience replay. It stores the agent's experiences as sequences with priorities and samples these experiences when needed to guide model training. By calculating the priority of each experience, prioritized experience replay plays a key role in balancing important and unimportant experiences. However, the frequent use of high-priority samples may lead to overfitting, making the model overly dependent on certain specific samples, thus introducing data bias. In addition, the frequent sampling of high-priority samples may suppress the exploration ability of the model, resulting in a decrease in sampling diversity. Moreover, priority sampling also increases the additional computational cost.

[0060] In contrast, the uniform random experience pool is a full-retention experience pool method that stores all the experiences generated from the interaction with the environment. In this way, the experience states in the experience pool are very rich, which helps the model better adapt to complex environmental changes. However, as the training progresses, the size of the experience pool will gradually increase, bringing problems such as high storage costs and decreased sampling efficiency. Especially when the size of the experience pool is large, the time complexity of randomly sampling samples increases significantly. In addition, as the experience pool expands, it may contain a large amount of redundant or outdated data, which has limited optimization effect on the current decision-making and may even introduce noise, affecting the training effect of the model.

[0061] The problems mentioned above are alleviated by combining the prioritized experience pool and the uniform random experience pool (Deep Q Network with Double Experiencepool, DQNDE).

[0062] Double experience pool replay balances the learning of different experiences by adjusting the size of the first-in-first-out experience pool and the sampling ratios of the two experience pools respectively, improves the sample utilization rate, and enhances the algorithm performance. The experiences in the uniform random experience pool D full can ensure that the experience pool has a wide state coverage, which helps to learn the states and actions that may appear in the future. While the prioritized experience pool D prioritizedThe experience can ensure better learning effects of the state and actions in the experience pool under the current policy because these experiences are closer to the current state distribution. By adjusting the experience pool size and sampling ratio, the state differences of samples for each sampling are changed. By combining the prioritized experience replay with the uniform random experience replay, the collected experience data can be utilized more effectively. All experience data are first stored in the uniform random experience replay, and then the data with higher values are screened out and stored in the prioritized experience replay, where the high-value experiences account for the top 70% sorted from largest to smallest in terms of the reward value. When sampling, the experience pool is selected with a certain probability. There is a 60% probability of sampling from the prioritized experience replay and a 40% probability of sampling from the uniform random experience replay. The uniform random experience replay adopts a simple sequential storage method, and samples are randomly drawn according to the batch size during sampling. Introducing the uniform random experience replay can effectively reduce the model's dependence on specific samples, simplify the sampling process, and reduce the computational overhead brought by sampling from the prioritized experience replay.

[0063] The network structure of DQNDE includes a main network and a target network. The main network is responsible for using the policy π as the decision criterion, calculating the Q-values of each action through the target network, and then adopting the ε-greedy method to generate a decision action a with a probability of ε t , as follows;

[0064]

[0065] The network parameters of the main network are initially set as θ. The structure of the target network is exactly the same as that of the main network, aiming to provide a stable update target for the parameters of the main network. The parameters of the target network are initially set as θ - , and every once in a while, the parameters of the main network are assigned to the target network to synchronize the parameters of the two. The parameters of the neural network are updated by minimizing the loss function. The loss function uses the mean squared error function, which is defined as follows:

[0066] J(θ) = E[y - Q(s t , a t ; θ)) 2

[0067] Among them, y is the target value. To reduce the overestimation bias, a double estimator is used to decouple the action selection and action evaluation of the target Q function, reducing the correlation between the action selection and the calculation of the target value y. The update of the Q-value can be expressed as:

[0068] y = r t + γQ(s t+1 , argmaxQ(s t+1 , a t+1 ; θ); θ - )

[0069] ​Among them, γ represents the discount factor. Between the last hidden layer and the output layer, the value function V(s) and the action advantage function A(s,a) in Dueling DQN are added. Among them, V(s) only depends on the state s and is used to evaluate the overall value of the current state; while the action advantage function V(s,a) is related to the state s and the action a and is used to measure the advantage of choosing a certain action over other actions in a specific state. The optimal action value function is obtained through the linear combination of the value function V(s) and the action advantage function A(s,a), which is expressed as:

[0070]

[0071] The advantage of such improvement is that it can improve the efficiency of the agent in learning the state value function. Each time it is updated, the state value function will be optimized. This optimization not only affects the value evaluation of the current state and action, but also affects the Q-values of other actions. This method can better utilize the existing data and enable the agent to adapt to the changes in the environment faster. At the same time, this method reduces repeated calculations, makes full use of the updated value function, and improves the stability and speed of learning.

[0072] We used four groups of models in the experiment for comparison. The basic algorithm of the four groups of models is the DQN algorithm, namely the DQN model using an equiprobable random experience pool, the DQN_PER model using priority sampling, the DQN_DOUBLE model using two equiprobable random pools, and the DQNDE model we proposed. The experiment was set within 500 rounds, and each model was experimented 10 times. The final total rewards were compared, and finally the average value of the 10 times was taken for comparison. The comparison results showed that the total cost was reduced by 50.35%, 35%, and 31% compared with all comparison methods. The classification performance of our method was clearly verified to be superior to other algorithms through the comparison results.

[0073] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention in any other form. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. An offloading decision method based on task collaboration, characterized by: The following steps are involved: Step 1, task decomposition: decompose the agricultural machinery cooperation task into several cooperative subtasks; Step 2, establishing a model, including: step 2a, building a task coordination model based on the dependency relationship between each coordinated subtask; step 2b, building a calculation model based on local execution energy consumption, edge node execution energy consumption, local execution delay and transmission delay required to the edge node; Step 3: Problem construction, construct the final optimization target as minλAT based on the weighted sum of latency and energy consumption 1:n +(1-λ)E; Step 4: Make a Markov decision based on the problem constructed in step 3; Step 5: Use the DQNDE model algorithm based on Markov decision making to solve the problem constructed in step 3.

2. The offloading decision method based on task collaboration as claimed in claim 1, characterized in that: The final optimization objective constructed in step 3 contains the following constraints: Constraint 1 stipulates that the task can only be executed locally or offloaded to an edge server; Constraint 2 limits the CPU frequency of the user device to not exceed its maximum value; Constraint 3 restricts the user device CPU frequency to be selected from a discrete set; Constraint 4 requires that the upload power of the user equipment also belongs to a discrete set; Constraint 5 stipulates that the weight factor λ of delay and energy consumption ranges from 0 to 1; Constraint 6 requires that the completion time of the last task of the application must not exceed the maximum tolerable delay.

3. The offloading decision method based on task collaboration as claimed in claim 1, characterized in that: The specific method of step 4 is: MDP is defined as a 4-tuple<S,A,T,R> , where S is the state space, A is the action space, T is the state transition function, and R is the reward function; The state space is represented as S = {D i ,C i ,Q j,i ,P up }; The state of the entire system at time t includes the amount of data generated by computing task i, D i , the number of CPU cycles required to process each bit of data task C i , the amount of data Q between predecessor task j and task i j,i , and the communication gain P of the uplink channel up ; The action space is represented as A = {0, 1}: A represents executing the current decision-making task on the mobile agricultural machine and offloading the current decision-making task to the edge server for execution; The state transition function is expressed as: T = (s t ,a t ): Its return is the state s at time t t Next, perform action a t The next state after that; The reward function R is expressed as: t =-λAT 1:n -(1-λ)E.

4. The offloading decision method based on task collaboration as claimed in claim 1, characterized in that: The specific method of step 5 is as follows: the DQNDE model algorithm combines the priority experience pool and the equal probability random experience pool; Dual experience pool playback balances the learning of different experiences by adjusting the size of the first-in-first-out experience pool and the sampling ratio of the two experience pools respectively; all experience data are first stored in the equal probability random experience pool, and then the data with higher value are screened out and stored in the priority experience pool, where high-value experience accounts for the first 70% of the reward value sorted from large to small; when sampling, the experience pool is selected with a certain probability; 60% of the probability is sampled from the priority experience pool, and 40% of the probability is sampled from the equal probability random experience pool; The equal probability random experience pool adopts a simple sequential storage method, and samples are randomly selected according to the batch size during sampling; The network structure of DQNDE consists of a main network and a target network. The main network is responsible for using the strategy π as the decision criterion, calculating the Q value of each action through the target network, and then using the ε-greedy method to generate a decision action a with probability ε t , as shown below; The network parameters of the main network are initially set to θ. The structure of the target network is exactly the same as the main network, and the parameters of the target network are initially set to θ. - , at regular intervals, the parameters of the main network are assigned to the target network to synchronize the parameters of the two; the parameters of the neural network are updated by minimizing the loss function, which uses the mean square error function and is defined as follows: J(θ)=E[yQ(s t ,a t ;i)) 2 ] Where y is the target value. In order to reduce the overestimation bias, a dual estimator is used to decouple the action selection and action evaluation of the target Q function, reducing the correlation between action selection and target value y calculation. The update of the Q value can be expressed as: y=r t +γQ(s t+1 ,argmaxQ(s t+1 ,a t+1 ;i);i - ) Where γ represents the discount factor. Between the last hidden layer and the output layer, the value function V(s) and the action advantage function A(s,a) in Dueling DQN are added. V(s) depends only on the state s and is used to evaluate the overall value of the current state. The action advantage function A(s,a) is related to the state s and the action a, and is used to measure the advantage of choosing an action over other actions in a specific state. The optimal action value function is obtained by a linear combination of the value function V(s) and the action advantage function A(s,a), which is expressed as:

Citation Information

Cited By

  • Aircraft fleet guarantee scheduling method based on discrete time Markov decision

    CN120525293A