Collaborative operation optimization method for unmanned aerial vehicle and robot
By building an inference acceleration framework and communication model for drones and robots, combining differential evolution and deep deterministic strategy gradient algorithm to optimize task offloading, the problems of uneven resource allocation and improper task scheduling in collaborative operations between multiple robots and drones are solved, and system efficiency improvement and efficient resource utilization are achieved.
Patent Information
- Application Number
- CN202510878472.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-27
AI Technical Summary
In the scenario where multiple robots and drones work together, the existing technology has problems such as static allocation of computing resources, uneven resource allocation, improper task scheduling, increased delay and low resource utilization, resulting in the system coordination performance not reaching the optimal level.
Build an inference acceleration framework for drones and robots, establish a communication model, optimize task offloading strategies based on delay and energy consumption calculation models, use differential evolution algorithms and deep deterministic strategy gradient algorithms to optimize task division and drone flight trajectory, and combine communication and perception integrated technology to improve system efficiency.
By dynamically adjusting the location and task allocation of drones, optimizing the inference performance of multiple robots, reducing task processing delays, improving resource utilization efficiency, enhancing system collaboration capabilities and data perception and analysis capabilities, and reducing hardware costs.
Smart Images

Figure CN120416931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and particularly to an optimization method for collaborative operation of unmanned aerial vehicles and robots. Background Art
[0002] In the scenario of collaborative operation of multiple robots and unmanned aerial vehicles, there are many deficiencies in computing resource allocation, task scheduling, and system collaborative efficiency.
[0003] First of all, most of the existing inference acceleration strategies focus on the optimization of single tasks or the allocation of local resources, ignoring the dynamic and global allocation of resources between multiple robots and unmanned aerial vehicles. As a result, the robot system cannot flexibly adapt to the workload changes in complex and multi-task environments, and the traditional methods are too static in computing resource allocation and cannot effectively respond to real-time task requirements, resulting in the computing power not being able to meet the load requirements of the robots in time, thereby reducing the task execution efficiency.
[0004] Secondly, in the collaborative operation system of multiple robots and unmanned aerial vehicles, each intelligent agent (including robots and unmanned aerial vehicles) has different computing capabilities, communication capabilities, and energy states. However, most of the existing resource scheduling algorithms cannot consider multiple factors such as inference calculation, communication, and energy consumption at the same time, resulting in uneven resource allocation. There are situations where some robots cannot complete tasks smoothly due to insufficient resources, which seriously affects the collaborative efficiency of the overall system.
[0005] Thirdly, there is a problem of improper task scheduling when multiple robots work together, which often leads to an increase in delay. Especially in scenarios that require a large amount of data transmission and real-time feedback, the delay problem is particularly prominent. It not only affects the efficiency of task completion but also may cause delays in system response, unable to meet the requirements of efficient work.
[0006] In addition, the optimization of inference performance in the existing technology usually stays at the level of a single intelligent agent, lacking global optimization of the entire system. In the collaborative operation of multiple robots and unmanned aerial vehicles, how to design a reasonable inference acceleration mechanism to ensure the coordination between each intelligent agent and the balance of task loads is a problem that the current technology has not fully solved.
[0007] Finally, the traditional resource scheduling strategies have poor flexibility and adaptability in dealing with complex task environments, and resources are often not fully utilized, resulting in the overall system efficiency not reaching the best level. Especially in scenarios with high loads and complex computing tasks, the optimization of resource configuration is insufficient, which affects the collaborative efficiency between robots and unmanned aerial vehicles and cannot maximize the resource utilization rate of the system. Summary of the Invention
[0008] The object of the present invention is to provide a collaborative operation optimization method for drones and robots, which can make full use of the high mobility and flexibility of drones, dynamically adjust the positions and task assignments of drones, so as to optimize the inference performance of multiple robots, improve resource utilization efficiency and relieve the computing bottleneck.
[0009] To solve the above technical problems, the present invention provides the following technical solutions: A collaborative operation optimization method for drones and robots, comprising the following steps: S1. Construct an inference acceleration framework including drones and robots; S2. Establish a first communication model between the drone and multiple robots, and establish a second communication model between the robots; S3. Obtain the total delay and total energy consumption required for the inference calculation of the DNN task based on the delay calculation model and the energy consumption calculation model; S4. Construct a joint optimization model, and minimize the total delay of the DNN task inference calculation based on the constraint conditions; And, S5. Complete the optimization of the DNN task partitioning and offloading strategy based on the differential evolution algorithm, and complete the drone flight trajectory planning based on the deep deterministic policy gradient algorithm.
[0010] Compared with the prior art, the present invention has the following technical effects: The collaborative operation optimization method for drones and robots of the present invention can make full use of the high mobility and flexibility of drones, dynamically adjust the positions and task assignments of drones, so as to optimize the inference calculation performance of multiple robots.
[0011] Specifically, the present invention searches for the optimal solution in the DNN task partitioning and offloading strategy based on the DE algorithm to efficiently optimize the partitioning of the DNN task, so that the task can be optimally offloaded between multiple vehicle robots and drones, thereby reducing the task processing delay, optimizing the utilization of computing resources. At the same time, by combining the DDPG algorithm, deep reinforcement learning and multi-agent cooperation are combined, and the system efficiency is maximized by the agent independently optimizing the decision-making strategy. It can construct a global resource perception model on the basis of real-time collecting information such as the inference requirements, communication status and drone energy status of multiple robots, and enhance the cooperation ability of multi-agents through an improved attention mechanism, so that the drone can adaptively adjust the inference task offloading and communication resource allocation in a heterogeneous resource environment, further improving the inference efficiency.
[0012] In addition, the present invention also introduces the communication and sensing integration technology, which can not only improve the spectrum utilization efficiency, reduce the hardware cost, but also enhance the data sensing and analysis capabilities of multiple robots in a complex working environment, thereby further improving the overall efficiency of the system. Description of the Drawings
[0013] Figure 1 is the step flow chart of the collaborative operation optimization method for the unmanned aerial vehicle and the robot in the present invention; Figure 2 is the structural schematic diagram of the inference acceleration framework in the present invention; Figure 3 is the change curve of the total reward with the number of training rounds in the present invention; Figure 4 is the flight path of the unmanned aerial vehicle obtained after planning by the DDPG algorithm in the present invention; Figure 5 are the total delay, total energy consumption and task completion rate processed by four algorithms under different numbers of robots and unmanned aerial vehicles in the present invention; Figure 6 are the total delay, total energy consumption and task completion rate processed by four algorithms under different communication network bandwidths, different numbers of robots and unmanned aerial vehicles in the present invention; Figure 7 are the total delay, total energy consumption and task completion rate processed by four algorithms under different computing capabilities of the unmanned aerial vehicle in the present invention. Detailed implementation manners
[0014] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described below in conjunction with the accompanying drawings. Embodiment 1:
[0015] As Figure 1 shown, this embodiment provides a collaborative operation optimization method for an unmanned aerial vehicle and a robot, which includes the following steps: S1. As Figure 2 shown, construct an inference acceleration framework including an unmanned aerial vehicle and a robot, wherein the unmanned aerial vehicle serves as a mobile edge computing node and carries computing and communication resources to provide inference acceleration services for the robot. The robot obtains the position information of the other robots at the current moment in real time through sensors, and based on a communication module (such as a 4G or 5G transmission module mounted on the robot), establishes a communication network between different robots; And in the inference acceleration framework, each robot can cooperate with the unmanned aerial vehicle, and can unload the DNN tasks of the robot layer by layer to the unmanned aerial vehicle for inference calculation, thereby improving the collaborative efficiency of the multi-robot system; There can be multiple unmanned aerial vehicles and robots, and each robot can move. For example, in this embodiment, each robot can be a vehicle-type robot (Mobile Robots with Vehicle Configurations), and an edge server is mounted on each unmanned aerial vehicle; S2. Establish the first communication model between the drone and multiple robots, and establish the second communication model between the robots; Among them, the first communication model includes one or more of a path loss calculation model, a signal-to-interference-plus-noise ratio (SINR) calculation model, and a data transmission rate calculation model, providing a comprehensive communication performance evaluation for task offloading. By dynamically adjusting communication channel parameters and task offloading strategies, it ensures efficient communication and data transmission between the robots and the drone, and reduces task processing latency; Among them, the path loss calculation model is as follows: ;
[0016] Among them, L path is the path loss; η is the path loss exponent, and its value range is 2 - 4; d is the distance (unit: m) between the signal transmitter and the signal receiver; d 0 is the reference distance (usually taken as 1 m); The signal-to-interference-plus-noise ratio calculation model is as follows: ;
[0017] Among them, P signal is the signal transmission power of the communication signal; P noise is the noise power; P interference is the power of the interference signal, and thus the signal quality can be evaluated through the above signal-to-interference-plus-noise ratio calculation model; The data transmission rate calculation model is as follows: ;
[0018] Among them, R is the data transmission rate, with the unit of bits per second (bps); B is the communication network bandwidth when the drone and the robots communicate; in this embodiment, the communication network can be a 4G or 5G network; Thus, a comprehensive communication performance evaluation for task offloading can be provided through the above first communication model. By dynamically adjusting communication channel parameters such as path loss, signal-to-interference-plus-noise ratio, and data transmission rate, it ensures efficient communication and data transmission between multiple robots and the drone, and reduces task processing latency; The second communication model is as follows: ;
[0019] Among them, R is the data transmission rate, with the unit of bits per second (bps); B 0 is the communication network bandwidth when robots communicate with each other. In this embodiment, the communication network can be a 4G or 5G network; P signal is the signal transmission power of the communication signal; N0 is the noise power spectral density; is the total power of the interference signal; thus, in this embodiment, the bandwidth of the communication network and the interference management strategy can be adjusted in real time according to the calculation result of the above second communication model to ensure smooth communication between robots; S3. Based on the delay calculation model and the energy consumption calculation model, respectively obtain the total delay required for the DNN (i.e., deep neural network) task inference calculation to be completed and the total energy consumption generated. Among them, the delay calculation model is as follows: ;
[0020] Among them, is the total delay required for the DNN task I i (t) inference calculation to be completed; is the delay of robot i during the inference calculation of the 1st to i (t) of the DNN task I l layer; is the delay generated when robot i offloads the DNN task I i (t) to drone j within time slot t; is the delay generated when drone j performs the inference calculation of the i (t) of the DNN task I received from robot i for the i (t) from the l +1st layer to the L layer; The energy consumption calculation model is as follows: ;
[0021] Among them, is the total energy consumption generated when the DNN task I i (t) inference calculation is completed; is the energy consumption generated by robot i during the inference calculation of the 1st to i (t) of the DNN task I within the t-th time slot; l is the energy consumption generated when robot i offloads the DNN task I i (t) to drone j within the t-th time slot; (t) to drone j within the t-th time slot; When the robot i offloads the DNN task I i (t) to the UAV j, the UAV j performs the DNN task I i (t) The energy consumption generated during the inference calculation from the l +1-th layer to the L -th layer; is the flight energy consumption generated by the UAV j in the t-th time slot; Specifically, the derivation process of the above delay calculation model is as follows: Assume that the total number of layers of the DNN task I i (t) generated by the robot i in the t-th time slot is L, and the DNN task I i (t) from the 1st layer to the l -th layer is locally calculated on the robot i. Thus, the processing delay during the inference calculation of the 1st layer to the i (t) from the 1st layer to the l -th layer of the DNN task I on the robot i can be obtained according to the following formula: ;
[0022] where, represents the amount of floating-point calculations required for each layer during the calculation of the 1st layer to the i (t) from the 1st layer to the l -th layer; f i (t) represents the amount of floating-point calculations provided by the robot i in the t-th time slot; Further, after the robot i finishes calculating the first i layers (i.e., the 1st layer to the l -th layer) of the DNN task I l (t), it can offload the DNN task I i (t) to the UAV j carrying the edge server and upload the calculation result output by the l -th layer to the UAV j through the uplink to continue executing the DNN task I i (t) from the l +1-th layer to the L -th layer of inference calculation. Therefore, the delay i generated when the robot i offloads the DNN task I to the UAV j in the time slot t is: ;
[0023] where, d l represents the size of the calculation result data output by the robot i after completing the inference calculation of the 1st layer to the i (t) from the 1st layer to the l -th layer; ri,j (t) represents the data transmission rate between robot i and drone j; After the drone j receives the DNN task I offloaded by the robot i i (t), it can execute the DNN task I through the edge server i (t) the l +1st layer to the L layer of inference calculation. Therefore, the drone j performs the DNN task I i (t) the l +1st layer to the L layer of inference calculation generated latency is: ;
[0024] Among them, f i,j (t) represents the computing resources allocated by the drone j for the i (t) the l +1st layer to the L layer of inference calculation during the t-th time slot, and there is , is the maximum computing resource of the drone j; Therefore, the total latency required for the DNN task I i (t) to complete the inference calculation is: ;
[0025] The derivation process of the above energy consumption calculation model is as follows: Since the battery energy carried by the drone and the robot itself is relatively limited, therefore, in this embodiment, the following several types of energy consumption are mainly considered: During the t-th time slot, the energy consumption generated by the robot i during the inference calculation of the 1st layer to the i (t) l layer , and there is: ;
[0026] Among them, represents the computing power of the robot i, and this computing power can be measured by the floating-point operation ability (FLOPS, Floating Point Operations Per Second). The higher the FLOPS value, the stronger the computing power; After the robot i offloads the DNN task I i (t) to the drone j, the drone j performs the DNN task I i (t) the lThe energy consumption generated during the inference calculations from layer +1 to layer L , and there is: ;
[0027] Among them, represents the computing power of drone j. Similarly, this computing power can also be measured by the floating-point operation ability FLOPS; During the t-th time slot, robot i offloads the DNN task I i (t) to drone j, and the energy consumption generated is , and there is: ;
[0028] Among them, P i is the signal transmission power when robot i offloads the task; In addition, the flight energy consumption generated by drone j during the t-th time slot is , and there is: ;
[0029] Among them, represents the weight of drone j; v j (t) is the flight speed of drone j, and there is , , are the positions of drone j at time k and time k - 1 respectively, which can be represented by the three-dimensional spatial coordinates of drone j, and δ represents the length of the time slot; S4. Construct a joint optimization model and minimize the total delay of the DNN task inference calculation based on the constraint conditions; among them, the joint optimization model is as follows: ;
[0030] Furthermore, the constraint conditions include one or more of the constraint conditions (a)-(k): ;
[0031] Among them, the constraint condition (a) is used to ensure that in each time slot, all robots can only request services from the same drone. Among them, a i,s (t) is a binary variable, indicating whether the edge server e carried on the drone decides to execute the DNN task i requested by the robot N represents the set of robots; ES represents the set of edge servers carried on the drone; T represents the set of time slots; Constraint (b) is used to ensure that the robot i cooperates with only one edge server e carried by a UAV for DNN task inference in the t-th time slot to ensure reasonable task division; where U represents the number of UAVs; Constraint (c) is used to ensure that the total amount of computing resources allocated to UAV j when processing task n does not exceed its maximum computing resources to avoid computing resource overload; a i,j (t) is a binary variable indicating whether UAV j chooses to provide services for robot i in the t-th time slot; Constraint (d) is used to ensure that the amount of computing resources allocated by the edge server carried on each UAV to each task m does not exceed its maximum computing resources, that is, the computing power of each UAV is also subject to corresponding constraints to ensure efficient task allocation; where f i,m (t) represents the computing power of robot i for task m in the t-th time slot; a i,m (t) represents the state variable of robot i processing task m; represents the maximum computing resources of the edge server carried on the UAV that can be allocated for task m; Constraints (e) and (f) are used to ensure that the UAV always flies within the preset flight area, avoiding the UAV exceeding the specified service area, thereby ensuring the UAV service coverage; 、 respectively represent the abscissa and ordinate of the projection of UAV j on the horizontal plane in the t-th time slot; X and Y respectively represent the length and width of the projection of the preset flight area on the horizontal plane; u represents the set of UAVs; Constraint (g) is used to ensure that the flight speed v of the UAV in any time slot j (t) does not exceed its maximum speed limit v max , thus ensuring flight safety; Constraint (h) is used to ensure that the distance between any two UAVs j and k at the same moment is not less than the minimum position interval loc min , thus preventing collisions between UAVs and avoiding mutual interference; 、 are the positions of UAVs j and k at the same moment respectively; Constraint (i) is used to ensure that the task division point of the DNN task is located between [0, L , where 0 means that the DNN task is completely offloaded to the UAV for inference calculation, L means that the DNN task is completely performed on the robot for inference calculation, to ensure the rationality and flexibility of task division, so that the task offloading strategy can be dynamically adjusted according to actual needs; Constraint (j) is used to ensure that the battery power of the UAV is greater than zero at the end of each time slot, avoiding mission interruption due to insufficient power. Among them, is the battery power consumption of UAV j in the t-th time slot; Constraint (k) is used to ensure that the total inference calculation delay of the DNN task does not exceed the maximum allowable tolerance delay, so as to ensure that the processing of each task does not time out and further improve the system performance and user experience. Among them, is the maximum tolerance delay for the robot i to complete all DNN tasks I i (t) inference calculation; And, S5. Optimize the DNN task partitioning and offloading strategy based on the differential evolution (DE) algorithm, and, complete the UAV flight trajectory planning based on the deep deterministic policy gradient (DDPG) algorithm; Among them, optimizing the DNN task partitioning and offloading strategy based on the differential evolution (DE) algorithm includes the following steps: S511. Randomly initialize S individuals to obtain an initial population pop = {V1, V2,..., V S} and define the chromosome encoding format of individual V S as V S = [A * , L * , F * ; Among them, A * = [a1(t), a2(t),..., a i (t),..., a U (t)], representing the task selection execution status vector of the UAV in the t-th time slot, including all DNN tasks generated by the robot, and a i (t) represents the flag indicating whether the UAV executes the i-th task (i.e., one of the above all tasks) in the t-th time slot, taking values of 1 or 0, and L * = l (t)], representing the partitioning point of the DNN task, that is, the inference calculation tasks from the first layer to the l -th layer are executed on the robot, and the inference calculation tasks from the l +1-th layer to the L -th layer are executed on the UAV, and F * = [f1(t), f2(t),..., f i (t),..., f U (t)]; representing the task function vector executed by the UAV in the t-th time slot, including the floating-point calculation amount assigned by the UAV to each task; S512. Generate mutant individuals based on the DE / best / 1 / bin mutation strategy and obtain an intermediate population; Specifically, in this step, mutant individuals are generated through the following formula: ;
[0032] where X iter+1 represents the mutant individual obtained through the DE / best / 1 / bin mutation strategy in the (iter + 1)-th iteration process; V best represents the optimal individual in the population generated after the iter-th iteration; V r1 (iter), V r2 (iter) represent two randomly selected individuals from the population {V k (iter) | k = 1, 2,..., S} generated after the iter-th iteration, and the chromosomes of these two individuals are different; Z is the proportionality factor that controls the influence of the differential vector; iter is a positive integer greater than or equal to 1; It should be noted that when the first iteration is performed, that is, iter = 1, the mutant individual X2 = V best0 + Z * (V r1 (1) - V r2 (1)), where V best0 is the initial optimal individual, which can be determined by evaluating all individuals in the initial population pop through algorithms such as the fitness function fitness(), for example, the individual with the optimal fitness value can be used as V best0 , V r1 (1), V r2 (1) represent two randomly selected individuals from the initial population pop, and the chromosomes of these two individuals are different; S513. Repeat the above step S512, and perform the DE / best / 1 / bin mutation strategy S times in total. After each mutation strategy is executed, the corresponding mutant individual is obtained, and an intermediate population {X k (iter + 1) | k = 1, 2,..., S} containing all mutant individuals after the S times of the DE / best / 1 / bin mutation strategy is executed is constructed, where S represents the number of individuals in the initial population pop, that is, the total number of times the mutation strategy is executed, and k represents the individual number in the population; S514. Based on the individual crossover operation, transfer the chromosomes of the optimal individual to the offspring to obtain a candidate population, which specifically includes the following steps: Perform binomial crossover calculation on the population {V k (iter) | k = 1, 2,..., S} generated after the iter-th iteration and the intermediate population {X k (iter + 1) | k = 1, 2,..., S} to obtain a candidate population, which specifically includes the following steps: Assume V kn (iter), X kn (iter + 1) respectively represent the n dimensions of individual V k (iter) and the n dimensions of individual X k (iter + 1), where V k (iter) represents the k-th individual in the population {V k (iter) | k = 1, 2,..., S} generated after the iter-th iteration, and X k (iter + 1) represents the k-th individual in the intermediate population {X k (iter + 1) | k = 1, 2,..., S}, where n = 1, 2,..., 2U + 1; Obtain the gene value of the k-th individual based on the following binomial crossover calculation formula: ; where CR represents the crossover probability; randn is a random integer drawn from the dimension index range {1, 2,..., 2U + 1} to ensure that the individual crossover operation occurs at least in one dimension; rand is a random decimal drawn from the interval [0, 1]; Y kn (iter + 1) represents the gene value of the k-th individual in the n-th dimension; Repeat the above steps of obtaining the gene value of the k-th individual based on the binomial crossover calculation formula until the gene values of the k-th individual in all dimensions are obtained, and splice the gene values of all dimensions to obtain the k-th individual Y k (iter + 1); Repeat the above steps of obtaining the k-th individual Y k (iter + 1) until S individuals are obtained, and construct the S individuals into a candidate population {Y k (iter + 1) | k = 1, 2,..., S}; S515. Determine the global optimal individual in the current iteration process based on the fitness function fitness() between individuals; Specifically, in this embodiment, the global optimal individual in the current iteration process is determined based on the following formula: ; where V k (iter + 1) represents the global optimal individual in the current iteration process (i.e., the (iter + 1)-th iteration); the fitness function fitness() includes a greedy algorithm, which can select the optimal vector as the target vector for the next generation; Steps S512 - S515 are the (iter + 1)-th iteration process; S516. Take V k (iter + 1) = V best and repeat steps S512 - S515 to obtain the globally optimal individual in the current iteration process after each iteration is completed; S517. Construct an optimal individual population {V k (iter + 1)|k = 1, 2,..., S} that includes all globally optimal individuals in the current iteration process, and determine the globally optimal solution V after all iterations are completed based on the following formula from the optimal individual population V k (iter + 1)|k = 1, 2,..., S}; best_m ; ;
[0033] where fitness() is the fitness function, and the globally optimal solution V best_m is used as the final DNN task partitioning and offloading strategy and output.
[0034] Thus, based on the differential evolution (DE) algorithm, through its global search feature, this embodiment finds the optimal solution in the task partitioning and offloading strategy. Through multi - generation iterative adjustment, the DE algorithm can efficiently optimize the partitioning of DNN tasks, enabling the tasks to be optimally offloaded among multiple vehicle - type robots and drones, thereby reducing the task processing delay and optimizing the utilization of computing resources.
[0035] Further, based on the DDPG algorithm, the UAV flight trajectory planning is completed, including the following steps: S521. Initialize the actor policy network μ θ (s), the target policy network , the critic policy network Q W (s, a) and the target critic policy network using random network parameters w and θ respectively; s is the environmental state, including UAV state information and DNN task information, where the UAV state information includes one or several of the UAV's position, speed, remaining battery power, etc., and the DNN task information includes information related to the DNN task partitioning offloading and resource allocation strategy, such as one or several of the DNN task partitioning points, the computing resources of the robot, the computing resources of the UAV, etc., and a is the vector output by the actor policy network μ θ (s) in the environmental state s; Initialize the replay buffer rpm, the reward weight, and the soft - update coefficient τ; And, initialize the random noise, which can be used for subsequent action selection and exploration process to ensure that the algorithm can conduct effective exploration; S522. Initialize a random process for action exploration and limit the output range of the action to [0, 1]; S523. Complete action selection and action execution at the current time t; Among them, the action selection includes: According to the current participant policy network μ θ (s t ) and the random noise to select the action a t , and the action a t is the output of the current participant policy network t under the current environmental state s μ θ (s t ), and add random noise for exploration; The action execution includes: Input the current environmental state s t and the action a t , and issue a command to execute the action a t , and this process will affect the environmental state and bring corresponding rewards; Since the environmental state in step S521 includes UAV state information and DNN task information, the action selection and execution in steps S522 - S523 are not limited to the DNN partitioning offloading and resource allocation strategy, but also carry the UAV flight trajectory planning instruction at the same time; S524. Obtain the current DNN task partitioning offloading and resource allocation strategy; S525. Calculate the reward r t according to the feedback after executing the action a t , the reward r t can be used to reflect the effect of the current operation, and, calculate the updated environmental state s t according to the feedback after executing the action a t+1 to complete the state transition, and the updated environmental state s t+1 is used to represent the change of the environmental state after the execution of the current action a t ; As described above, since the action selection and execution process includes the UAV flight trajectory planning instruction, the reward calculation process in this step can read the UAV state information in the environmental state s t+1 generated after the action execution, thereby reflecting the impact of the UAV flight trajectory strategy on flight energy consumption and communication delay; S526. The quadruple (st , a t , r t , s t+1 ) are stored in the experience replay pool rpm for subsequent use during training. This storage mechanism helps to implement experience replay; S527. Randomly sample G quadruples from the experience replay pool rpm, and calculate the weighted expectation y of each quadruple according to the formula ; where γ is the reward discount factor; i ; and use the formula to calculate the minimized objective loss Loss to update the critic policy network Q (s, a), and use the formula W to calculate the sampled policy gradient to update the actor policy network (s); μ θ (s); S528. Use the formula and the formula to correspondingly update the target critic policy network , the target policy network . At this time, the training episode at the current moment t ends; S529. Repeat steps S523 - S528 to end the training episode corresponding to this moment at different times; S530. After all training episodes end, return the final UAV flight trajectory policy and the DNN partition offloading and resource allocation policy, and calculate the total reward to evaluate the effectiveness of the current DNN task partition offloading and resource allocation policy and the UAV flight trajectory policy.
[0036] As Figure 3 shown, in this embodiment, the total reward reward gradually increases with the increase in the number of training episodes, and finally starts to converge and stabilize at around 3000 training episodes. The converged training reward is approximately around -115, indicating that the DDPG algorithm in this embodiment can significantly reduce the total delay and total energy consumption of thrust calculation.
[0037] Furthermore, in this embodiment, the DDPG algorithm is used to perform UAV flight trajectory planning for two UAVs, namely UAV 1 and UAV 2, and the initial three-dimensional positions of the two UAVs are known. From Figure 4It can be seen that after the DDPG algorithm planning of this embodiment, the flight paths of UAV 1 and UAV 2 not only ensure exclusive descent into the area with dense robots, enabling UAV 1 to hover in a lower coordinate range to shorten the communication distance with the robots, but also form a synergy between high-altitude bandwidth coverage and backhaul requirements, enabling UAV 2 to cruise smoothly above the base station for quick upload of inference results. Moreover, the flight trajectories of both UAVs are characterized by smooth trajectories, large turning radii, and small acceleration fluctuations. This not only shortens the total flight distance but also significantly reduces the additional energy consumption caused by sudden acceleration or sharp turns. Secondly, this embodiment can online sense the changes in the UAV task distribution and resource status through the DDPG algorithm, continuously fine-tune the flight path, and achieve dynamic adaptability under multi-robot concurrent services, thereby effectively balancing communication latency, computing efficiency, and flight energy consumption, enabling the UAV to exhibit excellent comprehensive performance in complex environments.
[0038] Furthermore, in this embodiment, the number of robots is set to 10, 20, 30, 40, 50, 60, the number of UAVs is correspondingly set to 1, 2, 3, 4, 5, 6, the computing resources of each UAV are 60G FLOPS, the available communication network bandwidth of the UAV is 20MHz, and the maximum tolerable task latency is 1 second. The total latency, total energy consumption, and task completion rate after being processed by the method of this embodiment (i.e., the "Proposed" algorithm), DQN algorithm, PSO_DDPG algorithm, and DE_AC algorithm are obtained respectively under different numbers of robots.
[0039] As Figure 5 shown, as the number of robots increases, the total latency and total energy consumption of the DNN task inference of the four algorithms also increase. The reason is that since a DNN inference task is generated by each robot in each time slot, when the number of robots increases, the number of DNN inference tasks and the total inference latency will increase. At the same time, the UAVs also gradually need to serve more and more robots. Therefore, the coverage range that the UAVs need to cover gradually becomes larger, and the flight speed also gradually increases, resulting in an increase in the flight energy consumption of the UAVs and the energy consumption generated during the DNN task inference calculation of the UAVs.
[0040] However, under different numbers of robots, the total delay and total energy consumption after the Proposed algorithm are lower than those of the other three algorithms. Especially when the number of robots is 60, the total task inference delay of the Proposed algorithm is 16.7%, 10.3%, and 5.7% lower than those of the DQN, PSO_DDPG, and DE_AC algorithms respectively, and the total energy consumption is 13.6%, 9.1%, and 6.4% lower than those of the DQN, PSO_DDPG, and DE_AC algorithms respectively. The reason is that, compared with the other three algorithms, this embodiment combines the advantages of the DE algorithm and the DDPG algorithm. It can efficiently search for task division points and offloading objects in the initial stage through the excellent global optimization ability of the DE algorithm, and continuously learn and optimize strategies in a dynamic environment through the reinforcement learning framework of the DDPG algorithm, finely adjusting the trajectories of the UAVs, computing resource allocation, and bandwidth scheduling strategies to adapt to the real-time changing task requirements of multiple robots. At the same time, in the face of multi-robot collaborative tasks, resource constraints, and trajectory optimization and other multi-dimensional constraints, it has stronger global search ability and dynamic adaptation ability.
[0041] At the same time, as the number of robots gradually increases, the task completion rates of the four algorithms gradually decrease. The reason is that with the fixed computing resources of the UAVs, the computing resources allocated to each task gradually decrease as the number of robots increases. Therefore, the number of tasks exceeding the maximum tolerable delay gradually increases, and the task completion rate gradually decreases.
[0042] However, when the number of robots is 60, the task completion rate of the Proposed algorithm is 6.7%, 3.8%, and 2.2% higher than those of the DQN, PSO_DDPG, and DE_AC algorithms respectively.
[0043] Therefore, the number of robots in this embodiment is preferably 40 - 60. Embodiment 2:
[0044] This embodiment mainly examines the influence of the communication network bandwidth B between the UAVs and the robots on the total delay, total energy consumption, and task completion rate.
[0045] Specifically, in this embodiment, the communication network bandwidth B between the UAVs and the robots is set to 10, 15, 20, 25, 30, 35 (the unit is MHz), the computing resources of the UAVs are 60 G FLOPS, the maximum tolerable delay of the tasks is 1 second, the numbers of UAVs and robots are correspondingly set to 1, 2, 3, 4, 5, 6 and 20, 30, 40, 50, 60, 70, and each UAV can serve 10 - 12 robots. After being processed by the method of this embodiment (i.e., the "Proposed" algorithm), the DQN algorithm, the PSO_DDPG algorithm, and the DE_AC algorithm, the total delay, total energy consumption, and task completion rate are obtained.
[0046] As Figure 6 shown, as the available bandwidth of the UAV (i.e., the communication network bandwidth B between the UAV and the robot) gradually increases, the total inference latency of the DNN tasks of the four algorithms gradually decreases. The reason is that after the UAV bandwidth increases, the downlink transmission rate also increases, more communication resources can be allocated, the task transmission rate is improved, thereby significantly reducing the task transmission latency and indirectly reducing the total inference latency of the task. The total inference latency of the task will also gradually decrease.
[0047] Furthermore, the total inference latency of the Proposed algorithm is less than that of the other three algorithms. Especially when the available bandwidth of the UAV is 35 MHz, the total inference latency of the Proposed algorithm is 19.3%, 10.3%, and 4.1% lower than those of the DQN, PSO_DDPG, and DE_AC algorithms, respectively.
[0048] At the same time, as the available bandwidth of the UAV increases, the total system energy consumption and the task completion rate gradually increase. The reason for the increase in the total system energy consumption is the same as that in Embodiment 1 and will not be elaborated here. However, at different available bandwidths of the UAV, the total system energy consumption of the Proposed algorithm is less than that of the other three algorithms. Especially when the available bandwidth of the UAV is 35 MHz, the total system energy consumption of the Proposed algorithm is 11.9%, 6.0%, and 3.1% lower than those of the DQN, PSO_DDPG, and DE_AC algorithms, respectively. The task completion rate of the Proposed algorithm is 6.7%, 2.6%, and 1.4% higher than those of the DQN, PSO_DDPG, and DE_AC algorithms, respectively.
[0049] Therefore, in this embodiment, the communication network bandwidth B between the UAV and the robot is preferably 20 - 35 MHz. Embodiment 3:
[0050] This embodiment mainly examines the impact of the computing power of the UAV on the total latency, total energy consumption, and task completion rate.
[0051] Specifically, in this embodiment, the computing power of the same UAV is set to 50, 60, 70, 80, 90, 100 (the computing power is measured by FLOPS, and the unit is G), the communication network bandwidth B between the UAV and the robot is 20 MHz, the number of robots is set to 20, and the maximum tolerable latency of the task is 1 second. Corresponding to obtaining the total latency, total energy consumption, and task completion rate of the same UAV under different computing powers after being processed by the method of this embodiment (i.e., the "Proposed" algorithm), the DQN algorithm, the PSO_DDPG algorithm, and the DE_AC algorithm.
[0052] As Figure 7As shown, as the computing power of the drone gradually increases, the inference latency of the DNN tasks of the four algorithms gradually decreases. The reason is that there is a direct connection between the computing power of the drone and the latency of the drone to process the DNN inference task. When the computing power increases, the latency generated by the drone to process the DNN task gradually decreases. Therefore, the total latency of the task inference gradually decreases.
[0053] Moreover, under different computing powers, the total latency of the Proposed algorithm is less than that of the other three algorithms. Especially when the computing power of the drone is 100G FLOPS, the total latency of the task inference of the Proposed algorithm is 18.4%, 15.2% and 6.5% lower than that of the DQN algorithm, PSO_DDPG and DE_AC algorithms respectively.
[0054] At the same time, as the computing power of the drone increases, the total system energy consumption and the task completion rate both gradually increase. The reason for the increase in the total system energy consumption is the same as that in Embodiment 1 and Embodiment 2, which will not be elaborated here. However, under different computing powers of the drone, the total system energy consumption of the Proposed algorithm is less than that of the other three algorithms. Especially when the computing resource of the drone is 100G FLOPS, the system energy consumption of the algorithm proposed in this paper is 8.5%, 3.5% and 2.0% lower than that of the DQN algorithm, PSO_DDPG and DE_AC algorithms respectively; Furthermore, due to the increase in the computing power of the drone, the computing inference latency of the DNN task will gradually decrease. Thus, when the maximum tolerable latency of the task remains unchanged, the task completion rate will increase. Therefore, under different computing powers, the task completion rate of the Proposed algorithm is the highest compared to the other three algorithms. Especially when the computing resource of the drone is 100G FLOPS, the task completion rate of the Proposed algorithm is 4.0%, 1.8% and 1.2% higher than that of the DQN algorithm, PSO_DDPG and DE_AC algorithms respectively.
[0055] Therefore, in this embodiment, the computing power of the drone is measured in FLOPS, and the value range of FLOPS is preferably 80 - 100G. Embodiment 4:
[0056] The difference between this embodiment and Embodiment 1 is only that the collaborative operation optimization method of the drone and the robot further includes the following steps: Integrate communication elements and sensing elements on the robot, and based on the Integrated Sensing and Communication (ISAC) technology, use the unified spectrum and integrated signal to simultaneously achieve communication and sensing functions. Thus, it can not only improve the spectrum utilization efficiency, reduce the hardware cost, but also enhance the data sensing and analysis capabilities of multiple robots in complex working environments, thereby further improving the overall efficiency of the system.
[0057] In summary, the collaborative operation optimization method for drones and robots of the present invention can make full use of the high mobility and flexibility of drones, dynamically adjust the positions and task assignments of drones to optimize the inference performance of multiple robots. For example, when a certain robot has an excessive computing load, the drone can provide computing support according to the task requirements or assist it in executing tasks, thereby improving resource utilization efficiency and alleviating the computing bottleneck.
[0058] Specifically, the present invention is based on the differential evolution (DE) algorithm to find the optimal solution in the DNN task partitioning and offloading strategy, so as to efficiently optimize the partitioning of DNN tasks, enabling the tasks to be optimally offloaded among multiple vehicle-type robots and drones, thereby reducing the task processing delay and optimizing the utilization of computing resources. At the same time, by combining deep reinforcement learning and multi-agent collaboration through the DDPG algorithm, and by the agents independently optimizing the decision-making strategy to maximize the system efficiency, it can build a global resource awareness model based on real-time collection of information such as the inference requirements, communication status, and drone energy status of multiple robots, and enhance the multi-agent collaboration ability through an improved attention mechanism, enabling the drone to adaptively adjust the inference task offloading and communication resource allocation in a heterogeneous resource environment, further improving the inference efficiency.
[0059] In addition, the present invention also introduces the communication-sensing integration technology, which can not only improve the spectrum utilization efficiency, reduce the hardware cost, but also enhance the data sensing and analysis ability of multiple robots in a complex working environment, thereby further improving the overall efficiency of the system.
[0060] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An optimization method for the collaborative operation of an unmanned aerial vehicle and a robot, characterized in that, It includes the following steps: S1. Construct an inference acceleration framework including drones and robots; S2. Establish a first communication model between the drone and multiple robots, and establish a second communication model between the robots; S3. Obtain the total delay and total energy consumption required for the inference calculation of the DNN task based on the delay calculation model and the energy consumption calculation model respectively; S4. Construct a joint optimization model, and minimize the total delay of the DNN task inference calculation based on the constraint conditions; And, S5. Complete the optimization of the DNN task partitioning and offloading strategy based on the differential evolution algorithm, and complete the UAV flight trajectory planning based on the deep deterministic policy gradient algorithm.
2. The collaborative operation optimization method according to claim 1, characterized in that, The first communication model includes one or more of a path loss calculation model, an interference signal-to-noise ratio calculation model, and a data transmission rate calculation model.
3. The collaborative operation optimization method according to claim 1, characterized in that, The second communication model is as follows: ; Among them, R is the data transmission rate; B 0 is the communication network bandwidth when robots communicate with each other; P signal is the signal transmission power of the communication signal; N0 is the noise power spectral density; is the total power of the interference signal.
4. The collaborative operation optimization method according to claim 1, wherein The delay calculation model is as follows: ; Among them, is the DNN task I i (t) The total time delay required for the inference calculation to be completed; is the robot i when performing the DNN task I i (t) The time delay during the inference calculation from the 1st layer to the l layer; is the time delay generated when the robot i offloads the DNN task I i (t) to the UAV j during the time slot t; is the UAV j after receiving the DNN task I offloaded by the robot i i (t), and the UAV j performs the DNN task I i (t) The l +1 layer to the L layer during the inference calculation.
5. The collaborative operation optimization method according to claim 1, wherein, The energy consumption calculation model is as follows: ; Among them, is the total energy consumption generated during the inference calculation of DNN task I i (t); is the energy consumption generated during the inference calculation of the 1st to i -th layer of DNN task I l (t) by robot i in the t-th time slot; is the energy consumption generated when robot i offloads DNN task I i (t) to drone j in the t-th time slot; is the energy consumption generated when drone j performs the inference calculation of the i -th layer to the i -th layer of DNN task I l +1 after robot i offloads DNN task I L (t) to drone j; is the flight energy consumption generated by drone j in the t-th time slot.
6. The collaborative operation optimization method according to claim 4, wherein The joint optimization model is as follows: ; Where, N represents the set of robots; T represents the set of time slots; The constraint conditions include one or more of constraint conditions (a)-(k): ; Among them, the constraint condition (a) is used to ensure that within each time slot, all robots can only request services from the same drone, where a i,s (t) is a binary variable, indicating whether the edge server e carried on the drone decides to execute the DNN task requested by the robot in the t-th time slot i ; ES represents the set of edge servers carried on the drone; Constraint (b) is used to ensure that the robot i cooperates with only one edge server e carried by a drone for collaborative inference of the DNN task in the t-th time slot, where U represents the number of drones; Constraint (c) is used to ensure that the total amount of computing resources allocated to UAV j when processing tasks does not exceed its maximum computing resources; a i,j x_{ij}(t) is a binary variable indicating whether UAV j chooses to serve robot i in the t-th time slot; Constraint (d) is used to ensure that the amount of computing resources allocated to each task m by each drone j does not exceed its maximum computing resource limit; where, f i,m (t) represents the computing ability of robot i for task m in the t-th time slot; a i,m (t) represents the state variable of robot i processing task m; represents the maximum computing resources of the edge server carried on the drone; Constraints (e) and (f) are used to ensure that the UAV always flies within the preset flight area; where and respectively represent the abscissa and ordinate of the projection of UAV j on the horizontal plane at the t-th time slot; X and Y respectively represent the length and width of the projection of the preset flight area on the horizontal plane; u represents the set of UAVs. Constraint (g) is used to ensure that the flight speed of the UAV does not exceed its maximum speed limit v at any time slot max ; Constraint (h) is used to ensure that the distance between any two drones j and k at the same time is not less than the minimum position interval loc min ; where , are the positions of drones j and k at the same time respectively; Constraint (i) is used to ensure that the task division point of the DNN task is located between [0, L layers. Here, 0 means that the DNN task is fully offloaded to the drone for inference calculation, L means that the DNN task is fully performed on the robot for inference calculation; Constraint (j) is used to ensure that the battery level of the UAV is greater than zero at the end of each time slot; where, is the battery power consumption of UAV j in the t-th time slot; Constraint (k) is used to ensure that the total inference calculation delay of the DNN task does not exceed the maximum allowable tolerance delay; among them, is the maximum tolerance delay for the robot i to complete all DNN tasks I i (t) inference calculation.
7. The collaborative operation optimization method according to claim 1, wherein Completing the optimization of the DNN task partitioning and offloading strategy based on the differential evolution algorithm includes the following steps: S511. Randomly initialize S individuals to obtain an initial population pop = {V1, V2,..., V S}, and define the chromosome coding format of individual V S as V S = [A * , L * , F * ; Among them, A * represents the task selection execution status vector of the UAV at the t-th time slot, and L * represents the division point of the DNN task, and F * represents the task function vector executed by the UAV at the t-th time slot; S512. Generate mutant individuals based on the DE / best / 1 / bin mutation strategy and obtain an intermediate population; Specifically, in this step, the mutant individuals are generated by the following formula: ; Among them, X iter+1 represents the mutant individual obtained by the DE / best / 1 / bin mutation strategy in the (iter + 1)-th iteration process; V best represents the optimal individual in the population generated after the iter-th iteration; V r1 (iter), V r2 (iter) represents two individuals randomly selected from the population {V k (iter) | k = 1, 2,..., S} generated after the iter-th iteration, and the chromosomes of these two individuals are different; Z is the proportionality factor that controls the influence of the difference vector; S513. Repeat the above step S512, and execute the DE / best / 1 / bin mutation strategy S times in total. After each execution of the mutation strategy, obtain the corresponding mutated individual, and construct an intermediate population {X k (iter + 1)|k = 1, 2,..., S} containing all the mutated individuals after S executions of the DE / best / 1 / bin mutation strategy. Here, S represents the number of individuals in the initial population pop, that is, the total number of times the mutation strategy is executed, and k represents the individual number in the population; S514. Transfer the chromosome of the optimal individual to the offspring based on the individual crossover operation to obtain a candidate population, which specifically includes the following steps: After the completion of the iter-th iteration, the population {V k (iter) | k = 1, 2,..., S} and the intermediate population {X k (iter + 1) | k = 1, 2,..., S} are subjected to binomial crossover calculation to obtain a candidate population, which specifically includes the following steps: Assume V kn (iter), X kn (iter + 1) respectively represent the n dimensions of individual V k (iter) and the n dimensions of individual X k (iter + 1), where V k (iter) represents the k-th individual in the population {V k (iter) | k = 1, 2,..., S} generated after the iter-th iteration, X k (iter + 1) represents the k-th individual in the intermediate population {X k (iter + 1) | k = 1, 2,..., S}, where n = 1, 2,..., 2U + 1; Obtain the gene value of the k-th individual based on the following binomial crossover calculation formula: ; where CR represents the crossover probability; randn is a random integer drawn from the dimensional index range {1, 2,..., 2U + 1}; rand is a random decimal drawn from the interval [0, 1]; Y kn (iter + 1) represents the gene value of the k-th individual in the n-th dimension; Repeat the above steps of obtaining the gene value of the k-th individual based on the binomial crossover calculation formula until the gene values of the k-th individual in all dimensions are obtained, and splice the gene values of all dimensions to obtain the k-th individual Y k (iter + 1); Repeat the above for the k-th individual Y k (steps for obtaining (iter + 1)) until S individuals are obtained, and construct the S individuals into a candidate population {Y k (iter + 1) | k = 1, 2,..., S}; S515. Determine the global optimal individual in the current iteration process based on the fitness function fitness() between individuals; Specifically, in this embodiment, the global optimal individual in the current iteration process is determined based on the following formula: ; Among them, V k (iter + 1) represents the globally optimal individual in the current iteration process; the fitness function fitness() includes a greedy algorithm; S516. Take V k (iter + 1) = V best and repeat steps S512 - S515 to obtain the globally optimal individual in the current iteration process after each iteration is completed; S517. Construct an optimal individual population {V k (iter + 1)|k = 1, 2,..., S} that includes all the globally optimal individuals in the current iteration process, and determine the globally optimal solution V k (iter + 1)|k = 1, 2,..., S} after all iterations are completed from the optimal individual population V best_m ; ; Among them, fitness() is the fitness function, and the global optimal solution V best_m That is, it serves as the final DNN task partitioning and offloading strategy.
8. The collaborative operation optimization method according to claim 1, wherein Completing the UAV flight trajectory planning based on the deep deterministic policy gradient algorithm includes the following steps: S521. Initialize the actor policy network μ θ (s), the target policy network , the critic policy network Q W (s,a), and the target critic policy network with random network parameters w and θ; s is the environmental state, including the UAV state information and the DNN task information, and a is the vector output by the actor policy network μ θ (s) under the environmental state s. Initialize the initialized experience replay pool rpm, the reward weight, and the soft update coefficient τ; And, initialize the random noise; S522. Initialize a random process for action exploration and limit the output range of the action to [0,1]; S523. Complete the action selection and action execution at the current time t; Wherein, the action selection includes: According to the current participant policy network μ θ (s t ) and random noise, select action a t , and the action a t is the output of the current participant policy network t under the current environmental state s μ θ (s t ); The action execution includes: Input the current environmental state s t and action a t and issue a command to execute action a t ; S524. Obtain the current DNN task partitioning offloading and resource allocation strategy; S525. Calculate the reward r based on the feedback after executing action a t and calculate the updated environmental state s based on the feedback after executing action a t to complete the state transition; t t+1 S526. Store the quadruple (s t+1 , a t , r t , s t+1 ) in the experience replay pool rpm; S527. Randomly extract G quadruples from the experience replay pool rpm, and calculate the weighted expectation y of each quadruple according to the formula ; where γ is the reward discount factor; i Calculate the weighted expectation y of each quadruple And, using the formula to calculate the minimized objective loss Loss to update the critic policy network Q W (s,a), and using the formula to calculate the sampled policy gradient to update the actor policy network μ θ (s); S528. Use the formula and the formula to update the target critic policy network and the target policy network respectively. At this time, the training episode at the current time t ends; S529. Repeat steps S523-S528 to end the training episode corresponding to this moment at different times; S530. After all the training episodes end, return the final UAV flight trajectory strategy and the DNN partitioning offloading and resource allocation strategy.
9. The collaborative operation optimization method according to claim 8, wherein, Step S530 further includes: calculating the total reward , for evaluating the effects of the current DNN task partitioning and offloading and resource allocation strategy and the UAV flight trajectory strategy.
10. The collaborative operation optimization method according to claim 1, wherein The collaborative operation optimization method of the drone and the robot further includes the following steps: Integrate communication components and sensing components on the robot, and based on the communication-sensing integration technology, use a unified spectrum and integrated signals to simultaneously implement communication and sensing functions.
Citation Information
Patent Citations
Deep neural network optimizing method based on coevolution and back propagation
CN106650933A
Edge-end collaborative AI model reasoning method based on energy consumption prediction
CN117331699A
Multi-unmanned aerial vehicle intelligent path planning method based on deep reinforcement learning
CN117553803A
Edge calculation unloading method based on unmanned aerial vehicle assistance
CN118200325A