Unmanned cluster task unloading method based on TRPO algorithm
By applying the TRPO algorithm and multi-agent Markov decision-making process in the unmanned cluster system, the problem of unbalanced resource utilization of unmanned cluster networking is solved, the system delay, energy and service quality are jointly optimized, and the fairness of resource allocation is improved.
Patent Information
- Application Number
- CN202510111583.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The unbalanced resource utilization of unmanned cluster networking collaborative resources leads to system delay, unoptimization of energy and service quality, and there are problems of unfair resource allocation.
The unmanned cluster task offload method based on TRPO algorithm is adopted to establish a three-layer computing framework for unmanned clusters, monitor the active status of ground equipment, build communication, delay and energy consumption models, make dynamic decisions through the multi-agent Markov decision-making process, and use the TRPO algorithm to jointly optimize delay and energy.
Effectively improve the fairness of resource allocation in unmanned cluster networking, optimize system delay and energy consumption, improve service quality, and extend the system's life cycle.
Smart Images

Figure CN119997267A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an unmanned cluster task unloading method based on a TRPO algorithm, and belongs to the technical field of wireless communications. Background Art
[0002] Mobile edge computing, as an emerging computing architecture, is used in digital healthcare, Internet of Vehicles, maritime communications, smart industry and other fields because it supports a large number of node interconnections and wide geographical distribution. Mobile edge computing, referred to as MEC, improves response delay and service performance by deploying computing servers in edge networks close to terminals without transmitting large amounts of data back to remote data centers. However, in areas far from the core Internet, setting up ground MEC base stations is a difficult task and expensive, so in emergency situations, areas often have limited resources. An unmanned cluster system is a system composed of multiple unmanned devices. Common unmanned devices include drones, unmanned vehicles, and embodied robots, which can collaborate to complete specific tasks. Due to the strong mobility and convenient deployment of unmanned devices, unmanned cluster networking as edge servers is considered to be an effective solution to improve user experience under MEC, especially for applications with high data real-time requirements. In addition, air communication relays operated by autonomous unmanned devices can quickly and securely transmit data collected by ground devices to remote command and control centers. This technology expands the communication range and can provide a higher degree of freedom.
[0003] Compared with traditional distributed hierarchical networks, the collaborative process of unmanned cluster networking has new problems. Ignoring the energy bottleneck of individual unmanned devices may cause failed nodes in the network, causing catastrophic disconnection and loss of connection. In addition to large-scale application scenarios, high-speed and frequent interactions are also constantly impacting performance bottlenecks. Optimizing resource allocation and providing low-latency services are urgent tasks. In recent years, deep reinforcement learning methods have attracted more and more attention in the field of wireless communications. The core idea of deep reinforcement learning is to use deep neural networks to approximate value functions and policy functions, thereby achieving learning and decision-making in complex environments. MEC assisted by unmanned cluster networking has considerable complexity in transmission control and offloading decisions. The design of offloading strategies needs to consider factors such as time-varying channel conditions, user mobility, energy supply, computing workload, and cache capacity. Therefore, methods based on deep reinforcement learning have become effective candidates for joint optimization. Summary of the invention
[0004] The present invention provides an unmanned cluster task offloading method based on the TRPO algorithm to solve the problem of unbalanced resource utilization in unmanned cluster networking collaboration, realize the joint optimization of system delay, energy and service quality, and effectively improve the fairness of unmanned cluster networking resource allocation.
[0005] An unmanned cluster task offloading method based on the TRPO algorithm comprises the following steps:
[0006] Step 1: Establish a three-layer computing framework for unmanned clusters, monitor the active status of ground equipment and receive tasks periodically generated by ground equipment, and represent the task information as a set based on the workload and maximum tolerable delay;
[0007] Step 2: Establish a Markov Poisson modulated random process to represent task arrival and quantitatively analyze the service delay of time-varying tasks in the model;
[0008] Step 3: Build the system's communication model, delay model, and energy consumption model;
[0009] Step 4: Introduce a multi-agent Markov decision process to describe the dynamic decision-making process of S2, and agents share information to make real-time responses;
[0010] Step 5: Use the TRPO algorithm to jointly optimize latency and energy, train the model to obtain the best unloading decision, and use the best decision to perform unloading.
[0011] The three-layer computing framework of the unmanned cluster in step 1 includes a task collection layer, a relay transmission layer, and an information control layer from bottom to top. The K small unmanned devices in the task collection layer are responsible for receiving tasks generated by the cell, which are represented by the set K = {1, 2, ..., K}. k ∈K, ensuring that the task can be received and maintaining a quasi-static state; the M unmanned devices in the relay transmission layer can act as relays to forward the task to the top layer in multiple hops, and also act as servers to perform local computing to achieve offload, which is represented by the set M = {1,2,…,M}, U m ∈M; large unmanned equipment U at the information control layer h With powerful computing efficiency and large-capacity battery, it acts as the leader to plan the entire system task offloading process.
[0012] The active state of the ground equipment in step 1 is determined by process A k (t) state space Characterization, constant elements Belongs to {0,1}, 0 indicates low activity and 1 indicates high activity.
[0013] The task information set in step 1 includes three parameters, which are expressed as Where V j Indicates the amount of data to be processed, O j Describes the total number of CPU cycles required to complete the task, Indicates the maximum delay that the task can tolerate, exceeding Task delivery failed; the task is operated in binary offload mode, and the unmanned device U kThe uninstall decision is α j ∈{0,…,N} represents; when α j = n, indicating that task j will be n The computing unit is executed and added to the offload queue Wait; that is, when α j = 0, the task is calculated locally, j joins The queue follows a first-come, first-served principle.
[0014] The random process of task arrival in step 2 follows Poisson distribution. For the process A where task J arrives at any unmanned device at the information collection layer k (t) Modeled as a two-state Markov modulated Poisson process.
[0015] The communication model, delay model, and energy consumption model of the system constructed in step 3 specifically include:
[0016] If the task is executed locally, task j is executed on the unmanned device U k Calculate the delay as
[0017] If the small unmanned device chooses to unload, since the unmanned devices in the community network are in a fixed position, the position of the small unmanned device is expressed as (x k ,y k ,z k ), another unmanned device U m The position at time t is (x m (t),y m (t),z m (t)), the distance between them is
[0018] According to the free space propagation model, the unmanned cluster network is modeled, and the channel gain is Where d0 represents the reference distance; according to Shannon's formula, the maximum transmission rate can reach Where W represents the channel bandwidth, P s represents the signal power of a small unmanned device, β0 represents the average channel power gain at the reference distance d0 = 1m, G km (t) represents the channel gain at time t, N0 represents the channel Gaussian white noise, and θ represents the azimuth angle; task j is in U k and U m The delay in the one-to-one transmission process is
[0019] If the offloaded object is not within the communication range, define the set N u Indicates the unmanned equipment associated with the unloading path. j =n, n∈Nu And it is the end point of unloading, the transmission time is the sum of multi-hop transmission, and the corresponding delay is:
[0020]
[0021] The total delay of a task is:
[0022]
[0023] The method for establishing the energy consumption model of the system in step 3 includes:
[0024] The energy loss includes the energy consumption cost of task j in the offloading and execution stages, including the energy consumed by the CPU to execute the task and the energy consumed by the offloading and forwarding task;
[0025] If the task is executed locally, task j is executed on the unmanned device U k The energy consumed by local computing is where κ is the effective switch capacitance that depends on the chip architecture; if the task is offloaded to the computing node U n The corresponding energy consumption is The total energy consumption of a task is:
[0026]
[0027] The basic elements of the Markov decision process in step 4 include:
[0028] State space: The states that a single agent can observe include the active state of the equipment in the coverage area, the congestion level of the local queue, and the state of other unmanned equipment in the environment. The state set of all agents at time slot t is represented by S(t) = {S A (t),S e (t),S o (t),S u (t)};
[0029] Action space: Each unmanned device makes a decision on the offloading object of task j based on observations, α j ∈{0,…,N},α j >0 represents the server's uninstall object, α j =0 means local calculation;
[0030] State transfer: After the agent takes an action, the action interacts with the environment, and the process of the entire system state transfer is Q Σ , whose generic element is defined as Q[s,s′]=Pr(S Σ (t) = s′|S Σ (t-1) = s), where s and s′ are two state constants of the Markov chain;
[0031] Reward: The goal of joint optimization is to minimize the delay and energy consumption of each unmanned device. The unmanned device will receive a reward after interacting with the environment in each time slot, ω k (t) represents the unmanned equipment U k Process the delay and energy cost of the task in time slot t; calculate the reward function ω = β*t total +(1-β)*E total , where β (0<β<1) is the weight of the delay; the reward function is Calculate, where E{ω i (t)}=∑ω*f k (ω), f k (ω) is the probability distribution function of the unmanned equipment overhead.
[0032] The training of the improved multi-agent TRPO algorithm in step 5 includes the following steps:
[0033] S5.1: Initialize state s1 and reset the environment;
[0034] S5.2: Before the maximum number of rounds is met, each agent obtains observations o(t) from the environment and follows the strategy π ′ (a t |o t )
[0035] θ
[0036] Perform actions and get immediate rewards t , observe the state s of the new environment t+1 , the trajectory Tr i (t) = {o i (t),a i (t),r i (t),s i (t+1)} is stored in the buffer pool, and the advantage function of each state is calculated;
[0037] S5.3: Estimate the sample policy model gradient and calculate the step size using the conjugate gradient algorithm;
[0038] S5.4: Update the Actor network parameters θ when the KL divergence is satisfied;
[0039] S5.5: Optimize the value network parameter φ to minimize the mean squared error between rewards and state values;
[0040] S5.6: Enter the decentralized execution phase and execute the trained network independently on each unmanned device.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The present invention designs a novel unmanned cluster networking computing framework, which not only clarifies the functions and roles of unmanned devices in the system and studies the operation of task computing offloading, but also focuses on the performance differences between individual unmanned devices in computing efficiency and communication distance, thereby solving the limitations of the flat role division of traditional unmanned cluster networking. Based on the improved multi-agent trust region optimization algorithm, the comprehensive optimal solution for unmanned device resource offloading is obtained, the system delay and energy are jointly optimized, the energy consumption of unmanned devices is constrained, and the heterogeneous network management is optimized, which effectively improves the fairness of unmanned cluster networking resource allocation and extends the life cycle of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1 The specific structure flow chart of the model training of the unmanned cluster task offloading method based on the TRPO algorithm of the present invention;
[0045] Figure 2 It is a schematic diagram of a three-layer computing framework of the unmanned cluster task offloading method based on the TRPO algorithm of the present invention;
[0046] Figure 3 It is the framework and structure diagram of the TRPO algorithm of the unmanned cluster task offloading method based on the TRPO algorithm of the present invention. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] Figure 1 : is a flow chart of an unmanned cluster task offloading method based on the TRPO algorithm according to an embodiment of the present invention, the method comprising the following steps:
[0049] Step 1: Establish a three-layer computing framework for unmanned clusters, monitor the active status of ground equipment and receive tasks periodically generated by ground equipment, and represent the task information as a set based on the workload and maximum tolerable delay;
[0050] Step 2: Establish a Markov Poisson modulated random process to represent task arrival and quantitatively analyze the service delay of time-varying tasks in the model;
[0051] Step 3: Build the system's communication model, delay model, and energy consumption model;
[0052] Step 4: Introduce a multi-agent Markov decision process to describe the dynamic decision-making process of S2, and agents share information to make real-time responses;
[0053] Step 5: Use the TRPO algorithm to jointly optimize latency and energy, train the model to obtain the best unloading decision, and use the best decision to perform unloading.
[0054] Preferably, in step 1, a three-layer computing framework for unmanned equipment is established, such as Figure 2 As shown in the figure, the framework includes the task collection layer, relay transmission layer and information control layer from bottom to top, covering unmanned devices with different functions. K small unmanned devices are deployed in the task collection layer to collect tasks, represented by the set K = {1, 2, ..., K}, U k ∈K; K unmanned devices are divided into G cells, covering the divided ground area, and the unmanned devices in the cells can communicate with each other; the small unmanned devices in the relay transmission layer are represented by the set M = {1,2,…,M}, U m ∈M; deploy large unmanned equipment with powerful computing efficiency and large-capacity batteries at the top layer, represented by U h ; Unmanned devices use binary offloading to allocate computing resources; Based on the above process, the physical activity range constraints of the three unmanned devices are: U k Deployed in a fixed position at low altitude to relay the U m The flight space width is limited to the area, the height is limited to below the information control layer, and large unmanned equipment U h In the center of the cluster;
[0055] The entire network acts as an MEC system, and the unmanned device acts as the MEC server. When the ground device periodically generates tasks, it dispatches the small unmanned device U corresponding to the cell task collection layer. k Receive tasks; Each unmanned device is equipped with a separate computing unit, referred to as CU, which is responsible for performing local calculations; Assume that the CPU performance level of unmanned devices other than large unmanned devices in the network is consistent, that is, the number of CPU calculation cycles per second is f; U k There are two queues in the internal CU. They represent the set of task execution queues and unloading queues at time t respectively, and the generated task j is expressed as Where V j represents the amount of data to be processed, and O j Describes the total number of CPU cycles required to complete the task, Indicates the maximum delay that the task can tolerate, exceeding Task delivery failed; the task is operated in binary offload mode, and the unmanned device U k The uninstall decision is α j ∈{0,…,N} represents; when α j = n, indicating that task j is handled by unmanned equipment U n The calculation unit is executed, adding Wait; that is, when α j = 0, the task is calculated locally, j joins The queue follows the principle of first come first served; the active state of the ground equipment is determined by process A k (t) State space Indication, its constant element belongs to {0,1}, 0 indicates low activity, and 1 indicates high activity; τ seconds are divided into decision cycles. Assuming that during τ seconds, the state of the entire system does not change. During this process, the leader U h Responsible for multi-dimensional evaluation of network status, achieving global coordination and making efficient decisions; the lower-level unmanned devices transmit beacon messages containing their own positions and queue status to the unmanned device leader U h .
[0056] Preferably, the specific steps of step 2 are as follows:
[0057] Assuming that the random process of task arrival follows Poisson distribution, for task J arriving at any information collection layer unmanned equipment U k Process A k (t) is modeled as a two-state Markov modulated Poisson process, referred to as MMPP, According to the above process, the state probability transfer matrix Q is established k , the rate matrix Λ k :
[0058]
[0059] Among them, the job arrival rate in state 0 is λ1, and the job arrival rate in state 1 is λ2; π represents the steady-state probability vector of the system, which satisfies the following equation in steady state:
[0060]
[0061] Where e is a unit vector; after determining the steady-state probability, the average job arrival rate of the system after steady-state is obtained:
[0062]
[0063] Let q e ,q oRespectively represent the maximum length of the execution queue and the unload queue, and define They respectively indicate the state of the underlying Markov chain at time t, indicating that at time t U k The number of cached tasks in the execution queue and the offload queue.
[0064] Preferably, the specific steps of step three are as follows:
[0065] Model the communication and latency of the system:
[0066] If the task is executed locally, task j is executed on the unmanned device U k Calculate the delay as
[0067] If the small unmanned device chooses to unload, since the unmanned devices in the community network are in a fixed position, the position of the small unmanned device is expressed as (x k ,y k ,z k ), another unmanned device U m The position at time t is (x m (t),y m (t),z m (t)), the distance between them is
[0068] Based on the free space propagation model, the unmanned equipment cluster network is modeled, and the channel gain is Among them, d0 represents the reference distance, which is equal to 1m; according to Shannon's formula, the maximum transmission rate can reach Where W represents the channel bandwidth, P s represents the signal power of a small unmanned device, β0 represents the average channel power gain at the reference distance d0 = 1m, G km (t) represents the channel gain at time t, N0 represents the channel Gaussian white noise, and θ represents the azimuth. According to the above analysis, task j is in U k and U m The delay in the one-to-one transmission process is
[0069] If the offloaded object is not within the communication range, define the set N u Indicates the unmanned equipment associated with the unloading path. j =n, n∈N u And it is the end point of unloading. At this time, the transmission time is the sum of multi-hop transmission, and the corresponding delay is:
[0070]
[0071] The total delay of a task is:
[0072]
[0073] Establish the energy consumption model of the system:
[0074] The energy loss includes the energy consumption cost of task j in the offloading and execution stages, including the energy consumed by the CPU to execute the task and the energy consumed by the offloading and forwarding task;
[0075] If the task is executed locally, task j is executed on the unmanned device U k The energy consumed by local computing is where k is the effective switch capacitance that depends on the chip architecture; if the task is offloaded to the computing node U n The corresponding energy consumption is The total energy consumption of a task is:
[0076]
[0077] The basic elements of establishing the Markov decision process in step 4 are expressed as follows:
[0078] State space: The states that a single agent can observe include the active state of the equipment in the coverage area, the congestion level of the local queue, and the state of other unmanned equipment in the environment. The state set of all agents at time slot t is represented by S(t) = {S A (t),S e (t),S o (t),S u (t)};
[0079] Action space: Each unmanned device makes a decision on the offloading object of task j based on observations, α j ∈{0,…,N},α j >0 represents the server's uninstall object, α j =0 means local calculation;
[0080] State transfer: After the agent takes an action, the action interacts with the environment, and the process of the entire system state transfer is Q ∑ , whose generic element is defined as Q[s,s′]=Pr(S ∑ (t) = s′|S ∑ (t-1) = s), where s and s′ are two state constants of the Markov chain;
[0081] Reward: The goal of joint optimization is to minimize the delay and energy consumption of each unmanned device. The unmanned device will receive a reward after interacting with the environment in each time slot, ω k (t) represents the unmanned equipment U k Process the delay and energy cost of the task in time slot t; calculate the reward function ω = β*t total+(1-β)*E total , where β (0<β<1) is the weight of the delay; the reward function is Calculate, where E{ω i (t)}=∑ω*f k (ω), f k (ω) is the probability distribution function of the unmanned equipment overhead.
[0082] Preferably, if Figure 3 As shown in Figure 2, the steps of training the improved multi-agent TRPO algorithm in step 5 are as follows:
[0083] S5.1: Initialize state s1 and reset the environment;
[0084] S5.2: Before the maximum number of rounds is met, each agent obtains observations o(t) from the environment and follows the strategy π θ′ (a t |o t ) Execute the action and get the immediate reward r t , observe the state s of the new environment t+1 , the trajectory Tr i (t) = {o i (t),a i (t),r i (t),s i (t+1)} is stored in the buffer pool, and the advantage function of each state is calculated according to the following formula, namely:
[0085]
[0086] in, is the advantage function of strategy π, that is, the advantage of action a over the average action in state s, Q represents the state-action value function, and V represents the value function in state s;
[0087] S5.3: Estimate the sample policy model gradient and calculate the step size using the CG conjugate gradient algorithm;
[0088] S5.4: When the KL divergence is satisfied, update the Actor network parameter θ to maximize the difference between the cumulative reward value obtained by using the new strategy and the old strategy. The optimization goal is,
[0089]
[0090] Among them, π θ′ is the new strategy, π θ It's an old strategy. is π θ The advantage estimate of D KL(·) is the KL divergence, which avoids drastic iteration of the strategy and limits the update step size; δ is the threshold of the expected KL divergence in the trust region. The conjugate gradient method is used to solve this optimization problem and update the strategy parameter θ i ;
[0091] S5.5: Optimize value network parameters To minimize the mean squared error between the reward and the state value, the loss function is defined as:
[0092]
[0093] S5.6: Enter the decentralized execution phase and execute the trained network independently on each unmanned device.
[0094] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.
Claims
1. An unmanned cluster task offloading method based on the TRPO algorithm, characterized by: The following steps are involved: Step 1: Establish a three-layer computing framework for unmanned clusters, monitor the active status of ground equipment and receive tasks periodically generated by ground equipment, and represent the task information as a set based on the workload and maximum tolerable delay; Step 2: Establish a Markov Poisson modulated random process to represent task arrival and quantitatively analyze the service delay of time-varying tasks in the model; Step 3: Build the system's communication model, delay model, and energy consumption model; Step 4: Introduce a multi-agent Markov decision process to describe the dynamic decision-making process of S2, and agents share information to make real-time responses; Step 5: Use the TRPO algorithm to jointly optimize latency and energy, train the model to obtain the best unloading decision, and use the best decision to perform unloading.
2. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 1, characterized in that: The three-layer computing framework of the unmanned cluster in step 1 includes a task collection layer, a relay transmission layer, and an information control layer from bottom to top. The K small unmanned devices in the task collection layer are responsible for receiving tasks generated by the cell, which are represented by the set K = {1, 2, ..., K}. k ∈K, ensuring that the task can be received and maintaining a quasi-static state; the M unmanned devices in the relay transmission layer can act as relays to forward the task to the top layer in multiple hops, and also act as servers to perform local computing to achieve offload, which is represented by the set M = {1,2,…,M}, U m ∈M; large unmanned equipment U at the information control layer h With powerful computing efficiency and large-capacity battery, it acts as the leader to plan the entire system task offloading process.
3. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 2, characterized in that: The active state of the ground equipment in step 1 is determined by process A k (t) state space Characterization, Belongs to {0,1}, 0 indicates low activity and 1 indicates high activity.
4. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 3 is characterized in that: The task information set in step 1 includes three parameters, which are expressed as Where V j Indicates the amount of data to be processed, O j Describes the total number of CPU cycles required to complete the task, Indicates the maximum delay that the task can tolerate, exceeding Failure to deliver the task; The task is operated in binary uninstall mode, and the unmanned device U k The uninstall decision is α j ∈{0,…,N} represents; when α j = n, indicating that task j will be handled by unmanned equipment U n The computing unit is executed and added to the offload queue Wait; that is, when α j = 0, the task is calculated locally, j joins The queue follows a first-come, first-served principle.
5. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 1, characterized in that: The random process of task arrival in step 2 follows Poisson distribution. For the process A where task J arrives at any unmanned device at the information collection layer k (t) Modeled as a two-state Markov modulated Poisson process.
6. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 1, characterized in that: The communication model and delay model of the system constructed in step 3 specifically include: If the task is executed locally, task j is executed on the unmanned device U k Calculate the delay as If the small unmanned device chooses to unload, since the unmanned devices in the community network are in a fixed position, the position of the small unmanned device is expressed as (x k ,y k ,z k ), another unmanned device U m The position at time t is (x m (t),y m (t),z m (t)), the distance between them is According to the free space propagation model, the unmanned cluster network is modeled, and the channel gain is Where d0 represents the reference distance; according to Shannon's formula, the maximum transmission rate can reach Where W represents the channel bandwidth, P s represents the signal power of a small unmanned device, β0 represents the average channel power gain at the reference distance d0 = 1m, G km (t) represents the channel gain at time t, N0 represents the channel Gaussian white noise, and θ represents the azimuth angle; task j is in U k and U m The delay in the one-to-one transmission process is If the offloaded object is not within the communication range, define the set N u Indicates the unmanned equipment associated with the unloading path. j =n, n∈N u And it is the end point of unloading, the transmission time is the sum of multi-hop transmission, and the corresponding delay is: The total delay of a task is:
7. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 1, characterized in that: The method for establishing the energy consumption model of the system in step 3 includes: The energy loss includes the energy consumption cost of task j in the offloading and execution stages, including the energy consumed by the CPU to execute the task and the energy consumed by the offloading and forwarding task; If the task is executed locally, task j is executed on the unmanned device U k The energy consumed by local computing is where κ is the effective switch capacitance that depends on the chip architecture; if the task is offloaded to the computing node U n The corresponding energy consumption is The total energy consumption of a task is:
8. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 1, characterized in that: The basic elements of the Markov decision process in step 4 include: State space: The states that a single agent can observe include the active state of the equipment in the coverage area, the congestion level of the local queue, and the state of other unmanned equipment in the environment. The state set of all agents at time slot t is represented by S(t) = {S A (t),S e (t),S o (t),S u (t)}; Action space: Each unmanned device makes a decision on the offloading object of task j based on observations, α j ∈{0,…,N},α j >0 represents the server's uninstall object, α j =0 means local calculation; State transfer: After the agent takes an action, the action interacts with the environment, and the process of the entire system state transfer is Q Σ , whose generic element is defined as Q[s,s']=Pr(S Σ (t) = s' | S Σ (t-1) = s), where s and s' are two state constants of the Markov chain; Reward: The goal of joint optimization is to minimize the delay and energy consumption of each unmanned device. The unmanned device will receive a reward after interacting with the environment in each time slot, ω k (t) represents the unmanned equipment U k Process the delay and energy cost of the task in time slot t; calculate the reward function ω = β*t total +(1-β)*E total , where β (0<β<1) is the weight of the delay; the reward function is Calculate, where E{ω i (t)}=∑ω*f k (ω), f k (ω) is the probability distribution function of the unmanned equipment overhead.
9. The unmanned cluster task offloading method based on the TRPO algorithm according to claim 1, characterized in that: The training of the improved multi-agent TRPO algorithm in step 5 includes the following steps: S5.1: Initialize state s1 and reset the environment; S5.2: Before the maximum number of rounds is met, each agent obtains observations o(t) from the environment and follows the strategy π θ' (a t |o t ) Execute the action and get the immediate reward r t , observe the state s of the new environment t+1 , the trajectory Tr i (t) = {o i (t),a i (t),r i (t),s i (t+1)} is stored in the buffer pool, and the advantage function of each state is calculated, that is, in, is the advantage function of strategy π, that is, the advantage of action a over the average action in state s, Q represents the state-action value function, and V represents the value function in state s; S5.3: Estimate the sample policy model gradient and calculate the step size using the conjugate gradient algorithm; S5.4: Update the Actor network parameters θ when the KL divergence is satisfied; S5.5: Optimize the value network parameter φ to minimize the mean squared error between the reward and the state value. The loss function is defined as: S5.6: Enter the decentralized execution phase and execute the trained network independently on each unmanned device.
Citation Information
Patent Citations
Internet of Things edge task unloading method and device
CN113225377A
Resource allocation and task unloading optimization method based on multiple agents
CN115175217A
Heterogeneous computing power-oriented multi-policy intelligent scheduling method and apparatus
WO2024060571A1
MEC unloading, resource allocation, and cache joint optimization method
WO2024240038A1
Agent policy learning method with privacy protection in mobile edge computing
WO2024254892A1