Metacosmic task unloading and resource scheduling method based on deep reinforcement learning

Through the multi-node collaborative optimization method of deep reinforcement learning, the dynamic allocation and energy consumption of edge computing systems in the metaverse is solved, flexible adjustment and efficient utilization of computing resources are achieved, and user experience and drone battery life are improved.

CN120335892AActive Publication Date: 2025-07-18JILIN UNIVERSITY

Patent Information

Application Number
CN202510838738.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-18
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing edge computing systems lack flexibility in the metaverse and cannot dynamically adjust the deployment of computing resources, resulting in dynamic allocation of computing resources, uneven resource utilization, and traditional strategies are difficult to adapt to complex and changeable user mobility and network state, which has energy waste.

Method used

Using a multi-node collaborative optimization method based on deep reinforcement learning, task offloading and resource scheduling between the drone cluster and fixed edge servers is performed through the PPO algorithm, and the deployment and allocation of computing resources are dynamically adjusted, resource sharing and intelligent scheduling between multiple nodes is realized, and energy consumption is optimized.

Benefits of technology

It realizes the adjustment of the flexible activity of computing resources, improves resource utilization and system efficiency, reduces latency and energy consumption, extends the battery life of the drone, and improves the real-time and sustainability of metacosmic services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335892A_ABST
    Figure CN120335892A_ABST
Patent Text Reader

Abstract

The invention discloses a meta universe task offloading and resource scheduling method based on deep reinforcement learning, and relates to the technical field of wireless communication, and the method comprises the steps: 1, putting an unmanned aerial vehicle cluster in a user dense region, and enabling a fixed edge server to obtain the position coordinates of the fixed edge server, each user and each unmanned aerial vehicle; 2, a user sends a rendering request to the fixed edge server through the unmanned aerial vehicle cluster; step 3, constructing an optimization target, converting the optimization target into a Markov decision process, solving the optimization target through a PPO algorithm, and obtaining a horizontal flight rate, a flight direction and a vertical direction rate of the unmanned aerial vehicle and a computing resource set available for edge equipment; step 4, the fixed edge server issues a decision result to the user and the unmanned aerial vehicle, and the user and the unmanned aerial vehicle execute a decision; and step 5, after the unmanned aerial vehicle completes the rendering task, returning a result to the user. The method has the characteristics of multi-node collaborative optimization, energy consumption optimization and resource utilization rate optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and more specifically, the present invention relates to a method for metaverse task offloading and resource scheduling based on deep reinforcement learning. Background Art

[0002] With the growth of metaverse and immersive experience demands, users' requirements for low-latency and high-bandwidth computing resources have increased significantly. Especially for users wearing head-mounted displays, when interacting in a virtual reality or augmented reality environment, real-time rendering and response speed are crucial for the user experience. However, the computing power and battery capacity of terminal devices (such as head-mounted displays) are limited and difficult to support high-intensity computing tasks. Therefore, it is considered to use edge computing to offload computing tasks to nearby drone computing nodes to solve this problem. In recent years, the "drone-assisted network" in mobile edge computing has emerged. Drones have shown significant advantages in providing edge computing services, expanding network coverage, and improving computing power.

[0003] The drone-assisted mobile edge computing network not only has flexible mobility and can provide computing resources anytime and anywhere, but also can cooperate with other edge servers or fixed base stations through an ad-hoc network to optimize resource allocation and computing offloading. In addition, as an intelligent decision-making algorithm, deep reinforcement learning can optimize the task offloading and resource allocation strategies adaptively in a dynamic environment, thus effectively coping with problems such as dynamic resource changes, user location uncertainty, and network environment fluctuations.

[0004] Currently, there are already some related technical solutions in the fields of metaverse, edge computing, and deep learning, which involve resource allocation and offloading strategies in computing: Chinese Patent Document CN114745736A discloses a method for processing metaverse services based on a base station edge computing node. This method deploys edge computing nodes at the base station and processes metaverse service requests through the nodes, reducing the computing delay and data transmission latency.

[0005] Chinese Patent Document CN115204505A discloses an optimal resource allocation method for a vehicle-mounted edge metaverse system. It provides computing support for AR vehicles through vehicle-mounted edge computing (VEC) servers. This method uses a deep learning model to optimize computing resources and transmission power, but it focuses on the application scenario of fixed vehicle-mounted edge servers.

[0006] Chinese Patent Document CN117255418A discloses a method for fog computing resource allocation based on deep reinforcement learning. Through the proximal policy optimization algorithm of deep reinforcement learning, it realizes resource allocation and computing offloading of mobile devices in a fog computing environment. This solution has strong adaptability in dynamic resource allocation and optimizes latency and energy consumption.

[0007] Although the above research has been strengthened in edge computing, resource allocation, and deep learning algorithms, there are still the following obvious disadvantages when applied to the dynamic scenarios of the metaverse: 1. Existing edge computing systems usually rely on fixed base stations or ground nodes (such as vehicle-mounted edge servers, fixed edge computing nodes), lacking flexibility and unable to dynamically adjust the deployment of computing resources according to the user's location. There are problems with the dynamic allocation of computing resources caused by mobility. This architecture of fixed nodes is difficult to meet the needs of mobile users in the metaverse for low latency and high real-time performance, especially when users wear head-mounted displays and move frequently.

[0008] 2. Most current technologies perform resource allocation on a single node, and few solutions involve the collaborative optimization of multiple mobile nodes (such as drones). Even in the case of vehicle-mounted edge computing networks, they mainly focus on local single-node optimization and fail to effectively achieve resource sharing and intelligent scheduling among multiple nodes, resulting in unbalanced resource utilization and affecting the overall performance of the system.

[0009] 3. Traditional edge computing offloading strategies and resource allocation methods are difficult to adapt to complex and variable user mobility, network status, and computing requirements. Many existing solutions adopt allocation strategies based on static rules or traditional optimization algorithms and lack the adaptive ability to adapt to highly dynamic environments. Therefore, it is difficult to achieve efficient task offloading and resource allocation. 4. Fixed base stations and ground servers usually have high power consumption, while drones have strong mobility but limited battery life. Therefore, continuing to operate in unnecessary scenarios may lead to energy waste. Existing technologies do not fully utilize intelligent algorithms to balance computing resources and energy consumption to adapt to the high-intensity and long-duration resource requirements in the metaverse environment. Summary of the Invention

[0010] The object of the present invention is to design and develop a metaverse task offloading and resource scheduling method based on deep reinforcement learning. Through multi-node collaboration combined with the PPO algorithm, it realizes intelligent task offloading and dynamic resource scheduling, effectively balances resource utilization and energy consumption while meeting the high-quality service requirements of the metaverse, extends the flight time of drones, and improves the overall resource efficiency of the system.

[0011] The technical solution provided by the present invention is as follows: A metaverse task offloading and resource scheduling method based on deep reinforcement learning, comprising the following steps: Step 1: Deploy a drone cluster in a user-dense area. Based on the Cartesian coordinate system, the fixed edge server obtains the position coordinates of itself, each user, and each drone. Step 2: The user sends a rendering request to the fixed edge server through the drone cluster. Step 3: Construct an optimization objective, convert the optimization objective into a Markov decision process, and solve the optimization objective through the PPO algorithm to obtain the horizontal flight speed, flight direction, vertical speed of the UAV, and the set of computing resources available to the edge device; The optimization objective is: ; ; ; ; ; wherein, is the frame delay weight, is the user 's frame delay, is the weight of energy consumption, is the user 's system energy consumption, is the proportion of computing resources allocated by the edge device to the user , is the allocation coefficient between the edge device and the user , is the distance between the user or edge device and the user or edge device , is the position coordinate of the user or edge device at time, is the boundary of the movable area of the user and the UAV, is the set of users, is the set of edge devices, and the set of edge devices includes the set of UAVs and fixed edge servers; Step 4: The fixed edge server sends the decision result to the user and the UAV, and the user and the UAV execute the decision; Step 5: After the UAV completes the rendering task, it transmits the result back to the user.

[0012] Preferably, the update of the user's position coordinate satisfies: ; wherein, is the position coordinate of the UAV at time, is the horizontal speed of the UAV at time, is the UAV At the direction at the moment, , is the vertical direction rate of the drone at the moment, is from the moment to the length of the time period at the moment, is the position coordinate of the drone at the moment.

[0013] Preferably, the frame delay of the user satisfies: ; In the formula, is the rendering delay, is the transmission delay.

[0014] Preferably, the rendering delay satisfies: ; In the formula, is the computing resource required to process unit data by the edge device, is the user unloaded to the edge device frame size, is the edge device available computing resources.

[0015] Preferably, the transmission delay satisfies: ; In the formula, is the communication rate between the edge device and the user ; The communication rate between the edge device and the user satisfies: ; In the formula, is the bandwidth of the edge device , is the power of the edge device , is the channel gain coefficient, is the edge device and the user distance between, is the noise power.

[0016] Preferably, the user The system energy consumption satisfies: ; Wherein, is the transmission energy, is the rendering energy.

[0017] Preferably, the transmission energy satisfies: ; Wherein, is the transmission power of the edge device .

[0018] Preferably, the rendering energy satisfies: ; Wherein, is the power factor of the edge device.

[0019] Preferably, converting the optimization objective into a Markov decision process specifically includes: State space : ; Action space : ; Reward function : ; Wherein, is the state space at time , is the set of user and drone positions at time , is the set of power of the user and the drone, is the set of task data volumes of the user, is the set of available computing resources of the edge device, is the action space at time , is the set of drone flight actions at time , is the set of offloading decisions for the user task, is the reward function at time .

[0020] Preferably, the PPO algorithm specifically includes the following steps: Step 1. Initialize the network parameters of the policy network , the network parameters of the value network and the experience pool ; Step 2. The fixed edge server interacts with the environment through the current policy and records the trajectory data and store it in the experience pool ; wherein, is the policy function of the -th iteration, indicating the policy output by the policy network with parameters , is the iteration round, is the environmental state at ; Step 3: For each trajectory in the experience pool , calculate the cumulative return of the actual observation: ; In the formula, is the number of time steps for a single sampling, is the -th power of the step discount factor, is the -th immediate reward obtained at ; Step 4: Use the advantage estimation method to calculate the advantage function based on the state value function directly predicted by the value network: ; In the formula, is the action value function and satisfies: ; In the formula, is the discount factor, is a hyperparameter, , used to control the bias-variance trade-off, is the temporal difference error at , where represents the current moment, represents the step offset from the current moment; Step 5: Use the gradient descent method to update the network parameters of the policy network, with the goal of maximizing the objective function; The objective function is: ; In the formula, is the objective function, is the expectation, is the probability ratio of the current policy and the old policy, is the advantage function, is the clipping function, is the clipping range; The network parameters of the policy network are updated to satisfy: ; In the formula, is the size of the experience pool at the th step; Step 6, update the value function parameters using the gradient descent method : ; In the formula, is the loss function of the value network, is the value network parameter updated after the th iteration.

[0021] Advantages of the present invention: (1) A method for metaverse task offloading and resource scheduling based on deep reinforcement learning designed and developed by the present invention uses drones as mobile edge computing nodes to provide metaverse services for users wearing head-mounted displays. This drone-assisted architecture makes up for the limitations of the coverage of fixed base stations or edge servers, enables the computing resources to be dynamically adjusted as the user moves, and ensures that users can obtain low-latency services at any location. Compared with traditional fixed nodes, the solution of the present invention has more advantages in terms of service coverage and flexibility.

[0022] (2) A method for metaverse task offloading and resource scheduling based on deep reinforcement learning designed and developed by the present invention, through a multi-node collaborative computing and resource sharing mechanism among drones, fixed edge servers and user terminals, combines deep reinforcement learning for task scheduling and load balancing, can intelligently allocate tasks among multiple nodes, avoid overloading of a single node or resource idleness, and achieve efficient utilization of resources. This multi-node collaborative optimization mechanism is more systematic than existing single-node or local optimization schemes, and can still provide high-quality metaverse services in real time when the network conditions, computing resources and user requirements change, improving the overall service quality.

[0023] (3) A method for metaverse task offloading and resource scheduling based on deep reinforcement learning designed and developed by the present invention, by introducing the deep reinforcement learning (DRL) algorithm, enables the system to intelligently adapt to changes in the positions of computing nodes (such as drones and edge servers) and users, optimize the offloading of tasks and the real-time allocation of computing resources, and can achieve efficient offloading and fast transmission of tasks, improve the dynamic resource allocation ability and real-time decision-making ability, significantly reduce the computing latency at the user end, and enhance the real-time experience of metaverse services.

[0024] (4)The method for metaverse task offloading and resource scheduling based on deep reinforcement learning designed and developed by the present invention intelligently analyzes the resource utilization rate and energy consumption status of different nodes, selects the optimal offloading path and computing nodes, reduces unnecessary energy consumption, extends the flight time of the unmanned aerial vehicle (UAV), and ensures service quality at the same time. In the UAV network with limited resources, this energy consumption management method effectively reduces unnecessary energy consumption, improves the sustainability and efficiency of the system, and is superior to the traditional high-energy-consuming edge computing scheme. Description of the Drawings

[0025] Figure 1 It is a schematic flowchart of the method for metaverse task offloading and resource scheduling based on deep reinforcement learning described in the present invention.

[0026] Figure 2 It is a schematic diagram of the total reward simulation curves of the algorithm described in the present invention and the greedy algorithm.

[0027] Figure 3 It is a schematic diagram of the frame number reward simulation curves of the algorithm described in the present invention and the greedy algorithm.

[0028] Figure 4 It is a schematic diagram of the energy consumption reward simulation curves of the algorithm described in the present invention and the greedy algorithm. Detailed Embodiment

[0029] The following further elaborates on the present invention in conjunction with the drawings of the specification, so that those skilled in the art can implement it with reference to the text of the specification.

[0030] As Figure 1 shown, a method for metaverse task offloading and resource scheduling based on deep reinforcement learning provided by the present invention is implemented through a system composed of three key components: a user terminal, a UAV, and a fixed edge server. The user terminal is equipped with a head-mounted display for receiving and presenting metaverse content, and the user terminal accesses edge computing resources through a wireless network; the UAV is a UAV carrying a computing module, serving as a mobile edge computing node, responsible for providing computing resources near the user, and capable of dynamic scheduling according to the user's location and computing task requirements; the fixed edge server is a fixed computing node deployed on the ground, forming a cooperative network with the UAV to provide support for greater computing requirements.

[0031] Therefore, the method for metaverse task offloading and resource scheduling based on deep reinforcement learning described in the present invention specifically includes the following steps: Step 1: Deploy a UAV cluster in a user-dense area. Based on the Cartesian coordinate system, the fixed edge server obtains the position coordinates of itself, each user, and each UAV. Among them, the user-dense area is an area greater than 0.05 users / m². The user set is , the drone set is , the edge device includes the drone set and a fixed edge server, and the set is defined as ; The position coordinate of user at moment is , the drone at moment has a position coordinate of , and the coordinate of the fixed edge server is ; The position of the user at moment can be expressed as: ; In the formula, is the position coordinate of the drone at moment, is the horizontal rate of the drone at moment, is the direction of the drone at moment, , is the vertical rate of the drone at moment, is the time period length from moment to moment, is the position coordinate of the drone at moment The distance between drones and between drones and users should meet the safety distance: ; ; In the formula, is the distance between the drone and the drone , is the distance between the user and the drone , is the minimum safety distance; The distance between the drone and the drone meets: ; ; Wherein, is the position coordinate of the UAV at moment, is the position coordinate of the UAV at moment.

[0032] Step 2: The user sends a rendering request to the fixed edge server through the UAV cluster; Step 3: Construct an optimization objective, convert the optimization objective into a Markov decision process, and solve the optimization objective through the PPO algorithm to obtain the horizontal flight speed, flight direction, vertical direction speed of the UAV, and the set of computing resources available to the edge device; Among them, the optimization objective is: ; ; ; ; ; Wherein, is the frame delay weight, is the frame delay of the user , is the weight of the energy consumption, is the system energy consumption of the user , is the proportion of computing resources allocated by the edge device to the user , is the allocation coefficient between the edge device and the user , is the distance between the user or edge device and the user or edge device , is the position coordinate of the user or edge device at moment, is the boundary of the movable area of the user and the UAV, is the set of users, is the set of edge devices, and the set of edge devices includes the set of UAVs and the fixed edge server; The frame delay of the user consists of two parts: rendering delay and transmission delay, that is: ; Wherein, is the rendering delay, is the transmission delay; The rendering delay satisfies: ; The transmission delay satisfies: ; Wherein, is the computing resource consumed by the edge device to process unit data, is the user offloaded to the edge device frame size, is the computing resource available to the edge device , is the edge device and the user communication rate between, and satisfies: ; Wherein, is the bandwidth of the edge device , is the power of the edge device , is the channel gain coefficient, is the edge device and the user distance between, is the noise power.

[0033] The system energy consumption of the user satisfies: ; Wherein, is the transmission energy, is the rendering energy; The transmission energy satisfies: ; The rendering energy satisfies: ; Wherein, is the transmission power of the edge device , is the power factor of the edge device.

[0034] Proximal Policy Optimization (PPO) is a reinforcement learning algorithm. In the present invention, the task offloading and resource scheduling strategies are optimized through the PPO algorithm to achieve efficient real-time computing allocation in a dynamic environment.

[0035] In the PPO algorithm, the task offloading problem is modeled as a Markov decision process (MDP) to formulate the decision-making process. The MDP is composed of a five-tuple which respectively represent the state space, observation space, action space, reward function, and state transition equation.

[0036] Among them, the state space : ; In the formula, is the state space at time, is the set of user and drone positions at time, is the set of power of the user and the drone, is the set of user task data volumes, is the set of available computing resources of the edge device.

[0037] The action space : ; In the formula, is the action space at time, represents the set of drone flight actions at time, including the horizontal direction, horizontal speed, and vertical speed of flight, is the set of user task offloading decisions, which is a set of binary variables. When at time, if the user task is offloaded to the edge device , then = 1; if the user task is not offloaded to the edge device , then

[0038] The reward function : ; In the formula, is the reward function at time.

[0039] The PPO algorithm mainly includes two networks, the policy network and the value network . Among them, are the network parameters of the policy network, are the network parameters of the value network. The specific steps are as follows: Step 1. Initialize the network parameters of the policy network, the network parameters of the value network, and the experience pool Step 2: The edge server fixes the trajectory data formed by interacting with the environment according to the current policy, collecting and recording information such as the state, action, reward, and next state related to each decision and storing it in the experience pool ; Among them, is the policy function of the th iteration, indicating the policy output by the policy network with parameters , is the iteration round, is the environmental state at time; Step 3: For each trajectory in the experience pool , calculate the cumulative return of the actual observation: ; In the formula, is the number of time steps for a single sample, is the th power of the step discount factor, , is the immediate reward obtained at time; Step 4: Using the advantage estimation method, calculate the advantage function based on the state value function directly predicted by the value network: ; In the formula, is the action value function, estimated using the generalized advantage estimation method, that is: ; In the formula, is the discount factor, is a hyperparameter, , used to control the bias-variance trade-off, is the temporal difference error at time, where represents the current time, represents the step offset from the current time; The temporal difference error at time satisfies: In the formula, is the temporal difference error at the current time , is the true value function value of the state The advantage function is used to measure the superiority of the current action relative to the baseline policy and guide the optimization direction of the policy.

[0040] Step 5: Use the gradient descent method to update the network parameters of the policy network to maximize the objective function; The objective function is: ; In the formula, is the objective function, is the expectation, is the probability ratio of the current policy and the old policy, is the advantage function, is the clipping function, is the clipping range, usually taken as 0.1, which is used to limit the policy update amplitude and avoid excessive policy updates; The probability ratio of the current policy and the old policy satisfies: ; In the formula, is the probability that the current policy network selects action in state , is the probability that the old policy network selects action in state ; The clipping function will limit the update amplitude within the range of to to ensure that the policy change amplitude is controlled, thereby avoiding excessive policy fluctuations. Through this process, the policy is continuously improved in multiple iterations to achieve stable improvement of the task offloading policy and dynamic adaptation to the environment. Therefore, the clipping function satisfies: ; The network parameters of the policy network are updated to satisfy: ; Step 6: Calculate , and use the gradient descent method to update the value function parameter : ; ; In the formula, is the size of the experience pool at the th step; In the formula, is the loss function of the value network, is the value network parameter updated after the th iteration; Step 4: The fixed edge server sends the decision result to the user and the UAV. The user and the UAV execute the decision, that is, the UAV adjusts its own trajectory according to the decision result, and the user unloads the rendering task to itself or the corresponding edge device (UAV or edge server) according to the decision result; Step 5: After the UAV completes the rendering task, it sends the result back to the user.

[0041] In this embodiment, the , and the other parameter values are shown in Table 1: Table 1 Parameter Table

[0042] As Figure 2 shown, compared with the traditional greedy algorithm, the algorithm of the present invention improves the total system reward by 66.7%, indicating that this dynamic decision-making ability based on DRL enables the system to more effectively meet the real-time requirements in the complex scenarios of the metaverse, which is significantly better than the traditional solutions that rely on fixed strategies or simple optimization algorithms.

[0043] As Figure 3 shown, in the simulation result graph, compared with the traditional greedy algorithm, the algorithm of the present invention reduces the energy consumption while increasing the number of frames by 58.4%. It can be seen that the algorithm of the present invention performs task scheduling and load balancing through deep reinforcement learning, can intelligently allocate tasks among multiple nodes, avoid overloading of a single node or resource idleness, and achieve efficient utilization of resources. This multi-node collaborative optimization mechanism is more systematic than the existing single-node or local optimization solutions and improves the overall service quality.

[0044] As Figure 4 shown, in the simulation result graph, compared with the traditional greedy algorithm, the energy consumption reward of the algorithm of the present invention is reduced by 7.2%. It can be seen that in the UAV network with limited resources, this energy consumption management method effectively reduces unnecessary energy consumption, improves the sustainability and efficiency of the system, and is better than the traditional high-energy-consuming edge computing solutions.

[0045] A method for metaverse task offloading and resource scheduling based on deep reinforcement learning designed and developed by the present invention uses the DRL algorithm to comprehensively optimize the offloading path and resource allocation according to the resource availability of drones and edge servers, user locations, and network status. According to the decision results, the drones dynamically adjust their movement trajectories, can be flexibly deployed according to the user's location, and provide low-latency and high-bandwidth computing support for users wearing head-mounted displays, making up for the limitations of traditional fixed edge servers, and is particularly suitable for scenarios where users move frequently in the metaverse. At the same time, the user's computing tasks are offloaded to selected edge nodes. For example, if a drone node is selected, the task will be transmitted to a nearby drone for processing. The drone quickly completes data processing through local computing and returns the processing results to the user terminal. The system realizes the optimal configuration of computing resources through the collaborative mechanism of multiple nodes to improve the resource utilization rate, latency performance, and reliability of the system, dynamically allocates tasks among multiple drone nodes and edge servers, effectively balances the load, and avoids resource bottlenecks of a single node. The system continuously optimizes and adjusts the strategy of the DRL decision module by monitoring user experience metrics (latency), energy consumption feedback, etc. The system continuously updates its strategy and real-time selects the optimal computing offloading strategy, not only realizing dynamic offloading but also optimizing the energy consumption of the drones. By balancing the energy consumption and computing resources, the system maximizes the flight time of the drones while providing efficient computing for users, so as to adapt to the needs of long-term and frequent computing tasks, adapt to the dynamic changes of the environment and task requirements, and realize efficient metaverse services.

[0046] Although the embodiments of the present invention have been disclosed as above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the embodiments shown and described herein.

Claims

1. A method for metaverse task offloading and resource scheduling based on deep reinforcement learning, characterized in that, Including the following steps: Step 1: Deploy a drone swarm in areas with dense user populations. Based on the Cartesian coordinate system, a fixed edge server obtains the position coordinates of itself, each user, and each drone. Step 2: The users send rendering requests to the fixed edge server through the drone swarm. Step 3: Construct an optimization objective, convert the optimization objective into a Markov decision process, and solve the optimization objective through the PPO algorithm to obtain the horizontal flight speed, flight direction, vertical speed of the drones, and the set of computing resources available to the edge devices. The optimization objective is: ; ; ; ; ; Wherein, is the frame delay weight, is the frame delay of the user , is the weight of the energy consumption, is the system energy consumption of the user , is the proportion of computing resources allocated by the edge device to the user , is the allocation coefficient between the edge device and the user , is the distance between the user or edge device and the user or edge device , is the position coordinate of the user or edge device at the moment , is the boundary of the movable area of the user and the UAV, is the user set, is the edge device set, and the edge device set includes the UAV set and the fixed edge server; Step 4: The fixed edge server distributes the decision results to the users and drones, and the users and drones execute the decisions. Step 5: After the drones complete the rendering tasks, they transmit the results back to the users.

2. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 1, characterized in that, The update of the position coordinates of the users satisfies: ; Wherein, is the position coordinate of the unmanned aerial vehicle at moment; is the horizontal speed of the unmanned aerial vehicle at moment; is the direction of the unmanned aerial vehicle at moment; , is the vertical speed of the unmanned aerial vehicle at moment; is the time period length from moment to moment; is the position coordinate of the unmanned aerial vehicle at moment.

3. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 2, wherein, The user has a frame delay that satisfies: ; Wherein, is the rendering delay, is the transmission delay.

4. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 3, wherein, The rendering delay satisfies: ; Wherein, is the computing resource consumed by the edge device to process unit data, is the user offloaded to the edge device frame size, is the edge device available computing resources.

5. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 4, wherein The transmission delay satisfies: ; wherein, is the communication rate between the edge device and the user; The edge device and the user The communication rate therebetween satisfies: ; In the formula, is the bandwidth of the edge device , is the power of the edge device , is the channel gain coefficient, is the distance between the edge device and the user , is the noise power.

6. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 5, characterized in that, The user has system energy consumption that satisfies: ; In the formula, is the transmission energy, is the rendering energy.

7. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 6, wherein The transmission energy satisfies: ; In the formula, is the transmission power of the edge device .

8. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 7, wherein, The rendering energy satisfies: ; In the formula, is the power factor of the edge device.

9. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 8, wherein, Specifically, converting the optimization objective into a Markov decision process includes: State space : ; Action space : ; Reward function : ; Among them, is the state space at time is the set of user and drone positions at time is the set of power levels of the user and the drone, is the set of user task data volumes, is the set of computing resources available to the edge device, is the action space at time is the set of drone flight actions at time is the set of user task offloading decisions, is the reward function at time 10. The method for metaverse task offloading and resource scheduling based on deep reinforcement learning according to claim 9, wherein, The PPO algorithm specifically includes the following steps: Step 1, initialize the network parameters of the policy network , the network parameters of the value network and the experience pool ; Step 2: The edge server fixes and interacts with the environment through the current policy to record trajectory data and stores it in the experience pool ; Among them, is the policy function of the th iteration, indicating the policy output by the policy network with parameters , is the number of iteration rounds, is the environmental state at time; Step 3. For each trajectory in the experience pool , calculate the cumulative return of the actual observations: ; In the formula, is the number of time steps for a single sampling, is the th power of the step discount factor, is the immediate reward obtained at the moment; Step 4: Use the advantage estimation method to directly predict the state value function based on the value network Calculate the advantage function: ; In the formula, is the action value function and satisfies: ; wherein, is the discount factor, is a hyperparameter, , used to control the bias-variance trade-off, is the time of the temporal difference error, where represents the current time, represents the step offset starting from the current time; Step 5. Use the gradient descent method to update the network parameters of the policy network to maximize the objective function; The objective function is: ; In the formula, is the objective function, is the expectation, is the probability ratio of the current policy and the old policy, is the advantage function, is the clipping function, is the clipping range; Network parameters of the policy network The update satisfies: ; In the formula, is the size of the experience pool at the th step; Step 6: Update the value function parameters using the gradient descent method : ; In the formula, is the loss function of the value network, is the value network parameters updated after the -th iteration.

Citation Information

Patent Citations

  • Metacosm service processing method and device, electronic equipment and storage medium

    CN114745736A

  • Optimal resource allocation method for AR enabling vehicle-mounted edge element universe system

    CN115204505A

  • Fog computing resource allocation method based on deep reinforcement learning strategy

    CN117255418A

  • Multi-unmanned aerial vehicle air charging and task scheduling method based on deep reinforcement learning

    CN114048689A

  • Multi-device edge video analysis system based on deep reinforcement learning

    CN114170560A

Cited By

  • Edge cluster self-organization and reconstruction method based on deep reinforcement learning

    CN122119756A

  • An edge cluster self-organization and reconstruction method based on deep reinforcement learning

    CN122119756B