Multi-user wireless channel resource allocation optimization method based on deep reinforcement learning
Patent Information
- Application Number
- CN202311813162.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-12-26
AI Technical Summary
这两个问题的基本假设是:每个用户或应用程序需要分配一个独立的无线信道,以便进行通信
[0019] 1. When allocating channel resources to users, this invention takes into account the user interaction experience and the timeliness of the experience, thereby improving the overall user experience quality.
Smart Images

Figure CN117793803B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a wireless channel optimization method, specifically, a multi-user wireless channel resource allocation optimization method based on deep reinforcement learning. Background Technology
[0002] In wireless networks, due to limited spectrum resources, multiple users need to communicate using these limited resources. Therefore, it is necessary to allocate these limited frequency resources rationally to avoid resource waste and channel conflicts. Thus, channel allocation is one of the most fundamental issues in wireless networks.
[0003] Traditionally, the channel allocation problem in communication networks has been approached from two main angles: fixed-frequency channel allocation and dynamic-frequency channel allocation. The fundamental assumption of both problems is that each user or application needs to be allocated an independent wireless channel for communication.
[0004] Deep Reinforcement Learning (DRL) is a novel machine learning technique that has emerged in recent years, combining Deep Learning (DL) and Reinforcement Learning (RL). Marked by AlphaGo's victory over top human Go players, developed by DeepMind, Deep Reinforcement Learning has gained increasing attention. Unlike traditional deep learning methods—supervised and unsupervised learning—Deep Reinforcement Learning learns through real-time interaction with the environment. By sampling the environment, the agent makes decisions based on the samples, and the environment then provides rewards. Through multiple iterations, the agent learns how to handle the current environment to maximize its gains. Therefore, Deep Reinforcement Learning is suitable for decision-making problems.
[0005] CN113518039B proposes a resource optimization method and system based on deep reinforcement learning under the SDN architecture. It can meet the bandwidth requirements of the flow while considering the overall load balancing of the network as much as possible, thus achieving the effect of congestion avoidance. At the same time, when the network is about to become congested, it reroutes some flows in the network, enabling the network to further avoid congestion and achieve further optimization of network resources.
[0006] With the development of technologies such as virtual reality, augmented reality, and mixed reality, the metaverse has gradually become a platform for providing people with more immersive virtual experiences. In the metaverse, network communication technology plays a crucial role; it is not only a bridge connecting users and the virtual world, but also the foundation for realizing functions such as virtual social interaction and virtual commerce.
[0007] In a wireless communication environment with limited resources, realizing a metaverse system with multiple users interacting with each other faces the following two main challenges: First, how to maintain a metaverse system with a huge amount of information to be transmitted and processed on a wireless network; second, how to optimize the frequent interaction requests between multiple metaverse users to ensure a good user experience in an environment with scarce channel resources.
[0008] In order to solve the above problems, people have been seeking an ideal technological solution. Summary of the Invention
[0009] The purpose of this invention is to address the shortcomings of existing technologies by providing a multi-user wireless channel resource allocation optimization method based on deep reinforcement learning.
[0010] To achieve the above objectives, the technical solution adopted by this invention is: a multi-user wireless channel resource allocation optimization method based on deep reinforcement learning, comprising the following steps:
[0011] Based on deep reinforcement learning, the multi-user wireless channel resource allocation problem is described as a partially observable Markov decision process by a single agent. The agent's state space S represents the possible states of a multi-layer heterogeneous network environment, its action space A represents the wireless channel resource allocation management strategy for each representative interactive user n, and its reward function y is... In the formula, t is the time slot. For a representative set of interactive users, For a non-representative set of interactive users; QoE n,t For the representative interactive user n in time slot t, K represents the quality of the metaverse service experience. n-,t+1 W4>0 represents the tolerance of non-representative interactive users at the beginning of time slot t+1, and W4>0 is the weighting parameter for the tolerance of non-representative interactive users.
[0012] The overall environment is set up, initializing the interaction data volume of representative interactive user n, metaverse scene update information, tolerance information of representative and non-representative interactive users n, channel bandwidth, transmit power, Gaussian white noise, and channel fading; a multi-agent allocation model is constructed based on the PPO reinforcement algorithm, the multi-agent allocation model is an Actor-Critic structure, and the policy optimization objective is:
[0013]
[0014] In the formula, M is the total number of channels owned by the metaverse service provider, z is the allocation of wireless channel resources for the representative interactive user n in time slot t, and T represents the maximum time length of a round.
[0015] Using a trained agent based on all representative interactive users Wireless channel resources are allocated based on the amount of interactive data in time slot t, scene update information, and the tolerance value of each representative interactive user.
[0016] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, the computer program being executed by the processor causing the processor to execute the aforementioned multi-user wireless channel resource allocation optimization method.
[0017] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned multi-user wireless channel resource allocation optimization method.
[0018] This invention has outstanding substantive features and significant progress compared to the prior art. Specifically,
[0019] 1. When allocating channel resources to users, this invention takes into account the user interaction experience and the timeliness of the experience, thereby improving the overall user experience quality.
[0020] 2. In terms of simulation scenarios, this invention considers actual dynamic scenarios with multiple user interactions, and pays more attention to the timeliness of interactions between multiple users. Based on the above, a reward function is designed. The first half of the reward function is a component of the QOE of the representative interactive user set. I(success) in QOE represents the number of successful interactions between the representative interactive user n and other users. In addition, the definition of I(success) also shows the timeliness. Specifically, if the user's interaction timeliness is not satisfied beyond the specified time slot duration, the interaction is considered to have failed.
[0021] 3. Considering the system uncertainties such as the huge system state space and unknown system dynamics, this invention proposes a DRL method for solving the optimal control strategy of Markov process (MDP), which uses the neural network in DRL to solve the dilemma faced by traditional RL in the huge state space.
[0022] 4. The PPO algorithm used in this invention is an improvement on the traditional AC policy gradient algorithm. To address the problems of instability and slow convergence of the algorithm, as well as the problem of multi-user wireless channel resource allocation, the internal network structure of the PPO algorithm is improved and appropriate parameters are adjusted. Empirical data is obtained from the environment to further update its own strategy, which also improves the performance of the algorithm. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the three-layer network structure of the Virtual Metaverse service.
[0024] Figure 2 This is a flowchart illustrating the present invention.
[0025] Figure 3 This is a schematic diagram of the interactive communication model of the multi-layer heterogeneous network of the present invention. Detailed Implementation
[0026] Metaverse service providers (MSPs) can quickly deploy virtual metaverse services in a Layer 3 network, provided they have already leased or purchased the corresponding hardware and resources from a network infrastructure provider (NINP).
[0027] like Figure 1 As shown, the infrastructure of each layer of the network and the functional description of each layer are as follows:
[0028] The third layer of infrastructure is the central cloud, which possesses powerful data computing, storage, and data analysis capabilities. This layer is more suitable for providing users with applications such as non-real-time, long-cycle data, and massive computing.
[0029] The second layer of infrastructure consists of edge base stations (EBS), designed to bring the central cloud closer to the user. It is suitable for applications such as high-response services and extended edge analytics.
[0030] The first layer of infrastructure consists of terminal devices, including AR sensing devices, rendering servers, smartphones, and VR devices that provide interaction for users of the metaverse. This layer is the last layer for transmitting interactive information to the user and also the layer with the weakest computing power.
[0031] How to maintain a metaverse system with a huge amount of information to be transmitted and processed on a wireless network, and how to optimize frequent interaction requests between multiple metaverse users to ensure a good user experience in an environment with scarce channel resources, are urgent problems to be solved.
[0032] To address this, the present invention proposes a multi-user wireless channel resource allocation optimization method based on deep reinforcement learning, in order to solve the problems of multi-user interaction and long-term overall user experience in a multi-layer network metaverse.
[0033] The technical solution of the present invention will be further described in detail below through specific embodiments.
[0034] Example 1
[0035] This embodiment provides a multi-user wireless channel resource allocation optimization method based on deep reinforcement learning, such as... Figure 2 As shown, it includes the following steps:
[0036] Configure the parameters required for the virtual metaverse service scenario: Input the number of users N and the number of wireless channel resources M to be allocated. Simultaneously, divide the users into representative interactive user sets. Unrepresentative Interactive User Sets Representative interactive users Non-representative interactive users
[0037] The virtual metaverse service operates on discrete time slots of equal duration Δ (in seconds), each indexed by a positive integer t∈N+. Clock synchronization is maintained centrally by the cloud. The channel bandwidth B, transmit power p, Gaussian white noise χ, and channel fading f are set.
[0038] Training agents based on deep learning algorithms:
[0039] Based on deep reinforcement learning, multi-user wireless channel resource allocation is described as a partially observable Markov decision process by a single agent. The state space S of the agent represents the possible states of a multi-layer heterogeneous network environment, expressed as follows:
[0040]
[0041] In the formula, S t The state of time slot t, This indicates whether a representative interactive user n has an interactive need with the metaverse environment in time slot t. This indicates whether a representative interactive user n has an interaction need with other representative interactive users, including itself. This indicates whether the metaverse itself has an interaction need with the metaverse users it serves in time slot t; K represents the scene update information for time slot t; n,t+1 This represents the tolerance level of the representative interactive user n in time slot t+1. K represents the sum of the tolerance levels of all representative interactive users n in time slot t+1; n-,t+1 n represents non-representative interactive users - Tolerance in time slot t+1 This represents the sum of the tolerance levels of all non-representative interactive users in time slot t+1.
[0042] Action space A represents the wireless channel resource allocation and management strategy for each representative interactive user n. After observing the system state in time slot t, the agent uses its strategy π to make a decision and determine the action. The agent's action is to allocate and manage wireless channel resources for each representative interactive user n, which can be divided into allocating channel resources and not allocating channel resources.
[0043] The reward function is
[0044] In the formula, QoE n,t Representative interactive users In the metaverse service experience quality of time slot t, W4>0 is a weighted parameter for the tolerance of non-representative interactive users.
[0045] Set up the overall environment, initialize the metaverse scene update information H, the amount of interaction data D of representative interactive users n, the tolerance information n, and the tolerance of non-representative interactive users n. - Tolerance.
[0046] Specifically, at the beginning of each time slot t, the central cloud senses all representative interactive users. Interaction requirements, interaction data volume, scene update information, and tolerance value K within this time slot n,t And all non-representative interactive users Tolerance value K n-,t This provides the intelligent agent with the necessary state information.
[0047] in,
[0048]
[0049] Among them, the amount of interaction data of the representative interactive user in time slot t is:
[0050]
[0051] In the formula, 1{·} is an indicator function; it equals 1 if the condition is met, and 0 otherwise. UM Duu is the amount of data that representative interactive user n needs to process when interacting with the metaverse; Duu is the amount of data that representative interactive user n needs to process when interacting with other representative interactive users, including itself; D MU It is the amount of data that the metaverse itself needs to process for the active interaction of a representative interactive user n.
[0052] For example, if a representative interactive user n has an interaction need with the metaverse environment in time slot t, then I UM =1, satisfying Conditions, therefore For example, if the metaverse itself has no interaction needs with the representative interactive user n it serves in time slot t, then I MU =0, not satisfied Conditions, therefore
[0053] A multi-agent allocation model is constructed based on the PPO reinforcement algorithm. This model is an Actor-Critic structure and is trained with the following policy optimization objective:
[0054]
[0055] In the formula, M represents the total number of channels owned by the metaverse service provider, z represents the allocation of wireless channel resources for a representative interactive user in time slot t, and T represents the maximum time length of a round.
[0056] Specifically, the steps for setting strategy optimization goals are as follows:
[0057] Generally, in a metaverse server, in a system model where the MSP owns a channel set of M∈[1,...,M] and N metaverse users, N is usually much larger than M. To alleviate the scarcity of channel resources, there are two allocation schemes for channel resources across the entire user group, such as... Figure 3 As shown, when the number of user-generated interactive tasks is small or the number of users occupying the channel simultaneously is small, a channel allocation scheme is adopted; otherwise, a channel non-allocation scheme is adopted.
[0058] Case 1: The user receives the allocation of channel resources.
[0059] After the agent perceives the state of the current time slot t at the central cloud, it makes a decision on the channel resources owned by the MSP, that is, whether to allocate channel resources to the representative interactive user n.
[0060] Furthermore, due to the scarcity of channel resources, a single channel may be allocated to multiple representative interactive users of the metaverse. Considering the mutual interference that may occur when multiple representative interactive users n occupy the same channel, the maximum transmission rate of representative user n can be expressed as:
[0061]
[0062] Among them, W m It is the bandwidth corresponding to channel m, m∈M,p m f is the transmit power corresponding to channel m. m χ represents channel fading at channel m, v is the number of users sharing channel m in time slot t, and χ is the number of users sharing channel m in time slot t. m It is the Gaussian white noise corresponding to channel m.
[0063] A representative user n uploads their interactive task to a nearby edge server via the channel in each time slot t. The upload latency in this process is:
[0064]
[0065] After receiving an interactive task uploaded by a representative user n, the edge server needs to analyze and process the task, and provide corresponding interactive feedback based on the received interactive data. The computational processing latency at the edge server can be expressed as:
[0066]
[0067] Where ρ1 is the number of cycles required by the edge server to process each bit of interactive data, and D... H ρ2 is the amount of data generated when the metaverse scene needs to be updated, ρ2 is the number of loops required by the edge server to process each bit of data updated in the scene, and Cedge is the computing power of the edge server.
[0068] Due to limited communication resources, edge servers struggle to meet the large data volume requirements when transmitting raw image feeds to multiple representative interactive users simultaneously.
[0069] To alleviate the current shortage of channel resources, semantic extraction technology was introduced to reduce the amount of data updated in the metaverse scenario and reduce the network burden under multi-layer heterogeneous networks.
[0070] The semantically extracted metaverse scene data and the user interaction feedback data processed by the edge server will be transmitted to the rendering server at the terminal device layer via channel resources. The downlink transmission latency generated in this process is represented as follows:
[0071]
[0072] Where Dse is the amount of data after semantic extraction, Dbase is the data that must be transmitted to maintain the realism of the metaverse scene in which the user is located, and Dfeedback is the feedback data of the edge server for processing various interactive requests.
[0073] After receiving semantic data and feedback information, the rendering server at the terminal device layer reconstructs the semantically extracted data packets and restores the corresponding scene image, which is ultimately displayed on the user device. The latency incurred at the rendering server is represented as follows:
[0074]
[0075] Where k1 represents the computational resources required per bit of data when parsing the interactive feedback information of representative interactive user n in time slot t; k2 represents the computational resources required per bit of data when reconstructing the semantic data of representative interactive user n; and Crs represents the computational power of the rendering server at the terminal device layer.
[0076] Case 2: Local processing (no channel allocation).
[0077] To alleviate the scarcity of wireless channel resources, the agent will forgo allocating channel resources to a subset of users. In this case, the interaction information of the representative user n, along with updates to the metaverse scene, will be processed at the terminal device layer closest to the user. Since the representative user n does not have channel resource support, there are no uplink or downlink transmission delays as in Case 1, nor are there processing delays for various data types at the edge server; all these delays are zero.
[0078] The interaction information of the representative interactive user n and the updates of the metaverse scene will be processed at the rendering server. The processing latency is expressed as:
[0079]
[0080] ρ3 is the number of loops required per bit of the interaction data of a representative interactive user n at the rendering server; ρ4 is the number of loops required per bit of the scene update data at the rendering server.
[0081] Since the rendering server updates the metaverse scene data and processes user interaction feedback based on the interaction feedback records and metaverse scene data cached by the device itself, the number of computing resources required for processing the interaction data and scene update data at this time is different from that when processing on the edge server, that is, the number of CPU loops required per bit of data is different.
[0082] The data generated by the rendering server will undergo final decoding at the terminal device (VR headset, etc.) of the representative interactive user n. After decoding, the metaverse scene and interactive information will be presented on the user device. The decoding latency of the user device in this process is expressed as:
[0083]
[0084] Where: k3 represents the computing resources required by the rendering server to parse each bit of the interactive data of the representative interactive user n in time slot t; k4 represents the computing resources required to decode each bit of the metaverse scene data of the representative interactive user n in time slot t; and Cvr represents the computing power of the terminal device of the representative interactive user n.
[0085] Furthermore, the wireless communication resource allocation in this invention also includes decision information regarding whether various types of data are generated and processed by edge servers or at the terminal device layer. For simplicity, we generally only refer to the channel allocation of wireless communication resources, without mentioning the location of data processing, because the former includes the latter.
[0086] Based on the preceding analysis, the total latency for completing the interaction and updating the scene in time slot t can be expressed as:
[0087]
[0088] Under the constraints of interaction effectiveness and scenario authenticity, it is necessary to further analyze the impact of the interaction between representative user n and other users in each time slot on the overall experience quality.
[0089] Specifically, unreasonable channel resource allocation and potential interaction loss due to local processing will reduce the tolerance value corresponding to the number of interaction losses and scene update failures. The number of successful interactions and the number of lost interactions for the representative user n in time slot t are expressed as follows:
[0090]
[0091]
[0092] z represents the allocation of wireless channel resources for a representative interactive user n in time slot t.
[0093] Taking the number of lost interactions as an example, when z≠0, in addition to the failure of interaction requests caused by excessive system latency, it is also necessary to calculate whether the interaction requests of other metaverse users who also have channel resources regarding the representative interaction user n meet the validity constraints of the interaction; when z=0, that is, the representative interaction user n has not obtained channel resources, so the interaction requests between the representative interaction user n and other users cannot be satisfied.
[0094] The tolerance value of a representative interactive user n at the beginning of the next time slot is predicted using the following formula:
[0095]
[0096] For a representative interactive user n in time slot t, the quality of service experience (QoE) with respect to the metaverse is expressed as:
[0097]
[0098] Where w1>0, w2>0, and w3>0 are weighted parameters for the number of successful interactions, scene updates, and failed interactions to satisfy the interaction validity requirement.
[0099] In a multi-layered heterogeneous network system model, non-representative interactive users n - They do not participate in the main interactions in this scenario. However, when the timeliness of the interaction of the representative user n cannot be guaranteed, the non-representative user n... - The tolerance level for the virtual metaverse service experience will also be affected. The following formula is used to predict the non-representative interactive user n.- The tolerance value at the beginning of the next time slot is:
[0100]
[0101] The optimization goal is to optimize the experience quality of user group N. Specifically, the goals can be divided into: 1) to satisfy the effectiveness of the interaction of all representative interactive users as much as possible; 2) to ensure the authenticity of the scenario for all representative interactive users; and 3) to maintain the tolerance value of user group N at an optimal standard as much as possible.
[0102] Based on the above discussion, the channel allocation optimization problem is defined as:
[0103]
[0104] In the formula, T represents the maximum time length of a round.
[0105] The first part of the objective function (i.e., the formula) aims to satisfy the effectiveness of interactions for each representative user and the updating of the metaverse scenario as much as possible. The remaining part considers all non-representative users n in this metaverse scenario. - The objective function aims to improve the user experience quality and maximize it as much as possible. The only constraint on the objective function is the integer optimization variable, representing the user's allocation of wireless channel resources in time slot t and the location selection for various data update processing.
[0106] The model training steps are as follows:
[0107] Specifically, based on the amount of interaction data of representative interactive user n, metaverse scene update information, and tolerance K... n,t1 unrepresentative interactive users n - Tolerance K n-,t The initial values form the initial state S1, and the agent obtains the set of observations S. At time t, the environmental state S is... t For the agent, action a is determined based on the reward function and the policy function. t Obtain a new environment S t+1 And the reward for this round y t The trajectory v = {S, a, y} generated in the current time slot is stored in the replay experience pool E to improve the performance of algorithm policy updates.
[0108] L groups of samples are randomly selected from the experience replay area E. The state parameters are input into the Actor network to obtain the action strategy for channel resource allocation. The action is then input into the Critic network to obtain the reward value y, and the network parameters of the Critic are updated.
[0109] Repeat the above steps to train continuously, stopping training when the set number of rounds is reached. After training, the average reward value after convergence, the number of successful interactions, and the remaining sum of tolerance values at the end of each round can be obtained. As needed, it can be applied to actual metaverse interaction scenarios (classrooms, conference rooms, etc.) to achieve good multi-user wireless channel resource allocation, ensure the effectiveness of interaction, and ensure the timeliness of metaverse scene updates.
[0110] Wireless channel resource allocation:
[0111] Using a trained agent based on all Wireless channel resources are allocated based on the amount of interactive data in time slot t, scene update information, and the tolerance value of each representative interactive user.
[0112] Example 2
[0113] This embodiment provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the aforementioned optimization method.
[0114] Example 3
[0115] This embodiment provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the aforementioned optimization method.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.
Claims
1. A multi-user wireless channel resource allocation optimization method based on deep reinforcement learning, characterized in that, Includes the following steps: Based on deep reinforcement learning, the multi-user wireless channel resource allocation problem is described as a partially observable Markov decision process by a single agent. The agent's state space S represents the possible states of a multi-layer heterogeneous network environment, its action space A represents the wireless channel resource allocation management strategy for each representative interactive user, and its reward function y is... In the formula, t is the time slot. For a representative set of interactive users, This is a non-representative set of interactive users; Representative interactive users The quality of service experience in the metaverse within time slot t. For non-representative interactive users At the beginning of time slot t+1, w4>0 is the weighted parameter for the tolerance of non-representative interactive user n_. The overall environment is set up, initializing metaverse scene update information, interaction data volume and tolerance information of representative interactive users, and tolerance information of non-representative interactive users; a multi-agent allocation model is constructed based on the PPO reinforcement algorithm. The multi-agent allocation model is an Actor-Critic structure, trained with the following policy optimization objective: ; In the formula, M is the total number of channels owned by the metaverse service provider, z is the allocation of wireless channel resources for a representative interactive user in time slot t, and T represents the maximum time length of a round. Using a trained agent, based on all representative interactive users ∈ Wireless channel resources are allocated based on the amount of interactive data in time slot t, scene update information, and the tolerance value of each representative interactive user. The tolerance of the non-representative interactive users in time slot t+1 for: ; The quality of the metaverse service experience ; In the formula, This represents the tolerance value of a representative interactive user in time slot t; The number of successful interactions for representative users in time slot t; The number of interaction loss for a representative interactive user in time slot t; The total latency for completing the interaction and updating the scene in time slot t; The latency for a representative interactive user to upload their interactive task to a nearby edge server via the channel in each time slot t; The latency for edge servers to analyze and process interaction tasks sent from the central cloud regarding representative interactive users, and to make corresponding interaction feedback based on the received interaction data; To extract semantics from the metaverse scene data, and to transmit the semantically extracted metaverse scene data and the interactive feedback data of the corresponding representative interactive users processed by the edge server to the rendering server of the terminal device layer through channel resources, the downlink latency will be reduced. This refers to the latency incurred by the rendering server at the terminal device layer when it receives semantic data and feedback information, reconstructs the data packets after semantic extraction, and restores the corresponding scene. The latency of the rendering processor processing interactive information of representative interactive users and updating the metaverse scene when no channel is allocated; The latency caused by the final decoding operation of the data generated by the rendering server on the terminal device of the representative interactive user; The maximum transmission rate for representative interactive users; in, This represents the tolerance value of non-representative interactive users at the beginning of time slot t. Let W be the amount of interactive data from a representative user in each time slot, p be the transmit power of channel m, f be the channel fading at channel m, and v be the number of users sharing channel m in time slot t. It is the Gaussian white noise corresponding to channel m; D is the number of cycles required for the edge server to process each bit of interactive data. H This refers to the amount of data generated when the metaverse scenario needs to be updated. Cedge represents the number of loops required by the edge server to process each bit of data for scene updates; Dse represents the amount of data after semantic extraction; Dbase represents the data that must be transmitted to maintain the realism of the metaverse scene in which the representative interactive user is located; Dfeedback represents the processing feedback data of the edge server for various interactive requests; k1 represents the computing resources required per bit of data when parsing the interactive feedback information of the representative interactive user in time slot t; k2 represents the computing resources required per bit of data when reconstructing the semantic data of the representative interactive user; and Crs represents the computing power of the rendering server at the terminal device layer. It refers to the number of loops required to analyze and process each bit of user interaction data at the rendering server. k1 represents the number of loops required to analyze and process each bit of scene update data at the rendering server; k2 represents the computational resources required by the rendering server to parse each bit of interaction data of a representative interactive user in time slot t; k3 represents the computational resources required to decode each bit of metaverse scene data of a representative interactive user in time slot t; Cvr represents the computing power of the user device; w1>0, w2>0, and w3>0 are weighted parameters for satisfying the interaction validity requirements, including the number of successful interactions, scene updates, and failed interactions.
2. The method for optimizing multi-user wireless channel resource allocation based on deep reinforcement learning according to claim 1, characterized in that: At the start of time slot t, the central cloud senses the amount of interaction data, scene update information, and tolerance value of all representative interactive users in that time slot. And the tolerance values of all non-representative interactive users. ; in, ; ; Among them, the amount of interaction data of the representative interactive user in time slot t is: ; In the formula, 1{•} is an indicator function; it equals 1 if the condition is met, and 0 otherwise. UM This indicates whether a representative interactive user has an interactive need regarding their metaverse environment in time slot t; I UU Does the representative interactive user in time slot t have an interaction need with other metaverse users, including themselves? MU Does the metaverse itself have an interaction need with the metaverse users it serves in time slot t? UM Duu is the amount of data that needs to be processed when a representative interactive user interacts with the metaverse; Duu is the amount of data that needs to be processed when a representative interactive user interacts with another representative interactive user; D MU It is the amount of data that the metaverse itself needs to process for representative interactive user active interactions.
3. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs an operation to implement the multi-user wireless channel resource allocation optimization method described in any one of 1 to 2.
4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multi-user wireless channel resource allocation optimization method described in any one of 1 to 2.
Citation Information
Patent Citations
Resource optimization methods and systems based on deep reinforcement learning under SDN architecture
CN113518039B