A Wireless Resource Allocation Method Based on Deep Reinforcement Learning (DQN) Algorithm under 5G Standard

By combining the deep reinforcement learning DQN algorithm with two-layer coding technology, the allocation of wireless resources is optimized, which solves the complexity problem of 5G networks in multi-user scenarios and maximizes the quality of panoramic video experience and improves resource utilization.

CN116939832BActive Publication Date: 2026-06-30FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUDAN UNIVERSITY
Filing Date
2023-07-07
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

The existing wireless resource allocation methods under the 5G standard fail to effectively combine dual-layer coding technology with the extended time-frequency frame structure under the 5G standard, which cannot meet the heterogeneous needs of multi-user scenarios, resulting in high network complexity and failing to maximize the sum of user experience quality within the system.

Method used

By employing the deep reinforcement learning (DQN) algorithm and combining it with two-layer coding technology, we model wireless resources and user experience quality, design a reasonable resource allocation order, and optimize resource allocation through a neural network architecture to maximize the overall quality of panoramic video experience.

Benefits of technology

While meeting basic user experience quality requirements, the system improved resource utilization, maximized panoramic video experience quality, and enhanced the scalability and usability of resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116939832B_ABST
    Figure CN116939832B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of 5G communication technology, specifically a wireless resource allocation method based on the deep reinforcement learning (DQN) algorithm under the 5G standard. The invention includes: combining wireless resources with dual-layer coding technology to model the panoramic video experience quality of individual users; fully considering user experience quality requirements and the heterogeneity of user channel states to determine the order of user resource allocation; modeling state information and user information; designing a neural network architecture and combining it with the deep reinforcement learning (DQN) algorithm to allocate appropriate optional parameter sets and minimum time slots to each user, thereby maximizing the overall panoramic video experience quality while meeting the basic experience quality requirements of all users. This invention can provide higher scalability and practicality for wireless resource allocation at the physical layer level, improve the utilization rate of limited communication resources, and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of 5G communication technology, specifically relating to a wireless resource allocation method based on the deep reinforcement learning (DQN) algorithm under the 5G standard. Background Technology

[0002] With the development of communication technology, the realization and application of immersive communication is expected to become a reality. The purpose of immersive communication technology is to enable interactive communication and exchange between remote participants and their surrounding environment, bridging the gap between the physical and virtual worlds. Panoramic video transmission, as an important component of immersive communication technology, is receiving increasing attention from academia and industry. Currently, commonly used panoramic video transmission technologies include chunked transmission and dual-layer coding. These technologies divide panoramic video into multiple video blocks according to regions, transmitting only video blocks within the user's area of ​​interest. Furthermore, video blocks within the same region can be encoded into base layer and enhancement layer video blocks using different bitrates, saving bandwidth while meeting diverse user requirements and making full use of limited communication resources.

[0003] Currently, 5G standards allow for flexible selection of optional parameter sets and minimum time slots, providing greater scalability and practicality for wireless resource allocation at the physical layer level. Common allocation methods include mixed-integer programming, random forests and other machine learning methods, as well as some classic deep learning methods.

[0004] With the development of the Internet and the exponential growth of the number of users, the issue of resource allocation for panoramic video transmission in multi-user scenarios has gradually attracted public attention. However, existing research has not combined dual-layer coding technology with the extended time-frequency frame structure design under the 5G standard for multi-user resource allocation. Furthermore, existing solutions have not combined the optional parameter set with the minimum time slot for allocation, and most assume that each user has only one optional parameter set setting, without considering that users can have multiple different optional parameter sets at the same time, ignoring the heterogeneity of user devices. At the same time, due to the high flexibility of heterogeneous user networks, the network is more complex, so traditional resource allocation methods can no longer meet the requirements, and it is necessary to consider using deep learning methods. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a wireless resource allocation method based on the deep reinforcement learning (DQN) algorithm under the 5G standard, which maximizes the sum of the experience quality of all users in the system, so as to solve the shortcomings of the existing wireless resource allocation technology under the 5G standard, and provide higher scalability and practicality for wireless resource allocation from the physical layer level.

[0006] Here, DQN is an algorithm in reinforcement learning, short for Deep Q-Network. DQN combines Q-learning algorithm and deep neural network to solve reinforcement learning problems with high-dimensional state spaces. Q-learning is a reinforcement learning algorithm whose core idea is to make decisions by learning a Q-value function. However, when the action or state space is large, a large amount of training data is needed to fully explore and learn the Q-value of each state-action pair. The core idea of ​​the DQN algorithm is to use a deep learning neural network to approximate the Q-value function based on the Q-learning algorithm. The Q-value function is used to evaluate the value of taking different actions in a given state and guide the agent to make decisions. The trained neural network can quickly output the Q-value of the input state-action pair without further exploration and learning. (Liu Quan, Zhai Jianwei, Zhang Zongchang, et al. A review of deep reinforcement learning [J]. Chinese Journal of Computers, 2018, 41(01):1-27.)

[0007] The wireless resource allocation method based on the deep reinforcement learning (DQN) algorithm under the 5G standard provided by this invention includes the following specific steps:

[0008] (i) First, wireless resources are combined with dual-layer coding technology to model the panoramic video experience quality of a single user.

[0009] (ii) Secondly, fully consider the user experience quality requirements and the heterogeneity of user channel status to determine the order of user resource allocation;

[0010] (iii) Then, the status information and user information are modeled;

[0011] (iv) Finally, the neural network architecture is designed, and the deep reinforcement learning DQN algorithm is combined to allocate a suitable set of optional parameters and a minimum time slot to each user, so as to maximize the overall panoramic video experience quality while meeting the basic experience quality requirements of all users.

[0012] The specific steps are as follows:

[0013] In step (i), the specific process of modeling the panoramic video experience quality for a single user is as follows:

[0014] Traditional panoramic video experience quality formulas generally correlate experience quality Q with normalized transmission rate. The relationship between them is described as a logarithmic relationship, and the formula for the quality of experience for a single user is:

[0015]

[0016] Here, coefficients a and b depend on the panoramic video content requested by the user; because the field of view covered by the base layer block and the enhancement layer block are generally different, the transmission rate R needs to be normalized. Where C is the field of view covered by the video block.

[0017] Dual-layer 360-degree panoramic video encoding divides the panoramic video into two layers: a base layer and an enhancement layer. The base layer encodes the entire 360-degree panoramic image at a low bitrate to provide basic quality. The enhancement layer covers a portion of the panoramic image and encodes that portion of the field of view at a high bitrate to provide a higher-quality panoramic viewing experience. The quality formulas for the base layer and the enhancement layer are as follows:

[0018]

[0019]

[0020] Among them, C BT and C ET G represents the video coverage of the base layer and enhancement layer, respectively. t and G r These represent the antenna gains at the transmitting and receiving ends, respectively. and p represents the squared power gain of the large-scale fading channel in the user's base layer and enhancement layer, respectively. n δ represents the downlink channel transmit power spectral density allocated by the base station to this user. 2 This represents the power spectral density of the noise.

[0021] Combining two-layer coding techniques, and considering the accuracy q of the enhanced layer's view prediction. n The formula for the panoramic video experience quality of the nth user is obtained as follows:

[0022]

[0023] Each user has their own basic experience quality requirements. The ultimate goal is, given the above parameters, to utilize deep reinforcement learning methods to optimize the basic layer wireless time-frequency resources available to each user. and enhanced layer wireless time and frequency resources By allocating resources appropriately, the overall quality of the panoramic video experience for all users within the system can be maximized.

[0024] The allocation of radio time-frequency resources (REs) is achieved by selecting a suitable set of optional parameters that conform to standards and a minimum time slot; according to existing 5G standards, the set of optional parameters is defined as Ω. μ The set of minimum time slots is defined as Ω. s If μ is a set Ωμ If the element is in the middle, then its corresponding subcarrier spacing is 2. μ ×15kHz; therefore, μ is defined as... max Ω μ The maximum value in the range corresponds to the subcarrier spacing. Define μ min =0 is Ω μ The minimum value in the range is defined as a subcarrier spacing of 15kHz. A physical resource block is a resource space with certain resources. Since a physical resource block contains 12 consecutive subcarriers, the frequency domain size occupied by a physical resource block is... Denoteed as the unit frequency domain; the time domain is divided into equal intervals. Therefore, the time domain size occupied by a physical resource block is [segment / unit]. This is denoted as a unit time domain; the amount of time-frequency resources occupied by a unit of physical resource is the product of the unit frequency domain and the unit time domain:

[0025]

[0026] The set of optional parameters and the minimum time slot value allocated to a user in each round will determine the number of physical resource blocks they receive, thereby determining the size of the radio time and frequency resources allocated to them.

[0027] In step (ii), the user resource allocation order is determined, and the specific process is as follows:

[0028] (1) Evaluate the channel quality of each user and rank the users from high to low according to the channel quality;

[0029] Specific evaluation method: First, divide all wireless resources equally according to the frequency domain, allocate all time domain resources to the basic layer, substitute them into the experience quality formula (2) to obtain the basic panoramic video experience quality of each user, and determine the order of user resource allocation from high to low.

[0030] (2) First provide services to users with the best channel quality, and then provide services to users whose channel quality gradually deteriorates in sequence;

[0031] (3) Repeat step (2) until communication resources cannot be allocated.

[0032] In step (iii), the modeling of state information and user information includes three parts: state design, action design, and reward design.

[0033] (1) State Design

[0034] First, the wireless resources are considered as a whole as a two-dimensional plane, with the horizontal axis representing the time domain and the vertical axis representing the frequency domain. The space is divided into time-frequency resource blocks of equal size, and then each resource block is numbered according to its coordinate position. A two-dimensional matrix L is used to record the allocation status of the time-frequency resource blocks. If L(m,n) = 0, it means that the resource block numbered (m,n) has not been allocated yet; if L(m,n) = 1, it means that the resource block numbered (m,n) has been allocated.

[0035] User information is recorded using a two-dimensional matrix O of size (N,5), where N represents the number of users. Each user contains 5 feature values, which record the following information: the number of basic layer resource blocks allocated to the user, the number of enhancement layer resource blocks allocated to the user, the ratio of the user's achieved experience quality to its basic experience quality requirements, the number of users other than the user who have not met the basic experience quality requirements, and the minimum experience quality achieved by other users other than the user.

[0036] When determining the location of resource blocks in each round of allocation, if the location of each user's resource block is added to the state, the action should also provide the location information of the resource block allocated to the user in this round. This would result in an excessively large action space, affecting the accuracy of decision-making and hindering convergence. To solve this problem, in each round of allocation, the total resource space allocation is directly obtained from the current state, and the starting position of the resource blocks for the next round of allocation is determined accordingly. Specifically, the location of the available resource blocks for the next round is added to the state returned from the current round. That is, after the action of the current round is completed, the search is performed from the starting position of the current round, first in the x-direction and then in the y-direction, to find the location of the first unallocated available resource block. This location is the starting position (x, y) of the resource blocks for the next round of allocation.

[0037] A two-dimensional vector, cur_location, is used to record the starting position of each round of resource allocation. After the current round of action is completed, cur_location is modified to the starting position of the next round and returned in the state. This is mainly because the resource space information and the feature information of each user device have been considered in the state. Moreover, the optimization goal is to maximize the overall system experience quality while meeting the basic experience quality requirements of all user devices, thereby maximizing the utilization of system resources as much as possible.

[0038] The advantage of the method proposed in this invention is that it only needs to focus on the utilization rate of resource blocks, without depending on the specific resource block location information of specific user devices. Therefore, the final state vector contains (L, O, cur_user, cur_location), where cur_user records the currently allocated user ID. However, it should be noted that cur_location and cur_user are only for data processing convenience during algorithm iteration and are not used as input to the neural network; the neural network only has (L, O) as input.

[0039] (2) Action Design

[0040] Each round of action allocation mainly includes information on the optional parameter set and the minimum time slot. Assume the size of the optional parameter set is A. μ The size of the minimum time slot set is A s The spatial size of the action is denoted as A. μ ×A s Once the set of optional parameters and the minimum time slot are determined for each user, time-frequency resource blocks can be allocated to them based on the values ​​of the optional parameter set and the minimum time slot, thereby affecting the quality of the user's panoramic video experience.

[0041] (3) Reward Design

[0042] An episode is defined as a round in the training process of a deep reinforcement learning algorithm. An agent is the main body that performs actions in the algorithm model and is responsible for interacting with the environment to obtain rewards. Specifically, the agent is a 5G base station that allocates wireless resources, the action is the set of optional parameters and the minimum time slot allocated to a user, and the environment is the allocation of wireless time and frequency resources.

[0043] Within an episode, the agent will receive a reward according to the following rules:

[0044]

[0045] First, define the conditions for ending the episode:

[0046] (1) When all resource blocks in the resource space have been allocated;

[0047] (2) There are still unallocated resource blocks, but no selectable actions;

[0048] (3) When the maximum number of iterations has been reached.

[0049] Once this episode ends, the rewards will be stored, and the next episode will begin.

[0050] If the basic experience quality requirements of all users in the system are met at the exact end of the episode, the reward is K, which is generally a positive value and serves as an incentive; otherwise, the reward is Z2, which is generally a negative value and serves as a penalty. This represents the quality of experience for the nth user at step t within an episode. This represents the quality of experience for the nth user at step t-1 within an episode, expressed as... The part representing improvement and enhancement is represented by Z1, which exists as a penalty term. The relationship between these two terms is balanced by different values ​​of α.

[0051] In step (iv), the neural network architecture is designed to maximize the overall quality of the panoramic video experience. Specifically, because the dimensions of resource space information and user information differ significantly, the neural network design needs to balance this dimensionality difference. This invention uses two independent convolutional layers to process resource space and user information respectively. Since the dimension of resource space is much larger than the dimension of user device information, a fully connected layer is added after the convolutional layer processing resource space information to reduce the dimensionality of the resource space information, aiming to make the resource space information have a similar dimension to the user device information. For the specific neural network architecture, see [link to specific details]. Figure 2 The main processing flow is as follows:

[0052] (1) Resource spatial information processing flow: The L matrix that records the allocation of resource blocks is used as the first input. After processing by convolution and pooling layers, the Flatten operation is performed, and then it goes through a fully connected layer.

[0053] (2) User information processing flow: The O matrix that records user-related information is used as the second input. After being processed by the convolutional layer, the Flatten operation is performed.

[0054] (3) After adding the two results obtained from the above processing together, the result is passed through a fully connected layer and finally output.

[0055] After completing state space modeling and neural network design, the DQN algorithm can be used to allocate wireless time-frequency resource blocks, thereby maximizing the overall panoramic video experience quality for all users in the entire system. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of the optional parameter set and minimum time slot allocation algorithm based on the DQN algorithm under the 5G standard of the present invention.

[0057] Figure 2 This is a neural network architecture diagram of the optional parameter set and minimum time slot allocation algorithm based on the DQN algorithm under the 5G standard of this invention. Detailed Implementation

[0058] Specific embodiments of the present invention will be given below, but it should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.

[0059] This invention provides an optional parameter set and a minimum time slot allocation algorithm based on the deep reinforcement learning (DQN) algorithm under the 5G standard. Specifically, it includes the following steps:

[0060] Step 1: Modeling the panoramic video experience quality for a single user;

[0061] Step two: Determine the order of user resource allocation;

[0062] Step 3, state-space modeling;

[0063] Step 4, Neural Network Design.

[0064] In this embodiment, step one is specifically as follows:

[0065] Consider a multi-user downlink panoramic video transmission system. A base station provides panoramic video wireless resources to N different user terminals in a cell. The goal is to maximize the overall system experience quality while meeting the basic experience quality requirements of all users. The system consists of one base station and N user terminals, each with its own basic experience quality requirements. The set of user terminals is defined as N = {1, ..., N}; in this embodiment, N is 4, meaning there are 4 users in the system; the communication distance d from the base station to the nth user is... n Given a uniform distribution within a 200m radius of the cell, the probability q of successfully predicting and transmitting enhancement layer resource blocks is as follows. n They follow a truncated normal distribution with a cutoff interval of [0.6, 1], and the mean and variance of the original normal probability density function are 0.8 and 0.49, respectively; all users satisfy a n =0, b n =1, the transmission power is equally distributed among all users. Other parameter settings in this embodiment are as follows:

[0066]

[0067] Once the above parameters are determined, the panoramic video experience quality for each user can be modeled using the experience quality formula.

[0068] In this embodiment, step two is specifically as follows:

[0069] In this example, the time domain size T = 1ms and the system bandwidth B = 4.32GHz. The bandwidth is equally divided among 4 users and substituted into the experience quality formula modeled in step one. The 4 users are sorted from highest to lowest experience quality. In each step, a user is selected for allocation according to the order, and the process is repeated until the end condition of the episode is met.

[0070] In this embodiment, step three is specifically as follows:

[0071] The time-domain size of a physical resource block is The frequency domain size is 12.15kHz = 180kHz. Based on T and B given in step two, a two-dimensional matrix L of size (56, 24) can be used to record the allocation of time-frequency resource blocks. Initially, all values ​​in matrix L are set to 0.

[0072] A two-dimensional matrix O of size (4,5) is used to record user information. There are a total of 4 users, and their basic experience quality requirements are... Each user has 5 feature values, recording the following information: the number of base layer resource blocks allocated to the user, the number of enhancement layer resource blocks allocated to the user, the ratio of the user's achieved experience quality to the basic experience quality requirement, the number of users other than the user who have not met the basic experience quality requirement, and the minimum experience quality achieved by other users other than the user. Initially, the fourth column of the O matrix is ​​all normalized to 1, and the rest are all set to 0.

[0073] A two-dimensional vector `cur_location` is used to record the starting position of each round of resource allocation. After the current round of actions is completed, `cur_location` is modified to the starting position of the next round, and the result is returned in the status. Initially, `cur_location` is set to (0,0).

[0074] The final state vector contains (L, O, cur_user, cur_location), where cur_user records the currently assigned user ID. In the initial state, cur_user is the first element of the user sequence.

[0075] (2) Action Design

[0076] Because the optional parameter set and minimum time slots need to be allocated at the physical layer level, according to the existing 5G standard, the set of optional parameters is defined as Ωμ={0,1,2}, and the set of minimum time slots is defined as Ω. s ={2,4,7}. Therefore, the space size for the action is 3×3.

[0077] (3) Reward Design

[0078] Given α = 0.5, Z1 = -0.01, K = 500, Z2 = -2, in one episode, the agent will receive a reward according to the following rules:

[0079]

[0080] In this embodiment, step four is specifically as follows:

[0081] Specific parameter settings in the neural network: For the first convolutional layer, in_channels=1, out_channels=1, kernel_size=(5,3), stride=(1,1), padding=(2,1); for the second convolutional layer, in_channels=1, out_channels=6, kernel_size=(3,3), stride=(1,1); for the pooling layer, kernel_size=(3,1), stride=(3,1); both fully connected layers have 64 nodes; the output layer has 9 nodes; all activation functions are ReLU functions; the Adam optimizer is used, and the learning rate is set to 0.001.

[0082] The remaining parameter settings for the DQN algorithm are: episode = 510, total number of steps in each episode time_step = 1000, and the parameters of the Q′ network are reset every 10 episodes. network.

[0083] Once the above parameters are determined, the neural network can be trained using the DQN algorithm, eventually reaching convergence in about 200 rounds.

[0084] In summary, the optional parameter set and minimum time slot allocation algorithm based on deep reinforcement learning (DQN) under the 5G standard of this invention, through reasonable design of state variables, reward functions, and neural network architecture, can achieve a higher overall system experience quality while satisfying as many basic user requirements as possible, and make resource allocation more reasonable and effective.

[0085] The above are merely preferred embodiments of the present invention. For those skilled in the art, it will be understood that any modifications, substitutions, and variations made without departing from the principles and spirit of the present invention should be included within the scope of protection of the present invention.

Claims

1. A wireless resource allocation method based on the deep reinforcement learning (DQN) algorithm under the 5G standard, characterized in that, The specific steps are as follows: (a) First, wireless resources are combined with dual-layer coding technology to model the panoramic video experience quality of a single user; (ii) Secondly, fully consider the user experience quality requirements and the heterogeneity of user channel status to determine the order of user resource allocation; (iii) Then, the status information and user information are modeled; (iv) Finally, the neural network architecture is designed, and the deep reinforcement learning DQN algorithm is combined to allocate a suitable set of optional parameters and a minimum time slot to each user, so as to maximize the overall panoramic video experience quality while meeting the basic experience quality requirements of all users. Step (iii) involves modeling state and user information, which includes three parts: state design, action design, and reward design. (1) State design First, the wireless resources are considered as a whole in a two-dimensional plane, with the horizontal axis representing the time domain and the vertical axis representing the frequency domain. This plane is divided into equal-sized time-frequency resource blocks, and each resource block is then numbered according to its coordinate position. A two-dimensional matrix is ​​then used to represent this. To record the allocation of time-frequency resource blocks; if The description number is The resource blocks have not yet been allocated, if The description number is The resource blocks have been allocated; Use a size of Two-dimensional matrix To record user information; among which This represents the number of users. Each user has 5 characteristic values: the number of basic layer resource blocks allocated to the user, the number of enhancement layer resource blocks allocated to the user, the ratio of the user's own achieved experience quality to its basic experience quality requirements, the number of users other than the user who have not met the basic experience quality requirements, and the minimum experience quality achieved by other users other than the user. When determining the location of resource blocks in each round of allocation, the total resource space allocation is directly obtained from the current state, and the starting position of the resource blocks for the next round of allocation is determined accordingly. Specifically, the available resource block locations for the next round are added to the returned state from the current round. That is, after the current round's actions are completed, the allocation is started from the current round's starting position and proceeded according to the previous round's allocation. Direction, back The search proceeds in the specified direction to find the location of the first unallocated available resource block; this location will be the starting point for the next round of resource block allocation. ; Use a two-dimensional vector It specifically records the starting position of each round of resource allocation, and will update the data after the current round of actions is completed. Modify it to the starting position of the next round and return it in the status; The final state vector contains ,in Record the currently assigned user ID; however, please note that... and This is only for easier data processing during algorithm iteration; it is not used as input to the neural network. The input to the neural network is only... ; (2) Action design Each round of allocation mainly includes information on the optional parameter set and the minimum time slot. Assume the size of the optional parameter set is... The size of the minimum time slot set is The spatial size of the action is denoted as Once the set of optional parameters and the minimum time slot are determined for each user, the time-frequency resource blocks can be allocated to them based on the values ​​of the optional parameter set and the minimum time slot, thereby affecting the quality of the user's panoramic video experience. (3) Reward Design definition As one round in the training process of a deep reinforcement learning algorithm, the agent is the main body that performs actions in the algorithm model and is responsible for interacting with the environment in the algorithm to obtain rewards. Specifically, the agent is a 5G base station that allocates wireless resources, the action is the set of optional parameters and the minimum time slot allocated to a certain user, and the environment is the allocation of wireless time and frequency resources. In one In this process, the intelligent agent will receive rewards according to the following rules: , (6) Firstly, the regulations Termination condition: (1) When all resource blocks in the resource space have been allocated; (2) There are still unallocated resource blocks, but no selectable actions; (3) When the maximum number of iterations has been reached; This round Once finished, store the rewards and start the next round. ; exist If the system ends at the exact moment, and all users' basic experience quality requirements are met, the reward is K, where K is a positive value and serves as an incentive; otherwise, the reward is... , It is a negative value, which serves as a penalty; Indicates the first One user in The first The quality of the experience while walking, Indicates the first One user in The first The quality of the walking experience is determined by... To indicate the parts that have been improved or enhanced. It exists as a penalty item, through Different values ​​are used to balance the relationship between these two items.

2. The wireless resource allocation method according to claim 1, characterized in that, The specific process for modeling the panoramic video experience quality for a single user, as described in step (I), is as follows: The panoramic video experience quality formula will determine the experience quality. With normalized transmission rate The relationship between them is described as a logarithmic relationship, and the formula for the quality of experience for a single user is: , (1) Among them, coefficient and Depending on the panoramic video content requested by the user; , which is the normalized transmission rate. It refers to the field of view covered by the video block; Panoramic video is encoded in two layers: a base layer and an enhancement layer, a process known as dual-layer 360-degree panoramic video encoding. The base layer uses a low bitrate to encode the entire 360-degree panoramic image to provide basic quality. The enhancement layer covers a portion of the panoramic image and uses a high bitrate to encode that portion of the field of view, providing a higher-quality panoramic viewing experience. The quality formulas for the base layer and the enhancement layer are as follows: , (2) , (3) in, and These represent the video coverage areas of the base layer and the enhancement layer, respectively. and These represent the antenna gains at the transmitting and receiving ends, respectively. and These represent the squares of the large-scale fading channel power gain for the user's base layer and enhancement layer, respectively. This represents the downlink channel transmit power spectral density allocated by the base station to this user. Represents the power spectral density of noise; Combining two-layer coding techniques, and considering the accuracy of enhanced layer view prediction. , obtained the Formula for panoramic video experience quality for individual users: , (4) Each user has their own basic experience quality requirements. The ultimate goal is to utilize deep reinforcement learning methods to optimize the basic layer of wireless time-frequency resources available to each user, given the available parameters. and enhanced layer wireless time and frequency resources Reasonable allocation is made to maximize the overall panoramic video experience quality for all users within the entire system; wireless time and frequency resources The allocation is achieved by selecting a suitable set of optional parameters that conforms to the standards and the minimum time slot; according to the 5G standard, the set of optional parameters is defined as... The set of minimum time slots is defined as ;like For set If the element is in the middle, then its corresponding subcarrier spacing is... Therefore, the definition is... for The maximum value in the range corresponds to the subcarrier spacing. ;definition for The minimum value in the range, the corresponding subcarrier spacing is defined as... A physical resource block is a resource space containing certain resources. Since a physical resource block contains 12 consecutive subcarriers, the frequency domain size occupied by a physical resource block is... , denoted as the unit frequency domain; the time domain is divided into equal intervals. A segment, the time domain size occupied by a physical resource block is denoted as unit time domain; the size of the time-frequency resources occupied by a unit physical resource is the product of the unit frequency domain and the unit time domain: , (5) The set of optional parameters and the minimum time slot value allocated to a user in each round will determine the number of physical resource blocks they receive, thereby determining the size of the radio time and frequency resources allocated to them.

3. The wireless resource allocation method according to claim 2, characterized in that, The specific process for determining the order of user resource allocation in step (II) is as follows: (1) Evaluate the channel quality of each user and rank the users from high to low according to the channel quality; Specific evaluation method: First, divide all wireless resources equally in the frequency domain and allocate all time domain resources to the basic layer. Substitute them into the experience quality formula (2) to obtain the basic panoramic video experience quality of each user. Determine the order of user resource allocation from high to low. (2) First provide services to users with the best channel quality, and then provide services to users whose channel quality gradually deteriorates in sequence; (3) Repeat step (2) until communication resources cannot be allocated.

4. The wireless resource allocation method according to claim 1, characterized in that, Step (iv) describes the design of the neural network architecture to maximize the overall quality of the panoramic video experience. Specifically, two independent convolutional layers are used to process the resource space and user information respectively. Since the dimension of the resource space is much larger than the dimension of the user device information, a fully connected layer is added after the convolutional layer that processes the resource space information to reduce the dimensionality of the resource space information, making the resource space information have a similar dimension to the user device information. The processing flow is as follows: (1) Resource spatial information processing flow: Recording resource block allocation status The matrix is ​​used as the first input, and after processing through convolution and pooling layers, it is then... The operation then passes through a fully connected layer; (2) User information processing flow: Recording user-related information The matrix is ​​used as the second input, processed by the convolutional layer, and then... operate; (3) After adding the two results obtained from the above processing together, pass them through a fully connected layer and finally output the result; After modeling the state information and user information and designing the neural network, the DQN algorithm can be used to allocate wireless time and frequency resource blocks, thereby maximizing the overall panoramic video experience quality while meeting the basic experience quality requirements of all users.