Dual-time-scale MuAR system resource allocation method under 6G network framework
Through a dual-time-scale resource allocation framework and an autonomously trained intelligent agent model, computing resources and channel configuration are optimized, the latency and consistency problems in multi-user augmented reality systems are solved, and low-latency and high-precision virtual object presentation is achieved.
Patent Information
- Application Number
- CN202510718073.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing 5G network architecture cannot effectively support real-time interaction in multi-user augmented reality systems, especially when faced with virtual object position offsets, asymmetric network traffic, high centralized scheduling and synchronization overhead, and insufficient interference coordination capabilities, leading to delays and consistency issues that affect the collaborative experience.
A dual-time-scale resource allocation framework is adopted, combined with a Bayesian optimization-enhanced deep deterministic policy gradient algorithm and a graph attention network, to optimize computing resource allocation and channel and subframe configuration in large and small time scales, respectively, enabling each intelligent agent to autonomously train a localized model, quickly adapt to environmental changes, and coordinate interference dependencies.
Under the premise of ensuring the delay requirements, the position offset rate is minimized, and the spatial alignment accuracy and collaborative experience quality of the multi-user augmented reality system are improved.
Smart Images

Figure CN120640420A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of 6G mobile communication technology, and specifically provides a resource allocation method for a dual-time-scale mobile multi-user augmented reality system under a 6G network framework. Background Art
[0002] Mobile augmented reality (AR) overlays virtual objects onto the physical environment through wireless optical or video see-through displays, extending the perception of the virtual world into the real world. Multi-user augmented reality (MuAR) systems facilitate real-time interaction with virtual objects in a shared augmented environment, revolutionizing the collaborative experience. Current 5G architecture resource allocation mechanisms are unable to support MuAR real-time interaction scenarios, particularly due to issues such as significant virtual object position offsets, inefficient asymmetric network traffic transmission, high centralized scheduling and synchronization overhead, and insufficient interference coordination capabilities. Therefore, next-generation mobile communication networks are needed to enable fast and accurate data sharing between AR users. MuAR systems also pose challenges to the underlying network: strict rendering latency (typically less than 20 milliseconds per frame) is required to prevent visual artifacts, while maintaining millisecond-level end-to-end latency consistency across all collaborating users. Timing differences between users can cause positional offsets of virtual objects, disrupting collaborative spatial perception and shared environment consistency. Addressing these two challenges is crucial to maintaining spatial alignment accuracy.
[0003] In the MuAR system, end-to-end latency is primarily generated during 3D registration and rendering. Offloading 3D registration and rendering tasks to the Mobile Edge Computing (MEC) server can effectively alleviate computational bottlenecks. However, asymmetric data traffic is generated when the native video captured by the terminal device is uploaded to the MEC server and when the rendered video is downloaded to the terminal device. This asymmetric traffic problem can be effectively addressed by adaptively configuring the uplink / downlink channel direction using Dynamic Time-Division Duplexing (DTDD) technology. In addition, any modifications initiated by users to shared virtual objects must be propagated to the views of all participants in real time, which can cause position offset problems caused by cross-device status updates.
[0004] Multi-Agent Deep Reinforcement Learning (MA-DRL) enables autonomous decision-making under time-varying network conditions and is often applied to resource allocation problems. However, traditional MA-DRL methods typically use a centralized training with decentralized execution (CTDE) approach, where agents share historical states and actions during centralized training and then execute pre-trained networks based on local environment observations. This model has serious limitations. 1) In a dynamic environment with real-time resource allocation across multiple base stations (BSs), frequent policy synchronization incurs excessive communication overhead and latency penalties. 2) Computational complexity increases exponentially with the number of BSs, leading to scalability issues. 3) The model over-relies on global information during centralized training, but during decentralized execution, each BS operates solely based on localized observations and lacks cross-cell state awareness, resulting in a mismatch between training and deployment and weakening the effectiveness of the policy. 4) Centralized training may implicitly assume a stable environment, but during decentralized execution, each agent's local environment may constantly change due to changes in the policies of other agents, further exacerbating the degradation of model performance. These limitations prevent the MA-DRL method from being directly applied to the multi-dimensional joint resource optimization problem in DTDD-enhanced MuAR systems. Summary of the Invention
[0005] In order to solve the problem of minimizing the position offset rate while ensuring the delay requirement in the mobile MuAR system under the 6G network framework, the present invention proposes a two-timescale distributed resource allocation (TDRA) framework to jointly optimize the long-term computing resource allocation (CRA) and the short-term channel and subframe configuration (CASC).
[0006] To achieve the above object, the present invention adopts the following scheme:
[0007] Due to the deep coupling of computing resources with channel and subframe resources and their different time-varying characteristics, this constitutes a cross-time domain resource management problem. From a practical point of view, computing resource allocation is a large-time-scale decision-making process, because the arrival time of computing tasks is relatively stable, with intervals in seconds. On the other hand, channel states and interference fluctuate on smaller time scales (such as milliseconds). In order to solve this time decoupling problem, the resource allocation framework proposed in this invention is a dual-time-scale one. Specifically, the large-time-scale window size is based on the video refresh cycle of the AR device (16 milliseconds), while the small-time-scale interval is based on the dynamic TDD subframe duration (1 millisecond). This hierarchical architecture can perform differentiated and synchronous management of heterogeneous resources in different time domains.
[0008] To overcome the limitations of traditional MA-DRL methods, this paper adopts a fully decentralized learning strategy, in which each agent autonomously trains a localized model, enabling rapid adaptation to transient environmental changes while maintaining decision-making responsiveness. Due to the limitation of partial observations, it has fundamental limitations in capturing global features. This limitation hinders effective coordination when resolving complex dependencies and may lead to decision conflicts between agents. To address this issue, this paper integrates a Graph Attention Network (GAT) into the fully decentralized learning strategy, explicitly modeling the hidden interference dependencies between agents to promote coordinated optimization. Furthermore, applying MA-DRL methods to multidimensional resource allocation problems with large state and action spaces can be computationally intensive and inefficient. To address this issue, this paper develops a Bayesian Optimization-Enhanced Deep Deterministic Policy Gradient (BO-DDPG) algorithm, which uses the joint prior distribution of computational requirements and computational resource allocation decisions to guide exploration, thereby improving the learning efficiency of DDPG.
[0009] The specific steps of the method of the present invention are as follows:
[0010] S1: Define the video refresh cycle of the AR device as a large time scale. Within each large time scale, apply a deep deterministic policy gradient algorithm enhanced by Bayesian optimization to make computing resource allocation decisions to minimize the computational latency of 3D registration and rendering tasks.
[0011] S2: The duration of a subframe is defined as a small time scale. Within each small time scale, a bidirectional graph attention network deep Q network is applied to make channel allocation and subframe configuration decisions to minimize the position offset rate of virtual objects presented on different display devices.
[0012] Specifically, in S1, the deep deterministic policy gradient algorithm enhanced by Bayesian optimization is applied, including:
[0013] S1-1: Statistical modeling of the conditional probability of events based on historical samples;
[0014] S1-2: Optimize the utility function related to the event occurrence, and output the optimal computing resource allocation decision through the gradient descent method based on the given posterior distribution
[0015] S1-3: Training DNN is a function approximator that obtains the optimal computing resource allocation decision based on the current strategy According to the Bellman equation, we can get and State-action value and
[0016] S1-4: By comparing two state-action values and The largest state-action value is selected as the final output action.
[0017] Specifically, S1-1: The process of statistically modeling the conditional probability of an event based on historical samples includes:
[0018] and
[0019] Bayesian optimization: Computational resource requirements are defined as triplets in and Represents the number of CPU cycles required to perform 3D registration, host rendering, and resolver rendering, respectively. Host rendering refers to the rendering of virtual objects initiated by the AR user when running as a host. Resolver rendering refers to the rendering of virtual objects initiated by other hosts when running as a resolver.
[0020] The distribution of the number of CPU cycles required for 3D registration is simulated as a Pareto distribution, expressed as:
[0021] Among them, a, b, and c represent shape, scale, and threshold respectively;
[0022] Then, its cumulative distribution function can be calculated as follows:
[0023]
[0024] The number of CPU cycles required for host rendering is expressed as the truncated log-normal distribution:
[0025] Where μ represents the mean and δ represents the standard deviation;
[0026] Then, its cumulative distribution function can be calculated as follows:
[0027]
[0028] The distribution of the total number of CPU cycles required for parser rendering is simulated as an exponential distribution as follows:
[0029] Where λ represents the interaction strength;
[0030] Then, its cumulative distribution function can be calculated as follows:
[0031]
[0032] Specifically, S1-2: Optimizing the utility function associated with the event occurrence and outputting the optimal computing resource allocation decision using the gradient descent method based on a given posterior distribution includes:
[0033] Define a function that maps CRA variables to computational resource requirements as follows:
[0034] where Ψ n,m is the CRA variable, namely BS B n For AR users m The decision variable of the allocated computing resources,∈,is the error term, which is regarded as Gaussian noise with mean 0;
[0035] According to Bayes' theorem, the posterior distribution and the prior distribution and likelihood function of g(Ψ) Related, as shown below:
[0036] in It is a historical sample set;
[0037] Define a utility function to describe the expected improvement of the function value g(Ψ), expressed as: g * (Ψ) represents the maximum function value in the past;
[0038] The gradient descent algorithm is used to solve the utility function maximization problem. By maximizing the utility function, the optimal CRA decision for the next time slot is found.
[0039] Specifically, the training DNN in S1-3 is a function approximator used to obtain the optimal computing resource allocation decision based on the current strategy The process includes:
[0040] The computational delay minimization problem is formulated as a Markov decision process consisting of a state space, an action space, and a reward function.
[0041] Taking the base station (BS) as the intelligent agent, B n The elements in the state space of Represents AR user U m The computing resource requirements of three types of tasks; AR user U m The action space includes in and denote the proportion of computing resources allocated to perform 3D registration, host rendering, and parser rendering respectively; the reward function is defined as the total negative value, i.e. in and represent the computational latency of performing 3D registration, host rendering, and resolver rendering, respectively.
[0042] The framework consists of a training DNN and a target DNN; the training DNN is a function approximator used to obtain the optimal action a according to the current policy π(θ) * (t) and the state-action value of the action As shown below:
[0043]
[0044] The target DNN is used to evaluate the best action based on the following factors:
[0045]
[0046] Among them, θ and ω are the weight parameters of the training DNN and the target DNN respectively; the weight parameters of the training DNN are updated through back propagation, and its error loss is:
[0047]
[0048] Among them, N R is the size of the random sampling batch of the replay memory; the target DNN copies the weight parameters of the training DNN every fixed step.
[0049] Specifically, in S2, the process of applying the bidirectional graph attention network deep Q network to make channel allocation and subframe configuration decisions includes: modeling the mobile multi-user augmented reality system as a system consisting of N B The directed graph composed of subgraphs captures the bidirectional interference dependencies between intelligent agents through a bidirectional graph attention network; and then generates a joint optimization decision of channel allocation and subframe configuration through a deep Q network to minimize the average position offset rate.
[0050] Specifically, the mobile multi-user augmented reality system is modeled as a B The process of capturing the bidirectional interference dependency between agents through a bidirectional graph attention network in a directed graph composed of subgraphs includes:
[0051] The mobile multi-user augmented reality system is modeled as a B The process of capturing the bidirectional interference dependency between agents through a bidirectional graph attention network in a directed graph composed of subgraphs includes:
[0052] The mobile multi-user augmented reality system is modeled as a B A directed graph composed of subgraphs, using Indicates that the nth subgraph is represented by Indicates that represents AR users, ε:={e m,m′} represents the interference received by node m from node m′;
[0053] The interference received by node m from other nodes and the interference caused by node m to other nodes are input into the MLP layer respectively. The MLP layer extracts the potential features of each node and converts them into high-dimensional features, which are expressed as:
[0054]
[0055] in and is the weight parameter that needs to be trained in MLP, σ(·) represents the ReLu activation function, I n,m and They are the interference received by node m from other nodes and the interference caused by node m to other nodes. The coefficients between nodes can be realized by multi-head attention as shown below:
[0056] α m,m′ =ATT(Wh m ,Wh m′ );
[0057] Among them, ATT(·) is the self-attention mechanism, W is the corresponding weight parameter, and the formula expresses the importance of node m′ to node m;
[0058] The masked and normalized attention coefficients can be implemented as follows:
[0059]
[0060] where softmax(·) is the softamx function and LeakyReLu(·) is the LeakyReLu function. The state embedding is obtained by connecting Q independent attention mechanisms in series, and its value is:
[0061]
[0062] Where ||(·) represents the concatenation operation and Q represents the number of attention heads;
[0063] The state embedding is passed to the fully connected layer for feature fusion, and its calculation formula is:
[0064]
[0065] Specifically, the process of generating a joint optimization decision for channel allocation and subframe configuration through a deep Q network includes:
[0066] The state space of partially observable Markov decisions is represented by express;
[0067] The training DNN in the deep Q network updates the weight parameters by minimizing the mean squared error loss function, which is:
[0068]
[0069] in, is the mini-batch size, and the weight vector μ(t+1) can be solved using the gradient descent method, i.e.:
[0070]
[0071] Where Γ represents the learning rate.
[0072] Beneficial effects: Compared with the prior art, the present invention can achieve at least the following technical effects:
[0073] 1. This paper proposes a dual-time-scale resource allocation framework to jointly optimize long-term computing resource allocation and short-term channel and subframe configuration, minimizing the position offset rate while ensuring the delay requirement.
[0074] 2. The present invention adopts a completely decentralized learning strategy, that is, each agent autonomously trains a localized model, so that it can quickly adapt to instantaneous environmental changes while maintaining decision-making responsiveness. The Graph Attention Network (GAT) is integrated into the completely decentralized learning strategy to explicitly simulate the hidden interference dependencies between agents to promote coordinated optimization.
[0075] 3. This paper develops the Bayesian Optimization-Enhanced Deep Deterministic Policy Gradient (BO-DDPG) algorithm, which uses the joint prior distribution of computational requirements and computational resource allocation decisions to guide exploration and improve the learning efficiency of DDPG. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 : Schematic diagram of VST AR system;
[0077] Figure 2 :MuAR system workflow diagram;
[0078] Figure 3 : D-TDD frame structure diagram;
[0079] Figure 4 : BiGATDQN framework;
[0080] Figure 5 : TD-RA working framework;
[0081] Figure 6 : Convergence performance graph of BODDPG, DDPG and BO algorithms for CRA problem;
[0082] Figure 7 : Relationship between the latency of BO-DDPG, DDPG, and BO algorithms and the number of AR users;
[0083] Figure 8 : Relationship between latency and computational power of BO-DDPG, DDPG, and BO algorithms;
[0084] Figure 9 : Relationship between latency and interaction intensity of BO-DDPG, DDPG and BO algorithms;
[0085] Figure 10 : Convergence performance of different algorithms in CASC sub-scheme;
[0086] Figure 11 : Relationship between the average mean offset rate and the number of AR users in different algorithms;
[0087] Figure 12 : Relationship between average mean offset rate and bandwidth in different algorithms;
[0088] Figure 13 : Relationship between average mean offset rate and mutual strength in different algorithms;
[0089] Figure 14: Variation of the position offset rate with the number of AR users under different resource allocation schemes (i.e., TD-RA, TD-RA woCRA, TD-RA woCA, and TD-RA w.o.SC);
[0090] Figure 15 : The relationship between the position offset rate and the computing energy for different resource allocation schemes (i.e., TD-RA, TD-RA woCRA, TD-RA woCA, and TD-RAw.o.SC).
[0091] Figure 16 : Variation of the position offset rate with bandwidth for different resource allocation schemes (i.e., TD-RA, TD-RA wOCRA, TD-RA wCA, and TD-RA w.o.SC). DETAILED DESCRIPTION
[0092] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention and not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0093] According to the visual presentation method, AR systems can be divided into two types: Optical See-Through (OST) and Video See-Through (VST). This paper studies the VST AR system, which works by using a camera to capture the real-world scene, then combining it with virtual objects and displaying it on a near-eye head-mounted display (HMD). The hardware composition and workflow of the VST AR system are as follows: Figure 1 As shown. Its specific workflow is (1) first capture video through the camera, and then the system generates a three-dimensional point cloud based on the visual data as a digital representation of the physical environment. (2) Perform a three-dimensional registration operation to accurately superimpose the virtual objects generated by the server into the real environment. The three-dimensional registration operation includes feature extraction, feature matching, and pose estimation. Pose estimation is to reposition and describe the objects in the coordinate system from the observer's perspective, determine which virtual objects the user can see, and define the precise spatial relationship between them. (3) In order to obtain a realistic visual experience, lighting information should be shared between the physical environment and virtual objects. The rendering process uses camera lens and position data to estimate the ambient lighting of the scene, and then performs virtual reality synthesis, including shadow generation, environmental occlusion, and reflection / refraction simulation. (4) Finally, the image is displayed to the AR user using the HMD.
[0094] The three-dimensional point cloud in the above process (1) is used to describe the high-fidelity details of the physical environment, but its massive data causes the transmission delay to exceed the acceptable threshold for real-time data sharing. Since AR users observe the physical environment in a limited area on the same side, the present invention adopts a lightweight multi-user augmented reality system framework, using a two-dimensional face-side representation to replace the three-dimensional point cloud, which can effectively reduce the problem of excessive delay caused by excessive data volume. In this MuAR system, there are two important roles, namely the host and the parser. It should be noted that the host and the parser are equivalent concepts, and the roles of the two can be interchanged.
[0095] like Figure 2 As shown in Figure 2, the MuAR workflow mainly includes seven steps: (1) the host captures video and generates two-dimensional face-side data of the physical environment based on a semantic-based quadtree method; (2) the host inputs virtual objects and performs a three-dimensional registration operation to create and maintain local real-word coordinates; (3) the host renders the virtual objects based on real-time lighting; (4) to share virtual objects, the host inserts the virtual objects into the quadtree data and transmits them to the parser; (5) the parser extracts semantic features from the received quadtree data and matches them with the local feature library; (6) the parser constructs a matrix to represent the spatial relationship of the field of view (FoVs) between the host and the parser; (7) the parser renders the virtual objects shared by the hosts. Position offsets caused by different device processing capabilities should be corrected to ensure that AR users have a unified experience. Therefore, position offset correction should be performed after step 7.
[0096] The present invention considers a DTDD-based edge computing system consisting of BSs and AR users. It is assumed that AR users are randomly distributed in the network and connect to the nearest BS. Each BS is equipped with a MEC server that processes complex computing tasks while also providing network services. BSs are connected to each other and to the central cloud server through a wired core network. Each BS caches the feature library of the physical environment and downloads the raw data of virtual objects from the central cloud server. The index set of the BS is used Indicates that N B is the number of BSs. In a small cell, BS B n Adopting the Orthogonal Frequency Division Multiple Access (OFDMA) structure, through N C The channels are N A In the nth small cell, the index sets of AR users and channels are respectively and Indicates. BS B nChannel assignment matrix Represents that the vector If BS B n For AR users U m Assign Channel C i , then c n,m,i =1, otherwise c n,m,i =0.
[0097] When an AR user acts as a host, it uploads its locally generated data to the relevant BS. When it acts as a resolver, it downloads data of other AR users from the BS. In order to handle the asymmetric traffic between uplink (UL) and downlink (DL) channels, the present invention adopts the DTDD frame structure, such as Figure 3 As shown in Figure 1, it defines a frame time as 10ms, consisting of 10 subframes, each subframe time is 1ms, and each subframe consists of 14 OFDM symbols, which are used for UL transmission, DL transmission and guard period (GP), respectively. GP is only configured for two adjacent symbols between subframes in different transmission directions. A complete subframe is completely used for UL or DL. Indicates that the subframe set of frame t is represented by where N F =10. BS B n In channel C i The subframe configuration set on If channel C i Subframe F on k BS B n Configured to send AR user U m DL, then f i,t,k =1, otherwise f i,t,k =0.
[0098] Since the execution of computing tasks is coherent, while the channel status and traffic patterns are time-varying, we propose a dual-time-scale resource allocation scheme, where the large time-scale is used for computing resource allocation and the small time-scale is used for channel allocation and subframe configuration. n For AR users m The joint optimization of channel allocation and subframe configuration is defined as the tuple Φ n,m :={c n,m,i ,f i,t,k}
[0099] (1) Transmission model
[0100] To improve spectrum efficiency, this experiment allows transmission links within different cells to reuse the same channel, which may lead to same-link interference (SLI) and cross-link interference (CLI) in the D-TDD system. SLI occurs when two AR devices in adjacent cells transmit data in the same direction; CLI occurs when two AR devices in adjacent cells operate in opposite directions.
[0101] As the host, AR user U m In subframe F k Through channel C i To BS B n Upload its registration data or virtual object model, the UL transmission rate is:
[0102]
[0103] where |C i | is channel C i Bandwidth, P n,m AR user U m Transmit power, g mn [i] is an AR user U m to BS B n The channel gain, σ 2 is the additive Gaussian noise power. It's BS n Received interference, including SLI from AR users in adjacent cells and CLI from neighboring BSs The calculation formula is:
[0104]
[0105] Among them, N′ M Is with AR user U m The number of AR users of the neighboring BS n′ using the same channel, P n′,m′ and P n′ They are AR users U′ m and the transmission power of the neighboring BS n′, g m′n is AR user U′ m to B n The channel gain, g n′n is BS n′ to BSB n channel gain.
[0106] As a parser, AR user U m In subframe F k Through channel Ci From BS B n Download the calculation results of your own tasks and the virtual object models of other AR users. The DL transmission rate is:
[0107]
[0108] Among them, P n It's BS n Transmit power, g nm [i] is BS B n To AR user U m channel gain. AR user U m The received interference is caused by the SLI from the neighboring BS and CLI from neighboring small cell AR users Composition, its value is:
[0109]
[0110] Among them, g n′m and g m′m are BS n′ and AR user U′ respectively m To AR user U m channel gain.
[0111] (2) Computational model
[0112] Since 3D registration and rendering are resource-intensive tasks, they are all performed by the MEC server. m First, a 3D registration is performed to obtain the current direction and position, and then the virtual objects initiated by it are rendered. m Render virtual objects initiated by other hosts based on their own direction and position information. Define triples As the BS B in the lth large time scale period n For AR users m Allocated computing resources decision variables. and The computational latency for performing 3D registration is represented by the ratio of computing resources allocated to performing 3D registration, host rendering, and resolver rendering. Computational latency of host rendering and parser rendering computation delay They are:
[0113]
[0114]
[0115] Computational resource requirements are defined as a triplet in and Represents the number of CPU cycles required to perform 3D registration, host rendering, and parser rendering, respectively. n Indicates BS B n The total number of CPU cycles of MEC servers in BS B n The decision variables for computing resource allocation should satisfy the constraints:
[0116]
[0117] (3) Delay model
[0118] Taking the HoloLens 2 AR device as an example, its display refresh rate and camera capture rate are both synchronized at 60Hz. This means that the end-to-end latency per video frame must not exceed 16ms, otherwise artifacts or jitter may occur. Therefore, we set 16 milliseconds as the large-scale period, and a subframe as the small-scale period. During the AR application initialization process, after the AR device captures the first video frame through the camera, it retrieves the current system frame number (SFN) through the Master Information Block (MIB) broadcast information. It then enters a standby phase until the large-scale lifecycle begins, ensuring synchronization with the MuAR system.
[0119] Assume that AR user U m In the Subframe starts to move towards BS B n Upload a frame of video and Subframe upload completed. Subframe and If the subframes are in the same frame, then:
[0120]
[0121] where |T s | represents 1ms, and the upload delay is:
[0122]
[0123] like The subframe is located in frame t+1, then:
[0124]
[0125] The upload delay is:
[0126]
[0127] like The subframe is located at frame t+2, and Subframe and The interval between subframes does not exceed 16ms, then:
[0128]
[0129] The upload delay is:
[0130]
[0131] Assume that the rendered virtual object is from The subframe starts downloading and The subframe ends. Subframe and If the subframes are in the same frame, then:
[0132]
[0133] The download delay is:
[0134]
[0135] like The subframe is located in frame t+1, then:
[0136]
[0137] The download delay is:
[0138]
[0139] like The subframe is located at frame t+2, and Subframe and The interval between subframes does not exceed 16ms, then:
[0140]
[0141] The download delay is:
[0142]
[0143] Therefore, AR user U m The end-to-end delay experienced by a video frame can be expressed as:
[0144]
[0145] (4) Position offset model
[0146] Latency consistency between AR users is crucial because inconsistent latency can reduce spatial alignment accuracy and cause position shifts in the same virtual object from different users’ perspectives. The relationship between the position shift rate and AR user latency differences can be modeled as:
[0147]
[0148] Where ρ1 acts as a global scaling factor, determining the overall magnitude of the position offset rate. ρ2 adjusts the initial amplitude of the exponential function that models the sensitivity to delay differences. ρ3 adjusts the nonlinear sensitivity to delay differences, where larger ρ3 values amplify PO n,m The responsiveness to delay changes. ρ4 controls the nonlinearity of the model response curve. Finally, ρ5 acts as a baseline offset term to ensure that PO n,m Even under zero delay conditions, the physically meaningful values are maintained. To meet the normalization requirements, the factors should satisfy ρ1≥0, ρ2>0, ρ3<0, ρ4>0, ρ5≥0 and
[0149] (5) Optimization goal
[0150] The end-to-end delay of video frames is a key indicator of AR user quality of experience (QoE), and keeping it below 16 milliseconds can meet user needs. In addition, the position offset rate is a key indicator for evaluating the consistency of interaction between AR users in the MuAR system. Therefore, the optimization goal of this paper is to minimize the position offset rate under the constraint of a delay below 16 milliseconds. To achieve this goal, we propose a dual-time-scale resource management framework in which channel allocation, subframe configuration, and computational resource allocation are jointly optimized. The position offset rate minimization problem can be expressed as
[0151]
[0152] Constraint C1 means that end-to-end latency cannot exceed 16 milliseconds. Constraints C2 and C3 ensure that the channel allocation and subframe configuration parameters are binary, respectively. Constraint C4 limits the computing resource allocation ratio of each base station to the range of 0 to 1, excluding 0. Constraint C5 shows the computing capacity limit of each base station.
[0153] (6) Solution
[0154] Aiming at the problem of minimizing the position deviation rate, a TD-RA scheme is proposed. The framework of the TD-RA scheme consists of two layers: the upper layer is the large time scale CRA sub-scheme, and the lower layer is the small time scale CASC sub-scheme, such as Figure 5As shown in Figure 2, each large time period consists of 16 small time periods, which are aligned with subframes. Each BS equipped with an MEC server acts as an agent to implement the resource allocation scheme in a distributed manner.
[0155] First, during each large-scale lifecycle, each AR device captures a video frame. The corresponding computational requirements are forwarded to the relevant base station (BS) and then shared between adjacent BSs via optical fiber. The computational requirements and CASC decisions of the last small time period of the previous large time period are input into the BO-DDPG network as state. The BO-DDPG network learns the characteristics of the computational requirements and then outputs the CRA decision with the lowest end-to-end delay based on the given CASC decision. During each small timescale interval, the CRA decision of the current large-scale lifecycle, the amount of uploaded video frame data, the output result downloaded after rendering, as well as signal strength and bidirectional interference, are input into the BiGAT-DQN network. Finally, the bidirectional interference characteristics are learned using BiGAT, and the DQN is used to make the CASC decision.
[0156] CRA sub-scheme based on BO-DDPG
[0157] At each large time scale, the BO-DDPG CRA subroutine is formulated as a computational delay minimization problem as follows:
[0158]
[0159] Problem II is a nonlinear programming problem involving strongly coupled polynomial functions and As a model-free algorithm for continuous action spaces, DDPG offers a general solution to the problem of minimizing computational latency. However, when acting as a parser, each AR user must collect the computational resource requirements of other users to initialize the rendering of virtual objects, and these requirements arrive in a random manner. Therefore, directly applying DDPG to the computational resource allocation sub-scheme is difficult. Integrating Bayesian Optimization (BO) into the DDPG framework improves DDPG's learning efficiency by providing informed action exploration directions to address this issue.
[0160] (1) Bayesian Optimization
[0161] The distribution of the number of CPU cycles required for 3D registration is simulated as a Pareto distribution (PD) as follows:
[0162]
[0163] Where a, b, and c represent shape, scale, and threshold, respectively. Then, its cumulative distribution function can be calculated as follows:
[0164]
[0165] The distribution of the number of CPU cycles required for host rendering has a severe tail characteristic, and its Truncated Lognormal Distribution (TLD) is shown below:
[0166]
[0167] Where μ is the mean and δ is the standard deviation; then, its cumulative distribution function can be calculated as follows:
[0168]
[0169] The distribution of the total number of CPU cycles required for parser rendering is simulated as an exponential distribution as follows:
[0170]
[0171] where λ is the interaction strength; then, its cumulative distribution function can be calculated as follows:
[0172]
[0173] Define a function that maps CRA variables to computational resource requirements as follows:
[0174]
[0175] Where ∈ is the error term, which is treated as Gaussian noise with mean 0. According to Bayes’ theorem, the posterior distribution and the prior distribution and likelihood function of g(Ψ) Related, as shown below:
[0176]
[0177] in is a set of historical samples. Then, a utility function is defined to describe the expected improvement of the function value g(Ψ) as follows:
[0178]
[0179] Among them, g * (Ψ) represents the maximum function value in the past. The gradient descent algorithm is used to solve the utility function maximization problem. By maximizing the utility function, the optimal CRA decision for the next time slot can be found.
[0180] (2)DDPG
[0181] The computational delay minimization problem is reformulated as a Markov decision process consisting of a state space, an action space, and a reward function. In the CRA sub-scheme, the base station is used as an intelligent agent, BS B n The elements in the state space of Represents AR user U m The computing resource requirements of three types of tasks; AR user U m The action space includes in and denote the proportion of computing resources allocated to perform 3D registration, host rendering, and parser rendering respectively; the reward function is defined as the total negative value, i.e.
[0182] The DDPG framework consists of a training DNN and a target DNN. The training DNN is a function approximator used to obtain the optimal action a according to the current policy π(θ) * (t) and the state-action value of the action As shown below:
[0183]
[0184] The target DNN is used to evaluate the best action based on the following factors:
[0185]
[0186] Where θ and ω are the weight parameters of the training DNN and the target DNN respectively. The weight parameters of the training DNN are updated through back propagation, and the error loss is:
[0187]
[0188] Among them, N R is the size of the random sampling batch of the replay memory. The target DNN copies the weight parameters of the training DNN every fixed step.
[0189] (3)BO-DDPG
[0190] Bayesian optimization provides a mathematical framework that statistically models the conditional probability of events based on historical samples and then optimizes the utility function associated with the occurrence of events to determine the best action. It enhances the role of semi-supervision by providing more informed directions for action exploration, especially in the early stages of training. Based on a given posterior distribution, the optimal CRA decision is output through gradient descent. The optimal CRA decision made by training DNN is used According to the Bellman equation, the state-action value is obtained. and Finally, by comparing the two state-action values, the final output action of DDPG is determined.
[0191] The specific process of BO-DDPG is as follows: (1) Parameters are initialized first; (2) CRA decisions are obtained using Bayesian optimization and training DNN respectively; (3) Based on the two actions, BS calculates their respective state-action values and selects the action output with the larger state-action value; (4) BS executes the selected action on the system and records the transition into the replay memory.
[0192] CASC sub-scheme based on BiGAT-DQN
[0193] Considering that the interaction between AR users is dynamic and each BS cannot obtain global environmental knowledge, we formulate the CASA problem as a distributed partially observable Markov decision process (PO-MDP). Each BS obtains local observation data and makes CASA decisions in a distributed manner. PO-MDP consists of<o,s,a,r> where o represents the observation value, s represents the state, a represents the action, and r represents the reward. Indicates that the action space is represented by a={c n,m,i , f i,t,k}. Given a CRA decision, the reward function can be formulated as a series of single-shot CASA problems as follows:
[0194]
[0195] The complex spatial relationships between base stations make solving the CASC problem in a distributed MuAR system based on DTDD extremely challenging. Specifically, the following challenges must be considered: 1) Each agent observes insufficient environmental information; 2) Agents update their strategies in a distributed manner, resulting in an uncertain and non-stationary learning environment; 3) Interference between agents is bidirectional, and the importance of different interferences varies. Specifically, when making decisions, agents need to consider the interference from neighboring cells and the interference their own decisions may cause to neighboring cells. Furthermore, given factors such as available channel capacity and channel status, the impact of the same amount of interference can vary.
[0196] In this work, we propose a cooperative mechanism between agents and use BiGAT-DQN to capture the relationship between agents. The MuAR system can be modeled as a B A directed graph composed of subgraphs, using Indicates that the nth subgraph is represented by Indicates that represents AR users, ε:={e m,m′} represents the interference received by node m from node m′. Figure 4 As shown in Figure 2, the interference received by node m from other nodes and the interference caused by node m to other nodes are input into the MLP layer respectively. The MLP layer then extracts the latent features of each node and converts them into high-dimensional features, which are represented as follows:
[0197]
[0198] in and is the weight parameter that needs to be trained in MLP, and σ(·) represents the ReLu activation function. n,m and are the interference received by node m from other nodes and the interference caused by node m to other nodes. The coefficients between nodes can be implemented through multi-head attention as shown below:
[0199] α m,m′ =ATT(Wh m ,,Wh m′ ) (41)
[0200] Where ATT(·) is the self-attention mechanism and W is the corresponding weight parameter. This formula expresses the importance of node m′ to node m. Then the masked and normalized attention coefficient can be achieved as follows:
[0201]
[0202] Where softmax(·) is the softmax function and LeakyReLU(·) is the LeakyReLu function. Coefficient α m,m′ and It can be obtained by equations (36) and (37). In addition, the state embedding can be obtained by connecting Q independent attention mechanisms in series, and its value is:
[0203]
[0204] Where ||(·) represents the concatenation operation and Q represents the number of attention heads. In addition, the state embedding is passed to the fully connected layer for feature fusion, which is calculated as:
[0205]
[0206] The state space of PO-MDP is expressed as Represented. Similar to the traditional DQN algorithm, the BiGAT-DQN algorithm uses a training DNN and a target DNN, where the training DNN takes a specific action in the current state according to (29) and (30), and the target DNN evaluates the action according to (31). The training DNN in the DQN algorithm updates the weight parameters by minimizing the mean square error loss function, which is:
[0207]
[0208] in, is the mini-batch size. The weight vector μ(t+1) can be solved using the gradient descent method, i.e.:
[0209]
[0210] Where Γ represents the learning rate.
[0211] TD-RA convergence analysis
[0212] From a lifecycle perspective, Problem I is a long-term cumulative function. However, Problems II and III are one-time problems, and decisions can only be made within a large or small time scale. Considering the stochastic process of other BSCAS decisions, the following theorem can be introduced to verify the convergence of the TD-RA scheme.
[0213] Theorem 1: For the long-term TD-RA problem, if the average position shift rate is ξ, then the probability that the traversal position shift rate is feasible is 1, that is,
[0214]
[0215] in,
[0216] Proof: If a positive integer sequence s t is convergent, then it satisfies the following conditions:
[0217]
[0218] in, is a sequence of positive integers s t The average value of .
[0219]
[0220] Among them, 0 <M0<∞。
[0221] Furthermore, Abel's theorem states that:
[0222]
[0223] Therefore, we have:
[0224]
[0225] On this basis, we define the long-term return as γ is the discount factor. Then, we can get:
[0226]
[0227] The proof is complete.
[0228] The present invention will be described in detail below with reference to the accompanying drawings and simulation examples.
[0229] The detailed steps of the BO-DDPG-based CRA solution are as follows:
[0230] Step 1: Determine and enter computing resource requirements and
[0231] Step 2: Initialize the number of BSs N B 、Number of AR users A , the computing power of each MEC server W n , parameters of Pareto distribution, truncated lognormal distribution and exponential distribution: a, b, c, δ, σ and λ, historical samples The weight parameters θ, ω of the training DNN and the target DNN.
[0232] Step 3: Agent n uses historical samples Estimating the posterior probability
[0233] Step 4: Maximize the utility function according to formula (28) and use the gradient descent algorithm to obtain the optimal CRA decision
[0234] Step 5: According to equations (29) and (30), use the state and Get the CRA decision made by the trained DNN
[0235] Step 6: Enter the action and to the target DNN.
[0236] Step 7: Evaluate the state-action value through the target DNN and
[0237] In step 8, the action that DDPG finally outputs is determined based on the higher state-action value.
[0238] Step 9, store (s(t), a(t), r(t), s(t+1)) into the playback memory.
[0239] Step 10: Update the weights of the trained DNN according to formula (32).
[0240] Step 11: Step 3 to Step 10 is one frame operation, repeat Step 3 to Step 10 T times, and repeat every T s Step 1, assign θ to ω.
[0241] Step 12: All the above steps are a fragment of the process, and the above steps are repeated until the end.
[0242] Next, we present simulation results to evaluate the effectiveness of the TD-RA scheme in a D-TDD MuAR system. In this system, 10 base stations are evenly distributed and 300 AR users are randomly distributed in an area of 2 × 2 square kilometers. The key parameters of the system are shown in Table I. The parameter value of the position offset rate model in Equation (20) is ρ1 = -5.247e -10 ,ρ2=5.1,ρ3=0.9521,ρ4=45.96,ρ5=-8.648e -5 A GAT consists of four attention heads, each of which is composed of a single-layer feedforward neural network activated by the LeakyReLU activation function. DDPG uses an input layer, three hidden layers, and an output layer with 512, 128, and 256 neurons, respectively. DQN consists of four fully connected layers with 128, 128, and 64 neurons, respectively.
[0243]
[0244] The convergence performance of BODDPG, DDPG and BO algorithms for CRA problems is as follows Figure 6 As shown, it is clear that BO-DDPG converges faster and achieves higher average cumulative rewards compared to other schemes. This is because BO improves DDPG's learning efficiency by estimating computational resource requirements. Furthermore, compared to both DDPG and BO, BODDPG's learning curve is more stable and smoother in both the early learning and convergence stages. This demonstrates that BO can effectively reduce random exploration in DDPG. It can also be seen that BO converges to the lowest average cumulative reward. This is because each base station uses BO to achieve the optimal CRA decision based on the distribution of historical CRA decisions over a large time scale of average cumulative rewards under the CASC sub-scheme. It does not take into account the dynamic nature of the network environment.
[0245] The relationship between the delay of BO-DDPG, DDPG and BO algorithms and the number of AR users is as follows: Figure 7As shown, as the number of AR users increases, the latency of all three algorithms increases, with BO-DDPG achieving the lowest latency. When the number of users exceeds 121, the computational latency of BO is zero, while that of DDPG is zero when the number of users exceeds 169. This is because in our approach, the computational latency is set to zero when the end-to-end latency of a user exceeds 16ms, indicating that the computational resource decision is invalid. This demonstrates that, given limited computational resources, our proposed CRA sub-scheme based on BO-DDPG can serve more AR users compared to DDPG and BO.
[0246] The relationship between the delay and computing power of BO-DDPG, DDPG and BO algorithms is as follows: Figure 8 As shown in Figure 3, the delays of the three algorithms are inversely proportional to the computing power, among which BO-DDPG achieves the lowest delay, which indicates that under the same computing power, BO-DDPG has the best delay performance.
[0247] The relationship between the delay and interaction intensity of BO-DDPG, DDPG and BO algorithms is as follows: Figure 9 As shown in Figure 3, as the interaction intensity increases, the delay of the three algorithms also increases, among which BO-DDPG achieves the lowest delay, which indicates that under the condition of limited computing resources, compared with DDPG and BO, our proposed CRA sub-scheme based on BODDPG meets the interaction requirements of more AR users.
[0248] Figure 10 The convergence performance of different algorithms for solving the CASC problem is demonstrated. It can be observed that the BiGAT-DQN algorithm almost converges in about 1800 rounds, and the GAT-DQN with MAH almost converges at 100, 200, 300, 400, and 500 computing power (MHz) and 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, and 1 rounds, while the traditional DQN algorithm almost converges in about 2100 rounds. The BiGAT-DQN algorithm without MAH almost converges in 1800 rounds, while the traditional DQN algorithm almost converges in about 2100 rounds.
[0249] The BiGAT-DQN algorithm achieved the highest average cumulative reward because ICI can be effectively reduced based on the bidirectional and multi-dimensional interference characteristics captured by the Bi-GAT mechanism. The convergence value of the BiGAT-DQN algorithm with the GAT-DQN algorithm is similar to that of the BiGAT-DQN algorithm without the algorithm, but the BiGAT-DQN algorithm converges faster. This is because the BiGAT-DQN algorithm without the algorithm allows the agent to distinguish whether changes in the wireless channel state are due to its own behavior or the behavior of other agents, which enhances its learning efficiency. Without the help of BiGAT and the multi-head attention mechanism, the traditional DQN algorithm would converge to the lowest average cumulative reward.
[0250] Figure 11-13 The relationship between the average offset rate and the number of AR users, bandwidth, and interaction intensity in different algorithms is shown, namely the relationship between BiGAT DQN, GAT-DQN and, BiGATDQN and DQN. Figure 11 It can be seen that the average deviation rate of all algorithms increases with the increase of the number of AR users. This means that the higher the density of AR users in the cell, the lower the spectrum efficiency. Figure 12 ,As the interaction intensity on each agent increases, the average ,drift rate of all algorithms decreases because the amount of interaction ,data to be processed and transmitted increases. Figure 13 As channel bandwidth increases, the average drift rate of all algorithms decreases, as higher bandwidth leads to faster transmission rates. Compared to other algorithms, our proposed BiGAT-DQN achieves the lowest position drift rate across various conditions, including the number of AR users, bandwidth, and interaction intensity. These results demonstrate that by analyzing the characteristics of mutual ICI between AR users, BiGAT-DQN can fully utilize channel resources and minimize the position drift rate caused by the processing power of heterogeneous devices.
[0251] Figure 14 The figure shows how the position offset rate varies with the number of AR users under different resource allocation schemes (i.e., TD-RA, TD-RAw.o.CRA, TD-RAw.o.CA, and TD-RAw.o.SC). It can be seen that the position offset of all methods increases with the number of AR users, but our proposed TD-RA consistently outperforms the other schemes, indicating that TD-RA is more suitable for dense network environments than other schemes.
[0252] Figure 15The figure shows how the position drift rate varies with computing capacity for different resource allocation schemes (i.e., TD-RA, TD-RA w.o.CRA, TD-RA w.o.CA, and TD-RA w.o.SC). It can be seen that under the same computing capacity, the position drift rate of our proposed TD-RA scheme consistently outperforms other schemes, especially when the computing capacity is small. This demonstrates that the TD-RA scheme is suitable for environments with a wide range of computing capacities, especially those with limited computing power.
[0253] Figure 16 The figure shows how the position deviation rate varies with bandwidth for different resource allocation schemes (i.e., TD-RA, TD-RA w.o.CRA, TD-RA w.o.CA, and TD-RA w.o.SC). The figure shows that the position deviation rate of our proposed TD-RA scheme is consistently lower than that of other schemes, regardless of whether the bandwidth is high or low. Its superior performance is particularly evident at low bandwidths. This demonstrates that the TD-RA scheme is suitable for a variety of bandwidth environments, especially those with limited bandwidth.
Claims
1. A dual-time-scale MuAR system resource allocation method in a 6G network framework, characterized in that: The following steps are involved: S1: The video refresh cycle of the AR device is set as a large time scale. Within each large time scale, a deep deterministic policy gradient algorithm enhanced by Bayesian optimization is applied to make computing resource allocation decisions to minimize the computational latency of 3D registration and rendering tasks. S2: The duration of a subframe is set as a small time scale. Within each small time scale, a bidirectional graph attention network deep Q network is applied to make channel allocation and subframe configuration decisions to minimize the position offset rate of virtual objects presented on different display devices.
2. The dual-time-scale MuAR system resource allocation method in the 6G network framework according to claim 1, characterized in that: In S1, the deep deterministic policy gradient algorithm enhanced by Bayesian optimization is applied including: S1-1: Statistical modeling of the conditional probability of events based on historical samples; S1-2: Optimize the utility function related to the event occurrence, and output the optimal computing resource allocation decision through the gradient descent method based on the given posterior distribution S1-3: Training DNN is a function approximator that obtains the optimal computing resource allocation decision based on the current strategy According to the Bellman equation, we can derive and State-action value and S1-4: By comparing two state-action values and The largest state-action value is selected as the final output action.
3. The dual-time-scale MuAR system resource allocation method in the 6G network framework according to claim 2, characterized in that: S1-1: The process of statistically modeling the conditional probability of an event based on historical samples includes: Bayesian optimization: Computational resource requirements are defined as triplets in and Represents the number of CPU cycles required to perform 3D registration, host rendering, and resolver rendering, respectively. Host rendering refers to the rendering of virtual objects initiated by the AR user when running as a host; resolver rendering refers to the rendering of virtual objects initiated by other hosts when running as a resolver. The distribution of the number of CPU cycles required for 3D registration is simulated as a Pareto distribution, expressed as: Among them, a, b, and c represent shape, scale, and threshold respectively; Its cumulative distribution function can be calculated as follows: The number of CPU cycles required for host rendering is expressed as a truncated log-normal distribution: Where μ represents the mean and δ represents the standard deviation; Its cumulative distribution function can be calculated as follows: The distribution of the total number of CPU cycles required for parser rendering is simulated as an exponential distribution as follows: Where λ represents the interaction strength; Its cumulative distribution function can be calculated as follows:
4. The dual-time-scale MuAR system resource allocation method in the 6G network framework according to claim 2, characterized in that: S1-2: The process of optimizing the utility function associated with the event and outputting the optimal computing resource allocation decision using the gradient descent method based on the given posterior distribution includes: Define a function that maps CRA variables to computational resource requirements as follows: where Ψ n,m is the CRA variable, namely BS B n For AR users m The decision variable of the allocated computing resources,∈,is the error term, which is regarded as Gaussian noise with mean 0; According to Bayes' theorem, the posterior distribution and the prior distribution and likelihood function of g(Ψ) Related, as shown below: in It is a historical sample set; Define a utility function to describe the expected improvement of the function value g(Ψ), expressed as: g * (Ψ) represents the maximum function value in the past; The gradient descent algorithm is used to solve the utility function maximization problem. By maximizing the utility function, the optimal CRA decision for the next time slot is found.
5. The dual-time-scale MuAR system resource allocation method in the 6G network framework according to claim 2, characterized in that: The DNN trained in S1-3 is a function approximator, which obtains the optimal computing resource allocation decision based on the current strategy. The process includes: The computational delay minimization problem is formulated as a Markov decision process consisting of a state space, an action space, and a reward function. Taking the base station as an intelligent agent, the nth base station B n The elements in the state space of Represents AR user U m The computing resource requirements of three types of tasks: 3D registration, host rendering and parser rendering; AR user U m The action space includes in and denote the proportion of computing resources allocated to perform 3D registration, host rendering, and parser rendering, respectively; the reward function is defined as the negative of the total computational latency, i.e. in and denote the computational latency of performing 3D registration, host rendering, and resolver rendering, respectively; The framework consists of a training DNN and a target DNN; the training DNN is a function approximator used to obtain the optimal action a according to the current policy π(θ) * (t) and the state-action value of the action As shown below: The target DNN is used to evaluate the best action based on the following factors: Among them, θ and ω are the weight parameters of the training DNN and the target DNN respectively; the weight parameters of the training DNN are updated through back propagation, and its error loss is: Among them, N R is the size of the random sampling batch of the replay memory; the target DNN copies the weight parameters of the training DNN every fixed step.
6. The dual-time-scale MuAR system resource allocation method in the 6G network framework according to claim 1, characterized in that: In S2, the process of applying the bidirectional graph attention network deep Q network to make channel allocation and subframe configuration decisions includes: S2-1: Model the mobile multi-user augmented reality system as a B A directed graph composed of subgraphs captures the bidirectional interference dependencies between agents through a bidirectional graph attention network; S2-2: Generate joint optimization decisions for channel allocation and subframe configuration through a deep Q network to minimize the average position offset rate.
7. The dual-time-scale MuAR system resource allocation method in the 6G network framework according to claim 6, characterized in that: In S2-1, the mobile multi-user augmented reality system is modeled as a B The process of capturing the bidirectional interference dependency between agents through a bidirectional graph attention network in a directed graph composed of subgraphs includes: The mobile multi-user augmented reality system is modeled as a B A directed graph composed of subgraphs, using Indicates that the nth subgraph is represented by Indicates that represents AR users, ε:={e m,m′ } represents the interference received by node m from node m′; The interference received by node m from other nodes and the interference caused by node m to other nodes are input into the MLP layer respectively. The MLP layer extracts the potential features of each node and converts them into high-dimensional features, which are expressed as: in and is the weight parameter that needs to be trained in MLP, σ(·) represents the ReLu activation function, I n,m and They are the interference received by node m from other nodes and the interference caused by node m to other nodes. The coefficients between nodes can be realized by multi-head attention as shown below: m,m′ =ATT(Wh m ,Wh m′ ); Among them, ATT(·) is the self-attention mechanism, W is the corresponding weight parameter, and the formula expresses the importance of node m′ to node m; The masked and normalized attention coefficients can be implemented as follows: Where softmax(·) is the softmax function and LeakyReLU(·) is the LeakyReLU function. The state embedding is obtained by connecting Q independent attention mechanisms in series, and its value is: Where ||(·) represents the concatenation operation and Q represents the number of attention heads; The state embedding is passed to the fully connected layer for feature fusion, and its calculation formula is:
8. The dual-time-scale MuAR system resource allocation method in the 6G network framework according to claim 6, characterized in that: The process of generating a joint optimization decision for channel allocation and subframe configuration through a deep Q network includes: The state space of partially observable Markov decisions is represented by express; The training DNN in the deep Q network updates the weight parameters μ by minimizing the mean squared error loss function, which is: in, is the mini-batch size, and the weight vector μ(t+1) can be solved using the gradient descent method, i.e.: Where Γ represents the learning rate.
Citation Information
Patent Citations
Deep reinforcement learning-based dual-time-scale resource allocation method combining service caching, communication and calculation
CN117156492A
Video streaming media double-time-scale wireless transmission method and system and storage medium
CN119946329A
System and method for variable time scale for multi-player games
US20110072094A1
System and method for resource management in heterogeneous wireless networks
US20160037387A1