Generative learning assisted cellular-free communication wireless resource management method for time delay certainty
By constructing a virtual CMDP module and pre-training a deep reinforcement learning algorithm, the resource allocation of the non-cellular MIMO system is optimized, solving the problems of low sampling efficiency and cold start of DRL in the non-cellular MIMO system, and achieving high energy efficiency and time-deterministic communication.
Patent Information
- Application Number
- CN202510888597.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In cellular MIMO systems, the low sampling efficiency of DRL and the violation of strict delay constraints during the cold start exploration phase make it difficult for radio resource management to meet the requirements of ultra-reliable low-latency communication.
A virtual CMDP module is constructed and pre-trained using a deep reinforcement learning algorithm. Combined with VAE-ChMDN and EA-CGMM models, the initial state distribution and state transition probabilities are generated to optimize the allocation of downlink resources in cellular-free MIMO, thereby maximizing energy efficiency while satisfying latency constraints.
It improves the sample efficiency and convergence speed of DRL in non-cellular MIMO systems, avoids the high time delay constraint violation rate during cold start, and improves energy efficiency and time delay determinism.
Smart Images

Figure CN120957232A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication technology, and particularly relates to a generative learning-assisted method for managing wireless resources in non-cellular communication with time-deterministic determinism. Background Technology
[0002] With the surge in mission-critical applications such as industrial automation, telesurgery, autonomous driving, and embodied intelligence, future wireless networks are placing unprecedented demands on latency-deterministic communication. This requires not only ultra-low latency but also stringent guarantees of bounded jitter and high reliability. Current ultra-reliable low-latency (ULR) communication technologies provide the physical layer foundation for deterministic communication using short packet transmission and finite block length coding techniques. However, limited wireless communication resources pose a challenge to the stringent requirements of ULR. Therefore, agile and adaptable wireless resource management is crucial, including techniques such as user scheduling, time and frequency resource allocation, and beamforming to reduce packet queuing and transmission latency, thereby improving latency deterministic assurance levels.
[0003] Cellular MIMO architecture is considered a promising technology for enhancing the deterministic communication capabilities of wireless networks, employing dense access point deployment and cooperative beamforming to achieve ubiquitous connectivity with enhanced spatial diversity. However, the high degree of freedom in resource allocation is central to the efficient operation of cellular MIMO architecture, leading to a sharp increase in the analytical complexity and computational overhead of mathematical optimization methods. These challenges become particularly pronounced with increasing UE density and antenna count. Dynamic Replication (DRL) offers significant potential for dynamic optimization of large-scale wireless networks by achieving model-free adaptation to complex system dynamics. However, directly deploying DRL in practical cellular MIMO systems faces two key limitations: 1) low sampling efficiency of DRL, requiring high sampling costs in practical implementation; 2) the cold start exploration phase, which inevitably violates strict latency constraints, posing a significant risk to QoS-sensitive applications. Summary of the Invention
[0004] Purpose of the invention: In order to solve the problems existing in the prior art, the present invention provides a generative learning-assisted method for non-cellular communication wireless resource management oriented towards time-deterministic determinism.
[0005] Technical Solution: This invention discloses a generative learning-assisted method for managing radio resources in non-cellular communication with time-deterministic determinism, specifically including the following steps:
[0006] Step 1: Construct a cellular-free MIMO downlink resource allocation scenario, focusing on optimizing the energy efficiency of the cellular-free MIMO-OFDM downlink system while satisfying latency violation rate constraints;
[0007] Step 2: Construct a virtual CMDP module, mapping the resource allocation scenario constructed in Step 1 to the virtual CMDP module: This involves mapping the channel state information R at time slot t... t and queue status information Q t As state s t ;Schedule consecutive TFUs Continuous beam selection The power ratio allocated by base station b to user u As the action at time slot t, where k represents the sub-band, k = 1, 2, ..., K; K represents the total number of sub-bands, b represents the base station, b = 1, 2, ..., B; B represents the total number of base stations, u represents the user, u = 1, 2, ..., U; U represents the total number of users; This is for discrete TFU scheduling. For binary variables, The time interval indicates that user u is scheduled in the TFU{t,k} of base station b, where TFU{t,k} represents the minimum schedulable time-frequency unit of the k-th sub-band in the t-th time slot. This indicates that user u was not invoked in base station b's TFU{t,k}. Given the power allocated by base station b to user u, construct the cost function and reward function of the virtual CMDP module based on the energy efficiency optimization problem in step 1;
[0008] Step 3: Pre-train the deep reinforcement learning algorithm based on the constructed virtual CMDP module, then apply the deep reinforcement learning algorithm to the real environment, adjust the parameters in the deep reinforcement learning algorithm, and obtain the final deep learning algorithm.
[0009] Step 4: Make a final decision using a deep learning algorithm.
[0010] Furthermore, the expression for the energy efficiency optimization problem in step 1 is as follows:
[0011]
[0012] in, Represents the energy efficiency function. The expression is as follows:
[0013]
[0014] In this system, a minimum schedulable time-frequency unit (MSFMU) of a sub-band consists of C consecutive subcarriers, and a minimum schedulable time unit consists of N OFDM symbols. The beamforming vector set by base station b for user equipment u. Let Δt be the number of bits transmitted to user u via code blocks in time slot t. sThe time slot length, Represents a set of base stations. Represents the set of sub-bands. Represents a set of users. Let Pr(.) represent the set of time slots, and d represent the probability. u,a (t) represents the delay of the a-th data packet arriving at user u in time slot t, D u For the deadline, η u It is the packet delay violation probability threshold, P max It is the base station's maximum power budget. A beam codebook with M codewords, f m Represents the m-th codeword, ||f m || = 1.
[0015] Furthermore, the number of bits transmitted to user u via code blocks in time slot t. The expression is as follows:
[0016]
[0017] in, For user u, the signal-to-interference-plus-noise ratio, Q -1 (·) is the inverse function of the Gaussian Q-function, ∈ u For block error rate, It is channel dispersion; The expression is as follows:
[0018]
[0019] in, This indicates that the signal is sent to the downlink channel of user u. The noise variance in the channel;
[0020] The expression is as follows:
[0021]
[0022] Furthermore, the channel state information R at time slot t t Including probe beam measurements Queue status information Q t This includes cache occupancy status and critical latency traffic status within the cache. The expression is:
[0023]
[0024] Where H represents the conjugate transpose operation. This indicates that the m-th codeword f mAs the beamforming vector set by base station b for user equipment u;
[0025] The reward function in the virtual CMDP module is the single-slot energy efficiency, and its expression is:
[0026]
[0027] Among them, s t Indicates the state of time slot t, a t Let r(.) represent the action in time slot t, and r(.) be the reward function.
[0028] The expression for the cost function is:
[0029] c(s t ,a t )=Pr(d u,a (t)>D u )
[0030] Where c(.) represents the cost function, and c(s) t ,a t )≤d, d=η u ;
[0031] The following formula will be used to... Convert to
[0032]
[0033] in, Indicates rounding down;
[0034] Discrete beam selection is achieved using the following formula. Convert to continuous beam selection
[0035]
[0036] The expression is:
[0037]
[0038] in,
[0039] Furthermore, during pre-training, Lagrange multipliers are used to multiply c1(s) by... t ,a t If )≤d1, it is converted to an unconstrained form.
[0040] Furthermore, the virtual CMDP module is constructed specifically as follows: random sampling is performed from the action space, and transition tuple data is collected through interaction with the environmental CMDP (s t,a t ,r t ,c t ,s t+1 The offline dataset is used to construct the reward and cost function module, the initial state distribution module, and the state transition probability module in the virtual CMDP module.
[0041] Furthermore, a reward and cost function module is constructed using a KAN network, specifically by using two KAN networks to predict the reward and cost values respectively.
[0042] Furthermore, the VAE-ChMDN model is used as the initial state distribution module, and the initial state is generated using the initial state distribution module, specifically as follows:
[0043] Step a: The VAE-ChMDN model includes an inference network and a generator network. The inference network uses a hidden Gaussian random variable as the latent representation of the initial state s0. Then, the hidden Gaussian random variable is input into the generator network to obtain Gaussian mixture parameters. The Gaussian mixture parameters include: weights π, π = (π1, π2, ..., π). G ), mean vector μ, μ = (μ1, μ2, ..., μ) G And the Cholesky factor U, U = (U1, U2, ..., U G G represents the total number of parameters in any mixture parameter. The expression for a Gaussian mixture model with G mixture components is as follows:
[0044]
[0045] Where p g (s0) is the probability density of the g-th multivariate Gaussian distribution;
[0046] Step b: The loss function of the VAE-ChMDN model is:
[0047]
[0048] in, For the prior distribution, For variational distribution, μ en , σ en Let represent the mean and standard deviation of the hidden Gaussian random variable, respectively, and KL represent the KL divergence.
[0049] Step c: Use the trained VAE-ChMDN model to call the generator network to generate the initial state distribution. The g-th mixture component is selected based on the mixture weight π, and then the initial state sample is sampled using the following formula.
[0050] ∈ represents the noise from the sampling.
[0051] Furthermore, the state transition probability module uses the EA-CGMM algorithm to infer the distribution of the next state given the current state-action pair, specifically:
[0052] Step 2.1: Using the VAE-ChMDN model, obtain the Gaussian mixture parameters. Each mixture parameter has J parameters. Then, convert the state transition tuple (s...) t+1 ,a t ,s t It is explicitly modeled as a Gaussian mixture model consisting of J mixture components;
[0053] Step 2.2: Calculate the j-th mean vector μ j Split into s t+1 Subvectors corresponding to the dimension and (a t ,s t Subvectors corresponding to the dimension Let the j-th Cholesky factor U j Split into AND and s t+1 Subvectors corresponding to the dimension With (a) t ,s t Subvectors corresponding to the dimension Will U j The remaining subvectors are denoted as j = 1, 2, ..., J; the conditional distribution of the j-th mixture component of the Gaussian model satisfies the following formula:
[0054]
[0055] in, For the current state action pair, This is the current state. For the current action, s t+1 For the state at the next moment, μ′ represents the j-th mixture component; j and Σ′ j The expression is as follows:
[0056]
[0057] Step 2.3: Update the weights of the Gaussian mixture model using a reweighting scheme based on dual evidence, obtaining the updated weights π′, where π′ = [π′1, π′2, ..., π′]. J ];
[0058] Step 2.4: Generate the Gaussian mixture model for the next state To obtain the next state distribution, This represents the j-th mixed component.
[0059] Furthermore, step 2.3 specifically includes:
[0060] Construct a statistical significance mask m = [m1, m2, ..., m J ]:
[0061]
[0062] in, For normalized residuals: dim(.) calculates the dimension, y t A random variable representing the marginal distribution of j mixed components;
[0063]
[0064] Update the weight vector according to the following formula:
[0065]
[0066] in This represents the marginal probability density.
[0067] Beneficial effects: Compared with the DRL algorithm that only interacts with the real environment, the pre-training framework based on virtual CMDP in this invention can provide a good pre-training environment for DRL, enhance the initial and final EE performance of DRL in real non-cellular MIMO downlink transmission scenarios, avoid the risk of high latency constraint violation rate during cold start, and improve its sample efficiency and accelerate the convergence speed. Attached Figure Description
[0068] Figure 1 This is a flowchart of an embodiment of the present invention;
[0069] Figure 2 This is a schematic diagram of the process of a generative learning-assisted non-cellular communication wireless resource management method based on time-deterministic determinism according to the present invention;
[0070] Figure 3 This is an architecture diagram of the VAE-ChMDN of this invention;
[0071] Figure 4 The graph shows the evolution of the EE performance of the DRL strategy over 1500 pre-training rounds.
[0072] Figure 5 The graph shows the evolution of the latency constraint violation rate of the DRL policy in 1500 pre-training rounds.
[0073] Figure 6A comparison of round reward performance between a pre-trained DRL deployed in a simulation environment for 1500 rounds and a DRL without pre-training.
[0074] Figure 7 A comparison of round cost performance between deploying a DRL with 1200 pre-trained rounds in a simulation environment for 1500 rounds of training and without pre-training. Detailed Implementation
[0075] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0076] like Figure 1 As shown, the present invention specifically comprises:
[0077] Step 1: System Model. Construct a cellular-free MIMO downlink resource allocation scenario, including transmission model, queue and latency model, and establish a system energy efficiency optimization problem with latency violation rate constraints.
[0078] Step S101: Consider a scenario where B base stations simultaneously serve U users in a cellular MIMO-OFDM downlink scenario, with system bandwidth B w It is divided into K sub-bands, each defining a cooperative base station set. Service users are grouped into Subband set Each base station is equipped with M antennas, and each user is equipped with a single antenna. Each subband is a minimum schedulable frequency unit, consisting of C consecutive subcarriers. A minimum schedulable time unit is defined as consisting of N OFDM symbols, where one OFDM symbol has C subcarriers, and the slot length is... Therefore, a minimum schedulable time-frequency unit (TFU) has CN resource elements (REs).
[0079] During the downlink signal transmission phase, the base station cooperatively transmits the signal on the RB to the scheduled user. This embodiment considers transmission within a time window consisting of T time slots, denoted as […]. binary variable This indicates that user u is scheduled in the TFU{t,k} of base station b, and is 0 otherwise. TFU{t,k} represents the minimum schedulable time-frequency unit (TFU) of the k-th sub-band in the t-th time slot.
[0080] For TFU{t,k}, RE{n,c}, where n∈{1,...,N} and c∈{1,...,C} represent the OFDM symbol index and subcarrier index respectively, given the beamforming vector set by base station b for user equipment u. in It is a beam codebook with M codewords, satisfying ||f|| m || = 1, In {f1,...,f M The selection is made within the range of}; and the power allocated by base station b to user equipment u, i.e. The transmitted signal of base station b can be obtained as follows:
[0081]
[0082] in, It is the unit energy modulation symbol for user u, i.e. Use the superscript H to denote the conjugate transpose operation. This indicates that the signal is sent to the downlink channel of user u. This represents additive white Gaussian noise in the channel, with a noise variance of: The received signal of user u can be represented as:
[0083]
[0084] The signal-to-interference-plus-noise ratio (SIR) of user u can be expressed as:
[0085]
[0086] Considering practical finite block length coding, given a block error rate ∈ u The number of bits transmitted to user equipment u via code blocks in time slot t is:
[0087]
[0088] Q -1 (·) is the inverse function of the Gaussian Q-function. It is channel dispersion, calculated using the following formula:
[0089]
[0090] Step S102, let A u (t) represents the number of data packets arriving at user u in time slot t, and the size of the a-th data packet arriving at user u in time slot t is denoted as . Where a∈{1,2,...,A} u (t)}. General assumption A u (t) is independently and identically distributed across all time slots, with an average arrival rate of λ. uAssume that for the same user u, all arriving data packets have the same deadline, denoted as D. u (in units of time slots). Each user u maintains a finite buffer with a maximum capacity of [missing information]. Let Q u (t) represents the queue length at the beginning of time slot t, and its change can be expressed as:
[0091]
[0092] in when Data packets arriving at Q will be discarded immediately. u (t+1) is always less than Existing data packets are not discarded, and scheduling is performed using an earliest deadline priority strategy, following a first-in, first-out (FIFO) principle within each user's buffer. For the a-th data packet arriving in time slot t, its delay is:
[0093]
[0094] The delay formula means finding the minimum τ such that the number of uncompleted old data bits + the total number of bits of the first a packets arriving ≤ the number of data bits sent from time slot t+1 to τ.
[0095] u
[0096] Step S103: This embodiment aims to maximize the energy efficiency (EE) of a cellular-free MIMO-OFDM downlink system by finding a joint resource allocation strategy for TFU, beamforming vector, and power allocation, while satisfying the time delay violation rate constraint. This problem can be described as follows:
[0097]
[0098] Where η u It is the packet delay violation probability threshold, P max The maximum power budget (EE) of the base station is calculated using the following formula:
[0099]
[0100] Step 2, Overall Scheme: Model the EE optimization problem of the system with delay violation rate constraints as CMDP, construct a pre-training environment based on virtual CMDP, use deep reinforcement learning algorithms (such as PPO algorithm) for pre-training, and then fine-tune in the real environment.
[0101] Step 2 describes the overall design of the pre-training framework based on virtual CMDP as follows: Figure 2 As shown, the specific steps include the following:
[0102] Step S201: Perform CMDP modeling on the resource allocation optimization problem with delay violation rate constraints:
[0103] State space: the state s of time slot t t From channel state information R t and queue status information Q t Composition. Channel state is determined by probe beam measurements. It means that among them This indicates that the m-th codeword f m As the beamforming vector set by base station b for user equipment u, it quantifies the beamforming efficiency of each subband between the base station and user pair. The queue state is represented as... in The cache occupancy status is indicated by the number of data packets. The critical latency traffic state in the buffer is represented by the number of packets with the maximum latency.
[0104] Motion space: Motion space Generated by the agent's policy network, these three continuous variables need to be transformed to produce the actual TFU scheduling, beam selection, and power allocation (the continuous actions output by the PPO's policy network). Specifically,
[0105] for definition As a floor function, the following transformation can be used to schedule discrete TFUs. Beam selection
[0106]
[0107] use This represents the power allocation ratio of base station b, where The actual power allocation for base station b is then:
[0108]
[0109] Reward Function: The reward function is designed as a single-slot EE.
[0110]
[0111] Cost function: The cost function is defined as the single-slot delay constraint violation rate, i.e.:
[0112]
[0113] The corresponding cost threshold is set to d = η u The cost value is less than or equal to d(c(s) t ,at The constraints of the base station maximum power, TFU, and beamforming vector have been satisfied in the design of the action space.
[0114] Step S202: The data acquisition module first uses a behavioral strategy. Interact with the real environment to collect transfer tuples (s t ,a t ,r t ,c t ,s t+1 An offline dataset is constructed, and then reward and cost prediction modules, initial state distribution modules, and state transition distribution modules are trained based on the offline dataset. These modules together constitute a virtual CMDP.
[0115] Step S203: Using the PPO algorithm based on the primal dual method, first use Lagrange multipliers to transform c(s) t ,a t If ≤ d, the policy is transformed into an unconstrained form, and then the PPO algorithm is used to achieve long-term optimization of the policy. During the pre-training phase, the agent interacts with the virtual CMDP and learns a policy that satisfies the constraints through safe exploration.
[0116] Step S204: Subsequently, the pre-trained policy is fine-tuned by interacting with the real environment a limited number of times, starting with a low-latency constraint violation rate to learn the optimal policy under the current environmental conditions.
[0117] Step 3, Virtual CMDP Modeling: Using randomized behavioral policies to interact with the real-world CMDP environment and collect offline data, driving rewards to build a virtual CMDP composed of a cost prediction module, an initial state distribution module, and a state transition probability module. This includes the following specific steps.
[0118] Step S301: Use a random behavior strategy, i.e., randomly sample actions from the action space and interact with the environment CMDP to collect transition tuple data (s). t ,a t ,r t ,c t ,s t+1 This constitutes an offline dataset.
[0119] Step S302: Use offline datasets to drive KAN to implement the reward and cost function module, VAE-ChMDN to implement the initial state distribution module, and EA-CGMM algorithm to implement the state transition probability module. Together, they constitute a virtual CMDP.
[0120] Step 4, Reward and Cost Prediction Module: Based on the offline dataset, use KAN to fit a deterministic mapping from state-action pairs to rewards and costs.
[0121] Step S401: The reward function and cost function are deterministic mappings from state-action pairs to scalar values. Two KANs are designed for fitting, with inputs of action-state pairs and outputs of predicted rewards and cost values, respectively, as follows:
[0122]
[0123] Step S402: Use the data from the offline dataset (s) t ,a t ,r t ,c t The tuples, using the mean squared error loss between predicted and true values, are used to train two KAN networks via backpropagation. The loss function is expressed as:
[0124]
[0125] Step 5: Based on the initial state data in the offline dataset, model the initial state distribution using VAE-ChMDN.
[0126] The VAE-ChMDN architecture described in step 5 is as follows: Figure 3 As shown.
[0127] Step S501: Use VAE-ChMDN to generate the initial state of the initial state distribution, where the inference network will use hidden Gaussian random variables. As a potential representation of the initial input state s0, where μ en , σ en These represent the mean and standard deviation of the latent variables, respectively. The framework further utilizes a mixture density network (generative network) to output Gaussian mixture parameters: mixture weights π = (π1, π2, ..., π). G The mean vector μ = (μ1, μ2, ..., μ) G ) and the Cholesky factor U = (U1, U2, ..., U G The initial state distribution ρ(s0) is represented as a Gaussian mixture model with G mixture components:
[0128]
[0129] Where π g ∈[0,1] satisfies And p g (s0) is the probability density of the g-th multivariate Gaussian distribution:
[0130]
[0131] in It is the Cholesky factor Ug The obtained full-rank covariance matrix, U g It is an upper triangular matrix with positive elements on the main diagonal, which guarantees Σ g Symmetry and positive definiteness.
[0132] Step S502, the loss function of VAE-ChMDN is derived from the variational distribution. With prior distribution The KL divergence between them, and in the generated Gaussian mixture model The negative log-likelihood of the observed initial state s0 is composed of:
[0133]
[0134] For one of the mixture components in the Gaussian mixture model, the logarithmic density of s0 can be calculated using the following formula:
[0135]
[0136] Step S503: The trained VAE-ChMDN calls the generator network to generate the initial state distribution. First, the g-th mixing component is selected with probability based on the mixing weight π. Then, the initial state sample is sampled using the sampling formula:
[0137]
[0138] This indicates sampling noise.
[0139] Step 6, State Transition Probability Module: Based on the state transition tuples in the offline dataset, use the EA-CGMM algorithm to infer the distribution of the next state given the current state-action pair.
[0140] Step S601: Use the EA-CGMM algorithm to model the state transition distribution. The goal is to model the action pair given the current state. Gaussian mixture model with J mixture elements for inferring the next state Where π′ j It is the weight of the j-th mixture component. The corresponding mean is μ′ j The standard deviation is Σ′ j The algorithm uses a multivariate Gaussian distribution. The algorithm process is as follows.
[0141] Step S602: First, use VAE-ChMDN to process the state transition tuples (s) in the offline dataset. t+1 ,a t ,s t ) is explicitly modeled as a Gaussian mixture model (GMM) consisting of J mixture components, with parameters expressed as follows: Represented as:
[0142]
[0143] Where for j = 1, 2, ..., J,
[0144] Step S603, corresponding to s t+1 and (a t ,s t ), and the j-th mean vector μ j Split into s t+1 Subvectors corresponding to the dimension and (a t ,s t Subvectors corresponding to the dimension Let the j-th Cholesky factor U j Split into AND and s t+1 Subvectors corresponding to the dimension With (a) t ,s t Subvectors corresponding to the dimension Will U j The remaining subvectors are denoted as
[0145]
[0146] Based on the mathematical derivation of the multivariate Gaussian distribution, given the current state and action pair... Under the following conditions, the conditional distribution of the j-th joint Gaussian distribution satisfies:
[0147]
[0148] The conditional mean and conditional variance are:
[0149]
[0150] Step S604, the edge distribution p in the mixed component j j (y t )satisfy To obtain the mixture weights π′=[π′1,π′2,...,π′] of the conditional Gaussian mixture model J ], given Using p j (y t The probabilistic properties of ) can be obtained from two sources: one is the marginal probability density. Another method is the normalized residual, which is calculated as follows:
[0151]
[0152] Step S605: Adaptive Behavioral Strategies with exploration strategies To address the severe distribution shift caused by inconsistencies between different data sets, this algorithm employs a reweighting scheme based on dual evidence. Inspired by the 3σ principle, a threshold is set for the normalized residuals to construct a statistical significance mask m = [m1, m2, ..., m...]. J This binary mask is defined as:
[0153]
[0154] The new weight π′ can be expressed as:
[0155]
[0156] This leads to the Gaussian mixture model of the next state. The next state value in the virtual CMDP can be sampled from it.
[0157] This embodiment collects 30,000 transition tuples (s) using a random behavior strategy. t ,a t ,r t ,c t ,s t+1 After randomly splitting the dataset into 80% for training and 20% for testing, this embodiment first evaluates the performance of each module in the virtual CMDP. The maximum mean difference (MMD) is used to evaluate the difference between the generated initial state samples and the actual initial state samples. The MMD calculated from two sets of samples from the same distribution is close to 0. The experimental results are as follows: the mean absolute errors (MAE) of the reward and constraint predictions implemented by KAN are 0.0219 and 0.0494, respectively; the MMD of the initial state distribution implemented by VAE-ChMDN is 0.002; and the MAE of the next state prediction implemented by the EA-CGMM algorithm is 0.0427. The evaluation results demonstrate the effectiveness and accuracy of the implementation of each module in the virtual CMDP driven by the offline dataset.
[0158] like Figure 4 and Figure 5 As shown, DRL improves EE and reduces latency constraint violation rate by interacting with the virtual CMDP in the constructed offline pre-training environment. Finally, it achieves stable EE convergence of 27 Mbit / s and stable latency constraint violation rate convergence of about 2% in 1200 rounds.
[0159] like Figure 6 and Figure 7 As shown, the DRL pre-trained for 1200 rounds achieved a 47% increase in initial EE compared to the unpre-trained DRL, and a 3.7% improvement in final EE performance with 50% fewer exploration steps. Furthermore, the pre-trained model achieved a low initial latency constraint violation rate of 2%, demonstrating a safe warm start. The pre-training results demonstrate the accuracy of virtual CMDP modeling and the effectiveness of the offline pre-training framework.
[0160] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.
Claims
1. A generative learning-assisted method for time-deterministic non-cellular communication radio resource management, characterized in that, Specifically, the steps include the following: Step 1: Construct a cellular-free MIMO downlink resource allocation scenario, focusing on optimizing the energy efficiency of the cellular-free MIMO-OFDM downlink system while satisfying latency violation rate constraints; Step 2: Construct a virtual CMDP module, mapping the resource allocation scenario constructed in Step 1 to the virtual CMDP module: This involves mapping the channel state information R at time slot t... t and queue status information Q t As state s t ; Continuous TFU scheduling Continuous beam selection The power ratio allocated by base station b to user u As the action at time slot t, where k represents the sub-band, k = 1, 2, ..., K; K represents the total number of sub-bands, b represents the base station, b = 1, 2, ..., B; B represents the total number of base stations, u represents the user, u = 1, 2, ..., U; U represents the total number of users; For discrete TFU scheduling. For binary variables, The time interval indicates that user u is scheduled in the TFU{t,k} of base station b, where TFU{t,k} represents the minimum schedulable time-frequency unit of the k-th sub-band in the t-th time slot. This indicates that user u was not invoked in base station b's TFU{t,k}. Given the power allocated by base station b to user u, construct the cost function and reward function of the virtual CMDP module based on the energy efficiency optimization problem in step 1; Step 3: Pre-train the deep reinforcement learning algorithm based on the constructed virtual CMDP module, then apply the deep reinforcement learning algorithm to the real environment, adjust the parameters in the deep reinforcement learning algorithm, and obtain the final deep reinforcement learning algorithm. Step 4: Make a final decision using a deep reinforcement learning algorithm.
2. The generative learning-assisted non-cellular communication radio resource management method based on time-deterministic determinism according to claim 1, characterized in that, The expression for the energy efficiency optimization problem in step 1 is as follows: in, Represents the energy efficiency function. The expression is as follows: In this system, a minimum schedulable time-frequency unit (MSFMU) of a sub-band consists of C consecutive subcarriers, and a minimum schedulable time unit consists of N OFDM symbols. The beamforming vector set by base station b for user equipment u. Let Δt be the number of bits transmitted to user u via code blocks in time slot t. s The time slot length, Represents a set of base stations. Represents the set of sub-bands. Represents a set of users. Let Pr(.) represent the set of time slots, and d represent the probability. u,a (t) represents the delay of the a-th data packet arriving at user u in time slot t, D u For the deadline, η u It is the packet delay violation probability threshold, P max It is the base station's maximum power budget. For a beam codebook with M codewords, f m Represents the m-th codeword, ||f m || = 1.
3. The generative learning-assisted non-cellular communication radio resource management method based on time-deterministic determinism according to claim 2, characterized in that, The number of bits transmitted to user u via code blocks in time slot t The expression is as follows: in, For user u, the signal-to-interference-plus-noise ratio, Q -1 (·) is the inverse function of the Gaussian Q-function, ∈ u For block error rate, It is channel dispersion; The expression is as follows: in, This indicates that the signal is sent to the downlink channel of user u. The noise variance in the channel; The expression is as follows:
4. The generative learning-assisted non-cellular communication radio resource management method for time-deterministic determinism according to claim 2, characterized in that, Channel state information R at time slot t t Including probe beam measurements Queue status information Q t This includes cache occupancy status and critical latency traffic status within the cache. The expression is: Where H represents the conjugate transpose operation. This indicates that the m-th codeword f m As the beamforming vector set by base station b for user equipment u; The reward function in the virtual CMDP module is the single-slot energy efficiency, and its expression is: Among them, s t Indicates the state of time slot t, a t Let r(.) represent the action in time slot t, and r(.) be the reward function. The expression for the cost function is: c(s t ,a t )=Pr(d u,a (t)>D u , Where c(.) represents the cost function, and c(s) t ,a t )≤d, d=η u ; The following formula will be used to... Convert to in, Indicates rounding down; Discrete beam selection is achieved using the following formula. Convert to continuous beam selection The expression is: in, 5. The generative learning-assisted non-cellular communication radio resource management method for time-deterministic determinism according to claim 4, characterized in that, During pre-training, Lagrange multipliers are used to multiply c(s) t ,a t )≤d is converted to an unconstrained form.
6. The generative learning-assisted non-cellular communication radio resource management method based on time-deterministic determinism according to claim 2, characterized in that, The virtual CMDP module is constructed as follows: it randomly samples from the action space and interacts with the environment CMDP to collect transition tuple data (s). t ,a t ,r t ,c t ,s t+1 The offline dataset is used to construct the reward and cost function module, the initial state distribution module, and the state transition probability module in the virtual CMDP module.
7. The generative learning-assisted non-cellular communication radio resource management method for time-deterministic determinism according to claim 6, characterized in that, The reward and cost function modules are constructed using KAN networks. Specifically, two KAN networks are used to predict the reward and cost values respectively.
8. The generative learning-assisted non-cellular communication radio resource management method for time-deterministic determinism according to claim 6, characterized in that, The VAE-ChMDN model is used as the initial state distribution module, and the initial state is generated using the initial state distribution module, specifically as follows: Step a: The VAE-ChMDN model includes an inference network and a generator network. The inference network uses a hidden Gaussian random variable as the latent representation of the initial state s0. Then, the hidden Gaussian random variable is input into the generator network to obtain Gaussian mixture parameters. The Gaussian mixture parameters include: weights π, π = (π1, π2, ..., π). G ), mean vector μ, μ = (μ1, μ2, ..., μ) G And the Cholesky factor U, U = (U1, U2, ..., U G G represents the total number of parameters in any mixture parameter. The expression for a Gaussian mixture model with G mixture components is as follows: Where p g (s0) is the probability density of the g-th multivariate Gaussian distribution; Step b: The loss function of the VAE-ChMDN model is: in, As a prior distribution, For variational distribution, μ en , σ en Let represent the mean and standard deviation of the hidden Gaussian random variable, respectively, and KL represent the KL divergence. Step c: Use the trained VAE-ChMDN model to call the generator network to generate the initial state distribution. The g-th mixture component is selected based on the mixture weight π, and then the initial state sample is sampled using the following formula. ∈ represents the noise from the sampling.
9. The generative learning-assisted non-cellular communication radio resource management method for time-deterministic determinism according to claim 8, characterized in that, The state transition probability module uses the EA-CGMM algorithm to infer the distribution of the next state given the current state-action pair, specifically: Step 2.1: Using the VAE-ChMDN model, obtain the Gaussian mixture parameters. Each mixture parameter has J parameters. Then, convert the state transition tuple (s...) t+1 ,a t ,s t It is explicitly modeled as a Gaussian mixture model consisting of J mixture components; Step 2.2: Calculate the j-th mean vector μ j Split into s t+1 Subvectors corresponding to the dimension and (a t ,s t Subvectors corresponding to the dimension Let the j-th Cholesky factor U j Split into AND and s t+1 Subvectors corresponding to the dimension With (a) t ,s t Subvectors corresponding to the dimension Will U j The remaining subvectors are denoted as The conditional distribution of the j-th mixture component of the Gaussian model satisfies the following formula: in, For the current state action pair, This is the current state. For the current action, s t+1 For the state at the next moment, μ′ represents the j-th mixture component; j and Σ′ j The expression is as follows: Step 2.3: Update the weights of the Gaussian mixture model using a reweighting scheme based on dual evidence, obtaining the updated weights π′, where π′ = [π′1, π′2, ..., π′]. J ]; Step 2.4: Generate the Gaussian mixture model for the next state To obtain the next state distribution, This represents the j-th mixed component.
10. The generative learning-assisted non-cellular communication radio resource management method for time-deterministic determinism according to claim 9, characterized in that, Step 2.3 specifically involves: Construct a statistical significance mask m = [m1, m2, ..., m J ]: in, For normalized residuals: dim(.) calculates the dimension, y t A random variable representing the marginal distribution of j mixed components; Update the weight vector according to the following formula: in This represents the marginal probability density.
Citation Information
Patent Citations
Non-cellular large-scale MIMO power distribution method based on deep reinforcement learning
CN114268348A
Microgrid space-time perception energy management method based on secure deep reinforcement learning
WO2024108817A1
Cited By
MAC layer data packet time delay prediction and scheduling optimization method based on generative learning
CN121486857A