Resource allocation method and system based on coordinated multipoint transmission and deep reinforcement learning, and storage medium
By adopting resource allocation methods of collaborative multi-point transmission and deep reinforcement learning in multi-cellular networks, URLLC resources are dynamically regulated, and the coexistence problem of eMBB and URLLC resource scheduling is solved, which achieves high reliability and low latency of URLLC, while protecting the performance of eMBB.
Patent Information
- Application Number
- CN202510131403.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The prior art is difficult to effectively schedule eMBB and URLLC resources in a multi-cellular network, which makes it difficult to meet the reliability and delay requirements of URLLC. At the same time, the eMBB resources are drilled in large quantities, resulting in waste of resources.
The resource allocation method based on collaborative multi-point transmission and deep reinforcement learning is adopted. By introducing the BLER interrupt probability of URLLC users, the joint resource allocation problem of eMBB and URLLC is modeled as a CMDP problem, and the twin latency depth deterministic strategy gradient algorithm is used for optimization, and the URLLC pilot length and data transmission symbol number are dynamically regulated.
It realizes the latency and reliability requirements of URLLC under a multi-cellular network, and meets the QoS indicators of URLLC in high mobility scenarios, while effectively reducing the attenuation of eMBB user performance due to URLLC preemption.
Smart Images

Figure CN119997237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular to a resource allocation method, system and storage medium based on coordinated multi-point transmission and deep reinforcement learning. Background Art
[0002] eMBB and URLLC are two typical scenarios under the fifth-generation (5G) and beyond 5G (B5G). Among them, eMBB is mainly aimed at the mobile Internet with explosive traffic growth, providing high-bandwidth services with peak rates of gigabit, such as 8K video services and augmented reality services. URLLC is mainly aimed at mission-critical applications in vertical industries, such as telemedicine, unmanned driving, smart grids and other applications, which have very strict requirements on reliability and latency. Now, URLLC requires 99.99% reliability and 5ms end-to-end (E2E) latency for 32-byte data packets transmitted from 5G, and evolves to 99.9999% reliability and 1ms end-to-end latency in B5G. At the same time, there are also a large number of multi-service converged applications (such as mixed reality) in B5G that require both low-latency and high-reliability URLLC services and high-bandwidth eMBB services. In order to effectively solve the coexistence problem of these two types of services and meet their respective performance requirements, URLLC needs to seize eMBB transmission resources (also known as "punching"). This mechanism aims to meet the stringent reliability and latency requirements of URLLC by sacrificing the eMBB user data rate, thereby achieving the coexistence of the two types of services.
[0003] On the other hand, the Constrained Markov decision process (CMDP) is very effective in handling resource scheduling in a dynamic environment, so this method is commonly used to model resource allocation problems in scenarios where eMBB and URLLC coexist. However, in actual scenarios, the CMDP problem often contains a state space and action space with huge dimensions, which makes the algorithm for solving this problem extremely complex. The resulting problem is that when the time it takes for the algorithm to obtain a resource scheduling strategy exceeds the coherence time of the channel, the strategy will expire and no longer apply, so it is not practical.
[0004] For resource allocation in scenarios where URLLC and eMBB services coexist, existing solutions are mainly based on single-cell networks, while more practical multi-cell networks are rarely considered. We also note that CoMP technology has received increasing attention for 5G / B5G multi-cell networks. CoMP technology uses multiple base stations to send the same information to the same user (often a cell edge user) to reduce the interference to the edge user, thereby improving its throughput performance. However, the existing multi-cell network resource allocation solutions based on CoMP technology only consider single-service scenarios, namely eMBB or URLLC, and do not consider the joint resource allocation of URLLC and eMBB. At the same time, due to the randomness of the wireless channel and the dynamic nature of URLLC users (such as mobility), it is difficult to guarantee the URLLC quality of service (QoS) requirements with probability 1 in practice. The existing URLLC resource puncturing solutions are more likely to meet the URLLC block error rate of each packet with probability 1, that is, the URLLC instantaneous reliability constraint. This solution will lead to a large number of eMBB resources being punctured in extreme cases, and the dilemma that URLLC reliability cannot be met, resulting in waste of eMBB resources. Summary of the invention
[0005] In order to solve the problems in the prior art, the present invention provides a resource allocation method based on coordinated multi-point transmission and deep reinforcement learning, comprising the following steps:
[0006] Step 1: Introduce the BLER outage probability for URLLC users and model the joint resource allocation problem of eMBB and URLLC as a CMDP problem;
[0007] Step 2: Adopt the twin delayed deep deterministic policy gradient algorithm based on constraint correction strategy optimization. Through offline training, the algorithm can use the deep neural network forward propagation to find the optimal resource allocation strategy with low complexity online, thereby reducing the processing delay of the algorithm.
[0008] As a further improvement of the present invention, in step 1, the following steps are included:
[0009] Step S1: Based on the imperfect channel mutual difference, the imperfection of the channel estimation delay, and the imperfection of the channel estimation error, a mathematical modeling of the imperfect channel is obtained as follows:
[0010]
[0011] in
[0012] Step S2: Calculate the signal-to-interference-to-noise ratio of the URLLC user under the imperfect channel, which is expressed as follows:
[0013]
[0014] in represents the effect of imperfect channel, is the precoding vector, κ[t] is the interference signal, and n[t] is Gaussian white noise;
[0015] Step S3: Assuming that the short packet length of the URLLC data packet is Z, when the RRH cluster compiles the URLLC packet into l through channel coding at time t x When there are [t] OFDM symbols, the achievable block error rate is approximated by the short packet formula as:
[0016]
[0017] Where W 0 ,T 0 are the subcarrier spacing and single symbol duration, R[t] = Z / (l x [t]W 0 T 0 ) is the actual transmission rate of the URLLC packet, is the channel divergence;
[0018] Step S4: Define the BLER outage probability and model the CMDP problem.
[0019] As a further improvement of the present invention, in the step S4, it also includes:
[0020] Step 1: Given the BLER violation indicator function c[t] at time t, defined as At the same time, to maximize the number of eMBB symbols, the reward function is defined as:
[0021] ξ[t]=L max -l p [t]-l x [t] (1.7);
[0022] Step 2: Under the deterministic strategy μ, the eMBB long-term reward is expressed as follows:
[0023]
[0024] The URLLC long-term reliability constraint is expressed as follows:
[0025]
[0026] Where Γ∈(0,1] is the reduction factor.
[0027] As a further improvement of the present invention, in the first step, if the BLER is higher than a given threshold ε th, indicating function c[t]=1, otherwise indicating function c[t]=0.
[0028] As a further improvement of the present invention, in step 2, it also includes:
[0029] Step 1: Initialize the parameters of two groups of neural networks, including the action network group and the evaluation network group. The action network group is responsible for learning and feeding back the URLLC pilot length and the number of data symbols, which is recorded as the action vector a[t] = [l p [t],l x [t]], the evaluation network group is responsible for evaluating the action vector;
[0030] Step 2: Based on the current actual uplink channel gain Get a set of observations According to the observed value, we get the sample transfer space consisting of the current state, the corresponding action, the corresponding reward and cost, and the next state, recorded as<o[t],a[t],ξ[t],c[t],o[t+1]> , and store it in the experience playback memory;
[0031] Step 3: Repeat step 2 until the experience playback memory is full, and then execute step 4;
[0032] Step 4: Randomly sample N in the buffer s The group transfers sample data, calculates the temporal difference error, and updates the evaluation group network;
[0033] Step 5: Use the evaluation group network to calculate the URLLC long-term reliability constraint. If the constraint is violated, the action group network uses stochastic gradient descent to minimize the long-term reliability constraint, otherwise it uses stochastic gradient ascent to maximize the long-term reward.
[0034] As a further improvement of the present invention, in step S1, the imperfect channel reciprocity:
[0035] In a TDD system, the channel vectors of the uplink and downlink are modeled as follows:
[0036]
[0037] in Denote the channel vectors of the uplink and downlink at time t, respectively, v j [t] describes the uncertainty of the imperfect channel reciprocity, φ∈[0,1] represents the channel reciprocity coefficient;
[0038] Imperfection of channel estimation delay:
[0039] In the CoMP system, the mathematical formula for describing the impact of channel estimation delay is as follows:
[0040]
[0041] in is the real channel at time (t-τ), τ represents the channel estimation delay, is a random vector, b is the channel correlation coefficient;
[0042] Imperfection of channel estimation error:
[0043] The channel estimation error is expressed as the difference between the true channel and the estimated channel, expressed as here represents the estimated channel at time (t-τ), and the channel estimation error e j,τ is a random vector that obeys a circularly symmetric complex Gaussian distribution, that is variance The mathematical form of is as follows:
[0044]
[0045] Here ρ j [t] = α j [t]β j , α j [t] and β j Represent the path loss and RRH maximum transmission power of time slot t, represents the noise variance, l p [t] represents the pilot length allocated to the URLLC user in time slot t.
[0046] The present invention also discloses a resource allocation system based on coordinated multi-point transmission and deep reinforcement learning, comprising: a memory, a processor, and a computer program stored on the memory, wherein the computer program is configured to implement the steps of the resource allocation method described in the present invention when called by the processor.
[0047] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the resource allocation method described in the present invention when called by a processor.
[0048] The beneficial effects of the present invention are: 1. The present invention can not only meet the URLLC delay and reliability requirements under multi-cellular networks, but also can still meet the URLLC QoS indicators in high mobility scenarios. At the same time, by dynamically adjusting the URLLC pilot length and the number of data transmission symbols, the present invention can effectively reduce the attenuation of eMBB user performance due to URLLC preemption; 2. The channel estimation delay of the present invention is within 2ms. Compared with the existing solutions, the present invention can still effectively meet the URLLC user's 1ms delay and 99.9999% reliability requirements, while obtaining good eMBB performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a CoMP network diagram with two RRH clusters according to the present invention;
[0050] Figure 2 It is a packet transmission mechanism of the coordinated multi-point transmission system of the present invention;
[0051] Figure 3 It is the twin delayed deep deterministic policy gradient algorithm based on CRPO of the present invention; DETAILED DESCRIPTION
[0052] Glossary:
[0053] URLLC: Ultra-reliable and low-latency communications;
[0054] eMBB: Enhanced mobile broadband;
[0055] CoMP: Coordinated multipoint, coordinated multipoint transmission;
[0056] DRL: Deep reinforcement learning, deep reinforcement learning;
[0057] RPS-U: Resource puncturing scheme for URLLC, URLLC resource puncturing scheme;
[0058] CRPO: Constraint-rectified policy optimization, constraint correction strategy optimization;
[0059] BBU: Baseband unit, baseband unit;
[0060] RRH: Remote radio head, remote radio head;
[0061] CMDP: Constrained Markov decision process, constrained Markov decision process;
[0062] The present invention is aimed at the coexistence scenario of ultra-reliable and low-latency communications (URLLC) and enhanced mobile broadband (eMBB) services with multiple cells and high mobility, and proposes a resource allocation scheme based on coordinated multipoint (CoMP) and deep reinforcement learning (DRL), also known as resource puncturing scheme for URLLC (RPS-U). By considering the imperfect mutual difference of channels, estimation error, estimation delay and interval interference, the present invention can dynamically adjust the number of URLLC pilot symbols and the number of data symbols, which not only realizes the ultra-low latency and ultra-high reliability requirements of URLLC, but also effectively reduces the resources occupied by eMBB and improves the block error rate interruption performance of URLLC. The present invention is mainly used for resource reuse of ultra-reliable and low-latency communications and enhanced mobile broadband services of 5G / B5G.
[0063] The present invention discloses a resource allocation method based on coordinated multi-point transmission and deep reinforcement learning, comprising the following steps:
[0064] Step 1: Considering the impact of deep fading of wireless channels and user dynamics, the BLER interruption probability for URLLC users is introduced, and the joint resource allocation problem of eMBB and URLLC is modeled as a CMDP problem; compared with the traditional solution that focuses on short-term benefits and instantaneous constraints, the solution of the present invention pays more attention to long-term eMBB throughput and aims to ensure the long-term reliability constraints of URLLC, allowing for extreme cases where URLLC user reliability cannot be met, thereby improving the long-term throughput performance of eMBB.
[0065] Step 2: Considering the real-time nature of the resource scheduling strategy, a twin delayed deep deterministic policy gradient algorithm based on constraint-corrected policy optimization (CRPO-based TD3 algorithm) is used. Through offline training, the algorithm can use the deep neural network (DNN) forward propagation online to find the optimal resource allocation strategy with low complexity, thereby reducing the processing delay of the algorithm.
[0066] Specific introduction:
[0067] The present invention considers a two-layer multi-cellular CoMP network, in which the first layer is composed of a baseband unit (BBU) and the second layer is composed of a remote radio head (RRH). These RRHs provide services for URLLC and eMBB users. The present invention considers that every two RRHs form an RRH cluster to jointly serve URLLC users, and each RRH is equipped with N x All RRHs and BBUs are connected via optical fiber. Figure 1 shown.
[0068] The time-frequency resources of each time slot are divided into L max Orthogonal frequency-division multiplexing (OFDM) symbols. Usually, eMBB user resources are scheduled once at the beginning of each time slot, and URLLC users perform resource puncturing for eMBB users at the beginning of each mini-slot. In the same RRH cluster of the CoMP system, two coordinated RRHs will use the same time-frequency resources to send the same data packet to the users they serve together. Figure 2 .
[0069] In order to achieve URLLC delay and reliability requirements under complex scenarios (interval interference, user dynamics, and real-time strategy) and protect eMBB user resources, the present invention takes into account imperfect channels and interval interference, and models the joint resource allocation problem of eMBB and URLLC into a CMDP problem, which is intended to improve the ability of the present invention to resist non-ideal environments. At the same time, the present invention combines the deep reinforcement learning algorithm with the CRPO algorithm, so that the present invention can not only solve the CMDP problem, but also has the ability to quickly feedback the environment.
[0070] Due to the widespread application of massive multiple-input and multiple-output (mMIMO) technology in future 5G / B5G, the time division duplex (TDD) communication system has also become the first choice in 5G / B5G, because the TDD system can significantly reduce the channel training overhead of the multi-antenna system. Based on this situation, the present invention considers imperfect channels based on the TDD system, and finally characterizes the URLLC signal-to-interference-plus-noise ratio (SINR) and BLER outage probability.
[0071] In high mobility scenarios, the perfect channel diversity assumed by TDD no longer holds. The channel vectors for uplink and downlink can be modeled as follows:
[0072]
[0073] in Denote the channel vectors of the uplink and downlink at time t, respectively, v j [t] describes the uncertainty of imperfect channel diversity, and φ∈[0,1] represents the channel diversity coefficient.
[0074] In an actual CoMP system, in order to achieve cooperative communication between RRHs, all RRHs in each RRH cluster will share channel information through the fronthaul link after channel estimation is completed. Therefore, there is a channel estimation delay, which is denoted as τ. The present invention uses mathematical language to express the influence of channel estimation delay as follows:
[0075]
[0076] here is the real channel at time (t-τ), is a random vector, b is the channel correlation coefficient. The channel correlation coefficient is mainly related to the maximum Doppler frequency shift and the channel estimation delay, and can be modeled as a 0th-order first-kind Bessel function.
[0077] In addition to the channel imperfections caused by imperfect channel reciprocity and channel estimation delay, channel estimation error can also cause channel imperfections. The present invention expresses the channel estimation error as the difference between the real channel and the estimated channel, expressed as here represents the estimated channel at time (t-τ). At the same time, the channel estimation error e j,τ is a random vector that obeys a circularly symmetric complex Gaussian distribution, that is variance The mathematical form of is as follows:
[0078]
[0079] Here ρ j [t] = α j [t]β j , α j [t] and β j Represent the path loss and RRH maximum transmission power of time slot t, represents the noise variance, l p [t] represents the pilot length allocated to the URLLC user in time slot t.
[0080] Taking the above three channel imperfections into consideration, the present invention can obtain the mathematical modeling of the imperfect channel as shown below:
[0081]
[0082] in
[0083] Precoding takes into account the maximum ratio transmission (MRT), so the signal-to-interference-to-noise ratio of URLLC users under imperfect channels is expressed as follows:
[0084]
[0085] in represents the effect of imperfect channel, is the precoding vector, κ[t] is the interference signal, and n[t] is Gaussian white noise.
[0086] The present invention takes into account that URLLC data packets are usually short packets. Assuming that the packet length is Z, when the RRH cluster compiles the URLLC packet into l through channel coding in the t time slot x [t] OFDM symbols, the achievable block error rate can be approximated by the short packet formula:
[0087]
[0088] Where W 0 ,T 0 are the subcarrier spacing and single symbol duration, R[t] = Z / (l x [t]W 0 T 0 ) is the actual transmission rate of the URLLC packet, is the channel divergence.
[0089] Based on the above statement, we will define the BLER outage probability. First, we give the BLER violation indicator function c[t] at time t, which is defined as If the BLER is higher than a given threshold ε th , c[t]=1, otherwise c[t]=0. Meanwhile, the present invention aims to maximize the number of eMBB symbols, and the reward function is defined as
[0090] ξ[t]=L max -l p [t]-l x [t] (1.7)
[0091] Since in the CMDP problem, the present invention generally considers long-term rewards and long-term constraints. Therefore, under the deterministic strategy μ, the eMBB long-term reward is given as follows:
[0092]
[0093] The URLLC long-term reliability constraint is expressed as follows:
[0094]
[0095] Where Γ∈(0,1] is the reduction factor.
[0096] Finally, if Figure 3 As shown, the present invention proposes a twin delayed deep deterministic policy gradient algorithm based on constraint correction policy optimization (CRPO-based TD3 algorithm). Next, we give the specific steps of the algorithm:
[0097] Step 1: Initialize the parameters of two groups of neural networks, including the action network group and the evaluation network group. The action network group is responsible for learning and feeding back the URLLC pilot length and the number of data symbols, which is recorded as the action vector a[t] = [l p [t],l x [t]], the evaluation network group is responsible for evaluating the action vector;
[0098] Step 2: Based on the current actual uplink channel gain Get a set of observations According to the observed value, we get the sample transfer space consisting of the current state, the corresponding action (i.e., resource allocation strategy), the corresponding reward and cost, and the next state, which is recorded as<o[t],a[t],ξ[t],c[t],o[t+1]> , and stored in the experience playback memory.
[0099] Step 3: Repeat step 2 until the experience playback memory is filled, and then proceed to step 4;
[0100] Step 4: Randomly sample N in the buffer s Transfer sample data, calculate temporal-difference error (TD-error), and update the evaluation group network;
[0101] Step 5: Use the evaluation group network to calculate the URLLC long-term reliability constraint. If the constraint is violated, the action group network uses stochastic gradient descent to minimize the long-term reliability constraint, otherwise it uses stochastic gradient ascent to maximize the long-term reward.
[0102] The present invention also discloses a resource allocation system based on coordinated multi-point transmission and deep reinforcement learning, comprising: a memory, a processor, and a computer program stored on the memory, wherein the computer program is configured to implement the steps of the resource allocation method described in the present invention when called by the processor.
[0103] The present invention discloses a resource allocation method based on coordinated multi-point transmission and deep reinforcement learning. By introducing the interruption probability of URLLC packet error rate (BLER), the joint resource allocation problem of eMBB and URLLC is modeled as a constrained Markov decision problem. The present invention dynamically adjusts the number of pilot symbols and the number of data transmission symbols allocated to URLLC users through the base station in each time slot, maximizes the number of eMBB user symbols under the premise of satisfying URLLC reliability and delay constraints, and realizes the coexistence of eMBB and URLLC users.
[0104] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the resource allocation method of the present invention when called by a processor.
[0105] The beneficial effects of the present invention are: 1. The present invention can not only meet the URLLC delay and reliability requirements under multi-cellular networks, but also can meet the URLLC QoS indicators in high mobility scenarios. At the same time, by dynamically adjusting the URLLC pilot length and the number of data transmission symbols, the present invention can effectively reduce the attenuation of eMBB user performance due to puncturing; 2. The channel estimation delay is within 2ms. Compared with the existing solutions, the present invention can still effectively meet the URLLC user's 1ms delay and 99.9999% reliability requirements, while obtaining good eMBB performance.
[0106] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A resource allocation method based on coordinated multi-point transmission and deep reinforcement learning, characterized in that: The following steps are involved: Step 1: Introduce the BLER outage probability for URLLC users and model the joint resource allocation problem of eMBB and URLLC as a CMDP problem; Step 2: Adopt the twin delayed deep deterministic policy gradient algorithm based on constraint correction strategy optimization. Through offline training, the algorithm can use the deep neural network forward propagation to find the optimal resource allocation strategy with low complexity online, thereby reducing the processing delay of the algorithm.
2. The resource allocation method according to claim 1, characterized in that: In step 1, the following steps are included: Step S1: Based on the imperfect channel mutual difference, the imperfection of the channel estimation delay, and the imperfection of the channel estimation error, a mathematical modeling of the imperfect channel is obtained as follows: in Step S2: Calculate the signal-to-interference-to-noise ratio of the URLLC user under the imperfect channel, which is expressed as follows: in represents the effect of imperfect channel, is the precoding vector, κ[t] is the interference signal, and n[t] is Gaussian white noise; Step S3: Assuming that the short packet length of the URLLC data packet is Z, when the RRH cluster compiles the URLLC packet into l through channel coding at time t x When there are [t] OFDM symbols, the achievable block error rate is approximated by the short packet formula as: Where W0 and T0 are the subcarrier spacing and single symbol duration respectively, R[t] = Z / (l x [t]W0T0) is the actual transmission rate of URLLC packets, is the channel divergence; Step S4: Define the BLER outage probability and model the CMDP problem.
3. The resource allocation method according to claim 2, characterized in that: In the step S4, it also includes: Step 1: Given the BLER violation indicator function c[t] at time t, defined as At the same time, to maximize the number of eMBB symbols, the reward function is defined as: ξ[t]=L max -l p [t]-l x [t]. (1.7); Step 2: Under the deterministic strategy μ, the eMBB long-term reward is expressed as follows: The URLLC long-term reliability constraint is expressed as follows: Where Γ∈(0,1] is the reduction factor.
4. The resource allocation method according to claim 3, characterized in that: In step 1, if the BLER is higher than a given threshold ε th , indicating function c[t]=1, otherwise indicating function c[t]=0.
5. The resource allocation method according to claim 1, characterized in that: In step 2, it also includes: Step 1: Initialize the parameters of two groups of neural networks, including the action network group and the evaluation network group. The action network group is responsible for learning and feeding back the URLLC pilot length and the number of data symbols, which is recorded as the action vector a[t] = [l p [t],l x [t]], the evaluation network group is responsible for evaluating the action vector; Step 2: Based on the current actual uplink channel gain Get a set of observations According to the observed value, we get the sample transfer space consisting of the current state, the corresponding action (i.e., resource allocation strategy), the corresponding reward and cost, and the next state, which is recorded as<o[t],a[t],ξ[t],c[t],o[t+1]> , and store it in the experience playback memory; Step 3: Repeat step 2 until the experience playback memory is full, and then execute step 4; Step 4: Randomly sample N in the buffer s The group transfers sample data, calculates the temporal difference error, and updates the evaluation group network; Step 5: Use the evaluation group network to calculate the URLLC long-term reliability constraint. If the constraint is violated, the action group network uses stochastic gradient descent to minimize the long-term reliability constraint, otherwise it uses stochastic gradient ascent to maximize the long-term reward.
6. The resource allocation method according to claim 2, characterized in that: In step S1, the imperfect channel reciprocity is: In a TDD system, the channel vectors of the uplink and downlink are modeled as follows: in Denote the channel vectors of the uplink and downlink at time t, respectively, v j [t] describes the uncertainty of the imperfect channel reciprocity, φ∈[0,1] represents the channel reciprocity coefficient; Imperfection of channel estimation delay: In the CoMP system, the mathematical formula for describing the impact of channel estimation delay is as follows: in is the real channel at time (t-τ), τ represents the channel estimation delay, is a random vector, b is the channel correlation coefficient; Imperfection of channel estimation error: The channel estimation error is expressed as the difference between the true channel and the estimated channel, expressed as here represents the estimated channel at time (t-τ), and the channel estimation error e j,τ is a random vector that obeys a circularly symmetric complex Gaussian distribution, that is variance The mathematical form of is as follows: Here ρ j [t] = α j [t]β j , α j [t] and β j Represent the path loss and RRH maximum transmission power of time slot t, represents the noise variance, l p [t] represents the pilot length allocated to the URLLC user in time slot t.
7. A resource allocation system based on coordinated multi-point transmission and deep reinforcement learning, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein the computer program is configured to implement the steps of the resource allocation method according to any one of claims 1 to 6 when called by the processor.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the resource allocation method according to any one of claims 1 to 6 when called by a processor.
Citation Information
Patent Citations
Slice resource allocation method and system for cognitive wireless network
CN111726811A
Power allocation method for packet prediction control system
CN111935667A
EMBB and URLLC service resource allocation method and system based on deep reinforcement learning
CN119110417A
Method and apparatus for machine learning based wide beam optimization in cellular network
WO2019231289A1
Systems, methods, and apparatus on wireless network architecture and air interface
WO2022205023A1