Network assisted full-duplex mode optimization method under low-altitude three-dimensional coverage scenario

By designing P-RZF downlink precoding and P-MMSE receivers in low-altitude three-dimensional coverage scenarios, and combining DQN reinforcement learning algorithm to optimize the duplex mode of RRU, the uplink and downlink interference problem in the network-assisted full-duplex CF-RAN system is solved, maximizing spectrum efficiency and efficient resource utilization.

CN116437370BActive Publication Date: 2026-05-01SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2023-04-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In low-altitude three-dimensional coverage scenarios, network-assisted full-duplex CF-RAN systems suffer from severe uplink and downlink cross-interference, which limits system performance and spectrum efficiency.

Method used

A user-centric association strategy is adopted to design P-RZF downlink precoding vectors and P-MMSE receivers. Combined with DQN reinforcement learning algorithm, the duplex mode selection of RRU is optimized. The neural network model is trained by an agent to maximize spectral efficiency.

Benefits of technology

While meeting power and quality of service requirements, the system's spectrum efficiency was improved, uplink and downlink interference was reduced, and flexible duplex mode selection and efficient resource utilization were achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116437370B_ABST
    Figure CN116437370B_ABST
Patent Text Reader

Abstract

The application discloses a network auxiliary full-duplex mode optimization method in a low-altitude three-dimensional coverage scene. The method is designed for the actual problem that the uplink and downlink data requirements of users are different in low-altitude three-dimensional coverage. The method designs an uplink P-MMSE receiver and downlink P-RZF precoding for scalable joint transmission, determines the uplink spectral efficiency based on Shannon channel theory and the downlink spectral efficiency based on a limited block length mechanism, and provides an RRU uplink and downlink selection method based on a DQN reinforcement learning algorithm. The duplex mode optimization method provided by the application realizes the maximization of system spectral efficiency under the constraints of the error probability and transmission power of downlink short packet transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Optimization method of network-assisted full-duplex mode in low-altitude three-dimensional coverage scenarios Technical Field

[0001] This invention relates to a duplex mode optimization method based on a network-assisted full-duplex CF-RAN system suitable for low-altitude three-dimensional coverage scenarios, belonging to the field of mobile communication technology. Background Technology

[0002] With the increasingly widespread application of mobile communication across various industries, higher demands are being placed on the spectrum efficiency and coverage of communication systems. In low-altitude three-dimensional coverage scenarios, drone users and ground users coexist. Compared to ground users with large downlink data transmission volumes, drone users have different and more rapidly changing uplink and downlink needs. Specifically, drone communication can be divided into two forms: payload communication and CNPC (Control and Non-Payload Communications). In uplink communication, drones transmit payload signals, mainly consisting of large amounts of data such as video and images, resulting in high data throughput and high spectrum efficiency requirements. In downlink communication, base stations send CNPC signals, primarily control signals, to drones to control their flight altitude, speed, and other behaviors. Therefore, downlink transmission needs to ensure low latency and high reliability. In low-altitude three-dimensional coverage scenarios, full-duplex communication is considered for simultaneous data transmission and reception on the same frequency band, doubling the channel capacity compared to traditional half-duplex communication. However, while full-duplex mode increases throughput, it also introduces severe uplink and downlink cross-interference, limiting system performance.

[0003] Cell-free Radio Access Network (CF-RAN), as a novel architecture for a non-cellular massive MIMO (Multiple-Input Multiple-Output) system, is physically divided into Remote Radio Units (RRUs), Edge Distributed Units (EDUs), User-Centric Distributed Units (UCDUs), and Central Control Units (CCUs). The RRUs handle the reception and transmission of radio frequency signals in each frequency band, the EDUs perform distributed precoding and reception, and the UCDUs handle user-centric data distribution and aggregation. Within the same time-frequency resource block, each RRU can flexibly choose between uplink reception and downlink transmission; the specific duplex mode is determined by the UCDU. When downlink RRU transmission interferes with uplink RRU reception, the UCDU sends interference information to the associated EDU, which then uses channel state information between RRUs to cancel the interference. By scheduling users within the UCDU, interference between uplink user transmissions and downlink user reception can be reduced. This coordinated duplexing method in CF-RAN is also known as Network-Assisted Full Duplex (NAFD). Therefore, a network-assisted full-duplex CF-RAN architecture is considered to meet user-centric service requirements in low-altitude three-dimensional coverage scenarios. By dynamically adjusting the duplex mode selection of RRUs, cross-link interference can be reduced, and the asymmetry between uplink and downlink transmissions can be resolved to achieve scalable and flexible duplex wireless communication. Summary of the Invention

[0004] Technical Problem: In view of this, the technical problem to be solved by the present invention is to provide a duplex mode selection optimization method for network-assisted full-duplex CF-RAN systems in low-altitude three-dimensional coverage scenarios, which can maximize spectrum efficiency.

[0005] Technical Solution: To achieve the above objectives, the present invention provides a network-assisted full-duplex mode optimization method for low-altitude three-dimensional coverage scenarios, employing the following solution:

[0006] Step S1: In a network-assisted full-duplex CF-RAN system, based on a user-centric association strategy and considering non-ideal channel state information, design a P-RZF downlink precoding vector and determine the spectral efficiency of transmission based on a finite block length mechanism under the downlink CNPC link.

[0007] Step S2: Based on the P-RZF downlink precoding used in step S1, and considering the presence of residual downlink interference, design an uplink P-MMSE receiver and determine the spectral efficiency based on Shannon channel theory under uplink payload communication.

[0008] Step S3: Based on the results in step S2, design a joint optimization problem for uplink and downlink spectral efficiency, and determine the system state function, action function, and reward function based on the DQN reinforcement learning algorithm principle;

[0009] Step S4: Based on the joint optimization problem designed in step S3, solve it using the intelligent DQN algorithm, and save the final state set and reward of the algorithm as the optimal RRU duplex mode and the maximized system spectral efficiency.

[0010] Step 1 specifically includes:

[0011] Step S101: Consider a low-altitude coverage CF-RAN system equipped with one CCU, several UCDUs and switches, M EDUs with buffering and computing capabilities, and N full-duplex RRUs configured with L antennas. Assume N... ul N are uplink receiving RRUs. dl There are K downlink transmission RRUs, K drones and ground users, and the user locations are randomly distributed, where K U K is the uplink sending user. D Each RRU in the system can select either uplink reception or downlink transmission according to user needs, using EDU and UCDU to select the appropriate mode. During the downlink transmission phase, the signal received by the k-th user is:

[0012]

[0013] In formula (1), D k This represents the association vector between the downlink RRU and the k-th downlink active user. This represents the channel vector between the downlink RRU and the downlink active user k. Let w be the channel vector between the nth downlink RRU and the kth downlink active user. k It is the channel precoding between the downlink transmitting RRU and the k-th downlink user, s k w is the signal sent by the downlink RRU to the kth active downlink user. j It is the channel precoding between the downlink transmitting RRU and downlink user j, s j p is the signal sent by the downlink RRU to the downlink active user j. ul,i U represents the transmission power of the i-th active uplink user. k,iFor the cross-link interference between the i-th uplink transmitting user and the k-th downlink receiving user, x i The signal sent by the i-th uplink active user satisfies z dl,k It is additive white Gaussian noise;

[0014] The power constraints that each downlink transmission RRU must satisfy are:

[0015]

[0016] In formula (2), W i Let P be the precoding matrix for the i-th downlink transmission RRU, P be the power constraint for each antenna, and E be the precoding matrix for the i-th downlink transmission RRU. i Let i be the identity matrix whose i-th column is not zero;

[0017] Step S102: Scalable P-RZF linear precoding is used on the Edge Distribution Unit (EDU) to eliminate inter-user interference. The P-RZF precoding vector transmitted from the RRU to the k-th downlink receiving user is:

[0018]

[0019]

[0020] In formulas (3) and (4), δ is a normalization coefficient obtained by satisfying the power constraint formula (2). S k This represents the set of downlink users k that are associated with some of the same RRUs. Then it represents set S k The channel estimation matrix from all users to the downlink RRU, where α > 0 represents the regularization coefficient. LN dl An identity matrix of order 1. For the i-th downlink transmission RRU, there is the P-RZF precoding matrix;

[0021] Step S103: In the downlink CNPC link of the UAV, short message control signaling is used to meet the transmission requirements of low latency and high reliability. Based on the P-RZF downlink precoding in formula (3), the spectral efficiency of the k-th downlink user under the finite block length communication (FBLC) mechanism is determined as follows:

[0022]

[0023]

[0024] In formulas (5) and (6), τ is the length of the pilot estimation sequence, and T is the coherent time slot. It is the P-RZF precoding vector of downlink user k. It is the P-RZF precoding vector of downlink user j. It is the signal-to-interference-plus-noise ratio (SIR) of downlink user k, V(γ)=1-(1+γ) -2 It is channel scattering, ε is the block error probability, e is the natural logarithm, and Q is the channel scattering. -1 (·) is the inverse function of the complementary cumulative distribution function Q-function of the standard Gaussian random variable, μ = BT is the number of bits used per channel, and B represents the system bandwidth. This indicates the downlink noise power.

[0025] Step 2 specifically includes:

[0026] Step S201: In the low-altitude three-dimensional coverage CF-RAN system, the uplink and downlink baseband signals are processed collaboratively by EDUs to mitigate downlink interference between RRUs to the uplink. When the downlink transmission uses the above P-RZF precoding, considering incomplete channel state information, the channel interference from the downlink RRU to the uplink RRU cannot be completely eliminated. During the uplink transmission phase, the signal received by the m-th EDU is:

[0027]

[0028] In formula (7), D represents the association vector between the m-th EDU and RRU. k D represents the association vector between uplink active user k and uplink RRU. i g represents the association vector between uplink active user i and uplink RRU. ul,k Let g be the channel vector between the k-th uplink active user and the uplink receiving RRU. ul,i Let p be the channel vector between the i-th uplink active user and the uplink receive RRU. ul,k p represents the transmission power of the k-th active uplink user. ul,i Let x represent the transmission power of the i-th active uplink user. k The signal sent by the kth active uplink user, x i The signal sent by the i-th active uplink user. This represents the estimation error of the channel between the downlink RRU and the uplink RRU. s is the P-RZF precoding vector of the j-th downlink user. j For the transmission signal of the j-th downlink active user, For the residual interference term, z ul It is additive white Gaussian noise;

[0029] Step S202: At the receiving end, a scalable P-MMSE receiver is used. The receiving vector of the k-th uplink user is represented as:

[0030]

[0031]

[0032] In formulas (8) and (9), the definition is... Σ k The covariance matrix representing residual interference and noise. S is the channel estimation matrix between uplink active user k and uplink RRU. k This represents the set of all uplink users associated with some of the same RRUs as uplink user k. Let set S k The channel estimation matrix S′ between uplink active user i and uplink RRU. k This represents the set of all downlink users associated with some of the same RRUs as uplink user k. LN ul An identity matrix of order 1. Indicates uplink noise power;

[0033] Step S203: The uplink transmission of UAV users is mainly payload communication, which has high requirements for channel capacity and data transmission rate. Therefore, the traditional Shannon channel theory is used to analyze the spectral efficiency. Using the P-MMSE uplink receiver in formula (8), the spectral efficiency of the k-th uplink user is determined as follows:

[0034]

[0035]

[0036] In formulas (10) and (11), τ is the length of the pilot estimation sequence, and T is the coherent time slot. It is the signal-to-interference-plus-noise ratio (SIR) of the k-th uplink user. This represents the correlation matrix between EDU and RRU.

[0037] Step 3 specifically includes:

[0038] Step S301: In a network-assisted full-duplex system, a single base station only needs to implement half-duplex functionality. Based on the real-time needs of UAVs and ground users, flexible scheduling of RRU uplink and downlink is performed to save resource overhead, improve system performance, and set optimization targets for uplink spectrum efficiency.

[0039]

[0040]

[0041] C2:α ul,n +α dl,n =1

[0042]

[0043] In formula (12), α ul =[α ul,1 ,…,α ul,N ], α dl =[α dl,1 ,…,α dl,N ], N is the total number of RRUs that can provide uplink and downlink options, α ul,n and α dl,n This represents the uplink / downlink selection of the nth RRU, when α ul,n When α = 1, dl,n If α = 0, then the nth RRU is responsible for uplink transmission; otherwise, when α = 0... ul,n When α = 0, dl,n =1, then the nth RRU is responsible for downlink transmission, and each RRU can only select one mode, γ ul,k γ represents the signal-to-interference-plus-noise ratio (SIR) of uplink user k. ul,k,min p represents the minimum signal-to-interference-plus-noise ratio (SIR) required for uplink users. ul,k P represents the transmission power of the k-th active uplink user. ul This represents the maximum uplink transmission power for the user.

[0044] Step S302, set the optimization objective for downlink spectral efficiency based on the finite block length transmission mechanism:

[0045]

[0046]

[0047] C5: ε≤ε max

[0048]

[0049] C7:p dl,k ≥0

[0050] C2-C3

[0051] In formula (13), γ dl,k γ represents the signal-to-interference-plus-noise ratio (SIR) of downlink user k. dl,k,min ε represents the minimum signal-to-interference-plus-noise ratio (SINNR) required for downlink users, and ε represents the block error probability. max w represents the maximum tolerable error probability. k This refers to the channel precoding between the downlink transmitting RRU and the k-th downlink user, where P is the threshold for downlink user transmission power. dl,k This represents the transmission power of the kth active downlink user;

[0052] Step S303: In a network-assisted full-duplex CF-RAN system, when uplink demand is high, the number of uplink serving RRUs can be increased to improve uplink spectrum efficiency; when downlink demand is high, the number of downlink serving RRUs can be increased to improve downlink spectrum efficiency. Since at a specific moment, the nth RRU can only choose either uplink or downlink mode, the optimization objectives of formulas (12) and (13) are contradictory. Therefore, a multi-objective optimization problem for RRU uplink / downlink mode selection is set up:

[0053]

[0054] C1-C7

[0055] Consider using multi-objective optimization to design an up-and-down scheduling scheme for RRUs to achieve a trade-off between the two problems and maximize the system gain;

[0056] Step S304, the intelligent mode selection algorithm based on DQN is as follows:

[0057] The DQN algorithm combines deep learning and reinforcement learning, offering high reliability for solving action selection problems based on discrete variables. In DQN, the agent uses a state-action value function Q and a greedy scheme based on a fixed probability ε. t The agent selects the action a(t) with the highest Q value, obtains the reward function r(t), and enters the next state s′. By storing the actions taken and the rewards obtained in memory, the agent can continuously train its own neural network model to obtain the optimal solution.

[0058] The loss function that minimizes the squared error of the neural network parameters of the agent is:

[0059]

[0060] In formula (15), r(t) is the reward function of action a(t) at step t, and l is the discount factor. To maximize the Q-value function of the next state s′, A is the set of actions, and Q(s(t),a(t)) is the Q-value function for choosing action a(t) in state s(t);

[0061] In the DQN-based intelligent uplink / downlink mode selection algorithm, the CCU is regarded as an intelligent agent, and the state-space function s(t) is defined as follows: A(t) represents the uplink and downlink selection vectors of the RRU in step t; the action space function a(t) is defined as This indicates the change in uplink / downlink selection of the RRU during step t;

[0062] The reward function r(t) is defined as:

[0063]

[0064] In formula (16), This represents the total spectrum efficiency for uplink users. γ represents the total spectral efficiency of downlink users under a finite block length transmission mechanism. a It is a regularization parameter that ensures network convergence. It is a constant parameter related to the sum of uplink and downlink user spectral efficiency.

[0065] Step 4 specifically includes:

[0066] Step S401, Initial setting t=0, initialize state s(0) and neural network parameters of the agent;

[0067] Step S402, based on probability ε t - A greedy strategy selects action a(t);

[0068] Step S403: Based on the current state s(t) and the selected action a(t), calculate the Q-value function Q(s(t), a(t));

[0069] Step S404: The current state jumps to the next state s′ based on the selected action;

[0070] Step S405: Calculate the reward function r(t) according to formula (16);

[0071] Step S406: Store the joint transition vector d(t)=[s(t),a(t),r(t),s′] in the memory pool, and train the neural network according to formula (15), t=t+1;

[0072] Step S407, if t = t max The output is the optimal state, i.e., the optimal RRU uplink / downlink selection scheme s(t). max ), and the maximum reward, which is the maximum benefit the system can obtain;

[0073] Otherwise, return to step S402.

[0074] Beneficial Effects: This invention considers the duplex mode optimization problem of a network-assisted full-duplex CF-RAN system in a low-altitude three-dimensional coverage scenario. By designing an uplink P-MMSE receiver and a downlink P-RZF precoding vector, the spectral efficiency based on Shannon channel theory for uplink payload communication and on a finite block length mechanism for downlink CNPC link transmission are determined. Furthermore, the uplink and downlink selection are jointly optimized based on the DQN reinforcement learning algorithm, maximizing the system's spectral efficiency while meeting power and quality of service requirements. Attached Figure Description

[0075] Figure 1 is a simulation scenario of the network-assisted full-duplex CF-RAN system under the low-altitude three-dimensional coverage scenario in Example 1.

[0076] Figure 2 is a simulation diagram showing the relationship between the average spectral efficiency and the number of RRU antennas based on the DQN optimization algorithm in Example 1. Detailed Implementation

[0077] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0078] Example 1

[0079] Referring to Figure 1, in this embodiment, it is assumed that within a three-dimensional cylindrical region with a radius of 1km and a height of 100m, the number of EDUs with a height of 3m is M=2, and the number of multi-antenna duplex RRUs with a height of 10m is N=10. U =10 uplink users and K D Ten downlink users are randomly distributed within the region. These users can be ground users or drone users. For the sake of simulation universality, we consider drone users based on the Ricean channel model. The path loss for drone users is defined as... Where d n,kLet d be the distance between n RRUs and the k-th user, a = 3.7 be the path loss exponent, and b be the reference distance d. n,k =Median of average path gain over a distance of 1km. Uplink user transmission power p ul,k =3w, downlink user transmission power p dl,k =5W, noise power is -174dBm, coherence time slot T=196.

[0080] For the network-assisted full-duplex CF-RAN system in the aforementioned low-altitude three-dimensional coverage scenario, the method in this embodiment specifically includes the following steps:

[0081] Step S1: Considering non-ideal channel state information, design downlink P-RZF precoding vectors and determine the spectral efficiency of transmission based on the finite block length mechanism under downlink CNPC links:

[0082] During the downlink transmission phase, the signal received by the k-th user is:

[0083]

[0084] In formula (1), D k This represents the association vector between the downlink RRU and the k-th downlink active user. This represents the channel vector between the downlink RRU and the downlink active user k. Let w be the channel vector between the nth downlink RRU and the kth downlink active user. k It is the channel precoding between the downlink transmitting RRU and the k-th downlink user, s k w is the signal sent by the downlink RRU to the kth active downlink user. j It is the channel precoding between the downlink transmitting RRU and downlink user j, s j p is the signal sent by the downlink RRU to the downlink active user j. ul,i U represents the transmission power of the i-th active uplink user. k,i For the cross-link interference between the i-th uplink transmitting user and the k-th downlink receiving user, x i The signal sent by the i-th uplink active user satisfies z dl,k It is additive white Gaussian noise;

[0085] The power constraints that each downlink transmission RRU must satisfy are:

[0086]

[0087] In formula (2), W iLet be the precoding matrix for the i-th downlink transmission RRU, and P be the power constraint for each antenna. Scalable P-RZF linear precoding is used on the edge distribution unit (EDU) to eliminate inter-user interference. The P-RZF precoding vector for the downlink transmission RRU to the k-th downlink receiving inter-user channel is:

[0088]

[0089]

[0090] In formulas (3) and (4), δ is a normalization coefficient obtained by satisfying the power constraint formula (2). S k This represents the set of downlink users k that are associated with some of the same RRUs. Then it represents set S k The channel estimation matrix from all users to the downlink RRU, where α > 0 represents the regularization coefficient. LN dl An identity matrix of order 1. Let E be the P-RZF precoding matrix of the i-th downlink transmission RRU. i Let i be the identity matrix whose i-th column is not zero;

[0091] In the downlink CNPC link of the UAV, short message control signaling is used to meet the transmission requirements of low latency and high reliability. The P-RZF downlink precoding in formula (3) is used to determine the spectral efficiency of the k-th downlink user under the finite block length mechanism (FBLC):

[0092]

[0093]

[0094] In formulas (5) and (6), τ is the length of the pilot estimation sequence, and T is the coherent time slot. It is the P-RZF precoding vector of downlink user k. It is the P-RZF precoding vector of downlink user j. It is the signal-to-interference-plus-noise ratio (SIR) of downlink user k, V(γ)=1-(1+γ) -2 It is channel scattering, ε is the block error probability, e is the natural logarithm, and Q is the channel scattering. -1 (·) The inverse function of the complementary cumulative distribution function (Q-function) of a standard Gaussian random variable, where μ = BT is the number of bits used per channel, and B represents the system bandwidth. Indicates downlink noise power;

[0095] Step S2: Based on the P-RZF downlink precoding used in Step S1, and considering the presence of residual downlink interference, design an uplink P-MMSE receiver and determine the spectral efficiency based on Shannon channel theory for uplink payload communication.

[0096] During the uplink transmission phase, the signal received by the m-th EDU is:

[0097]

[0098] In formula (7), D represents the association vector between the m-th EDU and RRU. k D represents the association vector between uplink active user k and uplink RRU. i g represents the association vector between uplink active user i and uplink RRU. ul,k Let g be the channel vector between the k-th uplink active user and the uplink receiving RRU. ul,i Let p be the channel vector between the i-th uplink active user and the uplink receive RRU. ul,k p represents the transmission power of the k-th active uplink user. ul,i Let x represent the transmission power of the i-th active uplink user. k The signal sent by the kth active uplink user, x i The signal sent by the i-th active uplink user. This represents the estimation error of the channel between the downlink RRU and the uplink RRU. s is the P-RZF precoding vector of the j-th downlink user. j For the transmission signal of the j-th downlink active user, For the residual interference term, z ul It is additive white Gaussian noise;

[0099] At the receiving end, a scalable P-MMSE receiver is used. The receive vector for the k-th uplink user is represented as:

[0100]

[0101]

[0102] In formulas (8) and (9), the definition is... Σ k The covariance matrix representing residual interference and noise. S is the channel estimation matrix between uplink active user k and uplink RRU. kThis represents the set of all uplink users associated with some of the same RRUs as uplink user k. Let set S k The channel estimation matrix S′ between uplink active user i and uplink RRU. k This represents the set of all downlink users associated with some of the same RRUs as uplink user k. LN ul An identity matrix of order 1. Indicates uplink noise power;

[0103] Using the P-MMSE uplink receiver in formula (8), based on the traditional Shannon channel theory, the spectral efficiency of the k-th uplink user is determined as follows:

[0104]

[0105]

[0106] In formulas (10) and (11), τ is the length of the pilot estimation sequence, and T is the coherent time slot. It is the signal-to-interference-plus-noise ratio (SIR) of the k-th uplink user. This represents the correlation matrix between EDU and RRU;

[0107] Step S3: Based on formulas (5) and (10), design a joint optimization problem for uplink and downlink spectral efficiency, and determine the system state function, action function, and reward function based on the DQN reinforcement learning algorithm principle:

[0108] Step S301: Set the optimization target for uplink spectral efficiency:

[0109]

[0110]

[0111] C2:α ul,n +α dl,n =1

[0112]

[0113] In formula (12), α ul =[α ul,1 ,…,α ul,N ], α dl =[α dl,1 ,…,α dl,N ], N is the total number of RRUs that can provide uplink and downlink options, α ul,nand α dl,n This represents the uplink / downlink selection of the nth RRU, when α ul,n When α = 1, dl,n If α = 0, then the nth RRU is responsible for uplink transmission; otherwise, when α = 0... ul,n When α = 0, dl,n =1, then the nth RRU is responsible for downlink transmission, and each RRU can only select one mode, γ ul,k γ represents the signal-to-interference-plus-noise ratio (SIR) of uplink user k. ul,k,min p represents the minimum signal-to-interference-plus-noise ratio (SIR) required for uplink users. ul,k P represents the transmission power of the k-th active uplink user. ul This represents the maximum uplink transmission power for the user.

[0114] Step S302: Set the optimization objective for downlink spectral efficiency based on the finite block length transmission mechanism:

[0115]

[0116]

[0117] C5: ε≤ε max

[0118]

[0119] C7:p dl,k ≥0

[0120] C2-C3

[0121] In formula (13), γ dl,k γ represents the signal-to-interference-plus-noise ratio (SIR) of downlink user k. dl,k,min ε represents the minimum signal-to-interference-plus-noise ratio (SINNR) required for downlink users, and ε represents the block error probability. max w represents the maximum tolerable error probability. k This refers to the channel precoding between the downlink transmitting RRU and the k-th downlink user, where P is the threshold for downlink user transmission power. dl,k This represents the transmission power of the kth active downlink user;

[0122] Step S303: Set up a multi-objective optimization problem for RRU uplink / downlink mode selection:

[0123]

[0124] C1-C7

[0125] Consider using multi-objective optimization to design an RRU up-and-down scheduling scheme to achieve a trade-off between the two problems and maximize the system gain.

[0126] Step S304: Determine the system state function, action function, and reward function based on the DQN reinforcement learning algorithm principle:

[0127] Specifically, step S304 includes:

[0128] The loss function that minimizes the squared error of the neural network parameters of the agent is:

[0129]

[0130] In formula (15), r(t) is the reward function of action a(t) at step t, and l is the discount factor. To maximize the Q-value function of the next state s′, A is the set of actions, and Q(s(t),a(t)) is the Q-value function for choosing action a(t) in state s(t).

[0131] In the DQN-based intelligent uplink / downlink mode selection algorithm, the CCU is regarded as an intelligent agent, and the state-space function s(t) is defined as follows: A(t) represents the uplink and downlink selection vectors of the RRU in step t. The action space function a(t) is defined as... This indicates the change in the uplink / downlink selection of the RRU during step t.

[0132] The reward function r(t) is defined as:

[0133]

[0134] In formula (16), This represents the user's total uplink spectrum efficiency. γ represents the total downlink spectral efficiency of the user under a finite block length transmission mechanism. a It is a regularization parameter that ensures network convergence. It is a constant parameter related to the sum of uplink and downlink user spectral efficiency.

[0135] Step S4: Solve the multi-objective optimization problem of formula (14) using the DQN reinforcement learning algorithm, and save the final state set and reward of the algorithm as the optimal RRU duplex mode and the maximized system gain:

[0136] In this embodiment, it specifically includes:

[0137] Step S401: Initialize t=0, initialize state s(0) and neural network parameters of the agent;

[0138] Step S402: Based on probability ε t - A greedy strategy selects action a(t);

[0139] Step S403: Based on the current state s(t) and the selected action a(t), calculate the Q-value function Q(s(t), a(t));

[0140] Step S404: The current state jumps to the next state s′ based on the selected action;

[0141] Step S405: Calculate the reward function r(t) according to formula (16);

[0142] Step S406: Store the joint transition vector d(t)=[s(t),a(t),r(t),s′] in the memory pool, and train the neural network according to formula (15), t=t+1;

[0143] Step S407, if t = t max The output is the optimal state, i.e., the optimal RRU uplink / downlink selection scheme s(t). max ), and the maximum reward, which is the maximum benefit the system can obtain;

[0144] Otherwise, return to step S402.

[0145] Table 1 shows the relationship between RRU uplink / downlink selection and antenna number in Example 1 based on the DQN optimization algorithm.

[0146] Table 1

[0147]

[0148] Specifically, Table 1 and Figure 2 show the changes in RRU uplink / downlink selection and average user spectral efficiency with the number of RRU antennas in a network-assisted full-duplex CF-RAN system under low-altitude three-dimensional coverage, when the uplink uses P-MMSE receivers and the downlink uses P-RZF precoding joint transmission. Table 1 shows that when the number of antennas L = 20, the agent selects half of the RRUs for uplink service. This is because uplink data transmission requires a large number of antennas, so when the number of antennas on the base station side is low, the system needs to be equipped with more uplink RRUs to serve uplink users. As the number of antennas increases, the number of RRUs selected for uplink service decreases, and more RRUs are used to support the ultra-reliable low-latency requirements of the downlink CNPC link. Therefore, in Figure 2, the downlink spectral efficiency increases from the number of antennas L = 20 to L = 40. When the number of antennas increases from L=40 to L=60, the average downlink spectral efficiency tends to plateau, while the average uplink spectral efficiency shows an increasing trend. This is because when the system is insufficient to simultaneously improve uplink and downlink performance, it prioritizes meeting the data transmission needs of uplink users while ensuring a certain level of downlink reliability. After the number of antennas reaches L=60, the uplink and downlink spectral efficiencies continue to increase. This is because, after meeting the uplink data transmission needs, the downlink spectral efficiency is further improved. Since the uplink uses a P-MMSE receiver, the spectral efficiency increases logarithmically. When the number of antennas increases to a certain level (L=80 to L=100), the average uplink spectral efficiency tends to plateau. From the above analysis, it can be seen that the uplink and downlink RRU selection optimization algorithm based on DQN proposed in this invention is reasonable and feasible, and effectively improves the spectral efficiency of joint uplink and downlink transmission in a low-altitude coverage network-assisted full-duplex CF-RAN system. It can be used for flexible scheduling of UAVs with different uplink concurrency and downlink CNPC requirements.

[0149] Any aspects of this invention not described in detail are well-known to those skilled in the art.

[0150] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A method for optimizing network-assisted full-duplex mode in low-altitude three-dimensional coverage scenarios, characterized in that, The process includes the following steps: Step S1, in a network-assisted full-duplex (CFRAN) system, based on a user-centric association strategy and considering non-ideal channel state information, design a PRZF downlink precoding vector to determine the spectral efficiency of transmission based on a finite block length mechanism under the downlink CNPC link; Step S2, based on the P-RZF downlink precoding used in Step S1, considering the presence of residual downlink interference, design an uplink PMMSE receiver to determine the spectral efficiency based on Shannon channel theory under uplink payload communication; Step S3, based on the results of Step S2, design a joint optimization problem for uplink and downlink spectral efficiency, and determine the system state function, action function, and reward function based on the DQN reinforcement learning algorithm principle; Step S4, based on the joint optimization problem designed in Step S3, solve it using the intelligent DQN algorithm, and save the final state set and reward of the algorithm as the optimal RRU duplex mode and the maximized system spectral efficiency; wherein: Step S1 specifically includes: Step S101, considering a low-altitude coverage CFRAN system equipped with one CCU, several UCDUs, and switches, An EDU with caching and computing capabilities, Each configuration has A duplex RRU with a root antenna, assuming One is an uplink receive RRU. One is a downlink transmission RRU, There are 10 drones and 10 ground users, with user locations randomly distributed. One is the uplink sending user. The first is a downlink receiving user; each RRU in the system can select uplink receiving or downlink transmitting according to user needs, and selects the appropriate mode through EDU and UCDU. During the downlink transmission phase, the first... The signal received by each user is: (1), In formula (1), Indicates the downlink RRU and the first The correlation vector between active users in the downlink, Indicates downlink RRU and downlink active users Channel vectors between For the first The downlink RRU and the first Channel vectors between downlink active users It is the downlink transmission RRU and the first Channel precoding between downlink users For the downlink RRU to give the first Signals sent by active downlink users It is the downlink transmitting RRU and the downlink user Channel precoding between, For downlink RRU to active downlink users The signal sent, Indicates the first The transmission power of each active uplink user For the first The uplink sent the user to the first Cross-link interference between downlink receiving users For the first The signal sent by an active uplink user satisfies , The noise is additive white Gaussian noise; the power constraint satisfied by each downlink transmission RRU is: (2), in formula (2), For the first The precoding matrix of each downlink transmission RRU, For the power constraint of each antenna, For the first A non-zero identity matrix; Step S102, employ scalable PRZF linear precoding on the edge distribution unit (EDU) to eliminate inter-user interference, and transmit the downlink RRU to the first... The P-RZF precoding vectors for each downlink receiving user are: (3), (4), in formulas (3) and (4), It is a normalized coefficient obtained by satisfying the power constraint formula (2). , Indicates connection with downstream users Associated with a set of downlink users sharing the same RRU. Then it represents a set The channel estimation matrix from all users to the downlink RRU, Represents the regularization coefficient. express An identity matrix of order 1. For the first The P-RZF precoding matrix of the downlink transmission RRU; Step S103, in the UAV downlink CNPC link, short message control signaling is used to meet the transmission requirements of low latency and high reliability. Based on the PRZF downlink precoding in formula (3), the P-RZF precoding matrix of the downlink transmission RRU is determined under the finite block length mechanism FBLC. The spectrum efficiency for each downlink user is: (5), (6) In formulas (5) and (6), To estimate the length of the pilot sequence, It is a coherent time slot. Downlink user P-RZF precoding vectors, Downlink user P-RZF precoding vectors, Downlink user Signal-to-interference-to-noise ratio, It is channel scattering. It is the probability of a block error. It is the natural logarithm. It is the inverse function of the complementary cumulative distribution function Q-function of a standard Gaussian random variable. It is the number of bits used per channel. Indicates system bandwidth. The downlink noise power is represented; step S2 specifically includes: step S201, in the low-altitude three-dimensional coverage CFRAN system, the uplink baseband signal and downlink baseband signal are processed collaboratively through cooperation between EDUs to alleviate the interference of the downlink between RRUs to the uplink; when the downlink transmission uses the above-mentioned P-RZF precoding, considering the incomplete channel state information, the channel interference from the downlink RRU to the uplink RRU cannot be completely eliminated; in the uplink transmission stage, the first The signals received by each EDU are: (7), In formula (7), Indicates the first The association vector between each EDU and RRU Indicates active users on the uplink The association vector between the uplink RRU and the uplink RRU. Indicates active users on the uplink The association vector between the uplink RRU and the uplink RRU. For the first Channel vectors between each uplink active user and the uplink receive RRU. For the first Channel vectors between each uplink active user and the uplink receive RRU. Indicates the first The transmission power of each active uplink user Indicates the first The transmission power of each active uplink user For the first Signals sent by active uplink users For the first Signals sent by active uplink users This represents the estimation error of the channel between the downlink RRU and the uplink RRU. It is the first P-RZF precoding vectors for each downlink user For the first Transmission signals of active downlink users For residual interference terms, The noise is additive white Gaussian noise; in step S202, an expandable P-MMSE receiver is used at the receiving end. The receive vectors of each uplink user are represented as follows: (8), (9), in formulas (8) and (9), define , The covariance matrix representing residual interference and noise. For active users The channel estimation matrix between the uplink RRU and the uplink RRU Indicates connection with upstream users A set of all uplink users associated with some of the same RRUs. For set Mid-to-upstream active users The channel estimation matrix between the uplink RRU and the uplink RRU. Indicates connection with upstream users A set of all downlink users associated with some of the same RRUs. express An identity matrix of order 1. Indicates uplink noise power; Step S203, the uplink transmission of UAV users is mainly payload communication, which has high requirements for channel capacity and data transmission rate. Therefore, the traditional Shannon channel theory is used to analyze the spectral efficiency; the PMMSE uplink receiver in formula (8) is used to determine the first The spectral efficiency for each uplink user is: (10), (11), in formulas (10) and (11), To estimate the length of the pilot sequence, It is a coherent time slot. It is the first The signal-to-interference-plus-noise ratio of each uplink user This represents the correlation matrix between EDU and RRU; step S3 specifically includes: step S301, in a network-assisted full-duplex system, a single base station only needs to implement half-duplex functionality, and performs flexible uplink and downlink scheduling of RRUs according to the real-time needs of UAVs and ground users, saving resource overhead, improving system performance, and setting an optimization target for uplink spectrum efficiency: (12), in formula (12), , , The total number of RRUs that can provide uplink and downlink options. and Indicates the first The uplink and downlink selection of each RRU, when hour, Then the first One RRU is responsible for uplink transmission, and conversely, when... hour, Then the first Each RRU is responsible for downlink transmission, and each RRU can only select one mode. Indicates uplink user Signal-to-interference-to-noise ratio, This represents the minimum signal-to-interference-plus-noise ratio (SIR) that uplink users need to achieve. Indicates the first The transmission power of each active uplink user Set the maximum uplink transmission power for the user; Step S302, set the optimization target for downlink spectral efficiency based on the finite block length transmission mechanism: (13), in formula (13), Indicates downlink users Signal-to-interference-to-noise ratio, This represents the minimum signal-to-interference-plus-noise ratio (SIR) that downlink users need to achieve. This represents the block error probability. This represents the maximum tolerable error probability. It is the downlink transmission RRU and the first Channel precoding between downlink users The threshold for downlink user transmission power. Indicates the first The transmission power of the downlink active users; Step S303, in the network-assisted full-duplex CF-RAN system, when the uplink demand is high, the number of uplink serving RRUs can be increased to improve uplink spectrum efficiency, and when the downlink demand is high, the number of downlink serving RRUs can be increased to improve downlink spectrum efficiency; because at a certain moment, the transmission power of the first active downlink user; Each RRU can only select either uplink or downlink mode. The optimization objectives of formulas (12) and (13) are contradictory. Therefore, we set up a multi-objective optimization problem for RRU uplink / downlink mode selection: (14) Considering multi-objective optimization, design an up-and-down scheduling scheme for RRUs to achieve a trade-off between the two problems and maximize the system gain; Step S304, the intelligent mode selection algorithm based on DQN is as follows: The DQN algorithm combines deep learning and reinforcement learning, and has high reliability for solving action selection problems based on discrete variables. In the DQN algorithm, the agent is based on the state-action value function Q and Greedy solution, based on a fixed probability Choose the action with the highest Q value. , obtain the reward function and enter the next state. By storing the actions taken and the rewards obtained in memory, the agent can continuously train its neural network model to obtain the optimal solution; the loss function for minimizing the squared error of the agent's neural network parameters is: (15), in formula (15), It is the first Step movement The reward function, As a discount factor, To maximize the next state Q-value function, For a set of actions, It is in state Select action The Q-value function; in the DQN-based intelligent uplink / downlink mode selection algorithm, the CCU is regarded as an intelligent agent, and the state-space function is... Defined as , Indicates in In-step RRU uplink / downlink selection vectors; action space function Defined as , Indicates in Changes in uplink / downlink selection of the RRU during the step; reward function Defined as: (16), in formula (16), This represents the total spectrum efficiency for uplink users. This represents the total spectral efficiency of downlink users under a finite block length transmission mechanism. It is a regularization parameter that ensures network convergence. It is a constant parameter related to the sum of uplink and downlink user spectral efficiency.

2. The network-assisted full-duplex mode optimization method for low-altitude three-dimensional coverage scenarios as described in claim 1, characterized in that, Step S4 specifically includes: Step S401, initial setup Initialization state and the neural network parameters of the agent; step S402, based on probability - Greedy strategy for choosing actions Step S403, based on the current state and the chosen action Calculate the Q-value function Step S404: The current state transitions to the next state based on the selected action. Step S405: Calculate the reward function according to formula (16). Step S406, the joint transition vector Stored in the memory pool, and trained into a neural network according to formula (15). Step S407, if The output of the optimal state is the optimal RRU uplink / downlink selection scheme. The maximum reward is the maximum gain the system can obtain; otherwise, return to step S402.

Citation Information

Patent Citations

  • Network-assisted full-duplex cellular-free large-scale MIMO duplex mode optimization method

    CN113078929A

  • Accuracy distribution method for analog-to-digital converter of network-assisted full duplex system

    CN115801072A