Multi-base station cooperative communication optimization method based on reinforcement learning

By employing local beamforming and Wyner-Ziv distributed source coding in a multi-base station cooperative communication system, and combining reinforcement learning agents to optimize the global precoding matrix, the problems of excessive fronthaul link bandwidth requirements and quantization noise impact were solved, achieving efficient multi-base station cooperative communication and improving network performance and user experience.

CN122513804APending Publication Date: 2026-08-04TIANYUAN RUIXIN COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANYUAN RUIXIN COMM TECH CO LTD
Filing Date
2026-07-07
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing multi-base station cooperative communication optimization schemes in C-RAN or centralized non-cellular architectures result in a linear increase in fronthaul link bandwidth requirements with the number of antennas and bandwidth, exceeding the fiber capacity limit. Furthermore, the quantization noise introduced by lossy compression erodes the multi-user interference suppression effect, leading to a dilemma where the more compression is applied, the worse the cooperation becomes.

Method used

A multi-base station cooperative communication optimization method based on reinforcement learning is adopted. By performing local beamforming in the remote unit, the high-dimensional antenna domain signal is reduced to a low-dimensional user data stream. The Wyner-Ziv distributed source coding joint quantizer is used for compression. The global precoding matrix is ​​generated by combining reinforcement learning agents, and the mutual information between the quantization noise covariance matrix and the user interference channel is optimized to minimize the impact of quantization noise on the desired signal.

Benefits of technology

It significantly reduces the amount of data in the fronthaul link, avoids the amplification of quantization noise, improves the robustness and real-time performance of the system in complex dynamic environments, maximizes user throughput, and reduces precoding feedback overhead through differential pulse code modulation and adaptive quantization, forming a stable distributed collaborative closed loop.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513804A_ABST
    Figure CN122513804A_ABST
Patent Text Reader

Abstract

The application discloses a multi-base station cooperative communication optimization method based on reinforcement learning and belongs to the technical field of wireless communication processing; through 7-2 layer function segmentation, beamforming is sunk to a remote unit, and local beamforming weight is calculated by using global precoding auxiliary information fed back by a central unit at a previous moment, so that high-dimensional antenna signals are reduced to low-dimensional user data streams; a joint quantizer based on Wyner-Ziv distributed source coding is adopted, spatial correlation between data streams of different remote units is used for compression, and a compression codebook and a global precoding matrix are jointly designed, so that mutual information between a quantization noise covariance matrix and a user interference channel is minimized; a multi-user channel matrix is reconstructed based on compressed data streams, and a reinforcement learning intelligent agent is introduced, an adaptive cooperative precoding strategy is output through intelligent decision-making, and differential pulse code modulation and adaptive quantization are adopted to efficiently feed back the global precoding matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication processing technology, and more specifically to a multi-base station cooperative communication optimization method based on reinforcement learning. Background Technology

[0002] Multi-base station cooperative communication aims to break the limitations of traditional base stations operating independently, allowing geographically separated transmission points (such as base stations in different cells) to work collaboratively to provide services to the same or a group of users. In traditional cellular networks, users at the cell edge often suffer from strong co-channel interference from adjacent base stations, resulting in weak signals, slow speeds, or even dropped connections. Multi-base station cooperative communication can effectively transform harmful inter-cell interference into useful signals, thereby significantly improving overall network performance, especially the experience for edge users. Multi-base station cooperative communication typically involves collaborative processing and coordinated scheduling / beamforming (CS / CB) methods.

[0003] When existing multi-base station cooperative communication optimization schemes are implemented in C-RAN or centralized non-cellular architectures, if baseband IQ samples are processed in a completely centralized manner, the bandwidth requirements of the fronthaul link (e.g., CPRI / eCPRI) will increase linearly with the number of antennas and bandwidth, quickly exceeding the fiber capacity limit. If lossy compression is forcibly performed, the introduced quantization noise will be amplified by the cooperative precoding matrix, severely eroding the multi-user interference suppression effect, forming a unique dilemma where the more compression, the worse the cooperation, resulting in the capacity-fidelity limitation of the fronthaul / backhaul link. Summary of the Invention

[0004] The purpose of this invention is to provide a multi-base station cooperative communication optimization method based on reinforcement learning, which can solve the technical problems mentioned in the background art.

[0005] The objective of this invention can be achieved through the following technical solutions: Multi-base station cooperative communication optimization methods based on reinforcement learning include: S1: Each remote unit performs local beamforming on the uplink multi-user signal according to the acquired local channel state information, reducing the high-dimensional antenna domain received signal to a low-dimensional user data stream, and sending it to the central unit through the fronthaul link; wherein, the local beamforming adopts 7-2 layer functional segmentation, and the beamforming weight matrix is ​​calculated based on the global precoding auxiliary information fed back by the central unit in the previous time step. S2: After receiving the low-dimensional user data streams sent by each remote unit, the central unit uses a joint quantizer based on Wyner-Ziv distributed source coding to compress each data stream and simultaneously generate a compressed codebook; the design of the compressed codebook is jointly optimized with the global precoding matrix in subsequent steps to minimize the mutual information between the quantization noise covariance matrix and the user interference channel. S3: The central unit reconstructs the multi-user channel matrix based on the compressed user data stream and generates a global precoding matrix using the cooperative precoding strategy output by the reinforcement learning agent; the state space of the reinforcement learning agent contains the rate-distortion performance index of the compressed codebook. S4: The central unit feeds back the global precoding matrix to each remote unit through the fronthaul link. Each remote unit updates its local beamforming weights for downlink transmission or uplink reception in the next time slot.

[0006] Furthermore, when acquiring local channel state information, in the first... At the start of each time slot For time slot index, the first Each remote unit performs channel estimation using the uplink reference signal. For the remote cell index, an exponentially weighted moving average is used to smooth the channel to reduce the impact of estimation errors on subsequent beamforming.

[0007] Furthermore, in the Time slot, remote unit It stores the global precoding matrix fed back from the central unit of the previous time slot. ,in, The total number of remote units, For the user data stream dimension, and Extract the submatrix corresponding to the remote cell from the global precoding matrix. .

[0008] Furthermore, local beamforming is performed to reduce the high-dimensional antenna signal to a user data stream for the remote unit. The received uplink antenna domain signal vector is , For remote cell indexing, The number of antennas for this remote unit is determined, and the calculated local beamforming weighting matrix is ​​used. To perform linear dimensionality reduction, the corresponding expression is: ;in, The user data stream after dimensionality reduction; For user data flow dimensions; The conjugate transpose of the local beamforming weight matrix.

[0009] Furthermore, the local beamforming weighting matrix The following feature projection method is used to ensure that the local dimensionality reduction matrix is ​​aligned with the global precoding direction: ;in, It projects the matrix onto the identity column orthogonal constraint set. Operators on, The matrix to be projected, such as the local beamforming weight matrix. for The conjugate transpose of . for Gram matrix, It is the identity matrix. This is the conjugate transpose of the smoothed channel matrix; This is the submatrix corresponding to the remote unit; This is a time slot index.

[0010] Furthermore, when constructing the joint data flow vector and distributed source coding model, the central unit collects the data flows from all remote units to form a global data vector. Define the compressed global data vector as Quantize noise vector Its covariance matrix is , This is the mathematical expectation symbol.

[0011] Furthermore, when constructing the coset quantizer structure based on Wyner-Ziv, for the ... Data stream of each remote unit The central unit maintains a global codebook. ,in, For total bit count budgeting; this codebook is shared by all remote units; in addition, each The quantification process is carried out independently.

[0012] Furthermore, the joint optimization objective of the compressed codebook and the global precoding matrix is ​​mathematically expressed as: ;in, For compressed codebook Find the minimum value of the precoding matrix W; For mutual information; To quantize noise; To interfere with the user channel matrix; For user symbol vectors; For compressed codebook; For constraint symbols; Maximum allowable mean square quantization distortion; This represents the feasible set of precoding matrices; The square of the Euclidean norm; This is the original global data vector; This is the compressed reconstructed global data vector; This is due to mean square quantization distortion; Let be the mathematical expectation.

[0013] Furthermore, during codebook iterative updates, the central unit uses a block coordinate descent method to alternately update the codebook. and precoding matrix .

[0014] Furthermore, let the uplink signal model be: ;in, This is the original global data vector; This is the global channel matrix; Emit symbols for the user, assuming ; It is additive white Gaussian noise with a covariance of , This represents noise power.

[0015] Furthermore, the compressed observation model is as follows: ;in, This is the compressed reconstructed data vector; This is the global channel matrix; Emit symbols for users; It is additive white Gaussian noise with a covariance of , For noise power, It is the identity matrix; Will and Combined into equivalent noise Its covariance matrix is ; To quantify the noise covariance.

[0016] Furthermore, using known pilot symbols , This represents the total number of currently active users. The length of the pilot signal is [length], and the corresponding compressed pilot received data is [data]. That is, the corresponding pilot time period value, The total number of remote units, For the user data stream dimension, the LMMSE estimate of the channel is: ;in, For the reconstructed multi-user channel matrix; for The conjugate transpose of .

[0017] Furthermore, when designing the state space of a reinforcement learning agent, time slots are defined. status for: ;in, For the reconstructed multi-user channel matrix; This is the output quantization noise covariance matrix; The instantaneous signal-to-noise ratio of the fronthaul link for each remote unit; This represents the remaining capacity of each fronthaul link; Rate-distortion performance index for compressed codebook is defined as the ratio of the actual distortion of the data stream of the i-th remote unit to the theoretical lower limit of rate-distortion. The reinforcement learning agent outputs continuous and discrete actions; the continuous action is the regularization factor. , This is the minimum value of the regularization factor. This represents the maximum value of the regularization factor. Discrete actions for precoding mode selection Where 0 represents regularization zero-forcing and 1 represents quantization noise alignment precoding in this embodiment.

[0018] Furthermore, the agent's policy network adopts an Actor-Critic architecture; the Actor network will store the state... Mapped to the mean and variance of a Gaussian distribution to generate continuous actions. Meanwhile, discrete actions are output through the Softmax layer. Probability distribution; Critic network output state value function , used for advantage estimation.

[0019] Furthermore, when generating the global precoding matrix using the agent's output, the central unit determines the precoding pattern based on the agent's output. and regularization factor Combined with the reconstructed channel Calculate the global precoding matrix Specifically: when hour, , This is a regularized zero-forcing precoding matrix; To reconstruct the conjugate transpose of the channel; when At that time, for reconstructing the channel Perform singular value decomposition: ;in, For channel The left singular matrix; It is a singular value diagonal matrix; For channel The conjugate transpose of the right singular matrix; Extract the channel subspace corresponding to the interfering users. Assuming all non-target users are interference sources, the interference channel matrix is ​​then obtained. for The remaining part after removing the target user column is used to construct the expression for the optimization problem as follows: ; ; in, Find the minimum value of variable W; W is the decision variable, i.e., the global precoding matrix; The objective function is to obtain the received power of interfering users; Align the precoding matrix to quantize noise; V is the inverse of the channel covariance matrix; V is the null basis matrix; is the conjugate transpose of the zero-space basis matrix; This is the compressed Gram matrix.

[0020] Furthermore, when implementing block and differential coding of the global precoding matrix, the central unit first divides the generated global precoding matrix into blocks according to the far-end units to obtain the precoding sub-matrices corresponding to the far-end units; Differential coding is performed on the precoding matrices of adjacent time slots. The difference matrix is ​​defined, the maximum absolute value of the difference matrix is ​​calculated, and each element of the difference matrix is ​​uniformly quantized. The quantized integer index is encoded as an unsigned integer and packed together with the sign bit.

[0021] Furthermore, when the remote unit decodes and recovers the precoding, after the i-th remote unit receives the feedback frame, it first performs a CRC check. If the check passes, it parses the differential index and quantization step size. It then recovers the estimated value of the differential matrix through inverse quantization and recovers the precoding submatrix of the current time slot using the precoding submatrix of the previous time slot stored locally.

[0022] Furthermore, when updating the local beamforming weight matrix, the far-end cell obtains the precoding submatrix of the current time slot. Then, its local beamforming weight matrix needs to be updated. ; For uplink reception, local beamforming weights The uplink received signal should be able to This is equivalent to the inverse process of downlink precoding; utilizing the reciprocity of TDD, the uplink channel is equal to the transpose of the downlink channel, therefore the optimal uplink combining matrix should be the conjugate of the downlink precoding matrix; thus, the update rule is: ;in, Apply the updated local beamforming weight matrix; This indicates that the matrix is ​​orthogonalized; right Perform QR decomposition: ;in, The orthogonality factor for QR decomposition; The upper triangular factor of QR decomposition; Pick .

[0023] Compared to existing solutions, the beneficial effects achieved by this invention are: This invention uses 7-2 layer functional segmentation to push beamforming down to the far-end unit and uses the global precoding auxiliary information fed back by the central unit in the previous moment to calculate the local beamforming weight, reducing the high-dimensional antenna signal to a low-dimensional user data stream, which significantly reduces the amount of data to be transmitted in the fronthaul link. At the same time, it enables the local processing to be pre-aligned to the global cooperative direction, providing a structured, high-quality input for subsequent joint compression and intelligent precoding. This invention employs a joint quantizer based on Wyner-Ziv distributed source coding, utilizing the spatial correlation between data streams from different remote units for compression. It also jointly designs a compressed codebook and a global precoding matrix to minimize the mutual information between the quantization noise covariance matrix and the user interference channel. While further compressing the amount of forward transmission data, it actively "guides" the energy of quantization noise to the user interference subspace, fundamentally avoiding the problem of quantization noise being amplified by cooperative precoding in traditional lossy compression, and breaking through the technical dilemma that the more compression, the worse the cooperation.

[0024] This invention reconstructs a multi-user channel matrix based on compressed data streams and introduces a reinforcement learning agent whose state space explicitly includes the rate-distortion performance index output in step S2. Through intelligent decision-making, it outputs an adaptive collaborative precoding strategy, enabling the system to perceive the cost of fronthaul compression and channel quality in real time, dynamically select the optimal precoding method, maximize user throughput under limited fronthaul capacity, and improve robustness and real-time performance in complex dynamic environments.

[0025] This invention employs differential pulse code modulation and adaptive quantization to efficiently feed back the global precoding matrix. The remote unit orthogonals the received precoding submatrix through QR decomposition and updates the local beamforming weights, significantly reducing the fronthaul overhead of precoding feedback. It also utilizes the reciprocity of the TDD channel to form a symmetrical closed loop of "downlink precoding-uplink merging". At the same time, through periodic full retransmission and verification mechanisms, it ensures the long-term stable convergence of distributed collaboration. Attached Figure Description

[0026] The invention will now be further described with reference to the accompanying drawings.

[0027] Figure 1 This is a flowchart of the multi-base station cooperative communication optimization method based on reinforcement learning according to the present invention.

[0028] Figure 2This is a flowchart of the sub-steps included in step S1 of the present invention.

[0029] Figure 3 This is a flowchart of the sub-steps included in step S2 of the present invention.

[0030] Figure 4 This is a flowchart of the sub-steps included in step S3 of the present invention.

[0031] Figure 5 This is a flowchart of the sub-steps included in step S4 of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] like Figures 1 to 5 As shown, this invention is a multi-base station cooperative communication optimization method based on reinforcement learning, comprising: S1: Each remote unit performs local beamforming on the uplink multi-user signal according to the acquired local channel state information, reducing the high-dimensional antenna domain received signal to a low-dimensional user data stream, and sending it to the central unit through the fronthaul link; wherein, the local beamforming adopts 7-2 layer functional segmentation, and the beamforming weight matrix is ​​calculated based on the global precoding auxiliary information fed back by the central unit in the previous time step. In C-RAN or centralized cellular architecture, traditional methods upload all raw IQ samples received by each remote unit (RRU / AP) to the central unit (CU), causing the fronthaul link bandwidth requirement to increase linearly with the number of antennas. This step adopts the 7-2 layer function partitioning defined by 3GPP to push beamforming / precoding functions down to the remote unit, so that the fronthaul only transmits the user data stream that has been reduced in dimensionality.

[0034] In addition, step S1 includes the following sub-steps: obtaining local channel state information, calculating local beamforming weights based on the global precoding auxiliary information from the previous time step, performing local beamforming to reduce the high-dimensional antenna signal to user data stream, data stream conversion and frame encapsulation, and periodically resetting and updating the local beamforming weight matrix; specifically: When obtaining local channel state information, at the first... At the start of each time slot For time slot index, the first Each remote unit performs channel estimation using the uplink reference signal. For remote unit indexing, uplink reference signals such as SRS are used to obtain the local instantaneous channel matrix. ,in, This represents the number of antennas for the remote unit. This represents the total number of currently active users. To reduce the impact of estimation errors on subsequent beamforming, an exponentially weighted moving average is used to smooth the channel; the corresponding expression is: ;in, This is the current smoothed channel matrix. As a smoothing factor, , This is the channel matrix after smoothing the previous time slot; this smoothing operation utilizes the time correlation of the channel to suppress the estimated jitter caused by transient noise and fast fading.

[0035] When calculating local beamforming weights based on the global precoding auxiliary information from the previous time step, traditional local beamforming often operates independently of the central unit's collaborative processing, leading to front-end and back-end mismatch. This embodiment feeds back global precoding auxiliary information from the central unit to the remote units, enabling local beamforming to pre-adapt to the global collaborative direction. Specifically: In the Time slot, remote unit It stores the global precoding matrix fed back from the central unit of the previous time slot. ,in, The total number of remote units, For the user data stream dimension, and Extract the submatrix corresponding to the remote cell from the global precoding matrix. ; Local beamforming weighting matrix The following feature projection method is used to ensure that the local dimensionality reduction matrix is ​​aligned with the global precoding direction: ;in, It projects the matrix onto the identity column orthogonal constraint set. The operator on the matrix is ​​specifically implemented through QR decomposition: for the intermediate matrix Perform QR decomposition ,Pick ; The matrix to be projected, such as the local beamforming weight matrix. for The conjugate transpose of . for The Gram matrix; this projection ensures that the data streams remain orthogonal after local beamforming, avoiding inter-stream interference within the far-end unit. It is the identity matrix. For column orthogonal matrices, , It is an upper triangular matrix. This is the conjugate transpose of the smoothed channel matrix. This is the symbol for the conjugate transpose.

[0036] When performing local beamforming to reduce the dimensionality of high-dimensional antenna signals to user data streams, the remote unit... The received uplink antenna domain signal vector is This includes the superposition of all user-transmitted signals after passing through the channel, as well as thermal noise, and utilizes the calculated local beamforming weight matrix. Perform linear dimensionality reduction: ;in, The user data stream after dimensionality reduction; The conjugate transpose of the local beamforming weight matrix; this operation reduces the antenna dimension from Compressed to user data stream dimension The amount of data transmitted from the front end has been reduced to the original amount. ;because It is usually equal to the number of users, or slightly greater than the number of users, while The values ​​can be very large, such as 64 or 128, resulting in significant dimensionality reduction. It should be noted that the local beamforming in this step does not completely eliminate interference between users, but rather retains enough information for subsequent global collaborative processing by the central unit.

[0037] When implementing data stream quantization and frame encapsulation, in order to adapt to the digital transmission format of the fronthaul link, the dimensionality-reduced data stream is... Analog-to-digital conversion and quantization are performed. Considering that the fronthaul link may have different capacities, this embodiment uses adaptive block floating-point quantization: In The data stream is grouped by symbol blocks, with each group sharing an exponential factor; the number of quantization bits. The remaining link capacity is dynamically determined by the link remaining capacity indication broadcast by the central unit at the previous moment, and the corresponding expression is: ;in, The maximum number of effective bits in the analog-to-digital converter of the remote unit; For the first Total capacity of the fronthaul links; The bandwidth occupied by overhead includes the bandwidth occupied by the overhead of non-user data stream transmission such as control signaling, synchronization, Ethernet frame header, CRC check, and co-tags. The signal sampling rate; The symbol period is the reciprocal of the signal sampling rate. The quantized data stream, along with the quantization index, is packaged into an Ethernet frame and sent to the central unit via the fronthaul link.

[0038] To prevent severe mismatch of local weights due to global precoding matrix feedback delay or channel changes during the periodic reset and update of the local beamforming weight matrix, this step also introduces a periodic reset mechanism: every [time period missing]... During a time slot, the remote unit temporarily disconnects from the network. Instead of relying on conventional maximum ratio beamforming, it uses conventional maximum ratio beamforming. Perform one uplink transmission and send the original estimated channel to the central unit via fronthaul. for The Frobenius norm is used; the central unit receives the data, recalculates the global precoding matrix, and feeds it back to correct the accumulated error; this mechanism ensures the long-term stability of the distributed cooperative closed loop.

[0039] In this embodiment, through the coordinated operation of the above steps, while maintaining the global optimization capability of the central unit, the bandwidth requirement of the fronthaul link is reduced from... Reduce to , Unlike existing simple antenna selection or fixed beamforming, the local beamforming weights in this step explicitly utilize the global precoding auxiliary information fed back by the central unit in the previous time step, so that local processing and global coordination form a prediction-correction closed loop. This provides a structured low-dimensional input data stream for the joint quantization compression in the subsequent steps S2 and the intelligent precoding generation in S3. Moreover, this data stream has been pre-aligned to the global coordination direction, which greatly reduces the processing complexity of the subsequent central unit and the risk of quantization noise amplification.

[0040] S2: After receiving the low-dimensional user data streams sent by each remote unit, the central unit uses a joint quantizer based on Wyner-Ziv distributed source coding to compress each data stream and simultaneously generate a compressed codebook; the design of the compressed codebook is jointly optimized with the global precoding matrix in subsequent steps to minimize the mutual information between the quantization noise covariance matrix and the user interference channel. After completing step S1, the central unit (CU) receives data from... Low-dimensional user data stream of each remote unit Each Although step S1 reduces the antenna dimension from [previous dimension] through local beamforming... Down to However, if all the data streams from remote units are directly spliced ​​together and processed centrally, the total amount of data transmitted to the frontend will still be [data missing]. Even with complex-valued symbols, the fronthaul capacity may still be exceeded in high user density or high bandwidth scenarios. To address this, this step introduces a joint quantizer based on Wyner-Ziv distributed source coding. By utilizing the spatial correlation between signals received by different remote units, joint compression is performed at the central unit side. Simultaneously, the design of the compressed codebook is coordinated with the global precoding matrix in the subsequent step S3 for optimization, guiding quantization noise to the user interference subspace and minimizing its impact on the desired signal.

[0041] Furthermore, step S2 includes the following sub-steps: constructing a joint data stream vector and a distributed source coding model; constructing a coset quantizer structure based on Wyner-Ziv; jointly optimizing the compressed codebook and the global precoding matrix; iterative codebook updates; coset partitioning and side information utilization; and outputting the compressed codebook and quantization noise covariance. Specifically: When constructing the joint data stream vector and distributed source coding model, the central unit collects the data streams from all remote units to form a global data vector. : ;in, The data stream for the i-th far-end unit comes from the local beamforming output in step S1; between different far-end units, due to their spatial proximity or overlapping line-of-sight paths, and There is statistical correlation between them; classical source coding theory shows that for distributed correlated sources, Wyner-Ziv coding can achieve the same rate-distortion performance as joint coding without exchanging side information; this step takes advantage of this characteristic, under the constraint of limited fronthaul link capacity, and uses the central unit as the decoding end to independently quantize and encode the data streams of each remote unit. The encoding ends do not communicate with each other, and the decoding ends use the correlation between all received compressed data streams for joint decoding, thereby reducing the total compression bit rate; Define the compressed global data vector as Quantize noise vector Its covariance matrix is , The mathematical expectation symbol; traditional methods will Treated as independent additive noise, this step actively shapes it through subsequent joint optimization. The structure is aligned with the user's interference channel, thereby reducing damage to the desired signal.

[0042] When constructing a coset quantizer structure based on Wyner-Ziv, this invention uses coset quantization as a specific implementation of Wyner-Ziv encoding; for the... Data stream of each remote unit The central unit maintains a global codebook. ,in, For total bit count budgeting; this codebook is shared by all remote units; in addition, each The quantification process is carried out independently and includes: Find the global codebook with Euclidean closest codeword and output the index. Due to different There are correlations between them, and redundancy exists among these indices. The idea behind Wyner-Ziv encoding is that at the encoding end, i.e., the remote unit side, the original index is not sent directly, but rather the "coset" label to which the index belongs is sent. At the decoding end, i.e., the central unit, the index information from other remote units is used to uniquely determine the original codeword. Specifically, the global codebook... Divided into Each coset contains a set of 1, and each coset contains 2, 3, 4, 5, 6, 7, 8, 9, 1, 1, 1, 1, 1, 2, 3 ... One codeword; the encoding end sends a co-collection tag. , The number of bits for the coset tag, and its bit rate The decoding end receives all co-set tags. Then, by utilizing correlation, the most likely codeword combination is selected within the coset to reconstruct the codeword. This is a conventional technical solution, and the specific implementation steps will not be elaborated here. In this embodiment, the coset partitioning method is dynamically adjusted according to the lower bound of the mutual information between the data streams of each remote unit, and the compression ratio allocation coefficient is output by the reinforcement learning agent in subsequent steps. The decision is made to form a closed-loop adaptive system.

[0043] The core innovation of this embodiment lies in integrating the design of the compressed codebook with the global precoding matrix in subsequent step S3 when implementing the joint optimization objective of the compressed codebook and the global precoding matrix. Alternatively, a precoding strategy can be used for joint optimization, rather than independent design. The optimization objective is to minimize the mutual information between quantization noise and user interference channels, that is, to ensure that the energy of quantization noise is mainly distributed in the user interference subspace while minimizing its impact on the desired signal subspace. Mathematically, this can be expressed as: ;in, For compressed codebook Find the minimum value of the precoding matrix W; For mutual information; To quantize noise; The channel matrix for interfering users is estimated by the central unit based on the uplink reference signal; For user symbol vectors; The codebook is a compressed codebook, i.e., the codebook of the designed joint quantizer, which contains all possible quantization reconstruction vectors; For constraint symbols; Maximum allowable mean square quantization distortion; This is a feasible set of precoding matrices, such as transmit power constraints or daily antenna power constraints; The square of the Euclidean norm; To mitigate mean square quantization distortion; since directly solving for minimizing mutual information is quite difficult, this embodiment derives its convex upper bound, transforming it into a more manageable form.

[0044] Based on the data processing inequalities and lemmas in information theory, specifically including: ;in, To quantize the noise covariance matrix, , for The inverse matrix; Let the user's symbolic covariance matrix be an identity matrix, i.e. ; It is the determinant of a matrix; It is a logarithm to the base 2; the joint optimization problem can be approximated as: ;in, These are Lagrange multipliers used to balance the mutual information and distortion terms; the objective function's bootstrap codebook The design tends to: under a given distortion budget, make the quantization noise covariance matrix... The inverse and interference channel covariance matrix The eigenvalues ​​of the product should be as small as possible, that is, the noise should mainly fall in the orthogonal complement space of the interference subspace.

[0045] When implementing iterative codebook updates, the central unit uses a block coordinate descent method to alternately update the codebook. and precoding matrix , The update is described in step S3, where the precoding is considered known and codebook optimization is performed; for fixed... The expression corresponding to codebook optimization is: ; in, For compressed codebook Find the minimum value; It is the statistical expectation of the data vector y; The square of the Frobenius norm of the quantization error; For codebook-based nearest neighbor quantizers, It is the noise covariance matrix generated by the quantizer, estimated by Monte Carlo sampling; The regularization parameters are those after transformation. Since the objective function is not differentiable with respect to the codeword, this embodiment adopts an annealing soft allocation strategy: replacing hard quantization with soft allocation based on Gaussian kernels, and using stochastic gradient descent to iteratively optimize the codebook on the channel sample set. Specifically, for a batch of training samples Where n is the sample index and N is the total number of samples, the soft quantization output is defined as: ;in, This is the soft quantization output of the nth sample. Let be the squared Euclidean distance between the sample and the codeword. Let k be the k-th codeword in the codebook, where k is the codeword index. It is an exponential function. The temperature parameter gradually decreases to zero with each iteration, simulating annealing; j is the summation index; Among them, quantization noise covariance This allows for soft allocation estimation and computation of the objective function for each codeword. The gradient is calculated and updated; after each iteration, the codeword is projected onto the transmit power constraint set, for example... , For vectors The square of the Euclidean norm, Maximum allowed codeword power; When implementing coset partitioning and edge information utilization, in the codebook After convergence, the central unit determines the coset partitioning method based on the mutual information matrix of the data streams from each remote unit, and calculates the coset partitioning method for any two remote unit data streams. and Mutual information estimate between Construct a correlation plot; For strongly correlated remote unit groups, i.e. Assign coarser cosets, i.e., smaller ones. ,For example Then there are only two cosets, each of which is very large, so that edge information can be used for joint decoding; for weakly correlated far-end units, i.e. To allocate finer cosets, that is, larger ones ,For example Then there is Each coset reduces error propagation; the coset partitioning scheme is communicated to each remote unit through the control channel of the fronthaul link. The threshold is estimated and determined by the system designer based on experience or offline simulation.

[0046] The decoding end receives all co-set tags Then, perform joint typical decoding: search all codeword combinations. , requiring each Belongs to the auxiliary group And the overall sequence Satisfies the joint probability density constraint; Select the combination that satisfies the conditions and minimizes the total distortion as the reconstruction. The decoding process is approximated on the correlation graph using a low-complexity message passing algorithm, such as belief propagation. When outputting the compressed codebook and quantization noise covariance, after completing the above iterative optimization, the central cell obtains the final compressed codebook. And record the current quantization noise covariance matrix. The current quantization noise covariance matrix The information will be passed to the reinforcement learning agent in step S3 as part of the state space to guide the selection of the global precoding strategy. At the same time, the central unit will feed back the coset partitioning table to each remote unit through the forward link so that the remote units can send coset tags in subsequent time slots according to the new coset rules.

[0047] In this embodiment, through joint optimization and distributed coding of the above sub-steps, compared with traditional independent quantization, such as antenna-by-antenna scalar quantization, joint quantization based on Wyner-Ziv reduces the total compression bit rate by utilizing spatial correlation. At the same time, through the joint design of codebook and precoding, quantization noise is guided to interference subspace, so that the precoding matrix in subsequent step S3 can effectively suppress the impact of this noise on the desired user signal, thereby breaking through the dilemma that the more compression, the worse the coordination.

[0048] S3: The central unit reconstructs the multi-user channel matrix based on the compressed user data stream and generates a global precoding matrix using the cooperative precoding strategy output by the reinforcement learning agent; the state space of the reinforcement learning agent contains the rate-distortion performance index of the compressed codebook. In step S2, the central unit completes the joint quantization compression of the low-dimensional user data stream based on Wyner-Ziv, and obtains the reconstructed data stream. and quantization noise covariance matrix The purpose of this step is to back-estimate the multi-user channel matrix based on the compressed reconstructed data stream, and at the same time design a reinforcement learning agent. This agent takes the rate-distortion performance index including the output of step S2 as its state, dynamically selects the optimal cooperative precoding strategy, and uses the strategy output by the agent to generate a global precoding matrix, which is then fed back to the remote unit in step S4.

[0049] Furthermore, step S3 includes the following sub-steps: reconstructing the multi-user channel matrix based on the compressed data stream, designing the state space of the reinforcement learning agent, designing the action space and policy network output, designing the reward function, generating the global precoding matrix based on the agent's output, training and updating the reinforcement learning agent, and outputting the global precoding matrix; specifically: When reconstructing a multi-user channel matrix based on compressed data streams, since the quantization and compression in step S2 introduces distortion, the reconstructed data streams are used directly. Channel estimation can introduce biases. This embodiment employs a Bayesian linear minimum mean square error (LMMSE) estimator, utilizing the quantization noise covariance obtained in step S2. As prior information, the channel is recovered from the reconstructed data stream; Let the uplink signal model be: ;in, This is the original global data vector; This is the global channel matrix; Emit symbols for the user, assuming ; It is additive white Gaussian noise with a covariance of , Noise power; The compression process in step S2 can be modeled as follows: ;in, This is the compressed reconstructed data vector; To quantify the noise, its covariance The result has already been obtained from step S2; therefore, the compressed observation model is: ;in, This is the compressed reconstructed data vector; Will and Combined into equivalent noise Its covariance matrix is ; Using known pilot symbols , The length of the pilot signal is [length], and the corresponding compressed pilot received data is [data]. That is, the corresponding pilot time period The LMMSE estimate for the channel is: ;in, To reconstruct the multi-user channel matrix, this estimation formula utilizes the quantization noise covariance to whiten the observation noise, thus avoiding the propagation of channel estimation bias caused by compression distortion.

[0050] When designing the state space of a reinforcement learning agent, this embodiment adopts a deep reinforcement learning framework, specifically a proximal policy optimization (PPO), whose state space... It not only includes conventional channel quality information, but also explicitly incorporates the rate-distortion performance index of the compressed codebook in step S2, thus explicitly including the cost of fronthaul compression in the decision-making process; it defines time slots. The status is: ;in, The reconstructed multi-user channel matrix reflects the current channel quality; The output is the quantization noise covariance matrix, with the diagonal elements representing the compression distortion power of each data stream; The instantaneous signal-to-noise ratio of the fronthaul link for each remote unit is obtained through physical layer monitoring. , where is the instantaneous signal-to-noise ratio of the i-th remote unit forward transmission; The remaining capacity of each fronthaul link. , where is the remaining capacity of the i-th fronthaul link; The rate-distortion performance metric for compressed codebooks is defined as the ratio of the actual distortion of the data stream in the i-th far-end unit to the theoretical lower bound of rate-distortion: ; ;in, For actual mean square distortion, To quantize the number of bits The lower bound of Wyner-Ziv rate distortion theory; This represents the actual number of quantized bits for the i-th remote unit. The data stream dimension for each remote unit; this metric reflects whether the current compression efficiency is close to the theoretical optimum, and is used to guide the agent whether it needs to adjust the precoding strategy to tolerate or compensate for additional distortion.

[0051] The reinforcement learning agent outputs actions in two dimensions to address the forward capacity-fidelity constraint, specifically including continuous actions and discrete actions; where: Continuous Actions: Regularization Factor The minimum value of the regularization factor is usually the Maximum regularization factor Used for RZF precoding; Discrete Actions: Precoding Mode Selection Where 0 represents regularized zero-forcing (RZF) and 1 represents quantization noise alignment precoding (QNA) in this embodiment. The agent's policy network adopts an Actor-Critic architecture; the Actor network stores the state... Mapped to the mean and variance of a Gaussian distribution to generate continuous actions. Meanwhile, discrete actions are output through the Softmax layer. Probability distribution; Critic network output state value function , used for advantage estimation.

[0052] The reward function guides the agent to achieve a balance between user throughput, forward overhead, and interference suppression fidelity; the designed reward function is as follows: ;in, The reward function value; Let u be the received signal-to-interference-plus-noise ratio for user u; u is the user index. This is the maximum capacity of the fronthaul link, used for normalization; The global precoding matrix generated for the current time slot. ; For the zero-forcing precoding matrix under ideal uncompressed conditions, ; These are the regularization coefficients for the forward compression ratio and the interference suppression bias, respectively. Typical values ​​were obtained through offline simulation tuning. ; It is the Frobenius norm.

[0053] When generating the global precoding matrix using the agent's output, the central unit uses the precoding pattern output by the agent. and regularization factor Combined with the reconstructed channel Calculate the global precoding matrix Specifically: when hour, , This is a regularized zero-forcing precoding matrix; This precoding method is used to reconstruct the conjugate transpose of the channel. It strikes a trade-off between suppressing inter-user interference and suppressing noise amplification, making it suitable for scenarios with low quantization noise.

[0054] when At that time, for reconstructing the channel Perform singular value decomposition: ;in, For channel The left singular matrix; It is a singular value diagonal matrix; For channel The conjugate transpose of the right singular matrix; Extract the channel subspace corresponding to the interfering users. Assuming all non-target users are interference sources, the interference channel matrix is ​​then obtained. for The remaining part after removing the target user column is used to construct the expression for the optimization problem as follows: ; ; in, Find the minimum value of variable W; W is the decision variable, i.e., the global precoding matrix; The objective function is to obtain the received power of interfering users; Align the precoding matrix to quantize noise; V is the inverse of the channel covariance matrix; V is the interference channel matrix. The matrix formed by the null space basis vectors can be decomposed using singular value decomposition. get, Let be the left singular matrix of the interference channel. The singular value matrix of the interference channel. It is part of the right singular matrix of the interference channel. This is another part of the right singular matrix of the interference channel; is the conjugate transpose of the zero-space basis matrix; The compressed Gram matrix; this precoding guarantees that the desired user's channel response is forced to zero while minimizing the signal energy received by interfering users; due to the quantization noise covariance Energy distribution and The structure is related, and this QNA precoding can simultaneously suppress inter-user interference and quantization noise.

[0055] During the training and updating of reinforcement learning agents, the central unit continuously collects trajectory data during runtime. ,in The PPO algorithm is used to update the agent. The loss function of PPO includes policy loss and value function loss, and the corresponding expression is: ;in, Let be the loss function of the PPO algorithm in time slot t; For sampling expectation; As importance weight, ; For policy network parameters; For generalized advantage estimation; This is the clipping function; This is the clipping hyperparameter, with a typical value of 0.2; These are the value function loss coefficient and entropy coefficient, respectively, with typical values ​​of 0.5 and 0.01. To accumulate the target value for the discount reward, , For instant rewards, This is the discount factor, with a value range of (0,1], and a typical value of 0.99. Discount factor Power of 1 For time offset index; This is the policy entropy, used to encourage exploration; For policy functions; The state value function; After each precoding decision, the central unit uses rewards Perform several PPO updates; since channel changes are relatively slow, the update frequency can be set to every Once per time slot, to reduce computational overhead.

[0056] After completing the above calculations, the central unit obtains the final global precoding matrix. , dimension This is equivalent to a precoded vector for each user; the matrix will be sent to step S4, where it will be differentially modulated and fed back to each remote unit via the fronthaul link for downlink transmission in the next time slot.

[0057] In this embodiment, quantized noise covariance is used for LMMSE channel reconstruction to suppress the impact of compression distortion on channel estimation. The constructed reinforcement learning state space explicit inclusion rate distortion performance index enables the agent to perceive compression efficiency and dynamically select RZF or QNA precoding. Quantized noise aligned precoding guides noise energy to the interference subspace, forming a complete closed loop with the preceding joint codebook design, significantly improving the multi-user interference suppression capability under compressed fronthaul. Through the synergy of the above steps, it is ensured that even when the fronthaul link is severely limited, the system can still maintain high spectral efficiency and reliability.

[0058] S4: The central unit feeds back the global precoding matrix to each remote unit through the fronthaul link. Each remote unit updates its local beamforming weights for downlink transmission or uplink reception in the next time slot.

[0059] In step S3, the central unit (CU) generates a global precoding matrix based on the cooperative precoding strategy output by the reinforcement learning agent. Or, equivalently, assign a precoding vector to each user. The core task of this step includes: The feedback is efficiently transmitted to each remote unit (RRU or AP) via the fronthaul link; each remote unit updates its local beamforming weight matrix based on the received feedback information. This allows for downlink transmission or uplink reception to be performed in the next time slot. Due to the limited bandwidth of the fronthaul link and the need for the feedback information to be aligned with the downlink transmission time slot, this embodiment uses a combination of low-bit differential pulse code modulation (DPCM) and adaptive quantization to compress the precoding matrix, and combines it with global precoding auxiliary information based on the previous time slot to form a closed-loop update mechanism. In addition, step S4 includes the following sub-steps: block and differential coding of the global precoding matrix, adaptive quantization and coding, fronthaul link frame encapsulation and transmission, remote unit decoding and precoding recovery, updating the local beamforming weight matrix, closed-loop verification and error correction, and downlink transmission preparation; specifically: When implementing block and differential coding of the global precoding matrix, the central unit first generates the global precoding matrix. Divide into blocks based on remote units: , ;in, Let be the precoding submatrix corresponding to the i-th far-end unit; in a C-RAN centralized cellular architecture, the total number of antennas in the M far-end units is much greater than the number of users, therefore The number of rows may be much greater than the number of columns.

[0060] To reduce feedback overhead, this embodiment utilizes the temporal correlation of the precoding matrix to perform differential coding on the precoding matrices of adjacent time slots, defining a differential matrix. : ;in, This is the precoding submatrix used by the i-th remote unit in the previous time slot. The precoding submatrix has been stored in the local buffer of the remote unit. Because channel changes are relatively slow, especially in low-mobility scenarios, The amplitude is much smaller than Therefore, it can be quantized using fewer bits; For difference matrix An adaptive step-size uniform quantizer is used, and the quantization step size is dynamically adjusted according to the remaining capacity of the fronthaul link. Maximum absolute value : ;in, For matrix vectorization operators, Let represent the real part and the imaginary part of the difference matrix, respectively; The maximum value function is based on the forward link remaining capacity vector observed by the reinforcement learning agent. The allocation strategy is as follows: prioritize links with high remaining capacity to receive more bits, as shown in the following formula: ;in, Allocate the number of quantization bits to the i-th remote unit; The total bit budget allocated to precoding feedback per time slot; Quantization step size for: ; Perform uniform quantization on each element of the difference matrix: ;in, These are the elements of the quantized difference matrix; It is a symbolic function; The r-th row of the difference matrix is ​​the first... Column elements; To round to the nearest integer; Quantized integer index It is encoded as an unsigned integer, and packed together with the sign bit; To further compress the data, the quantized integer indexes are sent using run-length encoding because a large number of elements in the difference matrix are close to zero.

[0061] When encapsulating and transmitting the fronthaul link frame, the central unit encapsulates the quantization differential index and quantization step size of each remote unit. The timestamp is encapsulated into an independent Ethernet frame and sent to the corresponding remote unit via the fronthaul link. At the same time, to enhance robustness, each frame header carries a cyclic redundancy check (CRC) code. If the remote unit detects a check error during decoding, it discards the frame and uses the precoding matrix of the previous time slot.

[0062] During the decoding and recovery of the pre-encoded data by the remote unit, after the i-th remote unit receives the feedback frame, it first performs a CRC check. If the check passes, it parses the differential index and quantization step size; then, it recovers the estimated value of the differential matrix through inverse quantization. : ; Using the precoding submatrix of the previous time slot stored locally Initially set as an identity matrix or a random matrix, the precoding submatrix of the current time slot is recovered. : ;in, This is the precoding submatrix of the previous time slot; Due to the existence of quantization error It may not meet the transmit power constraints, for example This may exceed the total transmit power; therefore, the remote unit performs power normalization on the recovered precoding matrix: ; in, This represents the upper limit of the system's total transmit power. The total power of the precoded matrix of all remote units is squared; this normalization implicitly assumes that the power allocation information between the central unit and the remote units has been synchronized through other control channels or a distributed consensus mechanism has been adopted.

[0063] When updating the local beamforming weight matrix, the far-end cell obtains the precoding submatrix of the current time slot. Then, its local beamforming weight matrix needs to be updated. In step S1, local beamforming is used for uplink reception dimensionality reduction, while in this step... Used for downlink transmit precoding. In time-division duplex (TDD) systems, leveraging channel reciprocity, the uplink beamforming weights and downlink precoding matrix have a dual relationship. This embodiment designs the following update rule to ensure consistency between downlink precoding and uplink beamforming: For uplink reception, local beamforming weights The uplink received signal should be able to This is equivalent to the inverse process of downlink precoding; utilizing the reciprocity of TDD, the uplink channel is equal to the transpose of the downlink channel, therefore the optimal uplink combining matrix should be the conjugate of the downlink precoding matrix; thus, the update rule is: ;in, Apply the updated local beamforming weight matrix; This involves orthogonalizing the matrix, for example, by obtaining the Q matrix through QR decomposition, to ensure... , It is a K×K identity matrix; specifically, for Perform QR decomposition: ;in, The orthogonality factor for QR decomposition; The upper triangular factor of QR decomposition; Pick This orthogonalization step ensures that the data streams are independent of each other after local beamforming, which is consistent with the operation of projecting to the orthogonal constraint set in step S1.

[0064] If the system uses Frequency Division Duplex (FDD), the uplink and downlink channels are not reciprocal. In this case, the remote unit needs to rely on the explicit channel state information (CSI) fed back by the central unit to perform uplink beamforming. To maintain versatility, in this embodiment, the central unit simultaneously feeds back differential coded versions of the downlink precoding matrix and the uplink combining matrix in FDD mode. For the sake of simplicity, this embodiment is described in TDD mode.

[0065] During closed-loop verification and error correction, to prevent closed-loop divergence due to quantization errors or channel aging, the remote unit, after updating its local weights, uses the uplink reference signal received in the next time slot, such as the SRS, to re-estimate the local channel and calculate the equivalent channel quality using the new weights. If a mismatch is found with the precoding matrix fed back by the central unit, for example, after projection... With the expected identity matrix If the deviation exceeds the threshold, the remote unit will send a retransmission request to the central unit, requesting a full retransmission of the precoding matrix, but without performing differential coding; simultaneously, the central unit will... The full precoding matrix is ​​broadcast once per time slot to reset the accumulated error of differential coding; This is the full retransmission period, and its value is a positive integer.

[0066] During downlink transmission preparation, after the local beamforming weights are updated, the remote unit can use [the technology] in the downlink phase of the next time slot. The user symbols are pre-coded and transmitted via the antenna; specifically, let the downlink user symbol vector be... Then the transmitted signal of the remote unit i is: ;in, This is the transmitted signal of the i-th remote unit; Because the transmitted signals from each remote unit naturally superimpose in the air, combined with the processing at the user receiver, spatial multiplexing and interference suppression through multi-base station collaboration can be achieved; simultaneously, this precoding matrix will also be stored as... It is used for differential coding and decoding in the next time slot.

[0067] In this embodiment, through the coordination of the above sub-steps, efficient precoding matrix feedback in the fronthaul link and closed-loop update of local beamforming weights in the remote unit are achieved. This step and the beamforming weight matrix calculation in step S1 based on the global precoding auxiliary information fed back by the central unit at the previous moment form a complete loop of distributed collaborative optimization. Compared with the conventional scheme, this step uses differential pulse coding and adaptive quantization to significantly reduce feedback overhead. The feedback precoding matrix is ​​mapped to uplink receiving beamforming weights through QR decomposition, making full use of TDD reciprocity to form a symmetrical closed loop of downlink precoding-uplink merging. At the same time, the periodic full retransmission and verification mechanism ensures long-term stability.

[0068] In the several embodiments provided by this invention, it should be understood that the disclosed methods can be implemented in other ways. For example, the embodiments of the invention described above are merely illustrative; for instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation.

[0069] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0070] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or in the form of hardware plus software functional modules.

[0071] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the essential characteristics of the present invention.

[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-base station cooperative communication optimization method based on reinforcement learning, characterized in that, include: S1: Each remote unit performs local beamforming on the uplink multi-user signal according to the acquired local channel state information, reducing the high-dimensional antenna domain received signal to a low-dimensional user data stream, and sending it to the central unit through the fronthaul link; wherein, the local beamforming adopts 7-2 layer functional segmentation, and the beamforming weight matrix is ​​calculated based on the global precoding auxiliary information fed back by the central unit in the previous time step. S2: After receiving the low-dimensional user data streams sent by each remote unit, the central unit uses a joint quantizer based on Wyner-Ziv distributed source coding to compress each data stream and simultaneously generate a compressed codebook; the design of the compressed codebook is jointly optimized with the global precoding matrix in subsequent steps to minimize the mutual information between the quantization noise covariance matrix and the user interference channel. S3: The central unit reconstructs the multi-user channel matrix based on the compressed user data stream and generates a global precoding matrix using the cooperative precoding strategy output by the reinforcement learning agent; the state space of the reinforcement learning agent contains the rate-distortion performance index of the compressed codebook. S4: The central unit feeds back the global precoding matrix to each remote unit through the fronthaul link. Each remote unit updates its local beamforming weights for downlink transmission or uplink reception in the next time slot.

2. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 1, characterized in that, Perform local beamforming to reduce the dimensionality of the high-dimensional antenna signal to a user data stream; remote unit The received uplink antenna domain signal vector is , For remote cell indexing, The number of antennas for this remote unit is determined, and the calculated local beamforming weighting matrix is ​​used. To perform linear dimensionality reduction, the corresponding expression is: ;in, The user data stream after dimensionality reduction; For user data flow dimensions; The conjugate transpose of the local beamforming weight matrix.

3. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 2, characterized in that, Local beamforming weighting matrix The following feature projection method is used to ensure that the local dimensionality reduction matrix is ​​aligned with the global precoding direction: ;in, It projects the matrix onto the identity column orthogonal constraint set. Operators on, The matrix to be projected, such as the local beamforming weight matrix. for The conjugate transpose of . for Gram matrix, It is the identity matrix. This is the conjugate transpose of the smoothed channel matrix; This is the submatrix corresponding to the remote unit; This is a time slot index.

4. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 3, characterized in that, The joint optimization objective of implementing the compressed codebook and the global precoding matrix is ​​mathematically expressed as: ;in, For compressed codebook Find the minimum value of the precoding matrix W; For mutual information; To quantize noise; To interfere with the user channel matrix; For user symbol vectors; For compressed codebook; For constraint symbols; Maximum allowable mean square quantization distortion; This represents the feasible set of precoding matrices; The square of the Euclidean norm; This is the original global data vector; This is the compressed reconstructed global data vector; This is due to mean square quantization distortion; Let be the mathematical expectation.

5. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 4, characterized in that, The compressed observation model is as follows: ;in, This is the compressed reconstructed data vector; This is the global channel matrix; Emit symbols for users; It is additive white Gaussian noise with covariance of , For noise power, It is the identity matrix; Will and Combined into equivalent noise Its covariance matrix is ; To quantify the noise covariance.

6. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 5, characterized in that, Using known pilot symbols , This represents the total number of currently active users. The length of the pilot signal is [length], and the corresponding compressed pilot received data is [data]. , The LMMSE estimate for the total number of remote units is: ;in, For the reconstructed multi-user channel matrix; for The conjugate transpose of .

7. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 6, characterized in that, When designing the state space of a reinforcement learning agent, time slots are defined. status for: ;in, For the reconstructed multi-user channel matrix; This is the output quantization noise covariance matrix; The instantaneous signal-to-noise ratio of the fronthaul link for each remote unit; This represents the remaining capacity of each fronthaul link; Rate-distortion performance index for compressed codebook is defined as the ratio of the actual distortion of the data stream of the i-th remote unit to the theoretical lower limit of rate-distortion. The reinforcement learning agent outputs continuous and discrete actions; the continuous action is the regularization factor. , This is the minimum value of the regularization factor. This represents the maximum value of the regularization factor. Discrete actions for precoding mode selection Where 0 represents regularization zero-forcing and 1 represents quantization noise alignment precoding in this embodiment.

8. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 7, characterized in that, When generating the global precoding matrix using the agent's output, the central unit uses the precoding pattern output by the agent. and regularization factor Combined with the reconstructed channel Calculate the global precoding matrix Specifically: when hour, , This is a regularized zero-forcing precoding matrix; To reconstruct the conjugate transpose of the channel; when At that time, for reconstructing the channel Perform singular value decomposition: ;in, For channel The left singular matrix; It is a singular value diagonal matrix; For channel The conjugate transpose of the right singular matrix; Extract the channel subspace corresponding to the interfering users. Assuming all non-target users are interference sources, the interference channel matrix is ​​then obtained. for The remaining part after removing the target user column is used to construct the expression for the optimization problem as follows: ; ; in, Find the minimum value of variable W; W is the decision variable, i.e., the global precoding matrix; The objective function is to obtain the received power of interfering users; Align the precoding matrix to quantize noise; V is the inverse of the channel covariance matrix; V is the null basis matrix; is the conjugate transpose of the zero-space basis matrix; This is the compressed Gram matrix.

9. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 8, characterized in that, When implementing block and differential coding of the global precoding matrix, the central unit first divides the generated global precoding matrix into blocks according to the far units to obtain the precoding submatrix corresponding to the far units; Differential coding is performed on the precoding matrices of adjacent time slots. The difference matrix is ​​defined, the maximum absolute value of the difference matrix is ​​calculated, and each element of the difference matrix is ​​uniformly quantized. The quantized integer index is encoded as an unsigned integer and packed together with the sign bit.

10. The multi-base station cooperative communication optimization method based on reinforcement learning according to claim 9, characterized in that, When the remote unit decodes and recovers the precoding, after the i-th remote unit receives the feedback frame, it first performs CRC check. If it passes, it parses the differential index and quantization step size. It recovers the estimated value of the differential matrix through inverse quantization and recovers the precoding submatrix of the current time slot using the precoding submatrix of the previous time slot stored locally.