Active Listening Method Based on Deep Reinforcement Learning

By using an active monitoring method based on deep reinforcement learning in large-scale MIMO-OFDM systems, the precoding matrix is ​​optimized and the power allocation factor is adjusted, and the problem of low monitoring efficiency under narrow beams is solved, and efficient data monitoring is achieved.

CN114884547BActive Publication Date: 2025-07-01NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210312148.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-28
Publication Date
2025-07-01
Estimated Expiration
2042-03-28

AI Technical Summary

Technical Problem

When narrow beams are used in 5G base stations, traditional monitoring methods become less efficient and cannot successfully monitor suspicious communication links outside the beam coverage range.

Method used

Adopting an active monitoring method based on deep reinforcement learning, the precoding matrix is ​​optimized by the monitor during the beam scanning stage, the transmitter is induced to select the beam index that is conducive to monitoring, and maximize the data monitoring rate by adjusting the power distribution factor and power gain factor in the data transmission stage.

Benefits of technology

Effectively induce the transmitter to select beams that are beneficial to the monitor, improve the data monitoring rate, and ensure that suspicious communication links can be successfully monitored in narrow beam scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114884547B_ABST
    Figure CN114884547B_ABST
Patent Text Reader

Abstract

The present invention discloses an active listening method based on deep reinforcement learning, belonging to the field of communications. In a large-scale MIMO-OFDM system, when the listener E and the suspicious receiver D are not within the coverage range of the same communication beam, traditional passive listening and active listening schemes become inefficient or even ineffective. To achieve legal listening in a large-scale MIMO-OFDM system, the listener is used as a pseudo-relay to achieve beam induction and data listening. When the transmitter S performs beam scanning, the listener E induces the transmitter to select a beam favorable for listening by optimizing the relay precoding matrix. In the data listening phase, the listener E improves the listening rate by optimizing the relay power allocation factor and the power gain factor. Since the channel state information of the suspicious communication link is unknown, the optimal precoding matrix and power allocation factor are found through the deep reinforcement learning algorithm - MADDPG. Computer simulations verify the effectiveness of the proposed design scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communications, and particularly relates to an active listening method based on deep reinforcement learning, and more particularly to an active listening method in a large-scale MIMO-OFDM (Multiple Input Multiple Output-Orthogonal Frequency Division Multiplexing) system based on deep reinforcement learning. Background Art

[0002] The MIMO-OFDM technology is considered a key technology for the fifth-generation (5G) mobile network. However, when advanced beamforming technology is adopted in a 5G base station, the directional narrow beam makes the efficiency of traditional listening methods low or even ineffective. Therefore, in order to achieve legal listening to suspicious links, it is crucial to study the listening scheme in the narrow beam scenario.

[0003] The existing literature on listening can be divided into three categories: passive listening, interference-based active listening, and spoofing relay-based active listening. In passive listening, the listener remains silent during listening, that is, only the data sent by the transmitter is listened to. This method is only effective when the listening channel is better than the suspicious channel. To overcome this shortcoming, the method of interference-based active listening is introduced, that is, the listener sends interference signals to the suspicious receiver, forcing the transmitter to reduce the rate, so that the information can be decoded by the listener. To flexibly implement active listening, a listening method called spoofing relay is proposed. When the listening channel is better than the suspicious channel, this method can maximize the listening rate by disguising the listener as a relay. However, when the transmitter sends information to the suspicious user with a directional beam, none of the above listening schemes can enable the listener outside the beam coverage to successfully listen. Summary of the Invention

[0004] Object of the Invention: Aiming at the defects of the above-mentioned existing technologies, the present invention studies a scheme for a listener outside the beam coverage in a MIMO-OFDM system to successfully listen to a suspicious communication link, and proposes an active listening method in a large-scale MIMO-OFDM system based on deep reinforcement learning to ensure that even if the transmitter uses a narrow beam to communicate with the suspicious receiver, the communication data can still be successfully listened to.

[0005] Technical Solution: An active listening method based on deep reinforcement learning includes the following steps:

[0006] (1) The transmitter S performs analog beam scanning in a time-division manner according to the beam precoding matrix.

[0007] (2) During the beam scanning phase when transmitter S is performing beam scanning, listener E determines the best beam index j that is beneficial to itself based on its own beam quality report and the beam report fed back to transmitter S by receiver D. * ;

[0008] (3) The listener E induces the transmitter S to select the best beam index j by optimizing the forwarding precoding matrix * ;

[0009] (4) In beam j * In the communication phase after confirmation, the listener E acts as a pseudo-relay for data forwarding, maintaining the communication beam and improving the data listening rate.

[0010] Furthermore, the step (2) includes the following steps:

[0011] 1) The receiver D and the listener E respectively receive the beam quality measurement reference signal sent by the transmitter S, and calculate the beam quality according to the received signal. The receiver D forms a beam quality report and feeds it back to the transmitter S as a reference for beam selection;

[0012] 2) The listener E determines the optimal beam index j based on its own beam quality report and the beam quality report fed back to the transmitter S by the listener receiver D, taking into account the power consumption factor, and finally determines the optimal beam index j according to the trade-off formula between the beam induction success rate and the power consumption. * .

[0013] The step (3) comprises the following steps:

[0014] 1. Listener formation optimization problem: Minimize the total transmit power of the listener under the constraint of successful beam guidance. The form of the optimal precoding matrix is ​​derived according to the optimization problem. It is found that the optimal precoding matrix is ​​related to the channel state information of the transmitter S and the receiver D.

[0015] ㈡The listener E uses the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) algorithm to train the first fitting network to determine the transmission parameters of the first forwarding matrix, and then uses the first forwarding matrix determined by the transmission parameters to forward the beam quality measurement reference signal to the receiver D, inducing the receiver D to send an erroneous beam measurement report, so that the transmitter S selects a beam that is beneficial to the listener E.

[0016] The step (4) comprises the following steps:

[0017] i The listener E receives the transmission data sent by the transmitter S and forms an optimization problem: maximize the data monitoring rate under the conditions of successful monitoring and the transmission power is less than the upper limit of the forwarding power;

[0018] ii. The listener uses the MADDPG algorithm to train the second fitting network to determine the power allocation factor and the power gain factor, allowing a part of the power to be used for decoding and a part for forwarding signals. Then, it forwards the communication data to receiver D using the second forwarding matrix to maintain the communication beam and improve the data listening rate.

[0019] Further, the step (ii) includes the following steps:

[0020] ① Model the beam induction problem as a first multi-agent collaborative MDP (Markov Decision Process) problem;

[0021] ② According to the form of the optimal precoding matrix, transform the problem of finding the optimal precoding matrix into the problem of finding a pair of constants, thereby accelerating the training process; at a specific moment, the action on a single subcarrier is the angle and amplitude of the precoding matrix. Therefore, the actions of all subcarriers are the set of actions on a single subcarrier;

[0022] ③ At a specific moment, the state on a single subcarrier is the beam report information obtained by listening to and analyzing the feedback channel plus the known channel information, and the global state is the union of non-overlapping information of all subcarrier states;

[0023] ④ Design the reward function at a specific moment to encourage successful beam induction while punishing behaviors that consume too much energy.

[0024] Further, the step ii includes the following steps:

[0025] I. Model the data listening problem as a second multi-agent collaborative MDP problem;

[0026] II. According to the form of the optimal precoding matrix, transform the problem of finding the optimal precoding matrix into the problem of finding a pair of constants, thereby accelerating the training process; at a specific moment, the actions on a single subcarrier are the power gain factor and the power allocation factor. Therefore, the actions of all subcarriers are the set of actions on a single subcarrier;

[0027] III. At a specific moment, the state on a single subcarrier is the signal-to-interference-plus-noise ratio obtained by listening to and the feedback channel information plus the known channel information, and the global state is the union of non-overlapping information of all subcarrier states;

[0028] IV. Design the reward at a specific moment to encourage the subcarrier to maximize the listening rate under the constraints of successful listening and power limitation.

[0029] Beneficial effects: The present invention is applicable to monitoring suspicious communication links in a narrow-beam large-scale MIMO-OFDM system. During the process of beam scanning and determining the beam at the transmitter, the listener realizes beam induction by optimizing the precoding matrix. During the process of the transmitter transmitting data, the listening rate is maximized by optimizing the power allocation factor and the power gain factor. Considering that it is difficult for the listener to obtain the channel information between suspicious nodes, the present invention proposes a learning scheme based on MADDPG to assist in beam induction and data listening. The active listening method in the large-scale MIMO-OFDM system based on deep reinforcement learning proposed by the present invention can not only effectively induce the transmitter S to select a beam beneficial to the listener E, laying a foundation for the subsequent data listening process, but also enable the listener E to readjust the power splitting factor and the power gain factor, effectively maintaining the communication link and improving the data listening rate. Description of the Drawings

[0030] Figure 1 is the active listening model diagram in the large-scale MIMO-OFDM system of the present invention;

[0031] Figure 2 is the action diagram of the listener E at different transmission stages of the transmitter S (BS and DT are abbreviations for the beam scanning and data transmission stages);

[0032] Figure 3 is the transceiver structure diagram of the listener E of the present invention;

[0033] Figure 4 is the relationship diagram of the beam induction success rate and the transmission power under different N te configurations of the present invention;

[0034] Figure 5 is the relationship diagram of the transmission power of the listener E and the transmission power of the transmitter S under different N te configurations of the present invention;

[0035] Figure 6 is the average listening rate diagram under different P S and N ts conditions of the present invention;

[0036] Figure 7 is the average listening rate diagram of various listening methods of the present invention. Detailed Embodiment

[0037] The present invention proposes a monitoring method in a large-scale MIMO-OFDM system based on traditional pseudo-relay monitoring, in which a legitimate full-duplex relay is used to achieve beam induction and data monitoring. The present invention assumes that analog beamforming is employed on the suspicious transmitter and uses beam scanning to select the optimal beam vector. Beam induction is completed in the beam scanning stage. The purpose of beam induction is to induce the suspicious receiver to select a beam that is beneficial to the monitor. To achieve this goal, the monitor acts as a relay, amplifying and forwarding the measurement reference signal of the desired beam to the suspicious receiver. In this stage, the goal of the present invention is to minimize the total transmission power of the monitor under the constraint of successful beam induction by optimizing the precoding matrix of the monitor. Through mathematical derivation, a closed-form expression of the optimal precoding matrix is calculated, which is related to the CSI (Channel State Information) of the suspicious communication pair. When the monitor does not know the CSI between them, the present invention uses the DRL (Deep Reinforcement Learning) algorithm - MADDPG (Multi-Agent Deep Deterministic Policy Gradient) to determine the transmission parameters of all subcarriers. Once beam induction is achieved, the monitor can implement data monitoring and improve the monitoring rate by continuing to act as a pseudo-relay. In this stage, the power splitting factor and the power gain factor are optimized to maximize the monitoring rate. Similarly, since the monitor does not know the CSI of the suspicious communication pair, the present invention still uses MADDPG to optimize the relay parameters of the monitor.

[0038] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings:

[0039] The application scenario of the present invention is as Figure 1 shown: The present invention considers a legitimate monitoring system, which consists of a pair of suspicious communication nodes (transmitter S and receiver D) and a legitimate monitor E. The transmitter S and the receiver D are respectively equipped with N ts transmitting antennas and N rd receiving antennas. The present invention assumes that both the transmitter S and the receiver D use large-scale MIMO-OFDM arrays with analog beamforming to transmit and receive information. The analog beam is selected from a predefined discrete codebook, and the present invention represents the codebook of the transmitter S as The monitor E acts as a full-duplex pseudo-relay, receiving signals from the transmitter S through N re antennas, and at the same time transmitting signals through N teThe antenna forwards the signal to the receiver D. To improve the monitoring quality, the listener E adopts digital beamforming technology on each subcarrier. The present invention assumes that all channels in the system remain unchanged in each RB (Resource Block), but may vary between different RBs according to the Markov model.

[0040] In the solution of the present invention, as Figure 2 shown, for the transmitter S, the entire process of each transmission block is divided into two stages: the BS (Beam Sweeping) stage and the DT (Data Transmission) stage. For the listener E, the monitoring process includes three stages: beam selection, beam-induced and deceptive data forwarding. In the beam selection stage, the listener obtains beam quality information by listening to the feedback channel. Specifically, when the transmitter S transmits with the beamforming vector , the signals received by the receiver D and the listener E on the k-th subcarrier can be expressed as

[0041]

[0042] and

[0043]

[0044] where s k is the transmission signal of the transmitter and denotes taking the expectation, f j is the beamforming vector and |f j (n)| = 1, n = 1,..., N ts , j is the beam index, and are the transmission powers on the k-th subcarrier, the channel matrix between the transmitter S and the receiver D, the channel matrix between the transmitter S and the listener E, denotes the matrix dimension. and are zero-mean additive white Gaussian noise with a covariance matrix of σ 2 I. At the receiver of the receiver D, the received signal is processed using 's analog beamformer and ‖v D ‖ 2 = N rd , where ‖·‖ represents taking the modulus of a vector or the F-norm of a matrix. At the listener E, the received signal on the k-th subcarrier uses the digital beamformer Processing the signal. In the BS phase of the transmitter S, the receiver D, and the listener E, the SNR (Signal to Noise Ratio) at the k-th subcarrier of the receiver D and the listener E is

[0045]

[0046] and

[0047]

[0048] The receiver D calculates the average SNR of all subcarriers where K is the number of subcarriers, and selects J with large values as candidate beams, and then and the corresponding indices are fed back to S, where J is the maximum number of feedback beams. The present invention assumes that the listener E can obtain this feedback information by listening to the feedback channel between the transmitter S and the receiver D. When the beam selected by the transmitter S results in low It is difficult for the listener E to listen to the communication information transmitted by the transmitter S. Therefore, the listener E acts as a pseudo-relay to induce the beam selection of the transmitter S. For the listener E, the ideal beam should provide a high SNR for both the listener E and the receiver D, because a low beam will consume more forwarding power of the listener E. Therefore, the listener E determines the required optimal beam index according to where δ is a trade-off factor for balancing the listening success rate and power consumption of the listener.

[0049]

[0050] where δ is a trade-off factor for balancing the listening success rate and power consumption of the listener.

[0051] After determining the required optimal beam index j * the listener E will induce the transmitter S to select the optimal beam index j during the next BS * . In the DT phase, the transmitter S transmits communication data to the receiver D, and the listener E acts as an AF (Amplify and Forward) spoofing relay and listens while forwarding the data.

[0052] As Figure 3 shown, α k and g k respectively represent the received beamforming vector, the transmitted beamforming vector, the power allocation factor, and the power gain factor of the listener E at the subcarrier k. In the beam induction phase, the received signal is amplified and passed through α k= 1 transmission, i.e., no need to decode the information. In the data forwarding stage, the received signal power is divided into two parts for decoding and forwarding. This invention will analyze how to optimize to achieve the maximum listening rate.

[0053] During the beam scanning of the transmitter S, the listener E will amplify and forward the pilot signal received from the transmitter S for measuring the beam quality between the transmitter S and the receiver D. This invention assumes that the delay of the AF relay used by the listener is much smaller than the symbol duration and can thus be ignored. Due to the full-duplex nature of the listener E, the received signal at the listener E on subcarrier k is

[0054]

[0055] where is the self-interference channel, is the precoding matrix of the listener E, is the signal received by the listener at the previous moment. It can be seen from (6) that if W k is in the null space of then Let be the right singular matrix corresponding to the zero singular value of

[0056]

[0057] where is the new matrix to be optimized. To ensure r0 > 0, this invention has N te > N re , i.e., the listener E needs more transmit antennas than receive antennas to suppress self-interference. After eliminating self-interference, the transmitted signal can be expressed as The transmission power of the listener E is The signal received by the receiver D after receive beamforming can be expressed as

[0058]

[0059] where is the channel matrix between the listener E and the receiver D on the k-th subcarrier, and and are the newly constructed equivalent channels. Then, the received signal-to-noise ratio of the receiver D on the k-th subcarrier can be expressed as

[0060]

[0061] To complete the induced transmitter S to select the optimal beam index j with the minimum transmission power * , this problem can be formulated as

[0062]

[0063] where and To obtain a closed-form solution of, the present invention decomposes problem (10) into K independent sub-problems, and let where Therefore, these K sub-problems can be formulated as

[0064]

[0065] To solve (11), the present invention first proves a lemma: the optimal solution of problem (11) can be expressed as

[0066]

[0067] where The proof of the lemma is as follows:

[0068] For the sake of simplicity, the identification of subcarrier k will be omitted in the following proof of the lemma. To prove the lemma, the present invention assumes that the feasible solution of (11) is where w ′ is the amplitude parameter of the feasible solution, then the power consumption P(W′) corresponding to the feasible solution is Then the present invention constructs the matrix where The inequality therein follows that for any matrix (vector) A and B, ||AB|| ≤ ||A||||B||. It will be shown below that the new matrix W″ is not only feasible for problem (11), but also obtains an objective value smaller than P(W′). Let By substituting W″ into the numerator and denominator of β in (9) k , thus there is

[0069]

[0070]

[0071] where, (13) and (14) follow the triangle inequality. Based on (13) and (14), the present invention infers that β(W″) ≥ β(W′) ≥ β D . The above results show that W″ is feasible for problem (11). By substituting W″ into the objective function of (11), the present invention obtains

[0072]

[0073] where (15) follows the Cauchy - Schwarz inequality In summary, for any solution W′ of problem (11), the present invention can always construct another to obtain a smaller objective value, which proves this lemma.

[0074] Substituting (12) into the objective function in (11), it can be seen that is an increasing function of w k . Therefore, in (12), gradually increasing w k from a small value until the constraint in (11) is satisfied, the unique unknown variable w k can be found. Since the lemma applies to any given Therefore, the optimal solution of (10) has the same form as (12). In theory, the present invention can obtain the optimal solution of (10) by considering all possible combinations . For a given the solution of (11) provides an upper bound for the solution of (10). Only when the listener E can know all channels can the listener E adopt the precoding matrix in (12). The present invention assumes that the listener E can obtain the equivalent channel vector by listening to the pilot signal. However, due to the non - cooperative relationship between the transmitter S and the listener E, it is difficult to obtain Therefore, the present invention adjusts D according to the feedback β and β of DRL to minimize P E . By adopting the learning framework of MADDPG, determine in real - time to induce the transmitter S to select the beam required by the listener E. Finally, W k in (7) can be expressed as the product of the column vector and the row vector , as shown in Figure 3 .

[0075] Successful beam induction does not mean that listening can be successfully carried out. In the data transmission phase, if the listener E does not forward the data to the receiver D, the bit error rate of the receiver D may be higher than the threshold and trigger the beam recovery process, thus switching the beam. Therefore, in order to achieve data relay and listening of the listener E under AF relay operation, the received signal is divided into two parts, one part is used to forward information to increase the signal - to - noise ratio of the receiver D, and the other part is used for information decoding to listen to the message sent by the transmitter S. Due to the introduction of α k ​ The power gain factor w in k needs to be re-optimized. Define as the normalized beamforming vector, then the transmitted signal of the listener E is expressed as

[0076]

[0077] where g k is the power gain factor, which is used to control the transmission power in the data listening stage, and α k is the power allocation factor. It should be noted that and are consistent with beam induction because in both stages, the present invention aims to improve the signal-to-noise ratio of the receiver D. Similar to (8), the received signal of the receiver D in the data transmission stage can be written as

[0078]

[0079] For a given and the signal-to-noise ratios of the received receiver D and listener E can be calculated as and Then, the goal of the listener E is to optimize so that the listening rate reaches the maximum under the constraint of the transmission power. Therefore, the optimization problem can be expressed as

[0080]

[0081] where, and P M are the total transmission power and power constraint of the listener E respectively. The present invention assumes that the listener E can only achieve listening when R E ≥R D and the corresponding listening rate is R D . If the listener E knows the global CSI, the solution of (18) can be derived by the Lagrange multiplier method. However, when is unknown, the present invention cannot obtain the optimal A more reasonable assumption than knowing is that can be obtained by listening to the uplink control channel between the transmitter S and the receiver D. Therefore, DRL is adopted to as the observation state and interact with the system to determine By using MADDPG to train the neural network, is given in real time,

[0082] Based on the above analysis, when the CSI between the transmitter S and the receiver D is unknown, the beam induction and data listening problems are formulated as an MDP (Markov Decision Process) problem. Considering all subcarriers as one agent and obtaining the policy through the Actor-Critic network of a single DDPG is the first-intuitive deep learning solution. However, in actual implementation, training a policy with a large action space is usually more difficult than training multiple policies with small action spaces. Therefore, in these two stages, the present invention regards each subcarrier as a separate agent, and they cooperate to achieve a common goal. Thus, the present invention adopts the learning architecture of MADDPG, which includes K Actors (policies) and a centralized Critic (value function). During the training stage, the Actors and the Critic are updated using global data, including the global state, shared reward, and all actions, which will be defined later.

[0083] The beam induction problem is modeled as a first multi-agent collaborative MDP problem. According to the form of the optimal precoding matrix, the problem of finding the optimal precoding matrix is transformed into the problem of finding a pair of constants (w k , θ k ), where θ k is the estimation of by the MADDPG algorithm, thus accelerating the training process. At time t, the action of the k-th subcarrier is represented by . Therefore, the actions of all subcarriers are . At time t, the state on each subcarrier k is β and β D are obtained by listening to and analyzing the beam reports on the feedback channel. The global state s t is the union of the non-overlapping information of all subcarrier states , that is, . The reward r t at time t is defined as r t = -a1P E - a2(β - β D - B) 2 + a3I(β, β D ), where is a positive coefficient used to balance the induction success rate and power consumption of the listener, B is a constant used to increase the probability of selecting the best beam index j * , and I(x, y) is a Boolean function, where I(x, y) = 1 when x ≥ y, otherwise I(x, y) = 0. The reward function encourages successful beam induction while punishing behaviors that consume too much energy.

[0084] Model the data monitoring problem as a second multi-agent collaborative MDP problem. According to the form of the optimal precoding matrix, transform the problem of finding the optimal precoding matrix into the problem of finding a pair of constants (g k , α k ), where g k and α k respectively represent the power gain factor and power allocation ratio of the listener on subcarrier k, thus accelerating the training process. At time t, the action of the k-th subcarrier is Therefore, the actions of all subcarriers are At time t, the state on each subcarrier k is where is obtained by listening to and feedback channel information. The global state s t is the union of non-overlapping information of all subcarrier states , that is The reward r t at time t is defined as where is a positive coefficient used to balance the listening rate and power consumption of the relay, and C is a constant used to improve the listening rate. The reward function encourages the subcarrier to maximize R E > R D and P E ≤ P M under the constraints of D .

[0085] As Figure 4 shows, in the beam induction phase, the relationship diagram between the beam induction success rate and the transmit power P te under different N S . The present invention allocates the same transmission power on each subcarrier, that is The induction rate is obtained by counting the number of β ≥ β 5 in 10 D Monte Carlo simulations. In the passive method, when the transmitter S performs beam scanning, the listener E remains silent. In this case, when the listener E and the receiver D are far apart, the receiver D will select the best beam index j * with a low probability. The results show that the success rate of the MADDPG-based method proposed by the present invention is close to 100%. These results verify the effectiveness of the method under different system configurations.

[0086] As Figure 5 shows, in the beam induction phase, the relationship diagram between P te and P E under different N S configurations. In the optimal scheme, P E is known in (10) Calculated optimal objective value. In the MADDPG-based scheme, P E is calculated using the parameters learned by MADDPG. It can be seen that although P E increases with the increase of P S , equipping more N te can effectively reduce P E . Combining Figure 3 and Figure 4 , it can be seen that even if is unknown, the present invention can still achieve beam induction using the beam induction strategy learned by MADDPG, and the transmission power is slightly higher than the theoretical minimum power.

[0087] As Figure 6 shown, in the data listening phase, the optimal solution is obtained by solving (18). The passive method with SBM (Successful Beam Misleading) means that the agent achieves beam induction in the BS phase but remains silent in the DT phase. As Figure 6 shown, after successful induction, the listening rate increases with the increase of P S or N ts , and the present invention can ensure R E ≥R D by adjusting the transmission parameters. At the same time, the listening rate of the method proposed by the present invention is close to the optimal solution and is significantly better than the passive listening method with SBM.

[0088] As Figure 7 shown, Figure 7 compares various listening schemes under different power constraints P M , and plots the traditional active interference scheme for comparison. The results show that the listening rate obtained by the MADDPG scheme proposed by the present invention is close to the optimal solution and increases with the increase of P M . When P M >55dBm, , the listening rate approaches the maximum R E . The listening performance of the passive listening scheme without SBM is independent of the transmission power of the listener E, and the listening performance of the passive listening with SBM is better than that of the method without SBM. The average eavesdropping rate of the interference scheme is limited by the power constraint of the listener E because it cannot ensure R M ≥R E ≥R D when the power limit value P M is relatively low.

[0089] Simulation proves that the active listening method in the large-scale MIMO-OFDM system based on deep reinforcement learning proposed by the present invention can not only effectively induce the transmitter S to select a beam favorable to the listener E, laying a foundation for the subsequent data listening process, but also enable the listener E to readjust the power allocation factor and power gain factor, effectively maintaining the communication link and improving the data listening rate. The combination of the two stages realizes a large-scale MIMO-OFDM system for listening to narrow-beam communications.

Claims

1. An active listening method based on deep reinforcement learning, comprising the following steps: (1) The transmitter S performs analog beam scanning in a time-division manner according to the beam precoding codebook; (2) During the beam scanning phase executed by the transmitter S, the listener E determines the optimal beam index j that is beneficial to itself based on its own beam quality report and the beam report feedback from the receiver D to the transmitter S. * ; (3) The listener E induces the transmitter S to select the optimal beam index j by optimizing the forwarding precoding matrix * ; (4) At the optimal beam index j * In the determined communication phase, the listener E acts as a pseudo-relay for data forwarding, maintains the communication beam, and improves the data listening rate; the following steps are included in the step (2): 1) The receiver D and the listener E respectively receive the beam quality measurement reference signals sent by the transmitter S, calculate the beam quality according to the received signals, and the receiver D forms a beam quality report and feeds it back to the transmitter S for beam selection reference; 2) The listener E determines the optimal beam index j according to its own beam quality report and the beam quality report fed back by the listening receiver D to the transmitter S, while considering the factor of power consumption, and finally based on the beam induction success rate and power consumption trade-off formula * ; The step (3) includes the following steps: (i) The listener forms an optimization problem: minimizing the total transmission power of the listener under the constraint of successful beam induction, derives the form of the optimal precoding matrix according to the optimization problem, and obtains that the optimal precoding matrix is related to the channel state information of the transmitter S and the receiver D; (ii) The listener E uses the MADDPG algorithm to train the first fitting network to determine the transmission parameters of the first forwarding matrix, and then uses the first forwarding matrix determined by the transmission parameters to forward the beam quality measurement reference signal to the receiver D, inducing the receiver D to send an incorrect beam measurement report, so that the transmitter S selects a beam favorable to the listener E; the step (4) includes the following steps: i The listener E receives the transmission data sent by the transmitter S and forms an optimization problem: maximizing the data listening rate under the conditions of successful listening and the transmission power being less than the forwarding power upper limit; ii The listener uses the MADDPG algorithm to train the second fitting network to determine the power allocation factor and the power gain factor, uses a part of the power for decoding and a part of the power for forwarding signals, and then uses the second forwarding matrix to forward the communication data to the receiver D to maintain the communication beam and improve the data listening rate; the step (ii) includes the following steps: ① Model the beam induction problem into a first multi-agent collaborative MDP problem; ② According to the form of the optimal precoding matrix, transform the problem of finding the optimal precoding matrix into the problem of finding a pair of constants, so as to speed up the training process; At a certain specific moment, the action on a single subcarrier is the angle and amplitude of the precoding matrix. Therefore, the actions of all subcarriers are the set of actions on a single subcarrier; ③ At a certain specific moment, the state on a single subcarrier is the beam report information obtained by listening to and analyzing the feedback channel plus the known channel information, and the global state is the union of the non-overlapping information of all subcarrier states; ④ The reward function design at a certain specific moment encourages successful beam induction and at the same time punishes behaviors that consume too much energy; the step ii includes the following steps: I Model the data listening problem into a second multi-agent collaborative MDP problem; II According to the form of the optimal precoding matrix, transform the problem of finding the optimal precoding matrix into the problem of finding a pair of constants, so as to speed up the training process; At a certain specific moment, the actions on a single subcarrier are the power gain factor and the power allocation factor. Therefore, the actions of all subcarriers are the set of actions on a single subcarrier; III At a certain specific moment, the state on a single subcarrier is the signal-to-interference-plus-noise ratio obtained by listening to and the feedback channel information plus the known channel information, and the global state is the union of the non-overlapping information of all subcarrier states; IV The reward design at a certain specific moment encourages the subcarrier to maximize the listening rate under the constraints of successful listening and power limitation.