Line-of-sight multi-antenna reliability and capacity cooperative improvement method and related equipment
By performing orthogonal mapping and reinforcement learning optimization on useful signals, the problem of improving the reliability and capacity of wireless communication systems under complex electromagnetic interference was solved, and the system adaptability and anti-interference capability were improved under unknown interference channel state information.
Patent Information
- Application Number
- CN202511859087.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-17
AI Technical Summary
In complex electromagnetic interference environments, existing technologies struggle to accurately adapt to interference characteristics in scenarios with unknown interference channel state information, and fail to effectively combine the orthogonal modes of discrete Fourier transform with the strong learning capabilities of artificial intelligence, thus limiting the improvement of reliability and capacity of wireless communication systems.
By modulating and orthogonally mapping the useful signal, a non-convex optimization problem is constructed with the goal of minimizing the system bit error rate. Combined with a reinforcement learning framework, the optimal transmit beamforming matrix and power allocation strategy are learned. The system monitors changes in interference channels in real time and dynamically iteratively optimizes the transmit strategy and receiver processing parameters to improve the system's communication reliability and capacity.
It achieves a synergistic improvement in the reliability and capacity of wireless communication systems under complex electromagnetic interference environments, solves the problem of insufficient adaptability of traditional methods in dynamic interference scenarios, and ensures long-term stable improvement in communication quality and capacity.
Smart Images

Figure CN121690299A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless communication, in particular to a line-of-sight multi-antenna reliability and capacity collaborative improvement method and related equipment. BACKGROUND
[0002] Due to the inherent broadcast characteristics and open channel environment, wireless communication systems are vulnerable to malicious interference, which seriously affects the communication reliability and effectiveness. Physical layer security technology has become an important research direction to cope with this challenge. Frequency hopping technology is the current mainstream anti-interference scheme, which reduces the interference probability by dynamically switching the carrier frequency of the communication parties according to the preset pseudo-random sequence. The introduction of artificial intelligence (AI), especially reinforcement learning technology, further improves the anti-interference performance of the frequency hopping system, but such frequency hopping methods based on reinforcement learning need to dynamically adjust the working frequency, resulting in significant synchronization overhead of the legal transmitting and receiving parties. Although some documents propose block offset pattern frequency hopping methods, which divide time blocks and dynamically generate frequency offset amount translation frequency hopping patterns based on reinforcement learning to balance the anti-interference performance and synchronization overhead, and some documents propose deep reinforcement learning wideband anti-interference frequency hopping strategies, which model the probabilistic interference mode of the jammer and realize optimal frequency hopping decision through Markov decision process to improve the throughput, but the traditional frequency hopping scheme and the existing artificial intelligence enabled frequency hopping strategy still have limited adaptability, relatively fixed pattern and other problems when facing complex electromagnetic countermeasures environment such as intelligent interference and wideband interference, and the overall anti-interference performance in the dynamic countermeasure scene is challenged.
[0003] In addition, there are documents that propose a beam frequency hopping anti-interference millimeter wave communication method based on reinforcement learning. Through the continuous interaction between the agent and the environment, the system is given spatial agility, which can fully utilize the high directional resolution of large-scale Multiple-Input Multiple-Output (MIMO) beamforming and the sparse characteristics of the millimeter wave channel, and realize autonomous decision-making against interference. The system adjusts the transmit and receive beam direction by dynamically selecting a row / column vector of the Discrete Fourier Transform matrix, effectively dealing with various unknown interference strategies including static, scanning and random, and realizing the communication capacity improvement of the MIMO system, thereby ensuring the reliability of the millimeter wave communication link in a complex counter-environment. However, this scheme can only select a single antenna element for signal transmission at each time, and this working mechanism causes most of the antenna elements in the uniform circular array to be in an idle state, which not only wastes array aperture resources, but also fundamentally limits the improvement of the system communication capacity. In addition, this method is mainly designed for sparse multipath channels. The MIMO channel in this scenario is usually in a non-full rank state, which seriously affects the effective channel dimension that can be utilized by the beam selection strategy, thereby directly affecting the final communication quality and reliability. Line-of-sight MIMO communication based on Discrete Fourier Transform (DFT) as an orthogonal basis provides a new way for anti-interference communication, as orthogonal mode signals can realize interference-free parallel transmission. Related documents propose a line-of-sight MIMO sensing integrated anti-interference method, which realizes interference source positioning through an enhanced multiple signal classification algorithm, and optimizes the suppression of various types of interference by combining transmit and receive beamforming and power allocation. However, in the actual scenario of unknown interference channel state information, the cooperative optimization of power adaptive allocation and anti-interference performance is still a key problem.
[0004] Currently, in a complex wideband interference environment, especially when the interference Channel State Information (CSI) is completely unknown, existing technologies cannot accurately adapt to the interference characteristics, and the orthogonality of the Discrete Fourier Transform orthogonal mode and the reinforcement learning ability of artificial intelligence are not deeply integrated. There is a lack of intelligent estimation algorithm for unknown interference channel state information, and the optimization of resource allocation and interference suppression is not realized with the help of artificial intelligence, which limits the reliability and capacity improvement of the wireless communication system. How to use the adaptive learning and reasoning ability of AI technology to accurately perceive the interference characteristics, intelligently optimize the power and beam in the unknown interference channel state information scenario, maximize the anti-interference gain of modal orthogonality, and cooperatively improve the reliability and capacity of the communication system, has become a core technical problem that needs to be solved. SUMMARY
[0005] In order to overcome the shortcomings of the existing technology, the present invention aims to provide a method and related equipment for synergistic improvement of reliability and capacity of line-of-sight multi-antenna transmission, so as to solve the technical problem of improving the reliability and capacity of wireless communication transmission in complex electromagnetic interference environments.
[0006] This invention is achieved through the following technical solution: In a first aspect, the present invention provides a method for synergistically improving the reliability and capacity of line-of-sight multi-antenna systems, comprising: The useful signal is modulated and orthogonally mapped to obtain the transmitted signal in orthogonal mode; The transmitted signal is transmitted after beamforming, and then received by the receiving end and preliminarily separated to obtain the deorthogonal mode received signal. A non-convex optimization problem with multiple coupled constraints is constructed, with the objective of minimizing the system bit error rate. Based on the reinforcement learning framework and combined with the non-convex optimization problem, the optimal transmit beamforming matrix and power allocation strategy are learned. The receive beamforming matrix is obtained by solving the optimal transmit beamforming matrix and power allocation strategy. Interference suppression and signal recovery are performed based on the receive beamforming matrix. The system monitors changes in the interference channel environment in real time and captures fluctuations in interference intensity. It dynamically iterates and optimizes the transmitting end strategy and receiving end processing parameters, and continuously improves the system's communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters.
[0007] The preferred orthogonal mapping process is as follows: use N × N Using the 3D inverse discrete Fourier transform matrix as an orthogonal basis, L The modulated useful signal is placed in the orthogonal base. L On orthogonal modes, the remaining ( N-L Set the signals of each mode to zero to generate N The transmitted signal is orthogonal to the mode domain of the road, and the expression of the transmitted signal is as follows:
[0008] Where x is N Dimensional transmission signal, N The total number of orthogonal modules supported by the system, and L≤N P is the transmitted power matrix; s is the power normalization matrix. N 3D modulation of useful signals; Let F be the conjugate transpose of the matrix, and let F be the discrete Fourier transform matrix.
[0009] Preferably, the transmitted signal is processed by beamforming and then transmitted. The receiving end then performs preliminary separation processing to obtain the deorthogonal mode received signal. The specific process is as follows: S1, the preset transmission beamforming matrix W T Loaded onto the transmitted signal, to obtain T x is transmitted through a uniform circular array antenna at the transmitting end; S2, the receiving end receives the transmitted signal through a uniform circular array antenna, where the expression for the received signal is as follows:
[0010] in, This is the channel matrix from the jammer to the legitimate receiver; The jamming signal transmitted by the jammer has a transmission power diagonal matrix of P. J ; With a mean of 0 and a variance of σ 2 Additive white Gaussian noise; F is the discrete Fourier transform matrix; H is the channel matrix; W T The beamforming matrix at the transmitting end; S3. Perform a Discrete Fourier Transform on the received signal to obtain the deorthogonal modulus. N The path receives the signal, wherein the original transmitted signal is recovered from the received signal with maximum suppression of interference. N The received signal obtained by loading the orthogonal mode signal onto the receiving beamforming matrix can be expressed as follows:
[0011] In the formula, W is an estimate of the received signal. R For the receiver beamforming matrix; Let F be the conjugate transpose of F; For power normalization N Dimensional modulation of useful signals.
[0012] Preferably, the specific expression for the non-convex optimization problem with multiple coupled constraints and the objective of minimizing the system bit error rate is as follows:
[0013] in, For the selected orthogonal mode l Allocated transmission power; The complete set of orthogonal modules supported by the system; For the selected set of orthogonal modules; For the first matrix n List; It is a modulo 2 norm; Orthogonal mode of the receiving end l The corresponding bit error rate; This represents the average bit error rate of a line-of-sight MIMO system.
[0014] The preferred design process for a reinforcement learning framework is as follows: The transmitter is considered the agent, and the state space of the reinforcement learning includes the interference channel matrix. , each orthogonal mode l Corresponding received signal-to-interference-plus-noise ratio and jammer transmit power diagonal matrix P J The action space of the intelligent agent includes the transmission beamforming matrix W. T and transmission power allocation ; Design a composite reward function with the bit error rate as the metric for a line-of-sight MIMO system; Construct a Q-value table and set the initial exploration rate ε and learning rate; The reinforcement learning framework uses the Q-learning algorithm to learn the optimal transmit beamforming matrix and power allocation strategy, and solves the receive beamforming matrix based on the optimal transmit beamforming matrix and power allocation strategy. The Q-learning algorithm incorporates an ε-greedy strategy and a dynamic learning rate adjustment mechanism.
[0015] Furthermore, in the optimal transmit beamforming matrix and power allocation strategy, the transmitter performs orthogonal mode selection and anti-interference strategy updates. The specific process includes: based on the estimated interference channel matrix... jammer transmit power diagonal matrix P J and the natural orthogonality of orthogonal modules, from the complete set of orthogonal modules Select the least affected by interference L A set of orthogonal modules serves as the transmission carrier; the agent obtains environmental feedback through a reward function, balances exploration and utilization using an ε-greedy policy, optimizes the convergence speed through a dynamic learning rate, and iteratively learns to obtain W. T With transmit power allocation The optimal combination is used to form an anti-interference strategy adapted to the current interference scenario.
[0016] Furthermore, the receive beamforming matrix is obtained by solving for the optimal transmit beamforming matrix and power allocation strategy. The specific process includes: The receiver acquires key parameters output by the transmitter in real time, including the orthogonal mode selection result. The transmit beamforming matrix W after self-learning T Power distribution The scheme, and the system's predicted interference channel state information. and the interference power P transmitted J ; Based on key parameters and using the minimum mean square error criterion, the receiving beamforming matrix is adaptively solved through matrix operations. W R ; Among them, the beamforming matrix W R The expression is as follows:
[0017] In the formula, , for N × N An identity matrix of dimensionality; Weighting matrix for the transmitter The conjugate transpose of; This is the conjugate transpose of the channel matrix H; This is the joint covariance matrix of interference and noise; Using matrices The received signal is processed to separate the useful signal from the interference signal, demodulate the useful signal, and suppress deliberate interference to the maximum extent.
[0018] Furthermore, the system continuously improves communication reliability and capacity by monitoring changes in the interference channel environment in real time, dynamically iterating and optimizing the transmitting end strategy and receiving end processing parameters, and continuously improving the system communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters. The specific process is as follows: The transmitter's reinforcement learning module continuously monitors the dynamic changes in the interference channel environment, captures fluctuations in interference intensity in real time, and dynamically updates the feedback value of the composite reward function based on environmental changes and communication quality data fed back from the receiver. This drives the Q-learning algorithm to continuously iterate and optimize the transmit beamforming matrix and power allocation strategy. The exploration probability and dynamic learning rate of the ε-greedy strategy are adaptively adjusted with environmental changes to ensure that communication reliability and communication capacity are continuously and synergistically improved in interference scenarios.
[0019] Secondly, the present invention also provides a system for collaboratively improving the reliability and capacity of line-of-sight multi-antenna systems, comprising: The signal processing module is used to modulate and orthogonally map the useful signal to obtain the transmitted signal in orthogonal mode; The beam transceiver module is used to transmit the signal after beamforming based on the transmitted signal, and the receiving end performs preliminary separation processing after receiving the signal to obtain the deorthogonal mode received signal. The modeling module is optimized to construct a non-convex optimization problem with multiple coupled constraints, aiming to minimize the system bit error rate. The anti-interference learning module is used to learn the optimal transmit beamforming matrix and power allocation strategy based on the reinforcement learning framework and the non-convex optimization problem, solve for the receive beamforming matrix based on the optimal transmit beamforming matrix and power allocation strategy, and perform interference suppression and signal recovery based on the receive beamforming matrix. The iterative optimization module is used to monitor changes in the interference channel environment and capture fluctuations in interference intensity in real time. It dynamically iteratively optimizes the transmitting end strategy and the receiving end processing parameters, and continuously improves the system's communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters.
[0020] Thirdly, the present invention also provides a mobile terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the line-of-sight multi-antenna reliability and capacity co-enhancement method as described above.
[0021] Compared with the prior art, the present invention has the following beneficial technical effects: This invention provides a method for synergistically improving the reliability and capacity of line-of-sight multi-antenna systems. By modulating and orthogonally mapping the useful signal, the useful signal is distributed across orthogonal modes to form the transmitted signal. The orthogonality of modes reduces crosstalk between signals at its source, providing a foundation for parallel transmission of multiple signals and effectively improving system communication capacity. Combining directional transmission with transmit beamforming and deorthogonal processing at the receiver focuses the transmission direction of the useful signal and initially separates the signal, reducing the probability of interference signals capturing the useful signal and improving transmission reliability. A non-convex optimization problem is constructed with the goal of minimizing the bit error rate while considering power and beamforming constraints, providing clear guidance for anti-interference strategy formulation and ensuring that the system prioritizes transmission quality while meeting resource constraints. Based on a reinforcement learning framework and using key state information such as the interference channel matrix and signal-to-interference-plus-noise ratio (SINR), the optimal transmit beamforming matrix and power allocation strategy are dynamically learned. Simultaneously, the receive beamforming matrix is solved based on the minimum mean square error (MMSE) criterion, effectively suppressing interference signals and efficiently recovering useful signals, achieving synergistic optimization of reliability and capacity under interference conditions. By monitoring the dynamic changes of the interference channel in real time and iteratively optimizing the transmit and receive parameters, the system can adapt to fluctuations in interference intensity and continuously adapt to complex electromagnetic environments, ensuring long-term stable improvement in communication reliability and capacity. This solves the problem of insufficient adaptability of traditional methods in dynamic interference scenarios. Attached Figure Description
[0022] Figure 1 This is a flowchart of the method for collaboratively improving the reliability and capacity of line-of-sight multi-antenna in an embodiment of the present invention; Figure 2A comparison of the average bit error rate between the anti-interference scheme proposed in this invention and existing mode hopping methods under different transmission signal-to-noise ratios; Figure 3 This is a schematic diagram of the system for collaboratively improving the reliability and capacity of line-of-sight multi-antenna in an embodiment of the present invention; In the diagram: 1. Signal processing module; 2. Beam transceiver module; 3. Optimization modeling module; 4. Anti-interference learning module; 5. Iterative optimization module. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0024] In complex broadband interference scenarios, when the state information of the interference channel is unknown, traditional anti-interference technologies often struggle to simultaneously meet the dual requirements of anti-interference performance and communication capacity. Therefore, under the premise of unknown interference channel state information, how to fully exploit the orthogonal characteristics between orthogonal modes to simultaneously improve the anti-interference capability and communication capacity of wireless communication systems has become a key technical bottleneck that urgently needs to be overcome. The purpose of this invention is to provide a method and related equipment for synergistically improving the reliability and capacity of line-of-sight multi-antenna transmission, in order to solve the technical problem of improving the reliability and capacity of wireless communication transmission in complex electromagnetic interference environments.
[0025] The present invention will now be described in further detail with reference to the accompanying drawings: Example 1 See Figure 1 In one embodiment of the present invention, a method for synergistically improving the reliability and capacity of line-of-sight multi-antenna systems is provided, comprising: Step 1: Modulate and orthogonally map the useful signal to obtain the transmitted signal in orthogonal mode; Specifically, the orthogonal mapping process includes, at the sending end, […]. L Useful signals are mapped to N On orthogonal dimensions. Specifically, utilizing N × N Using the 3D inverse discrete Fourier transform matrix as an orthogonal basis, L The road signal is placed in its L On orthogonal modes, the remaining ( N - L The signals of each mode are set to zero, thereby generating N Transmitted signals orthogonal in the modal domain , where x is N 3D transmitted signal, P is the transmitted power matrix, s is N Dimensional modulation of useful signals, Let F be the conjugate transpose of the matrix, and let F be the discrete Fourier transform matrix.
[0026] Step 2: Based on the transmitted signal, the transmitted signal is processed by beamforming and then transmitted. The received signal is then processed by the receiving end to obtain the deorthogonal mode received signal. Specifically, to effectively reduce the impact of interference on the useful signal, transmit beamforming needs to be performed at both the transmitting and receiving ends. T With receiving beamforming W R Preset. Load the preset transmit beamforming matrix into... N On the orthogonal signal of the road, we get T x, which is then transmitted via a Uniform Circular Array (UCA) antenna.
[0027] Receiver signal separation and processing: Transmitted signal T signal x is transmitted via signal H and received by the receiving UCA antenna. The signal y received at the receiving end can be expressed as:
[0028] in, This is the channel matrix from the jammer to the legitimate receiver; The jamming signal transmitted by the jammer has a transmission power diagonal matrix P. J ; The mean of the data received at the receiving end is 0, and the variance is... Additive white Gaussian noise. The received signal y is first subjected to a discrete Fourier transform to obtain the deorthogonal modes. N The signal is transmitted through a path. To maximize interference suppression and recover the original transmitted signal from the received signal, the obtained signal is... N The orthogonal mode signal is applied to the receiving beamforming matrix. The resulting received signal can be expressed as:
[0029] In the formula, This is an estimate of the received signal; For the receiver beamforming matrix; Let F be the conjugate transpose of F; For power normalization N Dimensional modulation of useful signals.
[0030] Step 3: Construct a non-convex optimization problem with multiple coupled constraints, aiming to minimize the system bit error rate; Specifically, to meet the communication requirements of line-of-sight MIMO systems, a method is constructed with the goal of minimizing the system bit error rate and the total transmit power. P T The selection of orthogonal modes and the constraint that the transmit beamforming matrix satisfies the constant modulus of the columns (i.e., the modulus 2 norm of each column vector in the matrix is strictly equal to 1) form a multi-constraint coupled nonconvex optimization problem. The transmit power expression is as follows:
[0031] in, For the selected orthogonal mode l Allocated transmission power; The complete set of orthogonal modules supported by the system; For the selected set of orthogonal modules; For the first matrix n List; It is a modulo 2 norm; Orthogonal mode of the receiving end l The corresponding bit error rate; This represents the average bit error rate of a line-of-sight MIMO system.
[0032] This embodiment employs different signal modulation methods such as Binary Phase Shift Keying (BPSK), Quadrature Phase Shift Keying (QPSK), and Quadrature Amplitude Modulation (QAM). There are different expressions, and the modulation method of the signal is not limited.
[0033] Step 4: Based on the reinforcement learning framework and combined with the non-convex optimization problem, the optimal transmit beamforming matrix and power allocation strategy are learned. The receive beamforming matrix is obtained by solving the optimal transmit beamforming matrix and power allocation strategy. Interference suppression and signal recovery are performed based on the receive beamforming matrix. Specifically, the reinforcement learning framework design includes setting the transmitter as an agent and setting the interference channel matrix. Orthogonal modules l The corresponding received signal-to-interference-plus-noise ratio (SINR) The jamming power diagonal matrix P transmitted by the jammer J As the state space for reinforcement learning, the beamforming matrix W will be sent. TTransmit power allocation This serves as the action space for the agent. A composite reward function, quantified by the bit error rate of the line-of-sight MIMO system, is designed to provide accurate feedback for the agent's policy learning. The reinforcement learning module employs the Q-learning algorithm and incorporates an ε-greedy policy and a dynamic learning rate adjustment mechanism.
[0034] The initialization of reinforcement learning parameters includes: constructing a Q-value table and setting the initial exploration rate ε and learning rate.
[0035] Specifically, the transmitter's orthogonal mode selection and anti-interference strategy update: The transmitter updates the anti-interference strategy based on the estimated interference channel matrix H. J Combining the natural orthogonality between orthogonal modules, the system can utilize the complete set of orthogonal modules. Select the one least affected by the current disturbance. L A number of orthogonal modes are determined as the carriers for signal transmission to achieve... L Parallel multiplexing and transmission of the channel signal. During this process, the agent obtains environmental feedback through a composite reward function and continuously interacts with the interference channel environment. By dynamically balancing the exploration of unknown anti-interference strategies with the utilization of known optimal strategies using an ε-greedy strategy, the convergence speed of the optimization algorithm is improved through dynamic learning rate adjustment, and the agent autonomously iteratively learns the design of the transmit beamforming matrix W. T With transmit power allocation The optimal combination is used to form an anti-interference strategy adapted to the current interference scenario.
[0036] Among them, the receiver beamforming matrix is solved: the receiver acquires key parameters output by the transmitter in real time, including the orthogonal mode selection results. The transmit beamforming matrix W after self-learning T Power distribution The scheme, and the system's predicted interference channel state information H J and the interference power P transmitted J Based on the above parameters, the mean square error minimization (MMSE) criterion is adopted, and the receiving beamforming matrix is adaptively solved through matrix operations. W R As shown below:
[0037] In the formula, , for N × N An identity matrix of dimensionality; Weighting matrix for the transmitter The conjugate transpose of; This is the conjugate transpose of the channel matrix H; This is the joint covariance matrix of interference and noise; Using matrices Processing the received signal can effectively separate the useful signal from the interference signal, achieve efficient demodulation of the useful signal, and maximize the suppression of deliberate interference.
[0038] Step 5: Monitor changes in the interference channel environment in real time, dynamically iterate and optimize the transmitting end strategy and receiving end processing parameters, and continuously improve the system communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters.
[0039] Specifically, during communication, the reinforcement learning module at the transmitting end continuously monitors the dynamic changes in the interference channel environment and captures fluctuations in interference intensity in real time. Based on environmental changes and communication quality data fed back from the receiving end, it dynamically updates the feedback value of the composite reward function, driving the Q-learning algorithm to continuously iterate and optimize the transmit beamforming matrix and power allocation strategy. Simultaneously, the exploration probability and dynamic learning rate of the ε-greedy strategy are adaptively adjusted with environmental changes, ensuring a continuous and synergistic improvement in communication reliability and capacity under interference scenarios.
[0040] The wireless communication system involved in this embodiment mainly consists of three parts: a transmitter, a receiver, and an interference source. The interference source can be configured with a single antenna or multiple antennas, attempting to block the legitimate communication link between the transmitter and receiver by densely transmitting interference signals. The interference source can generate various types of interference waveforms, specifically covering common forms such as single-tone interference, multi-tone interference, partial narrowband interference, broadband interference, and co-channel interference.
[0041] The sending end uses a method that includes N A uniform circular array UCA antenna with multiple elements generates multiple orthogonal beams to transmit effective and useful signals to the receiver; the receiver is also configured with... N Each array element of the UCA antenna is responsible for receiving useful signals from the transmitter and interference signals from interference sources. In a specific embodiment of the present invention, the UCA antennas at the transmitter and receiver are arranged in a non-parallel aligned manner. It should be noted that the types of transmitter and receiver antennas defined in this invention are not limited to UCA antennas; antenna structures with similar characteristics, such as spiral phase plate antennas and novel antennas based on metasurface materials, are also applicable to this technical solution.
[0042] The effectiveness of this embodiment can be further illustrated by the following simulation results: I. Simulation Conditions During the simulation, the carrier frequency was 2.4 GHz, the UCA radius of the transmitter was 0.5 m, the UCA radius of the receiver was 0.5 m, the transmit interference-to-noise ratio was 20 dB, and the transmit and receive UCAs were not parallel aligned.
[0043] II. Simulation Content and Results Under the above conditions, Figure 2 The average bit error rate (ABER) performance of this invention and existing methods (mode hopping, adaptive mode hopping, and this invention under interference channel estimation errors) at different signal-to-noise ratios (SNRs) is presented. It is evident that this invention exhibits the lowest ABER, demonstrating superior anti-interference capability. Traditional mode hopping and adaptive mode hopping have higher ABERs because they lack the optimized power allocation and interference suppression strategies of this invention. Even with interference channel estimation errors, this invention maintains competitive performance (though slightly inferior to ideal conditions). Overall, Figure 2 The superiority of this invention in reducing bit error rate under different signal-to-noise ratios has been verified.
[0044] In this embodiment, various forms of deliberate interference in wireless communication, such as single-tone, multi-tone, narrowband, and broadband, are addressed. Multiple parallel useful signals are transmitted after preprocessing by loading a Discrete Fourier Transform (DFT) matrix onto the transmitting antenna of a uniform circular array (UCA). The receiving antenna then receives and demodulates the signal by loading an Inverse Discrete Fourier Transform (IDFT) matrix.
[0045] Given the randomness of interference signals and the unknown nature of interference channel states, this invention deeply couples reinforcement learning mechanisms with the beam characteristics of orthogonal modes, constructing a non-convex optimization problem with multiple constraints coupled by using the average bit error rate (ABER) of the line-of-sight MIMO system as the objective function and the total transmitted power, orthogonal mode selection, and the constant modulus of the transmitted beamforming matrix (modulo-2 norm equal to 1) as the boundaries.
[0046] To address the challenge of optimization decision-making in scenarios with unknown interference channel indices (CSI), a composite reward function is designed with the bit error rate of the line-of-sight MIMO system as the objective. This function guides the transmitter, acting as an agent, to autonomously learn the optimal combination of transmit beamforming matrix and transmit power allocation in an environment where the state information of the interference channel is unknown. Specifically, based on the estimated interference CSI and the orthogonality between orthogonal modes, the transmitter selects the mode least affected by interference from the available orthogonal modes. L Using orthogonal modes as transmission carriers, to achieve LParallel multiplexing and transmission of the signal is implemented. At the receiver, based on the mode selection result, transmit beamforming matrix, power allocation scheme, and estimated interference channel state information output by the transmitter, the receive beamforming matrix is adaptively solved by minimizing the mean square error (MMSE) criterion, achieving accurate demodulation of the useful signal and effective suppression of interference signals. In the reinforcement learning mechanism design, the Q-learning algorithm is adopted, and an ε-greedy policy improvement and dynamic learning rate adjustment mechanism are introduced to effectively balance the contradiction between exploration and utilization, thereby achieving a synergistic improvement in communication reliability and capacity in complex electromagnetic environments.
[0047] In summary, traditional anti-jamming techniques (such as frequency hopping and fixed beamforming) heavily rely on pre-defined interference models or precise interference channel state information, resulting in a sharp decline in performance when facing unknown and dynamic intelligent interference. This invention introduces a reinforcement learning mechanism, leveraging the orthogonal mode characteristics of the UCA array, enabling the transmitter to act as an intelligent agent. Guided by the final communication quality (bit error rate), it autonomously learns the optimal anti-jamming strategy through real-time interaction with the environment. This method can effectively cope with various forms of dynamic and deliberate interference, including single-tone, multi-tone, narrowband, and broadband interference. Furthermore, this invention utilizes orthogonal mode selection (selecting the mode least affected by interference) to... L (modal) implementation L Parallel multiplexing of signals ensures transmission capacity; simultaneously, with bit error rate as the core objective, it combines reinforcement learning optimization and MMSE equalization to maximize anti-interference capability and demodulation accuracy, successfully breaking the trade-off between reliability and capacity and achieving synergistic improvement of both. This invention solves the problem of synergistic improvement of anti-interference and communication capacity in wireless communication systems with broadband interference under unknown interference information conditions, a challenge posed by traditional methods.
[0048] Existing MIMO anti-jamming technologies primarily exploit the spatial domain degree of freedom. This invention, through reinforcement learning, deeply couples two dimensions: orthogonal modes and UCA beamforming. The agent can not only dynamically adjust the beam in the spatial domain but also proactively and intelligently select the least interfered subset from all available orthogonal modes for signal transmission. This dual mechanism, compared to existing technologies that utilize only a single dimension, brings greater degrees of freedom and more flexible interference avoidance capabilities, thereby achieving a synergistic improvement in communication reliability and capacity at the same power level.
[0049] The optimization problem constructed in this invention, with bit error rate as the objective and constraints such as column constant modulus, is a complex non-convex problem that is difficult to solve directly and efficiently using traditional convex optimization methods. This invention innovatively uses reinforcement learning as an efficient approximate solution tool, transforming this mathematical problem into a learning process for an intelligent agent. By guiding the agent to directly search for a practical solution that approximates the global optimum in the policy space, it cleverly bypasses the obstacles of traditional mathematical solutions, providing a novel solution to this type of engineering problem.
[0050] Example 2 according to Figure 3 As shown, this embodiment also provides a system for collaboratively improving the reliability and capacity of line-of-sight multi-antenna systems, including: Signal processing module 1 is used to modulate and orthogonally map the useful signal to obtain the transmitted signal in orthogonal mode; The beam transceiver module 2 is used to transmit the signal after beamforming based on the transmitted signal, and the receiver performs preliminary separation processing on the received signal to obtain the deorthogonal mode received signal. Optimize modeling module 3 to construct a non-convex optimization problem with multiple coupled constraints, aiming to minimize the system bit error rate. The anti-interference learning module 4 is used to learn the optimal transmit beamforming matrix and power allocation strategy based on the reinforcement learning framework and the non-convex optimization problem, solve the receive beamforming matrix based on the optimal transmit beamforming matrix and power allocation strategy, and perform interference suppression and signal recovery based on the receive beamforming matrix. The iterative optimization module 5 is used to monitor changes in the interference channel environment and capture fluctuations in interference intensity in real time, dynamically iteratively optimize the transmitting end strategy and receiving end processing parameters, and continuously improve the system communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters.
[0051] Example 3 This embodiment also provides a mobile terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, such as a line-of-sight multi-antenna reliability and capacity co-enhancement program.
[0052] When the processor executes the computer program, it implements the steps of the above-described method for coordinating the improvement of reliability and capacity of line-of-sight multi-antenna systems, for example: The useful signal is modulated and orthogonally mapped to obtain the transmitted signal in orthogonal mode; The transmitted signal is transmitted after beamforming, and then received by the receiving end and preliminarily separated to obtain the deorthogonal mode received signal. A non-convex optimization problem with multiple coupled constraints is constructed, with the objective of minimizing the system bit error rate. Based on the reinforcement learning framework and combined with the non-convex optimization problem, the optimal transmit beamforming matrix and power allocation strategy are learned. The receive beamforming matrix is obtained by solving the optimal transmit beamforming matrix and power allocation strategy. Interference suppression and signal recovery are performed based on the receive beamforming matrix. The system monitors changes in the interference channel environment in real time and captures fluctuations in interference intensity. It dynamically iterates and optimizes the transmitting end strategy and receiving end processing parameters, and continuously improves the system's communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters.
[0053] Alternatively, when the processor executes the computer program, it implements the functions of each module in the above system, for example: Signal processing module 1 is used to modulate and orthogonally map the useful signal to obtain the transmitted signal in orthogonal mode; The beam transceiver module 2 is used to transmit the signal after beamforming based on the transmitted signal, and the receiver performs preliminary separation processing on the received signal to obtain the deorthogonal mode received signal. Optimize modeling module 3 to construct a non-convex optimization problem with multiple coupled constraints, aiming to minimize the system bit error rate. The anti-interference learning module 4 is used to learn the optimal transmit beamforming matrix and power allocation strategy based on the reinforcement learning framework and the non-convex optimization problem, solve the receive beamforming matrix based on the optimal transmit beamforming matrix and power allocation strategy, and perform interference suppression and signal recovery based on the receive beamforming matrix. The iterative optimization module 5 is used to monitor changes in the interference channel environment and capture fluctuations in interference intensity in real time, dynamically iteratively optimize the transmitting end strategy and receiving end processing parameters, and continuously improve the system communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters.
[0054] For example, the computer program can be divided into a signal processing module 1, a beam transceiver module 2, an optimization modeling module 3, an anti-interference learning module 4, and an iterative optimization module 5; The specific functions of each module are as follows: Signal processing module 1 is used to modulate and orthogonally map the useful signal to obtain the transmitted signal in orthogonal mode; The beam transceiver module 2 is used to transmit the signal after beamforming based on the transmitted signal, and the receiver performs preliminary separation processing on the received signal to obtain the deorthogonal mode received signal. Optimize modeling module 3 to construct a non-convex optimization problem with multiple coupled constraints, aiming to minimize the system bit error rate. The anti-interference learning module 4 is used to learn the optimal transmit beamforming matrix and power allocation strategy based on the reinforcement learning framework and the non-convex optimization problem, solve the receive beamforming matrix based on the optimal transmit beamforming matrix and power allocation strategy, and perform interference suppression and signal recovery based on the receive beamforming matrix. The iterative optimization module 5 is used to monitor changes in the interference channel environment and capture fluctuations in interference intensity in real time, dynamically iteratively optimize the transmitting end strategy and receiving end processing parameters, and continuously improve the system communication reliability and capacity based on the optimized transmitting end strategy and receiving end processing parameters.
[0055] The mobile terminal can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The mobile terminal may include, but is not limited to, a processor and memory.
[0056] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the mobile terminal, connecting various parts of the mobile terminal via various interfaces and lines.
[0057] The memory can be used to store the computer program and / or module. The processor implements various functions of the mobile terminal by running or executing the computer program and / or module stored in the memory and calling the data stored in the memory.
[0058] The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as sound playback, image playback, etc.); the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, SmartMediaCards (SMC), Secure Digital (SD) cards, FlashCards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for line-of-sight multi-antenna reliability and capacity coimprovement, characterized in that, The method comprises the following steps: modulating and quadrature mapping a useful signal to obtain a transmission signal in a quadrature mode; performing transmission beamforming processing based on the transmission signal, transmitting, and performing preliminary separation processing on a received signal to obtain a received signal in an inverse quadrature mode; constructing a non-convex optimization problem with multiple constraint conditions coupled to minimize the system bit error rate; learning an optimal transmission beamforming matrix and power allocation strategy based on the non-convex optimization problem in a reinforcement learning framework, solving a receiving beamforming matrix according to the optimal transmission beamforming matrix and the power allocation strategy, and performing interference suppression and signal recovery according to the receiving beamforming matrix; real-time monitoring of changes in the interference channel environment and capturing fluctuations in the interference intensity, dynamic iterative optimization of the sending end strategy and the receiving end processing parameters, and continuous improvement of the system communication reliability and capacity according to the optimized sending end strategy and the receiving end processing parameters. 2.The method of claim 1, wherein, The specific process of the quadrature mapping processing is as follows: use N × N Using the 3D inverse discrete Fourier transform matrix as an orthogonal basis, L The modulated useful signal is placed in the orthogonal base. L On orthogonal modes, the remaining ( N-L Set the signals of each mode to zero to generate N The transmitted signal is orthogonal to the mode domain of the road, and the expression of the transmitted signal is as follows: wherein x is N transmitting a signal, N is the total number of orthogonal modes supported by the system, and L≤N ; P is a transmit power matrix; s is a power normalized N dimensional modulated useful signal; is the conjugate transpose of the matrix, and F is a discrete Fourier transform matrix. 3.The method of claim 1, wherein, The specific process of the transmission beamforming processing based on the transmission signal and the preliminary separation processing on the received signal to obtain the received signal in the inverse quadrature mode is as follows: S1, load the preset transmit beamforming matrix W T onto the transmit signal, obtaining T x, transmitted by a uniform circular array antenna at the transmitting end; S2, the receiving end receives the transmission signal through a uniform circular array antenna, and the expression of the received signal is as follows: wherein, is a channel matrix from the jammer to the legitimate receiver; is a diagonal matrix of the transmission power of the interference signal transmitted by the jammer J ; is an additive white Gaussian noise with mean 0 and variance σ 2 ; F is a discrete Fourier transform matrix; H is a channel matrix; W T is a transmit beamforming matrix; S3, performing a discrete Fourier transform on the received signal to obtain a de-orthogonalized mode N of the received signal, wherein the obtained N of the received signal can be expressed as follows: wherein is an estimate of the received signal; W R is a receive beamforming matrix; is a conjugate transpose matrix of F; is a power-normalized N dimension modulation useful signal.
4. The method of claim 1, wherein, The specific expression of the non-convex optimization problem constructed to minimize the system bit error rate as the target and coupled with multiple constraint conditions is as follows: wherein, is the selected orthogonal mode l allocated transmit power; is the full set of orthogonal modes supported by the system; is the selected set of orthogonal modes; is the jthcolumn of matrix n is the jthcolumn of matrix is the modulo two norm; is the bit error rate corresponding to the received orthogonal mode l is the bit error rate corresponding to the received orthogonal mode is the average bit error rate for the line-of-sight MIMO system.
5. The method of claim 1, wherein, The specific process of designing the reinforcement learning framework is as follows: The sending end is an intelligent agent, wherein a state space of the reinforcement learning comprises an interference channel matrix , each orthogonal mode l , a corresponding received signal-to-interference-and-noise ratio , and an interference transmitter power diagonal matrix P J ; an action space of the intelligent agent comprises a sending beamforming matrix W T and a sending power allocation designing a composite reward function with the line-of-sight MIMO system bit error rate as the index; constructing a Q value table, setting an initial exploration rate ε and a learning rate; The reinforcement learning framework adopts a Q-learning algorithm to learn the optimal transmission beamforming matrix and the power allocation strategy, and solves the receiving beamforming matrix according to the optimal transmission beamforming matrix and the power allocation strategy, wherein the Q-learning algorithm is embedded with an ε-greedy strategy and a dynamic learning rate adjustment mechanism.
6. The method of claim 5, wherein, In the optimal transmit beamforming matrix and power allocation strategy, the transmitter performs orthogonal mode selection and anti-interference strategy updates. The specific process includes: based on the estimated interference channel matrix... jammer transmit power diagonal matrix P J and the natural orthogonality of orthogonal modules, from the complete set of orthogonal modules Select the least affected by interference L A set of orthogonal modules serves as the transmission carrier; the agent obtains environmental feedback through a reward function, balances exploration and utilization using an ε-greedy policy, optimizes the convergence speed through a dynamic learning rate, and iteratively learns to obtain W. T With transmit power allocation The optimal combination is used to form an anti-interference strategy adapted to the current interference scenario.
7. The method of claim 6, wherein, The specific process of solving the receiving beamforming matrix according to the optimal transmission beamforming matrix and the power allocation strategy includes: The receiving end acquires the key parameters output by the sending end in real time, wherein the key parameters include the orthogonal mode selection result , the sending beamforming matrix W learned autonomously T , the power allocation scheme, and the system-estimated interference channel state information and the interference power P transmitted thereby J ; Based on the key parameters with minimum mean square error criterion, through matrix operation adaptive solution receive beam forming matrix W R ; where the receive beamforming matrix W R The expression is as follows: wherein , is N x N a unit matrix of dimension is the conjugate transpose of the transmit-end weighting matrix is the conjugate transpose of the channel matrix H; is the conjugate transpose of the channel matrix H; is the joint covariance matrix of the interference and noise; Utilizing matrices The received signal is processed to separate the desired signal from the interfering signal, and the desired signal is demodulated while maximizing the suppression of the intentional jammer. 8.The method of claim 5, wherein, The specific process of real-time monitoring of changes in the interference channel environment, dynamic iterative optimization of the sending end strategy and the receiving end processing parameters, and continuous improvement of the system communication reliability and capacity according to the optimized sending end strategy and the receiving end processing parameters is as follows: The sending end reinforcement learning module continuously monitors the dynamic changes of the interference channel environment, real-time captures the fluctuations of the interference intensity, dynamically updates the feedback value of the composite reward function according to the environmental changes and the communication quality data fed back by the receiving end, and drives the Q-learning algorithm to continuously iterate and optimize the transmission beamforming matrix and the power allocation strategy; wherein the exploration probability of the ε-greedy strategy and the dynamic learning rate are adaptively adjusted with the environmental changes, which is used to ensure that the communication reliability and the communication capacity are continuously improved in the interference scenario.
9. A line-of-sight multi-antenna reliability and capacity coimprovement system, characterized by, The method comprises the following steps: a signal processing module for modulating and quadrature mapping a useful signal to obtain a transmission signal in a quadrature mode; A beam receiving module is configured to transmit a signal after beamforming processing based on the transmitting signal, and to receive a signal after preliminary separation processing by a receiving end to obtain a de-orthogonal module receiving signal; An optimization modeling module is configured to construct a non-convex optimization problem with multiple constraint conditions coupled to minimize the system bit error rate as a target; An anti-interference learning module is configured to learn an optimal transmitting beamforming matrix and power allocation strategy based on a reinforcement learning framework and the non-convex optimization problem, to solve a receiving beamforming matrix according to the optimal transmitting beamforming matrix and power allocation strategy, and to perform interference suppression and signal recovery according to the receiving beamforming matrix; An iterative optimization module is configured to monitor changes in an interference channel environment and capture fluctuations in interference strength in real time, to dynamically and iteratively optimize transmitting end strategies and receiving end processing parameters, and to continuously improve system communication reliability and capacity according to the optimized transmitting end strategies and receiving end processing parameters.
10. A mobile terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the line-of-sight multi-antenna reliability and capacity collaborative improvement method of any one of claims 1-8.