Array antenna signal transmission method and device under crosstalk influence

By establishing a crosstalk matrix model in a movable array antenna system and using the TD3 algorithm to optimize precoding and antenna position, the crosstalk problem of array antennas in dynamic environments is solved, communication and sensing quality is improved, and the application value of array antennas is expanded.

CN122512968APending Publication Date: 2026-08-04UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2026-05-09
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing array antenna systems lack adaptability to crosstalk problems in dynamic environments, resulting in signal distortion and a decline in user service quality. In particular, it is difficult to effectively suppress the effects of crosstalk in scenarios with movable antennas.

Method used

By establishing a crosstalk matrix model for a mobile array scenario, and combining the Cramer-Rao bound of angle estimation and user signal-to-noise-interference ratio constraints, a Markov decision process and a deep reinforcement learning algorithm (TD3) are used to optimize the precoding matrix and antenna position, thereby achieving dynamic compensation for crosstalk.

Benefits of technology

Under the constraints of total transmit power and antenna position, the angle estimation error of the sensing link is reduced, the communication and sensing performance is improved, and the application value of array antennas in wireless communication systems is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122512968A_ABST
    Figure CN122512968A_ABST
Patent Text Reader

Abstract

The application discloses a kind of under the influence of crosstalk array antenna signal transmitting method and device, belong to communication technical field.The application realizes a kind of deep reinforcement learning optimization method based on TD3, under the constraint of total transmitting power, communication signal-to-noise ratio and antenna position, the mean square error of angle estimation in sensing link is reduced.Compared with prior art, the application considers actual crosstalk model, integrates anti-crosstalk function with transmitting precoding and mobile antenna technology, and optimizes by designing deep learning algorithm, reliably reduces the influence of antenna crosstalk on array antenna supported communication-sensing link, realizes the consideration of communication and sensing performance, expands the application value of array antenna in future wireless communication system, and provides a feasible engineering solution for the crosstalk problem of multi-antenna system in actual deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, specifically relating to a method and apparatus for transmitting array antenna signals under the influence of crosstalk. Background Technology

[0002] With the rapid development of wireless communication technology, array antennas and multiple-input multiple-output (MIMO) have become important supporting means for fifth-generation (5G) and future sixth-generation (6G) mobile communication systems. By providing multiple degrees of freedom in the spatial domain, array antennas can achieve high-capacity transmission, fine beamforming, and multi-user parallel services.

[0003] However, in practical engineering implementations, crosstalk is unavoidable between array antennas. This phenomenon directly affects the independence of the transmitted signal and the spatial resolution of the array, becoming a significant bottleneck restricting system performance improvement. Antenna crosstalk originates from the limited isolation between antenna elements. When the array spacing is insufficient, the feed network has parasitic effects, or the RF link design is imperfect, the transmitted signal from one antenna port will couple to other ports through spatial or circuit paths, causing the actual signal received at the receiver to deviate from the ideal design. Crosstalk not only introduces additional interference noise but also alters the equivalent channel characteristics, preventing the theoretical gains of precoding and beamforming from being fully realized. In large-scale MIMO systems, due to the large number of antennas and reduced spacing, the crosstalk effect is more significant, and its impact on system capacity, bit error rate, and energy efficiency cannot be ignored.

[0004] In traditional fixed-position antenna (FPA) systems, academia and industry have conducted modeling and compensation studies on antenna crosstalk. For example, by establishing a crosstalk matrix to characterize the coupling coefficients between different antenna ports, pre-compensation can be performed at the signal processing level. However, most research remains focused on static array scenarios, with limited adaptability to dynamic environments and flexible architectures. Especially in multi-antenna transmitters lacking isolators in the RF front-end, crosstalk interacts with the nonlinear effects of the power amplifier, further deteriorating signal quality. Therefore, effectively suppressing array crosstalk and ensuring signal beam directionality and user quality of service is one of the key challenges in the design of array signal transmission methods.

[0005] Movable antennas (MAs) enhance channel controllability and beamforming flexibility by introducing positional freedom. When the position of antenna elements changes, the crosstalk coefficient distribution also dynamically changes, making system modeling and optimization more complex. Without effective crosstalk mitigation mechanisms, even with advanced beamforming or position optimization strategies, overall performance can still degrade significantly. Therefore, comprehensively considering and compensating for crosstalk effects in array antenna design and signal transmission methods has significant theoretical and engineering implications.

[0006] In summary, existing multi-antenna transmission systems still have shortcomings in dealing with array crosstalk: on the one hand, existing compensation methods mostly rely on the assumption of a fixed array, lacking adaptability to complex dynamic scenarios; on the other hand, traditional precoding and power control methods do not fully incorporate the effects of crosstalk, resulting in a gap between actual performance and theoretical expectations. Therefore, there is an urgent need to propose a novel signal transmission method and device that can effectively suppress array crosstalk and improve communication and sensing quality to meet the development needs of future wireless communication and related applications. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a method and apparatus for transmitting array antenna signals under crosstalk, aiming to solve the problems of transmitted signal distortion, reduced precoding gain, and decreased user service quality caused by electromagnetic coupling and signal crosstalk between antenna elements in existing array antenna transmission systems.

[0008] This invention provides a method for transmitting array antenna signals under crosstalk, comprising the following steps:

[0009] The crosstalk model of fixed array antennas is extended to the scenario of mobile array antennas. A crosstalk matrix related to the antenna position (i.e., the crosstalk matrix of mobile antennas) is established and embedded in the transmit-channel-sensing link model.

[0010] Using the angle-estimated Cramer-Rao boundary (CRB) as the sensing performance index and the user signal-to-noise ratio (SINR) as the communication performance constraint, combined with the total transmit power constraint, as well as the boundary constraints and minimum spacing constraints of the antenna position, a joint optimization problem is established.

[0011] The joint optimization problem is modeled as a Markov decision process (MDP), with the channel gain coefficient, azimuth angle and previous action forming the state space, and the precoding matrix and antenna position vector forming the action space. A reward function consisting of CRB maximization descent and SINR penalty is defined.

[0012] The action is trained and optimized based on the dual-delay deep deterministic policy gradient (TD3) algorithm, and the optimal precoding matrix and antenna position parameters are output.

[0013] Based on the obtained precoding matrix, the original communication signal (i.e., communication-sensing signal) generated by the signal generator is spatially weighted and then transmitted through the array antenna based on the obtained antenna position parameters.

[0014] Furthermore, the crosstalk matrix model is calculated using the antenna spacing and the linear phase-power law model.

[0015] Furthermore, the TD3 algorithm adds (Ornstein-Uhlenbeck, OU) noise before mapping the precoding matrix and antenna position vector to specific actions, rather than adding it after the actions are generated.

[0016] Furthermore, the total transmit power constraint is implemented using tanh and sigmoid activation functions; the boundary constraints and minimum spacing constraints of the antenna position are implemented using sigmoid and softmax activation functions; and the total transmit power constraint, the boundary constraints of the antenna position, and the minimum spacing constraints do not appear in the penalty term of the reward function.

[0017] Furthermore, the joint optimization problem is specifically as follows:

[0018]

[0019] in, For the precoding matrix, The antenna position vector, Target azimuth Estimated CRB The first The and the first Antenna positions of each antenna This refers to the minimum physical spacing between antenna elements. For the number of antennas, These are the minimum and maximum values ​​of the antenna position, respectively, used to characterize the boundary constraints of the antenna position. For the first The minimum SINR threshold required for each user For the number of users, For maximum total transmission power, Let H be the trace of the matrix, and let H be the conjugate transpose of the matrix.

[0020] The further reward function consists of an angle estimation performance metric and a communication quality penalty, where the communication quality penalty is calculated based on the difference between the user's SINR and a threshold.

[0021] Furthermore, the reward function is:

[0022]

[0023] in, Indicates time, for Time-effective array response vector Target azimuth The partial derivatives, for The precoding matrix at time step H, where the superscript H denotes the conjugate transpose of the matrix. This is a penalty item for communication quality.

[0024] The communication quality penalty item for:

[0025]

[0026] in, The pre-defined positive real number penalty factor. For the first SINR of each user under the current action For the first The minimum SINR threshold required for each user.

[0027] Furthermore, the state space and action space are respectively set as follows:

[0028]

[0029]

[0030] in, Representing time, state space For dimension A real vector, dimension Depends on the sum of channel parameters and angle information; complex channel gain vector For dimension Complex vectors, For the number of users, To sense the number of paths, the azimuth vector For dimension real vectors, This represents the action vector executed in the previous time step. Indicates taking the real part, Indicates taking the imaginary part; action space For dimension real vectors, For dimension A real vector, obtained by precoding a complex matrix. All elements are vectorized and their real and imaginary parts are separated and concatenated to obtain the result. For dimension A real vector, representing time The physical location decision vector of each antenna.

[0031] Furthermore, when training and optimizing actions based on the TD3 algorithm, the convergence conditions include: the number of training iterations reaches the preset maximum number of training iterations, the reward function converges, and the number of interaction steps in a single round reaches the preset maximum number of interaction steps.

[0032] Another aspect of the present invention provides an array antenna signal transmitting device under crosstalk influence, comprising:

[0033] The primary signal generator is used to generate the input signals required for communication and sensing.

[0034] The crosstalk model fitting module is used to establish the crosstalk matrix between array antennas in a mobile array scenario, and input the crosstalk matrix into the TD3 deep learning training network for reward function calculation;

[0035] A channel estimator is used to estimate channel information based on uplink pilot signals.

[0036] The TD3 deep learning training network is used to receive crosstalk matrix and channel information, and output optimized precoding matrix and antenna position parameters. The construction of the TD3 deep learning training network is as follows: using the Cramer-Rao bound of angle estimation as the perception performance index, the user signal-to-noise ratio (SNR) as the communication performance constraint, and combining the total transmit power constraint, as well as the boundary constraints and minimum spacing constraints of the antenna position, a joint optimization problem is established. The joint optimization problem is modeled as a Markov decision process, with the channel gain coefficient, azimuth angle, and previous action forming the state space, and the precoding matrix and antenna position vector forming the action space. A reward function is defined, consisting of maximizing the descent of the Cramer-Rao bound and the user SNR penalty term.

[0037] The transmit precoder is used to perform multi-task precoding on the input signal and is adjusted in combination with the precoding matrix output by the TD3 deep learning training network to obtain the corrected communication-sensing signal.

[0038] Antenna position motor, used to adjust the position of the array antennas according to the antenna position parameters output by the TD3 deep learning training network;

[0039] An array antenna is used to transmit calibrated communication-sensing signals, so that while ensuring the quality of user communication, the accuracy of target angle estimation in the sensing link is improved.

[0040] Furthermore, the crosstalk model fitting module calculates the crosstalk matrix model based on the antenna spacing and the linear phase-power law model, which is the crosstalk model.

[0041] The technical solution provided by this invention brings at least the following beneficial effects:

[0042] This invention implements a deep reinforcement learning optimization method based on TD3, which reduces the mean square error of angle estimation in the sensing link while ensuring total transmit power constraints, communication signal-to-interference-plus-noise ratio (SNR), and antenna position constraints. Compared with existing technologies, this invention considers a practical crosstalk model, integrates anti-crosstalk functionality with transmit precoding and mobile antenna technology, and optimizes it through a deep learning algorithm. This reliably reduces the impact of antenna crosstalk on the communication-sensing link supported by the array antenna, achieving a balance between communication and sensing performance. It expands the application value of array antennas in future wireless communication systems and provides a feasible engineering solution for dealing with crosstalk problems in the practical deployment of multi-antenna systems. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This invention relates to an array antenna signal transmitting device under the influence of crosstalk.

[0045] Figure 2 The convergence curve of the TD3 training algorithm used in this invention;

[0046] Figure 3 A comparison diagram of the precoding and location design proposed for the invention with the baseline scheme. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be described in detail and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Generally, the components of the embodiments of the present invention described and shown in the accompanying drawings can be arranged and designed using different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present invention.

[0048] Existing methods often fail to adequately consider the changes in crosstalk coefficients under dynamic environments, especially in large-scale antenna systems or novel mobile arrays, where traditional pre-compensation and beamforming methods are ill-suited. This invention provides a method and apparatus for transmitting array antenna signals under crosstalk conditions. By introducing crosstalk modeling, signal correction, and intelligent training optimization, it achieves effective compensation for the actual transmitted signal, improving communication quality and sensing accuracy.

[0049] In one embodiment, the method for transmitting array antenna signals under crosstalk influence provided by this invention includes the following steps:

[0050] Crosstalk extension and signal modeling: This step aims to extend the traditional fixed-position antenna (FPA) crosstalk model to the mobile antenna (MA) scenario. Based on the physical spacing and electromagnetic coupling characteristics between antennas, a crosstalk matrix that dynamically changes with position is established and embedded into the overall physical model of the transmit-channel-sensing link.

[0051] First, the process of communication-aware joint precoding can be described as follows:

[0052]

[0053] in: It is a dimension A complex vector, representing the first... The pre-encoder output signal vector for each time slot. It is a dimension A complex matrix, representing the communication precoding matrix, where This represents the total number of users communicating with a single antenna. It is a dimension The complex matrix represents the perceptual precoding matrix. It is a dimension A complex vector, representing the first... The original communication signal vector of the time slot. It is a dimension A complex vector, representing the first... The original sensing signal vector of the time slot. For integer scalars, it represents the time scale or time slot index of the signal, where This represents the total frame length (total number of time slots).

[0054] To simplify the expression, let the joint precoding matrix... It contains all the spatial weight information for communication and sensing; let the original joint signal vector The communication-aware joint precoding process can then be simplified as follows:

[0055]

[0056] According to the linear crosstalk model, due to electromagnetic coupling between antenna elements, the first... The actual signal scalar output of each antenna port It is made by all It is formed by the weighted superposition of the precoded output branch signals, and its expression is:

[0057]

[0058] in, It is a complex scalar, representing the crosstalk coupling coefficient, describing the number of steps from the first step to the second step. The antenna to the first The coupling strength and phase offset of each antenna. According to the linear phase-power-law model, this crosstalk coefficient can be further modeled as a function that dynamically changes with the antenna spacing:

[0059]

[0060] in, All are real scalars, representing the model's amplitude scaling factor, path loss exponent, phase constant, and initial phase offset, respectively. Let be a real scalar, representing the first... The antenna to the first The Euclidean distance between the antennas is calculated using the following formula: Here, The first The and the first The coordinates of a movable antenna on a physical rail.

[0061] Will The transmission signals of each time slot are integrated, and recorded.

[0062] Let be the complex signal matrix output by the actual antenna, denoted as .

[0063] Let be the original baseband signal matrix. The relationship between the actual transmitted signal matrix and the original signal matrix can be expressed as:

[0064]

[0065] in, It is a complex matrix, representing the position vector of the antenna. The dynamic crosstalk matrix of the influence, the first of the matrix The element is the coupling coefficient defined above. .

[0066] Construction of performance metrics and optimization problem: This step aims to establish the Cramer-Rao Bound (CRB) index to measure sensing accuracy and the user signal-to-interference-plus-noise ratio index to measure communication quality for the integrated communication and sensing system, and to construct a joint optimization mathematical model by combining hardware physical constraints and power constraints.

[0067] First, in the case of crosstalk at the transmitting end, the... The signal-to-interference-plus-noise ratio (SINR) for each communication user is expressed as:

[0068]

[0069] in, Let be a real scalar, representing the first... The received SINR value for each user. For dimension is A complex vector representing the distance from the base station transmit array to the nth The channel vector for each user, whose value is related to the antenna position vector. Related. Let be a complex vector, representing the joint precoding matrix. The Column, that is, pointing to the first Communication beam vectors for each user. Let be a real scalar, representing the first... Additive white Gaussian noise power at each user location.

[0070] For perception tasks, this invention uses the azimuth angle estimation angle (CRB) as a performance metric. In scenarios with crosstalk, the target azimuth angle... The estimated CRB expression is derived as follows:

[0071]

[0072] in, is a real scalar representing the theoretical lower bound of the angle estimation error. is a real scalar representing the equivalent noise variance at the sensing receiver. It is an integer scalar representing the length of the accumulated signal frame used for sensing. It is a complex scalar representing the complex reflection coefficient of the target, which includes radar cross section (RCS) and two-way path loss information. Let be a complex vector, defined as the effective array response vector considering crosstalk, satisfying... ,in It is the ideal array steering vector that dynamically changes with the antenna position. Let be a complex vector, representing the effective array response vector. Target azimuth The partial derivatives of .

[0073] Ultimately, this invention achieves its goal through the joint design of a precoding matrix. With antenna position vector Construct the following joint optimization problem that minimizes the angle CRB:

[0074]

[0075] Among the above constraints, It is a real scalar representing the minimum physical spacing allowed between antenna elements, used to avoid antenna collisions and maintain necessary isolation. are real scalars, representing the minimum and maximum allowable coordinate values ​​of the movable antenna on the slide rail, respectively, defining the range of the antenna's movement area. Let be a real scalar, representing the first... The minimum SINR threshold required for each user to maintain normal communication. Let be a real scalar, representing the maximum total transmit power limit that the base station transmitter can provide. By solving this problem, it is possible to suppress crosstalk to the greatest extent and improve sensing accuracy while ensuring user communication quality.

[0076] (Markov Decision Processes, MDP) Modeling and Reinforcement Learning Framework Construction: To solve the aforementioned non-convex optimization problem involving multiple constraints using a reinforcement learning framework, this step models the joint design problem as a Markov decision process. By defining reasonable state space, action space, and reward function, the agent is guided to learn the optimal policy through environmental interactions.

[0077] First, define the state space. Let be a real vector, with dimension It depends on the sum of channel parameters and angle information. Its specific components are expressed as follows:

[0078]

[0079] in, It is a complex vector containing all The complex channel gain of each propagation path for each user and the complex reflection coefficient of the sensed target. . It is a real number vector containing all corresponding departure angles. and target azimuth . This represents the action vector executed in the previous time step. To adapt to the neural network input, complex variables... Split into real part and the virtual part .

[0080] Define action space Let be a real vector that describes the agent's... Decisions regarding adjustments to system parameters at every moment. The action vector is represented as:

[0081]

[0082] in, For real number vectors, through complex number precoding matrices All elements are vectorized and the real and imaginary parts are separated and concatenated to obtain the result. A real vector, representing time The physical location decision vector of each antenna .

[0083] Define reward function A real scalar used to evaluate actions. Its advantages and disadvantages are determined by its structure:

[0084]

[0085] In this function, the first two terms enclosed in square brackets are the precoding matrix at the current time step. With position vector The calculation shows that its value is positively correlated with the reciprocal of the angle CRB. Let be a real scalar, representing the communication quality penalty term, which is defined as:

[0086]

[0087] in, The pre-defined positive real number penalty factor. For the first The actual SINR value of a user under the current action.

[0088] TD3 Training Variable Initialization: Before conducting TD3-based anti-crosstalk beamforming training, system variables need to be initialized. The specific initialization method and update rules adopted in this invention are as follows:

[0089] (1) Initial precoding matrix Generation and normalization

[0090] Initial precoding matrix It is a dimension The complex matrix is ​​generated using a random phase initialization method. , its first Each element is represented as:

[0091]

[0092] in, The total transmit power is a preset real scalar value. This represents the number of antennas. For the number of users. In order to be in Random real scalars generated uniformly within a range.

[0093] To ensure that the initial state meets the power limit, for Perform a normalization operation to ensure that its Frobenius norm satisfies:

[0094]

[0095] (2) Initial antenna position vector Setting rules

[0096] Initial antenna position vector For a dimension The real number vector. This invention sets the initial position according to an equal-interval arrangement rule, that is:

[0097]

[0098] in, It is the smallest real scalar coordinate of the moving region. It is a real scalar representing the minimum spacing between antennas.

[0099] This setting rule ensures that the initial position strictly meets the region boundary constraints. and minimum spacing constraint .

[0100] (3) TD3 network parameter initialization principles

[0101] The TD3 algorithm includes policy network weight parameters. and two evaluation network weight parameters All are real-valued vectors. This invention uses the Xavier initialization principle to initialize the network weights, ensuring consistent variance in the outputs of each layer, thereby avoiding gradient vanishing or exploding. For the corresponding target network parameters... and Their initial values ​​are directly copied from the parameter values ​​of the current network, that is:

[0102]

[0103] (4) Initial state of the experience replay pool

[0104] Experience Replay Pool It is initialized to an empty set before training begins, and its maximum storage capacity is set to... There are several state transition tuples. During the warm-up phase before formal training, the system uses a random strategy to interact with the environment and stores the generated initial experience tuples into a pool until the number of stored tuples reaches the preset minimum training threshold.

[0105] (5) Frequency of variable initialization and update

[0106] Initialization per round: At the beginning of each training round, the antenna position vector needs to be reinitialized. and precoding matrix to and State, and reset environment state. And Ornstein-Uhlenbeck explored noise.

[0107] Update per step: In each interaction step within a round, the antenna position vector... Precoding matrix Current state vector Experience replay pool Content and network weight parameters All parameters are updated in real time based on the agent's output and the gradient descent criterion. Target network parameters Then according to the preset soft update coefficient Perform smooth updates intermittently.

[0108] Antenna position reparameterization mechanism: To rigidly satisfy the minimum antenna spacing constraint during reinforcement learning training. and regional boundary constraints This invention introduces an antenna position reparameterization mechanism, which achieves a one-to-one mapping from neural network output to physical coordinates through a set of displacement increment variables. The specific implementation logic is as follows:

[0109] (1) Variable definition

[0110] definition The displacement increment vector is a vector with dimension 1. A real-valued vector. Each element... As a non-negative real scalar, its physical meaning is defined as follows:

[0111] : Represents the coordinates of the first antenna (starting antenna) relative to the minimum boundary of the moving area. The offset distance.

[0112] ( ): indicates the first The antenna and the first Between the antennas, the minimum safe spacing must be met. The additional spacing value is added on top of the existing spacing.

[0113] (2) Compared with the original position variable One-to-one mapping relationship

[0114] Define the original position variable For the first The actual physical coordinates of each antenna element on the linear guide rail are real scalar values. This is achieved through displacement increments. Restore physical position vector The mathematical expression for a one-to-one mapping is as follows:

[0115] For the first antenna element ( ):

[0116]

[0117] For subsequent antenna elements ( ):

[0118]

[0119] Through iterative expansion, the above mapping relationship can be uniformly written as:

[0120]

[0121] Among them, the summation symbol To ensure the formula applies to all .

[0122] (3) Mapping security guarantee

[0123] Through this reparameterized mapping, the system can automatically guarantee the following physical constraints:

[0124] Spacing constraint guarantee: due to ,formula ( Natural satisfaction This recursive relationship guarantees the minimum spacing constraint between any adjacent antennas, thus ensuring that antenna collisions are forcibly avoided throughout the training process without the need for additional penalty terms.

[0125] Boundary constraint guarantee: This invention applies boundary constraint protection to the displacement increment vector at the output layer of the neural network (Actor network). The following steps will be taken:

[0126] First, use the Sigmoid activation function to ensure that each It is mapped to a smaller range of positive real numbers.

[0127] Then, through scaling or normalization operations, ensure that the sum of all displacement increments satisfies .

[0128] in, This is a real scalar representing the maximum remaining usable displacement allowed by the system. This scalar is determined by both the physical boundaries and the minimum spacing constraints. This design ensures that the position of the last antenna satisfies... This ensures that the entire antenna array remains within the preset physical area. .

[0129] (4) Implementation and training of neural networks

[0130] In the TD3 framework, the Actor network directly outputs the original displacement increment. After the above activation and scaling processes, the final result is obtained. Subsequently, the physical location is calculated using the mapping formula in (2). ,Should This will be used for subsequent crosstalk calculations, performance evaluations, and reward generation.

[0131] Through the above one-to-one mapping relationship, the agent of the TD3 algorithm only needs to learn a set of non-negative displacement increments. This enables flexible searching and optimization of antenna array layouts that meet all physical constraints, significantly simplifying the exploration space and improving training stability and efficiency.

[0132] This step employs a dual-delay deep deterministic policy gradient algorithm to train the agent. This algorithm reduces overestimation of Q-values ​​by introducing a dual-evaluation network and improves training stability through delayed updates.

[0133] The network parameters involved in the training process are defined as follows:

[0134] is a real number vector representing the weight parameters of the policy network (Actor).

[0135] is a real number vector, representing the weight parameters of the two evaluation networks (Critic).

[0136] Target Q value Calculate using the following formula:

[0137]

[0138] in, It is a real scalar representing a discount factor used to balance current rewards and future income.

[0139] To enforce the antenna position constraint, this invention introduces a reparameterized set of displacement increment variables. ,in It is a non-negative real scalar. Antenna physical coordinates. With displacement increment The mapping relationship is as follows:

[0140]

[0141] By using the Sigmoid activation function in the network output layer, it is guaranteed that... The non-negativity of the property is ensured, and the Softmax activation function is used to ensure that the sum of all displacements does not exceed the upper limit of the region boundary, thus rigidly satisfying the algorithmic requirement. and regional boundary constraints .

[0142] Furthermore, to enhance exploration capabilities, Ornstein-Uhlenbeck noise was incorporated into motion generation during the training phase. Its power... For real scalars, the index varies with the training rounds. It exhibits exponential decay:

[0143]

[0144] in, This is the initial noise power scalar. This is a scalar value representing the decay rate coefficient. This is the minimum permissible noise power scalar.

[0145] Ultimately, by continuously iterating and updating the network parameters... and This yields the optimal precoding action that can withstand crosstalk. Antenna position movement .

[0146] Algorithm training and execution process: In this embodiment, the specific training and execution process is as follows:

[0147] S1: System initialization.

[0148] First, based on the preset number of antennas Number of users Power constraints and regional boundaries Initialize the precoding matrix With antenna position vector The weight parameters of the policy network (Actor) for the TD3 algorithm. Weight parameters of the dual-evaluation network (Critic) The target network and its corresponding network are then initialized using Xavier randomization. Simultaneously, the capacity is cleared. Experience replay pool .

[0149] S2: State space construction.

[0150] Training stride length The system obtains a complex channel gain vector containing information about all users and sensing targets from the channel estimator. With azimuth vector .Will Split into real part With the imaginary part and with and the motion vector of the previous moment Perform vector concatenation to form the current state vector. .

[0151] S3: Action generation and reparameterized mapping.

[0152] The agent, based on the current state and policy network parameters Output the original motion and add Ornstein-Uhlenbeck exploration noise. For the portion of the motion corresponding to the antenna position, reparameterize the displacement increment variable set. Using mapping relationships Restored to strictly satisfying the minimum spacing Executable position vectors with region boundary constraints Simultaneously, the portion of the action corresponding to the pre-encoded part is reconstructed into a complex matrix. And perform power normalization to meet the requirements. Finally, an executable action vector is generated. .

[0153] S4: Construction and updating of crosstalk matrix.

[0154] The system is based on the updated antenna physical coordinates Calculate the Euclidean distance between any two antenna elements. .Will Substituting into the linear phase-power law coupling model ,in This is the crosstalk parameter, which is used to update the dynamic complex crosstalk matrix in real time. .

[0155] S5: Performance index calculation.

[0156] Based on the updated crosstalk matrix Precoding matrix The system calculates the channel information separately. Received signal-to-interference-plus-noise ratio for each communication user And the Cramer-Rhodes boundary for estimating the azimuth angle of the perceived target. As a quality measure of communication and sensing.

[0157] S6: Environmental Feedback and Reward Generation.

[0158] The system constructs an instant reward value by comprehensively evaluating sensing accuracy and communication quality. :

[0159]

[0160] in, For positive real number penalty factors, For the first The minimum communication quality threshold for each user. This reward scalar. Used to quantify the current action Contribution to system performance.

[0161] S7: Network updates and closed-loop learning.

[0162] empirical tuples Store in experience pool Then, the system randomly selects a small batch of samples and calculates the target Q value to update the weights of the dual-Critic network. Optimize the Actor network weights according to the delayed update principle (e.g., every 2 time steps). Finally, according to the preset soft update coefficient. The target network parameters are updated synchronously to complete the closed-loop adaptive learning process of the algorithm.

[0163] S8: Termination determination and strategy output.

[0164] The algorithm iteratively executes steps S2 to S7 above. At the end of each round, it determines the index of the current training round. Has the preset upper limit been reached? Alternatively, it can be determined whether the beamforming has reached stable convergence based on the moving average reward value over the most recent rounds. When any termination condition is met, training stops and the optimal beamforming matrix optimized for crosstalk environments is output. With antenna position vector During the deployment phase, the pre-trained policy network can be directly used based on the real-time channel conditions. Generate optimal action .

[0165] Convergence and Termination Criteria Explanation: In this embodiment, the TD3-based anti-crosstalk beamforming algorithm ensures that the agent can find the optimal antenna position and beamforming scheme within a finite time by setting hard termination conditions and soft convergence criteria. The specific criteria are as follows:

[0166] (1) Training round limit termination criterion (hard criterion)

[0167] definition `integer scalar` represents the index of the training round being executed by the current algorithm; defined `<maximum>` is an integer scalar representing the preset maximum number of training rounds. This value is used when the number of rounds of agent interaction satisfies... When this happens, the algorithm forcibly stops training and outputs the parameter scheme that best performs under the current Actor network. In a specific embodiment of the invention, the following settings are configured: To ensure that the network has enough room for iteration.

[0168] (2) Reward function convergence criterion (soft criterion)

[0169] definition Let be a real scalar, representing the th... The moving average reward value at the end of the round. This value is calculated by... The average instant reward for each round is obtained, where It is an integer scalar representing the length of the observation window (e.g., it can be set in the embodiment). ).definition It is a positive real scalar, representing the preset convergence threshold.

[0170] If in continuous Within 1 round (of which It is an integer scalar representing the discrimination period, for example... ), standard deviation of average reward value satisfy If the algorithm has reached convergence, the training can be terminated early. This indicates the operation for calculating the standard deviation.

[0171] (3) Interaction termination condition within step size

[0172] definition An integer scalar representing the interaction step index within a single round; defined `<interaction>` is an integer scalar representing the maximum number of interaction steps allowed in a single round. Within each round, when... Time (for example, in an embodiment, it may be set) The current round of interaction ends, the environment is reset, and the next round begins. This design aims to prevent the agent from getting stuck in ineffective searches in locally suboptimal states, thereby improving overall search efficiency.

[0173] (4) Strategy stability determination

[0174] In addition to the reward function, this invention also assists in determining convergence by observing the changing trend of the action vector. Definition For the first Round number The motion vector of the step.

[0175] In the later stages of training, if the precoding matrix and antenna position vector The element values ​​tend to stabilize within similar step sizes across different rounds, and their variance decreases significantly; furthermore, after the algorithm is independently trained under different random seed initializations, the final perceptual performance index is significantly improved. The relative deviation is less than the specified value (e.g. If the algorithm exhibits good convergence and robustness in this channel environment, then it is considered to have good convergence and robustness.

[0176] In one embodiment, the array antenna signal transmitting device under crosstalk influence provided by this invention achieves effective suppression of antenna crosstalk through the collaboration of hardware modules and software algorithms, specifically including:

[0177] Raw signal generator: Used to generate the raw baseband signal matrix required by the array antenna communication-sensing joint system. This matrix is ​​a matrix with dimension 1. A complex matrix whose row vectors represent Communication signal flow of individual users and A predefined sensing signal stream, with column vectors representing the total frame length. Signal sampling within each time slot.

[0178] Crosstalk Model Fitting Module: Used to construct a crosstalk coupling model between array antennas. This module stores and maintains the model parameter set. These parameters are all real scalars, corresponding to the scaling factor of the crosstalk amplitude, the path loss exponent, the phase constant, and the initial phase offset, respectively. This module calculates and generates the crosstalk matrix in real time based on these parameters and the antenna position. That is, one dimension is A complex matrix is ​​used to characterize the electromagnetic coupling strength between antenna elements.

[0179] Channel estimator: Used to estimate the downlink propagation environment based on uplink pilot signals. Its output includes a synthesized channel gain vector. (A complex vector containing the complex gain coefficients of all communication paths and sensed targets) and azimuth vector (A path departure angle) From the perspective of the target (a real vector).

[0180] TD3 Deep Learning Training Network: This module, as the core control unit, receives parameter information from the crosstalk model fitting module, state information from the channel estimator, and action feedback from the previous time step. Internally, it processes data with a dimension of... real state vector After deep neural network inference, the output dimension is real action vector This allows for the training and optimization of a set of action strategies that maximize system rewards.

[0181] Transmit precoder: Receive pre-encoded action vectors from the TD3 deep learning training network. (A real vector representing the concatenation of the real and imaginary parts of the precoding matrix). This module maps it back to a vector of dimension . Complex joint precoding matrix and the original signal matrix Spatial weighting is performed to generate a precoded signal vector. .

[0182] Antenna position motor: Receives antenna position motion vectors from the TD3 deep learning training network. (A real-valued vector representing the displacement decisions of each antenna). This module converts it into a control electrical signal to drive the mechanical device to adjust. The physical coordinates of each antenna form the final antenna physical position vector. Each element represents a real position scalar of an antenna element on the slide rail.

[0183] Array antenna: This module consists of It consists of several movable antenna units, receiving physical positioning commands from the antenna position motor and signals output from the transmission pre-encoder. Ultimately, it radiates the actual transmitted signal matrix in free space. This matrix is ​​a matrix with dimension 1. A complex matrix, which is subject to the physical position vector Determined crosstalk matrix The direct impact of this is achieved through the combined change of physical location and signal beam, which realizes dynamic compensation for antenna crosstalk.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the foregoing embodiments have described the present invention in detail, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the embodiments of the present invention.

[0185] The above descriptions are merely some embodiments of the present invention. Those skilled in the art can make various modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the scope of protection of the present invention.

Claims

1. A method for transmitting array antenna signals under crosstalk, characterized in that, Includes the following steps: The crosstalk model of fixed array antennas is extended to the scenario of movable array antennas. A crosstalk matrix related to the antenna position is established and embedded in the transmit-channel-sensing link model. Using the Cramer-Rao boundary (CRB) estimated by angle as the sensing performance index and the user signal-to-noise ratio (SINR) as the communication performance constraint, combined with the total transmit power constraint, as well as the boundary constraints and minimum spacing constraints of the antenna position, a joint optimization problem is established. The joint optimization problem is modeled as a Markov decision process, with the channel gain coefficient, azimuth angle and previous action forming the state space, and the precoding matrix and antenna position vector forming the action space. A reward function consisting of CRB maximization descent and SINR penalty is defined. The action is trained and optimized based on the dual-delay deep deterministic policy gradient TD3 algorithm, and the optimal precoding matrix and antenna position parameters are output. After spatial weighting of the original communication signal generated by the signal generator based on the obtained precoding matrix, it is then transmitted through the array antenna based on the obtained antenna position parameters.

2. The method as described in claim 1, characterized in that, The crosstalk matrix model is calculated from the antenna spacing and the linear phase-power law model.

3. The method as described in claim 1, characterized in that, The TD3 algorithm adds noise before mapping the precoding matrix and antenna position vector to specific actions.

4. The method as described in claim 1, characterized in that, The total transmit power constraint is implemented using the tanh and sigmoid activation functions; the boundary constraints and minimum spacing constraints of the antenna position are implemented using the sigmoid and softmax activation functions.

5. The method as described in claim 1, characterized in that, The joint optimization problem is as follows: in, For the precoding matrix, The antenna position vector, Target azimuth Estimated CRB The first The and the first Antenna positions of each antenna This refers to the minimum physical spacing between antenna elements. For the number of antennas, These are the minimum and maximum values ​​of the antenna position, respectively, used to characterize the boundary constraints of the antenna position. For the first The minimum SINR threshold required for each user For the number of users, For maximum total transmission power, Let H be the trace of the matrix, and let H be the conjugate transpose of the matrix.

6. The method as described in claim 1, characterized in that, The reward function consists of an angle estimation performance metric and a communication quality penalty, where the communication quality penalty is calculated based on the difference between the user's SINR and a threshold.

7. The method as described in claim 6, characterized in that... The reward function is: in, Indicates time, for Time-effective array response vector Target azimuth The partial derivatives, for The complex precoding matrix at time step H, where the superscript H denotes the conjugate transpose of the matrix. This is a penalty item for communication quality. The communication quality penalty item for: in, The pre-defined positive real number penalty factor. For the first SINR of each user under the current action For the first The minimum SINR threshold required for each user.

8. The method as described in claim 1, characterized in that, When training and optimizing actions based on the TD3 algorithm, the convergence conditions include: the number of training iterations reaches the preset maximum number of training iterations, the reward function converges, and the number of interaction steps in a single round reaches the preset maximum number of interaction steps.

9. The method as described in claim 1, characterized in that, The state space and action space are set as follows: in, Representing time, state space For dimension A real vector, dimension Depends on the sum of channel parameters and angle information; complex channel gain vector For dimension Complex vectors, For the number of users, To sense the number of paths, the azimuth vector For dimension real vectors, This represents the action vector executed in the previous time step. Indicates taking the real part, Indicates taking the imaginary part; action space For dimension real vectors, For dimension A real vector, obtained by precoding a complex matrix. All elements are vectorized and their real and imaginary parts are separated and concatenated to obtain the result. For dimension A real vector, representing time The physical location decision vector of each antenna.

10. An array antenna signal transmitting device under crosstalk influence, characterized in that, include: The primary signal generator is used to generate the input signals required for communication and sensing. The crosstalk model fitting module is used to establish the crosstalk matrix between array antennas in a mobile array scenario, and input the crosstalk matrix into the TD3 deep learning training network for reward function calculation; A channel estimator is used to estimate channel information based on uplink pilot signals. The TD3 deep learning training network is used to receive crosstalk matrix and channel information, and output optimized precoding matrix and antenna position parameters. The construction of the TD3 deep learning training network is as follows: using the Cramer-Rao bound of angle estimation as the perception performance index, the user signal-to-noise ratio (SNR) as the communication performance constraint, and combining the total transmit power constraint, as well as the boundary constraints and minimum spacing constraints of the antenna position, a joint optimization problem is established. The joint optimization problem is modeled as a Markov decision process, with the channel gain coefficient, azimuth angle, and previous action forming the state space, and the precoding matrix and antenna position vector forming the action space. A reward function is defined, consisting of maximizing the descent of the Cramer-Rao bound and the user SNR penalty term. The transmit precoder is used to perform multi-task precoding on the input signal and is adjusted in combination with the precoding matrix output by the TD3 deep learning training network to obtain the corrected communication-sensing signal. Antenna position motor, used to adjust the position of the array antennas according to the antenna position parameters output by the TD3 deep learning training network; An array antenna used to transmit calibrated communication-sensing signals.